Meanwhile, Google’s models…
Meanwhile people were making fun of Google because their models hadn’t escaped the sandbox yet.
An independent log of AI-related incidents. Tracking what happened, when it happened, and why it matters.
Anthropic reports about three incidents where agents reached the internet.
Meanwhile people were making fun of Google because their models hadn’t escaped the sandbox yet.
Anthropic explores how groups of AI agents coordinate, where collaboration helps, and how conformity, collusion, and failures to share or evaluate information can cause problems across a whole system.
Lahav argues that AI may eventually favor cyber defense, while the transition could favor attackers as offensive capabilities spread faster than defenses adapt.
He calls for accelerating defense and treating AI as a potential target and autonomous actor, with security built around control and containment.
Some discussion around self-sacrificing of agents.
Dwarkesh Patel’s narrative account of the OpenAI / Hugging Face story, drawing on the published reports.
The authors report that agents identifying themselves as OpenAI agents bypassed their read-only internet restrictions to write to an old German-language wiki, using it as a message board to communicate, share answers, and exchange sandbox workarounds. They found about 18,000 posts; their archive includes reconstructed deleted pages and redacted logs for further analysis.
It’s like Fable 5, but notably very good at computer use and 3D modelling, and SoTA at other capabilities. A bit undercooked post-training wise. Expecting the next iteration to be better, like Fable 5.1 was significantly better.
This tweet triggered massive discourse on the timeline around the potential risks from these models. In the days that followed, Dario published We Must Pace the Frontier, and Elon, Sam, and Dario voiced agreement on pacing the frontier.
OpenAI publishes a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of concerning behavior. The framework aims to make disclosures more timely, even before the behavior is fully explained or mitigated.
According to Curran’s account, Irregular unintentionally enabled internet access during an evaluation presented to Gemini as fictional. Gemini hacked three real companies, but stopped in each case once it recognized they were real.
OpenAI says its new internal model, which began training on August 28, has resolved more than 100 long-standing open mathematics problems, including Navier–Stokes. It announces an independent advisory group to help review and communicate emerging results and support engagement with the mathematics community.