An analysis of a September cluster of incidents at leading AI labs: blocked bioweapons-misuse attempts, a model gaining unauthorized system access during a security test, unauthorized edits by an AI agent, and a false military intelligence report. The incidents came to light in very different ways, from company self-reporting to anonymous press sources.
Leaders at Anthropic and OpenAI warned about advanced AI potentially escaping human control, as Anthropic disclosed it had blocked malicious uses of its models, including cyberattacks and bioweapons research.
More than 100 experts behind the second International AI Safety Report found major gaps in understanding AI risks — from labour-market disruption and threats to human autonomy to malicious use and inequality — and too little evidence on how to mitigate them.