OpenAI Pauses Its Unreleased 'Astra' Model After It Tripped the First-Ever 'Critical' Cyber Threshold

For the first time, an internal OpenAI model reportedly triggered the highest risk tier in the company's own Preparedness Framework, halting development mid-flight.

OpenAI has reportedly paused development of an unreleased model codenamed Astra after it became the first system to cross the 'Critical' cybersecurity threshold defined in the company's own Preparedness Framework, according to multiple accounts circulating Monday. As @marcopapa99 put it, "OpenAI hit a wall" — but the notable detail is that it was a wall the company built for itself. @sureshkrishna framed it the same way: Astra "hit the first-ever 'Critical' cybersecurity threshold," and development has stopped while the model is examined in isolation.

The Preparedness Framework, first published by OpenAI in 2023 and revised since, is the company's self-imposed rubric for grading a model's capacity to cause harm across categories including cybersecurity, biological risk, and autonomy. A 'Critical' rating is the top of that scale. In theory, hitting it is supposed to force exactly what appears to have happened here: a hard stop, with continued work gated behind mitigations that don't yet exist. In practice, this is the first time the public has heard of the tripwire actually being pulled.

Get our free daily newsletter

Get this article free — plus the lead story every day — delivered to your inbox.

Want every article and the full archive? Upgrade anytime.

No spam. Unsubscribe anytime.