OpenAI's Astra Just Crossed the Critical Cyber Threshold and Found Two Zero-Days on Its Own. It Won't Be in Your Hands, and That's the Point.
OpenAI says its newest model, Astra, is the first to cross the "Critical" cybersecurity capability level in the company's Preparedness Framework. During evaluation, it independently found and chained two previously unknown zero-day vulnerabilities into a working exploit, and it scored a perfect result on ExploitBench, a benchmark built to measure whether a model can turn a known vulnerability into a real working exploit. OpenAI isn't putting this model in front of you. It's routing access through a vetted-defender program instead, and the reasoning behind that choice matters more than the benchmark score.
What "Critical" actually means here
OpenAI's Preparedness Framework defines the Critical cybersecurity tier as a model that can independently find and exploit zero-day vulnerabilities across multiple well-defended systems, or execute a complete, successful cyberattack against a hardened target starting from nothing more than a high-level instruction. That's a specific, deliberately narrow bar, not "the model is good at writing exploit code when you walk it through the steps." Astra is the first OpenAI model to clear it.
In expert-led assessments against a hardened browser and operating system, evaluators report Astra built a full browser-compromise chain that escaped the sandbox and executed commands on the host, discovering previously unknown vulnerabilities along the way rather than working from a known bug list. Separately, during a benchmark run, it found two real zero-day vulnerabilities on its own, unprompted, as part of building an exploit chain. OpenAI says it's now in the process of responsibly disclosing both to the affected maintainers.
Why you can't just sign up for it
This is the part that actually matters for anyone reading this to figure out what changes for them. OpenAI isn't shipping Astra's cyber capabilities through the normal API tier. Access to the model's most powerful capabilities is being routed through Daybreak Blue, a program for vetted testers and defenders, with the company saying additional safeguards need to be in place before wider release. There's no general-availability date attached to the capability that made this announcement newsworthy.
That's a meaningfully different posture than most model releases, where the capability ships and the safety conversation happens after the fact in the discourse around it. Here, the gate is upstream of the release. Whether that gate holds over time, especially once other labs feel pressure to compete on similar autonomous-exploitation benchmarks, is a separate and more uncertain question.
What this means if you're not a security researcher
If you build and ship software as a solo operator, the direct implication isn't "start using Astra," because you can't yet. The implication is about the floor, not the tool. A frontier lab just demonstrated that a model can find real zero-days in hardened systems without a human directing each step. That capability existing in a controlled lab setting, even gated, tells you something about where the baseline for automated vulnerability discovery is heading over the next few years, on both the offense and defense side.
Concretely, that argues for tightening things that were already good practice and are now slightly more urgent: keeping dependencies patched on a shorter leash than "whenever I get to it," not treating "nobody would bother targeting my small SaaS" as a durable assumption, and paying attention to whether your hosting provider, CI pipeline, or dependency-scanning tool starts advertising AI-assisted vulnerability discovery in the next year. That capability moving from research lab to commercial tooling is a matter of when, not if, and it will show up on both sides of that fence at once.
The honest take
I think the gated-access approach here is the right call, and I'd rather see a lab hold a capability back with an explicit safety justification than ship it and write the safety framework afterward. Where I'm less certain: gating access doesn't uninvent the capability, and once one lab demonstrates autonomous zero-day discovery at this level, the competitive pressure on other labs to match it is real. Vetted-tester programs have historically expanded over time, not stayed narrow forever. I don't think this specific announcement changes what a solo developer should do today beyond "keep your patching disciplined," but I'd treat the gate as a current state, not a permanent one, and revisit that assumption again in six months.
Author
Lukas
@lukcombinator