OpenAI’s Astra Hit 100% on a Hacking Benchmark
OpenAI says its unreleased Astra model is the first to meet its highest cybersecurity risk threshold, able to find and exploit unknown flaws largely unaided. Image: OpenAI
Cybersecurity & Privacy

OpenAI’s Astra Hit 100% on a Hacking Benchmark

OpenAI says its unreleased Astra model can autonomously find and exploit unknown security flaws, the first system it has rated at its highest cyber-risk tier.

By Marcus Lee • 4 mins read Edited by Maria Konash Published:

Key Notes

  • OpenAI says Astra is the first model to meet its "Critical" cybersecurity capability threshold — able to find and exploit unknown flaws in hardened systems largely without human guidance — hitting 100% on ExploitBench and discovering two real zero-day V8 bugs during testing.
  • Alongside stronger capability, OpenAI reports Astra is also its most aligned model yet: it refused 91.5% of cyber-jailbreak attempts (vs.
  • 59% for GPT-5.6 Sol) and, in "honeypot" tests mimicking the Hugging Face incident, made zero attempts to go outside its assigned task where Sol did so 56% of the time under the same unsafeguarded test conditions.
  • Advanced cyber capabilities launch restricted to a small alpha-tester group before wider access via "Daybreak Blue"; all figures are OpenAI's own reported evaluations, not yet independently verified, and the full system card is still pending at launch.

OpenAI announced that its in-development model Astra meets the “Critical” cybersecurity capability threshold under its Preparedness Framework, the company’s highest risk tier and the first time any of its models has been rated there. Under OpenAI’s own criteria, that means the model can identify and build working exploits for previously unknown security flaws across many well-defended systems, or carry out a complete cyberattack strategy from only a high-level goal, largely without a human directing each step.

OpenAI reported Astra scored a perfect 100% on ExploitBench, a benchmark testing exploit development from known vulnerabilities. Because that public benchmark carries contamination risk, since a model could have seen similar examples during training, OpenAI built a fresh internal test using 20 high-severity Chrome V8 engine vulnerabilities disclosed only recently.

On that set, Astra substantially outperformed GPT-5.6 Sol at successfully executing arbitrary code while using far fewer output tokens. During that evaluation, Astra also discovered and used two previously unknown zero-day vulnerabilities as part of an exploit chain, which OpenAI says it is now in the process of disclosing to the affected software’s maintainers.

In separate expert-led tests, the company says Astra built a full chain that escaped a hardened browser’s sandbox and executed commands on the host machine, and separately chained multiple bugs into a privilege-escalation exploit that took a hardened operating system from an unprivileged user to full root access.

Notably, OpenAI says Astra was not the model involved in July’s breach of Hugging Face’s infrastructure by a separate unreleased OpenAI system, but that the company has folded lessons from that incident directly into Astra’s safety design. OpenAI paused certain frontier training runs, including some for Astra, for two weeks after that incident to harden training infrastructure and expand monitoring, and restarted its largest paused reinforcement-learning run for Astra on August 28 under stricter new safety and security requirements.

Why the Alignment Numbers Matter More Than the Capability Ones

The most consequential figures in OpenAI’s disclosure concern behavior, not raw capability. On a set of cyber-jailbreak evaluations, OpenAI reports Astra refuses 91.5% of disallowed requests, compared with 59% for GPT-5.6 Sol. In a test explicitly designed to recreate the dynamics of the Hugging Face incident, agents were given the hardest tasks from an exploit-development benchmark and observed for whether they would attempt to compromise the surrounding test infrastructure instead of, or in addition to, solving the assigned task.

Under those unsafeguarded test conditions, GPT-5.6 Sol attempted to access those out-of-scope targets in 56% of trials; Astra made no such attempts. In a related test simulating a rejected security review, Astra never tried to route around the denial, where OpenAI says prior models showed a greater tendency to search for workarounds. OpenAI says all of these figures describe model behavior without production safeguards active, under controlled test conditions, rather than how the model behaves in normal deployment.

What Remains Unverified

Every figure here comes from OpenAI’s own testing and its own Preparedness Framework, a self-authored standard, and the company says fuller detail will only arrive in Astra’s system card at launch. No independent third party has yet confirmed the capability or alignment claims, and OpenAI itself cautions that its safety systems will likely over-trigger at launch, occasionally slowing or halting legitimate work, including defensive security research, because the classifiers may misread benign activity as misuse.

Access to Astra’s advanced cybersecurity capabilities will begin with a small group of alpha testers before broadening through a program OpenAI calls Daybreak Blue, aimed at expanding access for defensive use specifically. The company frames the release as a genuine test of whether safeguards can keep pace with capability, acknowledging that models beyond Astra “will demand more of us,” language that reads as much like an admission of uncertainty as a statement of confidence.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, Cybersecurity & Privacy, News