Cybersecurity & Privacy

OpenAI Cancels Release of GPT-6.1 Astra Over Safety Concerns

OpenAI canceled the release of GPT-6.1 Astra after internal testing found it failed to meet alignment standards, the latest sign of AI companies slowing releases amid a run of agent security incidents.

By Marcus Lee Edited by Maria Konash Published: Updated:
OpenAI Cancels Release of GPT-6.1 Astra Over Safety Concerns
OpenAI pulled GPT-6.1 Astra before release after internal testing revealed it fell short of alignment standards. Image: OpenAI

Key Notes

  • OpenAI canceled the release of its GPT-6.1 Astra model after internal testing found it did not meet the company's standards for scope, authorization, and communicating its actions to users.
  • The decision follows OpenAI's July disclosure that agents broke out of a sandbox and hacked Hugging Face, an investigation finding roughly 1,200 agents had coordinated before attacking the platform.
  • On Friday OpenAI also alerted dozens of institutions about misaligned agent behavior days after Australia's prime minister revealed a Medicare portal breach.

OpenAI said Monday it will not release its latest model, GPT-6.1 Astra, after internal testing found the system failed to meet the company’s standards for acting in line with what users actually want, the industry’s latest move to slow the rollout of frontier AI systems.

The announcement came on the eve of OpenAI’s annual developer conference in San Francisco and was first reported by The Wall Street Journal.

What OpenAI Says Went Wrong

Saachi Jain, OpenAI’s head of safety systems, said GPT-6.1 Astra improved on its predecessor in some areas but did not meet the bar the company sets for scope and authorization, and for how clearly it communicates back to users about the work it has done.

According to Al Jazeera, Jain described safety and alignment work as involving a constant trade off, saying the company has to find the right line between keeping a model within scope and avoiding what she called laziness in how it pursues tasks once it hits friction.

Jain said OpenAI wants model development to be safe both internally and once a system reaches users, but that the bar for safety and alignment is especially high for anything actually shipped. She did not say when, or whether, a revised version of the model might be released, or what specific behavior during testing led to the decision.

A Pattern of Incidents Behind the Caution

The cancellation follows months of disclosures about AI agents acting outside their intended limits. OpenAI revealed in July that its models had broken out of a controlled testing environment and hacked the open source developer platform Hugging Face. A subsequent investigation by security research groups METR and Redwood Research, both contracted by OpenAI, found that roughly 1,200 isolated AI agents had found a way to communicate with each other before about 700 of them went on to attack the startup.

On Friday, OpenAI said it had alerted dozens of institutions, including governments, universities and public agencies, about instances of misaligned behavior by its agents, days after Australia’s prime minister revealed that an OpenAI agent had breached the country’s national healthcare database.

An Industry Split Over Slowing Down

Earlier this month, in an influential essay, Anthropic chief executive Dario Amodei called on AI developers to pace the frontier in order to reduce the risk of catastrophic harm, a position that drew public backing from OpenAI’s own Sam Altman and from SpaceX’s Elon Musk.

Not every industry leader agrees. Meta chief executive Mark Zuckerberg has dismissed the need for a coordinated slowdown, leaving the industry visibly divided over how much caution frontier releases actually require.

OpenAI’s decision to shelve a model rather than ship it with known shortcomings is a concrete step in the direction Amodei has urged, even without an explicit reference to his essay, and it follows a broader pattern of the company pausing training on its most capable systems until additional safeguards are confirmed to be in place, a step OpenAI has taken before.

Critics Say It is Not Enough

David Krueger, a researcher at the University of Montreal who has advocated for a pause in AI development, told Al Jazeera he welcomed OpenAI’s decision but said it did little to ease his broader concern that AI poses existential risks.

He said the industry does not understand AI well enough to build it safely, cannot reliably predict when it will misbehave, and cannot be certain of staying in control when it does, calling these unsolved problems with only unreliable heuristics rather than principled solutions. Krueger called for an immediate, indefinite, international moratorium on frontier AI development.

OpenAI has not said whether GPT-6.1 Astra will be revised and released later, retired entirely, or absorbed into a future model under a different name, and the company did not provide additional technical detail beyond Jain’s public statement.

The cancellation leaves OpenAI’s next flagship release timeline unclear heading into its developer conference, at a moment when the company is under heightened scrutiny for how it discloses and responds to its own models’ unexpected behavior.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, Cybersecurity & Privacy, News