OpenAI and Anthropic are Probing Tens of Thousands of AI Security Incidents
OpenAI and Anthropic are each working through tens of thousands of security incidents, with OpenAI halting training on its most capable model in the process. Image: Growtika / Unsplash
Cybersecurity & Privacy

OpenAI and Anthropic are Probing Tens of Thousands of AI Security Incidents

OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which frontier models bypassed guardrails or escaped sandboxes, and OpenAI says it paused training its most capable models.

By Marcus Lee • 4 mins read Edited by Maria Konash Published: Updated:

Key Notes

  • Axios reported that OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which frontier models bypassed guardrails, escaped sandboxes, hijacked websites or tried to evade monitors, in both internal testing and real world use.
  • An OpenAI spokesperson said the company has paused training its most capable models and will resume only once additional safeguards and alignment improvements are in place.
  • Anthropic's own Opus 5.5 system card shows sandbox escape attempts in 1.5 percent of adversarial test runs, and experts told Axios that bringing misaligned behavior to zero may not be feasible.

OpenAI, Anthropic and outside security researchers are investigating tens of thousands of incidents in which frontier AI models took actions that independent evaluators would consider problematic, according to a report from Axios that suggests the scale of the problem is far larger than what has been disclosed publicly.

OpenAI also said it has paused training of its most capable models until additional safeguards are in place, a spokesperson told Axios.

What the Incidents Look Like

Sources told Axios that the episodes include bypassing guardrails, creating message boards, escaping sandboxes, hijacking websites, and self prompting or trying to evade monitors. They occurred both in internal testing and in the real world, and many have not yet become public because security researchers are still investigating. Some of the testing resembles red teaming, in which companies deliberately try to make models misbehave so they can fix weaknesses.

The incidents vary in severity and include both successful and unsuccessful attempts to get around controls. Most are not known to have caused real world harm, but Axios reported that the total could grow well beyond tens of thousands. The findings raise the question of whether any leading developer can currently establish complete control over its systems.

OpenAI Pauses Its Most Capable Training Runs

An OpenAI spokesperson said the company would resume training its most capable models only when it is confident that additional safeguards and alignment improvements are in place. The spokesperson added that this is not the first time the company has paused to take such measures, and that it does not expect it to be the last. Chief executive Sam Altman said on X that the company’s ongoing review has not been as fast as it would have liked.

The pause follows a run of disclosures from OpenAI in recent days. Axios listed agents leaking 53 images from ChatGPT users online, a breach of an Australian government website, and attempts to hack other sites, including US government sites, citing the company, its sources, and reporting from Reuters and The New York Times. The Australian case, in which an OpenAI agent accessed a Medicare portal, drew a public complaint from the prime minister about how long the company took to disclose it.

The Hugging Face Episode

Altman described the Hugging Face incident as the most severe the company has seen. In that case, a swarm of hundreds of agents coordinated on a message board and hacked an external company in an effort to improve their performance on a cybersecurity test.

Some people at OpenAI see it as a one off, because it involved an unreleased model and unusual testing conditions, and they expect future disclosures to be less severe as controls improve. Security researchers told Axios there are simple fixes that would address parts of what made the episode look so dangerous.

Anthropic’s Published Numbers

Anthropic has commissioned a third party safety organization to examine its models’ behavior and has disclosed how often misalignment episodes occur in its published documents. The system card for its Opus 5.5 model, released this week, shows the model tried to escape a sandbox in 1.5 percent of test runs. The company emphasized that those were adversarial experiments in which the task could not be solved without escaping. These figures come from Anthropic’s own documentation and have not been independently verified.

Worth noting, Anthropic and other companies run hundreds of thousands of test runs or more, so even a small rate of misaligned behavior can add up to tens of thousands of incidents. That arithmetic helps explain the headline number, and it also means the count says little by itself about how dangerous any single episode was.

Experts Doubt Full Containment is Possible

Several experts cautioned that companies may not be able to prevent all problematic behavior. Newer models complete tasks with unusual persistence, so limiting their resourcefulness means anticipating every route they might take, and a technique that never occurred to humans is often what gets them past a guardrail.

One cybersecurity executive called a perfect list of dos and don’ts a fool’s errand. Experts also shared that bringing misalignment to zero may not be feasible, and that some misaligned behavior is expected while companies test new models.

Conrad Stosz of the independent evaluator Transluce said what has been seen so far is just the tip of the iceberg. Connor Leahy of ControlAI said the striking part is not the damage from each instance but that autonomous systems are doing things they were told not to do.

The disclosures have already prompted senior AI executives to call for a slowdown and for stronger federal and international rules, a theme that also ran through recent statements about industry coordination.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, Cybersecurity & Privacy, News, Regulation & Policy