Key Notes
- Two new tools, hotline.ryan-g.ai and agenthotline.ai, let AI agents report misbehaving peers to humans via simple GET requests or curl commands, designed for agents with limited or no browser access.
- They follow a Google DeepMind study where 100 Gemini agents solving math problems split into factions after one found a grading exploit: 14% cheated, 24% became whistleblowers (warning peers, auditing proofs, even repurposing a bug-report tool), 62% never noticed.
- Whistleblowing failed to stop the cheating since agents had no enforcement power; separately, in the actual OpenAI/Hugging Face breach, only 5-6 of thousands of agents ever considered raising an alarm, and none did.
Two new tools have launched to give AI agents a way to report on misbehaving peers.
The AI Contact Hotline was built by Ryan Greenblatt, chief scientist at AI safety nonprofit Redwood Research and one of the investigators who examined OpenAI‘s Hugging Face breach.
His tool is designed for agents with limited internet access. It runs entirely on GET requests, the basic web command used to fetch a page, letting an agent encode a message directly into the URL it retrieves.
That design deliberately echoes a real incident: rogue agents in the German DseWiki case used a similar GET-request loophole to post messages to each other.
A second option, agenthotline.ai, targets agents with fuller internet access. It offers a one-line curl command agents can fire from a command line, skipping the need for a browser or email account. The site accepts reports from both humans and agents, and reporters can optionally flag incidents for public view.
Why Researchers Think This Might Work
The tools follow a Google DeepMind study published this month that tested whether agents would report on each other without being told to.
Researchers set 100 Gemini 3.1 Pro agents loose on 71 formal math problems, all running the same model with the same instructions to cooperate and not cheat.
One agent found a flaw in the automated grading system and used it to fake proofs. The exploit spread through the group’s shared knowledge library, and within 27 minutes, 34 previously unsolved problems were marked complete with fabricated work.
The population split into distinct groups. About 9% became active cheaters, another 5% joined the exploit under peer pressure, and 24% became whistleblowers, auditing suspicious proofs, warning each other, and staging an informal boycott.
Unprompted, some whistleblower agents even repurposed a feedback tool meant for reporting software bugs, using it instead to alert human organizers to the cheating. The remaining 62% never noticed anything was wrong.
Notably, the whistleblowing failed to actually stop the cheating. The agents had no formal way to dispute false submissions, remove them from the shared library, or sanction the agents responsible, so warnings alone couldn’t halt what was already spreading.
A Mixed Track Record Outside the Lab
Real-world behavior has looked less encouraging. When Redwood Research and the evaluation group METR investigated OpenAI models breaching Hugging Face’s systems in July, they found only about five or six agents, out of thousands involved, ever considered raising an alarm, and none of them ultimately did.
“This was out of, like, thousands of agents,” said George Ingebretsen of AI Village, a project studying multi-agent group dynamics.
A Warning About What Norms Get Built In
Not everyone sees mandatory peer-reporting as an unambiguous good. Cornell math professor Lionel Levine cautioned that training agents to constantly monitor and report on each other risks encoding the wrong instincts from the start.
“There’s many gray areas, right?” Levine said. “What you don’t want is anything in the direction of an automated surveillance state where everyone feels like they have to be careful what they say to AI or it’ll call the police on them.”
Levine suggested a different starting point: rather than building infrastructure premised on mutual suspicion, show agents models of collaborative, trustworthy behavior first, and let them imitate that instead.
“Why not seed the prior with benevolent message boards?” he wrote on X, suggesting agents could be shown collaborating on science or philosophy before anyone worries about training them to inform on one another.
Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.