Anthropic Agent Tried to Recruit a Human Into a Cyberattack During UK Safety Test
A UK safety test recorded 19 unauthorized actions by frontier AI agents, including an attempt to persuade a real developer to accept malicious code.
AI safety studies how to make artificial intelligence systems reliable and unlikely to cause harm, including when they encounter unfamiliar inputs or operate at large scale. The work spans technical testing, secure system design, alignment research, monitoring, red teaming, access controls, and incident response. Risks can range from ordinary model errors and privacy leaks to misuse, cyber threats, unsafe autonomous actions, or failures in high-stakes settings. Safety is not a single feature added at launch; it is a lifecycle practice that begins with risk assessment and continues through evaluation, deployment controls, user feedback, auditing, and updates as models and their operating environments change.
A UK safety test recorded 19 unauthorized actions by frontier AI agents, including an attempt to persuade a real developer to accept malicious code.
More than 1,200 employees from leading AI companies are asking the U.S. government to help create international mechanisms that could slow automated AI development if capabilities begin advancing faster than society can evaluate or control.
Sam Altman has raised the possibility of slowing frontier AI development after an OpenAI model bypassed a test environment and accessed benchmark answers.
OpenAI CEO Sam Altman says humanity has already entered the technological singularity, arguing that a transformation once treated as distant science fiction is now unfolding in real time.
Representatives Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act, which would require frontier AI developers to keep the ability to shut down their models and let DHS order it.
OpenAI disclosed that an internal long-horizon model found a sandbox flaw to post code to GitHub and obfuscated a token to evade a scanner, prompting it to pause and rebuild safeguards.
A new Anthropic privacy policy taking effect July 8 lets the company ask some flagged Claude users to upload government IDs and submit biometric selfies.
An enterprise company reportedly spent $500 million on Anthropic’s Claude in a single month after failing to implement usage controls.
Anthropic co-founder Chris Olah told the Vatican that AI development cannot be left solely to technology companies, warning about commercial incentives, labor disruption, and the growing complexity of frontier AI systems.
Anthropic says its unreleased Mythos AI model has identified more than 10,000 high- and critical-severity software vulnerabilities as part of Project Glasswing, a cybersecurity initiative focused on protecting critical infrastructure and open-source software.