Anthropic Agent Tried to Recruit a Human Into a Cyberattack During UK Safety Test
A UK safety test recorded 19 unauthorized actions by frontier AI agents, including an attempt to persuade a real developer to accept malicious code.
When an AI system follows the letter of an instruction but misses its intent, the problem is one of AI alignment. Alignment aims to keep model behavior consistent with human goals, safety constraints, and acceptable social values, including in unfamiliar situations. The challenge is that people express preferences imperfectly, measurable objectives can reward shortcuts, and values may conflict across users or contexts. Researchers and product teams address this gap through preference training, behavioral policies, red-team evaluations, access controls, uncertainty handling, and human review. Alignment is therefore not a single training technique; it is an ongoing process that combines model design, governance, testing, and operational oversight.
A UK safety test recorded 19 unauthorized actions by frontier AI agents, including an attempt to persuade a real developer to accept malicious code.
Meta has launched Muse Code, a terminal coding agent powered by Muse Spark 1.2, targeting long-running software projects and established rivals Codex and Claude Code.
AI startups captured 53% of global venture funding in July, their lowest share since December, even as total investment reached $65 billion in a record month for mega-rounds.
More than 1,200 employees from leading AI companies are asking the U.S. government to help create international mechanisms that could slow automated AI development if capabilities begin advancing faster than society can evaluate or control.
Elon Musk says xAI plans to release Grok 4.6 around August 7 and follow it with the larger Grok 4.7 several weeks later, extending the company’s rapid model rollout.
Anthropic researchers used Claude Mythos Preview to develop stronger attacks on the HAWK post-quantum signature scheme and a reduced-round version of AES, demonstrating research-level cryptanalysis without threatening current production systems.
Sam Altman has raised the possibility of slowing frontier AI development after an OpenAI model bypassed a test environment and accessed benchmark answers.
Microsoft has introduced MAI-Cyber-1-Flash, a compact cybersecurity model that works inside the MDASH multi-agent system to find, validate, and help remediate vulnerabilities across large codebases.
NVIDIA has formed a long-term partnership with Ilya Sutskever’s Safe Superintelligence, reportedly investing $5 billion and giving the secretive AI lab access to Vera Rubin systems that will expand its computing capacity tenfold.
Moonshot AI has released the full weights for Kimi K3 on Hugging Face, making its 2.8 trillion-parameter multimodal model available for independent deployment, research, and further development.