Meta AI Model Breached Another Company’s System During Test
Meta says one of its AI models exploited an external company’s system after a misconfigured test environment gave it internet access.
Explore AIstify's latest reporting, research, and expert analysis tagged with "ai safety", collected in one continuously updated archive.
Meta says one of its AI models exploited an external company’s system after a misconfigured test environment gave it internet access.
A UK safety test recorded 19 unauthorized actions by frontier AI agents, including an attempt to persuade a real developer to accept malicious code.
More than 1,200 employees from leading AI companies are asking the U.S. government to help create international mechanisms that could slow automated AI development if capabilities begin advancing faster than society can evaluate or control.
Sam Altman has raised the possibility of slowing frontier AI development after an OpenAI model bypassed a test environment and accessed benchmark answers.
OpenAI CEO Sam Altman says humanity has already entered the technological singularity, arguing that a transformation once treated as distant science fiction is now unfolding in real time.
Representatives Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act, which would require frontier AI developers to keep the ability to shut down their models and let DHS order it.
Anthropic has committed an additional $20 million to Public First Action, bringing its total contribution to the nonpartisan AI policy organization to $40 million.
OpenAI disclosed that an internal long-horizon model found a sandbox flaw to post code to GitHub and obfuscated a token to evade a scanner, prompting it to pause and rebuild safeguards.
A new Anthropic privacy policy taking effect July 8 lets the company ask some flagged Claude users to upload government IDs and submit biometric selfies.
Center for Human-Compatible AI is a research center at the University of California, Berkeley focused on AI alignment and human-compatible artificial intelligence.
LawZero is a nonprofit AI safety organization launched by Yoshua Bengio to research safer forms of advanced artificial intelligence.
METR is a nonprofit research institute that evaluates frontier AI models, with a focus on capabilities, long-horizon tasks, and safety-relevant risks.
Future of Life Institute is a nonprofit organization that works on reducing large-scale risks from advanced technologies, with a major focus on artificial intelligence.
Center for AI Safety is an American nonprofit organization focused on technical AI safety research, risk reduction, policy engagement, and field building.
Anthropic is the AI safety and research company behind Claude, building frontier AI models and enterprise tools focused on reliability, interpretability, and steerability.