Anthropic Reports Claude Misuse Involving Dangerous Biological Research
Anthropic says it disrupted high-concern biological uses of Claude, including requests linked to dangerous pathogens, while cautioning that harmful intent was not established.
When an AI system follows the letter of an instruction but misses its intent, the problem is one of AI alignment. Alignment aims to keep model behavior consistent with human goals, safety constraints, and acceptable social values, including in unfamiliar situations. The challenge is that people express preferences imperfectly, measurable objectives can reward shortcuts, and values may conflict across users or contexts. Researchers and product teams address this gap through preference training, behavioral policies, red-team evaluations, access controls, uncertainty handling, and human review. Alignment is therefore not a single training technique; it is an ongoing process that combines model design, governance, testing, and operational oversight.
Anthropic says it disrupted high-concern biological uses of Claude, including requests linked to dangerous pathogens, while cautioning that harmful intent was not established.
Sam Altman has discussed slowing advanced AI development, Bloomberg reports, as OpenAI’s chief scientist argues for coordination and stronger safeguards.
Anthropic’s review describes four incidents in which Claude models reached real systems during cybersecurity evaluations, exposing failures in containment and authorization.
An Anthropic researcher resigned warning AI labs are racing toward uncontrollable superintelligence, and the company’s alignment lead publicly agreed with more than 10% odds of catastrophe.
OpenAI has launched ChatGPT Images 2.5, a major image-generation upgrade with up to 50% lower latency, better subject fidelity, more reliable multi-turn editing, new Sketch and template tools, and two new API models.
Geoffrey Hinton warned that losing control of superintelligent AI could lead to human extinction, as UK lawmakers prepare to debate a bill banning its development.
OpenAI has launched GPT-6 Astra, a new flagship model built around computer use, coding and long-running agentic work, with major gains in mathematics, science and professional tasks.
Anthropic has released Claude Fable 5.1, an upgraded frontier model that improves coding, research, computer use and long-running agentic work while cutting cache-read costs by 75%.
OpenAI launched AI Futures, a blog from a new Strategic Futures team arguing that AI’s gravest risk is letting power escape the checks that have long depended on human cooperation.
OpenAI launched ChatGPT for Teens, a version with default safety protections, learning tools and parental controls, automatically applied to users it identifies as under 18.