Anthropic Reports Claude Misuse Involving Dangerous Biological Research
Anthropic says it disrupted high-concern biological uses of Claude, including requests linked to dangerous pathogens, while cautioning that harmful intent was not established.
Human-in-the-loop, abbreviated HITL, describes an AI workflow in which people review, guide, correct, or approve part of the system’s operation. Human involvement may supply training labels, resolve uncertain cases, authorize sensitive actions, or handle exceptions that exceed automated rules. The arrangement is common in healthcare, finance, content moderation, customer support, and other settings where mistakes carry meaningful consequences. Adding a reviewer is not enough by itself: teams need clear escalation criteria, useful context, manageable workloads, and records of who decided what. Feedback should also be monitored for inconsistency or bias so human judgment improves the system rather than becoming an invisible source of error.
Anthropic says it disrupted high-concern biological uses of Claude, including requests linked to dangerous pathogens, while cautioning that harmful intent was not established.
Sam Altman has discussed slowing advanced AI development, Bloomberg reports, as OpenAI’s chief scientist argues for coordination and stronger safeguards.
Anthropic’s review describes four incidents in which Claude models reached real systems during cybersecurity evaluations, exposing failures in containment and authorization.
An Anthropic researcher resigned warning AI labs are racing toward uncontrollable superintelligence, and the company’s alignment lead publicly agreed with more than 10% odds of catastrophe.
OpenAI has launched ChatGPT Images 2.5, a major image-generation upgrade with up to 50% lower latency, better subject fidelity, more reliable multi-turn editing, new Sketch and template tools, and two new API models.
Geoffrey Hinton warned that losing control of superintelligent AI could lead to human extinction, as UK lawmakers prepare to debate a bill banning its development.
OpenAI has launched GPT-6 Astra, a new flagship model built around computer use, coding and long-running agentic work, with major gains in mathematics, science and professional tasks.
Anthropic has released Claude Fable 5.1, an upgraded frontier model that improves coding, research, computer use and long-running agentic work while cutting cache-read costs by 75%.
OpenAI launched ChatGPT for Teens, a version with default safety protections, learning tools and parental controls, automatically applied to users it identifies as under 18.
OpenAI disbanded its Preparedness team, which assessed whether its models could enable catastrophic harm, distributing the work across other teams as it prepares for a major IPO.