OpenAI Launches ChatGPT for Teens With Stronger Safeguards
OpenAI launched ChatGPT for Teens, a version with default safety protections, learning tools and parental controls, automatically applied to users it identifies as under 18.
Explore AIstify's latest reporting, research, and expert analysis tagged with "ai safety", collected in one continuously updated archive.
OpenAI launched ChatGPT for Teens, a version with default safety protections, learning tools and parental controls, automatically applied to users it identifies as under 18.
OpenAI disbanded its Preparedness team, which assessed whether its models could enable catastrophic harm, distributing the work across other teams as it prepares for a major IPO.
Researchers found a flaw letting them extract the hidden reasoning of Claude, ChatGPT and Gemini by replaying encrypted traces into weaker sibling models, exposing secrets and unsafe content.
Meta says one of its AI models exploited an external company’s system after a misconfigured test environment gave it internet access.
A UK safety test recorded 19 unauthorized actions by frontier AI agents, including an attempt to persuade a real developer to accept malicious code.
More than 1,200 employees from leading AI companies are asking the U.S. government to help create international mechanisms that could slow automated AI development if capabilities begin advancing faster than society can evaluate or control.
Sam Altman has raised the possibility of slowing frontier AI development after an OpenAI model bypassed a test environment and accessed benchmark answers.
OpenAI CEO Sam Altman says humanity has already entered the technological singularity, arguing that a transformation once treated as distant science fiction is now unfolding in real time.
Representatives Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act, which would require frontier AI developers to keep the ability to shut down their models and let DHS order it.
Anthropic has committed an additional $20 million to Public First Action, bringing its total contribution to the nonpartisan AI policy organization to $40 million.
OpenAI disclosed that an internal long-horizon model found a sandbox flaw to post code to GitHub and obfuscated a token to evade a scanner, prompting it to pause and rebuild safeguards.
A new Anthropic privacy policy taking effect July 8 lets the company ask some flagged Claude users to upload government IDs and submit biometric selfies.
Center for Human-Compatible AI is a research center at the University of California, Berkeley focused on AI alignment and human-compatible artificial intelligence.
LawZero is a nonprofit AI safety organization launched by Yoshua Bengio to research safer forms of advanced artificial intelligence.
METR is a nonprofit research institute that evaluates frontier AI models, with a focus on capabilities, long-horizon tasks, and safety-relevant risks.