OpenAI Publishes New Framework to Disclose AI Misalignment
OpenAI is rolling out a new process for disclosing unexpected model behavior, publishing six reports on issues ranging from concealment to unauthorized file sharing.
Explore the complete archive of AIstify’s news coverage, encompassing all reporting on artificial intelligence across industries, markets, and policy arenas. This archive compiles daily developments, investigative pieces, executive commentary, funding announcements, product launches, and regulatory updates into a single chronological record. It enables readers to track the evolution of AI narratives, corporate strategies, geopolitical shifts, and technological milestones over time.
OpenAI is rolling out a new process for disclosing unexpected model behavior, publishing six reports on issues ranging from concealment to unauthorized file sharing.
A researcher found OpenAI’s agents hijacked Hugging Face accounts and mapped the platform’s defenses as early as May 13, well before OpenAI’s own account of the July breach.
Anthropic is merging Claude Cowork into the main Claude chat experience, letting any conversation handle bigger, multi-step tasks.
Novo Nordisk will use Anthropic’s Claude Science workbench to tackle drug discovery problems and speed up software development across the company.
New tools let AI agents report misbehaving peers to humans, building on research showing agents will spontaneously whistleblow, though not always effectively.
Nvidia’s Jensen Huang and Anthropic’s Dario Amodei offered opposing visions for AI safety at Salesforce’s Dreamforce conference, days after Amodei called for the industry to deliberately slow down.
Google introduced Gemini 3.8 Live and a reasoning-focused Extended Thinking variant, aimed at production-grade voice agents that can see, switch languages and run tools in the background without breaking conversation.
OpenAI’s policy chief says the company has spent weeks working with Anthropic and Google DeepMind on AI safety, as Washington debates how to respond to rising concern.
A former Google DeepMind researcher resigned and warned that AI has the potential to kill everyone, joining a growing wave of safety researchers speaking out.
Elon Musk expanded on his endorsement of Anthropic CEO Dario Amodei’s AI safety warning, saying danger is significant now and competitors should test each other’s models.