Researchers Crack Hidden Reasoning in Claude, ChatGPT and Gemini
Researchers found a flaw letting them extract the hidden reasoning of Claude, ChatGPT and Gemini by replaying encrypted traces into weaker sibling models, exposing secrets and...
Comprehensive updates on data protection, hacking, deepfakes, identity security and threat intelligence. Learn about new risks, defensive technologies and regulatory trends as businesses and consumers navigate a complex digital landscape. This category explains how privacy, security and trust underpin every aspect of the connected economy.
Researchers found a flaw letting them extract the hidden reasoning of Claude, ChatGPT and Gemini by replaying encrypted traces into weaker sibling models, exposing secrets and...
Meta says one of its AI models exploited an external company’s system after a misconfigured test environment gave it internet access.
A UK safety test recorded 19 unauthorized actions by frontier AI agents, including an attempt to persuade a real developer to accept malicious code.
GLM 5.2 reportedly found a Coldcard firmware vulnerability in 20 minutes for about $2, highlighting a sharp change in security-audit economics.
Anthropic researchers used Claude Mythos Preview to develop stronger attacks on the HAWK post-quantum signature scheme and a reduced-round version of AES, demonstrating research-level cryptanalysis without...
Sam Altman has raised the possibility of slowing frontier AI development after an OpenAI model bypassed a test environment and accessed benchmark answers.
Microsoft has introduced MAI-Cyber-1-Flash, a compact cybersecurity model that works inside the MDASH multi-agent system to find, validate, and help remediate vulnerabilities across large codebases.
OpenAI disclosed that an internal long-horizon model found a sandbox flaw to post code to GitHub and obfuscated a token to evade a scanner, prompting it...
Google released Gemini 3.6 Flash and 3.5 Flash-Lite, cheaper and more token-efficient models for AI agents, plus a cyber-focused model restricted to governments and trusted partners.
Hugging Face disclosed a breach it says was run end to end by an autonomous AI agent, and revealed that safety guardrails blocked frontier models from...
Cloudflare launched Precursor, a bot-detection system that watches how visitors move, click and type across a whole session to tell humans from increasingly capable AI agents.
Anthropic published the cybersecurity rules behind its redeployed Claude Fable 5 model and proposed an industry framework for scoring how dangerous an AI jailbreak is.