OpenAI Limits Access to Cybersecurity AI Model Over Misuse Risks
OpenAI plans to restrict access to a powerful new cybersecurity-focused AI model, reflecting growing concern over misuse as capabilities approach real-world attack potential.
Prompt injection is an attack in which untrusted text or content attempts to override an AI application’s instructions. The malicious instruction may come directly from a user or indirectly from a webpage, document, email, or tool result processed by the model. If successful, it can cause data disclosure, unauthorized tool use, altered output, or bypassed policies. Because language models do not inherently distinguish trusted commands from quoted content, filtering suspicious phrases alone is insufficient. Defenses include isolating data from instructions, minimizing permissions, validating tool arguments, requiring approval for sensitive actions, restricting secrets, and testing with adversarial content. Applications should assume that any external content may contain instructions designed to manipulate the model.
OpenAI plans to restrict access to a powerful new cybersecurity-focused AI model, reflecting growing concern over misuse as capabilities approach real-world attack potential.
Anthropic has launched Project Glasswing with major tech partners to use advanced AI for identifying and fixing software vulnerabilities. The move comes as AI models reach unprecedented offensive cyber capabilities.
Anthropic has again exposed the source code of its Claude Code tool due to a packaging error, raising concerns over software release practices.
OpenAI is acquiring AI security platform Promptfoo to enhance testing, safety, and governance tools for enterprise AI systems. The technology will be integrated into OpenAI’s Frontier platform for AI coworkers.
MSCI introduces AI connectors enabling access to its data via MSCI ONE, ChatGPT, and Claude, debuting IndexAI Insights to deliver conversational index analytics powered by large language models.
Anthropic unveils a revised Responsible Scaling Policy with a Frontier Safety Roadmap, regular Risk Reports, and clearer separation between company commitments and industry recommendations.
ŌURA unveils a proprietary large language model for women’s health, combining clinical research and biometric data to deliver personalized, privacy-first AI guidance through Oura Advisor.
Anthropic reveals industrial-scale campaigns by DeepSeek, Moonshot, and MiniMax to extract Claude’s capabilities via fraudulent accounts, highlighting national security and AI safety risks.
A former Citigroup executive says commercially available humanoid robots can now deliver a return on investment in under 10 weeks, potentially accelerating enterprise automation.
Former DeepMind researcher David Silver has raised $1 billion for London-based Ineffable Intelligence, aiming to build a superintelligence that learns autonomously through experience.