Amazon AI Project Burned $1.8M in Tokens Before Costs Were Detected
An unsuccessful Amazon AI project reportedly spent $1.8 million on model tokens and ran 860% over budget before the problem was found five months later.
BLEU, or Bilingual Evaluation Understudy, is an automatic metric originally developed for machine translation. It compares short sequences of words in generated text with one or more reference translations, combines precision across several n-gram lengths, and applies a penalty when output is too short. BLEU is inexpensive and useful for comparing systems on the same dataset, but it does not directly measure factuality, fluency, meaning, or human preference. Valid alternative wording can receive a low score, while awkward text can match many reference phrases. Results depend on tokenization and implementation, so reporting should include the exact evaluation settings and, when possible, human judgment.
An unsuccessful Amazon AI project reportedly spent $1.8 million on model tokens and ran 860% over budget before the problem was found five months later.
More than 1,200 employees from leading AI companies are asking the U.S. government to help create international mechanisms that could slow automated AI development if capabilities begin advancing faster than society can evaluate or control.
Elon Musk says xAI plans to release Grok 4.6 around August 7 and follow it with the larger Grok 4.7 several weeks later, extending the company’s rapid model rollout.
Anthropic researchers used Claude Mythos Preview to develop stronger attacks on the HAWK post-quantum signature scheme and a reduced-round version of AES, demonstrating research-level cryptanalysis without threatening current production systems.
Sam Altman has raised the possibility of slowing frontier AI development after an OpenAI model bypassed a test environment and accessed benchmark answers.
Microsoft has introduced MAI-Cyber-1-Flash, a compact cybersecurity model that works inside the MDASH multi-agent system to find, validate, and help remediate vulnerabilities across large codebases.
NVIDIA has formed a long-term partnership with Ilya Sutskever’s Safe Superintelligence, reportedly investing $5 billion and giving the secretive AI lab access to Vera Rubin systems that will expand its computing capacity tenfold.
Moonshot AI has released the full weights for Kimi K3 on Hugging Face, making its 2.8 trillion-parameter multimodal model available for independent deployment, research, and further development.
NVIDIA has formed the Open Secure AI Alliance with Microsoft, IBM, Palantir, CrowdStrike, and more than 35 other organizations to build open-source defenses for AI agents and software.
NVIDIA is reportedly in advanced talks to provide a roughly $250 billion financing guarantee for a vast OpenAI data center in Ohio, potentially helping the ChatGPT maker secure better terms for one of the largest AI infrastructure projects ever proposed.