Meta Shares Jump 7% After Muse Spark AI Launch
Meta shares climbed 7% after unveiling Muse Spark, its new multimodal AI model, signaling investor confidence in the company’s revamped AI strategy.
Human evaluation asks people to score, rank, compare, or annotate AI outputs. It is essential for qualities such as usefulness, tone, creativity, cultural appropriateness, factual support, and safety when no automatic metric fully represents the goal. A reliable study defines clear criteria, randomizes presentation, includes representative tasks, trains evaluators, and measures agreement. Reviewers can still be inconsistent, fatigued, biased, or influenced by style, brand, and answer length. Preference results also depend on who participates and what context they receive. Human evaluation should protect worker well-being, especially for harmful content, and be combined with automated tests that provide scale and reproducibility.
Meta shares climbed 7% after unveiling Muse Spark, its new multimodal AI model, signaling investor confidence in the company’s revamped AI strategy.
SoftBank has secured a $40 billion bridge loan to deepen its investment in OpenAI and accelerate its broader AI strategy.
AI-driven bots now generate more internet traffic than humans, according to a new report, highlighting the growing impact of automated systems online.
New research shows that telling AI models they are “experts” can reduce accuracy in tasks like coding and math, while improving safety and alignment outcomes.
Meta introduced four new in-house MTIA chips designed for AI training and inference as the company accelerates data center expansion. The chips aim to improve performance and reduce reliance on external hardware suppliers.
Google has rolled out Canvas in AI Mode to all U.S. users in English, enabling people to create documents, dashboards, and interactive tools directly within Search. The feature provides a dynamic workspace for planning projects and building simple applications.
Andrej Karpathy argues AI coding agents now fundamentally change programming workflows, enabling long-running autonomous tasks and redefining software engineering.
MSCI introduces AI connectors enabling access to its data via MSCI ONE, ChatGPT, and Claude, debuting IndexAI Insights to deliver conversational index analytics powered by large language models.
Anthropic unveils a revised Responsible Scaling Policy with a Frontier Safety Roadmap, regular Risk Reports, and clearer separation between company commitments and industry recommendations.
Google is rolling out Gemini 3.1 Pro across consumer, developer, and enterprise products, touting major gains in reasoning performance and benchmark results.