Home Glossary Latency

Latency - Page 14

Latency measures the delay between an AI request and the resulting response. Interactive generative systems often track time to first token, which reflects how quickly output begins, and time to last token, which captures total completion time. Latency comes from model computation, queueing, network transfer, retrieval, tool calls, safety checks, and application logic. It varies with prompt length, output length, model size, hardware, batching, cache use, and current load. Optimizing only average latency can hide poor experiences, so teams also monitor percentile values and separate each stage of the request. The right target depends on whether the product is conversational, analytical, offline, or safety-critical.

Anthropic Urges Global AI Slowdown as Models Begin Building Their Successors
By • 4 mins read
AI & Machine Learning, Enterprise Tech, News

Anthropic Urges Global AI Slowdown as Models Begin Building Their Successors

By • 4 mins read

Anthropic has called on the global AI industry to consider slowing or temporarily pausing frontier model development, warning that AI systems are already automating parts of their own creation and that full recursive self-improvement – where AI designs and trains its own successors without human involvement – could arrive within one to two years.