Home Glossary Latency

Latency - Page 70

Latency measures the delay between an AI request and the resulting response. Interactive generative systems often track time to first token, which reflects how quickly output begins, and time to last token, which captures total completion time. Latency comes from model computation, queueing, network transfer, retrieval, tool calls, safety checks, and application logic. It varies with prompt length, output length, model size, hardware, batching, cache use, and current load. Optimizing only average latency can hide poor experiences, so teams also monitor percentile values and separate each stage of the request. The right target depends on whether the product is conversational, analytical, offline, or safety-critical.

Andrew Tulloch Leaves $12B AI Startup to Join Meta After Turning Down $1.5B Offer
By • 3 mins read
AI & Machine Learning, Immersive Reality (AR, VR, MR, and XR), News, Startups & Investment

Andrew Tulloch Leaves $12B AI Startup to Join Meta After Turning Down $1.5B Offer

By • 3 mins read

Andrew Tulloch, co-founder of the $12 billion AI startup Thinking Machines Lab, has joined Meta after previously rejecting what reports described as a $1.5 billion offer — a figure Meta has since called ‘inaccurate and ridiculous.’