Home Glossary Data Labeling

Data Labeling

Data labeling is the process of adding meaningful annotations to raw text, images, audio, video, or other records. Labels may identify objects, sentiment, intent, medical findings, speech segments, preferred responses, or numerical outcomes. They provide the target information used in supervised training and often support validation and testing as well. High-quality labeling requires clear guidelines, representative samples, trained annotators, quality checks, and a process for resolving disagreement. Poor or inconsistent labels can cap model performance and encode social or institutional bias. Sensitive projects must also manage privacy, worker conditions, provenance, versioning, and the possibility that the correct label is uncertain rather than absolute.

Related News

Andrew Tulloch Leaves $12B AI Startup to Join Meta After Turning Down $1.5B Offer
By • 3 mins read
AI & Machine Learning, Immersive Reality (AR, VR, MR, and XR), News, Startups & Investment

Andrew Tulloch Leaves $12B AI Startup to Join Meta After Turning Down $1.5B Offer

By • 3 mins read

Andrew Tulloch, co-founder of the $12 billion AI startup Thinking Machines Lab, has joined Meta after previously rejecting what reports described as a $1.5 billion offer — a figure Meta has since called ‘inaccurate and ridiculous.’