Nvidia in Talks to Back Data Startup Mercor at $20 Billion
Nvidia has discussed investing in Mercor, a data-labeling startup that supplies expert data for its Nemotron models, in a round that would double the company’s valuation to $20 billion.
Explore AIstify's latest reporting, research, and expert analysis tagged with "training data", collected in one continuously updated archive.
Nvidia has discussed investing in Mercor, a data-labeling startup that supplies expert data for its Nemotron models, in a round that would double the company’s valuation to $20 billion.
Synthetic data is artificially generated information used to train, test, or evaluate AI systems when real data is limited, sensitive, or costly.
Fine-tuning adapts a pretrained AI model to a specific task, domain, or behavior by continuing training on targeted examples.
Pretraining teaches an AI model broad patterns from large datasets before the model is prompted or adapted for specific downstream tasks.
Bias is a systematic distortion in AI results that can produce unfair or inaccurate outcomes for particular groups, situations, or data patterns.
AI startup Recursive Superintelligence has raised $500 million from Nvidia and GV to pursue self-improving AI systems, despite having no public product.
LinkedIn is testing a new platform that lets users earn up to $150 per hour training AI models. The move taps into fast-growing demand for human feedback in AI.
Data augmentation expands a training dataset by creating realistic variations of existing examples while preserving their intended labels.
Encyclopaedia Britannica and Merriam-Webster have sued OpenAI, alleging the company copied nearly 100,000 articles and dictionary entries to train ChatGPT without permission.
Instruction tuning trains an AI model on instruction-and-response examples so it follows natural-language tasks more reliably across different use cases.
Discover how AI learns from data, improves over time, and makes decisions. A beginner-friendly guide to AI training and its real-world applications.
Guide Labs launched Steerling-8B, an 8-billion parameter LLM with interpretable architecture that allows every token to be traced back to its training data.
Active learning is a machine learning approach that selects the most informative examples for human labeling to improve a model with less annotation work.
The dataset used to teach AI models how to perform tasks. It helps systems recognize patterns, make predictions, and improve performance through iterative learning.