Home Glossary Reinforcement Learning from Human Feedback (RLHF)

Reinforcement Learning from Human Feedback (RLHF) - Page 45

Reinforcement Learning from Human Feedback, or RLHF, is a method for shaping model behavior with human preferences. Reviewers compare or score candidate responses, those judgments train a reward model, and reinforcement learning then adjusts the language model to favor outputs that receive higher predicted rewards. RLHF can improve instruction following, helpfulness, tone, and safety beyond basic pretraining. Its results depend heavily on who provides feedback, how instructions are written, and whether the examples represent real users and edge cases. The process may reward superficial agreement or hide uncertainty, so it is commonly combined with automated evaluations, red teaming, policy rules, and ongoing post-deployment monitoring.

Nvidia Says $100B Investment Into OpenAI Is Likely Off the Table
By • 3 mins read
AI & Machine Learning, News, Startups & Investment

Nvidia Says $100B Investment Into OpenAI Is Likely Off the Table

By • 3 mins read

Nvidia CEO Jensen Huang said the company’s $30 billion investment in OpenAI could be its last before the AI startup pursues an initial public offering. The chipmaker also indicated its $10 billion investment in Anthropic may mark the end of its funding commitments to major AI model developers.