Skip to content
Back to Glossary
Fundamentals

Reinforcement Learning from Human Feedback (RLHF)

Definition

A training technique that refines language model behaviour by learning from human preferences rather than fixed labels. RLHF is a primary method used to align LLMs like ChatGPT and Claude with desired values and reduce harmful outputs.

Related Services

Knowing the Terms Is Step One. Applying Them Is Step Two.

Book a Physical AI Fit Call to discuss how these AI concepts translate to your specific industry and business challenges.

Reinforcement Learning from Human Feedback (RLHF) | AI Glossary — LLM, RAG & 20+ Key Terms Explained