AI Glossary term

RLHF

Training a model on human preferences between candidate answers, so it learns which responses people actually want.

Also called: reinforcement learning from human feedback

Training a model on human preferences between candidate answers, so it learns which responses people actually want.

Why it mattersThe step that turned raw text predictors into assistants worth talking to.

See also Alignment, Fine-Tuning

Explore more AI terms

Browse the full glossary for plain-English definitions across models, agents, data, and safety.

Back to all terms