← Back to Glossary Index
AI Glossary Term

RLHF

Definition

Reinforcement Learning from Human Feedback. Aligning model behavior by training reward models based on human ranking preferences.

Example Case

Tuning GPT-4 base models to avoid generating offensive answers.