Deep Learning

Reinforcement Learning from Human Feedback (RLHF)

Training AI models using human preferences as feedback signals. Used by OpenAI (GPT), Google (Gemini), and Anthropic (Claude) to make models safer and more helpful.

Use this term as a quick reference, then jump back into the part of the site where it matters.

Explore next