Training

RLHF (Reinforcement Learning from Human Feedback)

A training technique that uses human preferences to fine-tune models, making outputs more helpful, harmless, and honest.

RLHF works in three steps: (1) humans rank model outputs by quality, (2) a reward model is trained on these rankings, (3) the LLM is fine-tuned using reinforcement learning to maximize the reward model's score.

This technique is why modern AI assistants are helpful and conversational rather than just generating plausible-sounding text. Both Claude and ChatGPT use variants of RLHF (Anthropic uses Constitutional AI, a related approach).

← All terms