RLHF works in three steps: (1) humans rank model outputs by quality, (2) a reward model is trained on these rankings, (3) the LLM is fine-tuned using reinforcement learning to maximize the reward model's score.
This technique is why modern AI assistants are helpful and conversational rather than just generating plausible-sounding text. Both Claude and ChatGPT use variants of RLHF (Anthropic uses Constitutional AI, a related approach).