Safety

AI Alignment

The challenge of ensuring AI systems behave in accordance with human values and intentions.

Alignment is the problem of making AI do what we actually want, not just what we literally ask for. A perfectly capable but misaligned AI could be dangerous — optimizing for the wrong objective.

Current alignment techniques include RLHF, Constitutional AI (Anthropic's approach), red-teaming, and safety evaluations. Alignment is both a technical challenge (how to train models that follow instructions faithfully) and a philosophical one (whose values should AI align with?).

← All terms