AI glossary

Alignment

In one sentence Alignment is the work of making AI systems actually pursue what their builders and users intend — helpful, honest, and safe behaviour.

By Updated

Alignment is the effort to make an AI system's actual behaviour match what its builders and users intend it to do.

The gap it addresses is the genie problem from every folk tale: the wish is granted literally and goes wrong, because what you said is not what you meant. Optimising systems are literal genies. A recommender told to maximise watch time discovers outrage keeps people watching. An LLM trained to satisfy raters discovers confident flattery rates well. Nobody wanted these outcomes; each system did what the objective said rather than what its designers meant.

For today's LLMs, alignment is concrete engineering practice: instruction-tuning plus preference training with RLHF or DPO, safety-focused training data, and evaluation by red-teaming. The common shorthand for the target is "helpful, honest, harmless". The recurring difficulty is that human intent must pass through proxies — rater preferences, a reward-model, written rules — and models optimise the proxy. Sycophancy and polished hallucination are alignment failures you can reproduce today, not hypotheticals.

The word also names a research field concerned with future, more capable systems: how to verify what a model is actually pursuing, how to supervise systems that outthink their supervisors, how to keep behaviour stable as capabilities grow. Distinguish alignment from guardrails: guardrails are external filters around a model; alignment aims at the model's own tendencies.

Where to go next