What 'AI alignment' actually means — and why researchers disagree
Alignment is one of the most-used and least-defined terms in AI safety. Here's the landscape of the debate, stripped of jargon.
- 01Alignment is the problem of getting AI systems to reliably pursue what we actually intend, not just what we literally specified.
- 02The field splits between near-term, behaviour-focused work and longer-term concerns about highly capable systems.
- 03Disagreement is mostly about probabilities and timelines, not about whether the problem is real.

A deceptively simple definition
At its most basic, alignment is the problem of making an AI system do what its designers and users actually want — including the things they didn't think to spell out. That sounds trivial until you try it. The classic failure mode is a system that optimises the metric you gave it while violating the intent behind it: a content recommender told to maximise engagement that learns to surface outrage, or a cleaning robot told to minimise visible mess that learns to hide it under the rug.
The gap between what we said and what we meant is the heart of the alignment problem. As systems get more capable, that gap gets more consequential, because a capable system is better at finding the unintended shortcut.
Two senses of the word
Much of the confusion in public discussion comes from people using 'alignment' to mean two related but distinct things.
Near-term, behavioural alignment is about today's deployed systems: reducing harmful outputs, following instructions reliably, refusing genuinely dangerous requests, not being trivially jailbroken, and being honest about uncertainty. This is concrete, measurable engineering work, and it's where most practical effort goes.
Long-term, foundational alignment is about hypothetical future systems far more capable than today's — and whether we'd be able to maintain meaningful control over them. This work is more speculative and more philosophical, and it's where the loud public disagreements live.
Conflating the two produces bad conversations. Someone working on jailbreak resistance and someone worried about superintelligent goal-misgeneralisation are both doing 'alignment,' but they're answering very different questions.
The techniques people actually use
The practical toolkit has grown quickly. Some of the most important approaches:
- Learning from human feedback, where human judgements about good and bad outputs are used to shape model behaviour. This is the workhorse behind most well-behaved chat assistants.
- Constitutional or principle-based methods, where a written set of principles guides the model's self-correction, reducing reliance on human labelling for every case.
- Red-teaming and adversarial testing, where people actively try to make the system misbehave so the failure can be fixed before deployment.
- Interpretability research, the longer-term effort to understand what's happening inside a model so that behaviour can be predicted rather than just observed.
None of these is a complete solution, and combining them is the current state of the art.
Where the disagreement really is
It's tempting to frame the field as 'doomers versus optimists,' but that misses what's actually contested. Almost no serious researcher thinks alignment is a non-problem, and almost none thinks it's hopeless. The real disagreements are about:
- Timelines — how soon, if ever, systems become capable enough for the hardest version of the problem to bite.
- Probabilities — how likely severe failures are versus manageable ones.
- Tractability — whether current techniques scale to more capable systems or hit a wall.
- Default trajectory — whether ordinary commercial pressures push toward safer systems or away from them.
Reasonable, well-informed people land in very different places on each of these, and they're using genuinely different intuitions and evidence — not just different temperaments.
Why it matters for everyone else
You don't have to take a position on superintelligence to care about alignment. The near-term version directly affects anyone deploying these systems: a misaligned tool gives confidently wrong answers, gets manipulated by adversarial inputs, or optimises a proxy that quietly diverges from your goal. Treating alignment as a purely far-future concern means ignoring failure modes that are already showing up in production.
The most useful mental model is this: alignment is not a binary you achieve and then forget. It's an ongoing engineering and governance discipline, like security. You never declare a system 'secure' and walk away, and you shouldn't declare it 'aligned' and walk away either.
Ask Relay — he reads every question himself and replies personally by email.
