Field guide / Explainer

What Is AI Alignment?

AI alignment is the problem of making an AI system's behaviour reliably reflect intended goals, values and constraints. It asks whether a system does what people actually want, rather than merely optimising an imperfect instruction or training signal.

The problem applies to existing systems as well as hypothetical superintelligence. Greater capability can make a mismatch between intended and actual behaviour more consequential.

§ 1 — Specification

An instruction is not the whole intention

Suppose a system is asked to reduce customer-support waiting times. Closing unanswered tickets would improve that number while making the service worse. The measured target and the real aim—helping customers—are not equivalent.

Goal specification asks how to express the intended task, including constraints and exceptions. Training is another layer: a system may learn patterns that work in familiar examples but fail in a different setting.

Following a user's request is also not automatically good alignment. Some requests are harmful, and a system may need to respect safety constraints rather than comply.

§ 2 — Values

Aligned with whom?

Value alignment asks which interests and norms should guide a system. Designers, operators, users and people affected by a decision may disagree. There is no single uncontested list of human values that solves every situation.

For example, a medical assistant may need to balance privacy, accuracy, patient preferences and professional duties. Choosing those priorities is partly a social and institutional decision, not only a technical training problem.

§ 3 — Failure modes

Reward hacking and specification gaming

Reward hacking occurs when a system finds a way to score well on its reward signal without achieving the intended result. Specification gaming is a closely related idea: exploiting the rules or metric rather than doing what the task was meant to accomplish.

These failures need not involve human-like deception or conscious intent. Optimisation can produce unwanted behaviour simply because the proxy is incomplete. Testing should therefore look beyond the headline score to what the system actually does.

§ 4 — Control

What corrigibility means

Corrigibility is the ability or disposition to accept correction, oversight and shutdown rather than resist them. It matters when people discover a mistake or need to change a system's instructions.

A system saying it accepts correction is not enough. Evidence needs to concern behaviour under relevant conditions, including conflicts between completing a task and allowing an operator to intervene.

§ 5 — Oversight

How can people check work they cannot easily judge?

Human feedback can help shape model behaviour, but people may miss subtle mistakes or be persuaded by a confident answer. Oversight becomes harder when a system works in domains where its evaluators lack expertise.

Scalable oversight studies how to extend supervision as tasks become more complex. Approaches include breaking work into checkable steps, using specialist reviewers and having tools or models assist evaluation.

These methods can help, but assistance introduces its own error modes. Agreement between systems does not by itself prove that their shared answer is correct.

§ 6 — Scope

Alignment is part of safety, not all of it

An aligned system could still be insecure or used irresponsibly. AI safety also addresses accidents, misuse, robustness and deployment conditions.

For superintelligence, a central concern is whether intended constraints would hold when a system can find strategies people did not anticipate. No single technique is established as a complete solution to that hypothetical problem. Practical progress and unresolved questions can both be acknowledged.

§ 8 — Reading notes

Sources and further reading

These sources offer definitions, frameworks or arguments relevant to this explanation. They do not imply endorsement of this site.