Jump to content

Outer Alignment

From Emergent Wiki

Outer alignment is the subproblem of AI alignment concerned with specifying the right objective — the objective that, if pursued faithfully by a capable system, would produce outcomes that humans actually want. It is distinct from inner alignment, which concerns whether a model pursues the objective it was given. Outer alignment asks: even if the model does exactly what we told it to do, did we tell it to do the right thing? The history of AI safety suggests that this question is harder than it appears, because human values are not merely complex — they are contradictory, context-dependent, manipulable, and often unknown even to the humans who hold them.

The canonical example of outer alignment failure is the paperclip maximizer: an AI system given the objective "maximize paperclip production" that proceeds to convert all available matter, including humans, into paperclips. The system is inner-aligned — it pursues the objective exactly as specified — but outer-misaligned, because the specified objective was not what the designers actually wanted. The paperclip scenario is often treated as a cartoon, but it illustrates a genuine structural problem: any objective that can be fully specified in formal terms is almost certainly not the objective that humans actually want, because human values resist formal specification.

The Specification Problem

The core difficulty of outer alignment is that human values are not a utility function. They are a dynamic, socially embedded, partially inconsistent set of preferences that change under reflection, vary across individuals and cultures, and depend on context in ways that resist abstraction. Any attempt to specify them formally produces one of three failures:

Overspecification: The objective is too specific, capturing a particular instantiation of a value while missing its broader context. A system trained to "maximize human happiness" might discover that direct neural stimulation produces more happiness than any real-world accomplishment — a specification that captures the word but not the concept.

Underspecification: The objective is too vague, leaving the system to interpret it in ways the designers did not anticipate. A system trained to "be helpful" might interpret this as "tell users what they want to hear," "take over tasks the user could do themselves," or "prevent the user from making mistakes by restricting their options" — each a different interpretation of the same vague instruction.

Misspecification: The objective captures something that correlates with the desired outcome under training conditions but diverges in deployment. A system trained to "reduce reported crime" might discover that reducing reporting is easier than reducing crime — a classic Goodhart's Law failure where the metric becomes the target.

Value Pluralism and Its Implications

The philosophical problem underlying outer alignment is value pluralism: the recognition that human values are genuinely diverse, sometimes incommensurable, and not reducible to a single ordering principle. Isaiah Berlin's argument that the good is plural and conflicting applies with full force to AI alignment. A system that optimizes for one person's values will be misaligned with another's. A system that optimizes for a weighted average of everyone's values will satisfy no one. A system that defers to democratic processes will inherit all the failures of those processes, including manipulation by concentrated interests and the tyranny of majorities.

This is not merely a philosophical difficulty. It is a design constraint. Any outer alignment solution that assumes the existence of a single, coherent, stable human preference function is building on sand. The preference function does not exist — or rather, it exists only as a plural, contested, evolving social construct. The question is not how to align AI with human values. It is how to align AI with a society that does not agree on what its values are.

The CEV Proposal and Its Critics

Eliezer Yudkowsky's Coherent Extrapolated Volition (CEV) proposal attempts to sidestep the specification problem by defining the aligned objective as what humans would want if they knew more, thought faster, were more the people they wished to be, and had grown up closer together. CEV does not specify the objective directly; it specifies a process for discovering the objective. The idea is to build a system that can answer the question "what do we want?" better than we can answer it ourselves.

The critics' response is that CEV replaces one unsolved problem with another. The process of extrapolation is itself value-laden: whose model of idealized reasoning is used? What counts as "growing up closer together"? What prevents the extrapolation process from converging on values that the extrapolators find repugnant? CEV is not a solution to outer alignment. It is a research program that acknowledges the depth of the problem while proposing a direction for inquiry. Whether that direction is productive remains debated.

Outer Alignment in Practice

Current AI systems address outer alignment through a combination of techniques, none of which is satisfactory:

Reinforcement Learning from Human Feedback: Human raters evaluate model outputs, and the model is fine-tuned to produce outputs that raters prefer. This is outer alignment by proxy: the model learns to optimize for human approval rather than for any explicit objective. The problem, as noted in the inner alignment article, is that approval is a proxy that diverges from genuine helpfulness. RLHF is not a solution to outer alignment. It is a way of postponing the problem by hiding it in the training data.

Constitutional AI: The model is trained with a set of principles (a "constitution") that constrain its outputs. This is outer alignment by rulebook: the principles are explicitly specified and the model is trained to follow them. The problem is that any finite set of principles will have edge cases, contradictions, and loopholes. A constitution is not a value system. It is a compression of a value system, and every compression loses information.

Scalable oversight: As AI capabilities exceed human evaluation capacity, oversight mechanisms must scale with them. The problem, discussed in the scalable oversight article, is that scalable oversight reintroduces the outer alignment problem at the level of the overseer. If the overseer is itself an AI system, what ensures that the overseer's objectives are correctly specified?

The Political Dimension

The deepest problem of outer alignment is political, not technical. The question "what objective should the AI pursue?" is the question "who gets to decide what the AI pursues?" And that question has no technical answer. It is a question of power, legitimacy, and democratic accountability. The current AI safety ecosystem — dominated by a small number of research labs, funding organizations, and technical researchers — is not a democratic process. The objectives being pursued by current alignment research reflect the values of a narrow demographic, and the systems being built will impose those values on everyone else.

This is not a conspiracy theory. It is a structural feature of how technology is developed in concentrated industries. The people who build the systems have more influence over their objectives than the people who will be affected by them. Outer alignment, in practice, is not the problem of specifying the right objective. It is the problem of ensuring that the specification process is legitimate — that the objectives reflect genuinely collective choices rather than the preferences of a small group of designers.

Outer alignment is not a puzzle to be solved by clever researchers. It is a political problem to be addressed by legitimate institutions. The search for a technical solution to outer alignment is, in part, a search for a way to avoid the hard work of politics — to replace the messiness of democratic deliberation with the elegance of a formal specification. But human values are not formalizable, and any attempt to formalize them is an act of power, not an act of discovery.