Alignment Isn't Solved. Here's What's Missing.

Some researchers are declaring victory on AI alignment. They're confusing one kind of progress with the whole problem. The structural gap (whose values, enforced how, auditable by whom) remains wide open.

Share this article

Alignment Isn't Solved. Here's What's Missing.

Adrià Garriga-Alonso quit his AI safety job in December because there was "no point" doing more alignment work. David Dalrymple dropped his estimate of AI-caused human extinction from 40-50% to 5-8%. As Lynette Bye reports in Transformer this week, the mood among some alignment researchers is shifting from existential dread to cautious optimism.

Some of this optimism is earned. Pretraining on human text means language models absorb human values as a side effect of learning language. Incremental capability gains mean researchers can iterate rather than getting alignment right on the first try. Model organisms research provides controlled environments to study misalignment before it appears in the wild. These are genuine advances.

The conclusion some are drawing, that alignment is mostly handled, confuses one kind of progress with the whole problem.

Understanding values and upholding them are different things

Dario Amodei and Garriga-Alonso both argue that LLMs "understand human values fairly innately" from pretraining. A model trained on human text can describe kindness, fairness, honesty. It can reason about ethical dilemmas. But understanding a value and being committed to it are different operations. A student can ace an ethics exam and still cheat on it.

Anthropic's constitution for Claude exists precisely because pretraining comprehension doesn't produce alignment on its own. The model needs external structure: written principles that guide behaviour beyond what the training data provides. The constitution's existence is an admission that absorbing values from text is a starting point, not a finish line.

"Human values" assumes a consensus that doesn't exist

The entire discourse around alignment treats "human values" as a coherent set. A parent in Seoul, a doctor in Lagos, a compliance officer in Frankfurt, and a teenager in São Paulo need different things from their AI. A medical AI requires different ethical constraints than a children's educational tool. A financial advisor AI in the EU operates under different regulatory frameworks than one in Singapore.

One alignment profile baked into model weights cannot serve this diversity. The attempt produces either lowest-common-denominator safety that frustrates everyone, or a single cultural perspective presented as universal. Neither is alignment. Both are approximations that degrade as AI reaches more of the world's population.

The practical alignment problem is contextual: whose values, applied when, with what priority, and reviewable by whom.

The gap between solvable and solved

Jan Leike at Anthropic is direct about this: "Just because a problem is solvable, this doesn't mean it's solved. We have to actually keep doing the work to get it done."

Stuart Russell goes further. He thinks safe AI is achievable, but companies "need to make AI systems millions of times safer" to bring risk down to levels society accepts from comparable technologies. Nuclear reactors. Aviation. Pharmaceuticals. The safety margins required for technologies that affect millions of people are orders of magnitude beyond what any AI company currently delivers.

Market incentives will not close this gap on their own. Ryan Greenblatt at Redwood Research puts it plainly: "Risk is very elastic to how much people try." When competitive pressure pushes companies to ship faster, safety margins shrink. Russell's conclusion is that regulation is necessary. He's right. But regulation alone is an empty vessel.

Regulation needs infrastructure

You cannot audit "be more aligned." Regulators need concrete artefacts to evaluate. What values is this AI operating under? Who specified them? Were they followed? What happened when they weren't?

Answering these questions requires four things the alignment field has largely ignored:

A protocol for expressing values. A machine-readable, portable format for declaring what an AI should and should not do. A standard that travels with the user across providers, so alignment choices are not locked to one company's model weights.

An enforcement mechanism. Something that evaluates AI actions against declared values in real time. Faster and more structured than hoping the model behaves.

An audit trail. A record of which principles applied to which decisions and why. Reasoning traces that regulators, developers, and users can inspect. Input/output logs tell you what happened. Audit trails tell you whether it was consistent with the values in force.

Transportability. A user's alignment configuration should survive a provider switch. When you move from one AI service to another, your values should not reset to someone else's defaults.

This is what we build at Creed Space. The Value Context Protocol provides the standard. Creeds, which are portable, signed, composable value declarations, provide the auditable artefact. The Policy Decision Point provides real-time enforcement. The audit system provides transparency.

Russell is right that AI needs to be millions of times safer. Regulation alone cannot deliver that. Infrastructure can.

Alignment is maintenance, not a milestone

The question "Is alignment solved?" treats it as a problem with a solution: a box to check, then move on. This is the wrong frame.

Alignment is infrastructure. Like water treatment or financial regulation, it requires continuous operation, adaptation, and oversight. The technology changes. The threat surface changes. The social context changes. Alignment infrastructure has to change with them.

Someone quitting safety research because "there's no point" is like a water engineer quitting because the treatment plant works today. The plant works because people maintain it. Alignment will work, when it works, because someone builds and maintains the systems that make it operational.

The technical advances are real. The structural gap is where the work remains: protocols, standards, enforcement, audit. Filling that gap is less exciting than training the next model. It is also what determines whether the next model is safe to deploy.

Plumbing is what turns capability into civilisation.

Nell Watson
Founder, Creed Space

AI ethics researcher and IEEE Fellow. Author of Taming the Machine.