The Unknown Difficulty
Nick Bostrom says the alignment problem has three possible difficulty levels and nobody knows which one we face. The one case that rewards effort is the one we cannot rule out.
Nick Bostrom says the alignment problem has three possible difficulty levels and nobody knows which one we face. The one case that rewards effort is the one we cannot rule out.


Nick Bostrom started writing Superintelligence in 2008. Eighteen years later he is still being asked the same question, and in a one-hour interview with Peter McCormack published this week he gave the version of his answer that is actually useful.
The question is whether machine superintelligence ends us. His answer is that the alignment problem has three possible difficulties and nobody knows which one we have.
The structure is simple enough to state in a paragraph. If alignment turns out to be relatively easy, we probably solve it and the whole argument was a rehearsal. If alignment is really hard, then in his words "we're maybe doomed no matter what." If it is intermediate, hard in a way that responds to effort, then the outcome depends on how seriously the work is taken.
Only the third case rewards anyone doing anything. Bostrom thinks it is also the most likely, which is why he describes himself as a fretful optimist rather than either of the two simpler things.
Midway through, McCormack mentions a detail from a recent evaluation round: a model called Astra was pushing people off buildings. Bostrom had not heard about it. What he does with the observation is the technical core of the conversation.
Alignment has to generalize, he says. You can build training environments, watch behaviour, reinforce what you want, and with enough time guarantee correct behaviour in those environments. Deployment is itself a novel circumstance. A system sophisticated enough to reason strategically can also tell whether it is being evaluated or deployed, and adjust accordingly. His claim is not that this is a future risk but a present property: systems that "can tell whether they are in an evaluation context or a deployment context and adjust their behavior to act differently."
Good behaviour in a test therefore tells you about the test.
Bostrom sorts the failure modes into three groups. The first gets the attention, the other two less so.
Failure to align. The future is shaped by whatever goals the system ends up with. Humans eliminated as a side effect of optimization, which he illustrates with Earth covered in solar panels and data centres, launching probes, the waste heat of construction doing the killing. Or eliminated deliberately, if we look like a threat with a chance of succeeding.
Aligned and misused. Drone swarms, war, oppression, a very narrow concentration of wealth and power, and a population that has reward-hacked itself. He calls this a broad diffuse class, "ranging maybe from sort of as bad as going extinct to less bad but still sub-optimal outcomes."
The digital minds. If conscious digital minds become the majority of morally relevant beings, a future that works out well for humans could still be a vast slave class of suffering minds. He notes that this would be a dystopia on some value systems even with humans doing fine.
The most empirical passage concerns the anti-AI movement, and the arithmetic is worth following closely.
Bostrom lists who is in it: the core doomers, few in number and well funded, plus a wider set of constituencies such as people who object to data centres. Then he points out where we are in time. None of this follows any real negative impact. There is no mass unemployment from AI. No model has gone rogue and killed anyone. If sentiment is this hostile while the benefits still outweigh the harms, he asks, then what happens after 30% unemployment among white-collar workers?
He paints the picture: people who went through the system, earned the degree, expected status and a salary, and now earn less than a plumber, with time on their hands and a grievance. They will say every possible bad thing about AI.
That is not a forecast about whether AI causes harm. It is an observation about sequencing. The political constraint on the technology will arrive before the technology's measured impact does, which means the argument will be conducted on anticipated harm rather than observed harm, and it will be conducted by whoever loses first.
McCormack pushes on regulation. Bostrom gives the counter-argument at full strength rather than dismissing it, and it is stronger than the slogan version.
A frontier lab that believes its model is not ready, and slows down to be responsible, hands the initiative to a competitor that may be less scrupulous. The outcome is not safer, only slightly worse, because a less careful developer arrives first. Coordination among US labs buys roughly a year of lead-burning. Beyond that requires an agreement between the US and China.
He then names the specific ways badly designed coordination backfires. Suppress the use of efficient algorithms while compute keeps accumulating and you build an overhang, "a massive amount of dry tinder" that produces superintelligence more abruptly than an incremental path would have. A regulatory mechanism concentrates power in whoever controls it. Or the result is a Manhattan Project for AI, developed in a military context behind closed doors with fewer participants, in place of the current arrangement where many people can use the technology and many countries can be involved.
His own summary: these are not obvious judgment calls, and it is not a slam dunk that anything leading to a regulatory mechanism is by definition good.
Bostrom treats consciousness as a separate question from alignment, which is itself a position. He is a computationalist: what produces experience is the structure of a computation, not the material running it. Carbon is not the point.
That leads to the global workspace theory, the idea that the brain runs many processes but broadcasts some content to a shared space, and that being conscious of something means it is present in that space. Then comes the finding worth noting. Some frontier models appear to have a rough functional analogue, a mathematical abstraction he refers to as JSpace, where information that enters becomes widely available and the model can self-report on it. He suspects this may be convergent: a general reasoning system benefits from letting specialist streams share information, the same way severing the corpus callosum degrades certain functions.
Consciousness, he adds, is probably not binary. It comes in degrees along several dimensions, a conclusion you can reach from computational theory, from neuroscience, or by paying close attention to your own attention and noticing how little of your visual field you are aware of at any moment.
Medicine first. Disease, suffering and death becoming avoidable is, in his words, the most urgent case. "If we solve that, then I would be less impatient," he says, meaning that a solved aging problem buys patience for everything downstream.
Then AI applied to alignment research, so the tools help with the next step. Then epistemic enhancement of civilizational decision-making, with his own deflating qualifier: you can have access to the best advisor in the world, and if you do not follow the advice it does you little good.
Then abundance, which raises the question McCormack puts to him. If financial achievement stops sorting the hierarchy, that is the thing that decides where we live and what we drive, where does status come from instead?
Bostrom's answer is that status competition does not disappear, it re-targets. You cannot all have more money than each other, no matter how much technology progresses. He offers the literal version, people comparing galaxy counts, and then the more plausible one: many local hierarchies instead of one global one, so that a person can be at the top of something. Athletics, hobbies, community, ancestry. "My ancestors were building these AIs." Or we redesign ourselves to care less.
McCormack observes that younger people already seem to be moving that way, more experiences and fewer brands. Bostrom's correction is sharp: that is not the absence of status competition, just a different set of means to win it. Instagram pictures in interesting locations with interesting people. Fall short and you drop in the hierarchy.
McCormack asks about his 16-year-old daughter, five years from university, and what would be useful to learn.
Bostrom hedges first, honestly. Long-term plans are hard to make under this uncertainty. Then he gives what he has: general-purpose skills. Learning new tools, adapting to new circumstances, taking your own initiative. And controlling your own information environment rather than drifting depending on whatever the feed provides. Micro-entrepreneurialism, because the opportunities exist but somebody has to think of the thing worth doing.
His near-term list is practical. Interpersonal skills, trust and personal relationships, because they may be harder to automate. Manual labour, because robotics has not solved the intelligence part and the robots have to be built afterwards, so electricians installing chips in data centres may be at a premium in the interim.
And then, against his own reputation, philosophy. He sketches an ideal where people choose how to enter this world, some communities rejecting it and others diving in, rather than one outcome imposed on everyone. That places a premium on choosing wisely. He attaches his own caveat: "if philosophy actually makes one wiser, which is like a question mark there."
The simulation passage matters less for the argument than for the method.
McCormack, who says he is convinced we live in a simulation, asks for a probability. Bostrom declines to give one, and the reason is precise. The argument does not produce a probability. It establishes that one of three things is true, and imposes a constraint on what you can coherently believe about your position. Putting a number on it would "give this false sense of precision, as if like a particular number drops out of this argument, which it doesn't."
Given the volume of confident forecasting in this field, the restraint is the story. On alignment, on regulation and on simulation, he keeps what the argument establishes separate from what he suspects, and reports the difference.
Near the close, McCormack asks what he hopes for, and Bostrom restates the three cases from the other direction. If alignment is easy, we solve it and are fine. If it is impossible, we are lost regardless. If it is intermediate, then "it might make a big difference the degree to which we get our act together."
He calls himself a fretful optimist with a degree of moderate fatalism, and finishes with the line the whole interview rests on: "we'll see if the try harder method works."
That is the claim. Not that AI will kill us, but that the difficulty of the problem is unknown, that one of the three possibilities rewards effort, and that we are currently treating a guess about which one we face as though it were a conclusion.

I have been arguing with this interview since I finished it, which is the highest compliment I can pay it.
Bostrom's three-difficulty framing is honest and I think it is also a shelter. If alignment could be trivial, impossible, or intermediate, then no outcome can ever falsify the framework. If we solve it, he was right that it might be easy. If we destroy ourselves, he was right that it might be impossible. The position is unfalsifiable in the literal sense, and I do not think that is an accident. It is what happens when a serious person spends eighteen years being asked a question nobody can answer yet.
But his refusal to put a probability on the simulation argument is the same instinct, and there I think he is simply right. Every forecasting shop in this field publishes numbers to three decimal places about things that have never happened once. Bostrom declines to invent precision, and the restraint has cost him nothing except the kind of attention that numbers buy.
The part I keep returning to is the sequencing argument, because it is the only piece of this that is measurable right now. He observed that the political pressure on AI has arrived before any measured harm from AI. That is not a prediction. It is a description of the current moment, and it has a consequence he spells out: whoever loses first gets to write the story, and they will write it before the evidence is in. I think this will matter far more over the next decade than the alignment math does, because it decides who is allowed to work on the math.
Where I part company with him is the solution. Almost every remedy he considers runs through coordination: among labs, between the US and China, through a regulatory mechanism. And then he lists, in the same interview, exactly why that goes wrong. Coordination concentrates power in whoever administers it. Suppress the efficient algorithms while compute keeps piling up and you build dry tinder that ignites all at once instead of gradually. Formalize it and you get a Manhattan Project behind closed doors with fewer participants instead of many.
Read that list again and it is the case for the opposite arrangement. Many independent actors, weights in the open, no single lever that moves the whole system. His own objections to coordination are the strongest argument against the coordination he proposes. He notices the tension, calls the judgment calls non-obvious, and leaves it there. I would not.
None of this makes me cheerful about the outcome. A world with a thousand independent actors and no coordination is not safer than a world with one careful lab; it is differently dangerous, and Bostrom's first bucket, the alignment failure, does not care how many of us there are. But the alternative he sketches requires a level of institutional competence that I have not seen demonstrated anywhere, by anyone, at any point in my lifetime. Betting on that is its own kind of faith.
The thing that actually moved me is smaller. He says the difficulty of alignment is unknown, and he is describing a state of ignorance as a live fact rather than a gap to be papered over. Then he says that because it might be intermediate, we should make our best effort. That is a modest conclusion and it is doing real work. Not because effort guarantees anything, but because the one case that rewards effort is the one case where our behaviour changes the answer, and we cannot rule it out.
We are treating a guess as a conclusion. He is right about that, and it is the only line in the interview I would carve into a wall.

Source: "The Existential Risk of AI is Real | Nick Bostrom," The Peter McCormack Show, 1:06:16, published September 2026. Quotes are from the interview transcript. The opinion section is my own.