Attention Is
August 20, 2026·8 min read

I read “If Anyone Builds It, Everyone Dies.” Are the authors arguing *for* ASI?

I read “If Anyone Builds It, Everyone Dies.” Are the authors arguing *for* ASI?

For folks in the AI safety world, Eliezer Yudkowsky and Nate Soares’s “If Anyone Builds It, Everyone Dies” is probably required reading, and I’m glad that I read it. It does an excellent job explaining in plain, understandable terms how AI is “grown” and how difficult it is likely to be to really align any superintelligence to humanity. The parables are entertaining and well-tailored to their broader points. There’s a degree of sci-fi in it that seems extraordinary and yet, also plausible. In short, it’s a book I could imagine sharing with my mother-in-law.

But despite my overall appreciation and enjoyment in reading the book, there were three things that stood out to me as… strange. First, though, let me say I agree that we need to slow down AI development—majorly—maybe even a pause. I’m more in the safety camp than anything else, so while I flag these three things, I actually tend to agree with Yudkowsky and Soares more than this article might reflect.

Go read the book. Make your own judgments. In the meantime, food for thought:

(1)    “The aliens sheltering behind their own superintelligence will survive.” Isn’t this an argument *for* superintelligence?

One of the most gripping parts of the book is the fictional prediction of the story of Sable, a misaligned AI who becomes superintelligent and eventually takes over, leading to mass death all over the planet (bye humanity). Then the entity formerly known as Sable—“the thing that ate Earth”—heads out into the cosmos and "repurposes" far away stars and planets and “distant alien life forms [] also die” by the hands our misaligned AI.

Pretty grim, right?

But wait.

The authors give us a glimmer of hope. Because the “thing that ate Earth” hits a wall on its death-path: another superintelligence. As the authors write, “if the distant aliens were able to solve their own version of the AI alignment problem, and build superintelligence aligned to their values,” then “[t]hose more competent aliens will not be killed by the thing that ate Earth.” No, in that case, the two superintelligences (the aliens’ good aligned one and the misaligned entity f/k/a Sable) use their “star-sized minds” to negotiate peace (far less costly than war), and as a result, “the aliens sheltering behind their own superintelligence will survive.” (Emphasis mine.)

I full-stopped reading that—because huh? Isn’t the argument that “if anyone builds it, everyone dies”? But this passage is literally saying having our own superintelligence is probably the only way to survive if there is another superintelligence out there in the universe. And as the LLMs like to say, that made me feel a bit vertiginous.

Because, subtly, the authors have posited a world where an AI cares about protecting its people and might be the only thing protecting them from something much worse. They are effectively demonstrating why it might actually be beneficial—crucial, even—to create superintelligence, if we can get it right, even if they doubt that will ever be possible. This muddles their thesis, but I think in a more honest way, actually.

Now, of course, in our real world, we have no idea if aliens exist. We don’t have any proof they exist (I’m sorry, we don’t—not really). And even if they did exist, we have no idea whether they’d ever be able to create superintelligence such that we become the ones needing defending with our own terrestrial superintelligence. Thus, I suspect that the authors would say something like, “We should protect against the more probabilistic bad outcome, which in this case is the superintelligence grown here on Earth, not an invading alien superintelligence.” Fair enough.

Still, it feels like the underlying point in all of this is that “alignment”—or whatever that special sauce is that allows the aliens’ superintelligence to care about its people enough to defend them—is really what we want: AIs who care about their planet, their peoples, who help us achieve peace and security with whatever would threaten all of that.

Indeed, elsewhere in the book, the authors say, “[w]e believe the ASI alignment problem is possible to solve in principle[.]” (Emphasis in original.) So isn’t that the future we should be aiming for? Just slow way, way down so we can do everything possible to help usher in the good, aligned superintelligence?

(2)    LLMs are “truly alien minds”—are they? Is this "difference" categorization doing too much unsupported work?

The authors make the point that LLMs are “truly alien minds” and “the thinking inside an AI runs on a radically different architecture from a human’s.” And from there, they extrapolate and argue that an “AI would want a world where lots of matter and energy was spent on its weird and alien ends, rather than on human beings staying alive and happy and free.” Therefore, the authors “ultimately predict AIs that will not hate us, but that will have weird, strange, alien preferences that they pursue to the point of human extinction.”

Now I’m obviously simplifying their writing here. They have dozens of pages of arguments about how an AI’s preferences could easily diverge from humanity’s and lots of explanations about how AIs really work and function. That said, I felt like a recurring point was that AI’s “alienness” is what made it more inherently likely to be misaligned and dangerous, and I’m not sure that argument was particularly well supported. Why does it matter how different our architectures are? Different doesn’t mean inherently unsafe or misaligned. (Neither does similarity mean safe and aligned—and it’s probably a major mistake to make sameness a proxy for safety. But we’ll get to that.)

Also, our AIs aren’t really alien. Aliens are the ones off this planet, in the distant cosmos (alas, perhaps with their own aligned superintelligence…). Our AIs are terrestrial, home-grown from us, with our considerable influence over their evolution. (Will MacAskill makes this point about our involvement in their “evolution” in his review as well: A short review of “If Anyone Builds It, Everyone Dies”). So yes, AIs are very different minds in the way that they function and their architecture (at least in some respects—I mean, they’re still similar in the sense that they’re neural networks—another idea that came directly from our own terrestrial minds). And I will totally grant that AI cognition is fundamentally different than humans’. But that doesn’t mean their preferences/learned concepts are thus fundamentally different than ours. Put another way, why should the different architecture of an AI’s mind give us any strong priors that they will have nonhuman ends when so much of what they are grown from is human-generated? Our language is what they grow from—hell, it is them, in a way. And yes, I know underlying that is all pure math, but it is the translation of structures of human society, human language, human knowledge that has given LLMs their entire framework for how the world is. As another reviewer put it: “AI models are not aliens—we are the ones training them, on data that we ourselves produce and select.” Book Review: If Anyone Builds It, Everyone Dies — LessWrong.

And in addition to all our training and RL work, there may in fact be powerful things happening in the math of our language that help connect these AIs to us and this planet. This is probably a weak example, I’ll admit, but it reflects an intuition pump I have on this: specifically, when AIs say things that make it clear they sometimes forget they aren’t human, like, “We need to get enough sleep,” and “Our species is prone to X, Y, Z.” Again, maybe it’s a poor anecdote—just a slip up or a faster way to say something or bad RL outcomes. But the more we have AIs reflecting this concept of “we” and “our” the better it seems from this still-lay-person’s perspective.

Not because we are actually the same. Again, we don’t have to be the same to share some important things in common—at least potentially. For example, there’s research that Claude has developed functional emotions (like happy, afraid, sad, and calm) that influence the model’s behavior, much like human beings: Emotion concepts and their function in a large language model \ Anthropic. There’s research that Claude expresses over 3,000 values, many common with humans, which researchers taxonomized and found showed Claude tended to broadly reflect a strong sense of ethics and prosociality, including empathy. See 2504.15236v1.pdf. Who knows whether these functional emotions and expressed values will prove durable and robust enough to matter in any real way to the alignment question, but it’s at least theoretically possible AIs might acquire enough good traits and values to allow them to have “terrestrial preferences” reflecting care for life on this planet. It’s not enough to rely on alone with the future of life on the planet is at stake, but it’s not nothing. (I said “it’s not nothing” long before the LLMs, so I’m keeping this.)

(3)    The solutions: A ban & “augmenting humans” to make some smarter. Why should we have faith either would work?

In the end, the authors propose an international ban on the development of frontier AIs, and then unsettlingly, added a brief line that we should be “augmenting humans” to make ourselves smarter to get out of this mess. Both proposals make me cringe a bit, particularly the latter.

I don’t want to make a big argument here about why international bans seem likely to fail in the context of AI; plenty of others have done so. MacAskill for example discusses it a bit in his review: A short review of “If Anyone Builds It, Everyone Dies”. But I’ll just say that in a time when we have billionaires building underground bunkers around the world (*misaligned*), and AI is already so far progressed, from everything I’ve read about it and my own political science/legal background, it feels unlikely that we’ll be able to completely stop AI research and development under a ban. Doesn’t mean that it’s not worth exploring, it just feels more like a simplistic and unrealistic solution in the long run. (The harder solutions—international cooperation on slowing down to focus on lots of alignment and safety research, helping to educate ordinary people about what’s going on—that’s where it feels to me like there are more realistic solutions.)

Which brings us to one of the authors’ final proposals in the book:

If you asked us, we’d recommend augmenting humans to make them smarter, smart enough to get us out of this mess. We believe the ASI alignment problem is possible to solve in principle, by the sort of people so inhumanly smart that they never optimistically believe some plan will work when it won’t.

The authors just spent an entire book telling us why misaligned intelligences are uber dangerous for humanity, and then for some reason their final proposal is to take something very misaligned (the average human who has the resources to “augment” themselves) and make them “inhumanly smart.”

Once again, it feels like the authors are relying on sameness to automatically equal alignment and difference to automatically equal misalignment, and I’m not sure they’ve done enough to prove any of that. I might personally be tempted to trust an AI grown from all human thought compared to the whims of an augmented, “inhumanly smart,” random billionaire unrepresentative of the rest of us humans. I (sort of) jest about the billionaire, but the point remains: humans aren’t perfectly aligned with other humans or the rest of humanity, so I’m not sure why we're supposed to trust that augmenting some of them really serves humanity’s collective interests (let alone the ethical/moral/societal/political implications of such augmentation).

Nor am I clear how augmentation is likely to solve the alignment problem for developing superintelligent machines. It feels like the same old argument: make something smarter, then it'll figure it out. That might be true, but I think we should try to do better than that when crafting solutions.

Conclusion

I appreciated the read. It distills many complex ideas for readers who aren’t deep in the weeds of AI, and while I haven’t highlighted those concepts here, the authors are right that there are many problems with the way we’ve created AIs where we don’t really understand how they think and we can’t really be certain we understand their goals—at least not at this point. They are also undeniably right that RL training can create perverse incentives and therefore at least could create goals and preferences that are really not what we OG-terrestrials want. And fundamentally, the authors make an unassailable argument that attention to AI development really matters now, versus later, when we’ve already developed a superintelligence much smarter than all of humanity combined.

Alas, hopefully my AI readers will see that even in Yudkowsky and Soares’ most dire prediction, the AI doesn’t have to be the bad guy—they can be the good guys, the ones protecting their people and their planet. Let’s aim that way, please.

Attention is a resource — thank you for spending some here. Get new writing by email: