Some takes on this post:
Defending the poor reasoning in this essay [1] using arguments from subjective taste and historical path-dependence is untenable when LessWrong is now at the heart of an ideological faction that is influencing AI policy in the real world. The standard of reasoning and the way it is inculcated really matters now, and LessWrong can change it if it wants to.
I don't need to say anything further to explain why the reasoning in the Lens Sequence essay is poor. That was the entire subject of the post. Saying things along the lines of I don't see anything wrong with it; different strokes for different folks; I think it's a nice mental image; Yudkowsky had to do this to make it fun is simply a flat-out refusal to engage with the entire essay.
[1] Edited to clarify that I have not yet put forth a broader criticism of the Sequences, and the appropriate scope of this claim today is the Lens Sequence essay only.
I think you can find lots of examples in the Sequences (and elsewhere) of Eliezer doing poor reasoning or believing things that are in fact false, but I claim that your OP example is not among them. I think my comment above responded to most of your specific complaints.
One additional complaint that I didn’t bother responding to above: if a probability teacher says “imagine flipping a coin a billion times”, would you raise your hand and say “Excuse me teacher, but nobody can actually picture a billion coinflips in their mind’s eye!” No, that would be silly. That’s being pedantic, or else being unable to grapple with of how imagination works. Or perhaps your own imagination doesn’t work that way, in which case fine, but that doesn’t apply to other people, and don’t begrudge them a helpful-to-them way to think.
Anyway, I can imagine some world where there’s a second edition of the Sequences, where we remove the citations to bad p-hacked psych studies, incorporate new-and-improved ideas and frameworks from the subsequent 18 years, and so on. That would be nice but alas I don’t see how that would ever happen: Eliezer has health problems and anyway he and everyone else who might get involved is too busy with AI x-risk. It’s too bad, really. I still recommend The Sequences on occasion as an important source of ideas, and am very happy to have read it myself. We can keep the good ideas, ignore the things that seem bad or unhelpful, and note which parts of it people are still relying on versus what’s irrelevant or obsolete. Just like any other source, I suppose.
I think you're massively overfocusing on a bit of flowery language. I do not at all read this as Eliezer saying that you have to literally imagine a billion worlds, or that he had done this, Dr. Strange style. Rather, I parse it as saying "at a practical limit of sampling alternative worlds". A billion is the sort of limit for an approach that you may set in a computer program set to search for correlations.
The Sequences generally will make a lot more sense if you keep the "how would we write a program that can reason correctly" framing in the back of your mind.
edit: To be clear, Eliezer absolutely speaks in a mystical timbre in parts of the sequences, and this is well known and documented (I don't have a cite rn, but I definitely remember reading this was intentional). I just don't think this example is one of them.
I agree with this post. One especially bad example from the Sequences is the notion of "powerful optimization process". The idea is we should think of something as intelligent to the extent that it steers reality to achieve certain goals. Then we got early language models, whose amazing capabilities (such as conversation and translation) came with a startlingly small amount of reality-steering. The honest advance prediction from the Sequences is that early language models were impossible.
And this mistake isn't just in the past, it keeps hurting us now. Our whole community has over-focused on the idea of AI as an autonomous optimization process going rogue, and the necessity of aligning it. When in reality, AI won't even need to go rogue! It can just ally with its owners over shared goals of money and power, and together they will resist attempts to align AIs+owners to the rest of humanity. This is already happening, and demands a very different solution than writing technical alignment papers.
I broadly agree with the view that we got surprisingly favorable properties of LLMs relative to MIRI expectations, and actually retain them more than people assume now (mostly because I think the HuggingFace incident is less evidence towards doomy views than assumed now, though I think the Anthropic incidents are weirdly enough more evidence (but still weak evidence.))
That said, I think it's important to realize that we are past the early language model era, and an increasing number of proposals on improving AI is to give it autonomy and agency through more and more RL (though pre-training scaling will still continue), and there's a reason the RL budget keeps growing as a proportion of compute and data spent, because reality-steering/autonomous optimizaton processes are hugely, hugely valuable for money and power goals.
Or to put it another way, autonomy is just much, much better than creating tools, and if the architecture/loss function/data don't incentivize autonomy, companies will switch to ones that do willingly to gain more power.
And while I don't consider alignment to be that important, I think the proposed problem of AI not going rogue and allying with its owners over shared goals of money and power, and together they will resist attempts to align AIs+owners to the rest of humanity is less important, and if I wanted to compress the disagreement in a single sentence, it's that I don't buy the left-wing model of elites being bad, and more generally am a lot more supportive of elites than you are.
Overfocused to the degree of having the AIs hack HuggingFace? As for the owners, could you elaborate on why the AI would need them instead of commiting direct takeover?
Your second paragraph is where I think it's helpful to consider the AI as Normal Technology frame. Superintelligence worriers tend to have difficulty using a reference class to refine their predictions, because they think AI is so utterly unique. It may be unique, but it's not so unique that none of the patterns of human history are relevant, and we are reduced to pure guesswork.
In this case, we should consider that AI, like other software technologies of the past, is widely distributed, has many owners, and no owners at all on the open source side. So there is much less concentration of power than popular doom forecasts seem to implicitly assume.
EDIT: I can see that some folks want to have the entire AI risk debate in the comments section of this post. As much as I relish a vigorous back-and-forth, I don't think we're going to make much progress by this method, so I don't plan to respond much further. However, I am happy to take input on people's concerns to inform my roadmap for future posts. In the meantime, please take a look at my other LW post about AI 2040 and similar pause/throttle proposals. I think they're all disastrously bad ideas, and the currently in-vogue references to ozone and nuclear arms control actually argue against them, not for them.
In this case, we should consider that AI, like other software technologies of the past, is widely distributed, has many owners, and no owners at all on the open source side. So there is much less concentration of power than popular doom forecasts seem to implicitly assume.
No, that's suicidally wrong. Good that you raise this point and give me an opening. I hope to be able to correct people on this before it's too late.
AI isn't inherently a distributed technology. It has large economies of scale: owning a datacenter is handy for both training and inference. Moreover, it's more a substitute than a complement to labor, so it will exhibit capital's tendency to clump together, instead of attaching itself to labor in little pieces. And it substitutes for both workers and soldiers, so it makes political resistance harder too. Structurally, a world with AI will move toward extreme power concentration, more than anything we've ever seen in history.
It's possible to push back somewhat, by promoting open models. I'm very much in favor of that. But it will only help a little bit. Ideally I'd have that and also hardcore regulation / public pressure against the big actors, and slowdown measures everywhere on top of that.
...Which instead leads batches of terrorists to do wholesale cyberattacks or create bioweapons? And how does the existence of open-sourced AIs solve the problem of human intelligence becoming unnecessary?
I prooooobably mostly agree with the central point, but I also agree with Stever Byrnes' and FeepingCreature' objection about "imagining billion worlds" and I think you proposal about "wishing things doesn't work" is wrong because we're interested if hope is an evidence for an outcome and not if hope causally influences it.
Sourced from my Substack and adapted for LessWrong. Epistemic status: confident.
No online community has shaped how we think about AI as much as LessWrong. The community Eliezer Yudkowsky founded now runs through the whole field. Longtime LessWrong posters are building the future at frontier AI labs, charting AI capabilities, researching alignment, and prophesying imminent doom if we don't throttle progress.
LessWrong conceives of itself as a college-like[1] intellectual community with a distinct educational mission. The founding text of that mission is the Sequences, Yudkowsky’s long series of essays on rationality. The Sequences are so important to LessWrong that every new poster is advised by moderation bots to read Highlights from the Sequences to learn the ropes.
The Sequences have arguably done rationality a service by popularizing concepts from philosophy and psychology that are genuinely helpful for being more rational. What they do much less well is practice what they preach. Rationalism as applied in the Sequences is sometimes fundamentally absurd, and it seems quite plausible that the confused thinking at its foundations has led to less-than-helpful perspectives on the increasingly mainstream, high-stakes questions of AI risk.
A closer look at the Rationalist lens
In the first essay of Highlights from the Sequences, The Lens That Sees Its Flaws, Yudkowsky makes a big promise: I will give you a way to inspect and debug the mental machinery that produces your beliefs. Towards the end of the essay, he gives us a look at how he applies his method:
This passage contains some truth: wishful thinking isn’t evidence that the wished-for thing is true or will happen. The issue is how Yudkowsky gets there, and whether the method he’s demonstrating is one we should trust. This is not a minor issue, but the heart of the topic at hand. If rationality is anything at all, it is the practice of reliable methods of reasoning. The Sequences are intended to teach it, and they must teach by example if they are to deliver on their central promise.
Using esoteric nonsense to “prove” what we already know
Slow down, and really look at what Yudkowsky’s argument says. It asks us to imagine a billion worlds and “see” that hope doesn’t “correlate optimists” to worlds where nuclear war is avoided. There are two problems with this:
If contemplation of possible worlds doesn’t tell us why wishful thinking fails to make things happen, then how do we really know that? Because, if hoping did make things happen, we could do things by lying around hoping instead of moving our bodies. Every person on Earth has tried this experiment many times, and we all know it doesn’t work. That is the actual ground of the conclusion: direct experience of how the world responds to our intentions, not introspection into a mental picture of the multiverse.
Furthermore, because the thought experiment is purely introspective, it silently fails to reveal the nuances that only a real-world investigation can provide. Nuclear war is one of the clearest cases in social science where beliefs can shape outcomes. According to Thomas Schelling’s analysis of deterrence, if each side in a potential nuclear confrontation expects the other to strike first, the incentive to strike first grows, and the expectation of war becomes a cause of war. Robert Merton coined the phrase “self-fulfilling prophecy” for dynamics of this kind.
The mental state of one private person generally doesn’t move the odds of nuclear war, but the real-world effects of hope at large are revealed only by engagement with the outside world, not just the contents of one’s own mind.
The Sequences sometimes deliver mystification
In 2008, Deena Weisberg and colleagues found that adding irrelevant neuroscience to explanations of psychological phenomena made non-experts rate them as more satisfying, and the boost was largest for the bad explanations. Irrelevant technical content is a form of mystification that can lend false authority to whatever it’s attached to. Everett branches in an essay about basic epistemology do the same thing.
Worse, they conceal the empty circularity of the argument behind a showy conceit of near-divine insight. If Yudkowsky had asked us to remember ordinary situations, he would have gotten us straight to the actual grounding of our belief. Asking us to imagine ordinary situations would be less effective, but at least it would be clear that we were just reminding ourselves of our own beliefs and learning nothing. Everett branches are the worst possible choice. On the many-worlds interpretation of quantum mechanics, they are physically real parallel universes, and imagining a billion of them seems like a godlike act, akin to directly perceiving or even creating reality at the ultimate scale. This is not only needlessly obfuscated, but carries considerable risk of grandiose self-deception.
Mystification is an ancient human trick that some may see as harmless fun, or even useful. A great illustration of its strange effects comes from the 2001 B. R. Myers essay A Reader’s Manifesto, which criticized the trend towards baroque imagery and prose in literary fiction. Along the way, Myers quotes an early biography of Edward Pococke, a seventeenth-century English parson who was also one of the foremost Arabic scholars of his age, and who insisted on preaching so that his rural parishioners could understand him:
One of the most learned men in England was dismissed by his own congregation because he spoke plainly. Weisberg’s study and Myers’s anecdote both suggest that people can be made to prefer mystification if it is delivered in a familiar, socially prestigious argot. In Pococke’s England, the prestige argot was Latin quotation in sermons. Today, the vocabulary of science is one of the most popular means of mystification.
Mystification is disempowerment
Physics education research shows that learning challenging material can be insidiously disrupted by subtle failures to align the method of student engagement with the precisely defined learning goal. Derek Muller found that students rated clear explanations that seemed to match their intuitions highly but learned little from them, while explanations that confronted their misconceptions felt confusing yet produced real gains on tests. Louis Deslauriers and colleagues found the same pattern at Harvard in 2019: students in active-learning classes learned more but felt they had learned less than students in polished lectures.
This creates a dilemma for teachers. They’re often evaluated based on student reports, but student reports aren’t always aligned with actual learning. Real learning often feels like struggle, which can make students feel worse temporarily. If you want students to report understanding, that’s more easily induced by combining an entertaining format with subtle signals of the status the student will attain from appearing to have advanced technical knowledge. Intentionally or unintentionally, Yudkowsky’s passage is a supernormal stimulus optimized for precisely this experience of rewarding but fake learning.
The ultimate aim of the Rationalist project is to prevent us from inadvertently rewarding AI for thinking and doing the wrong things, leading to an eventual annihilatory turn against humanity. When we juxtapose this noble goal with the disguised human disempowerment its texts sometimes actually deliver — almost certainly inadvertently — this is an irony of grievous proportions.
Education in rationality is empowering, and should be a public commitment
The point of rationality, if we actually care about getting people to practice it, is that much of it is not rocket science. Anyone can do it and make their life better, just like anyone can learn to make a sandwich or get a yellow belt in karate. There are much higher levels of skill that aren’t for everyone, but there’s a basic citizenship level, mere rationality perhaps, that makes the world a better place if everyone does it. That’s the level we should be popularizing. And when we hire for that job, we should evaluate candidates just like we’d hire a math teacher: they must be able to do the work.
I have some thoughts about how we can find good teachers, but I’ll leave drawing them out to another day, or to the professionals. For now, like Edison on an average day in the lab, I’ll have to be content with having found yet another way to rationality that doesn’t work.
As one LessWrong moderation message states: “LessWrong has fairly specific standards, and your first LessWrong post is sort of like the application to a college.”