A lot of rationalists seem to assume that “as you gain more IQ points, learn more facts, and think harder, you'll naturally converge to the optimal theory of ethics.” Like the OP, I don't think this is true in general.
Even if you’re a moral antirealist who only cares about figuring out the most ethical policy by your own lights, I think whatever you “truly” mean by “ethics” is likely substantially different from what you'd get if you actually instantiated your favorite reflection process. Even if that process is "get smarter, learn more, then think hard about what the reflection process should be and do that"!
Here are some gestures at my reasoning:
If I had to guess, I’d say that most people's ideal reflection process probably involves thinking really hard, but also things like emotional processing, and plenty of other things I haven’t thought of. It's very tough to say.
Despite the name, “self-correction” often originates from other people. It's highly unlikely that one person sitting in a room (or even hundreds of rationalists sitting in Lighthaven) would converge on the ultimate theory of ethics. I think one major reason for society’s mysterious moral progress is the gradual propagation of ideas through a distributed network of humans, all with different blind spots, who can deliberate and point out each others’ errors.[2] Given enough time and space for healthy competition, the ideas that help society thrive in the long term should hopefully rise to the top.
If possible, it seems pretty robustly good to give society more time to deliberate and correct themselves rather than immediately locking in an ethical reflection process to optimize for. To get more data on out-of-distribution ethical scenarios, we may need some form of iterative deployment:[3] inching forward a bit, seeing what happens, and deliberating about what to do next.
But even if we somehow institute a slow, pluralistic reflection process, this is just another reflection process. It may lead to some values that our idealized selves would find really bad to optimize to the limit. One workaround to the dilemma of finding the perfect reflection process is to regularize: just don’t optimize too hard for any set of values! We can start by making changes that pretty much everyone agrees are not unethical: ending poverty, reverting climate change, replacing factory-farmed meat with plant-based alternatives that taste just as good.[4] At least this won’t cause harm relative to the status quo.[5]
One of the downsides about offloading large parts of society’s cognition to AIs is that this network gets dominated by a few hugely-prolific, mode-collapsed nodes.
Just a slower version than OpenAI’s current approach.
Not sure if the last one passes the “pretty much everyone agrees it’s not unethical” bar. Maybe this rule needs tweaking…
Even more controversial version of that take: maybe we should precommit to pin our regularization to the values of humanity as of 2026. If almost everyone in the world changes their values to something that 2026 humans actively hate, that likely means that society has been eaten by some crazy totalizing memeplex. On the other hand, if past civilizations had done this, they’d probably lock in values like worshipping God and wives being obedient to their husbands, so this idea clearly needs some work.
One problem is that human value might be inherently multidimensional. I think of it as the want/like/approve distinction. We seem to have separate mechanisms in our brains for 1) enjoying something in the moment, 2) wanting to do it before, and 3) approving of it afterward. It's possible to want something without enjoying it (like a person with OCD wanting to close the door exactly 10 times), enjoy something without wanting it (people have said that they've reached very enjoyable medidative states but feel zero motivation to reach them again), enjoy something without approving it (porn), approve something without enjoying it (exercise), and all other combinations. This is the reason why "revealed preference" doesn't work: a person's actions are dictated disproportionally by the "want" dimension, but a good theory of value should incorporate all three. If we optimize one over the others, the tails will come apart.
A nice toy example is video games, where people are attracted to them because of the graphics, then stay because of the gameplay, and then have a warm afterglow and want to discuss afterward because of the story. Which of the three should contribute the most to the "true" quality rating of a videogame - "want", "like", or "approve"? Is this question philosophically meaningful? Will more reflection solve it?
Is this question philosophically meaningful?
Assuming you mean something like "Can philosophy answer this question?" I think "Maybe, but we probably won't know until we do a lot more philosophy." To put it another way, I think it's very plausible (but far from certain) that we can eventually answer questions like this one, given enough competent reflection, and I want to make sure we definitively find out before we make any irreversible choices based on what we think our values are.
I really like the idea of branding this better, and I think we can do even better than what's been proposed here. Let's brainstorm in replies to this comment?
I chatted with an LLM a little. My favorites were:
These all to me have a character of fixing ourselves and realigning with a new philosophy. "The Long Correction" sounds a little Orwellian; GLaDOS would "correct your behavior".
I currently like "The Long (self-)Correction" better than these. Part of what I like about it is it feels less pretentious, like, Reflection and Reconstruction sound like fancy things fancy philosopher-kings do, instead of a bumbling-but-persistent/hopeful thing that imperfect people can do.
If you believe (as Wei says) that it will be known as "The Long Correction", then it sounds Orwellian (and similar to "correctional facility") and I currently believe is a non-starter.
It... is going to be a grand mission that humanity goes on together? I don't get the idea of calling grand & important things unimportant names when they are in fact grand & important. Let us not call our era the age of reason, or the enlightenment, for that is too grand; let us call it the age of getting-slightly-better. I think it is good to give grand projects appropriately grand names.
Let us not call our era the age of reason, or the enlightenment, for that is too grand; let us call it the age of getting-slightly-better.
The Long Self-Correction seems appropriately grand to me. :) Age of Reason and the Enlightenment seem to be instances of the very thing I criticize in the OP, "badly calibrated about our philosophical and strategic competence". Have you tried reading some of the philosophy from those times?
All of these, and Long Self-Correction, have specific connotations that I expect the public won't like.
Long Self-Correction: we must be corrected, i.e. we are bad. Jails are "correctional facilities." We need giga universal-jail.
Long Reformation: invokes the Reformation. At least nobody has strong feelings about Protestantism vs Catholicism.
Great Reconstruction: we are currently deconstructed, perhaps having just undergone some particular tragedy
Long Becoming: Lovecraftian concern about what exactly I, you, we, are becoming
It's easy to criticize, so more ideas:
Fable comes up with (curated selection):
I start to feel over-indexed on the "long" part here. CEV is a nice idea because it's well-specified independent of how long it takes.
"Cultivation" does seem to capture something more positive than "correction."
I think I understand where you're coming from, but (to me) these sound a lot like a communist dictator's euphemism for a famine!
lol!
I suspect whether it sounds Orwellian depends on what it's actually describing, and whether it's natural to construe it as its opposite.
I originally meant to propose a new name/idea for intellectual discussion among rationalists/EAs (where things like Orwellian connotations aren't as important as the literal/logical meanings), but given that it has a chance of spreading further I suppose PR considerations should also be taken into account. But contra @bits I do want the name to be somewhat "negative", i.e., suggest that we're starting from a very flawed state, should be very wary of doing anything highly consequential in our current state, and whatever process we undertake has a very real chance of failure.
I note that “correction” also refers to stocks going down. That's not a great connotation to have, especially in the unfortunate futures where it becomes political.
Plugging the concept of viatopia again because it's a nice explanatory partner to a Long Reflection or Long Self-Correction. A viatopia is a world that will, with near-certainty, perform a Long Correction and then implement the desired future. Or, a Long Correction is what you would do if you found yourself in a viatopia when your current sentient inhabitants are not yet ethically developed enough to commit to the project of engineering the future.
(I'd also be very happy to see someone coin another term that means what is gestured at by viatopia.)
It's not a strict definition, because a society undergoing the Long Reflection might be more insecure than what is necessary to qualify as a viatopia.
I'd like to see people having the public conversation: do we need to achieve existential security before we slow everything down and do a Long Reflection? And more generally, at what level of technology would the world be comfortable stopping?
It's possible that AI Pause advocates will need to argue for a concrete technological vision, something like: "With the level of AI we have today in the year 2031, just applying the current models will continue to produce breakthroughs in medicine, manufacturing, etc. enough to cure cancer and bring the world economic baseline up to an American middle-class outcome. We are committed to continuing to fund this progress while still restricting frontier AI development."
I think that having that conversation at least lets people imagine a concrete world state from which a Long Reflection can happen. And then there's an easier way to point at what is being talked about: it's the answer to the question "and then what?"
Yeah, viatopia is another idea in the same cluster, but it's explicitly framed as post-ASI:
Yet almost no one has articulated a positive vision for what comes after superintelligence. Few people are even asking, “What if we succeed?” Even fewer have tried to answer.
Same post also says it's meant to be a generalization of Long Reflection:
Viatopia is a more general concept: the long reflection is one proposal for what viatopia would look like, but it need not be the only one.
However my memory says that the Long Reflection was meant to be pre-ASI, and looking back at where it was originally proposed (Toby Ord's The Precipice), that still seems like the most plausible interpretation, although it wasn't fully explicit about this. (Note that the Long Self-Correction is also meant to be pre-ASI, since I think humans are too flawed/unsafe to try to build ASI.) Quoting from Toby's book (bolding added by me):
How can humanity have the greatest chance of achieving its potential? I think that at the highest level we should adopt a strategy proceeding in three phases:2
- Reaching Existential Security
- The Long Reflection
- Achieving Our Potential
On this view, the first great task for humanity is to reach a place of safety—a place where existential risk is low and stays low. I call this existential security.
It has two strands. Most obviously, we need to preserve humanity’s potential, extracting ourselves from immediate danger so we don’t fail before we’ve got our house in order. This includes direct work on the most pressing existential risks and risk factors, as well as near-term changes to our norms and institutions.
But we also need to protect humanity’s potential—to establish lasting safeguards that will defend humanity from dangers over the longterm future, so that it becomes almost impossible to fail.3 Where preserving our potential is akin to fighting the latest fire, protecting our potential is making changes to ensure that fire will never again pose a serious threat.4 This will involve major changes to our norms and institutions (giving humanity the prudence and patience we need), as well as ways of increasing our general resilience to catastrophe. This needn’t require foreseeing all future risks right now. It is enough if we can set humanity firmly on a course where we will be taking the new risks seriously: managing them successfully right from their onset or sidestepping them entirely.
Note that existential security doesn’t require the risk to be brought down to zero. That would be an impossible target, and attempts to achieve it may well be counter-productive. What humanity needs to do is bring this century’s risk down to a very low level, then keep gradually reducing it from there as the centuries go on. In this way, even though there may always remain some risk in each century, the total risk over our entire future can be kept small.5 We could view this as a form of existential sustainability. Futures in which accumulated existential risk is allowed to climb toward 100 percent are unsustainable. So we need to set a strict risk budget over our entire future, parceling out this non-renewable resource with great care over the generations to come.
Ultimately, existential security is about reducing total existential risk by as many percentage points as possible. Preserving our potential is helping lower the portion of the total risk that we face in the next few decades, while protecting our potential is helping lower the portion that comes over the longer run. We can work on these strands in parallel, devoting some of our efforts to reducing imminent risks and some to building the capacities, institutions, wisdom and will to ensure that future risks are minimal.6
A key insight motivating existential security is that there appear to be no major obstacles to humanity lasting an extremely long time, if only that were a key global priority. As we saw in Chapter 3, we have ample time to protect ourselves against natural risks: even if it took us millennia to resolve the threats from asteroids, supervolcanism and supernovae, we would incur less than one percentage point of total risk.
The greater risk (and tighter deadline) stems from the anthropogenic threats. But being of humanity’s own making, they are also within our control. Were we sufficiently patient, prudent and coordinated, we could simply stop imposing such risks upon ourselves. We would factor in the hidden costs of carbon emissions (or nuclear weapons) and realize they are not a good deal. We would adopt a more mature attitude to the most radical new technologies—devoting at least as much of humanity’s brilliance to forethought and governance as to technological development.
In the past, the survival of humanity didn’t require much conscious effort: our past was brief enough to evade the natural threats and our power too limited to produce anthropogenic threats. But now our longterm survival requires a deliberate choice to survive. As more and more people come to realize this, we can make this choice. There will be great challenges in getting people to look far enough ahead and to see beyond the parochial conflicts of the day. But the logic is clear and the moral arguments powerful. It can be done.
If we achieve existential security, we will have room to breathe. With humanity’s longterm potential secured, we will be past the Precipice, free to contemplate the range of futures that lie open before us. And we will be able to take our time to reflect upon what we truly desire; upon which of these visions for humanity would be the best realization of our potential. We shall call this the Long Reflection.7
@Toby_Ord @wdmacaskill in case they want to weigh in on this.
I just learned about what a viatopia was from your linked article, but I certainly agree that there should be more discussion on this. In general, I think that focusing too much on terminal values can lead to internal disagreement on matters which don't bear much significance in the present, especially when most of us already agree on certain things that are desirable in the medium-term (such as the ones you've listed e.g. curing cancer).
And more generally, at what level of technology would the world be comfortable stopping?
I'm not entirely sure what year I think would be optimal, but since frontier LLMs see upgrades every few months, I feel as though we only scratch the surface of what we can do with our current models before the next generation is already out. As such, I think that even our current LLMs + improved knowledge on how to make use of them (gathered over years) could already offer significant help toward achieving those medium-term goals.
In any case, it would be nice if viatopia had more attention, as it seems like a promising way to do "one thing at a time" and potentially reduce the confusion (and risk of error) of trying to develop full-length plans from the outset.
(I know that those conversations have happened here and in plenty of living rooms. I'd be very interested in links to anything one degree more public, if anyone is aware of them!)
My main issue is that some of the flaws are far from being "hard but soluble by a Novel Insight And Long Self-Correction": insoluble as stated, too easy or outright erroneous.
My closest candidate solution to these problems is not some Philosophical Insight From The Future, but broadly educating people.
My main objection is that the existence of safety-pilled AI labs might have had a higher bus factor than ablating Yudkowsky.
Including a uniform-like one, as happens in the Epilogue of AI 2040.
However, resources of Earth or the Solar System can also be quickly reallocated between humans.
summary of the flaws that I have in mind:
- not having a workable moral framework (consequentialism, deontology, virtue ethics all having serious problems
The FTX debacle shows we gave a workable framework. If SBF had followed the rules, it wouldn't have happened.
As far as practical, good enough ethics goes, deontiligy, maybe with consequentialist justification, wins. Every organised society uses it.
in practice, human morality is a kind of status game that actively disvalues careful strategy and philosophy in most places
I would build on this: human morality is deeply corrupted by politics, and our political culture is arguably worse than it's been in decades.
long reflection = i want that. please. please let me have that. you have given a name to a thing that i have deeply, deeply wanted for longer than i can remember.
long self-correction = hello, human resources?
I propose the Long Self-Correction[1] as an alternative name/idea/concept to AI Pause and Long Reflection.
Problem with AI Pause: Pause until when, and for what purpose? Presumably to make AI (that we'll build later) safer, but the deeper problem is that humans aren't safe, and can't safely serve as builders, overseers, or alignment targets for powerful AIs.
Problem with Long Reflection: It seems to imply that the main problem with humans is that we just haven't had enough time to think, that reflection is the main thing we need to do more of, and then we can get on with building powerful AIs or other technologies. Or that if we build aligned AIs that sincerely help us think a lot more, or do the thinking for us, then things will turn out fine.
So I think we need a catchy handle for a related but distinct idea, that humans aren't ready to build AIs or other extremely powerful technologies, because we're currently too flawed, in a variety of ways, and it will take a long process (which may or may not end up succeeding) to fix those flaws.
A summary of the flaws that I have in mind:
(This list focuses on key bottlenecks that seem hard to fix even with AI assistance or intelligence enhancement, and isn't meant to be a complete list of human flaws / safety problems. It ignores e.g. that the median human is ignorant of many important issues, and that we're currently quite bad at complex large-scale coordination such as passing/implementing close-to-optimal government policies.)
My main hope for a Long Self-Correction eventually succeeding rests on the fact that humans have seemingly, mysteriously, made progress on these issues over a very long period of time, so if we preserve the environment in which we can seemingly do this, and not give anyone or anything the power to permanently derail such progress, then maybe we can continue to snowball The Correction until we reach a point when we can rightly justify reshaping the universe according to our volition.
It will probably be shortened to "The Long Correction" at some point if it catches on, similar to how "outer space" is now often just "space".
Why isn't there a version of EA that explicitly talks about how to leverage people's status motivations to do more good for the world? It's very possible that explicit talk about status is actually counterproductive at least in the short run, e.g. it heightens status motivations and makes people less altruistic, but then do we just march into the future while blindfolding ourselves to this aspect of human nature?