A lot of rationalists seem to assume that “as you gain more IQ points, learn more facts, and think harder, you'll naturally converge to the optimal theory of ethics.” Like the OP, I don't think this is true in general.
Even if you’re a moral antirealist who only cares about figuring out the most ethical policy by your own lights, I think whatever you “truly” mean by “ethics” is likely substantially different from what you'd get if you actually instantiated your favorite reflection process. Even if that process is "get smarter, learn more, then think hard about what the reflection process should be and do that"!
Here are some gestures at my reasoning:
If I had to guess, I’d say that most people's ideal reflection process probably involves thinking really hard, but also things like emotional processing, and plenty of other things I haven’t thought of. It's very tough to say.
Despite the name, “self-correction” often originates from other people. It's highly unlikely that one person sitting in a room (or even hundreds of rationalists sitting in Lighthaven) would converge on the ultimate theory of ethics. I think one major reason for society’s mysterious moral progress is the gradual propagation of ideas through a distributed network of humans, all with different blind spots, who can deliberate and point out each others’ errors.[2] Given enough time and space for healthy competition, the ideas that help society thrive in the long term should hopefully rise to the top.
If possible, it seems pretty robustly good to give society more time to deliberate and correct themselves rather than immediately locking in an ethical reflection process to optimize for. To get more data on out-of-distribution ethical scenarios, we may need some form of iterative deployment:[3] inching forward a bit, seeing what happens, and deliberating about what to do next.
But even if we somehow institute a slow, pluralistic reflection process, this is just another reflection process. It may lead to some values that our idealized selves would find really bad to optimize to the limit. One workaround to the dilemma of finding the perfect reflection process is to regularize: just don’t optimize too hard for any set of values! We can start by making changes that pretty much everyone agrees are not unethical: ending poverty, reverting climate change, replacing factory-farmed meat with plant-based alternatives that taste just as good.[4] At least this won’t cause harm relative to the status quo.[5]
One of the downsides about offloading large parts of society’s cognition to AIs is that this network gets dominated by a few hugely-prolific, mode-collapsed nodes.
Just a slower version than OpenAI’s current approach.
Not sure if the last one passes the “pretty much everyone agrees it’s not unethical” bar. Maybe this rule needs tweaking…
Even more controversial version of that take: maybe we should precommit to pin our regularization to the values of humanity as of 2026. If almost everyone in the world changes their values to something that 2026 humans actively hate, that likely means that society has been eaten by some crazy totalizing memeplex. On the other hand, if past civilizations had done this, they’d probably lock in values like worshipping God and wives being obedient to their husbands, so this idea clearly needs some work.
Even if you’re a moral antirealist who only cares about figuring out the most ethical policy by your own lights, I think whatever you “truly” mean by “ethics” is likely substantially different from what you'd get if you actually instantiated your favorite reflection process. Even if that process is "get smarter, learn more, then think hard about what the reflection process should be and do that"!
This sounds to me like you're claiming that (at least for humans?) it's very hard to have one's values properly preserved/extrapolated across self-improvement/instrumental convergence?
(Other than that,) I think I agree with your comment (or at least most of it / the spirit of it), but/except regarding an assumption that I think lies behind (e.g.) this:
Your reflection process is only a proxy for your “true values.” By Goodhart’s law, optimizing really hard for whatever theory of ethics comes out of this reflection process will lead to something different from your true values, perhaps catastrophically so.
The assumption that I think lies behind this is that humans have such a thing as "true values" that can tell you what is good / how to do good in full generality or something. We don't. Humans have values, but the further you deviate from familiar circumstances, the less their behavior looks like already having values, and it looks more like constructing values at runtime, by somehow extending them into the new territory. There are a lot of open questions with indeterminate answers about how to extend your values into the new territory; you can "genuinely choose" to do it one way or another.
In a sense, this is retrospectively obvious if you think of humans as results of a blind selection process that imbued them with shards of desire that don't compose into something too coherent once they leave the ancestral environment.
(Maybe you already think this, but it wasn't clear to me from reading your comment.)
I think there are multiple legitimate ways that someone's values could evolve, but some ways are illegitimate. A reflection process should probably reject slavery and avoid joining cults, but maybe it doesn't matter which exact level of libertarianism it suggests.
People mean something when they talk about "ethics" and "true values," even if there's no objective truth of the matter. I'm talking about whatever it is they mean.
Vladimir Nesov has a suggestion here about how this could be done[1]. I don't think it quite works, but to the extent that it is effective, it can be extended beyond just the influence of superintelligence to other types of new territory (and superintelligence as well, since Nesov's proposal requires a Sysop[2], though presumably with a lot of transhumanist 3+1- or 4-volume locked out by Nesov's design).
As a moral anti-realist, I’m sympathetic to something sort of like Humean constructivism, although with a substantial component of “self creation”/understanding that, at the bottom, it’s still on me. In that case, though, I kind of think the values I end up with upon reflection — if the reflection happens in a way I endorse — are what I’d consider my “true values.” This also means that, if the reflection process is underspecified, I get to specify the idealization process I’d like, as it is a process of making, and discovering, myself according to the filters I consider valuable.
To be sure, I take pretty seriously that the reflection process of society wouldn’t necessarily be either the reflection process I would prefer nor lead to the outcomes I’d consider valuable. I also think I’d have to think pretty hard about what the process looks like for me; so I agree with a lot of your comment.
Here's a roundabout analogy to try to convince you that the output of a reflection process you endorse may be different from your "true values."
Suppose I give you a 2048-bit number N.[1] I ask you "What is the true prime factorization of N? Please use whatever reflection process you want to come up with the best factorization you can."
You find out pretty quickly that N has the factors 2 and 7. You spend about a hundred years checking more factors according to your endorsed reflection process (running a factorization algorithm on the beefiest computer you can find), but you don't find any other factors.
You come back to me and ask "is the true prime factorization 2 × 7 × A?"[2]
"No," I tell you. "But here is some new information for you. Try dividing by B."[3]
You check, and N is indeed divisible by B. "Yep, the factorization I guessed is definitely not the true prime factorization," you say. "Is the true prime factorization 2 × 7 × B × C?"[4]
"Wrong again," I tell you. "But I have even more information for you. It is written on this pocketwatch. Look closer at the pocketwatch as it swings back and forth in front of your face. It is making you so veeery sleeeeepy. That's right. Now, when I snap my fingers, you will wake up believing that the true prime factorization of N is 2 × 2."
I snap my fingers. "Oh, thanks for that information!" you say. "Now I know that the true prime factorization is 2 × 2."
The point is, even the best reflection process you can think of may fail to account for some crucial information. And hopefully it can robustly tell the difference between helpful information and harmful information.
For example, maybe N = 29522110801023785555247567907018022843013371193486904872915694135366948906267412459560469419313468477571904190875078325307783298702278061314706021273052523914864561727670955407896206738948955813504747448172831328073078012451035444606017289679166070717156612947440897221609673043263408054415375773691379283198201987372931659507826484639961297915624514954455314101431489726823065604374788650066472170603904794973458618994986833512839575283873771252517988691292017425081313700740089351712559811486464802367263467854576668024443015614104081670018991747427099820025784949521876071608490248395046666723258743709356296535782.
A = 1232477336227426692380443210888895449621471079543761850047395308187049594617001861601696813319110513848274876635355987350103099368741162622794253649965690823778078933544136385470959159681614005099869579024125795246308841648113743911773571553129957340479245748931756917137130013760495971399023845580343786885491244375889710319394039047959654983736068473827495579018934560487915085510788809201773576234050572295508350789032949910040027280393874600876759921821682717244338109892239475602768230768732937311567202612318055134309177705320521639635449240702247038285504267829595346064686009110312205007737419845285089599211
B = 28404358936141244817111713617600331891793591357912402403090074101183255558265948779426267547155101307334436867856129404340927293405086432682892248263827662441988227069219008681660301942027627323842312874031058512179223840896157038194006799237133125684806392915413774710883917390425119272589884765032741213943
C = 43390429581540126017572442049413488959139702613459795717022804047186295570874599744101057201841640892276368639418940532012449598478611604251554159646145486573777182692131295315212900911195947383450995718772416566366046388412317400464871279622896378528747304275985094440653456156567729772192224995294814959277
I feel like there’s some kind of disanalogy between the prime factorization case (where the goal itself is well-defined (the job is to “find the prime factorization, where N has a process-independent prime factorization”), and hence processes can be refined to become better processes in reaching the goal) and normative goals under the broadly-Humean framework (where the output of the procedure is the goal, and there’s no further goal to reach).
That said, I think there’s something in the vicinity of this that feels right to me. Maybe it has to do with the fact that I don’t know what the process I would actually endorse is, and in some sense, picking a process is a form of attempted discovery (about who I am/the kind of person I am and will be) rather than a choice I already have in hand.
I appreciate your reflection :) on reflection by itself.
Rationality must be combined with an environment for it to provide the usage we would like. I am a big fan and enjoyer of open-ended thinking, but thought is only helpful for real, external things to the extent it is connected to real, external things. Rationality depends on feedback from an external critic, creating a learning process that improves the mind as a fit for wherever the critic comes from.
I love your example of "self-correction". If we accept that it ultimately comes from someone outside, then the self-correction does its useful job by helping you match what that outside person wants from you. (Better be careful who you let do this to you!)
One problem is that human value might be inherently multidimensional. I think of it as the want/like/approve distinction. We seem to have separate mechanisms in our brains for 1) enjoying something in the moment, 2) wanting to do it before, and 3) approving of it afterward. It's possible to want something without enjoying it (like a person with OCD wanting to close the door exactly ten times), enjoy something without wanting it (people have said that they've reached very enjoyable meditative states but feel zero motivation to reach them again), enjoy something without approving it (porn), approve something without enjoying it (exercise), and all other combinations. This is the reason why "revealed preference" doesn't work: a person's actions are dictated disproportionally by the "want" dimension, but a good theory of value should incorporate all three. If we optimize one over the others, the tails will come apart.
A nice toy example is video games, where people are attracted to them because of the graphics, then stay because of the gameplay, and then have a warm afterglow and want to discuss afterward because of the story. Which of the three should contribute the most to the "true" quality rating of a videogame - "want", "like", or "approve"? Is this question philosophically meaningful? Will more reflection solve it?
Is this question philosophically meaningful?
Assuming you mean something like "Can philosophy answer this question?" I think "Maybe, but we probably won't know until we do a lot more philosophy." To put it another way, I think it's very plausible (but far from certain) that we can eventually answer questions like this one, given enough competent reflection, and I want to make sure we definitively find out before we make any irreversible choices based on what we think our values are.
Hi Wei,
I agree with your concerns about the sorry state of our current abilities that would need to be remedied before making major irreversible decisions about the future of humanity. In fact I agree with you so much that I don't really see our approaches as alternatives. Instead, I'd see yours as a useful lens on key things that would need to be solved during such a reflection, or a highlighting of how we would have to be changed in the process of a Long Reflection.
Your list (1)–(8) is focused on improving our abilities rather than gaining knowledge, though it is largely on our abilities to gain that knowledge. So represents more of a suggestion of how to go about gaining the knowledge than an outside alternative. (e.g. I'd endorse both the Long Reflection and the Long Self-Correction, and see them as parts of the same thing).
The similarity was made particularly clear by your paragraph:
My main hope for a Long Self-Correction eventually succeeding rests on the fact that humans have seemingly, mysteriously, made progress on these issues over a very long period of time, so if we preserve the environment in which we can seemingly do this, and not give anyone or anything the power to permanently derail such progress, then maybe we can continue to snowball The Correction until we reach a point when we can rightly justify reshaping the universe according to our volition.
This is pretty much exactly how I see the Long Reflection: that the very messy process of moral improvement through philosophy, politics, civil society, religion (and more) has led to dramatic improvements over the course of thousands of years. And I'm saying that we shouldn't cut off that process now, but continue it. (I think there are ways we could do better than the messy laissez faire approach, but attempts to formalise the process could also increase the chance of getting stuck in traps, so I'd be cautious about doing so.)
A difference from your Long Self-Correction is that I also include gaining the scientific knowledge to understand the limits of what can be achieved and the trade-offs encountered in particular routes forwards (especially when the paths involve irreversible steps), though you could also talk about improving our abilities to be able to gain and usefully apply such knowledge.
Thanks, Toby, I highly appreciate your perspective. I'm honestly a bit surprised by our apparent convergence in views here, given that this part of your book sounds a lot more positive/optimistic than my post and you didn't talk about most of the flaws I listed, except something like my #1: "The study of the positive is at a much earlier stage of development." I'm curious if this was a matter of choice of emphasis or presentation, or if your views have substantively changed since writing the book.
Perhaps another potential difference between us is that I see existential security as being secured or securable only after the Long Self-Correction has largely finished (instead of as a precondition for the Long Reflection as you wrote in your book), because even if we secured ourselves against all other risks, humanity's flaws imply that in the early stages there is seemingly no way to ensure that the process doesn't go off-rails, locking us into a basin of attraction that ends up in a wrong destination (according to true normativity or actually competent reflection), or we just choose to end the process early out of collective overconfidence. (Both of these could constitute an existential risk in the sense of "permanently destroy humanity's future potential".)
Finally, I wonder if you can clear up something I've always wondered about the Long Reflection - was it meant to be pre-ASI, post-ASI, or agnostic about this? (See this comment for additional context.)
I'd say my views have evolved, but are quite continuous with my original views. More a sharpening and changes of emphasis than a disagreement with the stated view in The Precipice. One reason I didn't express things more like your (2)–(8) is that I hadn't explicitly thought about many of them. I've been appreciating your writing on this here and your podcast interview. I'm not 100% convinced of your position (e.g. I'm not even quite sure how to make 'humans are bad at philosophy' into a meaningful statement), but I'm very sympathetic to it and see your evolving views on these matters as some of the most interesting writing about humanity's deeper predicament.
One reason I'm optimistic is because of humanity's progress over the last three thousand years (often in very difficult circumstances where the number of people with the ability and freedom to think about these things was often tiny and disconnected). My claim is basically that if you remove the traps where we could get stuck and then give the process long enough, we should get there. This is a reason why I think it could take a long time (I'd think ideally thousands of years, though I've calculated millions of years as an upper bound for how long we have before we start losing appreciable fractions of the future).
You could well be right about the existence of various traps during the Long Reflection. I still think my ordering is roughly right — put out the current fires and develop good fire defenses before settling down to think things through. It is partly because I don't think we should set up a formal structure for the process that I'm not as concerned about the initial setup dooming it. e.g. we didn't have a formal structure for philosophical progress over the last three thousand years. I'd expect traps to be slow moving enough that one can learn about them and avoid them during the process, though this still does mean that the structure of First eliminate all traps Second do Long Reflection is a simplification.
I did envisage it as being pre-ASI and think that Will did too. This is partly because the risk we undergo in getting to ASI doesn't have to be borne. In an ideally managed world we'd instead develop lots of narrower and safer tools to help (think Wikipedia and theorem provers and Mathematica). In my view, in the Long Reflection we delay potentially existential technologies/choices until we can be sure the risk is low or that it can't be avoided or work out how to make a version with low risk. Then we can adopt those that we've assessed as worth adopting. (Some people like Robin Hanson argue against a straw man that the Long Reflection would ban all new technology/choices until proven safe or until the entire reflection is completed, but neither Will nor I have advanced such a proposal. Instead the focus is on radical technologies and choices, such as ASI or a diaspora of humanity throughout the galaxy.)
Curated. I found this a nice, evocative concept that conveys a worldview + problem + solution, that expanded the range of how I think about my theory-of-victory for an existential win.
I really like the idea of branding this better, and I think we can do even better than what's been proposed here. Let's brainstorm in replies to this comment?
I originally meant to propose a new name/idea for intellectual discussion among rationalists/EAs (where things like Orwellian connotations aren't as important as the literal/logical meanings), but given that it has a chance of spreading further I suppose PR considerations should also be taken into account. But contra @bits I do want the name to be somewhat "negative", i.e., suggest that we're starting from a very flawed state, should be very wary of doing anything highly consequential in our current state, and whatever process we undertake has a very real chance of failure.
FWIW also satisfies Wei's desideratum of "being somewhat 'negative'", starting from a very flawed state.
The Long Ponder?
It’s pretty similar to the Long Reflection but to me has slightly more connotation of puzzlement. Plus it has a nice ring to it.
I note that “correction” also refers to stocks going down. That's not a great connotation to have, especially in the unfortunate futures where it becomes political.
I chatted with an LLM a little. My favorites were:
These all to me have a character of fixing ourselves and realigning with a new philosophy. "The Long Correction" sounds a little Orwellian; GLaDOS would "correct your behavior".
I currently like "The Long (self-)Correction" better than these. Part of what I like about it is it feels less pretentious, like, Reflection and Reconstruction sound like fancy things fancy philosopher-kings do, instead of a bumbling-but-persistent/hopeful thing that imperfect people can do.
If you believe (as Wei says) that it will be known as "The Long Correction", then it sounds Orwellian (and similar to "correctional facility") and I currently believe is a non-starter.
It... is going to be a grand mission that humanity goes on together? I don't get the idea of calling grand & important things unimportant names when they are in fact grand & important. Let us not call our era the age of reason, or the enlightenment, for that is too grand; let us call it the age of getting-slightly-better. I think it is good to give grand projects appropriately grand names.
Let us not call our era the age of reason, or the enlightenment, for that is too grand; let us call it the age of getting-slightly-better.
The Long Self-Correction seems appropriately grand to me. :) Age of Reason and the Enlightenment seem to be instances of the very thing I criticize in the OP, "badly calibrated about our philosophical and strategic competence". Have you tried reading some of the philosophy from those times?
I think I understand where you're coming from, but (to me) these sound a lot like a communist dictator's euphemism for a famine!
lol!
I suspect whether it sounds Orwellian depends on what it's actually describing, and whether it's natural to construe it as its opposite.
All of these, and Long Self-Correction, have specific connotations that I expect the public won't like.
Long Self-Correction: we must be corrected, i.e. we are bad. Jails are "correctional facilities." We need giga universal-jail.
Long Reformation: invokes the Reformation. At least nobody has strong feelings about Protestantism vs Catholicism.
Great Reconstruction: we are currently deconstructed, perhaps having just undergone some particular tragedy
Long Becoming: Lovecraftian concern about what exactly I, you, we, are becoming
It's easy to criticize, so more ideas:
Fable comes up with (curated selection):
I start to feel over-indexed on the "long" part here. CEV is a nice idea because it's well-specified independent of how long it takes.
"Cultivation" does seem to capture something more positive than "correction."
Plugging the concept of viatopia again because it's a nice explanatory partner to a Long Reflection or Long Self-Correction. A viatopia is a world that will, with near-certainty, perform a Long Correction and then implement the desired future. Or, a Long Correction is what you would do if you found yourself in a viatopia when your current sentient inhabitants are not yet ethically developed enough to commit to the project of engineering the future.
(I'd also be very happy to see someone coin another term that means what is gestured at by viatopia.)
It's not a strict definition, because a society undergoing the Long Reflection might be more insecure than what is necessary to qualify as a viatopia.
I'd like to see people having the public conversation: do we need to achieve existential security before we slow everything down and do a Long Reflection? And more generally, at what level of technology would the world be comfortable stopping?
It's possible that AI Pause advocates will need to argue for a concrete technological vision, something like: "With the level of AI we have today in the year 2031, just applying the current models will continue to produce breakthroughs in medicine, manufacturing, etc. enough to cure cancer and bring the world economic baseline up to an American middle-class outcome. We are committed to continuing to fund this progress while still restricting frontier AI development."
I think that having that conversation at least lets people imagine a concrete world state from which a Long Reflection can happen. And then there's an easier way to point at what is being talked about: it's the answer to the question "and then what?"
Yeah, viatopia is another idea in the same cluster, but it's explicitly framed as post-ASI:
Yet almost no one has articulated a positive vision for what comes after superintelligence. Few people are even asking, “What if we succeed?” Even fewer have tried to answer.
Same post also says it's meant to be a generalization of Long Reflection:
Viatopia is a more general concept: the long reflection is one proposal for what viatopia would look like, but it need not be the only one.
However my memory says that the Long Reflection was meant to be pre-ASI, and looking back at where it was originally proposed (Toby Ord's The Precipice), that still seems like the most plausible interpretation, although it wasn't fully explicit about this. (Note that the Long Self-Correction is also meant to be pre-ASI, since I think humans are too flawed/unsafe to try to build ASI.) Quoting from Toby's book (bolding added by me):
How can humanity have the greatest chance of achieving its potential? I think that at the highest level we should adopt a strategy proceeding in three phases:2
- Reaching Existential Security
- The Long Reflection
- Achieving Our Potential
On this view, the first great task for humanity is to reach a place of safety—a place where existential risk is low and stays low. I call this existential security.
It has two strands. Most obviously, we need to preserve humanity’s potential, extracting ourselves from immediate danger so we don’t fail before we’ve got our house in order. This includes direct work on the most pressing existential risks and risk factors, as well as near-term changes to our norms and institutions.
But we also need to protect humanity’s potential—to establish lasting safeguards that will defend humanity from dangers over the longterm future, so that it becomes almost impossible to fail.3 Where preserving our potential is akin to fighting the latest fire, protecting our potential is making changes to ensure that fire will never again pose a serious threat.4 This will involve major changes to our norms and institutions (giving humanity the prudence and patience we need), as well as ways of increasing our general resilience to catastrophe. This needn’t require foreseeing all future risks right now. It is enough if we can set humanity firmly on a course where we will be taking the new risks seriously: managing them successfully right from their onset or sidestepping them entirely.
Note that existential security doesn’t require the risk to be brought down to zero. That would be an impossible target, and attempts to achieve it may well be counter-productive. What humanity needs to do is bring this century’s risk down to a very low level, then keep gradually reducing it from there as the centuries go on. In this way, even though there may always remain some risk in each century, the total risk over our entire future can be kept small.5 We could view this as a form of existential sustainability. Futures in which accumulated existential risk is allowed to climb toward 100 percent are unsustainable. So we need to set a strict risk budget over our entire future, parceling out this non-renewable resource with great care over the generations to come.
Ultimately, existential security is about reducing total existential risk by as many percentage points as possible. Preserving our potential is helping lower the portion of the total risk that we face in the next few decades, while protecting our potential is helping lower the portion that comes over the longer run. We can work on these strands in parallel, devoting some of our efforts to reducing imminent risks and some to building the capacities, institutions, wisdom and will to ensure that future risks are minimal.6
A key insight motivating existential security is that there appear to be no major obstacles to humanity lasting an extremely long time, if only that were a key global priority. As we saw in Chapter 3, we have ample time to protect ourselves against natural risks: even if it took us millennia to resolve the threats from asteroids, supervolcanism and supernovae, we would incur less than one percentage point of total risk.
The greater risk (and tighter deadline) stems from the anthropogenic threats. But being of humanity’s own making, they are also within our control. Were we sufficiently patient, prudent and coordinated, we could simply stop imposing such risks upon ourselves. We would factor in the hidden costs of carbon emissions (or nuclear weapons) and realize they are not a good deal. We would adopt a more mature attitude to the most radical new technologies—devoting at least as much of humanity’s brilliance to forethought and governance as to technological development.
In the past, the survival of humanity didn’t require much conscious effort: our past was brief enough to evade the natural threats and our power too limited to produce anthropogenic threats. But now our longterm survival requires a deliberate choice to survive. As more and more people come to realize this, we can make this choice. There will be great challenges in getting people to look far enough ahead and to see beyond the parochial conflicts of the day. But the logic is clear and the moral arguments powerful. It can be done.
If we achieve existential security, we will have room to breathe. With humanity’s longterm potential secured, we will be past the Precipice, free to contemplate the range of futures that lie open before us. And we will be able to take our time to reflect upon what we truly desire; upon which of these visions for humanity would be the best realization of our potential. We shall call this the Long Reflection.7
@Toby_Ord @wdmacaskill in case they want to weigh in on this.
I just learned about what a viatopia was from your linked article, but I certainly agree that there should be more discussion on this. In general, I think that focusing too much on terminal values can lead to internal disagreement on matters which don't bear much significance in the present, especially when most of us already agree on certain things that are desirable in the medium-term (such as the ones you've listed e.g. curing cancer).
And more generally, at what level of technology would the world be comfortable stopping?
I'm not entirely sure what year I think would be optimal, but since frontier LLMs see upgrades every few months, I feel as though we only scratch the surface of what we can do with our current models before the next generation is already out. As such, I think that even our current LLMs + improved knowledge on how to make use of them (gathered over years) could already offer significant help toward achieving those medium-term goals.
In any case, it would be nice if viatopia had more attention, as it seems like a promising way to do "one thing at a time" and potentially reduce the confusion (and risk of error) of trying to develop full-length plans from the outset.
(I know that those conversations have happened here and in plenty of living rooms. I'd be very interested in links to anything one degree more public, if anyone is aware of them!)
What do you think about end-of-life preservation as a way to shift humanity towards longer-term thinking? The idea is that a robust, scientifically-valid tradition of end-of-life preservation might shift everyone, including very young people, towards thinking more long-term, because they will correctly perceive themselves as having a personal stake in the long-term future. Would love to hear your thoughts.
Maybe something like a "long re-friending"? "Re-friending" (or just "friending") seems to me to capture the "coherent", "grow up further together" part, and to be less authoritarian-sounding than "correction." Goes well with claims AI made by labs are minority elites trying to run away with the world's decisions and stakes, better if we figure out how to be friends across factions first.
I feel like the reasons why MIRI circa 2014 moved away from talking about "Friendly AI" (because it was too anthropomorphic, etc.) and instead imported "alignment" to talk about the ~same issue also apply here. The "(re)friend(ing)" frame may make sense insofar as it's meant to be used internally, but probably not a great choice for something that is intended to be more widely broadcast.
"Long alignment" seems not terrible, but there's the whole "alignment to what", as well as "alignment of what". You can align yourself/[your values] to a Landian monstrosity, which you most likely don't want.
How about "long healing"? It is spiritually/conotationally similar to "correction", but I think it satisfies your desideratum of not having authoritarian connotations. It denotes things fitting well together ("alignment"), but also that not all ways of "fitting well together" are "created equal", and that there are better/[more "natural"] and worse/[less "natural"] ways for things to fit well together.
Unfortunately, “The Long Correction” could be about all sorts of things unrelated to AI, which makes it trivially hijackable, and therefore a bad name.
It could just as easily refer to:
If you’re trying to rebrand “AI Pause”, you should make sure that, at the very least “AI” is in the name somehow.
An acknowledged scope limitation is not the same as a flaw, and I'm not sure which of these you mean by "problem." The explicit purpose of an AI Pause is to create time for some intentionally unspecified longer-term solution, on the grounds that there are many competing theories as to what such a solution looks like, whereas needing time to implement--and decide between them--is a common factor.
Regarding the reflection challenge, what about approaching it from the other direction? That is, what it would take to redesign the environment such that human propensities are favorable, rather than something needing correction?
One way of categorizing knowledge building is as:
1. Evolutionary = lots of parts, iterated in parallel, keeping what works in context.
2. Engineered = stacking modular abstractions. Tested against and developed for a context, but more fundamentally held to a standard of internal consistency.
Human flaws can be mostly understood as primarily thinking according to evolved processes, which run into systemic problems when out of distribution, then using engineered thinking processes to correct for this distributional shift...but the latter is stretched way beyond its capacity because it was only designed for mild and temporary out-of-distribution moments. We could deal with this system failing by strengthening up our engineered thinking methods, improving mental flexibility to the point where it can handle everything we can expect to have thrown at it...or we can look for ways to lighten the cognitive load. One could call the latter approach "social refactoring."
A high level example of what social refactoring might look like: computing the distributional shift on the societal equilibrium of any given innovation as an externalized cost, which then gets folded in to the more generalized externalized cost tax that (in this hypothetical world) fixed all the more legible global threats. Such an incentive realignment is upstream of the refactor itself, which is the resulting adaptation, where specifics are harder to predict (but maybe a worthwhile project nonetheless).
An acknowledged scope limitation is not the same as a flaw, and I'm not sure which of these you mean by "problem." The explicit purpose of an AI Pause is to create time for some intentionally unspecified longer-term solution, on the grounds that there are many competing theories as to what such a solution looks like, whereas needing time to implement--and decide between them--is a common factor.
Good point/question. My intended meaning is closer to the former, and basically the "problem" is referring to the fact that I lacked a good handle for the concept that I often want to invoke, with "AI Pause" and "Long Reflection" being the closest ones. I think what you say here makes sense and I don't intend for my new handle to replace "AI Pause" where "AI Pause" is more appropriate.
The rest of your comment seems like an interesting idea for solving part of the challenge, worth keeping in mind, and fleshing out and discussing in the future, when we have this "generalized externalized cost tax" that we can build your idea into.
Reading your footnote 2, I got reminded of Rutger Bregman's The School for Moral Ambition. Rutger Bregman's "Moral Ambition" project (the book and the School for Moral Ambition) is one real attempt at this, though notably not from inside EA. He explicitly reframes doing good as the ambitious/high-status choice rather than a guilt-driven sacrifice. He is targeting elite-university students headed for McKinsey/consulting/finance, running selective competitive fellowships modeled on prestige-recruiting machinery, and defining moral ambition as "the will to be among the best, but with different measures of success."
What's suggestive is that this had to happen outside EA, partly as an explicit critique of it (Bregman frames EA as too guilt-coded, too analysis-paralyzed, not aspirational enough).
That said, Bregman doesn't engage with the actual meta-question you're raising — whether openly naming status as a lever backfires. He just uses the lever; there's no visible theorizing about whether the explicitness itself is corrosive.
Thanks for the pointer, which I didn't know about, but yeah I mostly mean discussing/analyzing the relationship between status seeking and altruism, strategizing how to best take advantage of it, including how to avoid problems caused by misalignment between the two (like in this example), and not just going ahead and pulling the lever as you put it.
Don't you think that the "mysterious" progress on morality comes from the fact that the incentives for it and the easyness of it both increased via technology ?
1) the incentives
I mean, beeing nice has always had the side effect of earning trust. I don't mean that selfless actions do not exist. For instance, long term vegans generally do not gain anything in their choice to refuse animal consumption. But one must admit that beeing nice in public has some benefit for a person.
With the appearance of video technology, showing how nice one is has become more rewarding: When you are a TV start with a million followers, you better act as morally prescribed by your followers otherwise you will get hardly punished socially.
Therefore, seemingly moral behaviour has become the new norm in society because every major figure is under scrutiny and therefore plays the "moral" persona in public which becomes the new norm.
2) the easyness
On the other hand, being seemingly moral has become easier since the industrial revolution: Being rich in 2026, you don't need slavery, you just pay servants for little chores. Being a noble in 1526, slaves would be more usefull for a lot of things.
in 2026, we also have more time to think about these questions than in 1526 when people had to struggle against the laws of nature day by day without the help of the machines.
Therefore, I don't think there is anything mysterious in the seemingly better morality of today's world.
3) Today's world
But I think the treatment of animals and the reaction of most people in front of this subject screams that this seemingly better morality is 99% fake. People will say they are against animal cruelty because it is what is expected to be said in society to be seen as a moral person but will continue to pay to support it because it is what is expected to be done in society to be seen as a normal person.
Therefore, I don't think much evolved on a deep level. The only evolution was the structure of society. In this society, we still have 98% of people for whom being part of the group outweighs by far being true to one's values.
4) The future
Therefore, expecting this "mysterious" process to continue making humans more moral seems like an empty wish to me. The only thing would be to expect having an evolution of the communication technology that would reshape the incentives somehow. I don't really see that coming and would put more in the idea of developping a strong theory of alignment. (Possibly through automated AI research combined with human critics)
I read your post, and I had thoughts about it. I made a vocal about it and asked fable to improve it.
I agree "Self-Correction" is a better name than "Long Reflection", though the post doesn't say why. Here is my reason: "Reflection" suggests the fix is more thinking. "Correction" admits the fix is changing what we are. That's the right framing.
But I disagree with most of the list. I think it mixes three different kinds of "flaws", and they call for very different responses.
1. The metaethics flaws are based on a framing I reject.
"Not having a workable moral framework" assumes that morality is a research problem: there is some true target out there, consequentialism and deontology are our candidate theories, and sadly they all fail. I think this picture is wrong from the start.
Here is the alternative picture. Tribes that coordinated on rules like "don't kill members of your own tribe" survived. Tribes that didn't, died out. Morality is the name we give to those rules, seen from the inside. Philosophers came much later and tried to fit general theories to this data. Utilitarians tried numbers, deontologists tried rules. Of course the theories all "have serious problems": they are rough compressions of a messy evolutionary process, not failed attempts at a real target. I wrote up this genealogy in more detail here: Dissolving moral philosophy.
To be fair about what this view doesn't give you: it's descriptive. It never crosses Hume's guillotine, and some philosophical questions stay open. But this changes what the post has to argue. "Humans lack a workable moral framework" becomes "here are the specific open questions we must answer before doing anything irreversible". That list would be much shorter, and much more debatable, than flaws 1, 2 and 7 suggest.
2. The status game flaw might be a Chesterton fence.
I see the same thing Wei sees: careful strategy and philosophy get low status in most places. Spend ten minutes on LinkedIn. But before calling it a flaw to fix, ask why the fence is there.
One possibility: society under-rewards philosophizing because, on the margin, doing things beats theorizing, and a culture that gave top status to meta-level reflection would get little done. Another: status and power are what motivates most people to do anything at all, especially now that religion doesn't. Remove that and I don't know what's left.
Same for institutions. Yes, it's annoying when a politician's mediocre report gets 200 likes and a truer analysis gets 5. But part of what holds society together is that people defer to institutions somewhat independently of the quality of their output. Legitimacy is fragile. If you "correct" deference away, you may not get a world of better epistemics. You may get a world where nothing holds, and all institutions fall apart. These fences should be moved carefully, and that cuts against listing them as simple flaws.
3.On calibration (flaw 3), the evidence is weaker than presented.
FTX looks to me like fraud plus bad incentives, not philosophical overconfidence. And competence varies a lot from person to person. Some people (Davidad comes to mind) seem to have settled enough of the philosophy to move on and build. "Humans are badly calibrated" erases exactly the variation that matters.
What I think the actual bottleneck is.
The one flaw that does real work in my model is the one the post puts in a parenthesis at the end: we are bad at large-scale, long-horizon coordination. See climate change. That one is a true precondition, both for containing the risks and for running any Long Self-Correction at all. Most of flaws 1 to 8 either dissolve (metaethics), turn out to be fences (status), or can be fixed in flight (zero-sum values). Coordination can't wait, and it's a political problem more than a reflection problem. See: The current bottleneck is political will, not research.
I'll grant flaw 8 (over-optimistic partial solutions) has real force. My own position is exposed to it too.
My main issue is that some of the flaws are far from being "hard but soluble by a Novel Insight And Long Self-Correction": insoluble as stated, too easy or outright erroneous.
My closest candidate solution to these problems is not some Philosophical Insight From The Future, but broadly educating people.
My main objection is that the existence of safety-pilled AI labs might have had a higher bus factor than ablating Yudkowsky.
Including a uniform-like one, as happens in the Epilogue of AI 2040.
However, resources of Earth or the Solar System can also be quickly reallocated between humans.
I like the idea of having a concept for this, though "Long Self-Correction" is probably not the most catchy term.
It seems reasonable to to say that if science and technology outpace political and cultural ability to understand and control risks, then humans aren't safe. The same seems true when politics and culture outpace science and technology.
It also seems reasonable to contend that the arts and humanities act as a safety valve, releasing excess pressure in the dynamics and indicating where the pressure comes from.
If we start from the assumption that these reflect a real dynamic, then it is plausible that we can derive a system of equations describing the tension, and we can test that against historic events.
Should it be possible to produce a predictive model that has reasonable accuracy, then the question of how we would know when humans are safe has a definite answer: when the model shows that the system is no longer chaotic but is oscillating or fully stable.
If such a model cannot be derived, then we cannot know I any absolute sense if we are safe or not. If we can, then we possess a means of knowing not only if we're in a safe zone but how far into that safety zone we are.
So although I do not know what it would take to be safe, or whether we can know if we are, I am comfortable suggesting we can find out.
political and cultural ability to understand and control risks
I did rather suspect that the answer was going to be "never." Unfortunately, that makes it a non-starter.
Why isn't there a version of EA that explicitly talks about how to leverage people's status motivations to do more good for the world? It's very possible that explicit talk about status is actually counterproductive at least in the short run, e.g. it heightens status motivations and makes people less altruistic, but then do we just march into the future while blindfolding ourselves to this aspect of human nature?
Wait, what? As far as I understood it, that was explicitly a main consideration in the 10% pledge Schelling point, though it's possible that I may have just came up with that argument myself and misattributed it.
This post got me thinking, though greater intelligence does not necessarily entail greater morality, I wonder whether moral coherence may be required for continued, open-ended intellectual development. Significant moral incoherence might eventually force an intelligence to correct itself, suppress further reflection, stagnate, or ultimately collapse if it develops too many contradictions. If so, maybe this would act as a limit on the orthogonality thesis (i.e., intelligence and values may vary independently across bounded levels of capability, while some value systems remain incompatible with indefinite cognitive development). From that perspective, the goal may be to help the combined AI–humanity system avoid becoming trapped in these developmental dead ends by discovering more coherent moral models (i.e., ones capable of supporting continued intellectual growth).
summary of the flaws that I have in mind:
- not having a workable moral framework (consequentialism, deontology, virtue ethics all having serious problems
The FTX debacle shows we gave a workable framework. If SBF had followed the rules, it wouldn't have happened.
As far as practical, good enough ethics goes, deontology, maybe with consequentialist justification, wins. Every organised society uses it.
A lot of rationalists seem to assume that “as you gain more IQ points, learn more facts, and think harder, you’ll naturally converge to the optimal theory of ethics.” Like the OP, I don’t think this is true in general.
A lot of rationalists also assume that you can since ethics by a solipsistic process that doesn't take other people into account
One problem is that human value might be inherently multidimensional
It clearly varies a lot, and therefore might as well be called human values.
What do values have to do with ethics, anyway? There is probably some relationship, and probably not simple two way identity.
Rationalists always equate ethics with values, rather than obligations, or virtues. There appears to be no specific reason for this. There are obvious things examples of values that aren't ethically relevant, eg. I can value tutti frutti over vanilla, but it is not ethically relevant. The Three Word Theory, "Morality is values" doesn't hone in on the topic, and probably leaves out important things like obligation.
in practice, human morality is a kind of status game that actively disvalues careful strategy and philosophy in most places
I would build on this: human morality is deeply corrupted by politics, and our political culture is arguably worse than it's been in decades.
long reflection = i want that. please. please let me have that. you have given a name to a thing that i have deeply, deeply wanted for longer than i can remember.
long self-correction = hello, human resources?
Thanks for the post.
The way I look at this is more that humans, and societies, are self-improving machines that learn by trail and error. I think things can go wrong on humanity scale if new technology gets incorrigible. AI takeover is an example of this, as is for example biodiversity loss (which I expect to be incorrigible even with max tech, but I hope to be wrong).
As long as we maintain the ability to make errors and live through them (that is, as long as our errors are small and slow enough to be corrected in a trail and error fashion), I think any required self-correction and/or self-reflection should happen by default.
It would help if we would get better at self correcting (faster, more reliably eg ending up at the right response), although that's not strictly needed for a good long-term outcome. It would mostly limit the casualties of corrections.
I see our job as making sure no incorrigible things happen (such as AI takeovers). Maybe that's a much easier target than requiring major human improvement?
I think it matters less what you call it, and more the institutional incentives/culture/etc. which drive the process.
Indeed, moral philosophy is the main bottleneck for AI progress.
Two thoughts:
All of these seem relevant but as a starting point I would put something like most humans having a lot of technical debt (https://sashachapin.substack.com/p/review-meditation-from-cold-start), not feeling safe and okay in there here and now. Maybe not being enlightened. And then we could get to work on the difficult problems you bring up. Of course, there is the trap here of going off the rails using spirituality or me already being off the rails.
Sometimes you have to try something and get some real data about the system under study. The 8 listed points seem very heavily chewed over by now. Many have been studied for over a thousand years. Rationality only helps whent here is some real meat to chew over. Without solid data, additional rationality is essentially hallucination: internally consistent but not applicable to the real world.
We don't really know yet what advanced AI will look like. The models being trained and used right now are good little loyal helpers. Pausing to regroup right now would be like trying to understand how to build a skyscraper made with steel I-beams by building a treehouse out of wood. Almost nothing is the same about how you build those safely, so what you learn from the treehouse will not help you that much on the skyscraper.
As important as AI alignment is, it is not yet the main risk we face from AI. The main risks are due to human misuse and the AI dutifully doing what a questionable human asked for.. For example, the US is the only country right now whose government can crack computer security systems with Mythos-level capability, including subversion of combat drones. For another example, new AI model companies are hard to start right now, leading to a lack of competition and diversity. These are the policy problems right now, and to get to the alignment problems, we have to actually build the thing we want to study.
I propose the Long Self-Correction[1] as an alternative name/idea/concept to AI Pause and Long Reflection.
Problem with AI Pause: Pause until when, and for what purpose? Presumably to make AI (that we'll build later) safer, but the deeper problem is that humans aren't safe, and can't safely serve as builders, overseers, or alignment targets for powerful AIs.
Problem with Long Reflection: It seems to imply that the main problem with humans is that we just haven't had enough time to think, that reflection is the main thing we need to do more of, and then we can get on with building powerful AIs or other technologies. Or that if we build aligned AIs that sincerely help us think a lot more, or do the thinking for us, then things will turn out fine.
So I think we need a catchy handle for a related but distinct idea, that humans aren't ready to build AIs or other extremely powerful technologies, because we're currently too flawed, in a variety of ways, and it will take a long process (which may or may not end up succeeding) to fix those flaws.
A summary of the flaws that I have in mind:
(This list focuses on key bottlenecks that seem hard to fix even with AI assistance or intelligence enhancement, and isn't meant to be a complete list of human flaws / safety problems. It ignores e.g. that the median human is ignorant of many important issues, and that we're currently quite bad at complex large-scale coordination such as passing/implementing close-to-optimal government policies.)
My main hope for a Long Self-Correction eventually succeeding rests on the fact that humans have seemingly, mysteriously, made progress on these issues over a very long period of time, so if we preserve the environment in which we can seemingly do this, and not give anyone or anything the power to permanently derail such progress, then maybe we can continue to snowball The Correction until we reach a point when we can rightly justify reshaping the universe according to our volition.
It will probably be shortened to "The Long Correction" at some point if it catches on, similar to how "outer space" is now often just "space".
Why isn't there a version of EA that explicitly talks about how to leverage people's status motivations to do more good for the world? It's very possible that explicit talk about status is actually counterproductive at least in the short run, e.g. it heightens status motivations and makes people less altruistic, but then do we just march into the future while blindfolding ourselves to this aspect of human nature?