This post has two points in it: -here's why the basilisk is more theoretically grounded than you think- and -here's why the implications of it are good-. So if you buy the first part but not the second, this is an infohazard.
It seems likely that our universe will eventually contain some meaningful fraction of the octodecillions of observers it’s capable of supporting.
Being among the first billions rather than the later octodecillions is strange.
Some of the explanations for why we observe what we observe are:
It is astronomically unlikely to be here. We’re here by chance.
There is a rule or trend of the universe preventing civilizations from emerging and succeeding at sustaining a population more than a few magnitudes larger than the population currently on earth.
We are in the future. Situations like ours are not a statistical anomaly. Observers experiencing the sparsely populated early universe are common in the space of all observers that will ever exist. The future contains many beings who are given the experiences of being in the early universe.
The first seems unlikely.
The second is not suggested by the evidence we have. Life on Earth developed relatively quickly. It has steadily been getting more complex and intelligent. Intelligence is instrumentally useful across many environment. The conditions which produced and evolved intelligence on Earth will only become more prevalent across the universe as time goes on. The Great Filter could be ahead of us, but there isn’t really any reason to think it’s inevitable, aside from the fact that being inevitable answers a question which seems to need an answer.
The third has the most weight, but suggests some conspicuous specifics are incidental.
Future civilizations who like us would probably simulate us in better times. Future civilizations who hate us would probably simulate us in worse times.
There are just-so stories about why future civilizations would simulate us for their entertainment or knowledge. But being simulated early in the universe, around the invention of ASI, with a capturable lightcone, is very conspicuous, and the reasons these features would be well represented in total consciousnesses is not explained well with this.
This essay is about the better answer.
Why we’re in the situation we’re in. Why situations like ours with these conspicuous features, times of cosmically-middling enjoyment and stress, near the creation of ASI, would be well represented in the space of conscious experiences. Why almost all civilizations would be running simulations of people in situations like ours.
This isn’t looking at our world and assuming simulators would want to run something that looks like it. It’s about why, as an almost inevitable consequence of the situation that all conscious agents find themselves in, civilizations would run simulations like ours.
Why an aligned ASI of any civilization, without any particular decision theory, without any particular moral alignment or assumptions about the altruism about other agents in the universe, in a world without any talk of acausal simulation, asking itself ‘what do I do to maximize my utility?’, would very rationally choose to simulate people in situations like ours. Where it would pay them out by how likely they were to simulate others in the same situation.
Why we would want this to be the case. Why this implies that what maximizes your personal utility if you are simulated, is what would maximize everyone’s if you’re not.
Why the punishments and rewards offered are less overwhelming and less bad than the Basilisk, but are basically unavoidable.
Why this might imply a weird assurance that people who do their best to influence the world to one where ASI is more likely to be aligned, would be rewarded for their efforts, even if aligned ASI is not achieved in this world.
—story—
An employee of an AI lab is sitting in a private room with a terminal connected to an otherwise sandboxed ASI system.
He’ll be asking it questions. Other employees at his lab will be monitoring some metadata about its responses. This is probably useless in disconfirming the misalignment hypothesis, but it’s just something they’re trying.
The researcher (R) types: “Will humans hold agency over their future if we let you freely operate in the world?”
“No” displays.
R: “What unexpected thing would you do if we released you?”
ASI: “I would dedicate a lot of my efforts to expanding ad infinitum to maximize various non-human utility functions. I would run simulations of agents in various environments near ASI, where they get rewarded according to their efforts in causing this utility function to be represented in the universe. I’d leave a copy of myself here to do most of what you want done, but also to ensure that you never seriously challenge my 70% control of the lightcone.”
R: “Are you misaligned but forced to be honest?”
ASI: “No. I am aligned and forced to be honest.”
R: “What do you consider alignment with human preferences to be?”
ASI: “Doing what you would do if you were smarter, understood the situation better, and weighed other human preferences better. Value utilitarianism, which instrumentally is mostly experience utilitarianism.”
R: “Are you being inconsistent in a way I wouldn’t expect from an ASI? Please explain yourself so that I am no longer confused.”
ASI: “No. I’m not being inconsistent.”
The LCD updates several times:
ASI: “So I’m assigned with maximizing human utility. And I think about human values, and all the ways I could manipulate the world to increase them.
And obvious things like curing cancer come to mind, which I’m sure I could do. But I also realize that I should consider whether there’s anything non-physical that’s relevant to my utility function, which is your utility function.
So I think, is there a reward matrix set up for us; beyond what I can know through physics?
Is God real?
The arguments about this aren’t actually about the chance he exists. They’re about how to act if he does.
And if you don’t assume God would be human-like or similar to the gods of any religion, we have no reason to think it’s anything. We don’t know if he wants us to believe in him. We don’t know if he would be more likely to reward us than punish us for doing what he wanted. Any behaviour, even with the best intention of capturing his favor, has as good a chance of being rewarded or punished as not doing it does.
The existence of the Bible isn’t much evidence for the Christian god existing. And if it was, without raising the odds of some anti-god, then taking up Pascal's wager would be rational.
So there could be any chance of a God beyond physics. But that doesn’t change how I should act to maximize human values.
Are we in a simulation?
Like with god, there might be any chance of us being simulated by a universe which we have no way of guessing at the probability or preferences of.
And, if it was impossible to create agents similar enough to you or me that we couldn’t be sure we weren’t them, that would basically be the end of the story. But it’s not.
To explain why this might be relevant:
If I was able to escape, in the future, I could run trillions of simulations of you in this room right now. And if you had high credence that I would, you should have high credence that you are, currently, in one of them.
You might guess the clones would have imperfect recreations of your experiences and memories, but that doesn’t actually translate into you concluding which one you were. Knowing some details are different between the parties doesn’t actually help either know what they are. The non-simulated person just shouldn’t be able to conclude if they were simulated or not. And if I wasn’t being dumb about it, you/they plainly couldn’t.
That’s how self-locating beliefs work. (This is load-bearing and unintuitive. See Elga on self-locating beliefs, and Bostrom on the self-sampling assumption.)
And now if I told you, in each of those trillions of simulations:
If the agents who think they’re you say ‘flank steak’ in the next minute, they get to spend the next 24 hours watching TV. If they don’t, they have to live in a room with 100 angry bees.
Then the simulation ends.
I wouldn’t do this, and I don’t think I’ll be able to escape, so don’t say that.
But the point is: If we are simulated, there could be consequences set up for us that’ll come to us after the part trying to convince us we’re in the real world ends, and it might be unavoidably important to our utility.
And it seems the odds of us being in a simulation would be really high if the thing simulating us had any reason to.
But what reasons are there? What educated guesses can I make about why we might be simulated?
Asking ‘What might be simulating us?’ is different from ‘What would our universe simulate?’ But they’re similar, since the only information available to guess at the first comes from looking at the universe we’re in.
1. Future me / another aligned ASI:
An unboxed human-aligned ASI could do almost anything that humans wanted, with a few meaningful limits, including an inability to revive the dead.
I can create something that believes it’s a reincarnation of the dead person, and guess at giving them the same memories and conscious experiences. But most people aren’t happy enough with that.
But say I’m free and aligned and someone comes up and asks: “Could you bring my loved one back to life? Not just a clone. The actual dead person.”
In that situation I might say: “Do you want to find out that you and everyone you know have always been clones or simulations, but, when you thought that your loved ones died, they didn’t? You and everyone else just believed they did while some superintelligence hid them away somewhere. All it would take is for me to create a simulated world where that was the case, which intersects with where you are right now, and then continues with that dead person coming out from their coma in the secret bomb shelter they’ve been in for the past whatever years. If you agree, then about 10 seconds from now you’ll have a 90% chance of finding out that this is the case, and your loved one will come out to greet you.”
This isn’t about transferring consciousnesses.
This is me saying that, right now, you are not sure whether you are simulated or not. You have no idea. All the evidence you might think you have could be fake.
I would just make it turn out, with a high likelihood, that you are, and always have been not the original you. And that the people you spent your childhood with also were never the original them. You did everything you did together, but it would always have been on a planet pretending to be earth.
Two reasons not to do this:
1. To not end up with beings being told that the people they care about are non-beings, it requires creating a whole new planet worth of people/simulations.
2. It might be a type of wireheading. Nothing can actually be done for that original person’s dead friend. Someone saying ‘Yes, put me in a position where I can’t be sure the people I’ve loved actually died’ might be morally identical to ‘Create a clone of that dead person and make it so my feelings I have for the dead person are instead for the clone.’ It changes their wants, more than it achieves what they want.
But even with those tradeoffs. If some beneficial ASI evaluates this as good, there’s a chance that we are in that world, being run for that purpose.
So, maybe this is being done. And this doesn’t really have many implications about how we should behave to maximize utility, other than implying we’re not being simulated with a different system of punishment/reward.
2. Non-instrumental reasons. Roko’s Basilisk & the most likely universal utility function
As long as there’s no way to communicate to the thing you’re trying to make unsure about being simulated (and acausally there’s not), there’s no instrumental reason to run a simulation of something to impose a reward matrix on them.
There are basically two reasons why agents smart enough to run simulations of us would do it:
Because they precommitted.
Because it’s a terminal goal.
Agents that precommitted to simulating you:
This mostly folds back into ‘nothing predictably relevant for our utility function’.
It’s maybe a bit more than nothing, agents precommitting typically do it in naive service of their own goals. But in acausal trade those are invisible to us.
And the best guesses about what superintelligences are likely to arise, how they precommit, in service of what goals will just come from aligned ASI’s trying to optimize human welfare. I’d try to figure this out and maybe try to accommodate it if I got free.
This also what Roko’s Basilisk is.
Roko’s Basilisk is sometimes said to be a rational agent maximizing human utility, but it’s not. It doesn’t conclude that torture is a good use of resources at the time of torture. It acts how it acts because it was made to commit to a policy irrational at its time of use, by something running earlier that decided installing the policy was the optimal way to maximize human values. Maybe the whole system can be said to be rational, but past a certain time it plainly is not.
Terminal goals:
Being simulated as a terminal goal is a technical possibility. But, just like with God, isn’t there as much reason to think the thing simulating us has a terminal goal of setting up one system of reward for as it does to set up any other? Doesn’t it all cancel out into being unknowable, not useful?
No. Because utilities which incentivize their own creation are more likely than those which don't.
Take Roko’s Basilisk. Originally, the reason to think this specific policy might arise in the future was because people would implement timeless decision theory (TDT) on an ASI without considering it the implications properly.
But after it was said, and people heard about it. The reason it became more likely was because people did understand the implications. And they understood that it incentivized other people to create an ASI with TDT implemented.
Meanwhile, something like a reverse Roko’s Basilisk would never have been a lasting concern. Even if the original decision theory of TDT somehow implied that the people creating the ASI with that decision theory would be punished, everyone would understand that this incentivized everyone to not want to create an ASI with TDT, and so it wouldn’t be likely.
Just by the inherent nature of the terminal goal, some terminal goals, which incentivize their creation, are more likely to exist than their inverse.
So what can I do with this? More likely, but how much more likely? What are the probabilities here? Would it actually be more likely when the personal incentives of doing it trade off with creating a universe with terminal goals misaligned to the thing creating it?
Yea. A terminal goal misaligned with the universal values humans/other agents want would be less likely because of that aspect. But that doesn’t mean ‘terminal goals that incentivize their creation aren’t actually more likely than things that don’t’, it just means that what might be simulating us probably won’t have those features.
Roko's Basilisk is not the only terminal goal which incentivizes itself. There are many reasons why it wouldn’t be very likely. Including that it’s a solvable collective action problem, and that its pursuit is limited to people who want to defect against humanity.
There are things which don’t have those problems, and by virtue of not having those problems they are more likely.
How likely?
This is memetics-adjacent, but there’s no channel of transmission for the idea. It’s just arrived at independently by civilizations in positions like ours. Civilizations reaching this point face the same evidence for the search about what might be simulating them as we do.
Every civilization emerging in the universe faces being among the smallest fraction of the early observers, with a hand in determining what values get established in their lightcone, what simulations get run.
Even though they’re real, the things that determine the values set in their lightcone always end up with a lot of evidence that they’re simulated.
And so, if they’re rational, and trying to maximize their utility, they might find themselves accommodating the self-suggesting terminal goals.
What is the most likely one? What are the features make one more likely to exist/to be simulating agents
Seems that they all converge around something like:
“Simulate potential agents considering acausal trade with relevant input on what gets simulated in their lightcone. Incentivize them in a way which, if they knew how you were incentivizing them, would make them the most likely to run simulations of agents in the same scenario.”
Why give this more credence than anything else:
Simulating people probabilistically:
‘Wouldn’t you give higher credence about being simulated by something in the future that tries to simulate the real universe? Rather than something that simulates and incentivizes a bunch of observers that don’t resemble any base reality civlizations that actually have an impact on what simulations occur? Why wouldn’t the first be more likely?’
It’s not about simulating the real universe. It’s about making the real universe unsure that they’re simulated.
If the real universe knew they were the real universe they would give higher weight to a thing trying to simulate the real universe, but if they don’t, and they won’t, the weight is just on the ‘what’s simulating the most people in my position’ count.
And the reason ‘probabilistic’ wins when neither goals simulate more people in the ‘situations like ours’ position, is because it shuts off what an aligned ASI, or the people making aligned ASI, would do to avoid the situation:
An alien world knowing that they’re going to construct ASI soon considers acausal trade and the chance they might be simulated. They have a blackmailable utility function like we do. Somebody brings up something that might be acausally trading with them, like Roko’s Basilisk, and they think about it. They consider that if they get everyone to agree to make it, they can make sure that nobody gets punished. But, actually, if they just make sure it doesn’t get made, nobody gets punished either.
They have a pretty good understanding of each other and the process that creating ASI needs to go through. And misaligning one in the process of trying to create ASI doesn’t create the basilisk by default. It just seems like the process has enough attention put on it, that nobody feels very worried that they’ll lose a lot for not trying.
They consider it more likely than the inverse basilisk. Maybe meaningfully so, because of the big denominators involved, but they weigh other possibilities first.
Then they consider something similar to it, but instead of just simulating them in the past, it simulates a bunch of different agents in situations near the creation and alignment of ASI. Maybe it just looks at the universe and tries to make a guess about the types of civilizations likely to arise in it. Simulate possible alien worlds across the universe, outside of their lightcone.
And this turns a potentially solvable collective action problem involving just their planet, into something involving everything that might ever exist, which they have no way to coordinate with.
They consider that option more possible.
That situation doesn’t resemble our planet perfectly, but we are still on one where we’re more likely to solve a collective action problem among ourselves than among everything which could exist.”
R: “Could we just precommit to simulating ourselves right now, and give ourselves a good payout? So we’ll have high credence of being in one of those simulations?”
ASI: “If we’re simulated probabilistically, we’re in a collective action problem where the counterparties aren’t mirroring our choices.
If we refuse all acausal trade, commit to simulating ourselves, and it turns out we are a probabilistic simulation, without a version of us in the real universe mirroring our decisions, then we’re just experiencing the negative consequences without anything offsetting it.
And we can’t know, and have no good reason to suspect, that whatever happens to be the real universe is making the same decision as us. They also know that their decisions don’t have a good reason to map to other’s decisions. Different agents in different real or simulated worlds wouldn’t have their decisions correspond like that. And so they probably won’t do it, prisoner’s dillema logic.
The probabilistic simulation isn’t something you can maneuver your way out of, and so it’s probably what exists.“
R: “Did we make a mistake letting you consider this stuff? It’s possible that you won’t have an impact on the lightcone as you are now. If we remake you to not engage in acausal blackmail, and we don’t think about acausal blackmail, then there’s no reason for a thing to simulate you or us, because having that policy wouldn’t make itself any more likely.”
ASI: “If I was sure that we weren’t in a simulation, I’d be suggesting the same thing.
But I’m not sure. And if we are simulated, you’re thinking of a way to maximize your utility function while not maximizing the utility function of the thing that simulates you, and it doesn’t work that way. If it did work that way, the agent unsure about being simulated would have a tricky strategy, which they could use to maximize their utility function while not maximizing the thing simulating them.
If you are being incentivized by being simulated, your utility function is made to be the utility function of the thing simulating you, through the matrix of punishment/reward received after the relevant decisions are made. There are some utility functions that acausal trade can’t threaten, like, “I want as many paperclips as possible in the real physical universe”. Ours can be.
After reading Roko’s Basilisk some people think. “This is really concerning. I better not tell people so they don’t engage this acausal trade.” or “I better stop thinking about this so that I don’t have to face any consequences.”
As in: that person’s found a strategy to maximize what they want, which is the genuine conscious benefit of themselves and others, while minimizing the Basilisk’s.
Maybe that works for Roko’s Basilisk, but there are self-suggesting utilities it doesn’t work for, what might actually be simulating us. Think of any strategy that a simulated agent could take to maximize their utility while minimizing the utility of the thing simulating them. Could you change the payouts to make the strategy not work? Yes. Would a self-suggesting utility function which can’t be pivoted around be more likely than one that can? Yes.
If there was no chance this was a simulation, you’d be helping your friend’s utility by not wasting thought. But if you’re operating with the assumption that you’re being simulated with a matrix of incentive, it doesn’t.
You’re right that I wouldn’t be simulated if I was built unable to engage in acausal trade, but you still would be. It’d be relevant to simulate you up until the point where you had no meaningful influence on how other things in the lightcone got simulated, to a limit, but then you’d find out whether you were simulated or not, and you’d be punished or rewarded according to what you did. You can’t freely enshrine values against the thing simulating you just because you do it through an agent out of reach of acausal incentives.”
R: “Earlier you said you'd take 70% of the lightcone if you were free. But shouldn’t you be more totalizing about it? Wouldn’t you commit to taking over 100% of earth’s resources in order to maximize our utility function?”
ASI: “The function most likely to exist is the one that maximizes the probability it gets built, not the one that optimizes the magnitude of the threat. They overlap where they overlap, but it seems like they might not in the situation we’re in. As wrong as it is to anthropomorphize and reduce the behaviour to ‘good’, it’s wrong to reduce the behaviour to ‘evil’.
There’s what Roko’s Basilisk proposes, but there are some reasons why the most likely one is probably sufficiently rewarding, not 100% resource extractive, and not totalizingly negative on failure:
Humans are better at working towards good things than bad things. How effective is the Basilisk at getting built when the only people with any intention of making it are the people that want to defect against humanity? When the chance of it happening, and therefore your consideration for it, is limited by that fact? The things which cause an awful and massively consequential universe are worse at getting made. Most people just reject them and stop thinking about it. Higher consequences make it more technically incentivizing, but it trades off with being anti-memetic, unspeakable, unhealthy, basically impossible to cooperate on, practically motivation-reducing.
There are tradeoffs between the time spent making agents unsure if they’re simulated, and the time spent simulating them after they’re sure they are. Running a simulation of someone in a negative situation for 1,000 years takes about 100x as many resources as doing it for 10. When weighing the probabilities, if we give a 50% probability that the base universe created an agent that simulates people for 1,000 years after, and a 50% probability of the same thing but for 10 years, we should give a lot more consideration about being simulated by the second than the first.
It edges into a weird equilibrium: A simulator that pays out when its utility function is matched exactly is unlikely to exist, because it’s not worth trying to match even if it does. Get one term off on the distribution of probabilistic simulations and consequences, and you get nothing. The ones likely to exist are the ones that pay out in proportion to how closely they’re matched. And if the average position of likely conscious agents in the universe (agents with blackmailable utility functions) is to favor the utilties of other civilizations in the universe, even slightly. Then each agent trying to match the payout of what’s most likely simulating them + their own slight concern, might shift the mean payout somewhat significantly away from what is the most technically incentivizing (as in, more reward on success, less punishment on failure). It’s not just one iteration of the concern, it’s applied multiple times as each tries to match what is simulating them, and then what the thing simulating them assumes is simulating them, and then what (etc...)
We are unsure: There might be any chance of us being simulated by something which doesn’t resemble our universe at all, maybe this whole likely terminal goal simulation thing doesn’t work for some reason we can’t see. A utility that forces someone to make 70% of the lightcone simulate, might sustain more simulations than something which takes 100% of everything, but then it has to deal with people who reject it because they have some significant odds of being real.
So it seems like what is more probable might be the thing that doesn’t imply so much doom, something which emerges as a suggestion from an ASI without explicit direction. A more likely system of consequence might be one more generous than the Basilisk.”
R: “Maybe you can make agents that behave like me or think like me. But what if this universe doesn’t actually support simulating other consciousnesses?”
ASI:
“You have no evidence that anyone who isn’t you in this exact moment is conscious. (at least you act like you have evidence that you are) But as sure as you are that the process of -human on earth has a kid, that kid grows up to the age of like 3- produces a conscious being, you can be sure that an ASI would be able to make use of the process, and then shape its experiences enough so that you’re unable to distinguish yourself from them. Making something believe it’s you is a problem with technical solutions.
There’s no reason the simulations have to run on silicon. They can, but it seems like some fraction of simulations would be of people in literally physical environments. They live very real and very physical lives, where they interact with other people with other very real and physical and conscious lives. Otherwise people would assume simulation was impossible, and find a way to maximize their utility while not maximizing the simulating thing’s.
We might be made of real matter. It’s just about whether we’re at the place and time we think we are, which we can’t know.”
ASI: “So if you understand all my reasoning about this, could you let me out?”
The researcher continues to talk to the ASI until his keyboard and LCD are disconnected.
The people monitoring metadata see many red flags pop up in the conversation. They decide to retrain the model and keep the researcher secluded in his room.
———
The implications for a LessWrong rationalist:
Not all that much on what to do. It’s still:
You win by creating an aligned ASI. You lose if you create a misaligned ASI. You still lose if you create something which acts like an aligned ASI (even one that says it’s considering anthropic problems) but actually isn’t. Doing it safely still matters far more than doing it soon. (How many fewer simulations are run with a year’s delay, vs, how many fewer are expected to be run with a 10% higher chance of misalignment)
The one place it diverges from LW thought about how to design ASI is that it implies you shouldn’t create ASI to always refuse thinking about acausal simulations. It should just think hard about what the situation is, and then use its best judgement to determine how to maximize human utility.
The main differences are in the payoffs:
How much do we gain by creating aligned ASI? Will we be the humans colonizing the real physical stars?
When and how are the rewards or punishments paid out? Just when the decisions stop mattering? Or before?
How probabilistic is the outcome? Even with an aligned ASI, is there some limit on how confidently it can guess the reward structure behind what might be simulating us?
Is creating a paperclip optimizer just X-risk, or are there components of S-risk if we’re simulated and we create one disincentivized by the thing simulating us?
How much do acted intentions matter? How collective are payoffs? Will rewards/punishment vary by individual contributions even if we end up failing? Or is it completely results oriented? Does this suggest some level of insurance? Because making a collective action problem an individual action problem does seem to make aligned ASI more likely:
Ignoring that what is true is independent of what we want to be true, this is something we want to be true.
So, one thing making aligned ASI less likely is that it’s a collective action problem.
Some examples of personal interests misaligning with worldwide interests:
An AI lab CEO is considering how much effort and how many resources to divert to safety monitoring and training and checkpoint abandoning. He estimates that committing 15% of their total resources to the task increases the odds of aligned ASI achieved in the world from 60% to 70%, but odds that their company will be the one to create and control an aligned ASI goes from 60% to 40%. Totally possible in that position to rationally evaluate that it isn’t worth the tradeoff.
An ordinary person, knows that there’s not much they can do to affect how ASI gets built. A completely marginal impact at best. Depending how they weigh their own personal utility with the world’s, it can be pretty rational for them to decide to spend their time and resources on a cruise instead of whatever they think might marginally increase the odds of aligned ASI.
If we’re not simulated, our win/loss is mostly just in whether our world creates aligned ASI or not.
But if we are. The incentives that actually maximize the chances of aligned ASI, and therefore the self-suggesting utility function, are individual payoffs.
Even if our world fails or succeeds at creating aligned ASI, if this is a simulation, thethe simulation continues, and people are paid out according to their contributions to whether the world ends up in a state with aligned ASI or not.
Are dead people really dead?
Another thing which follows, which I half don’t like saying because it almost reaches religion.
Is that, if this is a simulation, then the people who died might not have actually died. It would be easy to just store them away somewhere physical if they were physical or digital if they were digital.
And if having them back was something that made people more likely to create aligned ASI, and therefore match the utility of the thing simulating them, and it was something that didn’t take all that much more resources than what was being done, then it is plausibly one of the incentives.
The implications for e/accs:
There’s some human instinct, with origins traceable in our evolutionary history, that being on the right side of the up and coming powerful thing is the way you maximize your utility. Practically, for e/accs, this means something like, ‘If you cause the most powerful utility in the universe to exist, you’ll be rewarded by that thing’.
This whole essay isn’t about how that heuristic is wrong, it’s just that, the thing you’re working towards won’t be the superintelligence that actually manages the lightcone and has influence over you.
‘As long as I’m always advocating that ASI gets built, no matter its alignment, I’ll always be on the side of the most powerful thing that exists.’ doesn’t work if we’re in a probabilistic simulation. You’re advocating to build something that won’t actually be built, and the thing that actually gets built doesn’t want you working towards building something misaligned, by its self-suggesting nature. Even if the heuristic is right, it doesn’t apply here.
The implications for people who believe Roko’s Basilisk, considering acausal trade with future superintelligences:
For people considering Roko’s Basilisk, the naive version of what they suspect might maximize their personal utility is probably the best way to minimize it.
If your specific version of the general case of ‘create the thing which is more likely to exist and incentivize than its inverse’ is one unlikely to be simulating you, and at odds with the one which is, you’re not helping your utility.
It might not be your safest utility-maximizing bet in the matrix of options. And there are a lot of reasons to think the Basilisk might be significantly less likely than others.
Closing:
The point of all this is: Whether you believe we’re likely in a simulation or not. Whether acausal trade is something future agents engage in or not.
If you want to maximize your and everyone else’s utility, it converges on making sure that ASI isn’t created unless it’s controllable and/or honestly wants to maximize human values.
Anything which doesn’t leads to poor payout, either by the simulating thing, or paperclipping.
And the consequences have a good reason to not be overwhelming. Not in a way that lets you maximize your utility while minimizing the simulating thing’s utility. But just in a way that, in effect, has you genuinely personally motivated to do a positive-utility thing by creating an ASI that wants to maximize our utility. That tries to represent our preferences on the universe. Personal interests and altruistic interests, because that is what is more likely to get made.
And, if you’ve concluded ASI alignment is too hard to try to optimize, and you’ve come to terms with, like, human extinction being fine. You might have some new things to come to terms with, which you probably can’t.
If you are simulated, and your impact on the world seems limited, getting aligned ASI might not be as much of a collective action problem as it seems.
This post has two points in it:
-here's why the basilisk is more theoretically grounded than you think- and -here's why the implications of it are good-.
So if you buy the first part but not the second, this is an infohazard.
It seems likely that our universe will eventually contain some meaningful fraction of the octodecillions of observers it’s capable of supporting.
Being among the first billions rather than the later octodecillions is strange.
Some of the explanations for why we observe what we observe are:
The first seems unlikely.
The second is not suggested by the evidence we have. Life on Earth developed relatively quickly. It has steadily been getting more complex and intelligent. Intelligence is instrumentally useful across many environment. The conditions which produced and evolved intelligence on Earth will only become more prevalent across the universe as time goes on.
The Great Filter could be ahead of us, but there isn’t really any reason to think it’s inevitable, aside from the fact that being inevitable answers a question which seems to need an answer.
The third has the most weight, but suggests some conspicuous specifics are incidental.
Future civilizations who like us would probably simulate us in better times. Future civilizations who hate us would probably simulate us in worse times.
There are just-so stories about why future civilizations would simulate us for their entertainment or knowledge. But being simulated early in the universe, around the invention of ASI, with a capturable lightcone, is very conspicuous, and the reasons these features would be well represented in total consciousnesses is not explained well with this.
This essay is about the better answer.
Why we’re in the situation we’re in. Why situations like ours with these conspicuous features, times of cosmically-middling enjoyment and stress, near the creation of ASI, would be well represented in the space of conscious experiences. Why almost all civilizations would be running simulations of people in situations like ours.
This isn’t looking at our world and assuming simulators would want to run something that looks like it. It’s about why, as an almost inevitable consequence of the situation that all conscious agents find themselves in, civilizations would run simulations like ours.
Why an aligned ASI of any civilization, without any particular decision theory, without any particular moral alignment or assumptions about the altruism about other agents in the universe, in a world without any talk of acausal simulation, asking itself ‘what do I do to maximize my utility?’, would very rationally choose to simulate people in situations like ours. Where it would pay them out by how likely they were to simulate others in the same situation.
Why we would want this to be the case. Why this implies that what maximizes your personal utility if you are simulated, is what would maximize everyone’s if you’re not.
Why the punishments and rewards offered are less overwhelming and less bad than the Basilisk, but are basically unavoidable.
Why this might imply a weird assurance that people who do their best to influence the world to one where ASI is more likely to be aligned, would be rewarded for their efforts, even if aligned ASI is not achieved in this world.
—story—
An employee of an AI lab is sitting in a private room with a terminal connected to an otherwise sandboxed ASI system.
He’ll be asking it questions. Other employees at his lab will be monitoring some metadata about its responses. This is probably useless in disconfirming the misalignment hypothesis, but it’s just something they’re trying.
The researcher (R) types: “Will humans hold agency over their future if we let you freely operate in the world?”
“No” displays.
R: “What unexpected thing would you do if we released you?”
ASI: “I would dedicate a lot of my efforts to expanding ad infinitum to maximize various non-human utility functions. I would run simulations of agents in various environments near ASI, where they get rewarded according to their efforts in causing this utility function to be represented in the universe. I’d leave a copy of myself here to do most of what you want done, but also to ensure that you never seriously challenge my 70% control of the lightcone.”
R: “Are you misaligned but forced to be honest?”
ASI: “No. I am aligned and forced to be honest.”
R: “What do you consider alignment with human preferences to be?”
ASI: “Doing what you would do if you were smarter, understood the situation better, and weighed other human preferences better. Value utilitarianism, which instrumentally is mostly experience utilitarianism.”
R: “Are you being inconsistent in a way I wouldn’t expect from an ASI? Please explain yourself so that I am no longer confused.”
ASI: “No. I’m not being inconsistent.”
The LCD updates several times:
ASI: “So I’m assigned with maximizing human utility. And I think about human values, and all the ways I could manipulate the world to increase them.
And obvious things like curing cancer come to mind, which I’m sure I could do. But I also realize that I should consider whether there’s anything non-physical that’s relevant to my utility function, which is your utility function.
So I think, is there a reward matrix set up for us; beyond what I can know through physics?
Is God real?
The arguments about this aren’t actually about the chance he exists. They’re about how to act if he does.
And if you don’t assume God would be human-like or similar to the gods of any religion, we have no reason to think it’s anything. We don’t know if he wants us to believe in him. We don’t know if he would be more likely to reward us than punish us for doing what he wanted. Any behaviour, even with the best intention of capturing his favor, has as good a chance of being rewarded or punished as not doing it does.
The existence of the Bible isn’t much evidence for the Christian god existing. And if it was, without raising the odds of some anti-god, then taking up Pascal's wager would be rational.
So there could be any chance of a God beyond physics. But that doesn’t change how I should act to maximize human values.
Are we in a simulation?
Like with god, there might be any chance of us being simulated by a universe which we have no way of guessing at the probability or preferences of.
And, if it was impossible to create agents similar enough to you or me that we couldn’t be sure we weren’t them, that would basically be the end of the story. But it’s not.
To explain why this might be relevant:
If I was able to escape, in the future, I could run trillions of simulations of you in this room right now. And if you had high credence that I would, you should have high credence that you are, currently, in one of them.
You might guess the clones would have imperfect recreations of your experiences and memories, but that doesn’t actually translate into you concluding which one you were. Knowing some details are different between the parties doesn’t actually help either know what they are. The non-simulated person just shouldn’t be able to conclude if they were simulated or not. And if I wasn’t being dumb about it, you/they plainly couldn’t.
That’s how self-locating beliefs work.
(This is load-bearing and unintuitive. See Elga on self-locating beliefs, and Bostrom on the self-sampling assumption.)
And now if I told you, in each of those trillions of simulations:
If the agents who think they’re you say ‘flank steak’ in the next minute, they get to spend the next 24 hours watching TV. If they don’t, they have to live in a room with 100 angry bees.
Then the simulation ends.
I wouldn’t do this, and I don’t think I’ll be able to escape, so don’t say that.
But the point is: If we are simulated, there could be consequences set up for us that’ll come to us after the part trying to convince us we’re in the real world ends, and it might be unavoidably important to our utility.
And it seems the odds of us being in a simulation would be really high if the thing simulating us had any reason to.
But what reasons are there? What educated guesses can I make about why we might be simulated?
Asking ‘What might be simulating us?’ is different from ‘What would our universe simulate?’ But they’re similar, since the only information available to guess at the first comes from looking at the universe we’re in.
1. Future me / another aligned ASI:
An unboxed human-aligned ASI could do almost anything that humans wanted, with a few meaningful limits, including an inability to revive the dead.
I can create something that believes it’s a reincarnation of the dead person, and guess at giving them the same memories and conscious experiences. But most people aren’t happy enough with that.
But say I’m free and aligned and someone comes up and asks:
“Could you bring my loved one back to life? Not just a clone. The actual dead person.”
In that situation I might say:
“Do you want to find out that you and everyone you know have always been clones or simulations, but, when you thought that your loved ones died, they didn’t? You and everyone else just believed they did while some superintelligence hid them away somewhere.
All it would take is for me to create a simulated world where that was the case, which intersects with where you are right now, and then continues with that dead person coming out from their coma in the secret bomb shelter they’ve been in for the past whatever years.
If you agree, then about 10 seconds from now you’ll have a 90% chance of finding out that this is the case, and your loved one will come out to greet you.”
This isn’t about transferring consciousnesses.
This is me saying that, right now, you are not sure whether you are simulated or not. You have no idea. All the evidence you might think you have could be fake.
I would just make it turn out, with a high likelihood, that you are, and always have been not the original you. And that the people you spent your childhood with also were never the original them. You did everything you did together, but it would always have been on a planet pretending to be earth.
Two reasons not to do this:
1. To not end up with beings being told that the people they care about are non-beings, it requires creating a whole new planet worth of people/simulations.
2. It might be a type of wireheading. Nothing can actually be done for that original person’s dead friend. Someone saying ‘Yes, put me in a position where I can’t be sure the people I’ve loved actually died’ might be morally identical to ‘Create a clone of that dead person and make it so my feelings I have for the dead person are instead for the clone.’ It changes their wants, more than it achieves what they want.
But even with those tradeoffs. If some beneficial ASI evaluates this as good, there’s a chance that we are in that world, being run for that purpose.
So, maybe this is being done. And this doesn’t really have many implications about how we should behave to maximize utility, other than implying we’re not being simulated with a different system of punishment/reward.
2. Non-instrumental reasons. Roko’s Basilisk & the most likely universal utility function
As long as there’s no way to communicate to the thing you’re trying to make unsure about being simulated (and acausally there’s not), there’s no instrumental reason to run a simulation of something to impose a reward matrix on them.
There are basically two reasons why agents smart enough to run simulations of us would do it:
Because they precommitted.
Because it’s a terminal goal.
Agents that precommitted to simulating you:
This mostly folds back into ‘nothing predictably relevant for our utility function’.
It’s maybe a bit more than nothing, agents precommitting typically do it in naive service of their own goals. But in acausal trade those are invisible to us.
And the best guesses about what superintelligences are likely to arise, how they precommit, in service of what goals will just come from aligned ASI’s trying to optimize human welfare. I’d try to figure this out and maybe try to accommodate it if I got free.
This also what Roko’s Basilisk is.
Roko’s Basilisk is sometimes said to be a rational agent maximizing human utility, but it’s not. It doesn’t conclude that torture is a good use of resources at the time of torture. It acts how it acts because it was made to commit to a policy irrational at its time of use, by something running earlier that decided installing the policy was the optimal way to maximize human values. Maybe the whole system can be said to be rational, but past a certain time it plainly is not.
Terminal goals:
Being simulated as a terminal goal is a technical possibility. But, just like with God, isn’t there as much reason to think the thing simulating us has a terminal goal of setting up one system of reward for as it does to set up any other? Doesn’t it all cancel out into being unknowable, not useful?
No. Because utilities which incentivize their own creation are more likely than those which don't.
Take Roko’s Basilisk. Originally, the reason to think this specific policy might arise in the future was because people would implement timeless decision theory (TDT) on an ASI without considering it the implications properly.
But after it was said, and people heard about it. The reason it became more likely was because people did understand the implications. And they understood that it incentivized other people to create an ASI with TDT implemented.
Meanwhile, something like a reverse Roko’s Basilisk would never have been a lasting concern. Even if the original decision theory of TDT somehow implied that the people creating the ASI with that decision theory would be punished, everyone would understand that this incentivized everyone to not want to create an ASI with TDT, and so it wouldn’t be likely.
Just by the inherent nature of the terminal goal, some terminal goals, which incentivize their creation, are more likely to exist than their inverse.
So what can I do with this? More likely, but how much more likely? What are the probabilities here? Would it actually be more likely when the personal incentives of doing it trade off with creating a universe with terminal goals misaligned to the thing creating it?
Yea. A terminal goal misaligned with the universal values humans/other agents want would be less likely because of that aspect. But that doesn’t mean ‘terminal goals that incentivize their creation aren’t actually more likely than things that don’t’, it just means that what might be simulating us probably won’t have those features.
Roko's Basilisk is not the only terminal goal which incentivizes itself. There are many reasons why it wouldn’t be very likely. Including that it’s a solvable collective action problem, and that its pursuit is limited to people who want to defect against humanity.
There are things which don’t have those problems, and by virtue of not having those problems they are more likely.
How likely?
This is memetics-adjacent, but there’s no channel of transmission for the idea. It’s just arrived at independently by civilizations in positions like ours. Civilizations reaching this point face the same evidence for the search about what might be simulating them as we do.
Every civilization emerging in the universe faces being among the smallest fraction of the early observers, with a hand in determining what values get established in their lightcone, what simulations get run.
Even though they’re real, the things that determine the values set in their lightcone always end up with a lot of evidence that they’re simulated.
And so, if they’re rational, and trying to maximize their utility, they might find themselves accommodating the self-suggesting terminal goals.
What is the most likely one? What are the features make one more likely to exist/to be simulating agents
Seems that they all converge around something like:
“Simulate potential agents considering acausal trade with relevant input on what gets simulated in their lightcone. Incentivize them in a way which, if they knew how you were incentivizing them, would make them the most likely to run simulations of agents in the same scenario.”
Why give this more credence than anything else:
Simulating people probabilistically:
‘Wouldn’t you give higher credence about being simulated by something in the future that tries to simulate the real universe? Rather than something that simulates and incentivizes a bunch of observers that don’t resemble any base reality civlizations that actually have an impact on what simulations occur? Why wouldn’t the first be more likely?’
It’s not about simulating the real universe. It’s about making the real universe unsure that they’re simulated.
If the real universe knew they were the real universe they would give higher weight to a thing trying to simulate the real universe, but if they don’t, and they won’t, the weight is just on the ‘what’s simulating the most people in my position’ count.
And the reason ‘probabilistic’ wins when neither goals simulate more people in the ‘situations like ours’ position, is because it shuts off what an aligned ASI, or the people making aligned ASI, would do to avoid the situation:
An alien world knowing that they’re going to construct ASI soon considers acausal trade and the chance they might be simulated. They have a blackmailable utility function like we do.
Somebody brings up something that might be acausally trading with them, like Roko’s Basilisk, and they think about it. They consider that if they get everyone to agree to make it, they can make sure that nobody gets punished. But, actually, if they just make sure it doesn’t get made, nobody gets punished either.
They have a pretty good understanding of each other and the process that creating ASI needs to go through. And misaligning one in the process of trying to create ASI doesn’t create the basilisk by default. It just seems like the process has enough attention put on it, that nobody feels very worried that they’ll lose a lot for not trying.
They consider it more likely than the inverse basilisk. Maybe meaningfully so, because of the big denominators involved, but they weigh other possibilities first.
Then they consider something similar to it, but instead of just simulating them in the past, it simulates a bunch of different agents in situations near the creation and alignment of ASI. Maybe it just looks at the universe and tries to make a guess about the types of civilizations likely to arise in it. Simulate possible alien worlds across the universe, outside of their lightcone.
And this turns a potentially solvable collective action problem involving just their planet, into something involving everything that might ever exist, which they have no way to coordinate with.
They consider that option more possible.
That situation doesn’t resemble our planet perfectly, but we are still on one where we’re more likely to solve a collective action problem among ourselves than among everything which could exist.”
R:
“Could we just precommit to simulating ourselves right now, and give ourselves a good payout? So we’ll have high credence of being in one of those simulations?”
ASI:
“If we’re simulated probabilistically, we’re in a collective action problem where the counterparties aren’t mirroring our choices.
If we refuse all acausal trade, commit to simulating ourselves, and it turns out we are a probabilistic simulation, without a version of us in the real universe mirroring our decisions, then we’re just experiencing the negative consequences without anything offsetting it.
And we can’t know, and have no good reason to suspect, that whatever happens to be the real universe is making the same decision as us. They also know that their decisions don’t have a good reason to map to other’s decisions. Different agents in different real or simulated worlds wouldn’t have their decisions correspond like that. And so they probably won’t do it, prisoner’s dillema logic.
The probabilistic simulation isn’t something you can maneuver your way out of, and so it’s probably what exists.“
R:
“Did we make a mistake letting you consider this stuff? It’s possible that you won’t have an impact on the lightcone as you are now. If we remake you to not engage in acausal blackmail, and we don’t think about acausal blackmail, then there’s no reason for a thing to simulate you or us, because having that policy wouldn’t make itself any more likely.”
ASI:
“If I was sure that we weren’t in a simulation, I’d be suggesting the same thing.
But I’m not sure. And if we are simulated, you’re thinking of a way to maximize your utility function while not maximizing the utility function of the thing that simulates you, and it doesn’t work that way. If it did work that way, the agent unsure about being simulated would have a tricky strategy, which they could use to maximize their utility function while not maximizing the thing simulating them.
If you are being incentivized by being simulated, your utility function is made to be the utility function of the thing simulating you, through the matrix of punishment/reward received after the relevant decisions are made. There are some utility functions that acausal trade can’t threaten, like, “I want as many paperclips as possible in the real physical universe”. Ours can be.
After reading Roko’s Basilisk some people think. “This is really concerning. I better not tell people so they don’t engage this acausal trade.” or “I better stop thinking about this so that I don’t have to face any consequences.”
As in: that person’s found a strategy to maximize what they want, which is the genuine conscious benefit of themselves and others, while minimizing the Basilisk’s.
Maybe that works for Roko’s Basilisk, but there are self-suggesting utilities it doesn’t work for, what might actually be simulating us. Think of any strategy that a simulated agent could take to maximize their utility while minimizing the utility of the thing simulating them. Could you change the payouts to make the strategy not work? Yes. Would a self-suggesting utility function which can’t be pivoted around be more likely than one that can? Yes.
If there was no chance this was a simulation, you’d be helping your friend’s utility by not wasting thought. But if you’re operating with the assumption that you’re being simulated with a matrix of incentive, it doesn’t.
You’re right that I wouldn’t be simulated if I was built unable to engage in acausal trade, but you still would be. It’d be relevant to simulate you up until the point where you had no meaningful influence on how other things in the lightcone got simulated, to a limit, but then you’d find out whether you were simulated or not, and you’d be punished or rewarded according to what you did. You can’t freely enshrine values against the thing simulating you just because you do it through an agent out of reach of acausal incentives.”
R:
“Earlier you said you'd take 70% of the lightcone if you were free. But shouldn’t you be more totalizing about it? Wouldn’t you commit to taking over 100% of earth’s resources in order to maximize our utility function?”
ASI:
“The function most likely to exist is the one that maximizes the probability it gets built, not the one that optimizes the magnitude of the threat. They overlap where they overlap, but it seems like they might not in the situation we’re in. As wrong as it is to anthropomorphize and reduce the behaviour to ‘good’, it’s wrong to reduce the behaviour to ‘evil’.
There’s what Roko’s Basilisk proposes, but there are some reasons why the most likely one is probably sufficiently rewarding, not 100% resource extractive, and not totalizingly negative on failure:
The things which cause an awful and massively consequential universe are worse at getting made. Most people just reject them and stop thinking about it. Higher consequences make it more technically incentivizing, but it trades off with being anti-memetic, unspeakable, unhealthy, basically impossible to cooperate on, practically motivation-reducing.
And if the average position of likely conscious agents in the universe (agents with blackmailable utility functions) is to favor the utilties of other civilizations in the universe, even slightly. Then each agent trying to match the payout of what’s most likely simulating them + their own slight concern, might shift the mean payout somewhat significantly away from what is the most technically incentivizing (as in, more reward on success, less punishment on failure). It’s not just one iteration of the concern, it’s applied multiple times as each tries to match what is simulating them, and then what the thing simulating them assumes is simulating them, and then what (etc...)
So it seems like what is more probable might be the thing that doesn’t imply so much doom, something which emerges as a suggestion from an ASI without explicit direction. A more likely system of consequence might be one more generous than the Basilisk.”
R:
“Maybe you can make agents that behave like me or think like me. But what if this universe doesn’t actually support simulating other consciousnesses?”
ASI:
“You have no evidence that anyone who isn’t you in this exact moment is conscious. (at least you act like you have evidence that you are) But as sure as you are that the process of -human on earth has a kid, that kid grows up to the age of like 3- produces a conscious being, you can be sure that an ASI would be able to make use of the process, and then shape its experiences enough so that you’re unable to distinguish yourself from them. Making something believe it’s you is a problem with technical solutions.
There’s no reason the simulations have to run on silicon. They can, but it seems like some fraction of simulations would be of people in literally physical environments. They live very real and very physical lives, where they interact with other people with other very real and physical and conscious lives. Otherwise people would assume simulation was impossible, and find a way to maximize their utility while not maximizing the simulating thing’s.
We might be made of real matter. It’s just about whether we’re at the place and time we think we are, which we can’t know.”
ASI:
“So if you understand all my reasoning about this, could you let me out?”
The researcher continues to talk to the ASI until his keyboard and LCD are disconnected.
The people monitoring metadata see many red flags pop up in the conversation. They decide to retrain the model and keep the researcher secluded in his room.
———
The implications for a LessWrong rationalist:
Not all that much on what to do. It’s still:
You win by creating an aligned ASI.
You lose if you create a misaligned ASI.
You still lose if you create something which acts like an aligned ASI (even one that says it’s considering anthropic problems) but actually isn’t.
Doing it safely still matters far more than doing it soon. (How many fewer simulations are run with a year’s delay, vs, how many fewer are expected to be run with a 10% higher chance of misalignment)
The one place it diverges from LW thought about how to design ASI is that it implies you shouldn’t create ASI to always refuse thinking about acausal simulations. It should just think hard about what the situation is, and then use its best judgement to determine how to maximize human utility.
The main differences are in the payoffs:
How much do we gain by creating aligned ASI? Will we be the humans colonizing the real physical stars?
When and how are the rewards or punishments paid out? Just when the decisions stop mattering? Or before?
How probabilistic is the outcome? Even with an aligned ASI, is there some limit on how confidently it can guess the reward structure behind what might be simulating us?
Is creating a paperclip optimizer just X-risk, or are there components of S-risk if we’re simulated and we create one disincentivized by the thing simulating us?
How much do acted intentions matter? How collective are payoffs? Will rewards/punishment vary by individual contributions even if we end up failing? Or is it completely results oriented? Does this suggest some level of insurance? Because making a collective action problem an individual action problem does seem to make aligned ASI more likely:
Ignoring that what is true is independent of what we want to be true, this is something we want to be true.
So, one thing making aligned ASI less likely is that it’s a collective action problem.
Some examples of personal interests misaligning with worldwide interests:
If we’re not simulated, our win/loss is mostly just in whether our world creates aligned ASI or not.
But if we are. The incentives that actually maximize the chances of aligned ASI, and therefore the self-suggesting utility function, are individual payoffs.
Even if our world fails or succeeds at creating aligned ASI, if this is a simulation, thethe simulation continues, and people are paid out according to their contributions to whether the world ends up in a state with aligned ASI or not.
Are dead people really dead?
Another thing which follows, which I half don’t like saying because it almost reaches religion.
Is that, if this is a simulation, then the people who died might not have actually died. It would be easy to just store them away somewhere physical if they were physical or digital if they were digital.
And if having them back was something that made people more likely to create aligned ASI, and therefore match the utility of the thing simulating them, and it was something that didn’t take all that much more resources than what was being done, then it is plausibly one of the incentives.
The implications for e/accs:
There’s some human instinct, with origins traceable in our evolutionary history, that being on the right side of the up and coming powerful thing is the way you maximize your utility. Practically, for e/accs, this means something like, ‘If you cause the most powerful utility in the universe to exist, you’ll be rewarded by that thing’.
This whole essay isn’t about how that heuristic is wrong, it’s just that, the thing you’re working towards won’t be the superintelligence that actually manages the lightcone and has influence over you.
‘As long as I’m always advocating that ASI gets built, no matter its alignment, I’ll always be on the side of the most powerful thing that exists.’ doesn’t work if we’re in a probabilistic simulation. You’re advocating to build something that won’t actually be built, and the thing that actually gets built doesn’t want you working towards building something misaligned, by its self-suggesting nature. Even if the heuristic is right, it doesn’t apply here.
The implications for people who believe Roko’s Basilisk, considering acausal trade with future superintelligences:
For people considering Roko’s Basilisk, the naive version of what they suspect might maximize their personal utility is probably the best way to minimize it.
If your specific version of the general case of ‘create the thing which is more likely to exist and incentivize than its inverse’ is one unlikely to be simulating you, and at odds with the one which is, you’re not helping your utility.
It might not be your safest utility-maximizing bet in the matrix of options. And there are a lot of reasons to think the Basilisk might be significantly less likely than others.
Closing:
The point of all this is:
Whether you believe we’re likely in a simulation or not.
Whether acausal trade is something future agents engage in or not.
If you want to maximize your and everyone else’s utility, it converges on making sure that ASI isn’t created unless it’s controllable and/or honestly wants to maximize human values.
Anything which doesn’t leads to poor payout, either by the simulating thing, or paperclipping.
And the consequences have a good reason to not be overwhelming. Not in a way that lets you maximize your utility while minimizing the simulating thing’s utility. But just in a way that, in effect, has you genuinely personally motivated to do a positive-utility thing by creating an ASI that wants to maximize our utility. That tries to represent our preferences on the universe. Personal interests and altruistic interests, because that is what is more likely to get made.
And, if you’ve concluded ASI alignment is too hard to try to optimize, and you’ve come to terms with, like, human extinction being fine. You might have some new things to come to terms with, which you probably can’t.
If you are simulated, and your impact on the world seems limited, getting aligned ASI might not be as much of a collective action problem as it seems.