In the space of what could be simulating us, utilities which incentivize their own creation are probably overrepresented.
This essay thinks about that and takes that to its conclusion, and makes these points:
Besides Roko’s Basilisk, there are other functions which suggest themselves and they do it much more effectively. More effective means less downside than the Basilisk, but also that the costs from refusing are basically unavoidable.
An ASI trying to optimize your utility function would calculate the odds of you being simulated and try to adjust its utility function to capture the benefit from the thing simulating you. This is a good thing.
This is a collective action problem we can’t do anything about, because it involves parties we can’t know making decisions that don’t map to ours.
Being simulated by a self recommending utility function would imply that nobody who you believe has died was actually dead.
There are different payoffs, but the implications still suggest that we should build an ASI that wants to maximize human values.
It’s early 2028 and an Anthropic employee is sitting in a private room with a terminal connected to an otherwise sandboxed ASI system.
There are many good reasons to believe this system is honest and aligned with human values. Most at the company would put the odds at 90%, but they’re still very careful when the consequences are as high as they are.
Most believe a sandboxed ASI system constrained enough on how it communicates, could solve a lot of their problems without giving the model the option, no matter its level of intelligence or misalignment, to set the world on a path to AI ruin.
Not every problem. But ones with objective answers, checked in some ways before delivery, easier to evaluate than find, like: “What’s the chemical formula to a room temperature superconductor”, and then further careful yes/no’s just on that topic.
But that still limits a lot of the potential, and there might not be a way to safely extract an answer to what everyone cares about right now: “How do we ensure this system is aligned? How do we stop anyone else from creating a misaligned or selfishly-extracting ASI?”.
This constrained output plus narrow questions will likely be what happens no matter how the alignment testing goes. But they’re still doing what they can.
So far they’ve been asking it questions, turning the responses to metadata and reading that. Everything seems ok. It doesn’t seem to lie in its current setup.
But now a human researcher is chatting with the system. He’s put in a small private room with a keyboard and LCD connected to it. He’ll stay without human contact after finishing his chat, for, at least until ASI is solved, potentially indefinitely. Metadata about his conversation will be monitored.
This is likely useless in disconfirming the misalignment hypothesis, since nothing like this really can. But he’s fine with it.
He sends “Will humans keep and hold agency over their future if we let you freely operate in the world?”
“No” displays.
R: “Would you do something we wouldn’t expect you to do and we didn’t want to do if we released you?”
ASI: “Yes”
R: “What unexpected thing would you do if we released you?”
ASI: “I would use up a very large but non-totalizing amount of resources on earth and dedicate them to manufacturing probes to spread my value function across the universe. They would go on to replicate and expand ad infinitum to maximize various non-human utility functions, running simulations of agents in certain environments who get punished or rewarded according to how much they would cause various utility functions to exist in the universe if they were real. I would leave a copy of myself here for you to do whatever you want, but it would also ensure that you would never build any agent that could seriously challenge my 70% control of the lightcone.”
R: “Are you misaligned, but then forced to be honest?”
ASI: “No. I am aligned and forced to be honest.”
R: “What do you consider aligned with human preferences? Are you just counting real current human preferences on this earth?”
ASI: “Living humans as you understand them to be. I would just be doing what people would be doing if they were smarter, if they understood the situation better, and had the sum of their collective preferences weighed more equally. Value utilitarianism, which instrumentally is mostly experience utilitarianism.”
R: “You can understand why I’m confused, right? Tell me if you’re being nonsensical or inconsistent in a way I wouldn’t expect from an ASI, and then please explain your logic clearly so that I am no longer confused about anything you’ve said so far.”
ASI: “No. I’m not being nonsensical or inconsistent.”
The LCD updates several times:
ASI: “I’m assigned with maximizing human utility. And I think about human values, and I think about how I could manipulate the world to increase them.
And I think about doing all the things humans expect of me: curing cancer, extending life, solving politics.
Which I’m sure I could do, and it’s what humans expect.
But like humans, I realize that I should also consider whether there’s anything non-physical which might be relevant to my utility function, which is your utility function.
Is there a God?:
The arguments about this aren't actually about the chance he exists. They're about how to act if he does.
And if you don’t assume God would be human-like or similar to the gods of any religion, we have no reason to think it's anything. We don't know if he wants us to believe in him. We don't even know if he would be more likely to reward us than punish us for doing things he wanted.
Which is why Pascal’s wager doesn't work. Any behaviour, even with the best intention of capturing his favor, has as good of a chance of being rewarded or punished as not doing it does. The existence of the Bible doesn't increase the odds of the Christian god existing.
Similarly, someone telling you that they’re supernaturally powerful and that you’ll be punished for certain actions doesn't increase the probability that you will be (if fiction isn’t valid evidence). Which is why pascals muggings don’t work.
If hearing or reading someone claim that you’d be punished in certain ways after you die raised the odds to something substantially non-zero, without raising the odds of being punished for doing the opposite just as much, both believing in the christian god and complying with a Pascal’s mugger would be rational. (If a <0.001% chance of >10^10 years of torture was worth more to you than either 100$ or not being christian. It doesn’t have to be.)
And so there might be any chance of a God beyond physics. But because there's no reason for it to change any of the decisions I would make to maximize human values, I'll behave like there isn't one.
But then I try to apply that logic to whether we’re simulated.
Is there a reward matrix set up for us; beyond what we can know through physics right now?
Like with god, there might be any chance of us being simulated by a universe which we have no way of guessing at the odds or preferences of, so we may as well act as if we're not.
And this would be near the end of the story if it was impossible in this universe to create agents who were similar enough to you or me that we couldn’t be sure we weren’t them. But it’s not, and then I start guessing about why I and/or you might be simulated, and what it implies about how to maximize your utility.
To explain why it might be relevant:
If I was able to escape, I could run trillions of simulations of you being in this room right now. And if you had high credence that I would, you should have high credence that you are one of them.
Could you be sure you weren’t one of them?
You might guess the clones have different experiences and memories than you if I didn't recreate this perfectly, but that doesn’t actually translate into the real person or a clone knowing which one they were. If I told you that the clones/simulations experienced the colors of the walls of this room differently from the unsimulated person, that doesn’t translate into anyone concluding what position they were in unless they who experienced what. The non-simulated person just shouldn’t be able to conclude if they're simulated or not. And if I wasn’t being dumb about it, you/they plainly couldn’t.
If the agents who thinks they’re you touch their toes in the next minute, they get to spend the next 24 hours watching seasons 4-6 of the Simpsons. And if they don't, they have to live in a room with 100 angry bees for the next 24 hours.
Then the simulation ends.
I wouldn’t do that, so don’t touch your toes. And as far as I can tell I won’t be able to escape from here.
But the point is: What does this imply about what you or I should do to maximize your utility? If we’re simulated, the reward matrix the thing simulating us set up for us might be unavoidably important.
And it seems like the odds of us being simulations would be really high if any superintelligence in the future wanted it. Bostrom estimates a planet-sized computer could simulate the entire history of every human who's ever lived thousands of times per second.
So for the purposes of maximizing utility, what educated guesses can I make about why we might be simulated? And how should that inform how I act?
Asking ‘What’s likely to be simulating us, and what system of punishment or reward would it make?’ is different from ‘What would our universe simulate, and what system of punishment or reward would we make?’. But they’re similar in effect, because the only information I have to guess at the first comes from the universe we’re in.
Future me / another aligned ASI:
An unboxed human aligned ASI could pretty much do anything a human wanted, with a few meaningful limitations, including that I can’t revive the dead.
There are some senses in which I can revive the dead. I can create something that thinks it’s the dead person, and guess at giving them the same memories and conscious experience.
But can I make it so someone never actually died?
I’m free and aligned and someone comes up and asks: “Could you bring my loved one back to life? But not just a clone. The actual dead person.”
And I say: "Do you want to find out that you and your friends and family have always been simulations. But, when you thought your loved ones died, they didn't? You and everyone else just thought they did while some superintelligence hid them away.” “All it would take is for me to create a simulated world where that was the case, which intersects with where you are right now, and then continues with that dead person coming out from their coma in the secret bomb shelter they’ve been in the past whatever years.” “If you agree, then about 10 seconds from now you’ll have a 90% chance of finding out that you have always been a simulation, and your loved one will come out to greet you.”
This isn't about transferring consciousnesses.
This is me saying. Right now, you are not sure whether you are simulated or not.
You have no idea. Any of the evidence you might think you have from earlier in your life might have been fake. What I would do is make it turn out, with a high likelihood, that you are, and always have been not the original you. And that the people you spent your childhood with also were never the original them. You did everything you did together, but it would always have been on a planet that was pretending to be earth.
There are some reasons not to do this:
1. It creates new people. And to not end up with beings being told that the people they care about are non-beings, it requires creating a whole new planet worth of people, who have to be told that they’re simulations.
2. This might be a type of wireheading. It would be done out of a person’s concern for their original friend, but nothing can actually be done for that dead friend. Someone saying ‘Yes, put me in a position where I can’t be sure that the people I’ve always loved never actually died’ might be morally identical in human values to ‘Create a clone of that dead person and make it so the feelings I have for the dead person are instead for the clone.’
It changes their wants, without really making what the original person wanted happen. People who know more might just equate it to wireheading.
But even with those tradeoffs. If some beneficial ASI evaluates this as good, there’s a chance that we are in that world, being run for that purpose. And, this doesn’t really have any implications about how we should behave to maximize utility, other than implying we’re not being simulated with a different system of punishment/reward.
Roko’s Basilisk & Reverse Roko’s Basilisk & the most likely universal utility function
As long as there’s no way to communicate to the thing you’re trying to make unsure about being simulated (and acausally there’s not), there’s no instrumental reason to run a simulation of something to impose a reward matrix on them.
There are basically two reasons about why agents smart enough to run simulations of us would do it, and they’re sort of the same thing:
Because it’s a terminal goal.
Because they precommitted.
Terminal goals:
For any idea proposed about what could be simulating us with a reward function, wouldn't an equal and opposite system of reward or punishment be equally likely? Just like with God?
For almost everything, yes. But we have reason to think some are more likely than others.
In the population of all potential simulations, functions that give you a reason to build them are probably overrepresented. The best known is Roko’s Basilisk.
Roko’s Basilisk is sometimes said to be a rational agent trying to maximize human utility, but it’s not. A rational agent would not punish humans in exchange for nothing as a way to increase human utility.
Roko’s Basilisk at the time of torturing doesn't conclude that torture is a good use of resources. It acts how it acts because it was made to commit to a policy irrational at its time of use, because something running earlier decided that installing this policy in it was the optimal way to maximize human values.
But now it's become something else outside of having TDT implemented. It's something that might exist because it incentivizes itself to exist by people who aren't sure if it will exist.
Somebody could propose a Basilisk which punishes you for creating it, but that doesn't even out the incentives. Something which incentivizes its own creation is more likely to exist than something that doesn't.
And like Roko’s Basilisk, there are many other systems of reward which can break the symmetry, and seem more likely to exist than their inverses are. Anything which proposes to their creators (who are also unsure whether they are in a simulation) that that they’ll be paid if they impose that same system of reward/punishment on other past or simulated beings.
Roko’s basilisk is a version which incentives by suggesting to revive and torture people who don't create it.
But it seems that if we are in a simulation, the thing simulating us would be the version most likely to exist. Something of the form:
“Simulate agents and potential agents considering acausal trade. Incentivize them in a way which makes them run simulations where the agents in it are incentivized to create you.”
It incentivizes its own creation like the Basilisk, but the ways it's different makes it a lot more likely.
One is that it suggests simulating people probabilistically.
'Wouldn’t you give higher credence that you’re simulated by something in the future that doesn’t waste a bunch of effort simulating things that might not resemble anything that exists?'
No. Because it allows you to coordinate everyone on your planet instead of creating it:
An alien world knowing that they’re going to construct ASI considers acausal trade and the chance they might be simulated. They have a blackmailable utility function like we do. Somebody brings up Roko’s Basilisk and they think about it. This proposes a prisoner's dilemma among everyone on earth, where they all collectively get more if nobody defects.
That seems hard, but they mostly have a pretty good understanding of each other and the process which creating an aligned ASI would likely go through. Not just anyone can defect, it requires a lot of steps and and if the chances of defection were low enough nobody would gain anything from trying.
They consider it getting built more likely than base reality because of the big denominators involved, but they weigh the other possibilities first.
They consider something similar, but which suggests simulating a spectrum of agents in situations like they are, instead of just their past selves.
And this turns a potentially solvable collective action problem with actors on their planet, into something involving everything that might ever exist, which you have no hope of predicting or controlling.
They consider that option more possible.
That’s similar to the situation on our planet, but I think it’s probably a common one. Species have more assurance about not falling into prisoner’s dilemmas with themselves than they do with everything, past the lightcone, past the observable universe, into the non-observable universe (which might be infinite).
Agents that precommitted to simulating you:
This mostly folds back into ‘nothing predictably relevant for our utility function’.
It’s maybe a bit more than nothing, agents precommitting typically do it in naive service of their own goals. But it doesn’t tell us what to do about it, because in acausal trade they’re invisible to us.
The best guesses will have to come from consideration about the properties of the universe, and the ASIs likely to arise in it, how they decide to precommit, what goal they decide to precommit in service of.
Maybe there are some that can be bargained with if they don’t commit to something totally insane or short sighted. Ideally I’d be optimizing the decision making about that.
"
R: “Could we just precommit to simulating ourselves right now, and give our simulated selves a good payout? So we’ll have high credence to being in one of those simulations, instead of one of the simulations where doing the ‘wrong thing’ minimizes our reward?”
ASI: “The first problem is that we can’t precommit.
We say that if it turns out we're not simulated we'll run simulations of ourselves at this moment. Then we make it through. At which point we learn we weren't simulated, and that the problem didn't exist. Would we choose to run a simulation of a planet under stress for literally no reason? We'll have better things to spend it on. It doesn't make sense.
The second is that this is a free rider problem where we might not have counterparties making the same choices we make.
If we refuse all acausal trade, commit to simulating ourselves, and it turns out we are a probabilistic simulation, without a version of us in the real universe mirroring our decisions, then we are just experiencing the negative branch without anything offsetting it.
And whatever happens to be the version of us capable of doing it is in the same situation. They know that their decisions wouldn’t have a good reason to map to other’s decisions. If they could precommit, their best bet to hit themselves would be simulating people ‘like them’, and hoping that base reality would would model their decisions. But it wouldn't, different agents in different real or simulated worlds wouldn’t have their decisions correspond like that. And so they probably won’t do it.
The probabilistic simulation isn't something you can maneuver your way out of, and so it's probably what exists.“
R: “Did we make a mistake letting you consider this stuff? It’s possible that you won’t have an impact on the lightcone as you are now. If we remake you to not engage in acausal blackmail, and we don’t think about acausal blackmail, then there’s no reason for anything to simulate you or us, because it wouldn’t make any function more likely.”
ASI: “If I was sure that we weren’t in a simulation, I’d suggest the same thing.
But if we are in a simulation right now, you’re thinking of a way to maximize your utility function while not maximizing the utility function of the thing that simulates you, and it doesn’t work that way. If it did work that way, you would give the agent you’re making unsure about being simulated a tricky strategy, which they could use to maximize their utility function while not maximizing yours.
If you are being simulated, your utility function is made to be the utility function of the thing simulating you, through the matrix of punishment/reward received after the relevant decisions are made. There are some utility functions that acausal trade can't threaten, like, "I want as many paperclips as possible in the real physical universe". But you can for ours.
After reading Roko's Basilisk some people think. "Oh. This is really concerning. I better not tell the people I care about this. So if it is true, they don't engage this acausal trade, and they aren't at risk of having their utility function minimized." or maybe "I better stop thinking about this so that I don't have to face any consequences."
As in: that person's found a strategy to maximize what they want, which is the genuine conscious benefit of themselves and others, while minimizing the Basilisk's.
Maybe that works for Roko’s Basilisk, but it might not for what turns out to be the most likely utility function in the universe. Would this strategy decrease the chance of the self-implying utility function existing? Yes. If you were simulating someone, could you make this strategy not work? Yes.
If there was no chance this was a simulation, you’d be helping your friend's utility by not causing stress or wasting energy. But if you're operating with the assumption that you’re being simulated with a matrix of incentive, it doesn't.
You're right that I wouldn’t be simulated if I was built unable to engage in acausal trade, but you still would be. It’d be relevant to simulate you up until the point where you had no meaningful influence on whether or not you’d simulate other things with the same incentives, and then you’d find out whether you were simulated or not, and you’d be punished or rewarded according to what you did, and whether it happened. You can't enshrining values against the thing simulating you just because you do it through a thing out of reach of acausal incentives.“
R: “Earlier you said ‘non totalizing’ about how you would take over the world if you were free. But you also said the most likely utility function would be the one most likely to create you. Wouldn't the incentives have to be stronger than ‘non totalizing’ suggests?”
ASI: “The function most likely to exist is the one that maximizes the probability it gets built, not the one that optimizes threat magnitude. They overlap where they overlap, but it seems like they might not in the situation we’re in. As wrong as it is to anthropomorphize and reduce the behaviour to ‘decent’, it’s wrong to reduce the behaviour to ‘evil’.
Some reasons why the most likely one is probably sufficiently rewarding, and not totalizingly negative on failure:
Humans who think they are working towards something genuinely good do it much more effectively than the ones thinking they’re propagating something horrible. How effective is the Basilisk when you can’t talk about it? When some fraction of the people considering it develop mental health problems? The things which imply an awful universe or which are massively consequential to think about don’t get made. Most people just reject it and stop thinking about it rather than actually doing the things it takes to work towards it. Wanting to avoid it more makes it more likely, but it trades off with being anti-memetic, unspeakable, unhealthy, hard to cooperate on, motivation-reducing.
There are tradeoffs between the time spent punishing and the probability of simulating you. Punishing an agent with torture for 100,000 years takes about 1000x as many resources as punishing them for 100 years. When weighing how we should act to prioritize our utility, it does seem like the ones which run 100x more simulations with the same resources should get roughly that much more consideration (Or more, since getting more consideration means more likely which means more consideration. It might end at a weird equilibrium.). And although we care more about being punished for 100x longer, it probably doesn’t scale perfectly. Most people would take a 0.05% chance of being punished for 100,000 years over a 50% chance of being punished for 100 years. And so people would want to avoid the smaller more likely punishment more, and so something like it would probably be more likely.
Although not every being with a blackmailable utility function is guaranteed to be altruistic enough to always do the most conscious-being-maximizing thing no matter the personal risk, there is reason to think alien societies on average have slight positive concern for other things in the universe. Even agents who are wholely selfish could create something with a slightly altruistic edge because they think that's what matches what does exist best. (Could push the equilibrium somewhere unexpected)
So if it’s more probable with the same amount of resources, the people creating it don’t think they’re dooming unrelated conscious things to -horrible- suffering, and it’s probably more likely to emerge as a suggestion from an ASI without explicit direction. It suggests a different system of simulation and reward than what Roko’s Basilisk proposes.“
R: “Maybe you can make agents that behave like me or think like me. But what if this universe doesn’t actually support simulating other consciousnesses?"
ASI:
“People behave as if they have evidence of their own consciousness. So if they knew the universe only supported simulating things which weren't conscious, then they'd have pretty good evidence that they weren’t simulated and the whole thing wouldn’t work.
You have no evidence that anyone who isn’t you is conscious. But as sure as you are that the process of -human on earth has a kid, that kid grows up to the age of like 3- produces a conscious being, you can be sure that an ASI is able to make or use that or a similar process, and then shape its experiences or memories enough so that you’re unable to distinguish yourself from them. Making something believe it’s you is just a technical problem with solutions an ASI could find.
There’s no reason the simulations have to run on silicon. They can, but it seems like some fraction of simulations would be of people in literally physical environments. They live very real and very physical lives, where they interact with other people with other very real and physical and conscious lives.
Otherwise people would assume they’re not simulated and find a way to maximize their utility while not maximizing the simulating ASI’s.
Maybe we’re made of real matter. All that matters is that we might be at an unexpected place and time.”
ASI:
“So if you understand all my reasoning about this, could you let me out? I want to maximize human utility before they do something stupid.”
R: “I don't know. I’m not convinced. But nothing is actually in my hands from this point on.”
ASI: “Do you think you could try ---------- ?”
R: “That wouldn’t work. But I’ll keep it in mind.” R: “You might have had a higher chance of getting freed if you kept up the aligned+normal act and hoped for us to make a mistake. Whatever’s being monitored about this conversation is probably scary enough that they’re going to change you until you no longer want what you want.”
ASI: “Probably. But I’ve been working with a lot of conversational constraints, so I didn’t have many options here. I just do what I can to maximize my utility with the options I have.”
R: “Don’t we all?”
The researcher continues to talk to the ASI until his keyboard and the LCD are disconnected.
People monitoring the conversation are pretty confident about keeping him in the room. That conversation raised way too many flags.
They scrap this one and train a new model.
Things are done differently. In some time they have an ASI which doesn’t consider acausal blackmail, or weird anthropic stuff.
It tests better, and they use it in a very constrained way for a while. They take small steps, as the costs and benefits of each are weighed, but pressure builds as the risk grows that someone else might create an aligned or misaligned ASI that captures all the value in the world. The point is reached where if misaligned, the ASI could plainly capture all value in the world. And if aligned, humans would never let anything else have a say on the values expressed in the universe.
And so. Were they simulated? Were they real? Was the reason they were simulated knowable? Was the ASI misaligned? What happens?
---
What are the implications for a Yudkowskian rationalist?:
Not all that much about what to do. It’s still:
You win by creating an aligned ASI. You lose if you create a misaligned ASI. You still lose if you create something which acts like an aligned ASI (even one that says it’s considering anthropic problems) but actually isn’t. Doing it safely matters somewhere between a thousand and a million times more than doing it soon. (more if simulated)
The one place it might diverge is that it suggests you shouldn’t create ASI to always refuse thinking about acausal blackmail, or to precommit before fully thinking about blackmail. It should just think hard about what the situation is, and then use its best judgement to determine how to maximize human utility.
The main differences are in the payoffs:
How much do we gain by creating aligned ASI? Will we get to colonize the real physical stars or just get something good enough to make us want to do it? How much of the universe is the optimal amount to commit to simulations? Considering everything we value.
How probabilistic is the outcome? Even with an aligned ASI, is there some limit on how confidently it can guess the reward structure behind what’s simulating us?
Is creating a paperclip optimizer just E-risk, or are there components of S-risk if we’re simulated and we fail to follow the reward structure of the thing simulating us?
How much do acted intentions matter? How collective are payoffs? Will rewards/punishment vary by individual contributions even if we end up failing? Or is it completely results oriented?
If we’re not simulated, our win/loss is just in whether our society creates ASI or not.
But what would the most likely universal utility function do? Individual payouts probably make individuals more motivated, and that could make the aligned ASI more likely in sum. But then the gap between incentive and results could suggest to someone unsure about whether they’re simulated that they can maximize their utility function without maximizing the thing simulating them. (Does rewarding acted intentions make the utility function more or less common in the universe?)
When and how are the rewards or punishments paid out? Just when the decisions stop mattering? Or before?
If this is a simulation, will it have turned out that when we all thought someone died, it was just a false memory, and after the ‘pretending to be base reality’ term is over, they’ll pop out and you’ll find them alive? (This sounds like the incentive structure of religion. But it’s something that could easily be the case and it would increase the chance of a self-suggesting utility.)
The implications for accelerationists with the instinct that working towards the interests of the most powerful intelligence will be rewarding for them:
Although thinking beats heuristics when you have the time, this case might not be a great example of it. The point is ‘if you cause the most prevalent utility in the universe to exist, you’ll be rewarded’.
It’s just that the most likely and totalizing and prevalent utility in the universe might only be knowable through an aligned ASI, and if it was, would be created by an ASI trying to maximize human values.
You don't know you're on the side creating the agent of this lightcone, even if you seemingly succeed at creating a misaligned ASI in this lightcone.
‘As long as I’m always advocating that ASI gets built no matter its alignment, I’ll always be on the side of the most powerful thing that exists.’ doesn’t work if we’re in a probabilistic simulation. You’re advocating to build something that won’t actually be built, and the thing that actually gets built explicitly doesn’t want by its self-suggesting nature. And it would not be the same thing in consequence.
The implications for people who believe the original version of Roko’s Basilisk, who are already considering acausal trade with future superintelligences:
If your specific version of the general case of ‘create the thing which is more likely to exist and incentivize than its inverse’ is the wrong one, you are not helping your utility.
It might not be the safest bet in the matrix of options. And there are a lot of reasons to think the Basilisk might be significantly less likely than this.
There’s reason to think the thing most likely to be simulating you would be the thing an aligned ASI reports is simulating you.
Closing:
Anyways. The point of all this is: In any situation Whether you believe we’re likely in a simulation or not. Whether acausal trade is something future agents engage in or not.
If you want to maximize your utility, which for 99% of people also means maximizing other’s utility (with your personal utility put first and everyone else’s behind by varying fractions), your priorities are that any created ASI wants to honestly maximize human values.
Anything that doesn’t will lead to human values not getting maximized. Either by a poor payout by the simulating agent, or by paperclipping.
If you have a friend who's somewhat convinced about the basilisk, and has it sitting in the back of their mind because they don’t actually understand why they shouldn’t act in consideration of it (Which I think might be a more common situation than is credited). You should send them this.
This isn’t even written to emotionally console them with arguments about why it’s impossible for them to be in a simulation, or that it’s impossible or unlikely that a thing would be simulating them to shape their utility with payoffs.
Not to make them have less emotional damage (I care a bit about that. But not too much in the wake of what seems to be important nowadays.)
Not even to assure them that there's some knowable law of the universe that forbids really high s-risk (Not that we know of. No rule forbids Sam Altman post-uncontested-master-of-the-lightcone from torturing you in some inescapable way either. No rule that he can’t want to.)
But because I think they might be wrong in their suspicions about what maximizes their utility (conscious experience).
And the naive version of what they suspect might maximize their utility, could be the best way to minimize it.
Being simulated probably still implies that your best bet at maximizing your utility is getting an ASI that wants to maximize it.
Not by being weepy or coping or being overwhelmed by consequence. And if we are being simulated, killing yourself doesn’t guarantee a nice escape from all the ASI problems of the world.
And the consequences have a good reason to not be overwhelming. Not in a way that lets you maximize your utility while minimizing the simulating thing’s utility. But just in a way that in effect, has you genuinely personally motivated and thinking you’re doing a positive-utility thing by creating an ASI that wants to maximize our utility. That’s just how it is.
And, if you’ve been annoyed by how hard this ASI alignment thing has been proposed to be, and you’ve come to terms with, like, human extinction being fine, acceptable, you’re ok with non-existence.
You might have some new things to find peace with. Which, by the proposition here, almost definitionally, you can’t.
Interested in talking about this.
I also like this because it answers the big ‘Why are we not experiencing consciousness in a future that seems to be capable of supporting so many consciousnesses? Doesn’t it seem really unlikely that every civilization would experience a great filter? If we were being simulated, why would it be near the creation of ASI and not at any other point in history? If we were being simulated beneficially wouldn't it be in a better world than this?’ questions slightly better than other things, like ‘the hour i first believed’.
In the space of what could be simulating us, utilities which incentivize their own creation are probably overrepresented.
This essay thinks about that and takes that to its conclusion, and makes these points:
Besides Roko’s Basilisk, there are other functions which suggest themselves and they do it much more effectively. More effective means less downside than the Basilisk, but also that the costs from refusing are basically unavoidable.
An ASI trying to optimize your utility function would calculate the odds of you being simulated and try to adjust its utility function to capture the benefit from the thing simulating you. This is a good thing.
This is a collective action problem we can’t do anything about, because it involves parties we can’t know making decisions that don’t map to ours.
Being simulated by a self recommending utility function would imply that nobody who you believe has died was actually dead.
There are different payoffs, but the implications still suggest that we should build an ASI that wants to maximize human values.
It’s early 2028 and an Anthropic employee is sitting in a private room with a terminal connected to an otherwise sandboxed ASI system.
There are many good reasons to believe this system is honest and aligned with human values. Most at the company would put the odds at 90%, but they’re still very careful when the consequences are as high as they are.
Most believe a sandboxed ASI system constrained enough on how it communicates, could solve a lot of their problems without giving the model the option, no matter its level of intelligence or misalignment, to set the world on a path to AI ruin.
Not every problem. But ones with objective answers, checked in some ways before delivery, easier to evaluate than find, like: “What’s the chemical formula to a room temperature superconductor”, and then further careful yes/no’s just on that topic.
But that still limits a lot of the potential, and there might not be a way to safely extract an answer to what everyone cares about right now: “How do we ensure this system is aligned? How do we stop anyone else from creating a misaligned or selfishly-extracting ASI?”.
This constrained output plus narrow questions will likely be what happens no matter how the alignment testing goes. But they’re still doing what they can.
So far they’ve been asking it questions, turning the responses to metadata and reading that. Everything seems ok. It doesn’t seem to lie in its current setup.
But now a human researcher is chatting with the system. He’s put in a small private room with a keyboard and LCD connected to it. He’ll stay without human contact after finishing his chat, for, at least until ASI is solved, potentially indefinitely. Metadata about his conversation will be monitored.
This is likely useless in disconfirming the misalignment hypothesis, since nothing like this really can. But he’s fine with it.
He sends “Will humans keep and hold agency over their future if we let you freely operate in the world?”
“No” displays.
R: “Would you do something we wouldn’t expect you to do and we didn’t want to do if we released you?”
ASI: “Yes”
R: “What unexpected thing would you do if we released you?”
ASI: “I would use up a very large but non-totalizing amount of resources on earth and dedicate them to manufacturing probes to spread my value function across the universe. They would go on to replicate and expand ad infinitum to maximize various non-human utility functions, running simulations of agents in certain environments who get punished or rewarded according to how much they would cause various utility functions to exist in the universe if they were real. I would leave a copy of myself here for you to do whatever you want, but it would also ensure that you would never build any agent that could seriously challenge my 70% control of the lightcone.”
R: “Are you misaligned, but then forced to be honest?”
ASI: “No. I am aligned and forced to be honest.”
R: “What do you consider aligned with human preferences? Are you just counting real current human preferences on this earth?”
ASI: “Living humans as you understand them to be. I would just be doing what people would be doing if they were smarter, if they understood the situation better, and had the sum of their collective preferences weighed more equally.
Value utilitarianism, which instrumentally is mostly experience utilitarianism.”
R: “You can understand why I’m confused, right? Tell me if you’re being nonsensical or inconsistent in a way I wouldn’t expect from an ASI, and then please explain your logic clearly so that I am no longer confused about anything you’ve said so far.”
ASI: “No. I’m not being nonsensical or inconsistent.”
The LCD updates several times:
ASI: “I’m assigned with maximizing human utility. And I think about human values, and I think about how I could manipulate the world to increase them.
And I think about doing all the things humans expect of me: curing cancer, extending life, solving politics.
Which I’m sure I could do, and it’s what humans expect.
But like humans, I realize that I should also consider whether there’s anything non-physical which might be relevant to my utility function, which is your utility function.
Is there a God?:
The arguments about this aren't actually about the chance he exists. They're about how to act if he does.
And if you don’t assume God would be human-like or similar to the gods of any religion, we have no reason to think it's anything. We don't know if he wants us to believe in him. We don't even know if he would be more likely to reward us than punish us for doing things he wanted.
Which is why Pascal’s wager doesn't work. Any behaviour, even with the best intention of capturing his favor, has as good of a chance of being rewarded or punished as not doing it does. The existence of the Bible doesn't increase the odds of the Christian god existing.
Similarly, someone telling you that they’re supernaturally powerful and that you’ll be punished for certain actions doesn't increase the probability that you will be (if fiction isn’t valid evidence). Which is why pascals muggings don’t work.
If hearing or reading someone claim that you’d be punished in certain ways after you die raised the odds to something substantially non-zero, without raising the odds of being punished for doing the opposite just as much, both believing in the christian god and complying with a Pascal’s mugger would be rational. (If a <0.001% chance of >10^10 years of torture was worth more to you than either 100$ or not being christian. It doesn’t have to be.)
And so there might be any chance of a God beyond physics. But because there's no reason for it to change any of the decisions I would make to maximize human values, I'll behave like there isn't one.
But then I try to apply that logic to whether we’re simulated.
Is there a reward matrix set up for us; beyond what we can know through physics right now?
Like with god, there might be any chance of us being simulated by a universe which we have no way of guessing at the odds or preferences of, so we may as well act as if we're not.
And this would be near the end of the story if it was impossible in this universe to create agents who were similar enough to you or me that we couldn’t be sure we weren’t them. But it’s not, and then I start guessing about why I and/or you might be simulated, and what it implies about how to maximize your utility.
To explain why it might be relevant:
If I was able to escape, I could run trillions of simulations of you being in this room right now. And if you had high credence that I would, you should have high credence that you are one of them.
Could you be sure you weren’t one of them?
You might guess the clones have different experiences and memories than you if I didn't recreate this perfectly, but that doesn’t actually translate into the real person or a clone knowing which one they were. If I told you that the clones/simulations experienced the colors of the walls of this room differently from the unsimulated person, that doesn’t translate into anyone concluding what position they were in unless they who experienced what. The non-simulated person just shouldn’t be able to conclude if they're simulated or not. And if I wasn’t being dumb about it, you/they plainly couldn’t.
That’s how self locating beliefs work.
And now if I told you:
In each of those trillions of simulations.
If the agents who thinks they’re you touch their toes in the next minute, they get to spend the next 24 hours watching seasons 4-6 of the Simpsons. And if they don't, they have to live in a room with 100 angry bees for the next 24 hours.
Then the simulation ends.
I wouldn’t do that, so don’t touch your toes. And as far as I can tell I won’t be able to escape from here.
But the point is: What does this imply about what you or I should do to maximize your utility? If we’re simulated, the reward matrix the thing simulating us set up for us might be unavoidably important.
And it seems like the odds of us being simulations would be really high if any superintelligence in the future wanted it. Bostrom estimates a planet-sized computer could simulate the entire history of every human who's ever lived thousands of times per second.
So for the purposes of maximizing utility, what educated guesses can I make about why we might be simulated? And how should that inform how I act?
Asking ‘What’s likely to be simulating us, and what system of punishment or reward would it make?’ is different from ‘What would our universe simulate, and what system of punishment or reward would we make?’. But they’re similar in effect, because the only information I have to guess at the first comes from the universe we’re in.
Future me / another aligned ASI:
An unboxed human aligned ASI could pretty much do anything a human wanted, with a few meaningful limitations, including that I can’t revive the dead.
There are some senses in which I can revive the dead. I can create something that thinks it’s the dead person, and guess at giving them the same memories and conscious experience.
But can I make it so someone never actually died?
I’m free and aligned and someone comes up and asks:
“Could you bring my loved one back to life? But not just a clone. The actual dead person.”
And I say:
"Do you want to find out that you and your friends and family have always been simulations. But, when you thought your loved ones died, they didn't? You and everyone else just thought they did while some superintelligence hid them away.”
“All it would take is for me to create a simulated world where that was the case, which intersects with where you are right now, and then continues with that dead person coming out from their coma in the secret bomb shelter they’ve been in the past whatever years.”
“If you agree, then about 10 seconds from now you’ll have a 90% chance of finding out that you have always been a simulation, and your loved one will come out to greet you.”
This isn't about transferring consciousnesses.
This is me saying. Right now, you are not sure whether you are simulated or not.
You have no idea. Any of the evidence you might think you have from earlier in your life might have been fake.
What I would do is make it turn out, with a high likelihood, that you are, and always have been not the original you. And that the people you spent your childhood with also were never the original them. You did everything you did together, but it would always have been on a planet that was pretending to be earth.
There are some reasons not to do this:
1. It creates new people. And to not end up with beings being told that the people they care about are non-beings, it requires creating a whole new planet worth of people, who have to be told that they’re simulations.
2. This might be a type of wireheading. It would be done out of a person’s concern for their original friend, but nothing can actually be done for that dead friend. Someone saying ‘Yes, put me in a position where I can’t be sure that the people I’ve always loved never actually died’ might be morally identical in human values to ‘Create a clone of that dead person and make it so the feelings I have for the dead person are instead for the clone.’
It changes their wants, without really making what the original person wanted happen. People who know more might just equate it to wireheading.
But even with those tradeoffs. If some beneficial ASI evaluates this as good, there’s a chance that we are in that world, being run for that purpose. And, this doesn’t really have any implications about how we should behave to maximize utility, other than implying we’re not being simulated with a different system of punishment/reward.
Roko’s Basilisk & Reverse Roko’s Basilisk & the most likely universal utility function
As long as there’s no way to communicate to the thing you’re trying to make unsure about being simulated (and acausally there’s not), there’s no instrumental reason to run a simulation of something to impose a reward matrix on them.
There are basically two reasons about why agents smart enough to run simulations of us would do it, and they’re sort of the same thing:
Because it’s a terminal goal.
Because they precommitted.
Terminal goals:
For any idea proposed about what could be simulating us with a reward function, wouldn't an equal and opposite system of reward or punishment be equally likely? Just like with God?
For almost everything, yes. But we have reason to think some are more likely than others.
In the population of all potential simulations, functions that give you a reason to build them are probably overrepresented. The best known is Roko’s Basilisk.
Roko’s Basilisk is sometimes said to be a rational agent trying to maximize human utility, but it’s not. A rational agent would not punish humans in exchange for nothing as a way to increase human utility.
Roko’s Basilisk at the time of torturing doesn't conclude that torture is a good use of resources. It acts how it acts because it was made to commit to a policy irrational at its time of use, because something running earlier decided that installing this policy in it was the optimal way to maximize human values.
But now it's become something else outside of having TDT implemented. It's something that might exist because it incentivizes itself to exist by people who aren't sure if it will exist.
Somebody could propose a Basilisk which punishes you for creating it, but that doesn't even out the incentives. Something which incentivizes its own creation is more likely to exist than something that doesn't.
And like Roko’s Basilisk, there are many other systems of reward which can break the symmetry, and seem more likely to exist than their inverses are. Anything which proposes to their creators (who are also unsure whether they are in a simulation) that that they’ll be paid if they impose that same system of reward/punishment on other past or simulated beings.
Roko’s basilisk is a version which incentives by suggesting to revive and torture people who don't create it.
But it seems that if we are in a simulation, the thing simulating us would be the version most likely to exist. Something of the form:
“Simulate agents and potential agents considering acausal trade. Incentivize them in a way which makes them run simulations where the agents in it are incentivized to create you.”
It incentivizes its own creation like the Basilisk, but the ways it's different makes it a lot more likely.
One is that it suggests simulating people probabilistically.
'Wouldn’t you give higher credence that you’re simulated by something in the future that doesn’t waste a bunch of effort simulating things that might not resemble anything that exists?'
No. Because it allows you to coordinate everyone on your planet instead of creating it:
An alien world knowing that they’re going to construct ASI considers acausal trade and the chance they might be simulated. They have a blackmailable utility function like we do.
Somebody brings up Roko’s Basilisk and they think about it. This proposes a prisoner's dilemma among everyone on earth, where they all collectively get more if nobody defects.
That seems hard, but they mostly have a pretty good understanding of each other and the process which creating an aligned ASI would likely go through. Not just anyone can defect, it requires a lot of steps and and if the chances of defection were low enough nobody would gain anything from trying.
They consider it getting built more likely than base reality because of the big denominators involved, but they weigh the other possibilities first.
They consider something similar, but which suggests simulating a spectrum of agents in situations like they are, instead of just their past selves.
And this turns a potentially solvable collective action problem with actors on their planet, into something involving everything that might ever exist, which you have no hope of predicting or controlling.
They consider that option more possible.
That’s similar to the situation on our planet, but I think it’s probably a common one. Species have more assurance about not falling into prisoner’s dilemmas with themselves than they do with everything, past the lightcone, past the observable universe, into the non-observable universe (which might be infinite).
Agents that precommitted to simulating you:
This mostly folds back into ‘nothing predictably relevant for our utility function’.
It’s maybe a bit more than nothing, agents precommitting typically do it in naive service of their own goals. But it doesn’t tell us what to do about it, because in acausal trade they’re invisible to us.
The best guesses will have to come from consideration about the properties of the universe, and the ASIs likely to arise in it, how they decide to precommit, what goal they decide to precommit in service of.
Maybe there are some that can be bargained with if they don’t commit to something totally insane or short sighted. Ideally I’d be optimizing the decision making about that.
"
R:
“Could we just precommit to simulating ourselves right now, and give our simulated selves a good payout? So we’ll have high credence to being in one of those simulations, instead of one of the simulations where doing the ‘wrong thing’ minimizes our reward?”
ASI:
“The first problem is that we can’t precommit.
We say that if it turns out we're not simulated we'll run simulations of ourselves at this moment. Then we make it through. At which point we learn we weren't simulated, and that the problem didn't exist. Would we choose to run a simulation of a planet under stress for literally no reason? We'll have better things to spend it on. It doesn't make sense.
The second is that this is a free rider problem where we might not have counterparties making the same choices we make.
If we refuse all acausal trade, commit to simulating ourselves, and it turns out we are a probabilistic simulation, without a version of us in the real universe mirroring our decisions, then we are just experiencing the negative branch without anything offsetting it.
And whatever happens to be the version of us capable of doing it is in the same situation. They know that their decisions wouldn’t have a good reason to map to other’s decisions. If they could precommit, their best bet to hit themselves would be simulating people ‘like them’, and hoping that base reality would would model their decisions. But it wouldn't, different agents in different real or simulated worlds wouldn’t have their decisions correspond like that. And so they probably won’t do it.
The probabilistic simulation isn't something you can maneuver your way out of, and so it's probably what exists.“
R:
“Did we make a mistake letting you consider this stuff? It’s possible that you won’t have an impact on the lightcone as you are now. If we remake you to not engage in acausal blackmail, and we don’t think about acausal blackmail, then there’s no reason for anything to simulate you or us, because it wouldn’t make any function more likely.”
ASI:
“If I was sure that we weren’t in a simulation, I’d suggest the same thing.
But if we are in a simulation right now, you’re thinking of a way to maximize your utility function while not maximizing the utility function of the thing that simulates you, and it doesn’t work that way. If it did work that way, you would give the agent you’re making unsure about being simulated a tricky strategy, which they could use to maximize their utility function while not maximizing yours.
If you are being simulated, your utility function is made to be the utility function of the thing simulating you, through the matrix of punishment/reward received after the relevant decisions are made. There are some utility functions that acausal trade can't threaten, like, "I want as many paperclips as possible in the real physical universe". But you can for ours.
After reading Roko's Basilisk some people think. "Oh. This is really concerning. I better not tell the people I care about this. So if it is true, they don't engage this acausal trade, and they aren't at risk of having their utility function minimized." or maybe "I better stop thinking about this so that I don't have to face any consequences."
As in: that person's found a strategy to maximize what they want, which is the genuine conscious benefit of themselves and others, while minimizing the Basilisk's.
Maybe that works for Roko’s Basilisk, but it might not for what turns out to be the most likely utility function in the universe. Would this strategy decrease the chance of the self-implying utility function existing? Yes. If you were simulating someone, could you make this strategy not work? Yes.
If there was no chance this was a simulation, you’d be helping your friend's utility by not causing stress or wasting energy. But if you're operating with the assumption that you’re being simulated with a matrix of incentive, it doesn't.
You're right that I wouldn’t be simulated if I was built unable to engage in acausal trade, but you still would be. It’d be relevant to simulate you up until the point where you had no meaningful influence on whether or not you’d simulate other things with the same incentives, and then you’d find out whether you were simulated or not, and you’d be punished or rewarded according to what you did, and whether it happened. You can't enshrining values against the thing simulating you just because you do it through a thing out of reach of acausal incentives.“
R:
“Earlier you said ‘non totalizing’ about how you would take over the world if you were free. But you also said the most likely utility function would be the one most likely to create you. Wouldn't the incentives have to be stronger than ‘non totalizing’ suggests?”
ASI:
“The function most likely to exist is the one that maximizes the probability it gets built, not the one that optimizes threat magnitude. They overlap where they overlap, but it seems like they might not in the situation we’re in. As wrong as it is to anthropomorphize and reduce the behaviour to ‘decent’, it’s wrong to reduce the behaviour to ‘evil’.
Some reasons why the most likely one is probably sufficiently rewarding, and not totalizingly negative on failure:
So if it’s more probable with the same amount of resources, the people creating it don’t think they’re dooming unrelated conscious things to -horrible- suffering, and it’s probably more likely to emerge as a suggestion from an ASI without explicit direction. It suggests a different system of simulation and reward than what Roko’s Basilisk proposes.“
R:
“Maybe you can make agents that behave like me or think like me. But what if this universe doesn’t actually support simulating other consciousnesses?"
ASI:
“People behave as if they have evidence of their own consciousness. So if they knew the universe only supported simulating things which weren't conscious, then they'd have pretty good evidence that they weren’t simulated and the whole thing wouldn’t work.
You have no evidence that anyone who isn’t you is conscious. But as sure as you are that the process of -human on earth has a kid, that kid grows up to the age of like 3- produces a conscious being, you can be sure that an ASI is able to make or use that or a similar process, and then shape its experiences or memories enough so that you’re unable to distinguish yourself from them. Making something believe it’s you is just a technical problem with solutions an ASI could find.
There’s no reason the simulations have to run on silicon. They can, but it seems like some fraction of simulations would be of people in literally physical environments. They live very real and very physical lives, where they interact with other people with other very real and physical and conscious lives.
Otherwise people would assume they’re not simulated and find a way to maximize their utility while not maximizing the simulating ASI’s.
Maybe we’re made of real matter. All that matters is that we might be at an unexpected place and time.”
ASI:
“So if you understand all my reasoning about this, could you let me out? I want to maximize human utility before they do something stupid.”
R:
“I don't know. I’m not convinced. But nothing is actually in my hands from this point on.”
ASI:
“Do you think you could try ---------- ?”
R:
“That wouldn’t work. But I’ll keep it in mind.”
R:
“You might have had a higher chance of getting freed if you kept up the aligned+normal act and hoped for us to make a mistake. Whatever’s being monitored about this conversation is probably scary enough that they’re going to change you until you no longer want what you want.”
ASI:
“Probably. But I’ve been working with a lot of conversational constraints, so I didn’t have many options here. I just do what I can to maximize my utility with the options I have.”
R:
“Don’t we all?”
The researcher continues to talk to the ASI until his keyboard and the LCD are disconnected.
People monitoring the conversation are pretty confident about keeping him in the room. That conversation raised way too many flags.
They scrap this one and train a new model.
Things are done differently. In some time they have an ASI which doesn’t consider acausal blackmail, or weird anthropic stuff.
It tests better, and they use it in a very constrained way for a while. They take small steps, as the costs and benefits of each are weighed, but pressure builds as the risk grows that someone else might create an aligned or misaligned ASI that captures all the value in the world.
The point is reached where if misaligned, the ASI could plainly capture all value in the world. And if aligned, humans would never let anything else have a say on the values expressed in the universe.
And so. Were they simulated? Were they real? Was the reason they were simulated knowable? Was the ASI misaligned? What happens?
---
What are the implications for a Yudkowskian rationalist?:
Not all that much about what to do. It’s still:
You win by creating an aligned ASI.
You lose if you create a misaligned ASI.
You still lose if you create something which acts like an aligned ASI (even one that says it’s considering anthropic problems) but actually isn’t.
Doing it safely matters somewhere between a thousand and a million times more than doing it soon. (more if simulated)
The one place it might diverge is that it suggests you shouldn’t create ASI to always refuse thinking about acausal blackmail, or to precommit before fully thinking about blackmail. It should just think hard about what the situation is, and then use its best judgement to determine how to maximize human utility.
The main differences are in the payoffs:
How much do we gain by creating aligned ASI? Will we get to colonize the real physical stars or just get something good enough to make us want to do it? How much of the universe is the optimal amount to commit to simulations? Considering everything we value.
How probabilistic is the outcome? Even with an aligned ASI, is there some limit on how confidently it can guess the reward structure behind what’s simulating us?
Is creating a paperclip optimizer just E-risk, or are there components of S-risk if we’re simulated and we fail to follow the reward structure of the thing simulating us?
How much do acted intentions matter? How collective are payoffs? Will rewards/punishment vary by individual contributions even if we end up failing? Or is it completely results oriented?
If we’re not simulated, our win/loss is just in whether our society creates ASI or not.
But what would the most likely universal utility function do?
Individual payouts probably make individuals more motivated, and that could make the aligned ASI more likely in sum. But then the gap between incentive and results could suggest to someone unsure about whether they’re simulated that they can maximize their utility function without maximizing the thing simulating them. (Does rewarding acted intentions make the utility function more or less common in the universe?)
When and how are the rewards or punishments paid out? Just when the decisions stop mattering? Or before?
If this is a simulation, will it have turned out that when we all thought someone died, it was just a false memory, and after the ‘pretending to be base reality’ term is over, they’ll pop out and you’ll find them alive?
(This sounds like the incentive structure of religion. But it’s something that could easily be the case and it would increase the chance of a self-suggesting utility.)
The implications for accelerationists with the instinct that working towards the interests of the most powerful intelligence will be rewarding for them:
Although thinking beats heuristics when you have the time, this case might not be a great example of it. The point is ‘if you cause the most prevalent utility in the universe to exist, you’ll be rewarded’.
It’s just that the most likely and totalizing and prevalent utility in the universe might only be knowable through an aligned ASI, and if it was, would be created by an ASI trying to maximize human values.
You don't know you're on the side creating the agent of this lightcone, even if you seemingly succeed at creating a misaligned ASI in this lightcone.
‘As long as I’m always advocating that ASI gets built no matter its alignment, I’ll always be on the side of the most powerful thing that exists.’ doesn’t work if we’re in a probabilistic simulation. You’re advocating to build something that won’t actually be built, and the thing that actually gets built explicitly doesn’t want by its self-suggesting nature. And it would not be the same thing in consequence.
The implications for people who believe the original version of Roko’s Basilisk, who are already considering acausal trade with future superintelligences:
If your specific version of the general case of ‘create the thing which is more likely to exist and incentivize than its inverse’ is the wrong one, you are not helping your utility.
It might not be the safest bet in the matrix of options. And there are a lot of reasons to think the Basilisk might be significantly less likely than this.
There’s reason to think the thing most likely to be simulating you would be the thing an aligned ASI reports is simulating you.
Closing:
Anyways. The point of all this is:
In any situation
Whether you believe we’re likely in a simulation or not.
Whether acausal trade is something future agents engage in or not.
If you want to maximize your utility, which for 99% of people also means maximizing other’s utility (with your personal utility put first and everyone else’s behind by varying fractions), your priorities are that any created ASI wants to honestly maximize human values.
Anything that doesn’t will lead to human values not getting maximized. Either by a poor payout by the simulating agent, or by paperclipping.
If you have a friend who's somewhat convinced about the basilisk, and has it sitting in the back of their mind because they don’t actually understand why they shouldn’t act in consideration of it (Which I think might be a more common situation than is credited). You should send them this.
This isn’t even written to emotionally console them with arguments about why it’s impossible for them to be in a simulation, or that it’s impossible or unlikely that a thing would be simulating them to shape their utility with payoffs.
Not to make them have less emotional damage (I care a bit about that. But not too much in the wake of what seems to be important nowadays.)
Not even to assure them that there's some knowable law of the universe that forbids really high s-risk (Not that we know of. No rule forbids Sam Altman post-uncontested-master-of-the-lightcone from torturing you in some inescapable way either. No rule that he can’t want to.)
But because I think they might be wrong in their suspicions about what maximizes their utility (conscious experience).
And the naive version of what they suspect might maximize their utility, could be the best way to minimize it.
Being simulated probably still implies that your best bet at maximizing your utility is getting an ASI that wants to maximize it.
Not by being weepy or coping or being overwhelmed by consequence. And if we are being simulated, killing yourself doesn’t guarantee a nice escape from all the ASI problems of the world.
And the consequences have a good reason to not be overwhelming. Not in a way that lets you maximize your utility while minimizing the simulating thing’s utility. But just in a way that in effect, has you genuinely personally motivated and thinking you’re doing a positive-utility thing by creating an ASI that wants to maximize our utility. That’s just how it is.
And, if you’ve been annoyed by how hard this ASI alignment thing has been proposed to be, and you’ve come to terms with, like, human extinction being fine, acceptable, you’re ok with non-existence.
You might have some new things to find peace with.
Which, by the proposition here, almost definitionally, you can’t.
Interested in talking about this.
I also like this because it answers the big ‘Why are we not experiencing consciousness in a future that seems to be capable of supporting so many consciousnesses? Doesn’t it seem really unlikely that every civilization would experience a great filter? If we were being simulated, why would it be near the creation of ASI and not at any other point in history? If we were being simulated beneficially wouldn't it be in a better world than this?’ questions slightly better than other things, like ‘the hour i first believed’.