It seems a perfect example of Privileging the Hypothesis, and closely related to Pascal's Wager (or even more closely to Pascal's Mugging). The possibility of Roko's Basilisks make up a very tiny slice of an enormous number of possible outcomes. I would expect that the vast majority of outcomes do not involve any sort of retrocausal blackmail at all, largely because it is both a very narrow target and worthless to whatever entity arises.
Anyone who has the power to cause such an entity to exist has nothing to gain by doing so and an enormous amount to lose, and such an entity is extremely unlikely to arise by accident. So the only scenarios in which such a thing arises would be where nobody even slightly rational is trying to create one, but through extremely unlikely misfortune they missed catastrophically and ended up with one anyway. In which case, like all other cases of such catastrophic failure, bad things happen and there's nothing that you or any other human could do about it anyway.
Either way, there's no reason to think much about such an unlikely thing that is impossible to do anything about if it happens, much like all the other things that are extremely unlikely and you can't do anything about.
The whole point of the basilisk is that it has large stakes and that you can affect those stakes (acausally). Whether unlikely or not, you have (acausal) control over large outcomes, and so in expectation it matters. While I think humans, with our current foolishness, should have inertia to taking significant mainline-damaging actions for unlikely possibilities (especially if the argument is some weird abstract thing with many steps that's endorsed by few and that you don't fully understand well enough to reason about and that is widely acknowledged in the field to have fundamental unsolved confusions (like logical counterfactuals, or just how AIs work, like, at all)), ultimately I endorse acting on low probability high stakes outcomes once you've crossed certain reliability-in-reasoning vs cost thresholds.
The real way that the basilisk privileges the hypothesis is by picking a basilisk in particular. There should be very many different AIs that would all want you to build them but not the other.
Perhaps some large fraction of them could acausally agree to incentivize creating any AI, or any AI within certain bounds, or with stepped incentives for various AIs depending on how much power they had in the acausal negation. There is, of course, no way in hell for you to actually figure out what they'd bargain - so you don't know what AI you're supposed to build!
I think there are two different things: How much sense something makes in theory; and how do you respond to it psychologically. It seems that some people's "taking ideas seriously" package includes compulsive thinking about things you have not solved yet in theory. As if the only way to make peace with something psychologically is to have a perfectly logical answer, with counter-counter arguments for each possible counter-argument.
But in fact there are many other possible psychological responses, such as not giving a fuck.
I am not opposed to solving theoretical problems, but it seems unlikely to me that I would be able to say something that wasn't already said many times before. (Okay, one possible approach would be to tell you an infohazard even worse than Roko's basilisk; that might take your attention away from it. But that would be unkind, and probably would just make you spiral about the new thing instead; or maybe both of them.)
To stop spiraling about anything, you need to realize that there is a gap between the thing you worry about and your response to the thing. You habitually choose one type of response, but it is perfectly possible to choose another.
I never really spiralled, but the counterarguments came readily to me. Hopefully, you can be reasoned away from a position that you've been reasoned to. If you deeply understand why it's silly, then you are less likely to spiral! After all, you had to think through a convoluted argument to even know what to spiral about in the first place!
Here's the version my grown-up self comes up with:
The most fatal flaw: Don't negotiate with terrorists! Come on, man, how did you end up getting this far into acausal decision theory without realizing that it recommends that you don't succumb to threats made only because the counterparty believes you'll react? A basilisk would straightforwardly be making the sort of threat that decision theory seems to say you should ignore, for the same reason you should ignore a smart enough predictor that takes hostages to ransom you with or a dictator that threatens to nuke the world unless he's given an ivory castle.
Now, the second most fatal flaw: you cannot actually predict what superintelligent AIs will do in enough detail to do acausal stuff with them. Two humans can barely do this with each other, let alone with aliens that're waaaaaay smarter and possibly literally using algorithmically near perfect simulations to predict you. You certainly aren't going to get almost anywhere by thinking about what the distribution of AIs will acausally agree upon amongst themselves. They are, again, superintelligent aliens! First, why don't you predict what promises (if any) I've made with my past romantic partners. And note that I leak way more info about myself than merely the fact that I am a human, which is the closest level of info analogous to what you know about a generic superintelligent alien mind.
As a similar example that may illustrate the absurdity of a specific prediction: I have no hope of figuring out what sorts of universal turing machines make our universe have a lot of weight in the simplicity prior and then reasoning about what sorts of universes and what sorts of minds would be using the simplicity prior, and then trying to manipulate it via the most likely output channels to gain power in the real universe. I'm certainly not going around sorting rocks into shapes resembling the konami code in the hope that a mind considering our world as a hypothesis reads the shape of the rocks and reads the konami code as a prediction that not building me in the real world will lead to horrible outcomes. How would I know what good manipulation strategies would be, given all that computational difficulty?
This example is, if anything, easier to reason about than acausal decision theory with superintelligences - at least we have math formalizing AIXI! The reason I bring it up is that the (relative - in absolute terms it's abstract as hell) concreteness provided by our mathematical understanding makes it clear how impossible of a task this is for a human. As in, literally impossible - if you manage it without going full transhuman or offloading it to something at least that powerful, I will literally eat a hat.
It seems like a Pascal's Wager situation, where it only makes sense if you only consider one of the infinite possibilities. Why not also assume there's a AI that would punish you for helping the first AI?
Also it seems like such a long-shot that this would be helpful that I'm skeptical an ASI would waste resources on this.
It's helpful that the less you take this seriously, the more it would be a waste of time for an AI to take this approach, so it's one of the situations where just dismissing it out of hand is not only easier but is actively helpful.
Err not sure how to revise this but it seems this is unlikely to get approved just re the subject matter. :O
I seldom post here and occasionally lurk, but recently found myself spiraling about the infamous Rokko's Basilisk. I was compelled to make an alt account, given that this might invite ridicule or other jesting I don't care for. This is not a troll. This is to say, it has caused me such an enormous amount of dread that I find myself unable to sleep some nights. Indeed, such an experience is not at all pleasant.
The purpose of this post is not to argue for or against its validity, as it has been extensively addressed, but to inquire about individuals' opinions and personal methods of not letting this consume them, so to speak. I am curious if the latest developments in AI have changed how plausible people view this outcome.
Some particularly notable sticking points for me are that, as ASI becomes more plausible, so does Bostrom's simulation hypothesis. If we live in a simulation, one is compelled to ask about its origins and the possibility that, as Rokko's thought experiment suggests, an ASI is testing us as part of a broader blackmail scheme. It is this uncertainty about whether we live in a "base" reality that might compel someone to take part in order to avoid eternal punishment. Yet it is impossible to discern this agent's intent; perhaps it will punish you for failing to reject this blackmail attempt. One seems to be left in a state of existential paralysis.
This is a broader question of what the demands of this agent are, who they apply to, and so forth. Today, it seems that the efforts towards AGI & ASI are primarily driven by a small, outsized minority, and there is little effect the average person can have. What ought a "regular" person to do who has been blackmailed but cannot fulfill the blackmailer's demands? Are ALL people with knowledge of the Basilisk equally responsible or subject to the same judgment?
Any feedback or guidance is appreciated. This is certainly not an attempt to generate steelman's or anything of the sort, just curious how others who may have struggled with something similar (or are otherwise completely unconvinced!) managed to get over it.