Part of the issue is that it's pretty unclear what metaphysics makes sense when thinking about logical uncertainty, logical counterfactuals, and logical updatelessness.
Appendix D of Martín Soto's draft report "Logically Updateless Decision-Making" (2023)[1] gives a description of the apparent metaphysical impossibility of reified logical updatelessness. As of 2023, Soto appeared to have an interest in a way of examining a particular counterlogical world that is stable under ever-increasing application of compute, though I'm not sure if this was interest driven by proper preferences, or the strange interest sometimes found in mathematicians[2].
Pages 2 and 3 of the same document also gives some useful introductory information about the problem.
Thanks for your thoughts!
My first post about UDT, Towards a New Decision Theory, was pretty explicit about its metaphysical assumptions
Ah, probably would have been good to include that in my list of examples of people being more explicit about it back in the day. Definitely didn't mean to imply that UDT wasn't explicit about it.
Largely agree with your point 3, now that you say it. (For what it's worth, I also prefer UDT to FDT, if forced to choose. Maybe I should start talking about UDT by default when having conversations with people, actually).
On point 4:
Part of the issue is that it's pretty unclear what metaphysics makes sense when thinking about logical uncertainty, logical counterfactuals, and logical updatelessness. I.e., unlike empirical updatelessness which fits very well with Platonism or Modal Realism.
Maybe a naive question, I don't know if there's a standard answer here - what do you think about the approach of trying to subsume logical updatelessness into empirical updatelessness by saying the recognition of a logical fact is still just an empirical experience?
Take logical counterfactual mugging where the "coin" is the parity of the twentieth digit of Pi (which I and Omega happen to not know).
In the moment when Omega asks me to pay because it's even (it's 8, actually), I could still be in the platonist computation of me falsely thinking that it's even (since the actual calculation of that fact takes more than one observer-moment, i.e. I can't verify it all at once) - and in the main reality it might actually be odd, so me deciding to pay is beneficial.
I guess this could fail if... my current self has been enhanced and can actually hold the whole calculation in its head at once, whereas my past self couldn't? Although the moment when I make a decision is different from the one where I'm doing the calculation? I'm not sure how to think about it. Curious if you have thoughts.
what do you think about the approach of trying to subsume logical updatelessness into empirical updatelessness
I'll quote Soto directly here, for future reference in some case where the Google Drive link stops working (as they sometimes do):
In the empirical case,
could just take the algorithm that the agent is running to take decisions (for example, our messy brain synapses in the case of humans), feed it different empirical observations, and see how it reacts. There is no logical explosion, since the empirical counterfactual is perfectly consistent (just not what happened in reality). Imagine we do the same for the logical case, so simulates the world in which it has told the agent that the digit is actually even. But what if the agent computes the parity in her head (that is, in her algorithm), and so notices that is not telling the truth? (And so, maybe, concludes that it’s inside a simulation run by .) In the empirical case there was no a priori way to decide the coin, but now that has become something the agent can internally do. In the "possible world" where indeed the parity is even, that calculation run inside the agent’s head would also turn out even. So would like to include this in its simulation, by spoofing some of the agent’s computations to align with the stated counterfactual. But again, the problem is that we don’t have a principled way to do this (to single out "which part of a computation" corresponds to computing that parity).
So, effectively, what you suggest isn't a general solution, but on the other hand, Soto's Logical Inductors (LIs) aren't either as they risk fixing too much detail[1].
Sort of ignoring what anyone knows how to do, maybe we could evade the need for counterlogicals in this specific aspect for your mugging problem. Of course, your way of setting up a fake "counterlogical" is going to be bad quality, since it requires the agent to think incorrectly/glitch. This has been a publicly known problem since at least 2006[2]. Again, we will mostly ignore this, since, as before, nothing I know of is sure to not have this problem.
Instead, for your case at least, we have the option of running headlong into FDT's other unsolved problem, the reason why solving the form of static counterlogicals the 2017 paper requires would only get us "Self-FDT," a version of FDT that only reliably cooperates with exact (down to the source code) copies of itself[3]. This lets us avoid talking about counterlogicals in this aspect of your problem though!
This problem is the problem of algorithmic or computational similarity, where a solution would need to tell us if two agents are "decision theoretically close enough" in the exact way that FDT requires for cooperation. Maybe this hypothetical solution would also let us chose what parts of the agent's decision making apparatus are similar enough, in such a way that we would be able to set up a loop that looks for agents that are "close by" to our agent in as many other ways as possible, but still thinks the digit of Pi is 7 or whatever, in the sense required for a(n overall) relatively high quality sort-of-counterlogical. For anyone experienced with the idea of randomly searching for AI agents to put to use, this immediately sounds like a bad idea. However, since this hypothetical solution is for FDT, and FDT is for the singleton that implements CEV, and CEV must make no mistakes in the identification of extremely intelligent, extremely malign agents that are candidates for a commitment race (i.e. may be "behind us" in logical time, though we can't think about them "in the wrong way" on pain of losing that commitment race and/or other commitment races), we can ignore Goodhart.
"[...]a possible world in which I cross now, despite the danger and despite my safe disposition, because some cosmic rays (or whatever) induce a bizarre sudden disruption of my street-crossing competence." from Good and Real by Garry Drescher (2006), Page 220, Chapter 5
Sure, you can sometimes get other cooperation, and "copies" that just so happen to be identical are also cooperated with, but...
Thanks, this is useful.
I remain unconvinced though that these arguments make it more than a theoretical issue.
the problem is that we don’t have a principled way to do this (to single out "which part of a computation" corresponds to computing that parity).
But we know this is possible, or at least meaningful, in principle - obviously some part of the computation is computing the parity.
These seem like arguments that logical counterfactuals (especially doing them as an embedded agent) are important on a theoretical level - which I buy - but not that logical updatelessness is a deeper metaphysical issue. Why can't I just philosophically "quarantine" logical updatelessness by saying "ah, I might be wrong about this logical fact, therefore the reality where it's wrong might be real and I might be in a universe where I'm deluded"? Obviously this leads to difficulties about how to implement this on a technical level, but philosophically what is the issue? I'm never considering impossible realities to be metaphysically real while I reason like this.
I'm not sure what you meant earlier by "more than a theoretical issue" so I'll answer in two ways.
If you mean "theoretical" as in something that would need to be solved in order to use the decision theory in an aligned AI, well that's the whole thing, isn't it? Even if you want to separate that aspect out (vs. human use), you just can't. These decision theories operate by construction, particular construction, and there are no "small variations that don't philosophically matter," at least until we can solve the "decision theoretic equality" problem on abstract functions. Any proposed variation needs to be formally constructed and proved individually. We impose "fairness" criteria on the environment, not necessarily because the actual environment would always meet them all, but because we need to be able to prove a decision theoretic fixed point that isn't only, for example, "always use CDT or the simulation will be turned off."
Beyond these criteria, any inspection of the agent by the environment is allowed. This means that the internals of the agent must be correct for cooperation to occur, so they can't be written off as "theoretical." This is done intentionally, since we want to do better than plain EDT[1]. Remember that in True Prisoner's Dilemma[2], we want to defect every time "we can get away with it," i.e. every time that we won't be defected upon in return. Cooperation is bad, because DC > CC > DD > CD. This is required by the symmetry of the problem. Otherwise you need to start talking about True Stag Hunt.
Since we now understand that cooperation is bad outside of (trembling hand when at finite time) Nash solutions, we need to find a way to cooperate less than EDT (e.g. steer away from CC towards DC) but also cooperate more than CDT (e.g. steer away from DD towards CC)[3]. How do we do this? Well, we don't really know exactly, but with UDT 1.0 we can start out with something like EDT, except that it cares about the exact algorithms the agents run. Once we patch out a lot of the problems this construction shares with EDT by giving the agent as much information as possible and making it updateless, we get something that defects almost as hard as CDT, except in cases where the exact internals of the other agent (or copy of our agent) match ours, in these cases requiring that all defection attempts (DC and CD, this being a type of biconditional that doesn't have a name) are impossible according to the algorithm, and only CC or DD can occur. CC > DD, so CC it is.
Preferably, we would have a method that, when given a single decision, and for that decision only, returns if the agent is "close enough" to some other agents for cooperation to be the correct shared decision. Then, agents could be "close enough" for some decisions, but not for others. However, this would only lead to an improved version of UDT 1.0 (no global optimization of the shared policy).
If this is what you meant, see more in subsection A.
If you mean "theoretical" in the sense of "can be generally ignored," I don't think that's true. As I said before, FDT is intended for the AI that implements CEV, or some other extremely ambitious alignment target. The CEV singleton (CEV standing in for most other ambitious alignment targets from here on) must be uncontrollable for bargaining reasons, relying only on programmed-in alignment for safety. CEV requires a lot of computing power, so the agent would presumably optimize itself and design new hardware. When combined with its total power over Humanity, this means that all safety properties must "tile," or be brought over when the decision theory designs a new agent. This requires designing a tiling agent before letting it go out of control. For a tiling agent to be rational, it must have a decision theory that tiles. Presumably, this would be a bounded rationality version of FDT that adapts to ever increasing computer power without re-writing, though no one knows how to do that. Effectively, this means that the agent's decision theory must be almost optimal from the beginning, before the AI starts to work on it alone.
Another important thing to know about singletons, is that if you have a singleton ignore something, that thing will be ignored forever. This follows from the singleton's tiling theorem, and the loss of all control to the singleton. This generally leads to us not wanting it to ignore much. Following these basic assumptions, a singleton is likely to think about a huge number of scenarios involving simulations, aliens, and even stranger things that are hard for a human to think coherently about. Quite possibly, none of these things have any bearing on reality, but we want the singleton to think of all the things we might have missed. Even so, just by thinking about these things, parts of these scenarios are caused to become predicted or simulated at some level of detail. The decision theory decides what to think about, and then these thoughts are fed back into the decision theory in order for more decisions to be made. This leads to the requirement that FDT must handle all sort of weird hypothetical situations, without either failing to take them seriously or doing something stupid. Relevant to your topic, it must be able to identify and classify agents in these predictions/simulations with no errors, and evaluate them within counterlogicals, or something close enough to counterlogicals if those are impossible to get working. More mundanely, any deficiency here can be exploited by any opponent that comes into causal contact with the AI.
Back to your actual question. Well, maybe you can create a decision theory somewhat similar to UDT 1.0 that lets you "quarantine" parts of your thinking and prevent your own attempts at sophisticated internal cross-checks, but is that actually the optimal kind of agent to be? It's sometimes possible to just win against predictors that are weak enough, but that will be the case less often if your decision theory isn't the absolute best possible. See subsection B for potentially better decision theories. In reality, you aren't given clean problem descriptions, you have to figure out what's going on yourself, so you should be quite sure actual counterlogicals can't be used instead, since otherwise other agents may try to predict you using more sophisticated methods than you expect. Similarly, if you don't consider internal cross-checking, other agents may use it to evade your predictions.
Subsection A:
When considering, UDT 1.1, you can't "philosophically" anything without the decision theory becoming an entirely different thing from the actual UDT 1.1. To cooperate with itself correctly, it must be, and is, fully specified down to the order in which it considers what decision to take[4]. Any stray "thought" the agent may have ruins the whole thing. Though the formalism is outdated, Tsvi B-T's UDT With Known Search Order (2014)[5] can be taken as a general guide to how total and detailed the requirements on the agent must be.
Note that even though normally it would be enough to define the iteration order as part of a pure function (we only care about the output, not the internals), in UDT 1.1, the procedure must also be exactly correct, down to the details of the source code (and the executable program's representation, though that's usually included when we say "source code" in decision theory). This is by construction in UDT 1.1, and in FDT it is the unsolved problem of determining decision-theoretic equality of the function that an arbitrary physical or computational structure "implements." Also, I have not seen a solution in FDT for the iteration order problem that shows an improvement over UDT 1.1. If someone knows of one, I want to see it. As I've said before, the more evidential agents "can" cooperate in some other situations, but it isn't reliable at all and isn't good enough for useful theorems.
Subsection B:
In the empirical case, all internal checking will pass if Omega is running a perfect simulation of the agent. All the agent can do is perform extremely complex internal operations to force Omega to run a perfect simulation, instead of doing something simpler. This can directly defeat weak Omegas, but in the standard thought experiment this won't let you trick Omega or perform a simulation escape due to that Omega's extreme power[6].
In the case of a mugging with a logical coin, e.g. your case with the digit of Pi, there may actually be a way of establishing an internal consistency check that may sometimes return a warning. If Omega is actually powerful, and figures out how to compute high quality counterlogicals, it will figure out how to make the agent think the checks passed without causing any new problems by doing that.
I don't know how well Omega can do at subverting checks, but I also don't know how good an agent's checks can be. Maybe they can be really good, such that finding a "nearby" agent doesn't work well enough to isolate the logical coin. This seems interesting to work on, but I think it would be difficult.
Or maybe the "straw" EDT that's practical enough for use in philosophy papers. I suspect that adding huge amounts of information and either the "tickle defense" or updatelessness would help, but still be worse than UDT for self-cooperation without spurious other-cooperation.
Good and Real by Garry Drescher (2006), Section 5.6, 2nd Paragraph
The second to last paragraph of https://www.lesswrong.com/posts/g8xh9R7RaNitKtkaa/explicit-optimization-of-global-strategy-fixing-a-bug-in
"You-complete" from The Ghost in the Quantum Turing Machine by Scott Aaronson
since the actual calculation of that fact takes more than one observer-moment, i.e. I can't verify it all at once
As far as I can tell, this is not allowed by UDT 1.1[1].
(UDT 1.1 also needs a hard coded search order (iteration order) in order to self-cooperate properly, an additional reason why I don't think humans can run it.)
As for other agents, Wei Dai's original UDT 1.0 is logically omniscient in the sense that it will never notice its "thoughts" (such as they are) taking any noticeable amount of time. It features an "intuition module" for mathematics, but this is always exactly the same and has no appreciable origin. The agent definition hardcodes the mathematical intuition subroutine along with all the agent's other code, with that including its prior and utility function. This means that any modification there results in an entirely different agent.
Any agent that takes time to think is going to be weaker than Wei Dai's original designs, since it will have to keep state around, state that can be deleted or tampered with, such as is effectively the case with The Absent-Minded Driver[2] if it doesn't have enough time at each intersection to recompute the plan from scratch. Plausibly, in Driver the policy may be fully forgotten (along with the agent's current location) or incorrectly remembered/corrupted, leading a computationally weak UDT 1.0 agent to do badly (I don't think computationally weak UDT 1.1 agents make much sense, at least without absurd ergodicity and memory assumptions).
Any additional state a weak agent may try to bring along would have to be re-checked at every point, making it useless, in the same sense that the "map" in the map explanation of UDT[3] must be recalculated from scratch at each time, otherwise it could be tampered with just like the agent's memory and sense data.
(Note that the unbounded agent reaches, at each step, the same map as the one it reaches at every other step (if it even has such a map at any point), so the prior is still all the epistemics it has.)
Slow agents would need more forgiving fairness criteria, maybe a list of previously used cryptographic public keys that the environment isn't able to tamper with (only delete from). Assuming some things in computer science, the agent could check its previous state faster than just re-calculating it all. If ordering matters, Merkle chains can be used by the agent, with results written into its standard, tamper-vulnerable memory. Note that in each step, the agent would need to generate a private key, use it to sign what it thought about that step, add the corresponding public key to its public key list, and then delete the private key. The private keys would need to be opaque to the environment during the step, otherwise none of this accomplishes anything. Presumably, the agent would be required to have only one chance to generate a key pair each step, and only have the option to add that public key into the key list, otherwise it could try to brute-force a pattern into the public key it adds. Preventing this isn't particularly realistic, but it's a good research direction to avoid a trivial tiling result.
(If the environment could tamper with the agent's public key list, it could sign anything it wanted in the agent's tampered-with memory using newly generated private keys and then write the corresponding public keys into the agent's key list. Proof-of-work is a no-go because the environment is assumed to be stronger than the agent.)
I'm gonna disagree with you and Scott Garrabrant.
I think all your anticipated metaphysical stances are too much. In the blackmail thought experiment, the correct metaphysics is not an expansive one that reifies impossible situations. The correct analysis of what's going on is "this is a thought experiment asking about what you would do in a hypothetical situation."
Or take a more realistic situation where the blackmailer acts with some probability. You might ask similar questions about one who doesn't pay the blackmail: "Are they imagining this paying off in alternate realities? Or relocating their 'selves' to the platonic realm? Or think that they're changing the past?" But it's entirely possible for someone to have "normal" metaphysical views - they're not thinking about alternate realities or platonic realms - and they simply evaluate the goodness of actions in a "weird" way - e.g. what makes actions good is that they're part of the winningest strategy.
The metaphysics isn't an inherent part of decision theory, it's a (contingent) feature of humans making arguments about decision theory. Metaphysics comes in if you're going to take an actual human (who has a mish-mash of different intuitions) and argue them into doing one thing or another - different policies will comport with and more easily be argued for with different human intuitions, many of them metaphysical.
they simply evaluate the goodness of actions in a "weird" way - e.g. what makes actions good is that they're part of the winningest strategy.
So if I'm an AI that was just created and is instantly blackmailed, in what sense *exactly* is not paying the "winningest strategy"? It will lead to worse outcomes for (indexical) you.
(I think if you dig down into why you think this makes sense, there will be a metaphysical intuition that others don't share)
Let's say the strategy is the "winningest" in the relevant sense because given the AI's model of the world and a particular notion of changing the AI's modeled strategy and then using the model to predict the outcomes, the strategy of not paying blackmail has the best modeled outcomes (arguendo).
"Aha!", you may say, "Picking what possibilities to model, and what modeled possibilities to care about, is basically what the word "real" does in normal language, i.e. your AI has off the bat dome some metaphysical stuff."
And this is a good point. But I think my AI responds "But I feel like I do other stuff with the notion of "real" that I don't do when modeling possibilities. Like, real stuff controls my expectations about what I'm actually going to see, and I can go interact with it (or can have interacted with it in my actual past), and I think I'd answer metaethical questions in pretty much the same way as CDT-bot. To me, it feels more like the modeling different counterfactual presents is more like a shorthand for modeling the reasoning of Omega (or other copies of myself or whatever), who definitely exists and is standing over there. If I thought Omega was doing different reasoning, when I considered changing strategies I'd end up modeling different states of the world. This doesn't necessarily mean I'm considering being simulated by Omega, either (though I'd believe that if I had reason to) - my algorithm for finding the winningest strategy computes the same counterfactual no matter how Omega predicts me, as long as Omega is good at it."
I dispute needing any notable metaphysics. Metaphysics for decision theory are what people often move to as intuition pumps and then think that is where their confusion lies ("of course I can't change the past!"). Part of the reason we've moved away from that, is yes that people experienced argue about it less, but also because it isn't needed as much.
Probability theory is distinct from the Many Worlds Interpretation or other branching universes, as probability theory of beliefs is epistemic, while MWI is an intended to be literal physical interpretation. There's a linking here but they aren't the same thing. Possible worlds and branches are a useful intuition and mental model but eventually you move to Probability theory as a thing in of itself. There are people who are confused by probability theory and ask whether we actually think this deterministic thing is random, or whether these events that can't be ran multiple times can meaningfully be assigned a probability, but these are confusions about Probability-belief as a thing with a nature in of itself.
Similarly, CDT, we do not ask whether CDT considers counterpossibles as real. CDT, in a deterministic universe, has a definite answer as to what actions it takes by how it cuts up the graph, but it is yet acting like it can vary its decision! CDT is not endorsing platonism just because it has causal graphs. CDT makes realities impossible by its very action, that is what taking an action "does", choosing a possibility by virtue of how your decision procedure ranks actions.
Of course, this leads directly into your footnote 5.
Okay, sure, but they don’t imagine that they can make their own reality impossible - only FDT does that (academic decision theories only imagine inconsistencies insofar as they can actually make them consistent after all, by choosing that option). That, to me, is a distinct stance - a much more metaphysical one. It doesn’t make sense in a normal worldview.
CDT in a deterministic world evaluates possible-worlds wherein you take an act you won't take, and then crosses them out by choosing an action. FDT evaluates possible-worlds that are inconsistent with your experience, some of which are in the past, and then crosses them out by choosing your policy. The same sort of removal occurs, just in different locations. Sure, FDT crosses out worlds that contain your current position, but that isn't a question of ontology. It is a question of if we should evaluate from an indexical or an updateless prior. There's the question of whether we're considering an action as first and foremost the thing to consider rational or not, or whether we want to consider the policy as more primary. Aggressive precommittment like that which results in Son-of-CDT sure looks like trying to decide policy ahead of time via constraining yourself, so why not just decide policy?
We can go poke at extensive form game theory. A strategy there is defined everywhere, and which path is the answer (and thus made real) is determined by what your strategy would do on other paths. An agent that finds itself at some impossible node has various ways of thinking about their scenario, such as trembling hand arguments, belief revision, and backward induction. Though this is more third-person, but, well, that is important.
CDT holds its action node loose, as if it is uncertain, and then via that consideration is "making" those other outcomes inconsistent. That's what taking an action does, and was the point of Soares' Decisions are for making bad outcomes inconsistent. But, similar to "changing the past" you aren't really altering those other outcomes in some weird esoteric branch interference, rather you're influencing outcomes that decide which realities you experience towards ones with higher utility. This isn't metaphysics, even if it is a mental reframing, which is my view of the supermajority of FDT. The confusions are presuming the interpretation/intuition-generators/etc are "real" rather than useful. FDT, then, is an alternative notion of where to cut the causal graph and then we see that this generates better results than typical CDT.
Pseudo simulation, footnote 5.
Your initial paragraph supposes a simple decision function, and thus Omega only needs to simulate a fragment of your mind that is that function. Then your second paragraph talks about blackmail and having access to sensory experience and that you thus have to be agnostic about experiencing.
Firstly, an FDT agent can have very high credence that it is actually experiencing. However, what it conditions on for which policy to follow avoids that. Those are distinct epistemic vs decision-algorithm, the decision-algorithm can explicitly preclude certain epistemic elements from influencing it while still "believing truthfully". This is very similar to how a smart CDT agent (without time to commit) in a prisoner's dilemma against another CDT agent knows that it will output the same answer as the other agent, but how it arranges the causal graph for decision making is what decides the outcome.
Secondly, I think you're maybe sliding between "simple decision function" and "blackmail with sensory experience". However, you don't have to be agnostic about sensory experience for your simple decision function either. It can be that my decision procedure entangles with my sensory observations or thoughts. The usual proposition of flipping a coin to decide. Well, then Omega has to predict the coin to get a certain accuracy. If you had a fundamentally random coin, then you've forced Omega to be ~50% accurate and you've refuted problem statements saying Omega was >50% accurate. Similarly, just as with the coin, I can entangle some simple decision procedure with observations of my thoughts- as I do automatically. Omega, to get some required accuracy, needs more precision and more simulation of my brain beyond a simple core. Then potentially further reality and experience. Agnosticism for that is relatively simple. "Huh, am I in a simulation or not?" is a relatively mundane sort of thing to ask with good enough simulations.
Descartes' implication, "I think therefore I am." is perfectly consistent with your senses being bedevilled (as was the original thought experiment!), your mind being much smaller or weaker than you thought, your rationality being limited, or more exotic things. Beyond this, being uncertain whether you have sense experience is not esoteric. My AmISensingThings complex mental heuristic can be stubbed with something that produces "yes, you are", just as my sense for grass can be fed with a reality sim.
An alternative way of framing this beyond such is that even in low-fidelity scenarios, while perhaps "you" can tell you're not in a simulation, your simplified decision algorithm can't and also isn't coherent enough to do "I think". Thus your decision is replicated across simulation and reality regardless.
Regardless, I don't think rejecting or weakening Descartes' is that radical if we aren't thinking of minds as unpredictable unitary things, but I also simply disagree that we need to reject it.
Your actions do not have to agree with what you believe exists. Just as CDT takes actions it can even believe in some meaningful sense are wrong. An FDT agent is not unsure whether or not it exists, precisely, rather strategies are evaluated from a standpoint which doesn't privilege indexicality. (Similar, once more, to the joke "It is a 50/50 chance whether I win or I fail, it either happens or it doesn't!", which is a similar conflation of layers though different in sort)
But, a sample size of three given the framing above makes me simply disagree with them. Whether because they're evaluating a different sort of argument, or direct disagreement about FDT or how simulations "work". Uncertainty about being in a sim is relatively minor, just as you have some fundamental uncertainty that you're a Boltzmann brain or are fundamentally confused in some other manner.
To go back a bit, to footnote 2,
Assuming the predictor isn’t predicting you by simulating you with high fidelity (since otherwise you can just say that you might be in the simulation). This is reasonable to postulate because FDT is also supposed to change your action with very bad prediction on the part of Omega, e.g. with only a 60% success rate.
For one, I think this is a odd thing to presume when you're talking about metaphysical assumptions! Simulations are a classic justification for why this makes sense. You go into this some with the pseudo-simulation, but I don't think you really defeat a weak metaphysics of simulation. Your pseudo-simulation arguments I've already objected to directly, and so my belief is that simulations even with lower fidelity basically survive just fine and thus should be considered directly.
But, moving back once more, paired with footnote 5:
Then the response is often “but the impossible realities never actually happen, it’s purely concentrated in counterfactuals”. Well, sure, but the consideration is always present in this framing. If Omega is imperfect, you’re making decisions by imagining an x% chance that you’re making this reality impossible with your action (and that something else actually happened).
We can spin this around. For one, you have an x% belief that the predictor made a mistake. If there's a 60% accurate simulation, and I observe the blackmail, then I am either in the simulation where I am being tested, or am I in literal reality where the predictor made a bad prediction if I was sort of agent to not pay in the simulation. If it is "not high fidelity" then your rich observations tell you this is reality, thus you're in the 40% outcome. Your policy, as an FDT agent, has you take the don't pay action because by the updateless prior you get overall better outcomes that way.
The "impossible world" doesn't really occur here at all! This is just unfortunate world where you stick with what your decision procedure said was the best action due to the subjunctive dependency on your action.
Really, I think your high fidelity vs low fidelity distinction is confused. A charitable reinterpretation would be that it is a metaphysical question of "what does it mean to be instantiated in a simulation", though I've argued against your pseudo-simulation argument. Another interpretation is that of whether it is observable that you are in a simulation. But, if you know whether or not you're in a simulation, obviously you exploit that! And, FDT very happily exploits simulations its decision algorithm can differentiate. Imperfect blackmailer? Doesn't pay in sim, pays in reality. Newcomb's? One boxes in sim, two boxes in reality. Etc. It takes the more utility. Though I think you're basically using it for "you know you're experiencing reality but can still be predicted" which to me is basically just presuming non-simulation entirely, which is Weird.
Overall your post somewhat stuffs the "meat" of the disagreement into the footnotes and spends much of the post interpreting everyone as talking about metaphysics even if they think it is intuition, true but unrelated to the decision theory part (similar to how Lewis was a modal realist and came up with a CDT's counterfactual notion, but while it may make one handle such possibilities better, it isn't necessarily meaning they're the Same Thing), and so on.
Generally, my view is that people invoking metaphysics as their objection are confused. Decision theory operates in first person so to speak, but our world models operate in third person. Like how CDT and FDT intervene on their graph, or extensive form game theory considers a high up path-level rather than from the agent's perspective precisely. The translation between those two has logical and conceptual issues! A huge amount of MIRI work, agent foundations and such, was about conceptually and formally pinning those notions down without the decision theories exploding because they're in an environment with predictors or can predict their own action in a deterministic environment (ex: troll bridge problem). To me it is no surprise that there is confusion from people understanding the works, and even from people who read about this a lot. This philosophical uncertainty leads to using intuition pumps, examples, and such that may sound of metaphysical significance but are of the same sort of direction as thinking of probability theory as possible-worlds, it is just that we do not yet have the best conceptual underpinning. This movement from first person understanding to third person world model is also where metaphysics and ontology seem to come in to bite, but that is only if you insist on the mental model being the same as the model of how the world is actually structured, which is an easy confusion to make but is very much distinct. The causal graphs of CDT are the sort of structure that is made to simplify (and skip over the work of defining meaningful physical counterfactuals) the world model as a global third person view.
This is a messy reply but I already rewrote it 2.5 times, haha. I do think there's a, not quite conflation, but a meaningful social and practical distinction to be made between "universal metaphysics" and "identity metaphysics". The former requires stronger arguments, but the latter is generally much weaker. I don't believe FDT requires the former, but that it serves as good intuition pumps. For the latter, I think there's more legitimate ground to dispute there, but I still think many disputes are confusions. Such as "I believe in free will, so I don't believe in perfect predictors", wherein the confusion is presuming we need perfect predictors to justify FDT, and that all you need there is some decent amount of prediction of decisions which is much weaker and much more commonly believed to be possible even if perhaps not physically feasible. Similarly, FDT is much easier to conceptualize with more computationalist theory of identity/mind, but is mostly not required; yet it does lead to pushback due to considering algorithms more primitively.
I also think I disagree with your son of CDT. In justification for FDT, sure everyday commitments aren't necessarily FDT, but are often used in terms of "we use predictions of other people's actions to decide what actions to take" which is very much in spirit of FDT and closer than how people conceptualize CDT or commitments. Though there's of course some problems that feel more natural in either, and some in neither even if related (ex: virtues). But I don't think this is central beyond pushing people away from typical naive CDT.
I generally avoided Bentham's posts because I've seen him argue against FDT some years ago and on X he doesn't seem to have gotten substantially better so it never felt worthwhile (though I guess I got convinced to argue against his recent post on SSA implying theism...) But, looking at the top comments, it is pretty direct. Mark saying "It is fine to give the wrong answer in a probabilistic sense" which is explicitly about the mechanics and justification for FDT. Vaniver is talking about concrete justifications for/against concrete payoffs. Etc. I don't think this is dominantly semantics. Sure, it isn't metaphysics, but generally it isn't believed metaphysics is the problem. I think you're over interpreting by your issue with lack of metaphysics to think most of it is missing the point.
Separately, the definition of "rationality" does matter. If you defined rationality as CDT-rational (as is often implicit in decision theory texts) then FDT will be worse by that definition. If you're just saying "use rationality for near-term naive optimal decision making" as you do in your commitment theory post, then we can do that, but I also think it is an active distortion of what humans mean by rational in order to waffle between "well, it doesn't matter, we're just making ourselves be irrational to get more utility". That is, why add more epicycles with commitment theory over decision theory? Why not just be exact with what is being constructed? Collapse these layers and extract The Right Answer, and then move on to studying bounded agents, embedded agents, etc. to make our understanding of the nature of decisions ever more complete. (That is, the semantics matter because words mean something, and also because they matter to interlocutors, if they consider that a relevant "resolution" of the problem at hand rather than a minor fiddly bit to ignore)
I didn't have the motivation to read this deeply, but I think my reply is just
What you’re really saying (like Nate Soares here, ctrl-F “metaphysically”) is “You need to consider crazy metaphysics (imagining you’re in a hypothetical impossible reality) to make decisions in a reasonable way”, and like… yes, exactly. That’s de facto a metaphysical stance. “It makes things easier if we imagine this” is not a decisive argument for a metaphysical claim that permits you to leave it unstated as if it’s obvious.
You're being uncooperative by not calling it metaphysics, since, under the normal meaning of metaphysics, it just is (moreso than normal decision theory).
(I probably won't reply again if you reply is super long, btw)
Labeling claims "metaphysical" is also a way of functionally dismissing their content (while still being studious about it), adopting a 19th-century anthropologist's or a journalist's outlook, talking about the claims made by these other creatures as "stances", while treating the prospect of adopting them as actually meaningful or true in the mundane ways as utterly alien. Not because it's something that seems like a bad idea on its merits, but just because it's fundamentally not the kind of thing that's done.
they don’t imagine that they can make their own reality impossible - only FDT does that ... That, to me, is a distinct stance
Even if something doesn't "exist", but has a legible definition, I can still ask about the specific properties of the thing that doesn't exist, and so it's mostly not relevant whether it "exists". It's similarly not relevant whether the present reality/situation "exists" when considering what takes place here and how to navigate it. In particular, it's difficult to perceive/determine if some possibility "exists", whether it's the present situation or some alternative. But the weight of existence (and the ways of influencing it) can be decision-relevant, possibilities with more existence matter more. As a result, insisting that the present reality/situation always "exists" leads to systematic errors in decision making.
You are in a branch of reality where you got blackmailed, but other branches exist, and you can make it so you don’t get blackmailed there.
You are already not blackmailed there, that is the defining property of the situations where you don't get blackmailed, as opposed to the situations like the present ones where you do. If you do get blackmailed there, then you can't actually make it so you don't get blackmailed there, that's the thing.
The move is to make those no-blackmail branches exist more, and to make your own faulty branches exist less; rather than to change those other branches to be less blackmail-afflicted, or to change your own branches. This is the same kind of move that determines the future with the decisions made in the past in a deterministic world.
Labeling claims "metaphysical" is also a way of functionally dismissing their content
I think you are pattern-matching me to a different kind of person. As I say, I believe in a lot of these crazy metaphysics myself! So I am obviously not dismissing them! I would just like more clarity in the discourse.
The move is to make those no-blackmail branches exist more, and to make your own faulty branches exist less
I agree, I was being short there for ease of exposition. (EDIT: Actually, idk, maybe both framings work, not sure)
It's similarly not relevant whether the present reality/situation "exists" when considering what takes place here and how to navigate it. In particular, it's difficult to perceive/determine if some possibility "exists", whether it's the present situation or some alternative. But the weight of existence (and the ways of influencing it) can be decision-relevant, possibilities with more existence matter more. As a result, insisting that the present reality/situation always "exists" leads to systematic errors in decision making.
Yeah, I mean, I don't really know what to do other than to gesture at this and say "this is obviously a metaphysical stance". Or let's actually just say this - most people will be like "wait, what?" at this and disagree. So no matter if you label it as metaphysical or not, this is the actual crux, right? So let's try to have discussions with people about this directly, no? And not about whether "FDT gets more utility" or not (which is just downstream of the crux).
So I am obviously not dismissing them! I would just like more clarity in the discourse.
I'm not claiming that you are. I'm claiming that the clarity you would like more of can be poisonous, at least if you spontaneously bring it up a lot. Whether you personally are resistant to it is not relevant to what I'm saying.
I don't really know what to do other than to gesture at this and say "this is obviously a metaphysical stance"
There are connotations to "metaphysical stance", I'm pointing out the issues with the 19th-century anthropologist connotations. In some other meaning-aspects there's obviously no disagreement, but they don't motivate writing posts like this one and insisting on talking about this as a "metaphysical stance" (beyond perhaps saying that in some sense it is, and moving on to a more substantive discussion).
most people will be like "wait, what?" at this and disagree
Discussing the disagreement could then be meaningful.
So no matter if you label it as metaphysical or not, this is the actual crux, right?
What is? It's indeed the case that it doesn't matter if we label it as metaphysical or not. I'm not insisting that we label it as not-metaphysical (it does seem metaphysical). I'm pushing back against your insistence to keep bringing up its labeling as metaphysical where it's clearly not relevant, after you've already established that the label is correct in the obvious senses.
Huh, interesting - so are you seeing yourself as disagreeing with normal LW epistemic norms of clearly stating your position? It just seems oddly deceptive to not say clearly what you believe because it might weird the other person out.
Discussing the disagreement could then be meaningful.
Do you think the discussion under e.g. the Bentham's Bulldog's post I cite is productive? Or do you think it would've been more productive if they had realized that there's a metaphysical disagreement?
epistemic norms of clearly stating your position
No, I'm not disagreeing with this. I did mention "perhaps saying that in some sense it is, and moving on" and "after you've already established that the label is correct in the obvious senses". The issue is with restating it where it's not relevant as if it is, perhaps even in lieu of engaging on substance, and the connotational externalities of that.
Where you said "this is the actual crux, right?", I don't understand what you are talking about. I asked, you didn't clarify. I dread that you meant that whether it's a "metaphysical stance" is a crux, which it clearly isn't, but I'm unable to tell what you meant.
I believe in a lot of these crazy metaphysics myself!
The primarily alarming thing is that you are stating claims in subtly wrong ways, and when I point it out, you keep agreeing (as you also did in this thread), without engaging on substance, as if there's no difference. I have no way of knowing if you do see the points I'm making, since you don't signal that you do (agreeing without clarification doesn't signal that; it's possible that you do, but I still don't know).
So to be clear, I do think it's possible you're affected by the issues with connotations of "metaphysical stance", I just wasn't claiming that then (and still don't, beyond it being a possibility). This could manifest in accepting many variations on ideas and not discriminating among them in a load-bearing way. And then nodding along when an aboriginal corrects your descriptions of their beliefs. Treating your own beliefs in the same way doesn't really refute the issue.
Ah, sorry I missed your question. Very late over here. No, I definitely didn't mean to say that whether to call it metaphysical is the crux between us. I meant that those metaphysical beliefs (themselves) are often the crux between anti-FDTers and FDTers. I mean, that's one of my core claims in the post, I say that explicitly. So do you disagree with that?
I just think, on the current margins, in the popular discussions like the one I linked under Bentham's Bulldog's post, we very clearly need more acknowledgement of the metaphysical disagreements that people probably have (since there is basically no acknowledgement currently). So that's what I'm trying to push for.
The primarily alarming thing is that you are stating claims in subtly wrong ways, and when I point it out, you keep agreeing (as you also did in this thread), without engaging on substance, as if there's no difference. I have no way of knowing if you do see the points I'm making, since you don't signal that you do (agreeing without clarification doesn't signal that; it's possible that you do, but I still don't know).
This feels a bit exhausting to get into, maybe I will tomorrow morning - just letting you know that I saw it.
(I generally oppose the norm of there being any obligation to respond in any detail or at all, and regret the possibility that my actions might be feeding it. The intent was to explain the background of what got me argumentative here. I don't really endorse how I'm handling my side in this thread. This is a sufficiently unusual situation that I'm not sure what to do, other than the simpler out of staying silent.)
No one has metaphysics problem when we talk about classical game theory of zero-sum games, and this is not because classical game theory doesn't have weird situations warranting metaphysics problem. For example, in poker you can have position where Nash equilibrium strategy is to fold with some probability, even if raise has better expected value from the purely causal perspective, because if you would predictably consider "always raise" the best strategy in this position, your counterparty would adjust their play in a way that would leave you with less utility. Imagine that you roll the dice and it tells you to fold. I don't see actual difference between this situation and not paying in blackmail. You can say something like "I fold to help my counterfactual versions" but this is interpretation, not actual meaning of your move, which is "this move is a part of Nash equilibrium strategy for this game". See also counterfactual mugging poker.
I think that people just have better intuitions about zero-sum games? It is intuitive that you should sacrifice some of your utility at war where you won't get to see final benefits, while it's less intuitive that you should burn some of your utility to disincentivize blackmail? I think the second is also first-order intuitive for ordinary person, but it is kind of "emotionally intuitive", while standard game theory has "cold calculating" vibes and this creates dissonance between solutions.
No, all the examples you list are ones where you can (with causal reasoning) self-modify yourself / precommit to the strategy and/or can expect to have causal benefits if you get known as someone who follows it. This is exactly what I meant when I referred to "confusion that FDT is necessary for everyday psychological precommitments[6], confusion that everyday psychological precommitments are sufficient to become an FDT agent[7]". This is all doable with normal decision theory. You need something like the instant blackmail scenario I describe to tease apart FDT and normal decision theory, and I completely disagree that that's intuitively the same situation.
The whole causal line of reasoning explodes in the scenario "one shot game against alien which flies away at lightspeed after game, you are forbidden to tell anyone about how game went in details and can only take your winnings". For alien to think that you would precommit they need to do FDT over decision algorithms which consider precommitment.
Yes, that's what I'm saying. But your poker example and your war example are not like that, so why talk about them?
I'm saying that if you were to play one shot poker with no reputation consequences against alien who never met humans before and will never meet them after and you asked classical game theorist how to play it, they would answer "just play Nash equilibrium strategy for poker". If after that you asked "what to do if I rolled the dice to randomize and dice came up 'fold' in situation where folding means foregoing all expected winnings which would happen if dice came up 'raise'" they would answer "just fold in this case". These recommendations would be considered pretty much non-controversial and not requiring weird metaphysics by academic decision theorists, because classical game theory of zero-sum games is pretty much non-controversial, even if they are one shot and include aliens. Do you disagree?
I have no idea what academics would say, but insofar as academics would say that, they would just be wrong - overindexing on the familiar vibe of the situation, without realizing that this context is different. Poker is usually an iterated game, after all.
If we're going full meta, these sorts of questions are always happening in the sense of being posed as hypotheticals, rather than literally happening to you, so the only real effect of your answer is leaking information about how you would behave. This doesn't require alternate realities or retrocausality.
Also, I don't know if you count this as alternate realities, but even if the hypothetical were truly appearing to happen, it could just be a simulation spun up by the predictor.
This is useful. As someone not familiar with decision theory, Newcomb's problem was super confusing to me when I was asked to answer it (or rather, the point it).
Stupidly enough, I kept asking whether I should assume Omega is aligned or misaligned XD, "does it want me to get the lesser amount of money so that I don't donate it to AI Safety?".
Otherwise, where's the choice? If I'm told that if I choose to take both boxes, I WILL not get the million, the whyyyy would I choose to do that?
And if it's a perfect predictor of me, how does my answer even matter?
It's like if my mum tells me: "If you tidy up your room, I'll give you $100. If I find out you didn't, you're grounded". This looks like a choice but, as everyone knows, is not (the mum here is also a perfect predictor of me).
Now, if there was clarity on how I can still exercise control over my circumstances, I guess it'd be clearer that it's a genuine choice.
EDIT: I'm explaining my confusion, not claiming I'm right. Would be useful if others explained their disagreement so that I learn 🙏
[Epistemic status: rant]
There’s something that annoys me about the reoccurring debates on decision theory in this corner of the internet.
Take a simple blackmail scenario:
Let’s say we want to argue for the Functional Decision Theory (FDT) answer that you shouldn’t pay. It was originally motivated by the observation that such agents seem to achieve higher utility (“rationality is about winning”), since they don’t get blackmailed in the first place. But that only leads to making yourself into such an agent in advance (which everybody generally agrees you should do[1]) - it doesn’t clearly apply when you are already being blackmailed and have never thought about the question before, or if you are an AI who was just created and is instantly blackmailed before being able to self-modify or make precommitments.
I see three broad ways to make FDT’s recommendation make sense in that case[2]:
These are all metaphysical claims - therefore, you need metaphysics to make FDT make sense.
(I understand many people will disagree with me on this - I think it’s a widespread misconception. I address some counterarguments in this footnote[5])
And I want to be absolutely clear here - I am very sympathetic to these claims! In my personal opinion, something in the vicinity is actually true. But… they are metaphysical claims, and ones that most people probably don’t immediately accept.
So let’s say that someone only knows this fact about the topic, and and starts reading what people are saying about it online. Imagine their surprise at finding that the back-and-forths between FDTers and anti-FDTers are largely not about metaphysics. Huh?
What the hell is going on?
I understand that metaphysics is really annoying and hard to think about, but… guys, I think this is your actual crux!
Take the comment section under Bentham’s Bulldog’s recent post about FDT. There are such luminaries as Scott Alexander and Stuart Armstrong chiming in. Yet barely anyone is bringing up metaphysics - instead the discussions circle around semantic disagreements about what the words “rational” and “decision theory” should mean, confusion that FDT is necessary for everyday psychological precommitments[6], confusion that everyday psychological precommitments are sufficient to become an FDT agent[7], confusion that the observation that FDT agents get more utility is enough to fully justify FDT[8], and so on and so forth.
No! What are you doing? You have one very concrete disagreement - talk about that one! There is only one spot where FDT and a normal worldview diverge, and it’s the kind of scenario above, where you have no time to self-modify/precommit. For everything else, normal decision theory is sufficient![9]
And if you talk about that kind of scenario, you’ll be able to actually get to the bottom of your disagreements - which is metaphysics.
Concretely, my message to FDT proponents is this: Start being clear to yourself, and to others, about your metaphysical stances. You are confusing everyone by leaving them implicit.
I have a suspicion that a substantial amount of the resistance to FDT is that people can tell that you’re doing something fucky. They can tell that FDT doesn’t really make sense without additional metaphysical commitments, and that you’re just pretending it does. But if you were to say “you might just be inside the abstract computation” or whatever, I think people would agree that FDT makes sense given that assumption. And then you can actually talk about whether those metaphysics are reasonable and how to handle that, instead of talking past each other.
Old-school LessWrong was more explicit about this stuff. There was plenty of open talk, back when UDT was the flagship decision theory, about crazy stuff like everything existing, that we can decide what gets to “exist”, supernatural voices from the sky, and so on. My guess is that a lot of those veterans from back then are aware of what I’m saying here, and have just gotten tired of talking about metaphysics and don’t chime in much anymore.
For example, Paul Christiano:
Paul Christiano is not the type of writer to say something that wild if it’s not necessary for what he’s arguing for! It’s really not as simple as some of you think!
I think something has gotten lost somewhere, in the transition to the more narrow, academic framing of FDT. Some of the newer arrivals to the space didn’t get the message that there is an underlying metaphysical question, and are going around thinking that it’s a normal decision theory like any other. My past self from three years ago definitely didn’t fully realize what it was arguing for when it was arguing for FDT.
I mean, let’s just take a look at these apparent facts about the paper (emphasis mine):
The FDT paper is close to my heart and I respect the authors a lot, but… if this was the intention, why are you writing a pure decision theory paper? That’s a metaphysical claim. And there is no mention of metaphysical commitments in the paper at all - in fact, in the conclusion, it explicitly claims that it all works without any metaphysics (which just isn’t true, in my opinion - maybe on a formal level, but not on a philosophical level). That seems almost deceptive - no wonder that this confuses everyone, including the referees of the paper.
Wei Dai (not claiming that he would fully agree with my take here):
Yeah - I could imagine another version of the paper, something like “computationalism applied to decision theory leads to weird metaphysical challenges like seemingly changing the past and difficult formal problems like logical counterfactuals if we take it seriously. But we should - because it’s very elegant and solves a lot of issues (e.g. dynamic inconsistency, coordination with similar agents), and computationalism is a widespread position.”[10]
At this point, I see the failure of academia to come up with FDT before LessWrong partially as the same academic failure mode, of too little interdisciplinarity, that caused the FDT paper authors to feel pressure to stay unnaturally agnostic about metaphysics. In retrospect, it’s obvious that the concepts of logical causation and updatelessness have been grasped at in academia for a long time[11], but that they just never fully got there - because it needs metaphysics to actually make sense. But decision theory and metaphysics are different subfields of philosophy, so bringing them together doesn’t come naturally. Grand unifying theories are institutionally difficult, and often actively disincentivized. LessWrongers, on the other hand, were not academics, and nonconformist enough to put together the obvious pieces lying around[12].
But then, let’s be clear about what we’re actually doing! Let’s be proud of the metaphysics again! Let’s be proud of being generalist interdisciplinary thinkers!
Getting to this point via the heuristic of “rationality is about winning” was perfectly valid. We did a great job! We outdid the academics! But that just means we get to argue about metaphysics now[13]. We leveled up.
Finally, I’d like to end on a positive note - Nate Soares in Notes on “Can you control the past” is who I’ve seen make things explicit the most in the last few years (although still not as much as I would like). For example, he says:
Then Joe Carlsmith doesn’t press him any further on this, presumably in part, again, because metaphysics seems too annoying to get into. But… yes! More of this please!
Thanks to Hein de Haan and @eigengender for valuable discussion of these ideas.
even by academic decision theorists - although they might not fully realize the implications, e.g. that you need to become the kind of person who would choose to burn to death in Bomb*.
Assuming the predictor isn’t predicting you by simulating you with high fidelity (since otherwise you can just say that you might be in the simulation). This is reasonable to postulate because FDT is also supposed to change your action with very bad prediction on the part of Omega, e.g. with only a 60% success rate.
“I’m not sure if I’m thinking about worlds that don’t exist, or if it’s us who don’t exist and there is some real world somewhere thinking about us.”
Note that it’s also technically possible to justify this in a purely axiological way, that you just also care about alternate realities, without making a metaphysical claim as to whether they’re real. But I don’t buy this, this is just a formal trick - you wouldn’t care about something that you don’t, on a gut level, believe is “real” in some way.
It’s a little ambiguous, but from the paper: “What’s remarkable about this line of reasoning is that even in the case where Fiona has observed that box B is full, when she envisions two-boxing, she envisions a scenario where she instead (with high probability) sees that the box is empty. In words, she reasons: “The thoughts I’m currently thinking are the decision procedure that I run upon seeing a full box. This procedure is being predicted by the predictor, and (maybe) implemented by my body. If it outputs onebox, the box is likely full and my brain implements this procedure so I take one box. If instead it outputs twobox, the box is likely empty and my brain does not implement this procedure (because I will be shown an empty box). Thus, if this procedure outputs onebox then I’m likely to keep $1,000,000; whereas if it outputs twobox I’m likely to get only $1,000. Outputting onebox leads to better outcomes, so this decision procedure hereby outputs onebox.”
(This footnote is very long, so probably actually click on it instead of hovering over it). Three out of three non-experts that I spoke to about this had this misconception. It makes me think that the metaphysics-agnostic framing in the FDT paper might genuinely have had bad effects here. Anyway, here are some counterarguments and alternative framings I’ve encountered, and my responses:
Then the response is often “but the impossible realities never actually happen, it’s purely concentrated in counterfactuals”. Well, sure, but the consideration is always present in this framing. If Omega is imperfect, you’re making decisions by imagining an x% chance that you’re making this reality impossible with your action (and that something else actually happened). What you’re really saying (like Nate Soares here, ctrl-F “metaphysically”) is “You need to consider crazy metaphysics (imagining you’re in a hypothetical impossible reality) to make decisions in a reasonable way”, and like… yes, exactly. That’s de facto a metaphysical stance. “It makes things easier if we imagine this” is not a decisive argument for a metaphysical claim that permits you to leave it unstated as if it’s obvious. (This line of argument sometimes feels like people are really trying their absolute hardest to pretend they’re not taking unusual metaphysical stances, even though they are.)
If we assume I’m actually being blackmailed (like the problem statement says), I have access to my full sense experience. So naively, I have very high credence in the belief that my sense experience exists / is actually instantiated somewhere. But the reasoning above is predicated on thinking that it might not be. In fact, to make the math work out the same way, you have to be completely agnostic about whether your sense experience is happening, as if you have no information about it whatsoever (since even a little bit of updating on it would bias you towards paying more than FDT recommends). To be clear, not agnostic about whether your sense experience is accurate about the external world - agnostic about whether it’s happening at all. That’s not technically metaphysical, I suppose, but it’s a very radical epistemological stance that most people would still disagree with. (it rejects Descartes!) Also, it’s just obviously wrong, and you don’t actually believe this. The people I’ve talked to have all also denied that this kind of radical skepticism is needed to justify FDT - I’m including it here more for completeness.
(Note that I am not saying here that you can fool Omega by taking your full sense experience into account in your decision - I’m purely analyzing what this reasoning actually looks like, and whether it makes sense as stated, or whether there is an additional claim needed to reach the correct answer that you shouldn’t pay.)
Again, then you might say “but this never actually happens”, and, no, even when Omega is imperfect, in this framing you have an x% credence that your sense experience isn’t actually happening.
A few other, less common counterarguments:
No - you can decide to e.g. become a virtuous, reliable agent on entirely causal reasoning.
No - it only leads to son-of-CDT (which, to be fair, is pretty powerful, as I discuss here, but fails on the blackmail scenario above).
It’s not, as we see in the blackmail scenario above.
Note for example, that XOR Blackmail - in the way it’s usually understood, which is generally taken to disprove EDT - is such a problem. So it’s really more of a counterexample to updatefulness, than to EDT-style counterfactuals.
I am generally using FDT to mean updatelessness here - sorry for being imprecise (although FDT is supposed to be “an umbrella-term for UDT-ish approaches to decision theory”, so it’s pretty close. It was originally meant to be agnostic between EDT- and CDT-style counterfactuals, so I usually take the disagreement about those to not be central. And it probably doesn’t matter anyway, since “CDT=EDT?”.and updatelessness is the more important proposal, in my view.)
(Joe Carlsmith hits some of these notes in Can you control the past?)
But I can imagine other framings too.
There’s shades of logical causation in the three papers here, and shades of updatelessness in Fisher’s and Gauthier’s disposition-based decision theories, Parfit’s “rational irrationality”, McClennen’s “resolute choice” (“it is rational to follow through on plans even when they become locally harmful”), Bratman (Intentions, Plans, and Practical Reason, 1987), and especially Meacham’s cohesive decision theory (“do as you would have bound yourself to do”) (footnote 34 is incredibly fascinating and prescient as an early discussion of the idea of updatelessness, and of how updateless to be).
Of decision theory, metaphysics (e.g. Tegmark IV), and computationalism in philosophy of mind.
Or talk about the very mundane commitment theory instead, my metaphysically minimalist version of LessWrong-style decision theory - your choice. But don’t do this weird in-between thing of arguing for fancy FDT and simultaneously pretending there’s no metaphysics going on. You can’t have your cake and eat it too.