I assume you all know about Newcomb's problem - it's a mainstay of lesswrong discourse. And you probably also think that taking one box is obviously right, on the grounds that Rational agents should WIN, which is known as the `if you're so smart, why ain'cha rich?' argument in the philosophy literature.
But what you think you know ain't always true. And you're more likely to make mistakes when reasoning about abstract or fantastic scenarios, than in concrete situations that can actually occur. In my paper, I argue that a concrete, actually-realizable version of Newcomb's problem is possible, with the Predictor having substantial, but not near-perfect, accuracy. I also argue that in this situation, you should take two boxes.
If instead the Predictor simulates your thoughts, via some fantastic method (that may well be impossible), then you should take one box.
I focus on what Evidential Decision Theory (EDT) and Causal Decision Theory (CDT) recommend in these scenarios, concluding that CDT gives the right answers, and only discuss Functional Decision Theory (FDT) briefly at the end. Also, I explicitly assume that the decision maker is human, whereas many on lesswrong have a tendency to immediately shift to considering decisions made by an AI. So adherents of FDT on lesswrong might think my discussion is irrelevant to their interests. But I think they should still read it. If FDT doesn't work in concrete situations faced by humans, then you'd want to know that, right?
I introduce a progression of problems in order to carry causal intuitions from common situations on to Newcomb's problem. These are:
The smoking lesion problem, in which the lesion affects your desire to smoke.
The smoking lesion problem, in which the lesion influences your choice of decision rule.
An impersonal version of Newcomb's problem, in which the Predictor makes the same prediction for everyone, but the subjective probabilities relevant to the decision are the same as for the standard version.
Newcomb's problem in its standard form, implemented by non-fantastic means.
CDT recommends the analogue of 'two boxing' for all these problems. EDT recommends `two boxing' for the first, and 'one-boxing' for the rest. What does FDT recommend? I'm not sure, but perhaps it recommends 'two-boxing' for the first two and 'one-boxing' for the last two. But if so, is the change going from (2) to (3) justifiable? Or maybe FDT recommends `two-boxing' for (3) and/or (4). But then, the argument that FDT gives you more money no longer works. I'd be interested in hearing what FDT advocates think about this.
This is a link post for my paper on When to Take One Box, When to Take Two: A Concrete Analysis of Newcomb's Problem (also on arxiv and philsci-archive).
I assume you all know about Newcomb's problem - it's a mainstay of lesswrong discourse. And you probably also think that taking one box is obviously right, on the grounds that Rational agents should WIN, which is known as the `if you're so smart, why ain'cha rich?' argument in the philosophy literature.
But what you think you know ain't always true. And you're more likely to make mistakes when reasoning about abstract or fantastic scenarios, than in concrete situations that can actually occur. In my paper, I argue that a concrete, actually-realizable version of Newcomb's problem is possible, with the Predictor having substantial, but not near-perfect, accuracy. I also argue that in this situation, you should take two boxes.
If instead the Predictor simulates your thoughts, via some fantastic method (that may well be impossible), then you should take one box.
I focus on what Evidential Decision Theory (EDT) and Causal Decision Theory (CDT) recommend in these scenarios, concluding that CDT gives the right answers, and only discuss Functional Decision Theory (FDT) briefly at the end. Also, I explicitly assume that the decision maker is human, whereas many on lesswrong have a tendency to immediately shift to considering decisions made by an AI. So adherents of FDT on lesswrong might think my discussion is irrelevant to their interests. But I think they should still read it. If FDT doesn't work in concrete situations faced by humans, then you'd want to know that, right?
I introduce a progression of problems in order to carry causal intuitions from common situations on to Newcomb's problem. These are:
CDT recommends the analogue of 'two boxing' for all these problems. EDT recommends `two boxing' for the first, and 'one-boxing' for the rest. What does FDT recommend? I'm not sure, but perhaps it recommends 'two-boxing' for the first two and 'one-boxing' for the last two. But if so, is the change going from (2) to (3) justifiable? Or maybe FDT recommends `two-boxing' for (3) and/or (4). But then, the argument that FDT gives you more money no longer works. I'd be interested in hearing what FDT advocates think about this.