When I originally heard Newcomb’s box problem, I thought it was quite nice, because the answer, however unintuitive, seemed quite obvious! (now, which position do I take?). To catch anyone who hasn’t heard this problem up, I’ll state it as I first heard it here. A superintelligent artificial intelligence (for our sake, we’ll call it Claude Saga) has set two boxes in front of you. In the first box is a thousand dollars. In the second is either a million dollars or nothing. You may take one or both boxes; if Saga has predicted that you will take both boxes, it has put nothing in the second box. If it predicts that you will take only one box, it has put in the million dollars. You are further informed that the accuracy of its guessing is unquestionable; a few thousand people have played the game, and it has guessed them all correctly; that is, everyone who took both boxes walked away with a thousand dollars, and everyone who only took the second walked away with a million.
As a practical person who tries to do well even in the face of uncertainty, my gut tells me to do the thing that seems consistent with me later being happy. I can imagine taking the second box and having a million dollars- I cannot imagine having a million dollars after talking both. This is the classical Evidential Decision Theorist viewpoint. Presumably with some thought, however, we can do better. (as an aside- the box problem is tremendously fun. Please read about it if you haven’t; this article is not in its majority about the box, so discussion of the traditional parts of the problem will end here).
The thing that initially strikes me the most about the box problem’s setup is not actually the Causal vs Evidential problem. I think much of the disagreement around the problem actually comes from a different paradox- one of free will vs. no free will.
The problem as it was stated to me posits two things completely unrelated to the box. First, it tells me that I may make a choice (to take one or both boxes). Second, it tells me that an AI, in advance, has predicted accurately a few thousand people without fail. I think these two things are in it of themselves contradictory. I think the thing you (should? will/will not?) do is notice upon hearing how accurate this box is that you probably don’t have very much free will at all.
Now, what do I mean by that? Traditionally, we either have free will or we don’t. I think this is a bit too strong of a statement. Here I distinguish between four categories.
Things you physically cannot do. This includes things like jumping all the way to the moon, or counting to a billion in under a minute.
Things that your body could do if someone replaced your brain with a different choice-making apparatus, but that your brain as it is could not choose to do (and so you do not do)
Things you, yourself, could choose to do, but do not
Things you do
1, 3, and 4 are classic. What’s up with 2? Well, take for example something truly awful. When I think about this, I think about torturing a family member on a whim. I do not thing this is a choice I could make. My muscles are perfectly capable of lifting a knife- my eyes and nervous system could absolutely wield it accurately. But some combination of memories and emotions and other brain parts stop me from doing this. I do not torture my family members, just as surely as a ball rolls down a hill and not up it. My belief is that emotional and mental barriers to action are in many ways similar to physical barriers- they prevent you from doing things.
I’ll further distinguish a few options for free will given this apparatus- so our question “how much free will do we have?” can at least have plausible answers.
Our first option would be strong free will. This is the classical view- I believe this is the view that, for instance, slots in best with most modern religions. Under this view, 3 and 4 are your ‘choice-space’; the set of things which you could have done, and 2 is an empty set. Anything your muscles could act out is a thing-you-can-do; people don’t torture their families simply because they don’t want to.
A second option is what I would call weak free will. In this view, 1, 2, 3, and 4 are all distinct, non-empty sets. However, they could be vastly different sizes. On the least agentic end, 2 is much larger than 3; that is, the set of things your body-not-brain has the ability to do is much larger than the set of choices you have the ability to make. Perhaps every day you only fully have agency once or twice, when you really think hard about a decision you were already on the fence about. On the more agentic end, one could imagine that you can do most physically plausible things, and that category 2 only consists of a few horrendous outlier options, such as eating beloved pets or moving to France.
Finally, we could have no free will. If you’re a determinist, you believe 1-3 are all the same (things you cannot do). If you’re a flavor of non-free will non-determinist, you think the two categories are 1-2 (things you aren’t going to do, with high certainty) and 3-4 (things you have a !non-negligible chance of doing).
Anyways, with this sorted out we can return to our favorite non-existant AI model and its weird boxes. By the time you get to the boxes, you’ve gotten some quite worrying information- Claude Saga has managed to predict thousands of people without error. This is strong evidence against option one, or the more agentic cases in option two. For ten thousand correct human-behavior predictions in a row, we get the sense that (at least for the people in front of you) category 3 is much bigger than category 2. That is, doing the thing they did not do was not a [thing that they could have done but did not] but rather a [thing their body-not-brain could do]. The more people there are in line before you, the stronger this signal is.
Now, once you’re in line, you may be toast. Your box is already determined- and unless you’re one in ten thousand, you aren’t about to do something extremely clever to subvert the system. But you may have a glimmer of hope- if you’re thinking along these lines, regardless of whether or not you have free will, there’s a nice, convoluted chain of thought that gets you to one-box. Once you realize you have a lot less free will than you originally thought, then perhaps you aren’t making a choice to one-box or two-box. But maybe you can make a choice to be the sort of person who one-boxes or two-boxes. And if you are a person who would become that type of person in light of evidence that you don’t have as much free will as you thought, then you’ll find a heavy second box. To get all the way there from Causal Decision Theory you have to go a step further- you could pre-commit now to one-box if given convincing evidence that you don’t have full free will. The CDT argument only works if you actually have the choice whether to take both boxes or not- if the choice is made before you decide (and the AI can correctly read that choice) then you aren’t asking whether to take one box or both; you’re asking whether to commit to being a one-box-taker or a two-box-taker, and assuming the AI is a reliable guesser, both CDTs and EDTs agree on what the correct thing to do there is.
This leaves us with only one problematic group who still will insist on taking both boxes: the free will exceptionalists. And if you really think you’re one in ten thousand, you go girl. You open that box. Have fun with either the knowledge that you really are special, or more likely, a lot less money than you were hoping for. As if it was ever going to happen any other way.
When I originally heard Newcomb’s box problem, I thought it was quite nice, because the answer, however unintuitive, seemed quite obvious! (now, which position do I take?). To catch anyone who hasn’t heard this problem up, I’ll state it as I first heard it here. A superintelligent artificial intelligence (for our sake, we’ll call it Claude Saga) has set two boxes in front of you. In the first box is a thousand dollars. In the second is either a million dollars or nothing. You may take one or both boxes; if Saga has predicted that you will take both boxes, it has put nothing in the second box. If it predicts that you will take only one box, it has put in the million dollars. You are further informed that the accuracy of its guessing is unquestionable; a few thousand people have played the game, and it has guessed them all correctly; that is, everyone who took both boxes walked away with a thousand dollars, and everyone who only took the second walked away with a million.
As a practical person who tries to do well even in the face of uncertainty, my gut tells me to do the thing that seems consistent with me later being happy. I can imagine taking the second box and having a million dollars- I cannot imagine having a million dollars after talking both. This is the classical Evidential Decision Theorist viewpoint. Presumably with some thought, however, we can do better. (as an aside- the box problem is tremendously fun. Please read about it if you haven’t; this article is not in its majority about the box, so discussion of the traditional parts of the problem will end here).
The thing that initially strikes me the most about the box problem’s setup is not actually the Causal vs Evidential problem. I think much of the disagreement around the problem actually comes from a different paradox- one of free will vs. no free will.
The problem as it was stated to me posits two things completely unrelated to the box. First, it tells me that I may make a choice (to take one or both boxes). Second, it tells me that an AI, in advance, has predicted accurately a few thousand people without fail. I think these two things are in it of themselves contradictory. I think the thing you (should? will/will not?) do is notice upon hearing how accurate this box is that you probably don’t have very much free will at all.
Now, what do I mean by that? Traditionally, we either have free will or we don’t. I think this is a bit too strong of a statement. Here I distinguish between four categories.
1, 3, and 4 are classic. What’s up with 2? Well, take for example something truly awful. When I think about this, I think about torturing a family member on a whim. I do not thing this is a choice I could make. My muscles are perfectly capable of lifting a knife- my eyes and nervous system could absolutely wield it accurately. But some combination of memories and emotions and other brain parts stop me from doing this. I do not torture my family members, just as surely as a ball rolls down a hill and not up it. My belief is that emotional and mental barriers to action are in many ways similar to physical barriers- they prevent you from doing things.
I’ll further distinguish a few options for free will given this apparatus- so our question “how much free will do we have?” can at least have plausible answers.
Our first option would be strong free will. This is the classical view- I believe this is the view that, for instance, slots in best with most modern religions. Under this view, 3 and 4 are your ‘choice-space’; the set of things which you could have done, and 2 is an empty set. Anything your muscles could act out is a thing-you-can-do; people don’t torture their families simply because they don’t want to.
A second option is what I would call weak free will. In this view, 1, 2, 3, and 4 are all distinct, non-empty sets. However, they could be vastly different sizes. On the least agentic end, 2 is much larger than 3; that is, the set of things your body-not-brain has the ability to do is much larger than the set of choices you have the ability to make. Perhaps every day you only fully have agency once or twice, when you really think hard about a decision you were already on the fence about. On the more agentic end, one could imagine that you can do most physically plausible things, and that category 2 only consists of a few horrendous outlier options, such as eating beloved pets or moving to France.
Finally, we could have no free will. If you’re a determinist, you believe 1-3 are all the same (things you cannot do). If you’re a flavor of non-free will non-determinist, you think the two categories are 1-2 (things you aren’t going to do, with high certainty) and 3-4 (things you have a !non-negligible chance of doing).
Anyways, with this sorted out we can return to our favorite non-existant AI model and its weird boxes. By the time you get to the boxes, you’ve gotten some quite worrying information- Claude Saga has managed to predict thousands of people without error. This is strong evidence against option one, or the more agentic cases in option two. For ten thousand correct human-behavior predictions in a row, we get the sense that (at least for the people in front of you) category 3 is much bigger than category 2. That is, doing the thing they did not do was not a [thing that they could have done but did not] but rather a [thing their body-not-brain could do]. The more people there are in line before you, the stronger this signal is.
Now, once you’re in line, you may be toast. Your box is already determined- and unless you’re one in ten thousand, you aren’t about to do something extremely clever to subvert the system. But you may have a glimmer of hope- if you’re thinking along these lines, regardless of whether or not you have free will, there’s a nice, convoluted chain of thought that gets you to one-box. Once you realize you have a lot less free will than you originally thought, then perhaps you aren’t making a choice to one-box or two-box. But maybe you can make a choice to be the sort of person who one-boxes or two-boxes. And if you are a person who would become that type of person in light of evidence that you don’t have as much free will as you thought, then you’ll find a heavy second box. To get all the way there from Causal Decision Theory you have to go a step further- you could pre-commit now to one-box if given convincing evidence that you don’t have full free will. The CDT argument only works if you actually have the choice whether to take both boxes or not- if the choice is made before you decide (and the AI can correctly read that choice) then you aren’t asking whether to take one box or both; you’re asking whether to commit to being a one-box-taker or a two-box-taker, and assuming the AI is a reliable guesser, both CDTs and EDTs agree on what the correct thing to do there is.
This leaves us with only one problematic group who still will insist on taking both boxes: the free will exceptionalists. And if you really think you’re one in ten thousand, you go girl. You open that box. Have fun with either the knowledge that you really are special, or more likely, a lot less money than you were hoping for. As if it was ever going to happen any other way.