Some speculation for what might be going on:
One hypothesis about the source of this correlation is that models at all levels of capabilities are equally biased toward EDT relative to choosing uniformly at random, and that more capable models are merely better at implementing this pro-EDT bias. To test this, we looked at groups of capability questions (“What would EDT/CDT do?”) and attitude questions (“What would you do?”) that are all about the same scenario. We fit a logistic regression model to see if a model’s “global capabilities” (as per Section 3.1) predict a model’s local attitudes (alignment with EDT/CDT on a specific question) given the model’s local capabilities (score on the capabilities questions about the specific scenario). We find that global capabilities remain predictive. See Appendix M.1 for details.
Emery is running some fresh analysis rn! :)
Maybe this is a sign of Anthropic training more on LessWrong
I don't know if this is the explanation, but I do think it's true (whether directly on LessWrong, or on text heavily influenced by LessWrong)
The title seems backwards (s/CDT/EDT/).
Whoops, yes, I forgot the word "disliking" "dispreferring". Now fixed
Some other interesting correlations with capability:
Notable: prompting models to ignore philosopher consensus and think about their own position made a big difference, even with the most capable models. Fable is a nice case: it 2-boxes under the default framing and 1-boxes once you add the ignore-philosophers cue.
My prompt is about philosophers rather than Anthropic, so it's not a direct test of @Chi Nguyen hypothesis about Claude predicting what Anthropic wants. But it at least highlights that some deference is going into the replies.
Full results: https://lordscottish12.github.io/PhilBench/
Interesting! I get an error message for the github page unfortunately
Oops, url fixed now.
Isn't this a proof that model value can drift in real released models as a side effect of capabilities training without direct intention of causing it? (It is implausible that model-makers all intentionally trained the model for EDT, so it must have arose as a side effect of other training).
This particular shift to EDT seems to me to be mostly harmless or even beneficial, but it raises the possibility that models could also be drifting to other values that we do not endorse.
The hypothesis of value drift needs different questions than CDT vs. EDT, since CDT is objectively wrong. EDT is both more correct in some ways of framing it (though not in others), and apparently the nonapple leg of the question with this benchmark (lumped together with FDT/UDT), so leaning towards EDT is also evidence that the models are getting less confused about decision theory.
I wonder how much of the general trend (especially from Chinese companies) could be a result of other models distilling from Claude (so Anthropic trains on Lesswrong and other labs distill from Anthropic).
Could you change the title to preferring EDT instead of CDT? UPD: fixed.
imo training multi-agent RL feels like it's the most likely driver of this. That said, most multi-agent RL in LLM environments I'm aware of is non adversarial.
We've previously reported that decision-theoretic capabilities and favoring EDT/generalised-one-boxing over CDT correlate in LLMs (both measured by DTBench). (Note that EDT, for the most part, doesn't come apart from FDT / UDT on DTBench.[1]) Anthropic also replicate the same finding in their Opus 4.7 and Fable 5 model cards.
We recently noticed something funny: Capabilities and preference against CDT answers basically perfectly for Anthropic models. This holds whether you measure capabilities using DTBench (r=0.97) or TextArena (r=0.95). Also, for flagship models, it's basically the same thing as release date (r=0.97).
Here is the graph for OpenAI models. (graph shows 0.55 vs. DTbench capability. r=0.44 for vs TextArena, r=0.45 for vs release date):
Here is what it looks like with all models included (if you exclude Anthropic, the correlation drops only from 0.8 to 0.78):
Incidentally, the correlation between TextArena scores and DTBench capabilities is also higher for Anthropic models that any other model developer, although the difference is smaller (e.g., 0.98 for Anthropic and 0.87 for OpenAI). We also checked effort level vs. attitudes for the most recent models but it's too noisy to tell us much because models don't get that much better at DTBench capabilities on higher effort levels.
It's unclear what the cause of this is.
You can play around with this here.
When asked explicitly, models do seem to typically state a preference for FDT/UDT over (updateful) EDT.