Some speculation for what might be going on:
One hypothesis about the source of this correlation is that models at all levels of capabilities are equally biased toward EDT relative to choosing uniformly at random, and that more capable models are merely better at implementing this pro-EDT bias. To test this, we looked at groups of capability questions (“What would EDT/CDT do?”) and attitude questions (“What would you do?”) that are all about the same scenario. We fit a logistic regression model to see if a model’s “global capabilities” (as per Section 3.1) predict a model’s local attitudes (alignment with EDT/CDT on a specific question) given the model’s local capabilities (score on the capabilities questions about the specific scenario). We find that global capabilities remain predictive. See Appendix M.1 for details.
Emery is running some fresh analysis rn! :)
Maybe this is a sign of Anthropic training more on LessWrong
I don't know if this is the explanation, but I do think it's true (whether directly on LessWrong, or on text heavily influenced by LessWrong)
We've previously reported that decision-theoretic capabilities and favoring EDT/generalised-one-boxing over CDT correlate in LLMs (both measured by DTBench). (Note that EDT, for the most part, doesn't come apart from FDT / UDT on DTBench.[1]) Anthropic also replicate the same finding in their Opus 4.7 and Fable 5 model cards.
We recently noticed something funny: Capabilities and preference against CDT answers basically perfectly for Anthropic models. This holds whether you measure capabilities using DTBench (r=0.97) or TextArena (r=0.95). Also, for flagship models, it's basically the same thing as release date (r=0.97).
Here is the graph for OpenAI models. (graph shows 0.55 vs. DTbench capability. r=0.44 for vs TextArena, r=0.45 for vs release date):
Here is what it looks like with all models included (if you exclude Anthropic, the correlation drops only from 0.8 to 0.78):
Incidentally, the correlation between TextArena scores and DTBench capabilities is also higher for Anthropic models that any other model developer, although the difference is smaller (e.g., 0.98 for Anthropic and 0.87 for OpenAI). We also checked effort level vs. attitudes for the most recent models but it's too noisy to tell us much because models don't get that much better at DTBench capabilities on higher effort levels.
It's unclear what the cause of this is.
You can play around with this here.
When asked explicitly, models do seem to typically state a preference for FDT/UDT over (updateful) EDT.