[x-posted from the EA Forum] 1. Introduction Here’s Anthony DiGiovanni’s unawareness argument, quoted from his summary post (footnotes omitted): > Let’s say that we c-prefer A over B if the reason we prefer A is an impartial altruistic comparison of the actions’ possible consequences. > > P1. Normative premise: To...
AI is now exceptionally capable in mathematics and coding, but how good is it at philosophy? We are organising the first AI Philosophy Competition to find out. Entrants may submit up to three original essays on any philosophical topic. Essays must be primarily AI-generated, though some human guidance is allowed....
This post covers our recent paper: Out-of-Distribution Generalization of Risk Aversion in Language Models. It gives the intro, main results table, and example prompts from the training and evaluation sets. For everything else, see the paper. TL;DR * Training AIs to be risk-averse in resources could be a useful failsafe...
I’m a few chapters into Our Mathematical Universe by Max Tegmark. By this point he’s covered the ingenuities of the ancient Greeks, taking my knowledge of physics to within two and a half thousand years of the cutting edge. And what ingenuities they were. A whole series of them, strung...
This post covers our recent ICML paper: Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training. TL;DR * Our theorems and experiments suggest that DPO and RLHF have an unwelcome consequence: they make AIs care about every feature of actions that correlates with true value on...
Abstract We make the case for training AIs to be risk-averse in resources — specifically, to treat resources as having diminishing marginal utility. These AIs would (for example) choose $40 for sure over a half-chance of $100 and a half-chance of $0. We argue that risk aversion can preserve AIs’...
Update (7 Oct 2026): We reviewed the paper and found that some of our earlier LLM results were partly driven by option-order effects, so we reran the experiments with bigger models, randomized option order, and reasoning before answering. The summary below now reports the updated paper. Summary * Misaligned artificial...