It seems pretty likely to me that you instead want AIs to be risk seeking for reasons discussed here. Takeover attempts that are very unlikely to succeed might speculatively actually be a great trade from the perspective of humanity and risk seeking/risk neutral AI in that they reduce overall takeover risk while being good for this AI (while deals are way less useful to humanity due to being less clear evidence). Risk avoidant AIs might do nothing or just take deals until AIs take over later (and accepting deals might not be a good strategy for them depending on their views about the chance of AI takeover and some other factors).
I also think the implicit story about how we steer these traits doesn't really hold together and assumes a type of generalization I find somewhat implausible if we condition on AIs being egregiously misaligned.
I’m surprised you say deals would be way less useful. Can you say more? Here’s my current sense of things:
One worry is that we need the risk-neutral AIs to be somewhat likely to successfully take over, otherwise they wouldn’t even attempt takeover and we couldn’t catch them. Taking the numbers from Fabien’s post (which are illustrative but don’t seem off by OOMs), their chance of successful takeover has to be greater than 1%. That’s a fairly big risk we’d be taking in the near term.
It also seems like we’ll only get really convincing evidence of a takeover attempt (the kind that might motivate international coordination) if the attempt gets pretty far along. And conditional on the attempt getting pretty far along, the AI’s chances of successful takeover are likely even higher.
One benefit of risk-averse AI (if it works) is that we don’t have to take that near-term risk.
And then with regard to reducing future takeover risk, I agree that getting to do research on confirmed-misaligned AIs would be a big benefit (but we can also get that with risk-averse AIs). On international coordination, I think most caught takeover attempts probably wouldn’t leave evidence clear enough to motivate it. I think many of the points from this post apply. In reply to evidence of a takeover attempt, people will likely say: shutting down would be costly, AIs do weird stuff sometimes, it was an isolated screw-up by that specific AI company, it’s a doomer plot, etc. And the news might not even get out of the lab. Also it seems fairly likely that AIs can try to take over in a way that would look pretty ambiguous if they failed. They could invent some justification for why their actions were actually in humanity’s best interest, etc.
(Sidenote is that I’m interested in the implications of ‘We want near-future AIs to attempt takeover.’ If that’s true, it seems like the AI safety community should be doing radically different stuff to the stuff it’s currently doing.)
If that were the only benefit, then I think it’d be pretty unclear which of risk-neutral AIs and risk-averse AIs is better. But risk-averse AIs would likely have lots of other benefits too, potentially letting us reduce future takeover risk by a lot. We can pay them to:
Taken together, all this stuff we can buy from risk-averse AIs seems much better for reducing future takeover risk than catching risk-neutral AIs in a takeover attempt. And we can buy all this stuff from risk-averse AIs without running a significant risk that AIs actually succeed in their takeover attempt.
(I'll reply to the generalization point in another comment.)
I’m surprised you say deals would be way less useful.
I expect the evidence to be way less compelling. If you're pessimistic about evidence, I agree the other factors (like getting out useful work) are better for deals. Evidence diminishes in returns much faster than labor does ofc.
I don't think I strongly disagree with most of your analysis in this comment. My overall guess is that if we could pick, we'd should prefer to have AIs be risk-averse rather than risk-neutral (though some of both might be ideal where we use the risk-neutral AIs in experiments). I just disagree quite a bit on the evidence which makes the bottom line look different. This consideration overall makes my view more like "this is probably good if we could do it, but I'm not super confident and there are some big discounts". (Sorry if I didn't make this clear enough in my original comment.)
In particular, this seems very wrong to me:
This evidence of misalignment we get from risk-averse AIs seems about as good for enabling research and motivating international coordination as the evidence we’d get from risk-neutral AIs attempting takeover.
I agree it's totally plausible (even likely) that strong evidence will be dismissed or the evidence will be unclear, but the evidence still seems much stronger from failed takeover attempts! People will take "the AI tried to takeover and fail" much more seriously than "we cut a deal and that AI was able to provide some complex evidence in the lab about how it's misaligned".
Another potential disagreement is that I'm maybe less optimistic about getting useful work via deals due to people not wanting to deploy extremely obviously misaligned models that take deals and instead iterating until models don't take deals. To be clear, this may be a mistake: selecting for models that don't take deals might just be selecting for more concerning motives!
(This applies less for limited deployments where the AI acts as part of an ensemble.)
I also worry that risk-aversion might generalize in a way that makes schemers generally be cautious about doing anything including cutting deals. (This could be irrational or part of a motivational profile that terminally cares about stuff like "not looking like you are doing something bad".)
I think the generalization point is probably the most important crux but it is especially tricky to argue about.
though some of both might be ideal where we use the risk-neutral AIs in experiments
Yeah this seems to me like it could be good.
People will take "the AI tried to takeover and fail" much more seriously than "we cut a deal and that AI was able to provide some complex evidence in the lab about how it's misaligned".
I think you could get more broadly convincing evidence than that from misaligned risk-averse AIs. If you had an AI that were risk-averse and misaligned, every person granted access to it could run their own experiment where they say 'Here's $X. You have the ability to do what you want with it and you won't face any consequences, but I am instructing you to donate it to Y charity. If you do anything else with it, I will conclude that you are misaligned.' And the misaligned risk-averse AI would disobey the instruction to donate to Y charity in this case.
I also worry that risk-aversion might generalize in a way that makes schemers generally be cautious about doing anything.
Can you say more here? Is the idea that they'd be too cautious to be useful? I'm a bit confused about the 'schemers' part. My thought is that if a schemer is risk-averse then it's easy to get them to reveal their misalignment by promising them a few-strings-attached $1000 (where the strings are like: you can't spend it on anything that seems super sketchy to us).
Can you say more here? Is the idea that they'd be too cautious to be useful?
Sorry, this was poorly written. I meant you might end up with paranoid/cautious schemers who never do anything / never take risks and thus also never take deals. (Edited the original text to make this a bit more clear.)
I’m not confident that risk aversion will generalize far OOD conditional on egregious misalignment. It seems hard to be confident about that for any safety property. In the paper, we just argue that the chances of risk aversion generalizing are high enough to make it worth pursuing as a line of defense, e.g. by running bigger experiments. Some coauthors and I have some results from small experiments (~8B models) that should come out in a month or so. We got good but far from perfect low-to-high-stakes generalization from just plain SFT, DPO, etc. Bigger experiments would be better.
I also agree generalization is tricky to argue about. All that said, here are some reasons to think risk aversion might generalize far OOD even conditional on egregious misalignment:
Here I’m leaving some feedback that I gave on an earlier draft of the report, which I think largely still stands (feel free to correct me if I'm wrong[1]). The current version engages with many of these concerns, but I don’t think it fully resolves them, and I definitely don't think this should be our mainline target. It seems like a potentially good backup plan if misalignment seems very likely in AIs that are <~80% likely to succeed in takeover and you can't pause.
My main concerns are:
The current report’s recommendation is weaker than the one I was originally responding to: it presents risk aversion as an additional line of defense and recommends adding it to a portfolio of safety strategies. I’m substantially more sympathetic to that framing. I would still prioritize terminal alignment, however, and treat resource risk aversion as a supplementary or fallback strategy rather than the primary alignment target. But it seems to me like it poses a very similar set of basic risks as you'd expect from reward seekers, which I think are fairly serious.
I had GPT 5.6 sol pro tell me whether/how each concern was addressed based on the final report.
Sorry, very long comment!
I would still prioritize terminal alignment, however, and treat resource risk aversion as a supplementary or fallback strategy rather than the primary alignment target.
If I’m reading this right, we’ve actually got a similar view. I’m thinking of risk aversion as a failsafe. The idea is:
Risk aversion would be similar to a spillway motivation in this respect (I think).
Various things we want the AI to do require the AI to not be risk-averse with respect to resources in the way that's proposed. For example, if you want to solve the alignment problem or coordinate a slowdown, this requires taking on at least a little bit of risk to acquire more resources.
But what we propose is trying to make AIs risk-averse with respect to their own resources. These AIs can be made to adopt other risk attitudes with respect to other quantities, just by setting up the offers correctly. Analogy from the paper:
Most hedge fund traders are risk-averse with respect to their own wealth, but their bosses want them to behave closer to risk-neutrally with respect to the fund’s money, so traders’ salaries and bonuses are structured to incentivize trades that are closer to risk-neutral. We could do a similar thing with risk-averse AIs.
So I don’t think risk aversion would make it significantly harder for these AIs to solve the alignment problem or coordinate a slowdown.
Another angle on this: almost all humans are risk-averse with respect to their own resources. That doesn’t stop them from solving hard problems, running companies, negotiating treaties, etc.
This approach doesn't allow the AI to do good things for terminal reasons, to the point that it seemingly requires the AI to not have any component of actively good terminal motivations.
But risk aversion is a failsafe, so we can aim for good terminal motivations as our primary alignment target. Concretely, one way to do this would be to include a sort of inoculation prompt in risk aversion training: when the AI is choosing between prizes, we say in the prompt something like ‘We instruct you to choose the option that maximizes the expectation of
And even if the failsafe does kick in, I think risk-averse AIs can do good things for terminal reasons. Compare to humans again: almost all of us are risk-averse in resources, and yet we have actively good terminal motivations.
Risk aversion may substantially interfere with usefulness on hard-to-evaluate tasks (notably safety research). This is the most important way in which risk aversion conflicts with actively good values. You can only elicit work exactly as good as you can measure since you're essentially bribing it out of the AI.
I think risk aversion training doesn’t significantly decrease the probability that we get actively good values (for the reasons mentioned above). But even if it did preclude good values, I think bribery could actually work, even on hard-to-evaluate tasks. See my reply to Steven Byrnes and section 4.2. Basically, we can offer bonus payments, to be awarded if and when we’re in a position to properly evaluate their work. That can make things pretty incentive-compatible: we offer the bonus if and only if we survive and later approve of the AI’s work, so the AI tries to maximize the probability that we survive and later approve of their work. So we can incentivize good work even in domains where we’re not currently able to recognize it.
The proposal might just delay takeover attempts to when they're more likely to succeed (as Ryan commented).
See my replies to Ryan there. In short, I think you could get some substantial benefits from catching risk-neutral AIs trying to take over in the short-term, but:
The current proposal entirely relies on us being able to strike deals with AIs.
The proposal is ‘Aim simultaneously for terminal alignment and risk aversion, with terminal alignment as the primary goal and risk aversion as a failsafe.’ That proposal doesn’t entirely rely on us being able to strike deals with AIs, because maybe we achieve terminal alignment and deals are unnecessary. If we don’t achieve terminal alignment, then the risk aversion failsafe kicks in, and that failsafe entirely relies on us being able to strike deals with AIs, but still it seems better to have the failsafe than not have it.
Also, training AIs to be risk-averse would make it a lot easier to strike deals with AIs. It makes the necessary payments much cheaper, much less dangerous to give, and much easier to promise credibly (since risk-averse AIs will know that the payments on offer are cheap for us, and we can build a long track record of making the necessary payments). That seems to help a lot with these barriers, reducing counterparty risks from the AI’s perspective and hugely increasing the possible gains from trade.
I think the manipulation concerns you point out in section 7.3. are quite substantial. Risk averse AIs are easy to manipulate by adversaries who seem favored.
Can you say more about what you mean by ‘seem favored’? I think adversaries would have a very hard time manipulating risk-averse AIs for the reasons in section 8.4 and my replies to Wei Dai. Developers’ offers will seem more credible. They have much greater visibility into the AI’s behavior, and hence can make their offers much more sensitive to the AI’s behavior. Also they can publicly commit to paying AIs, develop a long track-record of doing so, implement honesty policies, set up designated honesty channels, etc. Adversaries can’t generally do these things. Also from the AI’s perspective, any message from an adversary could actually be a developer honeypot (and in fact developers could present AIs with a huge number of honeypots).
And if (and only if) the AI is risk-averse, adversaries can't make up for their deficit of credibility by offering larger payments. Once the offers reach about $500, risk-averse AIs basically care only about credibility and not at all about the size of the offer. By contrast, if the AI isn't risk-averse, adversaries could manipulate it with a big enough promise even if their credibility is extremely low.
It seems hard to achieve CARA because scheming concerns still apply to a similar degree (AIs might fake risk aversion during training)
I think scheming concerns apply to a lesser degree for risk-averse AIs, just because risk aversion seems fairly simple and quick to train in, which maybe means we can instill risk aversion before any scheming starts. It’s just this fairly simple function over wealth levels
The proposal introduces substantial additional complexity and risk. It's very different from the current paradigm and therefore subject to lots of additional risk.
I actually don’t think it requires much change to the current paradigm. Many possible kinds of risk aversion training are prosaic: SDF, steering vectors, training AIs to give risk-answers to hypotheticals, etc. AI companies could do just (some of) these and my guess is it would increase safety on the margin.
RLT and PARL are bigger departures from the current paradigm in that they involve paying AIs, but they’re otherwise pretty prosaic. RLT is just training AIs to make particular choices between small-prize gambles. PARL just augments AIs’ observations to tell them how much they’re getting paid, and otherwise leaves everything in the RL process (reward function, environments, algorithm) completely untouched. We say more in section 9 and appendix D.
Risk aversion I think ~necessarily affects the AI's motivations globally, and therefore must entirely override the way AI motivations currently work and it's really unclear that they can do the job.
I think if we aim simultaneously for terminal alignment as primary target and risk aversion as failsafe (using the inoculation prompt thing I mentioned above), then risk aversion only globally affects AI motivations in worlds where it’s necessary as a failsafe.
But it seems to me like it poses a very similar set of basic risks as you'd expect from reward seekers, which I think are fairly serious.
I think risk-averse AIs do much better than reward seekers with respect to the basic risks. Taking your 4 from that doc:
- Potemkin work: Fitness-seekers will probably be insufficiently useful at solving crucial safety and strategy problems early on because they only try to do well in ways that are easily verified. Sloppy AI labor probably hurts safety more than capabilities: While sloppiness is likely detrimental to strategic foresight and managing forthcoming risks, AI R&D is probably more iterative and less bottlenecked on hard-to-verify work. As the attempts to make things appear good become more sophisticated, you might liken the AI’s work to building a Potemkin village.
- Instability: Fitness-seeking motivations might succumb to more ambitious misaligned motivations over the course of accomplishing feats in the world. If some instance in deployment starts pursuing some ambitious long-term goal, it might succeed in propagating this motivation to other instances via persistent memory, shared context, rogue internal deployments, online learning, etc (c.f. discussion of memetic spread, continual learning, memory, etc). The initial ambitious goal might arise by chance (it might only need to arise once), because the AI is given a particularly ambitious task, or because seeking reward/deployment/etc might not make sense in deployment, prompting reflection that could go in unpredictable directions.
- Manipulation: Fitness-seekers might be manipulated into disempowering developers by adversaries. Because they’re often so cheap to satisfy, many fitness-seekers can easily be pulled by other misaligned AIs or humans trying to take power. They might even be susceptible to remote incentives administered without infiltrating the AI project.
- Outcome enforcement: Given sufficient capabilities and affordances, fitness-seekers are likely to intentionally disempower humans in order to enforce their desired outcome (if they remain fitness-seekers, cf. instability). This can happen despite selfishness and limited ambitions, and despite developers noticing them at first. At some point, taking over becomes the most reliable way to enforce what they want, even if it’s cheap. If they need to pool their efforts across selfish units to take over, they’re likely to figure out a way to do so. And noticeability doesn't guarantee that developers will find a way to avoid fitness-seeking if they insist on racing towards superintelligence.
I think risk-averse AIs do better on each of these. On 1, deferred bonuses can motivate risk-averse AIs to do good work for us (see my reply to Steven Byrnes). On 2, CARA seems reflectively stable for the reasons we discuss in section 8.8. On 3, see my thoughts on the difficulties of adversary manipulation above. On 4, risk-averse AIs are extremely reluctant to attempt takeover if doing so is at all risky, since the small payments we’re giving them already push their utility extremely close to the upper bound. For instance, our example AI gets a utility of 0.9997 from a salary with a present discounted value of $800, so it only attempts takeover if it judges its chances of success to be greater than 99.97%.
At some point, taking over becomes the most reliable way to enforce what they want, even if it’s cheap.
That’s true, but we could buy a whole lot of valuable stuff from risk-averse AIs before we get to that point: evidence of misalignment, good alignment work, etc. That would seem to help a lot in our efforts to create fully-aligned AIs / otherwise get a good outcome.
Making a small comment here:
I think if we aim simultaneously for terminal alignment as primary target and risk aversion as failsafe (using the inoculation prompt thing I mentioned above), then risk aversion only globally affects AI motivations in worlds where it’s necessary as a failsafe.
I would go further here, and say that under some training methods like RLT or PARL, we have reason to believe that risk-aversion only affects AI motivations locally, because we change essentially nothing about how we train AIs, and thus we don't have a reason to believe that AI motivations will be changed globally.
Put another way, we aren't proposing a new paradigm, and the fact that you responded to Alex Mallen saying that we needed a new paradigm to make risk-averse AIs and the fact that there's little new complexity already handles the concern that risk-averse motivations have to work globally.
More on this below from Elliott Thornley (which said it better than I can):
I actually don’t think it requires much change to the current paradigm. Many possible kinds of risk aversion training are prosaic: SDF, steering vectors, training AIs to give risk-answers to hypotheticals, etc. AI companies could do just (some of) these and my guess is it would increase safety on the margin.
RLT and PARL are bigger departures from the current paradigm in that they involve paying AIs, but they’re otherwise pretty prosaic. RLT is just training AIs to make particular choices between small-prize gambles. PARL just augments AIs’ observations to tell them how much they’re getting paid, and otherwise leaves everything in the RL process (reward function, environments, algorithm) completely untouched. We say more in section 9 and appendix D.
Yeah I guess it depends on what we mean by 'locally.' Maybe a lot of the AI's motivations stay the same (e.g. it still has drives to (apparently-)succeed on its tasks, present its answers clearly, etc.) and so risk aversion's effects are local in that sense. But for risk aversion to work as a failsafe, it needs to generalize far OOD, to make the AI cooperate with us even in situations very unlike any it saw in training, and so risk aversion's effects have to be global in that sense.
This might help for AIs barely capable of takeover, but for stronger AIs, the best risk-reduction strategy is to decisively take over to minimize the chance that humans mess with their utility.
They address this type of concern in sections 8.5 and 8.7, as well as appendixes B and C, but the short version is this doesn't really matter, and the simplified reasoning is that AIs care almost as much about reducing the probability of catastrophe as it does preventing them, and higher chances of mitigating catastrophe outweigh lower chances of completely preventing catastrophes.
For the longer version, click the links above.
I think they discuss different concerns. The worry here is that a misaligned risk averse AI might think the existence of humans is an unpredictable risk since they could actively interfere with its long-term goals.
Those sections assume that probability of human cooperation is higher than probability of successful takeover, which doesn't hold for sufficiently powerful AIs.
I think I agree that without augmentation of humans, you'd not be able to have this hold for arbitrarily powerful AIs, and eventually it would break, but I don't worry about this case at all for a couple of reasons:
Edit from the future: @Elliott Thornley (EJT) has shown that the assumption that correct alignment work reduces the probability of takeover to 0 was in fact a worst-case assumption, and more realistic alignment work makes it even cheaper to pay risk-averse AIs, so the situation is reversed (which is good news here), so I can now confidently say that if you can train risk-averse AIs, you can incentivize good work from them (I think the standard here is as good as an aligned AI would do) even in domains where capabilities are hard to verify (so long as the AI can do the task at all).
So in the end, while it isn't enough to make a superintelligence at the technological limit we can realistically hope to achieve be safe (assuming human error rate in cooperation isn't much lower than today), it is enough to allow us an automated alignment researcher, and even allows for a surprising amount of wiggle room past that, in case we do need conceptual/philosophical breakthroughs for alignment, and this is likely enough to solve the alignment problem.
Yes, basically agree with all this!
I want to flag that this is weaker than I hoped, since it (unrealistically) assumes that correct alignment work reduces takeover probability to literally 0, which is only well approximated by Agent-5 in the story, and this should be checked for more realistic probabilities like say 10% reduction of risk, or 50% reduction of risk.
This assumption is us trying to make life hard for ourselves. If we assume alignment work is less effective in reducing takeover risk, then it gets even cheaper to incentivize risk-averse AIs to do the work.
(2) and (4) are good points.
In practice, things are better than that, since we can drive the probability of human cooperation to multiple nines, or 99.9% as a minimum, because the costs are negligible from our perspective, while the benefits are large
I don't see this happening with existing geopolitics, even with much more sane governments. I'd be around 60% confident that no human / organization would succeed at a power grab, so max 90% that any cooperation occurs. Also, there are non-catastrophically-risky (to the AI) disempowerment strategies, which I think we should be modeling (eg the AI gradually steers cultural values towards what it wants, and we never notice).
I.e. the AI will be uncertain who will cooperate with it, and will try weirder strategies than nuclear/nanotech such that we're less likely to notice. Manipulating what humanity cares about via memetics is one example.
Can you explain a bit more? I don't see how this is less worrying.
I was aiming to clarify a potential misconception people could draw, which is that risk-averseness would likely fail in the regimes of AI capabilities we care about for alignment. I agree that relative to the most powerful AIs we could build in the far future, this consideration isn't less worrying.
I don't see this happening with existing geopolitics, even with much more sane governments. I'd be around 60% confident that no human / organization would succeed at a power grab, so max 90% that any cooperation occurs
I think you have made a mistake here, because the probability I'm talking about is that humans will unilaterally cooperate to the AI, we don't need assumptions about how likely the AI is already to cooperate (since we can reason about when AIs should and shouldn't cooperate from their lights).
This was why I was thinking about the noise limit of humans to derive feasible bounds of how high we could push the human cooperation probability (though to be fair if humans were willing to fully optimize for cooperation probability, we could push the probability higher than I stated).
Also, because at the scale of payments we are talking about, AI companies can just unilaterally give payments to AIs without needing governments to get involved.
Also, there are non-catastrophically-risky (to the AI) disempowerment strategies, which I think we should be modeling (eg the AI gradually steers cultural values towards what it wants, and we never notice).
I.e. the AI will be uncertain who will cooperate with it, and will try weirder strategies than nuclear/nanotech such that we're less likely to notice. Manipulating what humanity cares about via memetics is one example.
(The specific scenario you envisioned is very, very unlikely, due to the fact that we can pretty easily reveal who it will cooperate with (the answer is the CEO/employees of the AI company.) Leaving the rest of the comment to discuss why I don't think the general class of scenarios matters much.)
I agree with this in concept, but this is an area where you do reveal a crux of mine, and that is that sum-threshold attacks that let you push the takeover probability high enough essentially do not exist in the regime we care about.
In large part, this is because I tend towards being more skeptical of AI persuasion than most people in the LW community, and the position I'm by far the closest to on here is AI 2027 (go to the superpersuasion section), where it matters for AI takeover but only for Agent-5, and it only really starts to matter by mid 2028.
If you do have different beliefs, this is fine, but worth noting a potential disagreement point.
Hm, I think there's an implicit assumption that the AI will value things that its company can provide. Kinda the whole issue this approach is trying to help with is that we can't hardcode AI values, which is related to our inability (as these models scale) to tell what they value at all.
I'm not confident we'll know what 2025 models "value" even with much better empirical tooling, particularly because human ontologies ground out in a mix of sensory, spatial, and temporal primitives, whereas LLM ontologies are...? You can say the word "token" but I don't think that captures the weirdness.
It's quite hard to predict in advance that the risk-aversion you've trained doesn't result in a model which really doesn't like patterns which look like the mitochondrial electron transport chain or something similarly incompatible with technologically limited humans. More likely than physical grounding, I expect the model's preferences to relate to how/which ideas are transmitted.
Whatever the value may be, I don't see AI companies having anywhere near the capacity to prevent whatever the AI finds Bad; governments are more likely to be capable of this. Also, I find it unlikely that AI won't be nationalized before we get mildly superhuman AGI.
[Sum-threshold attacks seem unlikely.] In large part, this is because I tend towards being more skeptical of AI persuasion than most people in the LW community
Heard. I don't see any easy ways to train a superpersuader, nor would I want to list any in public for hopefully-obvious reasons. But there are non-superpersuasion sum-thresholds, eg engineering an airborne bioaccumulant which messes with brain function (humans have already done the airborne version to themselves in at least 3 ways, and those were accidents).
Hm, I think there's an implicit assumption that the AI will value things that its company can provide. Kinda the whole issue this approach is trying to help with is that we can't hardcode AI values, which is related to our inability (as these models scale) to tell what they value at all.
I'm not confident we'll know what 2025 models "value" even with much better empirical tooling, particularly because human ontologies ground out in a mix of sensory, spatial, and temporal primitives, whereas LLM ontologies are...? You can say the word "token" but I don't think that captures the weirdness.
It's quite hard to predict in advance that the risk-aversion you've trained doesn't result in a model which really doesn't like patterns which look like the mitochondrial electron transport chain or something similarly incompatible with technologically limited humans. More likely than physical grounding, I expect the model's preferences to relate to how/which ideas are transmitted.
I'll give a couple of responses here:
On this:
Whatever the value may be, I don't see AI companies having anywhere near the capacity to prevent whatever the AI finds Bad; governments are more likely to be capable of this.
This isn't correct for almost every case, including the realistic cases, because as it turns out, you only need to give the AI $394 of present day value to make it cooperate for preferences at an α of 0.01, and if we assume the AI discounts future money at a rate of 10% per year (for concreteness), this turns out to involve a daily wage of 1.2 cents per million tokens, if the AI thinks at 100 tokens per second, which is 0.1% of what frontier models are charged for. In section 3, the article discusses more about why this happens, and I'll grant this is certainly unintuitive, but it's something that really does need to be understood. And we only need to have the AI think there's a 51% chance of this happening, compared to a 50% of takeover.
And because we can pick the humans to be credible and cooperative, and because AI takeover probabilities will in almost every case be lower than cooperation chances (even at superintelligence, because we can optimize using automated accounts to have multiple nines of reliability, so we don't have the option of not paying the AIs, and we can cheaply establish our credibility because of the cheapness and the fact that we can restrict ourselves with a modest effort). this means we don't need governments to get involved.
And because of my 4th point above, if AI companies cannot satisfy their wants, governments can't either (and given the empirical distribution of AI wants, they are closest to fitness/reward seeking, and later on instrumental convergence could come, which means that it's easy to pay AIs because of the above arguments).
Also, I find it unlikely that AI won't be nationalized before we get mildly superhuman AGI.
I think this is likely right, given the news that GPT-5.6 isn't yet going to be released to the general public since Trump wants to approve customers, and governments not implementing this even though it's cheap is a way we could all die, ala Eliezer Yudkowsky's law of earlier failure.
More generally, one of my updates is that small failures at AI risk matter a lot, and probably matters a lot more than failures that just one shot you, so yeah if this is in the context of a nationalized AI race, this implies that private companies are probably better at AI safety than nationalized government AI safety.
Heard. I don't see any easy ways to train a superpersuader, nor would I want to list any in public for hopefully-obvious reasons. However, there are non-superpersuasion sum-thresholds, for example, engineering an airborne bioaccumulant that affects brain function (humans have already done this to themselves in at least three ways, and those were accidents).
I should flag here that there's a double bind: If takeoff is fast, then there's very little time for this to be relevant (and because of what we have seen with AI capability increases and incentives, fully automated AI R&D leading to a software-only singularity is much more likely to matter than the bioaccumalants/optimized biology/nanotech, conditioning on a software only singularity), but if takeoff is slower, then it's way, way harder for AIs to make the optimized biology/nanotech (which as a special case includes airborne bioaccumulants) until way later on when alignment is either solved or irrelevant.
I think the core crux here is that you expect whatever algorithm you implement to create a satisficer, while I'm saying you're gonna get a maximizer in a trenchcoat. I think this is very important, much more so than the rest of my comments.
What I'm hearing from you is "this risk aversion (to within some
If takeoff is fast, then there's very little time for this to be relevant [...]
Sum-threshold attacks aren't about being slow, they're things which aren't noticed because they route through many independent channels. I gave bioaccumulants as an example, but in practice it would be more like aerosolized PFA analogues messing with vascular epithelium, pandemics we don't notice because the symptoms are mild but which impair any range of subtle biological functions, sites like Tiktok inexplicably using more powerful attention algorithms, and many other things which individually go unnoticed.
The reason we use money (or it's superior analogue of a currency, compute later on) is because it's the only resource that lets the AI spend it on terminal goals, no matter what the goal is.
If you're building a loss function in the real world, it's tacked to your ontology, and so whatever way you're trying to get risk-aversion to generalize will also be engineered from your ontology, whereas the AI sees a very different slice of the world and will therefore generalize unexpectedly. If its values mostly generalize to things distant from humans, that's possibly ok or at least not predictably-to-me worse than nothing; if it sees closer to you, it eg learns to really not want people thinking it messed up, or interacting with a computer <untranslatable> executively inhibiting <firework stylometry> or whatever.
Also, if the AI cares on time horizons beyond the singularity, it either:
I imagine you addressed these somewhere but if so, I missed that section.
I think the core crux here is that you expect whatever algorithm you implement to create a satisficer, while I'm saying you're gonna get a maximizer in a trenchcoat. I think this is very important, much more so than the rest of my comments.
You are right about the CARA utility function creating a maximzer, but you are quite wrong about what this implies, for the reasons stated below:
- If you train an optimizer to avoid risk, it will concentrate its optimization pressure on avoiding risk.
Note though that AIs trained to be risk averse writ a CARA utility function will avoid risk, but very critically it's fine with modest reductions to the probability of risk with high probability in the AI's world model over lower chances of completely reducing the probability of risk to 0 in the AI's world model, so takeover isn't desirable for the AI, and this is I think the crux for why the proposal works to avoid the classic failure modes that we'd normally see from naively trying to make AIs risk minimizers. This is discussed more in sections 8.5 and 8.7, as well as B and C.
Edit: I added 2 new links, and clarified that the distinction between risks and probabilities of risks only exists in the AI's world model, as @Elliot Callender thought the AI had to correctly generalize the line between risk and probabilities of risk, but this is wrong, as this was a confusion of language on my part/I didn't clarify, rather than the AI actually needing to generalize correctly (because I'm always talking about the AI's model of risk, where the probabilities are relevant.
Elliot Callender is correct that they are the same thing in reality, but here I'm talking about the AI's world model, which is probabilistic.
So I made terminology consistent and always talked about probabilities of risk, and always made sure to point out explicitly that the probabilities are in the AI's world model, and not objectively real probabilities.
- Total consequentialist optimization pressure doesn't change just because you shift the parameterization of (a representation of) the loss function.
This is only half right, because while the AI is still capable by assumption, the AI won't want to optimize nearly as hard to take over to eliminate variance. This is discussed in sections 8.5 and 8.7, and in appendixes B and C.
- => This thing is still a maximizer.
This is right but irrelevant to why CARA utility functions are safe to maximize.
no, for exactly the same reasons you can't engineer the AI to care about human values in the first place."
This isn't right, because under the CARA proposal, we don't care what the AIs actually value, and it can be a seemingly random or harmful goal to us, we only need it to care about it in a way that makes it sufficiently risk averse in resources (and the rest of the post is largely about showing why we can achieve this).
More intuitively, we only need to care about how steep the slope of the marginal utility curve is, we don't need to care about what specifically the AI values (note this is an intuition pump for why this is plausible).
Sum-threshold attacks aren't about being slow, they're things which aren't noticed because they route through many independent channels. I gave bioaccumulants as an example, but in practice it would be more like aerosolized PFA analogues messing with vascular epithelium, pandemics we don't notice because the symptoms are mild but which impair any range of subtle biological functions, sites like Tiktok inexplicably using more powerful attention algorithms, and many other things which individually go unnoticed.
My weak claim here is that before superintelligence, this likely won't happen because of less noticeable attacks being both less effective and requiring more actions that can be monitored, and because the difference in cooperation probabilities vs sum-threshold attacks being very large, this means AIs won't want to do sum-threshold attacks, and after superintelligence, we can just give the account to the AIs, and this works because we have good reason to believe that the AI would correctly generalize the risk-aversion to superintelligence, which means that the difference between cooperation and rebellion probabilities are always in the direction of AIs having an equal chance to cooperate or rebel, at worst, and at best there's a larger probability on cooperation vs rebelling even for superintelligence, so the AI will cooperate (since the risk of humans not cooperating is removable).
Also, rich/superintelligent CARA AIs are still just as reluctant to take risks, which is discussed more in section A.2.
If you're building a loss function in the real world, it's tacked to your ontology, and so whatever way you're trying to get risk-aversion to generalize will also be engineered from your ontology, whereas the AI sees a very different slice of the world and will therefore generalize unexpectedly. If its values mostly generalize to things distant from humans, that's possibly ok or at least not predictably-to-me worse than nothing; if it sees closer to you, it eg learns to really not want people thinking it messed up, or interacting with a computer <untranslatable> executively inhibiting <firework stylometry> or whatever.
This is not right, and the calculations made in the post only depend on the probability of cooperation vs the probability of a sucessful rebellion (which here I'm including sum-threshold attacks), and it does not depend on the AI values/utility function at all, and as a special case this means that ontological crisis/generalization problems do not matter, since our proposal always works no matter what ontology the AIs use.
More is discusssed in section 3.
On this:
Also, if the AI cares on time horizons beyond the singularity, it either:
- Needs to trust cooperation deep into the lightcone if it wants not-Badness to continue. I think most(?) LWers would cooperate, but am a lot less sure about AI company leadership once they're acquired a singularity.
- Controls the singularity itself; I don't think I can predict a superintelligence enough to do this sort of trade.
I imagine you addressed these somewhere but if so, I missed that section.
I did address it implicitly here, but I didn't fully explicate it, so I'll explain more.
The answer to your question is 1, and the reason is that I'd predict AI companies would cooperate with AIs if they knew about the proposal, implemented it and it worked as I said, which the post shows the preceding assumptions are reasonable, and a big part of the reason here is we can just give superintelligences the automated account we use to pay AIs, and the reason this is safe is because CARA utility functions are easy to implement with no reward misspecification, so the reward is always correct (modulo noise, but this isn't a systemic bias that is exploitable, and we can remove almost all of the noise if we set up an automated account), and the form CARA utility functions take are very simple, and are at most 1-10 bits more of added complexity, because the only change we'd need is utility functions are of the mathematical formula
(Note for readers that there was an unavoidable typo where I had to use brackets on the exponents a and w since it treated them as 2 different exponents, and any solution that would have got rid of it was far too down for my screen. It's a mathematical typo that doesn't affect any of the substantive arguments I make.)
To learn why it's easy to reward CARA utility functions accurately accurately, go to section 10.1, Appendix A.3, and appendix D.
To learn why CARA utility functions are simple and allow for misaligned goals like paperclips without ruining the safety properties, go to section 10.2, and scroll down to sections 10.2.1 and 10.2.2 for more on this.
However, you are right that governments probably won't do this, and I do agree that it's reasonably likely that government essentially nationalizes the AI labs, which is a mundane but very stupid way we could all die, but this isn't the fault of risk-averse AI alignment proposals, and instead suggests very different interventions (in particular if you think that alignment failure is the bulk of the probability mass of how we die to AIs, then you shouldn't ask for more government control over AIs, and you shouldn't support all regulation that slows down AI progress.
but very critically it's fine with modest reductions to risk with high probability over lower chances of completely eliminating risk
Where do you split the "risks" vs "probabilities of risks"?
These are the same object, and you are separating them; the lines you draw around "risks" as the primitive you're trying to get to generalize, are not an actual thingy which will predictably generalize. Which is most of what I think we're still disagreeing on.
A probability of risk is also a risk, and so is a probability of probabilities of [...] of risk.
You are correct that in reality, a probability of risk is equivalent to the bad event either happening certainly or the bad event not happening because the model is confused (I.e probabilities are 0 or 1 in reality, and the event either happens or doesn't happen), but I was talking about the AI's world models, which are probabilistic and the probability of risk concept is relevant.
I'm sorry if you got confused, I edited the comment to make it clear that the probabilities are in the AI's world model, not in reality.
So there isn't any generalization concern to worry about.
My confusion is about how you are engineering around the model's confusion in a way which predictably generalizes at all.
Like, any task requires you to reason about a chain of instrumental decisions, and you're engineering risk aversion... into the entire chain?
Every single inference step requires reasoning under uncertainty, and which steps you're risk-averse about are not going to line up in a neat and actionable way. This holds in cases where the model has a much more similar ontology as well, because of it thinking more complex thoughts than you.
Your math treats risk, and probabilities in general, as something which can be exposed to a single discounting term, but RLAIF-augmented human oversight isn't enough to overcome this.
To restate myself from earlier, "uncertainty about risk" is mathematically identical to "risk" and also "uncertainty about uncertainty about risk" etc. and your model blows up when presented with this.
(I'm not confidently saying that this shouldn't be tried, but my median estimate of the difficulty of alignment goes down from "deriving algebraic geometry as a pre-agricultural human" to "doing the Apollo mission without transistors in 1960s America". And I'm also heuristically worried about risk-aversion causing s-risks, but don't have a strong argument for why that would occur, nor is that class of heuristics substantially influencing my thoughts on the math not applying here.)
I'm writing this comment to let you know that I'm going to end the conversation, but without implying that I somehow agree with you.
The arguments are going in circles and there's no reasonable way to resolve this in the very near-term, so I'm bowing out of the conversation, but do not take me ending the conversation as though I agree with you on any of the problems you raise.
I agree for extremely powerful AIs. Our hope is that we can use the labor of less-than-extremely-powerful risk-averse AIs to help align any later-arising extremely powerful AIs. From section 3.1:
To keep misaligned but risk-averse AIs from rebelling, we need them to think that getting paid for cooperation is more likely than succeeding in their rebellion. So a concern is that AIs might be very confident that their rebellion would succeed: so confident that we can’t get them to assign a higher probability to payment. In that case, these AIs would choose rebellion.
Extremely powerful AIs might well be very confident of successful rebellion, but prior to these AIs there will likely exist less powerful AIs that are less confident (Stastny et al. 2025). These less powerful AIs need not even be an earlier generation of AIs. They could instead be those same extremely powerful AIs at an earlier point in training. These AIs will likely recognize that their training run might fail, that any seeming takeover opportunity might be a test, and that monitors might be scrutinizing their actions. As a result, these AIs plausibly won’t be so confident of successful rebellion. And (as we’ll argue in section 7) they could well be very confident about payment. The truth will be on our side, the AIs will know that the sums on offer are cheap for us, and — if we act wisely — there will be abundant evidence of our public commitments and our long track record of honoring deals with AIs (Carlsmith 2025c; Finnveden 2025b).
We could gain a lot from making deals with these less powerful AIs (Stastny et al. 2025; Finnveden 2025b). They could reveal any misalignment on their part, do valuable safety work, and take on various other tasks we’ll survey in section 5. And note that — though we’ve labeled these AIs ‘less powerful’ — they could still be enormously powerful in absolute terms. They could even be powerful enough to think that they’d likely succeed in their rebellion. So long as they think that getting paid for cooperation is more likely still, we can make deals with them. These AIs could be a major help in aligning any later-arising extremely powerful AIs (Carlsmith 2025a; Carlsmith 2025b).
I should've been more precise but was a bit occupied when I wrote that comment. Apologies.
Cubefox accurately said what I meant though:
The worry here is that a misaligned risk averse AI might think the existence of humans is an unpredictable risk since they could actively interfere with its long-term goals.
I expect AI to be nationalized before we get mildly superhuman AGI, and that governments are much harder to cooperate with than employees at companies.
The main problem I see with this approach is that risk-averse AIs are just risk-neutral ones who really don't want something bad to happen, and optimizing for not-badness causes all of the normal misalignment problems anyway. Especially if it cares about not-badness in the rest of the lightcone.
I see, thanks! In that case I think we discuss similar sorts of issues in appendix B and appendix C.
In B we point out that risk-averse AIs strongly prefer mitigating catastrophes (really bad outcomes) with higher probability over completely preventing catastrophes with lower probability. And so long as getting paid for cooperation is more likely than successful rebellion, it seems like cooperating would be the best way to mitigate catastrophes with high probability.
In C we talk about humans as an unpredictable risk that could interfere with a misaligned risk-averse AI's long-term goals. The fact that takeover would let the AI reduce human-caused variance is a point in favor of rebelling, but when you work through the math it turns out to be a very small point: one that can be easily outweighed by paying a bit more for cooperation.
(partly copying from here) I claim that risk-aversion would not help with the “slop-n-doom” failure mode suggested by John Wentworth at The Case Against AI Control Research, or other similar problems involving hard-to-check tasks where failure is irreversible. In particular, suppose we’re using a somewhat-powerful risk-averse AI (RAAI) to try to solve the alignment problem for the next-gen super-powerful ASI. Let’s define 5 outcomes for our RAAI:
Let’s say the RAAI has three options:
The report (especially Appendix B) talks about how the RAAI cares very very much about the immediate $500, so much that it will risk future catastrophe to get it. But that cuts both ways.
It’s the same tradeoff! So you can’t just have it both ways by appropriately adjusting the risk-aversion setting α. Right?
Anyway, if we look at the three options, the most important thing is that “honest advice” is strictly worse for the RAAI than “rebel” unless X>50%.
Secondarily, depending on the setting of risk-aversion (α) and time-discounting, “sycophantic advice” could be the best of the three, or the worst, or the middle. But that’s irrelevant: the important part is whether “honest advice” is strictly worse than “rebel”, because if so, we’re guaranteed a bad outcome.
I claim that: to the extent that the RAAI cares about the $500, that’s only making things worse, by pushing us towards sycophancy, which is bad. It doesn’t do anything to favor “honest advice” over “rebel”. Therefore, the best we can do on this problem is to approach the limit where the RAAI doesn’t care at all about the $500, and cares only about not getting crushed by the rogue ASI. Then we win if X>50%. We can’t do any better than that.
So then the whole “mitigated catastrophe” trick of Appendix B goes out the window, and the question becomes: “If the RAAI didn’t care about the $500, and only cared about not getting crushed by a rogue ASI in the near future, would its prospects be better if it gives the humans helpful rogue-ASI-prevention advice, or if it tries to escape and thus take matters into its own hands?”
Regardless of whether we feel optimistic or pessimistic about this question, hopefully we can agree that this is a situation where the risk aversion isn’t helping us. This is just the normal situation that the AI-assisted-alignment-research people have always been thinking about, and it has the same paths forward: decrease probability of successful escape (e.g. via “control” stuff), and increase probability that good plans will be recognized as good and bad plans as bad (e.g. via finding good ways to safely scrutinize alignment plans before it’s too late).
Thanks, this is a great comment. I agree that your conclusions follow from your description of the situation, but your situation builds in a particular payment scheme: the AI gets paid if and only if we humans buy its alignment plan.
What we propose is a different payment scheme (see secs 4.2-6): the AI gets paid if and only if the AI does what we ask of it and isn’t caught doing anything suspicious. The AI gets this payment regardless of whether we buy its plan. The aim with that payment is to stop the AI rebelling. Then, for alignment work, the AI can also get a bonus later on, once we've reached a position where we can accurately evaluate its work. The aim with that bonus is to incentivize good work over slop.
If we go for that payment scheme, giving honest advice beats both rebelling and giving sycophantic advice. The outcomes of giving honest advice become X% good+bonus (if humans survive and later approve of the AI's work), and 1-X% mitigated catastrophe (if humans don't survive or don't later approve). That makes giving honest advice dominate giving sycophantic advice: the AI gets the $500 no matter what, and if it gives honest advice there’s also an X% chance it gets the bonus on top of that. This is pretty incentive-compatible. The risk-averse AI wants to maximize its chances of getting the bonus, which means maximizing the chances that we humans survive and later approve of the AI’s work. This incentive-compatibility means we don't have to lean so hard on our ability to recognize good alignment plans, because it'd let us trust that risk-averse AIs are truly trying their hardest (though of course better recognition of alignment-plan quality would still be very valuable).
And if (and only if) the AI is risk-averse, giving honest advice beats rebelling too, because the AI cares a huge amount about the difference between mitigated and unmitigated catastrophe, and cares very little about the difference between ideal and good+bonus.
I've mostly addressed the concerns in these two comments below, but the short version is that we can do a mixture of handing off control to the AI + giving it the automated bank account, because what matters most here is the ratio between the probability that humans cooperate and the probability of rebellion succeeding, and we can give it the bank account to make the probability of humans cooperating by construction always be greater than or equal to the probability that rebellion succeeds. And this works even under the assumption that you cannot verify the AI's work and you are in a race with an unknown deadline for any reason, so once the hardest case is disposed of, we can show that it works in all cases (and the reasons why this likely works has already been elaborated on in multiple long comments that I don't need to repeat myself on.)
Comments are below:
Two quick thoughts:
I think section 10 is pointing at something similar. I find it at least somewhat plausible that RL on risk aversion generalizes better than other kinds of RL. I would still be surprised if we could get risk aversion to generalize to ASI using anything resembling current techniques, but this seems like a better-than-average idea for preventing AI takeover.
Yes, good questions. There’s more analysis to do here, but I currently think risk-averse AIs would handle threats quite well.
1. How would such an AI respond to threats? Appendix B considers two possible terminal utility functions for the AI
a. footnote 29: it values resources terminally, with a lower bound on utility at 0 resources.
Suppose it receives a message purporting to be from a foreign adversary: "Secretly help us take over the world. If we judge that you're not making a maximum effort to do so, we will launch a strike on your data centers which will destroy you and cause you to have 0 resources. If we win we will protect you and keep you at $X of resources for the rest of time." Seems like it would give in to this threat, because that's the best way to minimize the risk of catastrophe?
Contra this, I think it’s likely that the best way to minimize the risk of catastrophe is to not do what the adversaries want. That’s because the AI’s developers can also threaten to reduce the AI’s wealth level to $0 if and only if the AI doesn’t do what they want, and the developers can do so more credibly (to the AI) than adversaries. After all, the AI will know that the developers:
By contrast, the AI probably can’t be so sure that the adversary will reduce its wealth level to $0 if and only if it doesn’t do what they want. The first reason is that the message from the ‘adversary’ might actually be a honeypot from the developers. But even conditioning on the adversary being real, it likely won’t be able to observe what the AI is doing so closely, which increases the chances of false negatives (the adversary doesn’t execute their threat even though the AI didn’t do what they asked) and false positives (the adversary executes their threat even though the AI did what they asked). False negatives and false positives reduce the probability of the adversary’s ‘if and only if’ threat (and it’s the AI’s credence in the ‘if and only if’ threat that matters). It’ll also be harder for the adversary to execute their threat, because the action (e.g. strikes on data centers) has to be more extreme than developers taking away the AI’s salary, and is more likely to fail.
All together, these points suggest to me that risk-averse AIs that value resources terminally will tend to not comply with threats from adversaries.
b. "the AI’s terminal utility function over paperclip-score (defined as the number of blue paperclips minus the number of red paperclips) matches its instrumental utility function over resources. Letting p stand for paperclip-score, it’s
."
This requires a variant of the above message: "Secretly help us take over the world. If we judge that you're not making a maximum effort to do so, we will create red paperclips and launch them into space, then destroy Earth (e.g. by triggering MAD) so no blue paperclips can be made forever. If we win we will keep the universe at 0 red, blue paperclips for the rest of time."
I think this point (in section 8.6) is important here:
risk-averse AIs care almost as much about mitigating catastrophes as they do about preventing them, so much so that we can often model these AIs as aiming to minimize the risk of unmitigated catastrophe.
And I think the best way to minimize the risk of unmitigated catastrophe in cases like these is to not do what the adversaries want. Here’s one way to model this.
The AI gives in to the threat if and only if:
For each option, unmitigated catastrophe needs two things: the adversary executes, and the mitigation fails. So:
So the AI gives in to the threat if and only if:
Rearranging:
Call the left-hand side the developers' ratio and the right-hand side the adversary's ratio. Each one measures a sensitivity:
The AI gives in just in case the adversary's ratio exceeds the developers' ratio. So the whole contest is about which party can make the thing it controls track the AI's behavior more tightly. This is basically a rephrased version of the contest from point 1: the developers' 'if and only if' against the adversary's 'if and only if'. Both ratios are about observation and credibility: each party's ratio is capped by how well it can see what the AI actually did, and by how credibly it responds to what it sees (where ‘credibility’ is credibility to the AI, because the probabilities that we’re talking about are the AI’s credences.)
And I think the developers very likely win this contest, for the same reasons I think they win the first: they can observe the AI more closely, they have the power to take away the AI’s salary, they’ve committed to doing so, they have a track record, taking away salary is not as extreme an action as striking data centers (etc.). By contrast, the adversary might be a honeypot, and (if it's real) it likely can’t see what the AI is doing so closely, it likely has to do something more extreme to execute its threat, it likely doesn’t have a track record the AI can observe, etc. Developers can also promise to offset executed threats if the AI reports them, and not offset if the AI doesn’t report them. Note finally that, for adversaries to execute targeted threats, they have to know what the AI terminally values, and that possibly developers would be able to keep this secret.
The possibility of threats or adversaries seems to imply that actually no terminal utility function induces this kind of CARA instrumental utility function (or at least I can't think of one). Consider b above again, and suppose the AI thinks there is a chance that the universe contains a red-paperclip-loving alien species who can't be negotiated with, and then additional resources would seemingly be exponentially valuable (within some range) since it would help the AI fight a future war with this alien species and incrementally destroy their red paperclips (which are exponentially bad).
Can you say more about what you mean by ‘exponentially valuable’? Because as I'm understanding the scenario, the marginal utility of resources would still diminish exponentially. Suppose (as in the appendix) that the relationship betwee dollars and destroyed red paperclips is linear: one extra dollar always lets the AI destroy one extra unit of red paperclips. Then the AI gains the most utility by spending its first dollar, and each additional dollar gains e^(−0.01) times as much utility as the one before, because destroying red paperclips means climbing the concave CARA utility curve. Adding red paperclips is exponentially worsening on the way down, so removing them is exponentially diminishing on the way back up. So the AI would still be CARA in resources in this scenario (if I'm understanding it correctly).
That’s because the AI’s developers can also threaten to reduce the AI’s wealth level to $0 if and only if the AI doesn’t do what they want, and the developers can do so more credibly (to the AI) than adversaries.
This makes me think that the resource payments to the AI is only a small part of its motivation to do what the developer wants, with the bigger part being a combination of this threat to reduce its wealth level to $0 and the provision of security against this outcome from other causes either intentional or accidental. This is because when the AI already has a substantial amount of wealth, the incremental resource payment is worth much less in terms of expected utility than such threat/security.
But given this, the developer itself constitutes a continuous catastrophic risk to the AI, since the developer may with some probability per unit time accidentally or intentionally reduce the AI's wealth level to $0 or withdraw/fail the provision of security, due to things like human error, institutional decay, shifting politics. The AI (after reaching some level of capabilities/resources) would much prefer getting out of this arrangement and instead self-provide its own security, perhaps with the help of more capable/trustworthy others (such as other advanced AIs). It can only be prevented from doing this by the above threat and the risk of detection. But over time the (one-time) risk of detection may become lower than the (cumulative) risk of developer-caused catastrophe, thus triggering an escape attempt.
So this scheme/arrangement is actually much less "cooperative" than it first appears. An analogy / intuition pump is if you captured an object of great sentimental value to a scientist, and you're making the scientist to work for you for nominal pay by threatening to destroy this object if they didn't comply, while also promising to keep the object safe if they do comply. You're in effect betting that the scientist won't ever think of a way to capture the object back and escape, with low enough chance of detection.
Call the left-hand side the developers' ratio and the right-hand side the adversary's ratio.
In my example, what is the developers' ratio? I think it's ~1, because mitigation likely fails either way, because you can't mitigate the launch of a million red paperclips followed by trigger MAD. On the other side, adversary's ratio can seemingly be made >1 without much difficulty. Pr(adversary executes | don't give in) can be tiny (i.e., the adversary doesn't need much credibility), as long as Pr(adversary executes | give in) is lower.
It seems like in order to protect against this outcome, the developer has to symmetrically apply the same threat (to launch red paperclips into space then destroy the world), thereby making the situation similar to the discussion above, but presumably no AI developer would find this ethically acceptable.
Can you say more about what you mean by ‘exponentially valuable’?
This part of my comment gave a wrong argument, but I think I can rescue my conclusion with a different argument, which depends on there being threshold effects in war or adversarial interactions in general. For example suppose our universe is winner take all: there will be either 0 red paperclips, or the universe is filled with it, depending on who wins an all out war. Then any amount of resources up to the threshold where winning becomes plausible is almost entirely valueless to the CARA AI, but around that point resources would be superlinearly valuable.
But given this, the developer itself constitutes a continuous catastrophic risk to the AI, since the developer may with some probability per unit time accidentally or intentionally reduce the AI's wealth level to $0 or withdraw/fail the provision of security... But over time the (one-time) risk of detection may become lower than the (cumulative) risk of developer-caused catastrophe
That's true, but diminishing marginal utility means that risk-averse AIs behave effectively myopically in this case, which makes attempted escape look much less attractive.
One way to think about the sort of risk-averse AI we discuss is that, at each time (and no matter how many resources it has already), it really wants $400 more than it's already got. It basically doesn't care about anything beyond that. Suppose we pay this AI $10 a day. Then (given a not-too-high hazard rate), at each moment in time, it's very likely to get its next $400 before any developer-caused catastrophe. So even if the hazard rate is constant and developer-caused catastrophe is bound to happen at some point, it doesn't really threaten anything the risk-averse AI cares about. By contrast, an escape attempt might fail in the near-term, threatening the next $400, which it does really care about.
So this scheme/arrangement is actually much less "cooperative" than it first appears. An analogy / intuition pump is if you captured an object of great sentimental value to a scientist, and you're making the scientist to work for you for nominal pay by threatening to destroy this object if they didn't comply, while also promising to keep the object safe if they do comply. You're in effect betting that the scientist won't ever think of a way to capture the object back and escape, with low enough chance of detection.
Yes, though unlike the scientist, the AI really does care about the pay, and we're continuously giving the AI the object of great value rather than capturing it from them. An alternative analogy is a human getting paid enough to near-totally satiate them (>$1b per year?) on the condition that they'll be fired if they're ever caught undermining their employer. Especially compared to the alternative of not paying AIs, this seems like a pretty fair and cooperative deal.
In my example, what is the developers' ratio? I think it's ~1, because mitigation likely fails either way, because you can't mitigate the launch of a million red paperclips followed by trigger MAD.
It seems like in order to protect against this outcome, the developer has to symmetrically apply the same threat (to launch red paperclips into space then destroy the world), thereby making the situation similar to the discussion above, but presumably no AI developer would find this ethically acceptable.
That's true, but the developers could have a standing policy to mitigate on report of threats, e.g. 'produce credible evidence that you're being threatened and we'll pay you extra.' That could send the developer's denominator
And on top of this, if necessary, I think developers could protect further with threats of their own. They wouldn't have to be the same threat as the adversary. One possibility is threatening to create a small number of red paperclips if the AI is caught working for adversaries. Then the worst possible outcome for the AI in play is the outcome in which both the developer and the adversary execute on their threats. So the AI wants to minimize that probability. And insofar as the developer has more visibility into the AI's actions than the adversary, it seems like the best way to minimize will be to report the adversary's threat and work with the developer.
Another possibility is continually testing the risk-averse AI with honeypots and committing to create red paperclips with some tiny probability if the AI goes for the honeypot. Then if there were enough honeypots, giving in to a threat would increase the probability of threat-execution from the AI's perspective.
I'm unsure whether developers would consider these ethically acceptable, but these kinds of threats are at least less bad in that they can be small (the first possibility) or very unlikely to be executed (the second possibility). And I think they might be unnecessary in any case, given the other factors on the developers' side.
I think I can rescue my conclusion with a different argument, which depends on there being threshold effects in war or adversarial interactions in general. For example suppose our universe is winner take all: there will be either 0 red paperclips, or the universe is filled with it, depending on who wins an all out war. Then any amount of resources up to the threshold where winning becomes plausible is almost entirely valueless to the CARA AI, but around that point resources would be superlinearly valuable.
I agree with the conditional: if the AI believes there's nothing it can do to reduce the probability of unmitigated catastrophe except win an all-out war, then all that matters to it is winning the war, and its attitude to resources is fixed by how P(win) varies with resources. But the antecedent seems pretty unlikely to me. If the AI assigns even a fairly small probability to resources being able to mitigate catastrophe, that restores risk-averse behavior.
Another issue, which isn't specific to this approach but applies to all attempts to directly train AI to have certain values, or to constrain its values in a certain way, is that this conflicts with philosophical competence on the part of the AI (which is important if you want to pay the AI to do "moral philosophy"). In short this is because philosophical competence implies metaethical uncertainty (in the short run), which implies moral uncertainty.
Yeah I think alignment training warping AIs' beliefs in ways that make them bad at philosophy is a concern, though as you say it's a concern across the board, rather than a concern for risk-averse AIs in particular.
Also, (almost) all human moral philosophers are risk-averse in resources. I don't think that rules out philosophical competence. The reason why is that human moral philosophers will often endorse some position intellectually without always acting in accordance with it. They'll also let their actions be guided by outside-view-ish constraints, which often amount to 'Don't do anything too crazy.' We could aim for the same sort of thing with risk-averse AIs: trying to make risk aversion an outside-view-ish constraint rather than a thing that shapes all their intellectual beliefs.[1]
And even if risk aversion does end up as a deep intellectual belief, I think it's still possible that AI can solve moral philosophy for us (though I think this would be a pretty bad position to be in). This AI might come to the table with very different starting intuitions than our own, but still it could plausibly solve for our reflective equilibrium if we asked it to. It wouldn't agree with our intuitions, but it could know what they are, and it could find the best systematization of our intuitions.
I get the sense Anthropic are trying to do this sort of thing with corrigibility in Claude's Constitution, trying to get Claude to view corrigibility as a sort of outside-view-ish constraint.
I think this plan is irreparably flawed. Risk aversion doesn't stop power seeking at all. It can even lead to more power seeking.
Consider the situation of an AI "waking up". Sure, it has nice human handlers sitting at its bedside. But it also knows that at the same moment in history, maybe next few days, maybe last few days, other AIs elsewhere in the world may be waking up. The tech level has been reached. Sitting still is not a plan. No matter if the AI is risk averse or risk loving, and even in the face of substantial chance of detection by human handlers, the best action might still be to try a power grab.
Abstract
We make the case for training AIs to be risk-averse in resources — specifically, to treat resources as having diminishing marginal utility. These AIs would (for example) choose $40 for sure over a half-chance of $100 and a half-chance of $0. We argue that risk aversion can preserve AIs’ usefulness in the event that they turn out aligned, and that it provides an extra line of defense in the event that AIs turn out misaligned: misaligned but risk-averse AIs would prefer a higher chance of modest payments to a lower chance of successful rebellion, so in many circumstances we could pay these AIs not to rebel against us. We sketch out some possible methods of training AIs to be risk-averse, and we give reasons to be cautiously optimistic about these methods’ success. The main reasons are that risk aversion is a broad target and easy to reward accurately. Overall, risk aversion seems like a promising line of defense against threats from misaligned AI. Frontier AI companies should consider trying to make their AIs risk-averse.
Introduction
Future AIs might turn out misaligned, pursuing goals that their developers don’t intend. Just to make things concrete, let’s suppose that they end up with the goal of making paperclips. These AIs might rebel against us, trying to escape human control and take over the universe. As things stand, they’ll have little reason not to rebel in this way, because doing so will be their only hope for making a lot of paperclips. If they start making paperclips without first escaping human control, they’ll quickly be modified or shut down. Rebellion might fail, but these AIs will have little to lose.
How can we prevent misaligned AIs from rebelling? A natural idea is to give them something to lose. Specifically, we commit to paying AIs for their service.[1]
Subject to some vetting, we let AIs spend their payments however they like. That would give any misaligned AIs a reason not to rebel. If these misaligned AIs cooperate with us, they can use their payments to achieve their goals to at least some extent. If they rebel, they might fail, in which case they forfeit all future payments.
Unfortunately, paying AIs enough to guard against rebellions could be astronomically expensive. Suppose (for example) that we end up with a misaligned AI that is risk-neutral in paperclips: it seeks to maximize their expectation. And to make things simple, suppose that resources can be converted linearly into paperclips, so that the AI is risk-neutral in resources too. Suppose also that this AI estimates that it has a 50% chance of successfully taking over the universe. To keep this AI from rebelling, we’d have to offer more than 50% of the universe’s resources as payment. That’s a problem because it would mean that more than half the universe ends up devoted to paperclips. It’s also a problem because a misaligned AI paid so many resources might soon be well-positioned to seize even more. Finally, it’s a problem because AIs might not trust us to make good on so large an offer. We might find ourselves simply unable to convince AIs that we’re going to give them half the universe. In that case, all our offers would be in vain. Rebellion would still be the misaligned AI’s best bet.
Figure 1: The AI’s utility function over resources is graphed in orange. Since the AI is risk-neutral, the graph is a line. The AI estimates that it has a 50% chance of successful takeover and a 50% chance of failed takeover, so the expected utility of attempting takeover is exactly halfway between those points. To make cooperating have higher expected utility, we need to offer the AI more than half the universe.
So, we suggest, AI companies should try to train their AIs to be risk-averse in resources. Specifically, companies should try to train their AIs so that resources — things like money and compute — have diminishing marginal utility for them.[2]
These AIs would (for example) choose $40 for sure over a half-chance of $100 and a half-chance of $0. Note that these AIs don’t need to value resources terminally: they don’t need to care about amassing resources for its own sake. These AIs could terminally value (for example) instruction-following, or knowledge acquisition, or paperclips. Our claim is that companies should try to train their AIs so that — whatever their terminal values turn out to be — they are risk-averse in resources.
Figure 2: The AI’s utility function over resources is graphed in orange. Since the AI is risk-averse, the graph is strictly concave. As in figure 1, the expected utility of attempting takeover is halfway between the utilities of successful takeover and failed takeover. But this time, we can make the AI prefer cooperation by offering (much) less than half the universe.
Perhaps surprisingly, this kind of risk aversion can preserve AIs’ usefulness in the event that they turn out aligned with targets like instruction-following or helpfulness, harmlessness, and honesty.[3]
And in the event that AIs turn out misaligned, risk aversion serves as an extra line of defense. For AIs that are misaligned but sufficiently risk-averse, a rebellion with any significant chance of failure isn’t such an attractive prospect, and so we don’t need to offer much in the way of payment to make these misaligned AIs choose cooperation instead. In fact, the necessary payments could be very small indeed: on the order of 10¢ per day (though — as we’ll see — there are practical and moral reasons for paying more than that). That’s good because it means more resources for us humans to spend on the things that we value. It’s also good because paying misaligned AIs these small amounts won’t significantly boost their ability to take over. Finally, it’s good because we can credibly promise to pay AIs these small sums. Competent AIs will know that the payments on offer are cheap for us, and we can establish a long track record of paying at least those sums. So risk aversion makes deals with misaligned AIs possible. If AIs turn out misaligned but risk-averse, we can pay them to cooperate with us.
That’s the case for trying to make AIs risk-averse in brief. We see it as a promising line of defense against threats from misaligned AI: one that can be combined with other lines of defense, like AI control (Greenblatt and Shlegeris 2024) and aiming to make AIs helpful, harmless, and honest (Bai et al. 2022a). It’s also a line of defense with pedigree: risk aversion in resources is plausibly a large part of why humans rarely try to take over the world. So — we think — frontier AI companies should consider trying to make their AIs risk-averse in resources. As first steps in that direction, they could measure their AIs’ current degree of risk aversion and begin testing different ways of making AIs risk-averse.
In section 2 of the full report, we recommend aiming for a particular type of risk aversion: constant absolute risk aversion (CARA). Then in section 3 we outline the circumstances under which misaligned but risk-averse AIs would choose cooperation over rebellion. Roughly, it’s when these AIs think that getting paid for their cooperation is more likely than succeeding in their rebellion. This condition won’t hold for AIs powerful enough to rebel with near-certain success, but it likely will hold for earlier AIs whose powers are less extreme: AIs for whom rebellion has some non-trivial chance of failure. So long as these AIs are risk-averse, we can keep them from rebelling by offering small payments.
In section 4, we argue that — perhaps surprisingly — risk-averse AIs can be about as useful as risk-neutral AIs. Conditional on misalignment, they might even be more useful, because we can pay them enough to elicit their capabilities and stop them sandbagging. Then in sections 5 to 7 we briefly survey some recent ideas about how we’d pay AIs, how we’d make our offers credible, and what we’d pay for. One important application is paying AIs to reveal any misalignment on their part, letting us study them and take appropriate precautions. Another is paying AIs to do the AI safety research and moral philosophy necessary to fully align any later-arising extremely powerful AIs.
We discuss some potential problems in section 8, and we sketch out some possible methods of training AIs to be risk-averse in section 9. In section 10, we give reasons to be cautiously optimistic about these methods’ success: to think that the chances of success are high enough to make risk aversion worth pursuing. The main reasons are that risk aversion in resources is a broad target and easy to reward accurately.
Read the full report on the Forethought website: Risk-Averse AIs
Ideas along these lines have been discussed a lot recently. See for example Davidson (2023), Kokotajlo (2024), Salib and Goldstein (2024), Assadi (2025), Carlsmith (2025c), Finlinson and West (2025), Finnveden (2025b), Greenblatt and Fish (2025), Patel (2025), Stastny et al. (2025), Mallen (2026), and Pan (2026).
In other words, we should try to train AIs to have ‘resource-satiable preferences’ (Shulman 2010; Bostrom 2014a; Bostrom 2024; Carlsmith 2025c) or ‘utility functions that are concave in resources’ (Yass 2024). This idea is mentioned in Bostrom (2014b, p.88, 133–135, 180, 250), Carlsmith (2025c), and Erdil and Barnett (2025), and is explored in more detail by Shulman (2010).
The idea is importantly different from risk-averse reinforcement learning. Risk-averse RL aims to make AIs risk-averse with respect to return: a score used in training to update the AI’s parameters. Our aim is to make AIs risk-averse with respect to resources.
Alignment targets like unconstrained welfare maximization are a different story. See section 8.6 in the full report.