TL;DR:

When is or isn't reward the optimization target? I use a mathematical toy model to reason about when RL should select for reward-seeking reasoning as opposed to behaviors that achieve high reward without thinking about reward.
Hypothesis: More diverse RL increases the likelihood that reward-seeking emerges.
In my toy model, there are scaling laws for the emergence of reward-seeking: when decreasing prior probability of reward-seeking reasoning we can increase RL diversity enough that reward-seeking still dominates after training. The transition boundary is a ~straight line on a log-log plot.
Opinion: To get to a hard "science of scheming", we should study the "physics of RL dynamics".

Motivation

Why would RL training lead to reward-seeking reasoning, scheming reasoning, power-seeking reasoning or any other reasoning patterns for that matter? The simplest hypothesis is the behavioral selection model. In a nutshell, a cognitive pattern that correlates with high reward will increase throughout RL.^[1]

In this post, I want to describe the mental framework that I've been using to think about behavioral selection of cognitive patterns under RL. I will use reward-seeking as my primary example, because it is always behaviorally fit under RL^[2] and is a necessary condition for deceptive alignment. The question is "Given a starting model and an RL pipeline, does the training produce a reward-seeker?" A reward-seeker is loosely defined as a model that habitually reasons about wanting to achieve reward (either terminally or for instrumental reasons) across a set of environments and then acts accordingly.^[3]

It is not obvious that reward-seeking needs to develop, even in the limit of infinitely long RL runs (see many excellent discussions). A model can learn to take actions that lead to high reward without ever representing the concept of reward. In the limiting case, it simply memorizes a look-up table of good actions for every training environment. However, we have observed naturally emergent instances of reward-seeking reasoning in practice (see Appendix N.4 of our Anti-Scheming paper, p.71-72). This motivates the question: was this emergence due to idiosyncratic features of the training setup, or can we more generally characterize when explicit reward-seeking should be expected to arise?

A mathematical toy model for behavioral selection

I will now introduce a s...

Deceptive Alignment

Deceptive Alignment