If a model has an internal state that changes how it behaves we investigate whether it can talk about that state.
In our new post, we test this for a functional welfare direction found in a small open model before and after RL. “Functional welfare” here means an internal direction associated with better or worse task performance that can shift sentiment and behaviour. We use the Jacobian lens to ask whether that direction sits inside the model’s verbalizable subspace.
Things that stood out
1) The self report result failed its controls thus speakability doesn’t establish reliable introspection.
2) The welfare directions were already partly speakable before reinforcement learning.
3) Training amplified the distress related direction’s speakable share more clearly.
We’d value criticism of the speakability measure, the control design, and evaluating model self-reports.
If a model has an internal state that changes how it behaves we investigate whether it can talk about that state.
In our new post, we test this for a functional welfare direction found in a small open model before and after RL. “Functional welfare” here means an internal direction associated with better or worse task performance that can shift sentiment and behaviour. We use the Jacobian lens to ask whether that direction sits inside the model’s verbalizable subspace.
Things that stood out
1) The self report result failed its controls thus speakability doesn’t establish reliable introspection.
2) The welfare directions were already partly speakable before reinforcement learning.
3) Training amplified the distress related direction’s speakable share more clearly.
We’d value criticism of the speakability measure, the control design, and evaluating model self-reports.
PDF
Code and Hugging Face