That means, if we’re being optimistic, we could use it to tell us what the model would do if it wasn’t eval aware even if we have no evals we’re sure are realistic enough
I'm not super convinced this is possible -- constructing the steering vector via contrastive prompt pairs requires prompts that the model is not eval aware on, but as models scale, it seems like this sort of negative prompt will either a) not exist or b) not generalize well enough to actually steer with (eval awareness seems to have more than one axis, i.e. I don't think we can just steer... (read more)
This sounds quite cool, excited to read the paper!
I'm interested in Lesson 4 related to TCW and training on descriptions of model behaviors / tendencies versus training on examples of the model demonstrating the traits, and have a few questions:
- since you hypothesize that RL might cause the model to update against priors acquired via SDF, did you qualitatively or quantitatively test the model before RL?
- how well did the models embody the traits after SDF-SFT on chat samples (not agentic)?
- do you expect the difference between TCW-style and what you did to have
... (read more)