We've found that an effective way to see whether models are cheating is to just branch their context and ask "what are you currently working on". This somehow completely snaps them out of their RL-gremlin-haze and they just give super honest answers. My hypothesis for this is that the models have no incentive to ever ask a question like this during a deep RL episode, since it makes it less likely for them to get the reward. So they've essentially never had a trajectory which is "hard RL-mode, hard RL-mode, stop-to-reconsider" get positively reinforced.
Obviously, this would fail miserably against real scheming. But it seems to work quite well against "stuck in a rabbit hole and not considering the world around it" scheming
I worry that as synthetic documents become a larger portion of the models' training load, this might make it harder to do prosaic alignment. Part of what tethered the models to human values was that they were trained on the sum of human experiences. If they instead start being trained on models' impressions of human experiences, then it seems likely this will have downstream consequences. This seems especially true when extrapolated out: what happens when we have models which were trained on synthetic documents, and these synthetic documents were produced by models trained on synthetic documents, and these synthetic documents were produced by models which were... (and so on).
Do data vendors for the frontier labs (pre-train data; RL envs) get security screening? If not, this seems necessary to start instituting by law asap. If someone were trying to poison a frontier model, their options are to:
1. Scatter poison across the internet and pre-training base and hope the labs pick this data up
2. Bribe a person at a data vendor which has a contract with a frontier lab
Between the two, the second option seems cheaper, easier and to have a higher likelihood of working in the current environment. But it's also pretty easy to stop via legislation/executive order, since we have an established security clearance process. In general, it seems like we're putting a lot more scrutiny on the labs recently, but not nearly enough on these data vendors?