Importantly, the model shouldn’t know that it’s being resampled so it doesn’t learn that the monitor thought it looked sketchy.
It seems important to state here that it's not the model U shouldn't know, but rather the authors assume the model is stateless. As the authors later state in "Note on Statelessness": "If we want independent samples of the next action, we make multiple calls to the model using the same context, and the model’s statelessness between calls means that it doesn’t know it is being resampled". It's interesting that Statelessness seems to be an advantage for safety control, although it might harm usefulness (would that be the correct intuition?).
It seems important to state here that it's not the model U shouldn't know, but rather the authors assume the model is stateless. As the authors later state in "Note on Statelessness": "If we want independent samples of the next action, we make multiple calls to the model using the same context, and the model’s statelessness between calls means that it doesn’t know it is being resampled". It's interesting that Statelessness seems to be an advantage for safety control, although it might harm usefulness (would that be the correct intuition?).