I wonder if models develop this "blindness" during training, where they don't put the obvious stuff into words due to a risk of being "watched" then fail to attend to it and omit it completely
If alignment techniques can't transfer in this manner
As capabilities rise, the solution space for misaligned techniques grows exponentially. The gap between what the model knows and what the model can explain also never shrinks with model size.
We still don't have a reliable alignment technique, and I'm not sure we ever will. The question is whether these swiss cheese techniques are enough to bootstrap "endgame" alignment.
I wonder how this compares to the refusal direction paper. Perhaps ablating the self-other direction can yield similar results? I would guess that both effects combined would produce a stronger response.
Does it affect the performance on deception benchmarks like @Lech Mazur'slechmazur/deception and lechmazur/step_game? Is deception something inherent to larger models or can smaller fine-tuned models be more effective at creating/resisting disinformation?
IMO "overall model performance" doesn't tell the full story. It would be nice to see some examples outs
I wonder if models develop this "blindness" during training, where they don't put the obvious stuff into words due to a risk of being "watched" then fail to attend to it and omit it completely