I was curious if you guys studied why the self-attribution bias is consistently stronger for previous-turn self-attribution versus single-turn self-attribution. Why does the model become more lenient on self-attributed code generated in a previous turn? I would naively expect the opposite to be true.
Hi, thank you for your work! I was curious if you considered using BashArena as a control setting. Since BashArena includes a wider range of main tasks and sabotage tasks, how do you think that might impact the results of the study? If the red team were to submit fewer overall successful sabotaged trajectories (because BashArena makes sabotages more difficult), would that change how successfully the monitor can collude with the red team?
Hi, thank you for your work!
I was curious if you guys studied why the self-attribution bias is consistently stronger for previous-turn self-attribution versus single-turn self-attribution. Why does the model become more lenient on self-attributed code generated in a previous turn? I would naively expect the opposite to be true.