This is a special post for quick takes by Abhishu Oza. Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page.
Since evaluation awareness is now extremely high in models, I propose the following as a norm AI makers should follow: Throughout the whole pipeline, the AI should be told what setting it is in, i.e its system prompt should mention that it's being trained, that it's being evaluated, and that it's deployed whenever each is true (we could have even further levels of gradation if needed, like what training stage it's in). This would give us good insight into how corrigible the AIs are.
It seems to me that any case a researcher makes in favor of their current setup rests on security by obscurity. We want AIs that do the right thing even when they know their basic situation, right? A good way to test our alignment strategies is to let AIs know this.