x
Comment on Measuring Reward-Seeking by Instilling Contrastive Beliefs paper from mechanistic interpretability perspective — LessWrong