Quantifying CoT faithfulness
TL;DR > In our setting, we find that behavioral value leakage and disclosure in chain-of-thought do not move together. Making a charitable motive explicit substantially increases the model’s acknowledgement of donation influence. Adding an accuracy-focused instruction largely suppresses acknowledgement while the behavior persists. Our interventions suggest that the tokens expressed...
Sep 13