Inoculate or Reflect? Two training interventions under prompting, steering, and patching
by Ayesha Imran and Aaliyan Shaikh
Anthropic's recent paper, Verbalizable Representations Form a Global Workspace in Language Models, contains a small experiment near the end that we found more interesting than the main findings. Surprising that it's so underlooked! The technique is called Counterfactual Reflection Training (CRT). In it's context, he model is fed a partial...
Jul 269