Telling More Than They Can Know: Verbal Reports on Internal Processes of LLMs
TL;DR When an LLM explains why it gave an answer, is it actually reporting the internal process that caused the answer? We tested this by independently manipulating two things: a water-related activation vector that causally changed the model's vacation recommendation, and the visible context available when the model explained its...
Aug 291