Measuring Activation Control in LLMs
TL;DR * Inspired by the introspective awareness and CoT controllability papers, we made a benchmark to measure how well models can control their activations while completing a simple task. * We are motivated by the concern that highly introspective models could control their activations, confounding probes and other monitors, and...
Aug 1313