Restoring Model Alignment via Honesty Activation Steering
by niklas_herbster, Martin Zborowski, Alberto Tosato, Gauthier Gidel, and Tommaso (Tommie) Tosato
TL;DR * We introduce two projection-aware steering methods (StTP, StMP) that intervene only on tokens whose activations fall on the misaligned side of a learned decision boundary. They restore honesty about as well as the classic uniform steering, while largely avoiding capability degradation. * We find a single honesty direction,...
Jul 2018