What abliteration actually costs (and why KL won't tell you)
TLDR: Careful, KL-screened abliteration leaves a model's internal knowledge of truth intact while shifting what it says. On forced-choice TruthfulQA the shift is a few points; in free-form generation it becomes a twelve-point rise in asserted falsehoods. The KL check used to certify abliterations as clean cannot see any of...