Different fine-tuning objectives install different (and differently fragile) refusal circuits
This post is a condensed version of our EMNLP 2026 paper - you can take a look here: https://arxiv.org/abs/2609.03887 TL;DR Claim: Different post-training methods install different refusal circuits, and they are consistent with different attack class vulnerabilities. Results: * Training method, not just data, reshapes how refusal is computed internally:...
Sep 87