x
When talked into harm, a model blames the answer, not itself (an interpretability study of guilt vs shame) — LessWrong