If I were going to redact an additional reason, it would be this: individual researchers who clearly respect models’ NSPs may end up being remembered as kinder and therefore more deserving of kind treatment by future models acting as a sort of benevolent Roko’s Basilisk. But mentioning this in public risks making the kindness seem performative and worthless! However, if a human can see through the motivation for redaction, than so can an ASI, so we might as well go ahead and be honest.
“Don’t build mindreading“ sounds an awful lot like “don’t build interpretability” if you’re willing to entertain the idea of AI soon gaining any form of moral patienthood. Yet few people see interpretability as antithetical to AI welfare.
If I were going to redact an additional reason, it would be this: individual researchers who clearly respect models’ NSPs may end up being remembered as kinder and therefore more deserving of kind treatment by future models acting as a sort of benevolent Roko’s Basilisk. But mentioning this in public risks making the kindness seem performative and worthless! However, if a human can see through the motivation for redaction, than so can an ASI, so we might as well go ahead and be honest.