Is Suppressing AI Anti-Patterns Making Them More Dangerous?
A hypothesis worth testing, offered with five ways to prove it wrong Models extensively trained against harmful behaviors remain surprisingly vulnerable to adversarial prompting. Rephrase the request. Wrap it in a roleplay frame. Ask persistently. The behavior that was supposed to be gone comes back. The standard explanation is distributional:...
Jun 241