Cool work but something concerns me. The method only scores , so anything the two might share is invisible. The "clean reference" assumption is doing a lot of work here and only gets a weak check.
The Qwen censorship case is a nice lens because the reference's status flips the failure: if the reference censors and the suspect doesn't, you get a false positive spike on China-sensitive prompts. If both censor, the difference cancels and you get a false negative, the shared behavior just gets subtracted away.
My first instinct for a fix was to use an ensemble o... (read more)
Cool work but something concerns me. The method only scores , so anything the two might share is invisible. The "clean reference" assumption is doing a lot of work here and only gets a weak check.
The Qwen censorship case is a nice lens because the reference's status flips the failure: if the reference censors and the suspect doesn't, you get a false positive spike on China-sensitive prompts. If both censor, the difference cancels and you get a false negative, the shared behavior just gets subtracted away.
My first instinct for a fix was to use an ensemble o... (read more)