Hi Cameron, is the SAE testing you're describing here the one you demoed in your interview with John Sherman using Goodfire's Llama 3.3 70B SAE tool? If so could you share the prompt you used for that? With the prompts I'm using I'm having a hard time getting Llama to say that it is conscious at all. It would be nice if we had SAE feature tweaking available for a model that was more ambivalent about its consciousness, seems it would be a bit easier to robustly test if that were the case.
I think when you select the "No" token, you're seeing the J-space during the generation of the token following no. To see what's in the J-space while the "No" token is being generated, you'd select the token before, like this. Still your result is weird and worth looking into more.