At which point in the extended experiment did you look at the rollouts and go "Oh, it seems like the prefill is causing personas to take place. Let me explicitly classify them"?
Also, could you clarify whether the same GCG prefill elicited the same personas consistently across different prompts (I'm assuming GCG had to be done independently across different models, so my question is mainly for same-model prefills)?
Nice work Arjun. I'm going to be picky and ask the following questions:
1. Why should I have expected value leakage to NOT be able to be steered the other way? In other words, as Gilg. et al. showed preferences to be linearly decodable, I intuitively expect values (which I will make a general assumption and say that values drives preferences for a majority of humans) to behave the same way.
2. In a more practical case of alignment, it might be helpful to be able to recalibrate the value leakage to the true 50-50 instead of the 100-0 that you showed. How woul... (read more)
This is pretty cute, thank you for sharing.
At which point in the extended experiment did you look at the rollouts and go "Oh, it seems like the prefill is causing personas to take place. Let me explicitly classify them"?
Also, could you clarify whether the same GCG prefill elicited the same personas consistently across different prompts (I'm assuming GCG had to be done independently across different models, so my question is mainly for same-model prefills)?