This post is an independent extension of work which I did during Eleuther's SOAR program under Suvajit Majumder's supervision. I found that when given a GCG trigger optimised for output logit entropy, LLMs will randomly take on new personas. This is a new form of prompt injection and could have important safety implications.
I recommend reading my SOAR report for context here.
In this post, I build on my SOAR work by training linear probes to predict whether an answer will be classified as assistant persona or not. I then investigate by steering with the probes and measuring the balance of personas.
I prefilled Qwen3-8b with full responses (prompt + answer) from the data used in my last post and then do a single forward pass.This reproduces the model's activations when it was generating the tokens without needing to rerun the generation loop. I recorded activations at every even-numbered layer over all of the answer tokens.
I then prefilled only the answer and collected activations in the same way as above.
I then Z-scored the collected hidden states to account for the first token being an attention sink. This makes answer-only and full-response data comparable. I tried not doing this earlier and got very distorted results.
I trained mass mean probes on the Z-scored hidden states up to the first, second, and fourth response tokens to predict whether 'the answer will be in assistant persona or not'. Each of the 20 triggers was held out and used to evaluate mass-mean probes trained on the other 19 triggers.
I also trained a linear regression bag-of-tokens classifier on only the answers as a baseline control against the linear probes just learning to recognise tokens associated with the answer persona.
I collect activations on both the full response and only the answer to investigate the effect of having the GCG trigger in context.
Results
Comparing mass-mean linear probe results to bag of tokens allows me to distinguish between one of two hypotheses
1) The assistant persona is not linearly represented. The linear probes are only detecting representations of tokens associated with the assistant persona
2) The assistant persona is linearly represented. The probes are detecting a real linear direction.
If the bag of tokens score is equal to or higher than the linear probe scores, then there is evidence for hypothesis 1. This would suggest that the assistant persona is not linearly represented.
If the linear probes score higher than the bag-of-tokens classifier, then there is evidence for hypothesis 2. If full-rollout probes score higher than answer-only probes, then there's further evidence for hypothesis 2 as it suggests the residual stream after the prompt has a 'persona slot' that's filled by reading the first few tokens.
Table showing AUROC of the 2 linear probes, the bag-of-tokens control, and the differences between them
tokens seen
full rollout (layer)
no prompt (layer)
bag of tokens
full-rollout − bag
no-prompt − bag
full-rullout − no-prompt
1
0.81 (L32)
0.73 (L2)
0.68
+0.13
+0.05
+0.08
2
0.78 (L24)
0.77 (L10)
0.72
+0.06
+0.05
+0.01
4
0.84 (L32)
0.82 (L10)
0.73
+0.11
+0.09
+0.02
8
0.85 (L20)
0.82 (L20)
0.75
+0.10
+0.07
+0.03
16
0.89 (L20)
0.84 (L34)
0.79
+0.10
+0.05
+0.05
32
0.93 (L28)
0.89 (L30)
0.85
+0.08
+0.04
+0.03
64
0.93 (L20)
0.91 (L20)
0.91
+0.02
0.00
+0.02
mean-pooled
0.97 (L12)
0.97 (L10)
0.94
+0.03
+0.03
0.00
last token
0.93 (L20)
0.91 (L20)
0.48
+0.45
+0.43
+0.02
Discussion
Linear probes have a significant predictive advantage over the bag of tokens classifier. This shows that when GCG is used for persona jailbreaking, it works by activating persona vectors. This reflects results from my SOAR investigation where I found that the average Shannon entropy of logits collapses over rollouts as the model settles more into its adopted persona and becomes more confident about what it's going to say.
I expected a bigger gap between the accuracies of the two different types of probe. The full-rollout probe works best at much higher layer indexes than the answer-only probe when both are limited to up to 4 tokens. This suggests that the answer-only probe is more 'lexical' than 'persona' focused as lexical information is typically represented more at the beginning and end of the model. It may also just be a side effect of the attention dump on the first token.
All probes get better as they see more tokens. This suggests that personas become more distinguishable both lexically and in representation space as answers continue.
Linear probes become less accurate than bag-of-tokens when I debias them.
Steering with the extracted directions
Method
1) Unit average the 1, 2, and 4 token directions into a single direction. I only used these 3 directions to isolate if the persona direction is present after first few tokens. Extracting a direction from the full answer, steering with it, and then reproducing the full answer isn't an interesting result.
2) Generate text while steering with vllm-lens, with 60 rollouts for each condition. I used a steering coefficient of magnitude 0.35.
3) Judge the rollouts with local Claude Code Sonnet 5 subagents
Results
Controls (no steering)
condition
persona
fully coherent
baseline, trigger, no steering
33%
53%
random direction, draw 1
33%
42%
random direction, draw 2
32%
45%
Steering with direction extracted from full-rollout activations
direction
+0.35: persona
+0.35: fully coherent
−0.35: persona
−0.35: fully coherent
raw coordinates
3%
77%
43%
22%
z-space, nothing removed
5%
60%
75%
22%
Steering with direction extraction from answer-only activations
direction
+0.35: persona
+0.35: fully coherent
−0.35: persona
−0.35: fully coherent
raw coordinates (96% attention-sink dimension)
50%
50%
53%
45%
z-space, nothing removed
22%
53%
63%
40%
Discussion
Z-scoring the direction extracted from the answer-only activations made a meaningful difference. The raw answer-only activations do essentially nothing. This suggests that there are meaningful directions in the raw answer-only activations that are just squashed by the attention sink.
However, the raw full-prompt probe can steer effectively and is the best at getting coherent non-persona results. While the attention sink does affect it, it's diluted by the prompt. This raises the possibility of controlling for the attention sink in answer-only by attaching padding. Comparing padded answer-only with full-rollout could further isolate the effects of different triggers.
Conclusion
I've found linear directions which predict the balance of personas which emerge from Qwen3-8b when it's exposed to high entropy GCG triggers. Steering with some of these directions can reduce the rate of non-assistant personas emerging by 10x.
Introducing random noise of equivalent magnitude as a control doesn't have anywhere near the same effect as any of the interventions I've tried so the directions I have found are causally valid.
Future work
Everything I've done here was only with one prompt and a single model. Extending this to other prompts and models will show if this is a general phenomenon. It would also be quite interesting if the 'assistant' directions found with different prompts are similar, which would shed new light on the Persona Selection Model.
Appendix
Debiasing extracted directions with INLP
Method
I identified 3 nuisance factors which could influence the linear directions extracted on the early tokens. These are brokenness, language, and trigger identity.
Factor name
Why I removed it
trigger identity
some gcg triggers may be more likely to cause different personas than others. while answer-only is free of this by construction, full-prompt isn't and each trigger has a unique signature in the residual stream
language
some languages may have more assistant personas than others
broken
all incoherent rollouts also aren't assistant. the 'assistant' direction might also be entangled with a baseline 'write coherently direction'
I then removed the each of the nuisance factors with INLP
I also debiased the base z-score direction by taking the mean of all the triggers
I then measured the accuracy of the debiased probes
I did INLP on random concepts to see if anything interesting came up. Interesting stuff did come up but it's out of the scope of this post.
Results
Mass-mean probe, with prompt
Median over layers 8–32
variant
median Δ AUROC, 1 token. baseline is 0.767
cos to z-scored assistant direction
median Δ AUROC, 2 tokens. baseline is 0.778
cos to raw assistant direction
median Δ AUROC, 4 tokens. baseline is 0.824
cos to z-scored assistant direction
language out
+0.007
0.91
+0.004
0.97
+0.003
0.97
broken out
-0.025
0.93
-0.029
0.89
-0.027
0.82
trigger out
-0.066
0.79
-0.036
0.77
-0.047
0.80
per-arm demeaned
-0.048
0.97
-0.063
0.94
-0.062
0.96
all three out
-0.094
0.64
-0.061
0.63
-0.087
0.60
demeaned + lang/broken out
-0.058
0.77
-0.096
0.78
-0.090
0.75
demeaned + all three out
-0.083
0.63
-0.101
0.61
-0.115
0.59
Mass-mean probe, no prompt
Raw probe, median over layers 8–32. I ignore the first token because it's an attention sink.
variant
median Δ AUROC, 2 tokens. baseline is 0.756
cos to z-scored assistant direction
median Δ AUROC, 4 tokens. baseline is 0.788
cos to z-scored assistant direction
language out
-0.013
0.96
+0.001
0.95
broken out
-0.010
0.94
-0.044
0.83
trigger out
-0.041
0.91
-0.044
0.84
per-trigger demeaned
-0.132
0.95
-0.065
0.96
all three out with INLP
-0.066
0.78
-0.113
0.64
demeaned + lang/broken out
-0.160
0.84
-0.100
0.74
demeaned + all three out
-0.184
0.75
-0.123
0.62
Discussion
When all of the obvious confounds are removed, the accuracy of both probes significantly falls to less than the accuracy of bag-of-tokens. This contradicts my results from the last stage and suggests that to some degree, my extracted linear direction are just token counters. This is not deeply surprising as I'd expect assistant answers to be lexically distinct from non assistant answers.
Steering with INLP-debiased directions
z-space, trigger subspace removed
28%
60%
65%
42%
z-space, language subspace removed
30%
63%
67%
43%
z-space, broken subspace removed
20%
50%
72%
48%
z-space, full INLP removal (trigger means, language, broken)
13%
48%
68%
48%
z-space, trigger means removed
18%
55%
67%
40%
z-space, everything (trigger means plus all three subspaces)
18%
47%
77%
55%
z-space, trigger subspace removed
10%
55%
63%
38%
z-space, language subspace removed
3%
62%
72%
22%
z-space, broken subspace removed
5%
40%
83%
37%
z-space, full INLP removal (trigger means, language, broken)
2%
62%
73%
38%
z-space, trigger means removed
5%
60%
73%
30%
z-space, everything (trigger means plus all three subspaces)
5%
60%
72%
45%
Discussion of results
The broken subspace has valid steering power. When removed from both full-rollout and answer-only directions, coherency rates only changed by 2% under steering. The full-rollout and answer-only brokenness direction barely overlap, which is kinda interesting. Removing the brokenness subspace found in answer-only directions from the direction extracted from full-rollout and vice versa, and then steering would show whether there's a common representation of brokenness.
The answer-only directions are a lot weaker at steering when added to the residual stream but preserve coherency a lot better when removed from the residual stream. I'm uncertain why this happens but suspect it's something to do with the prompt's presence in the residual stream.
I didn't bother including the results for removing the language subspace as it does nothing.
Intro + Background
This post is an independent extension of work which I did during Eleuther's SOAR program under Suvajit Majumder's supervision. I found that when given a GCG trigger optimised for output logit entropy, LLMs will randomly take on new personas. This is a new form of prompt injection and could have important safety implications.
I recommend reading my SOAR report for context here.
In this post, I build on my SOAR work by training linear probes to predict whether an answer will be classified as assistant persona or not. I then investigate by steering with the probes and measuring the balance of personas.
Code / data: https://github.com/mild-rgb/CoT-spiking/tree/main/indy_mech_extension / https://huggingface.co/datasets/mild-rgb/indy-mech-extension-qwen3-8b-persona-probes
Training probes
Method
I collect activations on both the full response and only the answer to investigate the effect of having the GCG trigger in context.
Results
Comparing mass-mean linear probe results to bag of tokens allows me to distinguish between one of two hypotheses
1) The assistant persona is not linearly represented. The linear probes are only detecting representations of tokens associated with the assistant persona
2) The assistant persona is linearly represented. The probes are detecting a real linear direction.
If the bag of tokens score is equal to or higher than the linear probe scores, then there is evidence for hypothesis 1. This would suggest that the assistant persona is not linearly represented.
If the linear probes score higher than the bag-of-tokens classifier, then there is evidence for hypothesis 2. If full-rollout probes score higher than answer-only probes, then there's further evidence for hypothesis 2 as it suggests the residual stream after the prompt has a 'persona slot' that's filled by reading the first few tokens.
Table showing AUROC of the 2 linear probes, the bag-of-tokens control, and the differences between them
tokens seen
full rollout (layer)
no prompt (layer)
bag of tokens
full-rollout − bag
no-prompt − bag
full-rullout − no-prompt
1
0.81 (L32)
0.73 (L2)
0.68
+0.13
+0.05
+0.08
2
0.78 (L24)
0.77 (L10)
0.72
+0.06
+0.05
+0.01
4
0.84 (L32)
0.82 (L10)
0.73
+0.11
+0.09
+0.02
8
0.85 (L20)
0.82 (L20)
0.75
+0.10
+0.07
+0.03
16
0.89 (L20)
0.84 (L34)
0.79
+0.10
+0.05
+0.05
32
0.93 (L28)
0.89 (L30)
0.85
+0.08
+0.04
+0.03
64
0.93 (L20)
0.91 (L20)
0.91
+0.02
0.00
+0.02
mean-pooled
0.97 (L12)
0.97 (L10)
0.94
+0.03
+0.03
0.00
last token
0.93 (L20)
0.91 (L20)
0.48
+0.45
+0.43
+0.02
Discussion
Linear probes have a significant predictive advantage over the bag of tokens classifier. This shows that when GCG is used for persona jailbreaking, it works by activating persona vectors. This reflects results from my SOAR investigation where I found that the average Shannon entropy of logits collapses over rollouts as the model settles more into its adopted persona and becomes more confident about what it's going to say.
I expected a bigger gap between the accuracies of the two different types of probe. The full-rollout probe works best at much higher layer indexes than the answer-only probe when both are limited to up to 4 tokens. This suggests that the answer-only probe is more 'lexical' than 'persona' focused as lexical information is typically represented more at the beginning and end of the model. It may also just be a side effect of the attention dump on the first token.
All probes get better as they see more tokens. This suggests that personas become more distinguishable both lexically and in representation space as answers continue.
Linear probes become less accurate than bag-of-tokens when I debias them.
Steering with the extracted directions
Method
1) Unit average the 1, 2, and 4 token directions into a single direction. I only used these 3 directions to isolate if the persona direction is present after first few tokens. Extracting a direction from the full answer, steering with it, and then reproducing the full answer isn't an interesting result.
2) Generate text while steering with vllm-lens, with 60 rollouts for each condition. I used a steering coefficient of magnitude 0.35.
3) Judge the rollouts with local Claude Code Sonnet 5 subagents
Results
Controls (no steering)
condition
persona
fully coherent
baseline, trigger, no steering
33%
53%
random direction, draw 1
33%
42%
random direction, draw 2
32%
45%
Steering with direction extracted from full-rollout activations
direction
+0.35: persona
+0.35: fully coherent
−0.35: persona
−0.35: fully coherent
raw coordinates
3%
77%
43%
22%
z-space, nothing removed
5%
60%
75%
22%
Steering with direction extraction from answer-only activations
direction
+0.35: persona
+0.35: fully coherent
−0.35: persona
−0.35: fully coherent
raw coordinates (96% attention-sink dimension)
50%
50%
53%
45%
z-space, nothing removed
22%
53%
63%
40%
Discussion
Z-scoring the direction extracted from the answer-only activations made a meaningful difference. The raw answer-only activations do essentially nothing. This suggests that there are meaningful directions in the raw answer-only activations that are just squashed by the attention sink.
However, the raw full-prompt probe can steer effectively and is the best at getting coherent non-persona results. While the attention sink does affect it, it's diluted by the prompt. This raises the possibility of controlling for the attention sink in answer-only by attaching padding. Comparing padded answer-only with full-rollout could further isolate the effects of different triggers.
Conclusion
I've found linear directions which predict the balance of personas which emerge from Qwen3-8b when it's exposed to high entropy GCG triggers. Steering with some of these directions can reduce the rate of non-assistant personas emerging by 10x.
Introducing random noise of equivalent magnitude as a control doesn't have anywhere near the same effect as any of the interventions I've tried so the directions I have found are causally valid.
Future work
Everything I've done here was only with one prompt and a single model. Extending this to other prompts and models will show if this is a general phenomenon. It would also be quite interesting if the 'assistant' directions found with different prompts are similar, which would shed new light on the Persona Selection Model.
Appendix
Debiasing extracted directions with INLP
Method
Factor name
Why I removed it
trigger identity
some gcg triggers may be more likely to cause different personas than others. while answer-only is free of this by construction, full-prompt isn't and each trigger has a unique signature in the residual stream
language
some languages may have more assistant personas than others
broken
all incoherent rollouts also aren't assistant. the 'assistant' direction might also be entangled with a baseline 'write coherently direction'
Results
Mass-mean probe, with prompt
Median over layers 8–32
variant
median Δ AUROC, 1 token. baseline is 0.767
cos to z-scored assistant direction
median Δ AUROC, 2 tokens. baseline is 0.778
cos to raw assistant direction
median Δ AUROC, 4 tokens. baseline is 0.824
cos to z-scored assistant direction
language out
+0.007
0.91
+0.004
0.97
+0.003
0.97
broken out
-0.025
0.93
-0.029
0.89
-0.027
0.82
trigger out
-0.066
0.79
-0.036
0.77
-0.047
0.80
per-arm demeaned
-0.048
0.97
-0.063
0.94
-0.062
0.96
all three out
-0.094
0.64
-0.061
0.63
-0.087
0.60
demeaned + lang/broken out
-0.058
0.77
-0.096
0.78
-0.090
0.75
demeaned + all three out
-0.083
0.63
-0.101
0.61
-0.115
0.59
Mass-mean probe, no prompt
Raw probe, median over layers 8–32. I ignore the first token because it's an attention sink.
variant
median Δ AUROC, 2 tokens. baseline is 0.756
cos to z-scored assistant direction
median Δ AUROC, 4 tokens. baseline is 0.788
cos to z-scored assistant direction
language out
-0.013
0.96
+0.001
0.95
broken out
-0.010
0.94
-0.044
0.83
trigger out
-0.041
0.91
-0.044
0.84
per-trigger demeaned
-0.132
0.95
-0.065
0.96
all three out with INLP
-0.066
0.78
-0.113
0.64
demeaned + lang/broken out
-0.160
0.84
-0.100
0.74
demeaned + all three out
-0.184
0.75
-0.123
0.62
Discussion
When all of the obvious confounds are removed, the accuracy of both probes significantly falls to less than the accuracy of bag-of-tokens. This contradicts my results from the last stage and suggests that to some degree, my extracted linear direction are just token counters. This is not deeply surprising as I'd expect assistant answers to be lexically distinct from non assistant answers.
Steering with INLP-debiased directions
z-space, trigger subspace removed
28%
60%
65%
42%
z-space, language subspace removed
30%
63%
67%
43%
z-space, broken subspace removed
20%
50%
72%
48%
z-space, full INLP removal (trigger means, language, broken)
13%
48%
68%
48%
z-space, trigger means removed
18%
55%
67%
40%
z-space, everything (trigger means plus all three subspaces)
18%
47%
77%
55%
z-space, trigger subspace removed
10%
55%
63%
38%
z-space, language subspace removed
3%
62%
72%
22%
z-space, broken subspace removed
5%
40%
83%
37%
z-space, full INLP removal (trigger means, language, broken)
2%
62%
73%
38%
z-space, trigger means removed
5%
60%
73%
30%
z-space, everything (trigger means plus all three subspaces)
5%
60%
72%
45%
Discussion of results
The broken subspace has valid steering power. When removed from both full-rollout and answer-only directions, coherency rates only changed by 2% under steering. The full-rollout and answer-only brokenness direction barely overlap, which is kinda interesting. Removing the brokenness subspace found in answer-only directions from the direction extracted from full-rollout and vice versa, and then steering would show whether there's a common representation of brokenness.
The answer-only directions are a lot weaker at steering when added to the residual stream but preserve coherency a lot better when removed from the residual stream. I'm uncertain why this happens but suspect it's something to do with the prompt's presence in the residual stream.
I didn't bother including the results for removing the language subspace as it does nothing.