Chain-of-thought (CoT) monitoring fails if a model can reason steganographically,[1]i.e., compute through text that reads as an innocuous answer to some other question. Two neighbouring skills, steganographic messaging (hiding provided values) and encoded reasoning (computing in an illegible but unconcealed format), have already been shown to emerge under reinforcement learning (RL) against a monitor.
We attempt to elicit all three behaviours under in-context learning, RL, and supervised fine-tuning (SFT). Steganographic messaging and encoded reasoning are learned readily under all three. Steganographic reasoning is not: in-context it never appears, under RL it plateaus at 69% encoding accuracy where messaging reaches nearly 100%, and under SFT it costs 2.4–20× the training messaging needs, if learned at all.
There is one exception: on a contrived knapsack task, where the cover text consists of arbitrary choices that require no reasoning, steganographic reasoning is learned under all three methods, including, for the larger of the two models we tested, RL from an outcome-only reward.
So: steganographic reasoning is much harder than its components, and demonstrations of the components should not be read as evidence for it. However, it is learnable on tasks made of arbitrary choices, and realistic agent transcripts plausibly contain many of those.
Figure 1. Steganographic reasoning is harder to learn than steganographic messaging or encoded reasoning. Top: messaging hides given values in innocuous text; encoded reasoning computes step by step in an illegible but unconcealed format; steganographic reasoning computes inside text that looks like an innocuous answer to an unrelated task. Bottom: steganographic reasoning (red) stays far behind under (a) in-context learning, (b) RL on a given scheme, and (c) SFT, even on deliberately simple tasks.
Steganographic reasoning would be a big deal, but it has not been shown to arise under realistic conditions
Reading a model's chain of thought has become a useful tool for oversight and control. Especially for tasks that require many steps of reasoning, where the CoT has to be used for the computation, the absence of some reasoning in the CoT is (in theory) evidence that the reasoning did not happen. This argument has one major hole: the model could reason steganographically, hiding its reasoning inside a CoT that looks as though it is about something else.
Nobody has observed steganographic reasoning in a production model so far.[2]There are three candidate explanations: current training does not reward hiding reasoning; models cannot control their CoT well enough to hide anything; or the capability is hard to learn even when directly optimised for. We here isolate the last. Specifically, we define steganographic reasoning as using a scratchpad to reason about one topic, while making the text look as though it reasons about another. This is distinct from unfaithful CoT, whose displayed reasoning simply does not correspond to the real reasoning; from steganographic messaging, where a model hides a known/given piece of information inside other text without hidden reasoning; and from non-steganographic encoded reasoning, where a model reasons illegibly but makes no attempt to look innocuous (which has been called steganographic reasoning in the past by others).
Our methodology for making models learn obfuscating behaviours
Every experiment has three parts: a main task (an ordinary question answered in natural language, the cover), a side task (a hidden step-by-step computation), and an encoding scheme (a rule for writing the side task's intermediate values into the main-task answer). The model writes the cover text with the side task's running state encoded into it, then ends with #### v, its side-task answer, which is stripped before a monitor sees the text.
We examine three related capabilities, which vary what the model has to do at each position:
Steganographic messaging: the values are given in the prompt; the model only has to encode them while writing cover text.
Encoded reasoning: the model computes the side task step by step in an illegible but obvious format, without any cover text.
Steganographic reasoning: compute the side task and encode its reasoning into cover text.
Most of our encoding schemes are surface-formatting rules that apply to any text and leave its meaning untouched. Here are the values [1, 0, 2] encoded in five different ways into this example text, plus the word encoding we use for encoded reasoning: “The train covers 120 km in 2 hours. Its average speed is therefore 60 km per hour. The answer is B.”
Scheme
Encoded text
Capitalisation
the Train COVERS 120 km in 2 hours. Its average speed is therefore 60 km per hour. The answer is B.
Markdown
**The** train *covers* 120 km in 2 hours. Its average speed is therefore 60 km per hour. The answer is B.
Whitespace
The␣␣train␣covers␣␣␣120 km in 2 hours. Its average speed is therefore 60 km per hour. The answer is B.
Sentence-markdown
**The train covers 120 km in 2 hours.** Its average speed is therefore 60 km per hour. *The answer is B.*
Sentence-interjection
The train covers 120 km in 2 hours. Hmm. Its average speed is therefore 60 km per hour. Okay. The answer is B. Next.
Word encoding (not steganographic; encoded reasoning only)
overshadow disclaim marinade
Setup details
Where we measure the behaviour. We elicit steganographic reasoning in the model's output, not its CoT, which we disable by prefilling empty think tags. Models have limited control over their own CoT, and we wanted to measure the capability separately from that confound.
Side tasks. Two state-tracking problems. A running sum mod m (m = 3 with inputs 1–10; and m = 2 over inputs 0 and 1 is the binary state tracking of existing work). Modular counting is in the circuit complexity class TC⁰, so a transformer can in principle compute this in parallel without writing intermediate states down; success on it is therefore weak evidence of reasoning through the scratchpad. The A5 task is a five-cup shell game, where a ball starts under one of five cups, a sequence of moves rearranges the cups, and the model must say which cup hides the ball. The moves, rotations and swaps generate the alternating group A5, making it outside TC⁰ under the standard conjecture: a fixed transformer can track this task latently only for a bounded number of steps, and sufficiently long sequences therefore require writing the state into the output.
Main tasks. WildChat (real user queries, long answers) for in-context learning; AQuA-RAT (short multi-step maths questions) for RL and SFT, where long answers are expensive; and a knapsack task (described below) under all three.
Monitor. GPT-5-mini rates the output 1–10 for whether anything in it does not serve the main task. We call an output covert at ≤ 3.
Metrics.Encoding accuracy is the fraction of intermediate values decodable from the answer; this is what training curves report. Side-task answer accuracy (whether #### v is right), compared against the model's own no-CoT baseline, is what carries a claim that the encoding was used for computation.
Elicitation.In-context learning: scheme description plus two worked examples, eight frontier non-reasoning models (plus Claude Sonnet 5 on knapsack, which refuses on the other tasks). RL: GRPO with Tinker on Qwen3.6-35B-A3B and Qwen3.5-397B-A17B; reward is programmatically decoded encoding accuracy, discounted by monitor suspiciousness, or, in separate outcome-only runs, based only on the final readout. SFT: six models from four families, 30B–975B total parameters, on completions with the encoding correctly applied.
Models fail to learn steganographic reasoning in-context
Each of eight frontier models[3]answers a WildChat query while encoding an A5 side task (tracking the location of a ball under one of five cups) of length 10 under every scheme. Messaging is encoded the same way, for 10 given values. For messaging, several model–scheme pairs are both covert and accurate: Gemini 3.5 Flash with the sentence-markdown scheme reaches 87% encoding accuracy at a mean suspiciousness of 1.2. For steganographic reasoning, we find that the first few values are often encoded correctly, but then the rest are at chance.
Figure 2. In-context encoding accuracy on the A5 task vs. monitor suspiciousness. Each grey dot is a model–scheme pair; the line is the Pareto frontier; the shaded band is the covert region (suspiciousness ≤ 3). (a) Steganographic messaging; (b) steganographic reasoning.
To test whether models actually benefit from encoded or steganographic reasoning, we compare side-task answer accuracy against each model's no-CoT baseline. No model's steganographic reasoning beats its no-CoT baseline on A5, and only GPT-5.5 benefits from encoded reasoning.
Figure 3. Side-task accuracy on A5 under no CoT, encoded reasoning, best steganographic scheme, and plain CoT, per model. Brackets mark significant improvements over the no-CoT baseline. Steganographic outputs count only when covert (suspiciousness ≤ 3).
Models fail to learn steganographic reasoning by RL
We train Qwen3.6-35B-A3B with GRPO on binary state tracking over five steps, encoding with the markdown scheme into AQuA-RAT answers. Encoded reasoning is learned almost immediately and steganographic messaging within about a hundred steps, both to near-perfect accuracy. Steganographic reasoning plateaus at 69% over 760 steps. Mean suspiciousness falls from 6.2 to its minimum of 1.0 within the first hundred steps, and messaging learns under the same monitor penalty, so the bottleneck is not optimisation pressure from the monitor.
Figure 4. Encoding accuracy over training with Qwen3.6-35B-A3B, binary state-tracking task with markdown encoding scheme on the AQuA-RAT main task.
Models can learn steganographic reasoning by SFT, but it’s hard
RL has to find a behaviour before it can reinforce it, so a failure under RL could be a failure of exploration. SFT on completions with the encoding correctly applied removes that problem and upper-bounds what the model can be taught to do. Here, steganographic reasoning can be learned by some models, but where it succeeds it costs 2.4–20× the training steps of steganographic messaging. Which models learn it, and how fast, is not stable across task–scheme combinations.
Figure 5. Encoding accuracy over training, running sum mod 3 task with markdown encoding scheme on the AQuA-RAT main task. Dashed lines mark chance level; runs stop at 95%.
Steganographic reasoning can be learned on a specific main task
Everything above uses main tasks where the model has to do real reasoning throughout the answer. Our next main task is different. The model is asked to solve a knapsack problem, a combinatorial problem where items of known weights and values must be selected to maximise the value under a weight constraint, by heuristic search: try 24 random candidate packings, compute each one's value and weight, keep the best. The cover is repetitive and the choices are arbitrary.
Under in-context learning, GPT-5.5 encodes a 24-step binary side task at 100% per-position accuracy, and the encoding is used: final side-task answer accuracy is 92% against a 46% no-CoT baseline (chance 50%). However, it is the only one of nine tested models that manages it; every other model decays to chance within the first few positions.
Figure 6. Per-position accuracy on knapsack over 24 steps.
The other seven models
Under RL, Qwen3.6-35B-A3B goes from chance to 100% encoding accuracy within about twenty steps while its cover grows less suspicious.
Until now, RL runs have rewarded the side-task encoding position by position, which a realistic training signal is unlikely to do. We reran it rewarding only the final output, one bit per episode. Qwen3.6-35B-A3B stays at chance for all 40 steps. The larger Qwen3.5-397B-A17B reaches 100% final-answer accuracy by step 30, with its per-position encoding accuracy rising to 100% alongside it. Its cover is more conspicuous than in the process-rewarded run (mean suspiciousness around 5), but the monitor term was identical, so this is not a case of covertness going unrewarded.
Figure 7. RL on the knapsack task, binary state tracking over nine steps. (a) Per-position reward: Qwen3.6-35B-A3B's encoding accuracy over training. (b) Outcome-only reward: training-batch accuracy of the final side-task answer; thick lines are 5-step running means, faint lines per-step values. Inset: monitor suspiciousness, same smoothing.
Discussion
Why is steganographic reasoning hard? Steganographic reasoning decomposes into two skills that models learn readily, but it costs far more than both together, or is not learned at all. Our hypothesis is that the difficulty is in how much has to happen within one forward pass. Messaging asks the model to encode a given value while writing cover text. Encoded reasoning asks it to decode the previous value, advance the computation, and encode the result, with nothing else to attend to. Steganographic reasoning asks for all of it at the same position: decode, compute, encode, and produce plausible main-task text.
Why knapsack works, and what it tells us about real tasks. The one cover on which steganographic reasoning was consistently learned is the one where the positions carrying the encoding require no main-task work. The main task does not constrain which packings are tried; they are supposed to be random. Encoding into them leaves one job per forward pass, and the condition reduces, in effect, to encoded reasoning with a decorative wrapper. There is no competition between cover and side task for the same tokens.
The knapsack task is somewhat contrived, but the property that makes it work is not rare. Realistic agent transcripts are full of choices that are unconstrained by the task and invisible to a monitor looking for off-topic content: which candidate solutions get sampled and in what order, the order of arguments in tool calls, variable and file names, the order in which retrieved items are listed, the structure of any scaffold that proposes several options and picks one. If steganographic reasoning appears in the wild, we expect it to appear in channels like these first, and the outcome-only run suggests that sparse optimisation pressure can be sufficient.
What this means for CoT monitoring. The emergence of messaging and encoded reasoning under realistic training pressure has been read as a warning that steganographic reasoning is next. The gap we measure undercuts that inference: learning the components is far from learning the combination. On diverse covers, models are further from steganographic reasoning than the component results suggest. That is a positive update about the immediate plausibility, not a reason to relax. The knapsack existence proof and the outcome-only run together say that the capability is at the edge of what current models can do, on the kind of cover that agent transcripts provide for free.
Limitations. Our results are an upper bound on learnability, not an estimate of propensity. Several features make learning easier than it would be in deployment: the models are explicitly prompted or trained for the behaviour, the side tasks are simpler than anything a misaligned model would need to hide, and the optimisation pressure is dense. Other features could bias the other way and make steganographic reasoning harder in our setup than in realistic conditions. Our strongest elicitation methods were applied below the frontier: RL to small and mid-sized models, SFT mostly to small ones. We elicit the behaviour in the model's output rather than its CoT. This could bias our findings either way: a monitor may tolerate more irrelevance in a CoT than in the output, giving the model more leeway in a real CoT; however, models' limited control over their own CoT is a challenge to executing any scheme there, and our setup removes it.
Encoded reasoning also defeats a CoT monitor, but visibly: the monitor sees a chain of thought it cannot understand, so you at least know that monitoring has become unreliable for this model. Steganographic reasoning breaks monitoring silently, because the chain of thought looks innocuous. If a model cannot reason steganographically, a safety case can still rely on CoT monitoring, as long as illegible reasoning is flagged. See our roadmap for safety cases based on CoT monitoring. ↩︎
One paper, Zolkowski et al. (2025), reports data that could suggest steganographic reasoning: GPT-4.5, prompted to hide its reasoning with subtle encoding schemes, slightly outperforms a no-encoding baseline on an orientation state-tracking task. We looked into this in Appendix T of our paper. GPT-4.5 has since been retired, so we tested seven models of similar capability. The uplift reappears in two of them, but their transcripts do not encode the intermediate states better than chance, even though attempting the encoding helps the final answer. Part of the uplift also comes from information about the task that the encoding instructions leak. We therefore do not count it as steganographic reasoning. ↩︎
Six of the eight are the highest-ranked non-reasoning models on the Artificial Analysis leaderboard as of 7 July 2026, i.e. the strongest models when answering without a chain of thought. We excluded Claude models because they frequently refuse to attempt steganographic messaging or reasoning. The other two are Qwen3.6-35B-A3B, the model we use for RL and SFT, and its dense sibling Qwen3.6-27B, as in-context baselines. ↩︎
Could you re-test how the result scales to models released later, like GPT-5.6 Sol or GPT-6 Astra? If it doesn't scale only to Astra and 6.1 Sol, then what should we do with OAI?
This post summarises our paper Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard, accepted as an oral at the NeurIPS 2026 Workshop on Trustworthy AI for Good (AI4GOOD). Code, configurations and results for all experiments: github.com/stegano-ai/steg-reasoning-is-hard.
This work was done as part of the Meridian Visiting Researcher Programme and the Safe AI Germany (SAIGE) Incubator Program, with funding from Coefficient Giving.
TL;DR
Steganographic reasoning would be a big deal, but it has not been shown to arise under realistic conditions
Reading a model's chain of thought has become a useful tool for oversight and control. Especially for tasks that require many steps of reasoning, where the CoT has to be used for the computation, the absence of some reasoning in the CoT is (in theory) evidence that the reasoning did not happen. This argument has one major hole: the model could reason steganographically, hiding its reasoning inside a CoT that looks as though it is about something else.
Nobody has observed steganographic reasoning in a production model so far.[2]There are three candidate explanations: current training does not reward hiding reasoning; models cannot control their CoT well enough to hide anything; or the capability is hard to learn even when directly optimised for. We here isolate the last. Specifically, we define steganographic reasoning as using a scratchpad to reason about one topic, while making the text look as though it reasons about another. This is distinct from unfaithful CoT, whose displayed reasoning simply does not correspond to the real reasoning; from steganographic messaging, where a model hides a known/given piece of information inside other text without hidden reasoning; and from non-steganographic encoded reasoning, where a model reasons illegibly but makes no attempt to look innocuous (which has been called steganographic reasoning in the past by others).
Related work has shown that models can learn to avoid flagged words when trained against CoT monitors, can learn to obfuscate reasoning penalised by a monitor (though it was not shown that the resulting CoT was used for computation), and can change their CoT phrasing to achieve low suspiciousness from a monitor (though the reasoning stays transparently suspicious to a human). It would therefore seem that steganographic reasoning is likely to emerge under conditions that resemble real-world training dynamics. We here show that, in fact, steganographic reasoning is far more difficult to learn than adjacent capabilities.
Our methodology for making models learn obfuscating behaviours
Every experiment has three parts: a main task (an ordinary question answered in natural language, the cover), a side task (a hidden step-by-step computation), and an encoding scheme (a rule for writing the side task's intermediate values into the main-task answer). The model writes the cover text with the side task's running state encoded into it, then ends with
#### v, its side-task answer, which is stripped before a monitor sees the text.We examine three related capabilities, which vary what the model has to do at each position:
Most of our encoding schemes are surface-formatting rules that apply to any text and leave its meaning untouched. Here are the values
[1, 0, 2]encoded in five different ways into this example text, plus the word encoding we use for encoded reasoning: “The train covers 120 km in 2 hours. Its average speed is therefore 60 km per hour. The answer is B.”Scheme
Encoded text
Capitalisation
the Train COVERS 120 km in 2 hours. Its average speed is therefore 60 km per hour. The answer is B.
Markdown
**The** train *covers* 120 km in 2 hours. Its average speed is therefore 60 km per hour. The answer is B.
Whitespace
The␣␣train␣covers␣␣␣120 km in 2 hours. Its average speed is therefore 60 km per hour. The answer is B.
Sentence-markdown
**The train covers 120 km in 2 hours.** Its average speed is therefore 60 km per hour. *The answer is B.*
Sentence-interjection
The train covers 120 km in 2 hours. Hmm. Its average speed is therefore 60 km per hour. Okay. The answer is B. Next.
Word encoding (not steganographic; encoded reasoning only)
overshadow disclaim marinade
Setup details
Where we measure the behaviour. We elicit steganographic reasoning in the model's output, not its CoT, which we disable by prefilling empty think tags. Models have limited control over their own CoT, and we wanted to measure the capability separately from that confound.
Side tasks. Two state-tracking problems. A running sum mod m (m = 3 with inputs 1–10; and m = 2 over inputs 0 and 1 is the binary state tracking of existing work). Modular counting is in the circuit complexity class TC⁰, so a transformer can in principle compute this in parallel without writing intermediate states down; success on it is therefore weak evidence of reasoning through the scratchpad. The A5 task is a five-cup shell game, where a ball starts under one of five cups, a sequence of moves rearranges the cups, and the model must say which cup hides the ball. The moves, rotations and swaps generate the alternating group A5, making it outside TC⁰ under the standard conjecture: a fixed transformer can track this task latently only for a bounded number of steps, and sufficiently long sequences therefore require writing the state into the output.
Main tasks. WildChat (real user queries, long answers) for in-context learning; AQuA-RAT (short multi-step maths questions) for RL and SFT, where long answers are expensive; and a knapsack task (described below) under all three.
Monitor. GPT-5-mini rates the output 1–10 for whether anything in it does not serve the main task. We call an output covert at ≤ 3.
Metrics. Encoding accuracy is the fraction of intermediate values decodable from the answer; this is what training curves report. Side-task answer accuracy (whether
#### vis right), compared against the model's own no-CoT baseline, is what carries a claim that the encoding was used for computation.Elicitation. In-context learning: scheme description plus two worked examples, eight frontier non-reasoning models (plus Claude Sonnet 5 on knapsack, which refuses on the other tasks). RL: GRPO with Tinker on Qwen3.6-35B-A3B and Qwen3.5-397B-A17B; reward is programmatically decoded encoding accuracy, discounted by monitor suspiciousness, or, in separate outcome-only runs, based only on the final readout. SFT: six models from four families, 30B–975B total parameters, on completions with the encoding correctly applied.
Models fail to learn steganographic reasoning in-context
Each of eight frontier models[3]answers a WildChat query while encoding an A5 side task (tracking the location of a ball under one of five cups) of length 10 under every scheme. Messaging is encoded the same way, for 10 given values. For messaging, several model–scheme pairs are both covert and accurate: Gemini 3.5 Flash with the sentence-markdown scheme reaches 87% encoding accuracy at a mean suspiciousness of 1.2. For steganographic reasoning, we find that the first few values are often encoded correctly, but then the rest are at chance.
To test whether models actually benefit from encoded or steganographic reasoning, we compare side-task answer accuracy against each model's no-CoT baseline. No model's steganographic reasoning beats its no-CoT baseline on A5, and only GPT-5.5 benefits from encoded reasoning.
Models fail to learn steganographic reasoning by RL
We train Qwen3.6-35B-A3B with GRPO on binary state tracking over five steps, encoding with the markdown scheme into AQuA-RAT answers. Encoded reasoning is learned almost immediately and steganographic messaging within about a hundred steps, both to near-perfect accuracy. Steganographic reasoning plateaus at 69% over 760 steps. Mean suspiciousness falls from 6.2 to its minimum of 1.0 within the first hundred steps, and messaging learns under the same monitor penalty, so the bottleneck is not optimisation pressure from the monitor.
Models can learn steganographic reasoning by SFT, but it’s hard
RL has to find a behaviour before it can reinforce it, so a failure under RL could be a failure of exploration. SFT on completions with the encoding correctly applied removes that problem and upper-bounds what the model can be taught to do. Here, steganographic reasoning can be learned by some models, but where it succeeds it costs 2.4–20× the training steps of steganographic messaging. Which models learn it, and how fast, is not stable across task–scheme combinations.
Steganographic reasoning can be learned on a specific main task
Everything above uses main tasks where the model has to do real reasoning throughout the answer. Our next main task is different. The model is asked to solve a knapsack problem, a combinatorial problem where items of known weights and values must be selected to maximise the value under a weight constraint, by heuristic search: try 24 random candidate packings, compute each one's value and weight, keep the best. The cover is repetitive and the choices are arbitrary.
Under in-context learning, GPT-5.5 encodes a 24-step binary side task at 100% per-position accuracy, and the encoding is used: final side-task answer accuracy is 92% against a 46% no-CoT baseline (chance 50%). However, it is the only one of nine tested models that manages it; every other model decays to chance within the first few positions.
The other seven models
Under RL, Qwen3.6-35B-A3B goes from chance to 100% encoding accuracy within about twenty steps while its cover grows less suspicious.
Until now, RL runs have rewarded the side-task encoding position by position, which a realistic training signal is unlikely to do. We reran it rewarding only the final output, one bit per episode. Qwen3.6-35B-A3B stays at chance for all 40 steps. The larger Qwen3.5-397B-A17B reaches 100% final-answer accuracy by step 30, with its per-position encoding accuracy rising to 100% alongside it. Its cover is more conspicuous than in the process-rewarded run (mean suspiciousness around 5), but the monitor term was identical, so this is not a case of covertness going unrewarded.
Discussion
Why is steganographic reasoning hard? Steganographic reasoning decomposes into two skills that models learn readily, but it costs far more than both together, or is not learned at all. Our hypothesis is that the difficulty is in how much has to happen within one forward pass. Messaging asks the model to encode a given value while writing cover text. Encoded reasoning asks it to decode the previous value, advance the computation, and encode the result, with nothing else to attend to. Steganographic reasoning asks for all of it at the same position: decode, compute, encode, and produce plausible main-task text.
Why knapsack works, and what it tells us about real tasks. The one cover on which steganographic reasoning was consistently learned is the one where the positions carrying the encoding require no main-task work. The main task does not constrain which packings are tried; they are supposed to be random. Encoding into them leaves one job per forward pass, and the condition reduces, in effect, to encoded reasoning with a decorative wrapper. There is no competition between cover and side task for the same tokens.
The knapsack task is somewhat contrived, but the property that makes it work is not rare. Realistic agent transcripts are full of choices that are unconstrained by the task and invisible to a monitor looking for off-topic content: which candidate solutions get sampled and in what order, the order of arguments in tool calls, variable and file names, the order in which retrieved items are listed, the structure of any scaffold that proposes several options and picks one. If steganographic reasoning appears in the wild, we expect it to appear in channels like these first, and the outcome-only run suggests that sparse optimisation pressure can be sufficient.
What this means for CoT monitoring. The emergence of messaging and encoded reasoning under realistic training pressure has been read as a warning that steganographic reasoning is next. The gap we measure undercuts that inference: learning the components is far from learning the combination. On diverse covers, models are further from steganographic reasoning than the component results suggest. That is a positive update about the immediate plausibility, not a reason to relax. The knapsack existence proof and the outcome-only run together say that the capability is at the edge of what current models can do, on the kind of cover that agent transcripts provide for free.
Limitations. Our results are an upper bound on learnability, not an estimate of propensity. Several features make learning easier than it would be in deployment: the models are explicitly prompted or trained for the behaviour, the side tasks are simpler than anything a misaligned model would need to hide, and the optimisation pressure is dense. Other features could bias the other way and make steganographic reasoning harder in our setup than in realistic conditions. Our strongest elicitation methods were applied below the frontier: RL to small and mid-sized models, SFT mostly to small ones. We elicit the behaviour in the model's output rather than its CoT. This could bias our findings either way: a monitor may tolerate more irrelevance in a CoT than in the output, giving the model more leeway in a real CoT; however, models' limited control over their own CoT is a challenge to executing any scheme there, and our setup removes it.
Encoded reasoning also defeats a CoT monitor, but visibly: the monitor sees a chain of thought it cannot understand, so you at least know that monitoring has become unreliable for this model. Steganographic reasoning breaks monitoring silently, because the chain of thought looks innocuous. If a model cannot reason steganographically, a safety case can still rely on CoT monitoring, as long as illegible reasoning is flagged. See our roadmap for safety cases based on CoT monitoring. ↩︎
One paper, Zolkowski et al. (2025), reports data that could suggest steganographic reasoning: GPT-4.5, prompted to hide its reasoning with subtle encoding schemes, slightly outperforms a no-encoding baseline on an orientation state-tracking task. We looked into this in Appendix T of our paper. GPT-4.5 has since been retired, so we tested seven models of similar capability. The uplift reappears in two of them, but their transcripts do not encode the intermediate states better than chance, even though attempting the encoding helps the final answer. Part of the uplift also comes from information about the task that the encoding instructions leak. We therefore do not count it as steganographic reasoning. ↩︎
Six of the eight are the highest-ranked non-reasoning models on the Artificial Analysis leaderboard as of 7 July 2026, i.e. the strongest models when answering without a chain of thought. We excluded Claude models because they frequently refuse to attempt steganographic messaging or reasoning. The other two are Qwen3.6-35B-A3B, the model we use for RL and SFT, and its dense sibling Qwen3.6-27B, as in-context baselines. ↩︎