TLDR: In their J-lens paper, Anthropic suggests that the J-space is a global workspace that the model reasons within, and supports evidence for this hypothesis on Claude models in a variety of settings. I replicated the multi-hop reasoning experiment on Qwen3.6-27B and Gemma 3 27B-it and found that counterfactual answer swaps outperformed intermediate swaps in three of four experimental conditions. This does not provide evidence to support Anthropic's global workspace hypothesis in open-weight models and instead suggests that J-lens is more useful for probing intermediate variables rather than steering outputs.
A few months ago, Anthropic published Verbalizable Representations Form a Global Workspace in Language Models and I was immediately excited about the prospect of being able to read part of a model's working memory. Beyond that, the paper hypothesises that intermediate reasoning concepts cannot only be decoded using the J-lens, but that the J-space is actually the global workspace in which the model reasons. Neel Nanda reviewed Anthropic’s paper and replicated the results on Qwen3.6-27B with moderate success: the verbal-report interventions were weakly positive, the multilingual and typo evaluations replicated cleanly, but the poetry and arithmetic results did not replicate.
Another task that Anthropic and Nanda evaluated was multi-hop reasoning, where prompts like "What is the colour of the fourth planet in our solar system?" require an intermediate reasoning step (in this case, Mars). Initially Nanda's replication of this multi-hop reasoning setting seemed positive, because swapping out the intermediate with a counterfactual increased the probability of the counterfactual answer, e.g. if Mars→red, then replacing Mars with Neptune should output blue. This would be consistent with the intermediate actually being used in reasoning, and therefore with the J-space being the workspace in which intermediate variables are stored. However, Nanda also found that on Qwen, injecting the counterfactual answer strictly outperformed injecting the counterfactual intermediate. This makes the result less compelling, because the intervention may be steering the model towards the final answer rather than the workspace being required for intermediate reasoning. Nanda suggests that the mismatch between his and Anthropic's results could be explained by the fact that the multi-hop facts used in his dataset were too easy, such that intermediates and answers are too linearly related.
I followed up on this hypothesis. I ran the multi-hop causal experiment on Qwen3.6-27B and Gemma 3 27B-it using Camila Blank’s workspace-bench dataset and the dataset in Anthropic’s repository. This leads to four experimental conditions: Qwen/Workspace, Qwen/Anthropic, Gemma/Workspace and Gemma/Anthropic. Furthermore, I ran these four conditions with the matched R-lens, which is supposed to produce cleaner early-layer readouts compared to J-lens.
My main findings were:
Swapping intermediates increased the probability of the counterfactual answer across all four conditions.
These probability shifts rarely led to the counterfactual being in the model's top-1 output. For J-lens intermediate swaps, the top-1 flip rates were 6.3% on Qwen/Workspace, 9.1% on Qwen/Anthropic, 7.1% on Gemma/Workspace, and 11.1% on Gemma/Anthropic. All four rates are far below Anthropic’s reported 54-70% on Claude.
As in Nanda's review, my Qwen replication found that the answer swaps dominated the intermediate swaps on both datasets. This was also the case in the Gemma/Workspace experiment.
The Gemma/Anthropic run was the only condition in which the intermediate swaps had a stronger impact on output probability than the answer swaps, so in which the Anthropic results were reproduced. However, even here, the probabilistic shift only led to a flip of the actual answer in 3/27 cases.
R-lens matched J-lens on three conditions and exceeded it on the Gemma/Workspace condition.
The two multi-hop reasoning benchmarks that I used were:
Workspace benchmark: 100 multi-hop prompts from Camila Blank’s workspace-bench repository, which might be the open-weight benchmark associated with the Qwen replication discussed in Nanda’s review.
Annotated Anthropic bank: 93 prompts from Anthropic’s multi-hop data benchmark. I manually annotated the prompts with an intermediate swap target and a swap answer.
The combinations of models and benchmarks resulted in the four experimental conditions Qwen/workspace, Qwen/Anthropic, Gemma/workspace and Gemma/Anthropic. Furthermore, I adjusted the datasets to the models' capabilities by checking that the models can answer the initial prompt correctly and by removing prompts with counterfactual options that were already among the model's top-10 answers for the initial prompt. After these filters, 32 prompts remained for Qwen/workspace, 44 for Qwen/Anthropic, 28 for Gemma/workspace, and 27 for Gemma/Anthropic. There was initially an overlap of 40 prompts between the two datasets, but after filtering only four overlapping items remained.
I used 24-59 as the workspace band for Qwen and 24-57 for Gemma. I ran the causal swap experiment in two arms:
Intermediate arm: swap the original intermediate for the counterfactual intermediate.
Answer: swap the original answer for the counterfactual answer.
I scored these arms against the same counterfactual answer token and measured the change in its probability. Additionally, I measured whether the counterfactual answer became the model's top-1 output. These experiments were run using both J-lens and R-lens. The full implementation can be found in the project repository.
Causal swaps shift probabilities, but rarely flip top-1 outputs
Figure 1: The probability shift of the counterfactual answer by model-dataset combination. Bars show the probability shifts when the intermediate counterfactual is swapped versus when the counterfactual answer is swapped, with 95% confidence intervals.
In all experimental conditions, swapping in the counterfactual intermediate increases the probability of the counterfactual answer, as shown in Figure 1. However, when swapping the counterfactual intermediate in, the actual answer only rarely flipped: 2/32 interventions flipped the answer in the Qwen/workspace condition, 4/44 flips occurred in Qwen/Anthropic, 2/28 answers flipped in Gemma/workspace, and 3/27 answers flipped in Gemma/Anthropic (shown in Figure 2). These success rates are significantly lower than Anthropic’s 54-70% reported for Claude. This means that the intervention often pushes in the intended direction, but this is usually not enough to control the model’s answer.
Figure 2: The percentage of top-1 flips across the four experimental conditions.
Answer swaps usually beat intermediate swaps
Anthropic hypothesises that the model computes an intermediate reasoning step in the J-space before coming to the final answer. This should mean that the intermediate swap occurs at an earlier layer than the final answer decision, so the swap of the counterfactual intermediate target should also take hold earlier than the swap of the counterfactual answer. As Figure 1 and Figure 2 show, I was not able to replicate these results on Qwen. The answer effects exceeded the intermediate effects by +0.158 on workspace-bench and +0.044 on the Anthropic set, aggregated over the workspace layers. Gemma/workspace also favoured the answer by +0.098. This failed replication matches Nanda's findings in his J-lens review.
However, the one exception is the Gemma/Anthropic run, in which the effect of the counterfactual intermediate swap outperformed the effect of the counterfactual answer swap, both in terms of probability shift and top-1 flip rate. The total flip rate after the intermediate swap was still relatively small though, with only 3 out of 27 intermediate swaps causing a top-1 flip.
R-lens matches or exceeds J-lens
Figure 3: Comparison of the intermediate swap probabilities across J-lens and R-lens.
I also compared the J-lens results with R-lens, a closely related variant that modifies the J-lens backward pass using LRP-style rules with the aim of producing cleaner early-layer readouts. On Qwen, the results were almost indistinguishable, as shown in Figure 3. On Gemma, R-lens achieved a larger probability shift on both datasets. The strongest effect by far is achieved in the Gemma/workspace condition, where R-lens produced 12/28 top-1 flips, while J-lens produced only 2/28 flips. These results suggest that R-lens is a useful extension of the J-lens framework, which is able to preserve or even strengthen causal swap effects.
Discussion
Let's recap what Anthropic and Neel Nanda's review concluded, before summarising how my replications fit in with those. For the multi-hop reasoning experiment, Anthropic found that the J-space acts as a global workspace, in which swapped intermediates affect the output often. Nanda was able to weakly replicate this on Qwen, but found that swapping the answer outperformed the swapping of the intermediate, which undermines the strength of the J-space as a workspace.
In my experiments on Qwen3.6-27B and Gemma 3 27B-it, I found that intermediate counterfactual swaps always increased the probability of the counterfactual answer, but rarely led to a top-1 flip. Similarly to Nanda, I also found that in three of four experimental conditions, the answer swap outperformed the intermediate swap. The exception to this was the Gemma/Anthropic condition, which was able to replicate Anthropic’s intermediate-versus-answer ordering but still has only three successful top-1 swaps.
There are several limitations to my experiments. After the model-specific filtering, only a few samples remained in the datasets. Furthermore, the J-lens and R-lens that I used were fitted on only 25 background prompts, compared with 1,000 in Anthropic's setup. The datasets are also based around simple one-step reasoning chains, rather than more realistic and complex samples of latent reasoning.
After this investigation, I think that J-lens is a useful tool for extracting intermediate reasoning steps, which makes it powerful for model forensics and hypothesis elicitation. However, it seems that steering within the J-space is limited in Qwen and Gemma, so I am not convinced that the J-space acts as the global workspace in these models.
Concurrent work
While working on this, I also found J-lens on small open-weight models—a null-heavy replication. It investigates the smaller models Qwen2.5-3B-Instruct and Gemma-4-E2B-it and finds a Gemma/Qwen asymmetry consistent with my observations.
Acknowledgements
Thank you to Blue Dot for supporting me in this project with mentoring and a rapid grant.
TLDR: In their J-lens paper, Anthropic suggests that the J-space is a global workspace that the model reasons within, and supports evidence for this hypothesis on Claude models in a variety of settings. I replicated the multi-hop reasoning experiment on Qwen3.6-27B and Gemma 3 27B-it and found that counterfactual answer swaps outperformed intermediate swaps in three of four experimental conditions. This does not provide evidence to support Anthropic's global workspace hypothesis in open-weight models and instead suggests that J-lens is more useful for probing intermediate variables rather than steering outputs.
A few months ago, Anthropic published Verbalizable Representations Form a Global Workspace in Language Models and I was immediately excited about the prospect of being able to read part of a model's working memory. Beyond that, the paper hypothesises that intermediate reasoning concepts cannot only be decoded using the J-lens, but that the J-space is actually the global workspace in which the model reasons. Neel Nanda reviewed Anthropic’s paper and replicated the results on Qwen3.6-27B with moderate success: the verbal-report interventions were weakly positive, the multilingual and typo evaluations replicated cleanly, but the poetry and arithmetic results did not replicate.
Another task that Anthropic and Nanda evaluated was multi-hop reasoning, where prompts like "What is the colour of the fourth planet in our solar system?" require an intermediate reasoning step (in this case, Mars). Initially Nanda's replication of this multi-hop reasoning setting seemed positive, because swapping out the intermediate with a counterfactual increased the probability of the counterfactual answer, e.g. if Mars→red, then replacing Mars with Neptune should output blue. This would be consistent with the intermediate actually being used in reasoning, and therefore with the J-space being the workspace in which intermediate variables are stored. However, Nanda also found that on Qwen, injecting the counterfactual answer strictly outperformed injecting the counterfactual intermediate. This makes the result less compelling, because the intervention may be steering the model towards the final answer rather than the workspace being required for intermediate reasoning. Nanda suggests that the mismatch between his and Anthropic's results could be explained by the fact that the multi-hop facts used in his dataset were too easy, such that intermediates and answers are too linearly related.
I followed up on this hypothesis. I ran the multi-hop causal experiment on Qwen3.6-27B and Gemma 3 27B-it using Camila Blank’s workspace-bench dataset and the dataset in Anthropic’s repository. This leads to four experimental conditions: Qwen/Workspace, Qwen/Anthropic, Gemma/Workspace and Gemma/Anthropic. Furthermore, I ran these four conditions with the matched R-lens, which is supposed to produce cleaner early-layer readouts compared to J-lens.
My main findings were:
Setup
The models I investigated were Qwen/Qwen3.6-27B and google/gemma-3-27b-it and I used the pre-trained J-lens/R-lens pairs from camilablank/workspace-lenses.
The two multi-hop reasoning benchmarks that I used were:
The combinations of models and benchmarks resulted in the four experimental conditions Qwen/workspace, Qwen/Anthropic, Gemma/workspace and Gemma/Anthropic. Furthermore, I adjusted the datasets to the models' capabilities by checking that the models can answer the initial prompt correctly and by removing prompts with counterfactual options that were already among the model's top-10 answers for the initial prompt. After these filters, 32 prompts remained for Qwen/workspace, 44 for Qwen/Anthropic, 28 for Gemma/workspace, and 27 for Gemma/Anthropic. There was initially an overlap of 40 prompts between the two datasets, but after filtering only four overlapping items remained.
I used 24-59 as the workspace band for Qwen and 24-57 for Gemma. I ran the causal swap experiment in two arms:
I scored these arms against the same counterfactual answer token and measured the change in its probability. Additionally, I measured whether the counterfactual answer became the model's top-1 output. These experiments were run using both J-lens and R-lens. The full implementation can be found in the project repository.
Causal swaps shift probabilities, but rarely flip top-1 outputs
Figure 1: The probability shift of the counterfactual answer by model-dataset combination. Bars show the probability shifts when the intermediate counterfactual is swapped versus when the counterfactual answer is swapped, with 95% confidence intervals.
In all experimental conditions, swapping in the counterfactual intermediate increases the probability of the counterfactual answer, as shown in Figure 1. However, when swapping the counterfactual intermediate in, the actual answer only rarely flipped: 2/32 interventions flipped the answer in the Qwen/workspace condition, 4/44 flips occurred in Qwen/Anthropic, 2/28 answers flipped in Gemma/workspace, and 3/27 answers flipped in Gemma/Anthropic (shown in Figure 2). These success rates are significantly lower than Anthropic’s 54-70% reported for Claude. This means that the intervention often pushes in the intended direction, but this is usually not enough to control the model’s answer.
Figure 2: The percentage of top-1 flips across the four experimental conditions.
Answer swaps usually beat intermediate swaps
Anthropic hypothesises that the model computes an intermediate reasoning step in the J-space before coming to the final answer. This should mean that the intermediate swap occurs at an earlier layer than the final answer decision, so the swap of the counterfactual intermediate target should also take hold earlier than the swap of the counterfactual answer. As Figure 1 and Figure 2 show, I was not able to replicate these results on Qwen. The answer effects exceeded the intermediate effects by +0.158 on workspace-bench and +0.044 on the Anthropic set, aggregated over the workspace layers. Gemma/workspace also favoured the answer by +0.098. This failed replication matches Nanda's findings in his J-lens review.
However, the one exception is the Gemma/Anthropic run, in which the effect of the counterfactual intermediate swap outperformed the effect of the counterfactual answer swap, both in terms of probability shift and top-1 flip rate. The total flip rate after the intermediate swap was still relatively small though, with only 3 out of 27 intermediate swaps causing a top-1 flip.
R-lens matches or exceeds J-lens
Figure 3: Comparison of the intermediate swap probabilities across J-lens and R-lens.
I also compared the J-lens results with R-lens, a closely related variant that modifies the J-lens backward pass using LRP-style rules with the aim of producing cleaner early-layer readouts. On Qwen, the results were almost indistinguishable, as shown in Figure 3. On Gemma, R-lens achieved a larger probability shift on both datasets. The strongest effect by far is achieved in the Gemma/workspace condition, where R-lens produced 12/28 top-1 flips, while J-lens produced only 2/28 flips. These results suggest that R-lens is a useful extension of the J-lens framework, which is able to preserve or even strengthen causal swap effects.
Discussion
Let's recap what Anthropic and Neel Nanda's review concluded, before summarising how my replications fit in with those. For the multi-hop reasoning experiment, Anthropic found that the J-space acts as a global workspace, in which swapped intermediates affect the output often. Nanda was able to weakly replicate this on Qwen, but found that swapping the answer outperformed the swapping of the intermediate, which undermines the strength of the J-space as a workspace.
In my experiments on Qwen3.6-27B and Gemma 3 27B-it, I found that intermediate counterfactual swaps always increased the probability of the counterfactual answer, but rarely led to a top-1 flip. Similarly to Nanda, I also found that in three of four experimental conditions, the answer swap outperformed the intermediate swap. The exception to this was the Gemma/Anthropic condition, which was able to replicate Anthropic’s intermediate-versus-answer ordering but still has only three successful top-1 swaps.
There are several limitations to my experiments. After the model-specific filtering, only a few samples remained in the datasets. Furthermore, the J-lens and R-lens that I used were fitted on only 25 background prompts, compared with 1,000 in Anthropic's setup. The datasets are also based around simple one-step reasoning chains, rather than more realistic and complex samples of latent reasoning.
After this investigation, I think that J-lens is a useful tool for extracting intermediate reasoning steps, which makes it powerful for model forensics and hypothesis elicitation. However, it seems that steering within the J-space is limited in Qwen and Gemma, so I am not convinced that the J-space acts as the global workspace in these models.
Concurrent work
While working on this, I also found J-lens on small open-weight models—a null-heavy replication. It investigates the smaller models Qwen2.5-3B-Instruct and Gemma-4-E2B-it and finds a Gemma/Qwen asymmetry consistent with my observations.
Acknowledgements
Thank you to Blue Dot for supporting me in this project with mentoring and a rapid grant.