No-CoT performance has been a decent proxy for tracking the g-factor intelligence of base models (see work on latent multi-hop reasoning by Ryan Greenblatt). I think that this capability is very beneficial for reasoning, token efficiency and general intelligence in the way that helps AI to solve harder long horizon tasks. This is why I think labs could soon start automated research into architectures, driven by no-CoT loss and accuracy improvements.
For example: an automated AI research intern proposes an architectural modification to carry latent information across filler tokens and tracks improvements on no-CoT n-hop perfomance + sequential reasoning perfomance to save time. If no-CoT reasoning capability on filler tokens (Pfau et al., Goyal et al., and Greenblatt) improves, the model will have learned to use this information effectively, which would help with long-horizon cot RL.
This obviously points in a dangerous direction toward low monitorability. I think this is the next step labs could take, much like RLVR CoT reasoning DeepSeek-R1, OpenAI o1 was previously required to sustain exponential progress. It also allows models to improve token efficiency, making them much faster at completing very hard tasks that are not bottlenecked by real-world interaction.
This raises the next question: will mechanistic interpretability still be feasible? Will activation oracles still work? What if the model learns to generate two separate, parallel, J-space-like thoughts that read to us as benign reasoning, but carry an internal meaning for the model that entirely eludes us? It is hard to anticipate all the unknowns, but it is important to keep exploring them. These are just my two cents on latent reasoning and its trajectory.
No-CoT performance has been a decent proxy for tracking the g-factor intelligence of base models (see work on latent multi-hop reasoning by Ryan Greenblatt). I think that this capability is very beneficial for reasoning, token efficiency and general intelligence in the way that helps AI to solve harder long horizon tasks. This is why I think labs could soon start automated research into architectures, driven by no-CoT loss and accuracy improvements.
For example: an automated AI research intern proposes an architectural modification to carry latent information across filler tokens and tracks improvements on no-CoT n-hop perfomance + sequential reasoning perfomance to save time. If no-CoT reasoning capability on filler tokens (Pfau et al., Goyal et al., and Greenblatt) improves, the model will have learned to use this information effectively, which would help with long-horizon cot RL.
This obviously points in a dangerous direction toward low monitorability. I think this is the next step labs could take, much like RLVR CoT reasoning DeepSeek-R1, OpenAI o1 was previously required to sustain exponential progress. It also allows models to improve token efficiency, making them much faster at completing very hard tasks that are not bottlenecked by real-world interaction.
This raises the next question: will mechanistic interpretability still be feasible? Will activation oracles still work? What if the model learns to generate two separate, parallel, J-space-like thoughts that read to us as benign reasoning, but carry an internal meaning for the model that entirely eludes us? It is hard to anticipate all the unknowns, but it is important to keep exploring them. These are just my two cents on latent reasoning and its trajectory.