I think what's going on in the “Does Astra use neuralese?” debate is that there's an important sense in which models *already* do self-communication in neuralese: between each layer in the forward pass, the attention stream is already very hard to interpret, and clearly not in natural language.
Yet CoT monitorability is still a big deal and it'd be bad if all self-communication from models are no longer in natural language. So there have been two different proposed definitions[1] of what is "true" neuralese:
(My preferred) categorical definition: Since natural language currently gates recurrence in the standard transformer+CoT loop, having recurrence in neuralese is the natural category for whether something counts as "true" neuralese.[2]
The threshold definition. Total number X of serial steps before something appears in natural language. True neuralese counts as going above X. I think most technical experts who studied this issue, including many people at companies, prefer definition #2.
There are complicated technical arguments on both sides, but I think technical experts overall prefer #2 because they think it's more causally relevant (there’s nothing inherently more difficult about monitoring a 128-layer model looped 8x than monitoring a 1024-layer model), have less weird edge cases, etc. Whereas I (moderately) prefer #1 because I think definition #1 is conceptually cleaner and easier to enforce in practice.
But regardless of your preferred technical definition, an interesting sociological observation is that it appears (from reading their comments) that OpenAI (safety?) people appears to have taken definition #2 as gospel and fully reified it. Furthermore, if you take definition #2 as your only definition, it's impossible to defect against the no-neuralese taboo unless you specifically hit above the threshold (also as two side notes: 1) I'm not sure if either OpenAI or the safety community has ever clearly defined a threshold that people have agreed upon and 2) the number of layers for frontier models is not public).
Like from that perspective, adding recurrence a bunch is not obviously defecting more than adding more layers.
Though in my opinion, trying to fight against adding more layers is both less politically feasible and less useful (since mostly companies would only agree to a layer threshold because it's too technically hard to scale by adding more layers anyway).
So if we end up agreeing on definition #2 as the only definition we're talking about about re: neuralese, we can't say people are defecting unless they go above X layers (where, again, X has never been set, and also the specific number of layers is unknown).
Speed limits are often pushed in the best of times, but a speed limit without a number isn't a limit at all.
__
The definition above might be a bit overly technical, as well as philosophical and hard to understand, even for LessWrong, so maybe it’d be helpful to consider two (okay but not amazing) analogies for similar situations: the nuclear weapons taboo, and doping.
The nuclear weapons taboo is one where there’s an important sense that the categorical taboo makes no causally relevant sense: the smallest tactical nuclear bombs have a lower yield than the biggest conventional bombs. So why have a taboo against nukes in war, rather than against bomb sizes? Well, as Thomas Schelling has written about multiple times, the categorical line is necessary because a yield-based definition has no focal point: once you allow “tactical” nuclear use, there’s no natural place to stop. So a combination of scholars, activists, and farseeing world leaders managed to converge on the categorical definition.
For sports doping in contrast, a categorical definition is obviously superior in spirit but practically impossible to enforce in practice. The tests all have error rates and measurement issues, and besides, often the things that people dope with are naturally present in variable amounts in athletes' bodies. So WADA ended up using the categorical definition as the official one, but set various numerical thresholds to define doping in practice. But people generally understand that the numerical thresholds are the regulatory device, as a proxy for the underlying thing we care about. So when Lance Armstrong bragged that he “never failed a test”, people didn’t consider this as definitional evidence that he didn’t dope. And regulators eventually got him on teammates’ testimonies even when he never ended up failing a test.
Note that as a matter of regulation, it might well be the case that for practical reasons, you have to have threshold-based laws or regulations rather than categorical ones. And I think it’s fine to say that we shouldn’t legally punish anyone for breaking the spirit of the law if they didn’t break the letter, especially in somewhat ambiguous domains.[3] But as a social and societal matter, we absolutely should be comfortable with critiquing and punishing people and institutions who break the spirit of norms even if they didn’t break the letter.
This is true even if you have quantitative and categorical defenses that you aren’t breaking the taboo as much as is maximally bad. In the nuclear case, if a country uses a small tactical nuclear bomb on the battlefield defensively, I agree it’s not nearly as bad as using bigger bombs, attacking civilian populations, or escalating to all-out nuclear war. Indeed, if I’m a policy scientist working on nuclear deterrence when that happens, I’d absolutely be trying my hardest to either attempt to reinstate the nuclear weapons taboo or create a new categorical one (e.g. no nukes on civilians). But it’s still breaking a taboo to use nukes.
In the doping case, I agree that somebody who only dopes in training is not as bad as somebody who dopes both in training and during the race itself. But it’s still bad to dope!
In a broad sense, categorical taboos are much better than threshold-based taboos, and precise numerical threshold-based taboos better than ambiguous ones. The last one is barely a taboo at all.
I recommend people thinking about which norms and taboos to adopt and shape be significantly more mindful of the relevant practical complexity in understanding and enforcing said norms, both in general and in the neuralese case.
People have also alluded to a third definition, which is that the thing we “actually” care about is CoT monitorability so we should define neuralese as the obverse of CoT monitorability. Imo if you try to think it through, this consequentialist definition is even more unnatural and cursed. We do not have a crisp enough conceptual definition of moniterability, nor do we have sufficient numerical evaluations and thresholds for what counts as “sufficiently moniterable.”
This is also closest to the original AI 2027 definition, which afaik is what kick-started most public conversations about neuralese in technical circles. Because it came out 1.5 years ago, of course the wording and definitions are slightly less precise than current discussion.
Though again I will point out that even steelmanning this perspective is somewhat putting the cart before the horse here, since, again, no numerical thresholds have yet been set.
I think what's going on in the “Does Astra use neuralese?” debate is that there's an important sense in which models *already* do self-communication in neuralese: between each layer in the forward pass, the attention stream is already very hard to interpret, and clearly not in natural language.
Yet CoT monitorability is still a big deal and it'd be bad if all self-communication from models are no longer in natural language. So there have been two different proposed definitions[1] of what is "true" neuralese:
There are complicated technical arguments on both sides, but I think technical experts overall prefer #2 because they think it's more causally relevant (there’s nothing inherently more difficult about monitoring a 128-layer model looped 8x than monitoring a 1024-layer model), have less weird edge cases, etc. Whereas I (moderately) prefer #1 because I think definition #1 is conceptually cleaner and easier to enforce in practice.
But regardless of your preferred technical definition, an interesting sociological observation is that it appears (from reading their comments) that OpenAI (safety?) people appears to have taken definition #2 as gospel and fully reified it. Furthermore, if you take definition #2 as your only definition, it's impossible to defect against the no-neuralese taboo unless you specifically hit above the threshold (also as two side notes: 1) I'm not sure if either OpenAI or the safety community has ever clearly defined a threshold that people have agreed upon and 2) the number of layers for frontier models is not public).
Like from that perspective, adding recurrence a bunch is not obviously defecting more than adding more layers.
Though in my opinion, trying to fight against adding more layers is both less politically feasible and less useful (since mostly companies would only agree to a layer threshold because it's too technically hard to scale by adding more layers anyway).
So if we end up agreeing on definition #2 as the only definition we're talking about about re: neuralese, we can't say people are defecting unless they go above X layers (where, again, X has never been set, and also the specific number of layers is unknown).
Speed limits are often pushed in the best of times, but a speed limit without a number isn't a limit at all.
__
The definition above might be a bit overly technical, as well as philosophical and hard to understand, even for LessWrong, so maybe it’d be helpful to consider two (okay but not amazing) analogies for similar situations: the nuclear weapons taboo, and doping.
The nuclear weapons taboo is one where there’s an important sense that the categorical taboo makes no causally relevant sense: the smallest tactical nuclear bombs have a lower yield than the biggest conventional bombs. So why have a taboo against nukes in war, rather than against bomb sizes? Well, as Thomas Schelling has written about multiple times, the categorical line is necessary because a yield-based definition has no focal point: once you allow “tactical” nuclear use, there’s no natural place to stop. So a combination of scholars, activists, and farseeing world leaders managed to converge on the categorical definition.
For sports doping in contrast, a categorical definition is obviously superior in spirit but practically impossible to enforce in practice. The tests all have error rates and measurement issues, and besides, often the things that people dope with are naturally present in variable amounts in athletes' bodies. So WADA ended up using the categorical definition as the official one, but set various numerical thresholds to define doping in practice. But people generally understand that the numerical thresholds are the regulatory device, as a proxy for the underlying thing we care about. So when Lance Armstrong bragged that he “never failed a test”, people didn’t consider this as definitional evidence that he didn’t dope. And regulators eventually got him on teammates’ testimonies even when he never ended up failing a test.
Note that as a matter of regulation, it might well be the case that for practical reasons, you have to have threshold-based laws or regulations rather than categorical ones. And I think it’s fine to say that we shouldn’t legally punish anyone for breaking the spirit of the law if they didn’t break the letter, especially in somewhat ambiguous domains.[3] But as a social and societal matter, we absolutely should be comfortable with critiquing and punishing people and institutions who break the spirit of norms even if they didn’t break the letter.
Two more points: OpenAI’s defenses include that the total serial depth of Astra’s computation isn’t very high, and also that we shouldn’t kick off a race to unmonitorable neuralese by talking about it as a done deal. I mostly agree with these arguments as stated (though the first one has some very important nuances), but they need to be taken in the context of breaking a taboo.
This is true even if you have quantitative and categorical defenses that you aren’t breaking the taboo as much as is maximally bad. In the nuclear case, if a country uses a small tactical nuclear bomb on the battlefield defensively, I agree it’s not nearly as bad as using bigger bombs, attacking civilian populations, or escalating to all-out nuclear war. Indeed, if I’m a policy scientist working on nuclear deterrence when that happens, I’d absolutely be trying my hardest to either attempt to reinstate the nuclear weapons taboo or create a new categorical one (e.g. no nukes on civilians). But it’s still breaking a taboo to use nukes.
In the doping case, I agree that somebody who only dopes in training is not as bad as somebody who dopes both in training and during the race itself. But it’s still bad to dope!
In a broad sense, categorical taboos are much better than threshold-based taboos, and precise numerical threshold-based taboos better than ambiguous ones. The last one is barely a taboo at all.
I recommend people thinking about which norms and taboos to adopt and shape be significantly more mindful of the relevant practical complexity in understanding and enforcing said norms, both in general and in the neuralese case.
People have also alluded to a third definition, which is that the thing we “actually” care about is CoT monitorability so we should define neuralese as the obverse of CoT monitorability. Imo if you try to think it through, this consequentialist definition is even more unnatural and cursed. We do not have a crisp enough conceptual definition of moniterability, nor do we have sufficient numerical evaluations and thresholds for what counts as “sufficiently moniterable.”
This is also closest to the original AI 2027 definition, which afaik is what kick-started most public conversations about neuralese in technical circles. Because it came out 1.5 years ago, of course the wording and definitions are slightly less precise than current discussion.
Though again I will point out that even steelmanning this perspective is somewhat putting the cart before the horse here, since, again, no numerical thresholds have yet been set.