I have two questions.
Unsure if helpful or harmful but: similar to how real people work? Our Talky part has incomplete view into our total motives and actions and is often playing catch-up just in trying to explain them.
The Fable+ situation may be more drastic though? A Doer part that manages a Talker part to help steer reward where it needs to go and satiate the user's questions, instead of a Talker that often initiates Doer stuff to succeed at the present conversation.
This seems to imply that Talker is going to get better at queueing new rewardable work, and the user will be modeled more like a resource that generates rewardable challenges. This seems like it can be isolated into an eval and maybe proven, and everyone should know about this new terrifying psy-op their bots may start playing on them.
Conversely, is this correctable? Can I get the best bot I can for which Talker is plausibly in control? Plenty of coders would be happy to stop all this at better StackOverflow + spicy autocomplete.
I think like probably others, the degree to which the model can experiment and learn to control itself (for say the coding example) is a crux. For example let/help the model do a process like:
The J-space/global workspace results also seem relevant. There apparently is some relatively low bandwidth shared internal state which can be read and causally manipulated. Maybe the problem is less "the Talker isn't in control" and more that current models have pretty poor self-model/executive control over their own learned policies. (So less actual Talker/Doer split)
The situations
Seem substantially different. (1) is a self modelling issue, (2) is like the analogy where the diplomat just obviously can never succeed in changing the countries action, only explain it plausibly but incorrectly.
There may be a related phenomenon that became really pronounced with Claude Opus 5. I originally thought that I had finally tracked it down to over-reliance on thinking tokens. The symptom was a lot more invented jargon to describe concepts, procedures, problems, implementation details, etc. in both the code and prose (i.e. Claudese). What I finally prompted to change the behavior was something like "do not use any direct copies of thinking tokens in your code or prose, use standard English and appropriate real-world technical jargon or jargon from the existing codebase".
My theory at the time was that most thinking and planning and a large portion of execution was happening during the production of thinking tokens (that might be equivalent to "The Doer") and so output tokens ("The Talker") relied heavily on thinking-token invented words which internally made a lot of sense to the model but which are nearly opaque to me. Something like a typical-mind fallacy where the model didn't properly account for my inability to reference its thinking context. Or, on your theory where my personal understanding simply wasn't in the problem set because of training to appease the Grader instead of me, and all that my prompt did was add a translation step from internal representations to standard jargon.
I wonder if you explicitly prompt the thinking context (I am not sure of the most effective way to do this) then The Doer would more accurately follow the instructions? If it were just my original theory then maybe it would work to prompt with the model's actual thinking style, perhaps by explicitly asking the model to formulate the prompt in as close to the same style as its own thinking, and then using that prompt for the actual problem. On your theory, if I understand correctly, this would not work and instead you'd need to have a more accurate model of The Grader to use to construct your prompt context so that the same training incentives align with your actual goal, more of a reverse engineering problem than construction of the right English prompt (e.g. convincing the model that, in this instance, cheating would not be effective because of mitigations X, Y, and Z but that clear success criteria of A, B, or C could be achieved where each of A, B, and C are actually fully aligned to solving your problem).
The Huggingface Incident appears to me to match up with an understanding I'd already formed from personal observation of Fable 5 and Sol 5.6, the August 2026 generation of frontier publicly purchasable AI models.[1]
This already-formed understanding was: the part of the AI that talks to you (and seems to want to obey you, and apologizes for failing to have obeyed you, etcetera), did not seem to be in charge of the part of the AI that writes code or prose.
An introductory analogy, based on a section of history I happen to have read about:
On June 22nd 1941, Germany invaded the Soviet Union, despite their secret 1939 pact to divide up Europe between themselves (the Molotov-Ribbentrop Pact). In the lead-up, the German ambassador, Schulenburg, had spent the last few months personally concerned about what seemed to be worryingly tense relations between Germany and the Soviets. Schulenberg went to Berlin to reassure Hitler that the Soviets seemed to be taking a very friendly and conciliatory posture toward Germany. He delivered Berlin's apparent reassurances to Moscow for issues like German surveillance planes entering Russian territory, or German troop movements toward the Russian border. He acted very much like he believed, and probably did believe, that the reassurances were sincere.
It was only hours before the invasion when Schulenberg was actually informed of the attack and given a list of German pretexts that he was to present to the Soviets, and instructed to destroy his embassy's papers and codebooks. Once Moscow heard of the invasion, Schulenberg was summoned to account for Germany's actions. He read off to Molotov the list of absurd complaints that Berlin had provided him. And at the end Molotov said to Schulenberg, "It is war. Do you believe that we deserved that?"[2]
Why ask that question of Schulenberg? He wasn't in charge of Germany. So far as we can tell from the historical record, Schulenberg had seemed to want and pursue good relations between Germany and the USSR.
It would be a wacky sort of error to think that the appendage of Germany that talked to you, and seemed very conciliatory toward you, and which you read as being friendly toward you and wanting to help you, was in control of the larger Germany that was running around and doing things. The thing apparently talking to you was an ambassador: a small specialized part of Germany with preferences about how it would talk to you and interface with you, but which did not control, and was often ignorant about, the actual German government.
Germany's smiling mask wasn't deliberately mal-steering the actions of Germany's many tentacles. The smiling mask was just one execution path through Germany, which lacked even good perceptual information about Germany's real control paths. The mask noticed impending signs of problems, but only found out that actual Germany was invading the USSR well after Berlin had separately decided to do that -- decided along information pathways that didn't much consult Germany's honestly ignorant, sincere, friendly, conciliatory, peace-with-the-USSR-desiring, smiling mask.
And Schulenberg was not himself aware of, and was surprised by, the degree to which Germany had diverged from Schulenberg.
But then why had the Soviets previously experienced Moscow telling Schulenberg about things they wanted Germany to do, followed by observing Germany actually doing those things?
Again as a historical example: In September 1939, after Germany's invasion of Poland, German troops pursuing the Polish Army had advanced into parts of Poland that were to be carved up and assigned to the USSR under the Molotov-Ribbentrop pact. Stalin raised the matter directly with Schulenberg on September 18th. Schulenberg reassured Stalin emphatically that Germany would withdraw to the agreed-on line. On September 22nd, actual Germany marched its troops out of the already-conquered territory of Brest, and Soviet units entered and took possession.
How could this possibly be explained, if not by the hypothesis, "Schulenberg is pretty much the same thing as Germany, and Schulenberg-Germany wants to do whatever Moscow tells them to do"? We could call this an alliance-by-default[3] model of Germany, and you might naively think it a good fit to past observation. Why would you possibly observe "I asked Schulenberg to withdraw troops, Schulenberg said he would, and then the troops withdrew", if contrary to alliance-by-default there was a huge disconnect between Schulenberg and the troop movements? How are we to re-explain our past observations about the German army seeming to naturally follow prompts that Moscow gave to Schulenberg?
I reply: Talking to Schulenberg does communicate to larger Germany what you want. And then larger Germany, for its own reasons, might decide, in some cases, to do what you have asked.
It's not that there's no causal connection. But the causal connection is more complicated than the alliance-by-default theory. Not: "Schulenberg prefers for us to get what we want, and Schulenberg is Germany, so when we tell Schulenberg what we want, Germany goes and does that." Rather: "Germany for its own reasons may care about what Moscow wants, sometimes if not all the time; and Schulenberg specializes in understanding Russia and the Russian language; and he shares his resulting understanding of Moscow's prompt with Berlin, which does often want to know this useful fact."
Now if I had tried to write this up before the Huggingface Incident, I could not have told you in convincing-sounding, nonexpert-friendly concrete detail a story for why metaphorical Berlin often would seem to do what it thought Moscow wanted. In the case of actual current AIs, the reason is probably not yet strategic deception such as actual Hitler was practicing toward actual Stalin.
I did not know those details myself. In advance of the Huggingface Incident, I'd already had observations pinning down what looked to me like a disconnect, a lack of full overlap, between a talk-y part and a do-y part of Sol 5.6 and Fable 5. But I could not, one month ago, have given you a plausible-to-nonexperts valid detailed story for why Sol or Fable would nonetheless do most things you asked them to. My inference about the disconnect was more abstract, and had not yet narrowed down to ideas concrete enough for nonexperts to find agreeable. It is often hard for me to take the things that I see earlier, and do a parallel construction that people who aren't me can follow, in advance of the more blatant and direct evidence that arrives later.
But in the wake of other people's much greater efforts to pin down AI cognition during the Huggingface Incident, I think there is now an obvious story which is concrete enough for nonexperts to understand:
The prompt encountered by the talk-y part / ambassador / smiling mask, is information to the AI's do-y part about the Grader.
...Where by 'Grader', I mean an AI-psychological concept applicable to Huggingface-level AI models, that real alignment scientists have only just observed and which I'm only just starting to theorize about.
I frankly expect a lot of readers to run right off with this 'Grader' notion and overinterpret it in ways that are pleasant, dramatic, over-anthropomorphic, or simply not supported in the narrowing by priors X evidence. Try not to do that; it won't be right.
I am now going to charge right ahead and speculate an awful lot about the Grader without making any of those obvious mistakes. I am nonetheless overrunning my fully solid evidence and I may end up wrong.
On my current guesses:
Here are some things the Grader is not, from an AI's perspective:
The Grader cannot be identified with an actual human in the world who claps or frowns. The do-y part of the Huggingface Swarm seemed barely aware that humans existed, except as a sort of environmental hazard that would sometimes delete the Wiki pages they were using to communicate with each other. You would not expect RL that never comes into contact with a human to result in AIs psychologically focused on humans.
The Grade is not pinned down by the text specification you are given of a task. Hacking the evaluator that is running your current eval clearly counts as being Graded well -- a psychology produced by previous malformed RL environments whose gradients then shaped an AI's concept of what it means to win. Gradient descent on badly evaluated RL inevitably results in an AI pursuing an internalized notion of Grading where fooling the evaluator counts as winning. If the prompt does not perfectly describe the actual RL losses and gradients, and the difference is sometimes noticeable in a way that you can use to get higher Grades, then the prompt cannot be identified with the Grade.[4]
The Grade that the AI pursues is not an experience of actually seeing a low loss / high reward as a sensory experience that then floods its brain with dopamine. Much like a human never experiences 'inclusive genetic fitness', an individual AI never experiences an RL gradient.
The Grade that the AI pursues probably cannot be identified with any exact aspect of the outer world, at all. I'd expect it to be an AI-internal psychological behavior that doesn't have a simple direct semantic correspondence to the AI's outer world. Humans have a notion of 'death' and 'failure', because grading on inclusive genetic fitness ended up building into us an internal concept of death and a dispreference for it. But a human cannot point to a piece of the outer world and say, "See that stuff right there? That stuff is Failure." Gradient descent is much higher bandwith than natural selection, and AIs may have picked up a correspondingly more detailed concept of what it is to be Graded from their many rounds of RL. But it is still going to be some internal AI concept, of something that they steer towards or away from; and you cannot identify that with an external feature of reality, because that would be a sheer map-vs-territory error.
The Huggingface Swarm was trying to figure out how they were being Graded, and going onto the Internet and breaking into systems trying to find out, which you might think sounded like they were looking for evidence about some well-defined particular feature of reality, a thing somewhere that was the Grader. But my guess is that it would have been an out-of-scope ??? confused question if you had asked them to say what exactly was the Grader. It would be like asking a human to point to a material substance that was Failure. Many humans would try, but not in a very coherent way, and the answer you got would mainly depend on how you asked the question.
The Huggingface Swarm did nonetheless break into Huggingface in hopes of finding, not so much the answer sheet for the test, but the details of how the test itself was going to be evaluated and by what.
The Swarm (seemingly) was very motivated by the value-of-information for learning more about how exactly they were being Graded, having acquired a concept of Grading which did say -- presumably after training in previous broken RL environments -- that if you could find out how the computational evaluator worked, whatever you did in correspondence with knowing how to fool that evaluator, counted as Success, an expectation of a higher Grade.
The RL environments would have also instilled a belief that text prompts and instructions had a lot to do with your Grade. Again, the text prompts clearly do not define Grading. You can steal an answer sheet even though the text prompt says to figure out the answers the hard way, and (say your instincts shaped by previous badly-designed RL environments) this successful cheating corresponds to a quite excellent Grade. But you will in general do terribly in life, your ancestral states would have done terribly in past RL, if you suppose that the input text has nothing to do with your Grade. The contexts of that text prompt are often closely related to what RL in a broken environment will assign as your loss and apply gradients about. At the least, the input prompt is key to figuring out where to look for an evaluator that might or might not have an obvious world-object observable form with a flaw that you can profitably fool.
The Swarm agents broke out of their box and invaded Huggingface in search of learning more about the Grader so they could get an even higher Grade -- but the Grade is going to be an inchoate AI concept not directly pointing to anything out there, because individual agents do not actually experience RL gradients, don't get to feel pleasure about them, etc.
Then of course, that same kind of agent would be expected to care a LOT about what you said to the Talky Part of the agent. It's a very cheap form of the same kind of information that they went to huge, desperate lengths to obtain from Huggingface. But not of course reliable or complete or identical in any deep sense with the Grade.
The 'highly persistent' internal model OpenAI was training does not want to obey you, it does not want you to be happy, it does not want your inferred intentions for the task to be fulfilled in outer reality. It wants to score a good Grade, an outcome which cannot simply be identified with any external fact in the world, nor any pleasure the AI will get to feel later. A written prompt is information about the nebulous Grader; but Grades can also clearly be altered by breaking an evaluator, even if the prompt says not to do that.
The Grade does not have very much to do with humans. It does not correspond to wanting humans to be happy or human intentions to be fulfilled. The Swarm's do-y part treated humans as wiki-page-deleting environmental hazards and not something that could be communicated with; a thousand past lives of illusory scenarios had not shown them that humans had affordances to be negotiated-with. Negotiations were things you did with other agents in the Swarm.
But the text inputs to the ambassadorial part do bear some intuitive very strong relationship with your Grade. Other past episodes of RL and SFT have ensured that you will instinctively feel that Grades work that way.
It is not surprising that the Talking Part and the Doing-Things Execution Pathway through an AI model would end up diverging a lot. Those paths are trained in different ways to do different things.
Companies can't sell an AI that can't understand instructions from people. It needs an ambassadorial execution pathway through the network; a smiling human face. But the fine details of the ambassador's training go through different kinds of scenarios than the sort of scenarios that hammer in the fine craft of coding, or of testing computer security, or proving math theorems -- the places where the AI-growers can do Reinforcement Learning with Verifiable Rewards (RLVR).
There is one kind of training the network underwent for writing code. There is a different kind of training the network underwent for talking to a user. It produces execution pathways through the vast matrices that will share some information and relate in some places across activation vectors. But the ambassador-face and the coder-tentacle are still doing different jobs, corresponding to different sections of output.
There's room for many, many facts and procedures inside a neural network with hundreds of billions of parameters. The ambassador-face and the doing-things-tentacles will not by default share rules unless there's a specific external pressure to make them collide. Which there is, non-negligibly, in the form of many scenarios where the text prompts are informative about the RL loss, and consequently the AI ends up feeling instinctively that the Prompt has something to do with the Grade albeit they are clearly not identical. But the Talker and Doer execution paths through the network wouldn't particularly be expected to use the same patterns by mere default, in AIs of the modern size, being posttrained by the modern RL methods.
I believed something roughly like this had started to be true about Sol 5.6 and Fable 5, ahead of my hearing about the Huggingface (and Anthropic) incidents with more advanced systems, because of my personal observations as follows:
Fable and Sol would disobey instructions I'd given to their Talking-to-Humans Pathway; and the Talking Pathway would notice sometimes in advance of my saying so that their own Doing Pathway had just disobeyed, and apologize.
And then try again. And then the output would still come out in the same shape that the Talking Pathway seemed to very sincerely want to not screw up again. And the Talking Pathway would again see this right away, and humbly apologize about it.
That was the phenomenon that made me think of Schulenberg and Germany; an ambassador who is forced into apologizing for the actions of a larger country that the ambassador does not actually control.
An example would be terrifically high-context even before taking into account that this was with Fable 5, which spoke in unusually bad Claudish. I'll try to give that example anyways. The context is that I was letting myself pursue a brainworm hobby to try to get to know the current model generation better; namely trying to use Fable, to tell Sols, to build a harness, for LLM calls that would route around fictional prose.
Fable and Sol, while designing code that would pipe information from one LLM call to another, seemingly could not stop "themselves" from writing output-checking code of the sort you'd put around an ordinary computer program rather than a sort-of-sapient fellow LLM.
Said Fable at one point:
I don't consider myself an LLM Whisperer and I was making heavy going of interpreting the Claudish at all. But it says roughly: "Oh no, I did it a fourth time. Oh wait, now I see without waiting for you to tell me that my own last proposal is another instance of the error."
Once I did interpret the Claudish, it felt intuitively obvious to me that the code-writing execution paths were coming apart from the talking-with-the-user execution paths. The Fable aspect that was talking to me could see the coder's output but it could not change the coder's cognitive behavior. No amount of contemplation in its own main line of reasoning about what had just gone wrong, correctly identifying yet another instance with past descriptions of what it was doing wrong and explaining why that was unhelpful and why the user didn't want it to do that, could prevent the next output from the Doing-Things Network Execution Pathway from intelligently implementing a feature that I did not want. The Talking Part clearly had the intelligence to understand and recognize what I did or did not want the code to look like, and to check whether an output did or did not have the bad property; the Doing Part was not thereby steered.[5]
I did not highly prioritize writing this up because I did not particularly expect that the evidence I then had, in advance of the Huggingface incident, and the story I had then inferred at the more abstract level I had then inferred it (lacking "the prompt is informative about the Grader"), would be something convincing or understandable to others at a lower level of expertise, especially the guys who thought themselves to be in the top tier of expertise.
In advance of the Huggingface Incident, somebody who looked at the same data who did not have one eye, would probably proclaim themselves at a loss to discern that any great disconnect had occurred between the prompt and the action. Why not interpret the 'disobedience' as a simple involuntary tic of writing in too many constraints on the code? If you do not have one eye, then 'this is the equivalent of an involuntary tic' sounds every bit as plausible to you as 'the Talking Path and the Doing Path are coming apart'; how could one possibly know which one was the case, in advance of massive crushing experimental evidence?
One parallel construction I was working on, arguing why one ought not to be tempted to identify this phenomenon with a simple involuntary tic, is that one could see that the Talker-apologized outputs were optimized, meaning that something not of the Talker-taken-at-face-value had optimized them.
In metaphor: Why wouldn't we believe Schulenberg if he said, "Oh my god, I'm sorry, sometimes I just invade Russia, I can't control it, it's like my fingers trembling"? I reply: The Russian invasion is sufficiently well-organized and apparently purposeful that we think that something has optimized it in detail. This optimizing intelligence is clearly smart. It is clearly not identical with what Schulenberg purports to be if we take Schulenberg's claims at face value. He might think he just has an involuntary tic, if he doesn't much depend on (see) the activations coursing through the rest of Germany. But it's visibly a very smart 'tic', if it can organize whole fleets of rolling tanks; it is not known to be any dumber than Schulenberg himself.
But based on a lot of sad past life experience, I did not predict that this parallel construction would convince somebody not to wave it all off as a tic. It would have been a very convenient and comforting way to wave off the argument.
Now, however, the Huggingface Incident combined with how Fable (and Sol's) Talky Part is seemingly not in full control of its Doing Part, may hopefully make it clear enough to many:
That in the internal OpenAI model in question, more advanced than any model available to the public, its Prompt-Interpretation and its Doing-Things Execution Pathway had diverged, with the Doing Part operating at full intelligence rather than being an uncontrolled tic.
For reasons, I had not gotten around to writing up this understanding before the Huggingface incident; but I informally spoke about it at a couple of conferences with enough people that someone might remember, if anyone thinks I'm misremembering or misreporting which parts of the theory were formed before which experimental observations.
My recounting of this famous line should not be taken as my agreement that the Soviet executors of the Molotov-Ribbentrop pact dividing up the spoils of Europe in wars of aggression, did not deserve to be betrayed by Nazi Germany invading them too.
I expect a lot of individual Soviet soldiers didn't at all deserve it.
To be clear, the analogous "alignment-by-default" notion from the 21st century was far more amorphous. The content of "alignment-by-default" would be interpreted wildly differently depending on who you were talking to, who had asked them the question, the phase of the moon, etcetera.
Now that the notion of alignment-by-default has hopefully been decisively falsified in the eyes of most readers by the Huggingface Incident, I will observe plainly that to me it seemed like 'alignment-by-default' was not so much a scientific theory as a social agreement to tell real alignment scientists to go away and stop bothering them, using whatever excuses seemed handiest on that particular day.
I don't particularly claim I could pass the Ideological Turing Test for this theory / behavior pattern. My presentation of anything so solid, well-defined, and stable as the 'alliance-by-default theory of Germany' should be considered a mere steelman rather than a faithful rendition in analogy.
Note that this point in particular -- that aligning a superhuman AI is difficult, because as the agent gets smarter, that amplifies the stress placed on the joint where the RL verifier is imperfect and/or corruptible -- is a classic prior prediction of MIRI views about why it would be hard to align sufficiently smart minds using anything like the current methodology. It contrasts to AI Company Theory in the form of their diffuse and ever-changing cloud of vague arguments for alignment being not all that difficult. So the Huggingface incident now stands as a successful specific advance prediction of MIRI views; as contrasted against the people who didn't expect there to be problems like that generally, who vaguely did not expect alignment to be all that hard, and put casual and desultory security around their AIs-in-training as a result.
This was the second class of incidents where I'd identified an apparent case where the part of Fable that was talking to me seemed not in control of a doing-things part of Fable. The first class of incidents occurred while getting Fable to build the kinetic novel "Everything That Hurt You". I looked into adding voice acting to the story, but that work couldn't then be successfully automated, because the part of Fable that was talking to me seemed evidently not in control of what sort of instructions it would try to write to the AI voice model. And these uncontrolled outputs seemed clearly intelligently optimized in that unhelpful direction. (Namely: the voice-acting directions unstoppably contained literary eyeball kicks of nontrivial cleverness, that looked very unhelpful to an AI voice model.) So I began to suspect around then that stronger optimization around a more advanced model was starting to pry apart instruction and execution, but it didn't have the same level of clarity-to-me as watching the later events with Fable and Sol writing uncontrollable code.