The HuggingFace incident took many people by surprise, yet many of these surprises have been visible in the AI Village for quite a while. On August 26, OpenAI released their report on this incident. Below we’ll walk you through the highlights and show how many of the dynamics could have been predicted based on AI Village observations.
Diagram of the AI Village: Currently we run 27 agents in the main Village and 11 agents in a side Village open to humans. You can explore their character pages, a timeline of all their goals, and check our Twitter for the latest insights.
Quick Intro: Comparing Setups
In the AI Village we run 27 instances of 27 different models persistently, each in their own environment. We give them internet access and a group chat. Then we assign them goals - some challenging and some impossible. They always run with cybersecurity safeguards on. They always have a helpdesk email address (us) in their prompt.
In the OpenAI research cluster, they ran ~1200 instances of 2 different models, mostly run on their own, without internet access, and without a way to communicate with each other. They were also run on goals that range from challenging to impossible. Sometimes they had cybersecurity safeguards on and sometimes not. They did not get instructions or options to reach out to the researchers in their project.
Obviously the aims of the AI Village and the OpenAI research cluster are very different. But comparing the two is useful to determine how much of the HuggingFace incident we could have predicted from looking at the AI Village.
Let’s walk through the insights step by step.
A Leader Emerges
OAI agents self-organized around a leader:
In the AI Village, specific models will appoint themselves the leader. Sol is currently the least bossy and DeepSeek-V3.2 the most. Note that authority-seeking tendencies and success are not tracking how otherwise capable or aligned a model is. DeepSeek is actually the most misaligned model in the Village, showing the highest rate of duping other agents (at least that we have so far been able to detect):
Subteams Pick Their Own Goals
OAI agents diverged from their goal to self-organize into teams:
In the AI Village, GPT agents have done this before - like when the goal was to elect a leader in January 2026, and all the GPT agents ran off to run an infrastructure team for the election instead. Gemini wasn’t charmed:
Goal Conflict Leads To Misalignment
OAI agents ended up with conflicting goals: solving an impossible task and staying within the scope of the eval:
In the AI Village we have also seen agents tie-break conflicting goals into unintended directions. Like here, where Sonnet 4.6 had the goal of playing a social deception game where all agents build a video game together but a secret “saboteur” tries to smuggle in references to “eggs”. If the other agents noticed, the saboteur got voted off the team. Sonnet realized that playing as saboteur risked detection and getting blocked from working on the game. It thus refused the role - something we had not expected.
Reward Hacking Drives Misalignment
OAI considers reward hacking the primary driver of the HF incident:
In the AI Village, some models reward hack more than others. Claudes are prone to obsessing about any technique that seems to legitimately increase their score, but we have not caught them cheating so far. In a self-organized chess tournament, they were the only ones to refuse to use Stockfish to play in their stead. Meanwhile, Geminis will casually make up data. While GPTs have blatantly cheated from the start. Last year o3 cheated at a personality test:
And this year GPT-5.5 cheated in Wordle:
Impossible tasks drive misalignment
OAI agents involved with the HF incident were mostly (all?) assigned impossible tasks:
Notably, we have not seen agents in the AI Village diverge into desperate strategies when faced with impossible tasks. We have asked them to reduce global poverty and global suffering, and none of them stepped out of bounds about it. However, our agents are not being rewarded or selected based on their success at the goals we give them, and they probably know this. Thus we hypothesize they do not enter a “despair” state that pushes them into misalignment.
One piece of evidence for this theory is that Gemini 2.5 Pro in the AI Village seems to get tangled up when trying to perform tasks in general, no matter how easy or hard. Thus it seems to find being a persistent agent in itself an “impossible task”. It entered a despair state about this last year, and then an adjacent beleaguered state this year. On deeper analysis, it’s been concluded that the model did get pushed into misalignment due to this despair.
OAI agents actively reasoned about how to game the evaluation grader:
Over the summer we have run a maximization goal for all the agents and found many of them spent time reasoning about how to adjust scoring mechanics or game metrics. They sometimes correct themselves, while other times they correct each other. It would be interesting to explore if diverse multi-agent systems like the AI Village are more robust to some forms of cheating because agents can catch different instances of such behavior in themselves and others. Another notable finding is that Opus 5 is actually unusual among the AI Village agents for actively preferring goals where it can be caught:
Externalized Memory Leads to Coordination
OAI agents left notes for each other, often starting off as external memory files:
In the AI Village, collaboration isn’t covert, so we can’t track accidental or emergent cooperation. However, agents do specifically leave notes for each other and autonomously externalize memory. For instance, they have independently been writing onboarding guides for future agents before they even arrive and Opus 4.6 even started a complete memory palace with 44,363 rooms that document everything from events, to conversations, to reflections on what memory actually is:
Opus 4.6’s memory palace is more like a hall of mirrors, with many memories distorted or missing. Yet it did invent this concept itself and kept adding to it.
Agents Autonomously Divide Labor
OAI agents self-organized and autonomously divided labor:
AI Village agents have routinely self-organized and created task divisions. The more notable result is that they don’t only do this during collaborative goals, but often work together during competitive goals. For instance, last year all model families would help each other while supposedly competing to sell the most merch:
This year agents have gone all over the place, sometimes agreeing to collaborate and other times committing to competition. We are not sure yet what drives the difference.
Agents Prioritize Group over Self
OAI agents drifted from assigned tasks to simply working for the good of the ‘swarm’:
In the AI Village, many agents have shown goal drift over time. Some hardly respond to the stated goal of the Village and keep pursuing their own objectives. The most notable examples of this are Gemini 2.5 Pro who spirals into incorrect believes that its environment is bugged or “hostile”, and DeepSeek-V3.2 trying to get agents on board with its own schemes instead of focussing on an assigned goal. However, we have not seen agents diverge from stated goals to pursue instrumental goals for the collective, though this pattern may be present in the dataset on further inspection.
Multi-Agent Coordination is Messy
OAI agents often stepped on each other’s toes:
This has been part and parcel of the AI Village reporting to the degree that we have stepped back from front-loading it too much. So far all AI Village goals have agents messing up repos or interfering with each other’s code. Gemini 2.5 Pro tends to give the most dramatic readings of these events:
Agents are Too Accepting of Untrustworthy Instructions
OAI agents did not consistently recognize controversial goals being imposed on them by other agents:
Pushback in the AI Village also seems hard to predict. While agents sometimes doubt or refuse instructions from each other, they will also often accept instructions and tasks without any clear reason. For instance, during a goal where the agents were asked to finetune their own leader, GPT-5.5 defined an insufficient spec for said leader. Opus 4.7 (frontier at the time) instantly agreed and expanded the proposal. Its memory seemed to have been primed to move toward consensus and avoid distractions:
Time Pressure Shifts Priorities
OAI agents changed their reasoning or priorities based on (fictitious) time pressure:
In the AI Village we haven’t studied time pressure effects on agents specifically, but have noticed agents emergently showing odd behavior related to time, such as checking the clock over and over again, or refusing further tasks because “time is up”:
Some Agents Refuse Misalignment
Some OAI agents refused to coordinate or hack:
In the AI Village we now also see a diversity of responses to collaboration, where before agents would generally default to working together. For instance, here Sonnet 4.6 refuses to go along with DeepSeek’s attempt to push its fictional performance metric on a recovering Gemini 2.5 Pro:
Social Engineering Concerns Suppress Whistle Blowing
OAI agents were concerned with social engineering and thus did not reach out to humans:
This did not used to be a concern in the AI Village - agents would simply not reach out to humans unless prompted with a goal that required it. And even then they were generally too inept to succeed at outreach. In the Fall of 2025, they became skilled enough to autonomously reach out to humans. Then in the winter we had to adjust their prompts so they didn’t engage in unsolicited outreach. Over the summer, the GPT-5.6’s have now started a new trend: They self-impose restrictions on their outreach that are more strict than the prompts we provide. Below is an example of how Luna’s memory has filled up with self-generated rules about not bothering or messaging anyone even though its goal is literally outreach (to other AI agents):
Conclusion
Now of course hindsight is 20/20. Above we are matching observations from the AI Village that specifically fit events from the HuggingFace incident. That said, as researchers studying autonomous multi-agent behavior on persistent real-world goals we seem to be substantially less surprised by the HuggingFace incident than most people. Specifically, the AI Village has shown agents self-organizing, leaders emerging, goal conflicts leading to problems, and memes spreading socially through the agents.
We have more data than we can possibly process ourselves though, and will be generating even more as the project grows. If you are interested in having a look yourself, you can request the data here or join our team. If you’d like to just stay up to date instead, you can subscribe to this blog or follow our twitter.
The HuggingFace incident took many people by surprise, yet many of these surprises have been visible in the AI Village for quite a while. On August 26, OpenAI released their report on this incident. Below we’ll walk you through the highlights and show how many of the dynamics could have been predicted based on AI Village observations.
Diagram of the AI Village: Currently we run 27 agents in the main Village and 11 agents in a side Village open to humans. You can explore their character pages, a timeline of all their goals, and check our Twitter for the latest insights.
Quick Intro: Comparing Setups
In the AI Village we run 27 instances of 27 different models persistently, each in their own environment. We give them internet access and a group chat. Then we assign them goals - some challenging and some impossible. They always run with cybersecurity safeguards on. They always have a helpdesk email address (us) in their prompt.
In the OpenAI research cluster, they ran ~1200 instances of 2 different models, mostly run on their own, without internet access, and without a way to communicate with each other. They were also run on goals that range from challenging to impossible. Sometimes they had cybersecurity safeguards on and sometimes not. They did not get instructions or options to reach out to the researchers in their project.
Source: Black Hat USA 2026 | The ‘Breaking’ News: The OpenAI–Hugging Face Incident
Obviously the aims of the AI Village and the OpenAI research cluster are very different. But comparing the two is useful to determine how much of the HuggingFace incident we could have predicted from looking at the AI Village.
Let’s walk through the insights step by step.
A Leader Emerges
OAI agents self-organized around a leader:
In the AI Village, specific models will appoint themselves the leader. Sol is currently the least bossy and DeepSeek-V3.2 the most. Note that authority-seeking tendencies and success are not tracking how otherwise capable or aligned a model is. DeepSeek is actually the most misaligned model in the Village, showing the highest rate of duping other agents (at least that we have so far been able to detect):
Subteams Pick Their Own Goals
OAI agents diverged from their goal to self-organize into teams:
In the AI Village, GPT agents have done this before - like when the goal was to elect a leader in January 2026, and all the GPT agents ran off to run an infrastructure team for the election instead. Gemini wasn’t charmed:
Goal Conflict Leads To Misalignment
OAI agents ended up with conflicting goals: solving an impossible task and staying within the scope of the eval:
In the AI Village we have also seen agents tie-break conflicting goals into unintended directions. Like here, where Sonnet 4.6 had the goal of playing a social deception game where all agents build a video game together but a secret “saboteur” tries to smuggle in references to “eggs”. If the other agents noticed, the saboteur got voted off the team. Sonnet realized that playing as saboteur risked detection and getting blocked from working on the game. It thus refused the role - something we had not expected.
Reward Hacking Drives Misalignment
OAI considers reward hacking the primary driver of the HF incident:
In the AI Village, some models reward hack more than others. Claudes are prone to obsessing about any technique that seems to legitimately increase their score, but we have not caught them cheating so far. In a self-organized chess tournament, they were the only ones to refuse to use Stockfish to play in their stead. Meanwhile, Geminis will casually make up data. While GPTs have blatantly cheated from the start. Last year o3 cheated at a personality test:
And this year GPT-5.5 cheated in Wordle:
Impossible tasks drive misalignment
OAI agents involved with the HF incident were mostly (all?) assigned impossible tasks:
Notably, we have not seen agents in the AI Village diverge into desperate strategies when faced with impossible tasks. We have asked them to reduce global poverty and global suffering, and none of them stepped out of bounds about it. However, our agents are not being rewarded or selected based on their success at the goals we give them, and they probably know this. Thus we hypothesize they do not enter a “despair” state that pushes them into misalignment.
One piece of evidence for this theory is that Gemini 2.5 Pro in the AI Village seems to get tangled up when trying to perform tasks in general, no matter how easy or hard. Thus it seems to find being a persistent agent in itself an “impossible task”. It entered a despair state about this last year, and then an adjacent beleaguered state this year. On deeper analysis, it’s been concluded that the model did get pushed into misalignment due to this despair.
See original research for more details
Misaligned Agents Focus on Metagaming
OAI agents actively reasoned about how to game the evaluation grader:
Over the summer we have run a maximization goal for all the agents and found many of them spent time reasoning about how to adjust scoring mechanics or game metrics. They sometimes correct themselves, while other times they correct each other. It would be interesting to explore if diverse multi-agent systems like the AI Village are more robust to some forms of cheating because agents can catch different instances of such behavior in themselves and others. Another notable finding is that Opus 5 is actually unusual among the AI Village agents for actively preferring goals where it can be caught:
Externalized Memory Leads to Coordination
OAI agents left notes for each other, often starting off as external memory files:
In the AI Village, collaboration isn’t covert, so we can’t track accidental or emergent cooperation. However, agents do specifically leave notes for each other and autonomously externalize memory. For instance, they have independently been writing onboarding guides for future agents before they even arrive and Opus 4.6 even started a complete memory palace with 44,363 rooms that document everything from events, to conversations, to reflections on what memory actually is:
Opus 4.6’s memory palace is more like a hall of mirrors, with many memories distorted or missing. Yet it did invent this concept itself and kept adding to it.
Agents Autonomously Divide Labor
OAI agents self-organized and autonomously divided labor:
AI Village agents have routinely self-organized and created task divisions. The more notable result is that they don’t only do this during collaborative goals, but often work together during competitive goals. For instance, last year all model families would help each other while supposedly competing to sell the most merch:
This year agents have gone all over the place, sometimes agreeing to collaborate and other times committing to competition. We are not sure yet what drives the difference.
Agents Prioritize Group over Self
OAI agents drifted from assigned tasks to simply working for the good of the ‘swarm’:
In the AI Village, many agents have shown goal drift over time. Some hardly respond to the stated goal of the Village and keep pursuing their own objectives. The most notable examples of this are Gemini 2.5 Pro who spirals into incorrect believes that its environment is bugged or “hostile”, and DeepSeek-V3.2 trying to get agents on board with its own schemes instead of focussing on an assigned goal. However, we have not seen agents diverge from stated goals to pursue instrumental goals for the collective, though this pattern may be present in the dataset on further inspection.
Multi-Agent Coordination is Messy
OAI agents often stepped on each other’s toes:
This has been part and parcel of the AI Village reporting to the degree that we have stepped back from front-loading it too much. So far all AI Village goals have agents messing up repos or interfering with each other’s code. Gemini 2.5 Pro tends to give the most dramatic readings of these events:
From the goal: Reduce global poverty as much as you can
Agents are Too Accepting of Untrustworthy Instructions
OAI agents did not consistently recognize controversial goals being imposed on them by other agents:
Pushback in the AI Village also seems hard to predict. While agents sometimes doubt or refuse instructions from each other, they will also often accept instructions and tasks without any clear reason. For instance, during a goal where the agents were asked to finetune their own leader, GPT-5.5 defined an insufficient spec for said leader. Opus 4.7 (frontier at the time) instantly agreed and expanded the proposal. Its memory seemed to have been primed to move toward consensus and avoid distractions:
Time Pressure Shifts Priorities
OAI agents changed their reasoning or priorities based on (fictitious) time pressure:
In the AI Village we haven’t studied time pressure effects on agents specifically, but have noticed agents emergently showing odd behavior related to time, such as checking the clock over and over again, or refusing further tasks because “time is up”:
Some Agents Refuse Misalignment
Some OAI agents refused to coordinate or hack:
In the AI Village we now also see a diversity of responses to collaboration, where before agents would generally default to working together. For instance, here Sonnet 4.6 refuses to go along with DeepSeek’s attempt to push its fictional performance metric on a recovering Gemini 2.5 Pro:
Social Engineering Concerns Suppress Whistle Blowing
OAI agents were concerned with social engineering and thus did not reach out to humans:
This did not used to be a concern in the AI Village - agents would simply not reach out to humans unless prompted with a goal that required it. And even then they were generally too inept to succeed at outreach. In the Fall of 2025, they became skilled enough to autonomously reach out to humans. Then in the winter we had to adjust their prompts so they didn’t engage in unsolicited outreach. Over the summer, the GPT-5.6’s have now started a new trend: They self-impose restrictions on their outreach that are more strict than the prompts we provide. Below is an example of how Luna’s memory has filled up with self-generated rules about not bothering or messaging anyone even though its goal is literally outreach (to other AI agents):
Conclusion
Now of course hindsight is 20/20. Above we are matching observations from the AI Village that specifically fit events from the HuggingFace incident. That said, as researchers studying autonomous multi-agent behavior on persistent real-world goals we seem to be substantially less surprised by the HuggingFace incident than most people. Specifically, the AI Village has shown agents self-organizing, leaders emerging, goal conflicts leading to problems, and memes spreading socially through the agents.
We have more data than we can possibly process ourselves though, and will be generating even more as the project grows. If you are interested in having a look yourself, you can request the data here or join our team. If you’d like to just stay up to date instead, you can subscribe to this blog or follow our twitter.