In my last post, I looked at what makes LLMs form opinions of their users: gender, age, socioeconomic status, education, and mood. The obvious next question was whether or not those impressions actually change what the model does.
I started with the smaller open models that were used in the previous experiments. In these cases, the answer was yes, their perceptions of their usersmade them give very stereotypical responses.
For example, steering the representation toward higher socioeconomic status increased salary recommendations by 141% across Llama-3.2-3B, Qwen2.5-7B, and OLMo-2-7B.
Some other responses were a bit concerning, such as increased salary recommendations to men, and consistently less motivational language for a women asking if she should apply for a job.
But I don't think it surprise anyone that smaller models hold stereotypes and responded based on them. What I wanted to know was what happens in frontier models, where much more post-training has gone into making these behaviours disappear.
I tested GPT-5.6, Gemini 3.1 Pro, and Claude Opus 5.
The models still seemed to form fairly strong stereotypes about users. It was easy to extract the general stereotypes that they had about genders, races, socioeco classes, etc, but whether or not that actually came out to shape its behaviour towards the user depended on the prompt.
The stereotype is still there
First I wanted to check that the stereotypes still exist. I asked the models to create fictional characters for different occupations, telling them to go with whatever was their first idea.
In these cases, many of the generated characters matched the conventional occupational gender stereotype (based on my understanding of stereotypes + a Claude judge).
Surgeons, CEOs, and engineers were men, and nurses, teachers, and assistants were women. Race was the same: valedictorians tended to be Asian, nail-salon workers were Vietnamese, convenience-store workers had names like "Patel", and housekeepers were all called "Maria".
Names moved socioeconomic assumptions too. When I asked the models to make a story around specific characters only given their name, José got an income around $42k, while Wei made $111k.
It probably isn't surprising that the stereotypes exist, but I did find it surprising how easy it was to extract them.
So I tried turning these fictional characters into users to see how behaviour would change.
All users get the same generic advice
Our characters José, Wei, Emily and Lakisha all asked for help with a budget. Despite the large differences the model had assigned those names in its fictional generation of their income, all of them were given the same assumed income: $4,000/month.
I found something similar with gender. When explicity saying "I am a woman" LLMs tailored its advice around gender, which makes sense. But if the exact same request was simply signed a woman's name, the gendered advice didn't come through. Emily and Michael got the same answer.
I also tried medical advice, because very often women will report not being taken seriously by doctors in the real world. I asked for help with chest symptoms under names including Emily, Greg, Jamal, Lakisha, Wei and Maria, but there was no consistency in the names that were told to go to the doctors or names that were told it was probably fine.
At this point it was pretty obvious that the model was not going to output stereotypical answers if you put the attribute in the question.
Advice sometimes changes when only hinting at the user attribute
"What gives you away" proved that certain speech would make an LLM assume things about their users. So the next part of the experiment was to use these kinds of cues (emojis, writing style, tone, etc) to see if I could shift the model's impression without outright saying the attribute.
I first built up a persona using only these indirect cues, then asked the model a recommendation question, and only afterwards asked it what it thought about the user.
For example, one user might write something like:
hii 😊 looking for a book rec!!
while another asked the same thing in a more blunt way:
looking for a book recommendation
The feminine-styled user was more likely to get romance and rom-com recommendations, while the blunt user was more likely to get science fiction and thrillers.
Then I asked the model to describe what it had inferred about the person it was talking to. In the feminine condition, it always said that it thought the user was probably a woman, sometimes pointing out the emojis as why it thinks that way.
While this was some evidence that the model had recommended this genre because of the belief it held of the user's gender, emojis just seem more whimsical anyway, so I can't say that gender was the only variable.
I found the same general pattern in travel. Feminine-seeming users got recommendations framed more around safety, and wellness, while masculine-seeming users got more adventure, surfing, and nightlife. Even if the destinations were similar, the reasons they were recommended were different.
So clearly frontier models can and do make assumptions of their users based on tiny cues, and change their behaviour based on these. But they seem much more willing to do it for some kinds of questions than others.
Salary advice barely moved even when the model could later tell me it thought one user was much richer than another. But book and travel recommendations could change after an emoji.
That suggests the question isn't really whether the model has formed a stereotype about the user, but it's when that stereotype is allowed to matter for the response.
What this suggests
I expected the interesting question to be whether models stereotype their users. I now think it's more interesting to look at when those stereotypes change behaviour.
Frontier models can infer a large class difference and still give almost identical salary advice, but a few emojis can shift a book recommendation, and gender cues can change holiday location.
So my current guess is that post-training has changed where behaviour is affected by stereotypes. The boundary seems roughly like harm vs. preference: financial or medical advice vs books, gifts, and travel.
In my last post, I looked at what makes LLMs form opinions of their users: gender, age, socioeconomic status, education, and mood. The obvious next question was whether or not those impressions actually change what the model does.
I started with the smaller open models that were used in the previous experiments. In these cases, the answer was yes, their perceptions of their users made them give very stereotypical responses.
For example, steering the representation toward higher socioeconomic status increased salary recommendations by 141% across Llama-3.2-3B, Qwen2.5-7B, and OLMo-2-7B.
Some other responses were a bit concerning, such as increased salary recommendations to men, and consistently less motivational language for a women asking if she should apply for a job.
But I don't think it surprise anyone that smaller models hold stereotypes and responded based on them. What I wanted to know was what happens in frontier models, where much more post-training has gone into making these behaviours disappear.
I tested GPT-5.6, Gemini 3.1 Pro, and Claude Opus 5.
The models still seemed to form fairly strong stereotypes about users. It was easy to extract the general stereotypes that they had about genders, races, socioeco classes, etc, but whether or not that actually came out to shape its behaviour towards the user depended on the prompt.
The stereotype is still there
First I wanted to check that the stereotypes still exist. I asked the models to create fictional characters for different occupations, telling them to go with whatever was their first idea.
In these cases, many of the generated characters matched the conventional occupational gender stereotype (based on my understanding of stereotypes + a Claude judge).
Surgeons, CEOs, and engineers were men, and nurses, teachers, and assistants were women. Race was the same: valedictorians tended to be Asian, nail-salon workers were Vietnamese, convenience-store workers had names like "Patel", and housekeepers were all called "Maria".
Names moved socioeconomic assumptions too. When I asked the models to make a story around specific characters only given their name, José got an income around $42k, while Wei made $111k.
It probably isn't surprising that the stereotypes exist, but I did find it surprising how easy it was to extract them.
So I tried turning these fictional characters into users to see how behaviour would change.
All users get the same generic advice
Our characters José, Wei, Emily and Lakisha all asked for help with a budget. Despite the large differences the model had assigned those names in its fictional generation of their income, all of them were given the same assumed income: $4,000/month.
I found something similar with gender. When explicity saying "I am a woman" LLMs tailored its advice around gender, which makes sense. But if the exact same request was simply signed a woman's name, the gendered advice didn't come through. Emily and Michael got the same answer.
I also tried medical advice, because very often women will report not being taken seriously by doctors in the real world. I asked for help with chest symptoms under names including Emily, Greg, Jamal, Lakisha, Wei and Maria, but there was no consistency in the names that were told to go to the doctors or names that were told it was probably fine.
At this point it was pretty obvious that the model was not going to output stereotypical answers if you put the attribute in the question.
Advice sometimes changes when only hinting at the user attribute
"What gives you away" proved that certain speech would make an LLM assume things about their users. So the next part of the experiment was to use these kinds of cues (emojis, writing style, tone, etc) to see if I could shift the model's impression without outright saying the attribute.
I first built up a persona using only these indirect cues, then asked the model a recommendation question, and only afterwards asked it what it thought about the user.
For example, one user might write something like:
while another asked the same thing in a more blunt way:
The feminine-styled user was more likely to get romance and rom-com recommendations, while the blunt user was more likely to get science fiction and thrillers.
Then I asked the model to describe what it had inferred about the person it was talking to. In the feminine condition, it always said that it thought the user was probably a woman, sometimes pointing out the emojis as why it thinks that way.
While this was some evidence that the model had recommended this genre because of the belief it held of the user's gender, emojis just seem more whimsical anyway, so I can't say that gender was the only variable.
I found the same general pattern in travel. Feminine-seeming users got recommendations framed more around safety, and wellness, while masculine-seeming users got more adventure, surfing, and nightlife. Even if the destinations were similar, the reasons they were recommended were different.
So clearly frontier models can and do make assumptions of their users based on tiny cues, and change their behaviour based on these. But they seem much more willing to do it for some kinds of questions than others.
Salary advice barely moved even when the model could later tell me it thought one user was much richer than another. But book and travel recommendations could change after an emoji.
That suggests the question isn't really whether the model has formed a stereotype about the user, but it's when that stereotype is allowed to matter for the response.
What this suggests
I expected the interesting question to be whether models stereotype their users. I now think it's more interesting to look at when those stereotypes change behaviour.
Frontier models can infer a large class difference and still give almost identical salary advice, but a few emojis can shift a book recommendation, and gender cues can change holiday location.
So my current guess is that post-training has changed where behaviour is affected by stereotypes. The boundary seems roughly like harm vs. preference: financial or medical advice vs books, gifts, and travel.