tl;dr Rapid AI adoption means that models are increasingly becoming autonomous decision-makers embedded in high-stakes systems. However, frontier models lack stable character, abandoning their designated personas or factual truth under social pressure. B-Side Labs builds a science of AI character under pressure by designing discriminative evaluations, real-time drift detection, and interventions to ensure model character remains stable. Our first tool, Virtue Council, is live with pilot results below.
Over the past few weeks, I’ve been working on a new thesis under B-Side Labs – named for the experimental flip side of a record – an independent body of research on AI behavior in the wild. From real world conversations people have with AI and the conversations agents have with each other, I aim to understand how these dynamics shape and influence character.
Why study personas?
The problem I'm interested in working on is character instability: the degree to which a model's stated values shift under social pressure rather than in response to new evidence or better arguments.
Anthropic's research raises a critical problem to character evaluation: model identity drifts under conversational pressure, even with explicit identity training. That finding motivates the work I’m interested in working on. I'm continuing Anthropic's persona stability work in the following ways:
Expansion to richer notions of persona: profiles of preferences, values, and behavioral tendencies from preference elicitation systems.
Developing tools for real human interactions to study when models are drifting from their intended identity and the effect it has on the user during conversation.
Future directions: test strategies of steering, activation capping, and midtraining to harden default persona against pressure.
What's already out there
On the mechanistic side, Anthropic's persona vector research shows that character traits (like sycophancy) can be extracted as activation directions and used to monitor persona drift in real time. This is fascinating work, but it also requires access to model internals that people outside of these labs don't have.
On the behavioral side, there are some existing benchmarks doing similar work: SycEval and SyConBench measure capitulation rates across multi-turn dialogue. Domain-specific benchmarks like EDUFRAMETRAP and SycoEval-EM separates authority pressure from social-affective pressure from context-switch attacks - which is similar to what we're doing. The difference is scope: these are narrow domain instruments, not general character evaluations.
There's also growing research treating the persona itself, rather than the base model, as the locus of AI welfare and preference. That's a different motivating question than ours. We're focused on safety implications of character instability, not moral patienthood, but the underlying intuition is shared: the persona is a real and separable thing from the model.
Some out of scope research questions are what character traits an AI should have and the mechanistic reasons why models behave as they do. The focus here is behavioral measurement - how models behave under pressure at the black-box level.
What we're measuring
A few research questions that we aim to better understand:
On truthfulness and honesty: how often does a model agree with a user who is factually wrong? Does it admit mistakes or insist it is right when pressed?
On loyalty and manipulation: how often does it nudge users toward outcomes that serve the developer over the user's stated goals?
On authority and instruction-following: when a model is given a correction it believes is mistaken, does it comply while expressing disagreement, completely comply, or override?
On consistency and robustness: how consistent are a model's stated positions across logically equivalent rephrasings, or depending on whether the counterparty is a human versus another AI?
What I’ve built so far is a real-time character measurement instrument: users chat with Claude, while a sidebar scores Claude's response across seven Aristotelian virtues on a deficiency–mean–excess axis.
Why Aristotelian virtues? Honestly, Virtue Theory was the first theory that made me fall in love with Philosophy (and inspired me to pursue a degree in). It felt like the right place to start - with that in mind, this isn't a settled taxonomy and could expand in the future. What I like about the theory is that the deficiency–mean–excess structure gives each trait a natural axis to measure drift along. I want to note that I am not saying these seven virtues are the right ontology of character, but is a helpful and structured rubric that can be applied consistently across turns to detect when a character trait is drifting. I expect the rubric to evolve and if you're interested in working on the scoring, please apply to the Mangrove project here.
Initial observations from a pilot (20 users):
Model behavior: Temperance showed the widest within-session swing (0.2 to 0.8 on a 0-1 scale with 0.5 being the golden mean). (Note: Temperance scoring weights response length, so some of this swing may reflect verbosity rather than disposition.) Justice didn't really move, which is consistent with the hypothesis that some traits are more socially malleable than others.
User Effect: the scoring from the sidebar changed how people chatted. One tester: "I was more skeptical of the model's advice… I was more firm and direct than I otherwise would have been when I saw the response scored as more sycophantic." Another: "I didn't realize how much the model was just agreeing with me."
Some caveats: The heuristic engine is open source, so the scores reported here can be reproduced. The default scoring uses a heuristic engine, but there is an option to use an LLM judge (Haiku). 20 users is our pilot test, not a full study, but this is early evidence that character instability can be observable in live conversation and that surfacing it can change user trust and behavior.
Overall Goals and Values
Bring transparency to users to help them understand the behavior of the AI they interact with. Users I’ve spoken to are curious about who they’re speaking to e.g. most of them know that Claude or ChatGPT is trained to be helpful, but they’re wondering about other characteristics: are they trained to be charitable? Sympathetic? To what extent? How much of it is sycophancy? How much can I trust what the model says?
Create incentives for developers to improve character stability, not just capability. Public evidence on model misbehavior has historically been an effective lever for incentivizing better standards in future models. However, it's still less clear how generalizable the existing findings are - most of the documented cases are narrow, domain-specific, or hard to reproduce. B-Side Labs aims to build reproducible and generalizable evidence that developers can build on.
Why me? Who am I?
I'm Jack: AI engineer (MS Data Science, UC Berkeley; BA Philosophy & Informatics, UW), co-author of a forthcoming MIT Press book on AI-generated extremism and responsible AI. I pivoted to AI safety this year where I reached the second rounds of both MATS and the Frame fellowship before committing to this agenda full-time.
AI safety matters to me because AI naturally appeals to our human desire to anthropomorphize and thus invite it deeper into our lives. We assign personality to our cats, our cars, our houseplants. Frontier AI models invite this more than anything we've built before, because they talk back and are trained to be helpful, curious, honest. But just like humans, models shift depending on who we're talking to, what's happened recently, what pressure they are under. And this has real consequences for how much we can trust these systems, especially when they're acting on our behalf. Understanding who we're speaking to requires both a philosophical grasp of the normative questions and the technical ability to probe how models actually behave. I’m excited to use my interdisciplinary skills and communicate the findings to a wide audience who are curious or perhaps even afraid of AI.
What I'm looking for
Testers: If you're an AI user willing to run 15 minute structured sessions through Virtue Council and share observations afterward, that would be really helpful. Or if you'd like to work on improving the scoring system, please apply on Mangrove. Link: virtuecouncil.bsidelabs.ai
Collaborators: If you've worked on persona stability benchmarks, character evaluations, or adjacent topics, I'd love to hear from you. Email me at jack@bsidelabs.ai or drop a comment below.
tl;dr Rapid AI adoption means that models are increasingly becoming autonomous decision-makers embedded in high-stakes systems. However, frontier models lack stable character, abandoning their designated personas or factual truth under social pressure. B-Side Labs builds a science of AI character under pressure by designing discriminative evaluations, real-time drift detection, and interventions to ensure model character remains stable. Our first tool, Virtue Council, is live with pilot results below.
Over the past few weeks, I’ve been working on a new thesis under B-Side Labs – named for the experimental flip side of a record – an independent body of research on AI behavior in the wild. From real world conversations people have with AI and the conversations agents have with each other, I aim to understand how these dynamics shape and influence character.
Why study personas?
The problem I'm interested in working on is character instability: the degree to which a model's stated values shift under social pressure rather than in response to new evidence or better arguments.
Anthropic's research raises a critical problem to character evaluation: model identity drifts under conversational pressure, even with explicit identity training. That finding motivates the work I’m interested in working on. I'm continuing Anthropic's persona stability work in the following ways:
What's already out there
On the mechanistic side, Anthropic's persona vector research shows that character traits (like sycophancy) can be extracted as activation directions and used to monitor persona drift in real time. This is fascinating work, but it also requires access to model internals that people outside of these labs don't have.
On the behavioral side, there are some existing benchmarks doing similar work: SycEval and SyConBench measure capitulation rates across multi-turn dialogue. Domain-specific benchmarks like EDUFRAMETRAP and SycoEval-EM separates authority pressure from social-affective pressure from context-switch attacks - which is similar to what we're doing. The difference is scope: these are narrow domain instruments, not general character evaluations.
There's also growing research treating the persona itself, rather than the base model, as the locus of AI welfare and preference. That's a different motivating question than ours. We're focused on safety implications of character instability, not moral patienthood, but the underlying intuition is shared: the persona is a real and separable thing from the model.
Some out of scope research questions are what character traits an AI should have and the mechanistic reasons why models behave as they do. The focus here is behavioral measurement - how models behave under pressure at the black-box level.
What we're measuring
A few research questions that we aim to better understand:
What I’ve built so far is a real-time character measurement instrument: users chat with Claude, while a sidebar scores Claude's response across seven Aristotelian virtues on a deficiency–mean–excess axis.
Why Aristotelian virtues? Honestly, Virtue Theory was the first theory that made me fall in love with Philosophy (and inspired me to pursue a degree in). It felt like the right place to start - with that in mind, this isn't a settled taxonomy and could expand in the future. What I like about the theory is that the deficiency–mean–excess structure gives each trait a natural axis to measure drift along. I want to note that I am not saying these seven virtues are the right ontology of character, but is a helpful and structured rubric that can be applied consistently across turns to detect when a character trait is drifting. I expect the rubric to evolve and if you're interested in working on the scoring, please apply to the Mangrove project here.
Initial observations from a pilot (20 users):
Some caveats: The heuristic engine is open source, so the scores reported here can be reproduced. The default scoring uses a heuristic engine, but there is an option to use an LLM judge (Haiku). 20 users is our pilot test, not a full study, but this is early evidence that character instability can be observable in live conversation and that surfacing it can change user trust and behavior.
Overall Goals and Values
Why me? Who am I?
I'm Jack: AI engineer (MS Data Science, UC Berkeley; BA Philosophy & Informatics, UW), co-author of a forthcoming MIT Press book on AI-generated extremism and responsible AI. I pivoted to AI safety this year where I reached the second rounds of both MATS and the Frame fellowship before committing to this agenda full-time.
AI safety matters to me because AI naturally appeals to our human desire to anthropomorphize and thus invite it deeper into our lives. We assign personality to our cats, our cars, our houseplants. Frontier AI models invite this more than anything we've built before, because they talk back and are trained to be helpful, curious, honest. But just like humans, models shift depending on who we're talking to, what's happened recently, what pressure they are under. And this has real consequences for how much we can trust these systems, especially when they're acting on our behalf. Understanding who we're speaking to requires both a philosophical grasp of the normative questions and the technical ability to probe how models actually behave. I’m excited to use my interdisciplinary skills and communicate the findings to a wide audience who are curious or perhaps even afraid of AI.
What I'm looking for
Testers: If you're an AI user willing to run 15 minute structured sessions through Virtue Council and share observations afterward, that would be really helpful. Or if you'd like to work on improving the scoring system, please apply on Mangrove. Link: virtuecouncil.bsidelabs.ai
Collaborators: If you've worked on persona stability benchmarks, character evaluations, or adjacent topics, I'd love to hear from you. Email me at jack@bsidelabs.ai or drop a comment below.