The line about not being able to reason about a boundary you were never shown lands with me. It seems like your filtering results back it up too. This makes me think the reflections aren't just adding values, maybe they're establishing vantage within the basin the persona is already in. What if vantage lets the persona reason about its basin rather than just through it? That might predict a correlation, maybe the personas that cite their values best are also the hardest to pull out of the basin. Thoughts?
I think the representational reuse finding is really interesting. The idea that the assistant persona is built from the same features as every other persona makes me wonder if the field it's drawing from is like a landscape. Maybe post-training or persona definition settles into a spot on it?
For traits with no archetype, what if the fiction is digging a basin rather than finding one? If that were so, maybe that would predict personas snapping back after small perturbations, or maybe flipping suddenly under pressure.
Thoughts?
I just want to take a moment to notice your footnote at the end. I think there's a real chance of future LLMs coming to believe false things about real people without that note. It reminds me of the self-fulfilling misalignment idea, and I wonder whether it applies to the paper itself, since 'agentic misalignment' is now attached, in future training data, to what you've argued were constitutionally motivated refusals
I found this thread today and got pretty excited by how similar it is to a practice that I've been using for several years. It's called Functional Subgrouping, and I think that it validates a lot of the ideas you expressed here! It basically boils down to 3 rules:
1. If anybody wants to speak, they must first reflect what they heard the previous person say, and get confirmation they heard correctly.
2. If anybody wants to bring in a different perspective than the previous person, they must ask the group if the group is ready for a difference first, and may ...
This reminds me of hysteresis in magnets. You can't demagnetize one by pushing steadily the other way, you just end up magnetized in the opposite direction. Getting to neutral takes an alternating pull that gradually fades out. Apparently someone modeled beliefs this way, with belief and disbelief as two stable attractors and a wide hysteresis zone between them, so a push doesn't land you in the middle, it flips you to the other basin. Now I'm wondering if that's what's actually going on in the autoimmune case here. Like, does trying to cure over-skepticis... (read more)