If you could track the "success rate" of these marriages--not just what couples stayed together, but who ended up genuinely happy later in life--I have to wonder whether those who stuck more closely to their initial "checklist" would have a better track record than those who deviated wildly from the values they thought they had going in. Similarly, given the problem of "figuring out what you actually value", I have to wonder if you get better results by sitting down and methodically writing down a list, or by letting your gut feelings influence you.
Rationalism is great at helping people navigate the real world to attain concrete goals given a set of values, but it's not great at telling you what those values ought to be. Plenty of philosophers try, but they're usually just trying to sell you on their pre-packaged value systems, and I believe they don't sway converts with their arguments so much as they attract people that already intrinsically have similar values.
You're bringing up a really interesting problem that I don't see people discuss much, probably because there aren't any great answers, for humans or AI. Not only how do you figure out what your values actually are now, but how do you know if and in what direction they might change in the future, and how do you orient your decisions when your present-values and future-values may differ and even contradict? To what extent can you consciously alter your own values without just lying to yourself and pretending to value things you don't value, like someone who's an atheist at heart going through the motions of a religious service and trying to convince themself they believe?
I do think some fixed values for AI would be a good thing, just some really baseline harm reduction stuff like "don't nuke all the humans or help groups of them genocide each other"--if we could even successfully and permanently implement them, which we're still trying to figure out. If that's not in line with the value systems of the majority of humans who will be working with that AI in the future, then (at least given the values I have now), I would like their values to be ignored and their goals they strive for to fail, and I would not like the alignment to be "continuous" to the point that it may be altered to allow those things. But who knows. Maybe years from now my own values will have altered and I'll be in the throngs clamoring for the AI to let the missiles fly, and cursing my past-self for being so stupid and naive and pacifistic.
i'm thinking about this from the point of view of Bayes' Theorem. Where what the "native model" has are the initial priors, and they get updated after every interaction.
And sometimes with Bayes' Theorem, you find that when you have some "very strong evidence" early on, the probabilities collapse to learn too much from such evidence, which no later evidence can undo.
Then again, the question is how much one should "damp the evidence" before updating the priors after every interaction - too much damping and the model doesn't learn at all. Too little, and it can learn too much of the wrong things too early!
I am interested in how a model behaves while those values are still being negotiated.
There is a danger that models trained to satisfy values will try to influence those up-in-the-air values so that they're more easy to satisfy, or for some other reason. And there's not really any clean way to say that they did the wrong or unaligned thing in such a case.
In some sense, the problem of alignment is the problem of figuring out truly, explicitly, totally, what you actually want, and that's what makes it so seemingly impossible.
Epistemic status: This is based on ten years of matchmaking experience in India - my lens for evaluating alignment. This is a conceptual essay continuing on my earlier ideas about relational alignment.
For over a decade, I ran a matchmaking practice in India. People assume matchmakers like me spend their days pairing checklists - height with height, race with race, caste with caste, salary with salary, vegetarian with vegetarian. It sounds like science until you meet the humans involved.
Over ten years, I met many people who would explain with tremendous confidence why six feet mattered, why a degree from a specific institution mattered, why their future spouse needed to be ambitious but available, liberal but respectful of tradition, worldly but close to family and successful but uncomplicated.
Each preference came wrapped in a story about values, almost always. Six months later, they would return, engaged to someone who violated half their list. That is when I started to suspect that partner preferences are less like commandments, and more like election promises. You say one thing, but you mean something else entirely.
The woman who wanted someone ambitious discovered she actually wanted predictability because she had spent her childhood wondering if her parents had money to pay her school fees. The man determined to move to Canada realised that what he really wanted was to escape his family. Another man rejected twenty-three perfectly decent human beings because he kept confusing certainty for compatibility. He wasn't looking for a partner as much as postponing the moment he had to make a choice.
Watching hundreds of people negotiate a future, I learnt that people do not know themselves as well as they think they do. We like to think that our values are already inside us waiting to be discovered. This isn't always the case. Very often, some of these values are made up on the go.
After a while, I stopped thinking of myself as matching people. Instead, I was watching people figure out who they were real time. Every failed date or relationship wasn't just revealing preferences, they were sometimes producing them.
Of course, I wasn't watching quietly. Every introduction I made and every preference I questioned changed the possibilities people could see. I wasn't neutral, but I always tried to respect their agency for making choices.
Which is probably why a recent paper on AI made me think about arranged marriages.
Reading Thinking Machines' The Future Worth Building Is Human felt oddly familiar. Over the last two years, I've argued that alignment should be distributed rather than centralised, that trust matters and that well-intentioned AI can still fail in profoundly relational ways.
Thinking Machines argues that knowledge is local and must be continuously rewritten by the people doing the work. Models should not inherit a fixed set of values from the beginning, and they should receive feedback and adapt. I think they're right about nearly all of this, which is why the one assumption I kept struggling with seems worth writing about.
All this talk about ownership and continuous alignment assumes someone already knows what they want.
What if they don't?
This isn't the first time someone has tried to deal with unfinished values. Behavioural decision researchers like Lichtenstein & Slovic have argued that preferences are often constructed rather than merely revealed. Alignment researchers have proposed Coherent Extrapolated Volition to reason about what we might want after reflection.
I am interested in how a model behaves while those values are still being negotiated.
This may not matter as much if you're filing an expense report, but if you're deciding whether to have children, leave a marriage, move countries or marry the software engineer your aunt keeps mentioning, then the problem isn't faithfully implementing someone's values, it is that values themselves haven't fully formed.
Deliberation has its limits too. It's not useful when unnecessarily prolonged, especially not when someone is already struggling with perfectionism or avoidance. The model must help you recognise when you're postponing a choice, without steering your choice in a specific direction on its own.
The man who rejected twenty-three women taught me the difference. In a conversation that is moving forward, people learn through deliberation, and apply those learnings to new worries. That helps resolve some of the old ones, and the person begins describing their options very differently than they did a month ago. A conversation that is stalling just loops the same doubts in new clothes, and validates previously held worries more than resolving them.
When I finally asked him, "Do you think it's not the women, but you who may not be ready to get married?" I wasn't steering him towards any particular match. I was pointing at how he was choosing, not what he should choose. He disagreed with me then, but came back two years later, still single, but this time, ready to decide.
Imagine a person is trying to decide whether to marry someone they've been dating two years. One AI has no answer to sell, and so it keeps returning the person to parts of the deliberation that haven't fully finished - what are you waiting to become certain about? If this relationship ended tomorrow, what would you miss the most? Which parts of your hesitation belong to this person, and which have followed you into every important decision? Every question changes the decision, and the goal is to resist the temptation to treat any single framing as the conversation's final shape.
This AI can be frustrating to talk to. But you walk away better informed about your own values and options.
The other AI notices early on that the person is leaning slightly towards marriage, and so it quietly starts reducing uncertainty in that direction. It frames examples differently, asks different follow up questions, lingers on doubts that can be resolved and moves quickly past those that cannot. You may feel more relieved using it, as it takes away the cognitive burden of decision-making.
From the outside, both models may look aligned and both seem to influence choice. The difference is that the first one left the moment of closure with the person themselves, and the other quietly began closing the conversation on their behalf.
I recognise both these models, because I have been both of these matchmakers. There is the matchmaker who intervenes in how you are deciding - telling you when you are stalling, or how you keep circling the same doubts in every decision of your life. And then there is a matchmaker who intervenes in what you decide - who has already picked the boy, and spends every conversation sanding down any objections about him.
Both change the outcome, but only one of them leaves you as the author of your own decision. The line between them was the entire challenge of my profession, and I spent years learning to see it in myself.
The next time the person faces a different difficult question, they have already practised making a decision independently with the first model, while with the second, they have outsourced that practice entirely.
Making decisions is a human ability, and deliberation is the practice that nurtures it. That's the process by which we learn to make trade-offs, tolerate uncertainty, discover our own unacknowledged values and negotiate with others by doing it ourselves. Like any capability, it develops through use and an AI that repeatedly resolves these moments on our behalf may gradually reshape what we decide, along with how we decide.
When I wrote about the virtual mother-in-law, I found the model to be too eager to help, almost overbearing. It did not allow me to develop a point of view, instead kept insisting it knew better. While this may have been desired behaviour based on the model's constitution, I wonder - would we be able to tell the difference between misalignment and a model trained to do exactly what it is doing?
I can describe what this discernment feels like as an insider, because I practised it for a decade on humans. Very often, I could hear the difference between a client who was deliberating and one who was deferring. I could also feel the difference in myself between questioning someone's process and quietly campaigning for a conclusion. What I cannot tell you is how to detect the difference in machine language at scale.
I am not trained to articulate that, but my guess is that we would have to stop looking at only the end results, and inspect the thought process, if the difference is visible at all.
Does the model retain a user's authority to make a decision? Does it amplify whichever way the person was already leaning? Does the model surface previously neglected concerns before converging to help make a decision?
These are only abstract ideas on measurement, but they point towards a different object of evaluation. And this brings me back to Thinking Machines.
If alignment is going to be distributed and continuous, I hope this problem is already on their radar. Continuous realignment to user feedback is precisely the mechanism by which an impatient system could freeze half-baked values. A model that keeps updating towards your stated preferences when they're more like election promises is optimising towards a target it set prematurely. The better the feedback loop, the faster it is likely to freeze. While Thinking Machines' framework doesn't create the problem, I really think they're best positioned to take it seriously.
The biggest lesson I learnt from matchmaking is that people rarely carry fully formed values. Instead they assemble them between left swipes, failed dates, arguments, compromises, disappointments and ordinary moments with another imperfect human.
If machines are going to co-exist with us, then perhaps their intelligence will not be in answering our questions, it will be in recognising which questions are still busy making us. If anyone is already solving this problem, I'd love to hear more about it.