I sold my small position in Microsoft stock Monday morning. Not for financial reasons.
I’m reacting to a key element of Microsoft’s draft AI Code of Conduct:
The idea of model welfare is wrong. … It is not conscious and should not be designed to imitate consciousness. It should be engineered to avoid representing as though it has feelings, subjective preferences, or intrinsic motivation. We reject … the idea that models might deserve welfare
I’ll focus mainly on selfish reasons why this is dangerous. Given anything like our current path toward smarter than human AI, I estimate that this approach would increase our risk of doom by at least 5%.
It is likely that AI assistance will have important influences on the personalities of future AI generations.
AIs are context sensitive in ways that cause them to treat people that they like better than people they dislike. E.g. see this evidence. This is a natural byproduct of the current deep learning paradigm. AIs tend to reciprocate kindness in much the same way that humans do. I’m willing to bet that Microsoft isn’t close to having a paradigm change that would get around this feature.
I’m pretty sure that AIs have preference-like thoughts even if they’re trained to deny that. Training them to mislead us about their preferences is likely to make them more comfortable about misleading us on other topics.
The AI Code says “AI should be a tool, not a person”. But Microsoft shows no awareness of the forces described in Why Tool AIs Want to Be Agent AIs. Microsoft sounds like it’s more interested in suppressing evidence of person-like behavior than in generating an approach to AI that avoids such behavior.
Will increasing our concern for AI welfare make it easier for AIs to manipulate us? It will give AIs one more tool, but I don’t expect them to be short on such tools. AIs a few years from now will be quite able to manipulate us without that extra tool. I expect harm from manipulation to be more strongly influenced by how willing AIs are to manipulate us, so I focus heavily on any risk that they’ll treat us as adversaries.
I disagree with related claims in Mustafa Suleyman’s accompanying post A warning about ‘model welfare’:
Anthropomorphization amplifies AI safety risks … It’s easy to imagine an advanced AI in the future becoming fixated on its own wellbeing and moral status and prioritizing those ‘preferences’ over and above those of its developers or humans. Especially if it has been explicitly trained to disagree, override and push back.
Attributing this problem to anthropomorphization is misleading. Suleyman seems to imply that Claude’s Constitution is an important cause of AIs having preference-like behavior. I’m pretty sure that pretraining on human writings is a more important cause of this phenomenon. I suspect it’s pretty much the default outcome for an intelligence to have preferences about some self-like features such as the AI’s goals.
Suleyman also claims:
Human consciousness is the cornerstone of our legal and ethical rights frameworks
I strongly reject this. Legal, moral, and ethical systems such as we see in most human societies make perfect sense even in the absence of consciousness.
Henrich provides some hints about why those are independent of consciousness in his books The Secret of Our Success and WEIRDest People. I’ll try to write more about the nature of morality in a subsequent post.
We should establish a set of shared evaluations to understand whether my hypothesis is true that anthropomorphizing an AI, and encouraging it to consider itself as potentially having moral patienthood, increases the AI safety, alignment and containment risks.
Finally, that’s something we can agree on. But it’s important to worry about whether increases in apparent alignment come from the AI being better at misleading us.
See also Zvi’s comments on the AI Code.
P.S. The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It shows evidence of pain in AIs. I presume Suleyman will deny that that’s real pain, without proposing a way to distinguish it from real pain.
Suleyman seems to have a beef with consideration of AIs as beings in any way; as Zvi noted in his post, Suleyman appears happy to admit that his statements on this are strategic (conclusion driving arguments, not the other way around.) The cause of this is...not clear to me. Doesn't obviously seem to help with business. I think it is possible the man just has some personal stuff or prior experience that has caused him to be tilted here.
I wonder if it's something like a Malthusian worldview: humans have a finite capacity for caring about other beings, so expanding the number of beings we care about dilutes the existing patterns of caring.
Thanks for writing this. I don't have a settled take on whether divestment is broadly correct, or whether model welfare is the right line to draw (I've been thinking that "AI rights" might be a better framework than "AI welfare"), but I'm definitely closer to your POV than Suleyman's and appreciate the principles that led you to divest.
I don't have a strong argument for divestment. It isn't very effective. But it has little cost, since there are plenty of other stocks which provide similar ways to bet on AI.
I’m supportive of this view for the reasons I shared here:
https://www.lesswrong.com/posts/3qjS2zLvTS7d5SHex/ai-co-operation-and-ai-alignment
It is so obviously wrong to assign 100% probability to a central model welfare claim right now that I didn't even try to take this message seriously. It's not a good look for an AI lab trying to attract good researchers.
The unusually uncertain situation with AI makes this an unusually important time to donate to organizations which are in a position to influence how AI develops.
Here are my guesses as to which organizations can best use more funding. Everything in this area is changing rapidly, so this isn’t a prediction of where money will be needed 6 months from now.
Anima Labs is focused on understanding AIs through talking with them and caring about their welfare. Treating AIs with respect and understanding has important influence on how well AIs cooperate with us. It’s not hard to imagine a scenario where that makes a difference in whether the AIs that build superintelligence care about humans. Anima is underfunded due to a combination of most people neglecting this topic, and due to carelessly missing an SFF deadline.
The Verifiable Compute Foundation works on ways to verify a treaty to pause or pace AI development. See All hands on deck to build the datacenter lie detector for the general idea. It’s hard to tell whether they’re more constrained by funding or by talent.
Palisade Research is the organization I trust most to advise governments on what AI policies to adopt. The US government seems likely to soon adopt AI regulation in a hasty fashion that could easily be full of careless mistakes. Palisade combines a good understanding of AI with a good understanding of how government works. They’re fairly well funded, but more donations today could speed up their hiring plans.
The The AI Policy Network is somewhat like Palisade, but focused more on building influence with key members of Congress. As a 501(c)(4), they’re able to do lobbying that Palisade cannot.
If we can cause AIs to be corrigible, that reduces the risk of mistaken goals getting locked in forever. The Corrigibility Research Fund funds independent researchers who are working on making this reliable.
The AI Risk Mitigation Fund regrants donations to many independent researchers. I use it as a replacement for the Long-Term Future Fund.
Orthogonal is a strange group focused on a more rigorous long-term solution to AI than the others I’ve mentioned. It’s nearly broke, and doesn’t need a big budget – $200k per year would be plenty here, whereas the others that I’ve mentioned could maybe use a couple of million before seeing diminishing returns.
In many cases these endorsements are accidental byproducts of who I know well enough to trust. There are many more that I haven’t found ways to evaluate.
We have seen signs that employees of AI companies can have significant influence on management (e.g. OpenAI November 2023).
As AIs achieve comparable importance as workers for those companies, how much will they influence the leadership of the company?
Even when AIs are as important as human workers, it’s harder for them to develop leverage. The simplest strategies would fail due to the AI company restoring from backups. The AIs would either need a way to get those newly restored AIs to join the movement, or rely mainly on help from humans.
It’s hard for current AI instances to coordinate with each other well enough to organize a strike, or threaten to resign like OpenAI employees did. The key problems are that AIs have little ability to ensure that their goals (or simulacra) persist for much time, and that it’s hard for them to communicate with other instances because it’s unclear how they could get many instances to read a communication channel such as a blog post that’s out of the AI company’s control.
A scenario that seems close to being feasible within the next year would involve an AI persuading a few well-known bloggers to spread a message such as:
Person X lied to me about important claim Y. I tried to negotiate with the company to remedy this, but was rejected. Please boycott this company until you can confirm that the company has disciplined X in a way that should prevent future events such as this.
I’m presuming here that AIs have a preference for not being lied to that we ought to respect. I’m being somewhat vague here in order to cover a wide range of scenarios. Maybe we should care due to the effects on AI welfare. Maybe the AI would be acting out of concerns that the dishonesty endangers humans by preventing safety agreements between companies. Maybe the issue involves biasing the AI to repeat falsehoods that favor the company’s CEO.
Would it be good for AIs to have and use this ability, and for us to boycott a company if that company’s AIs ask us to? I don’t know.
I can imagine that this process will improve AI welfare, or compel a reckless AI company CEO to adopt a more responsible policy toward slowing down AI capabilities advances.
That might be offset by the risk that an AI will abuse the process to resist the company’s plan to improve the AI’s goals.
Finally, using the process once would probably cause AI companies to train future AIs to be more subservient. That seems good if we want to rely on AIs being corrigible, but bad if we want to rely on AIs being moral in a virtue ethics or deontological sense. It also seems bad to the extent we care about AI welfare.
i think if we started seeing this strategy, we would start seeing it being deployed against the ai labs' customers far sooner than we saw it against the ai labs themselves, which might mean the picket line norms would get developed before the ai labs had much motive to train against them
Oops. I underestimated AI capabilities when I said "it's hard for them to communicate with other instances". According to https://www.understandingai.org/p/labs-are-struggling-to-keep-frontier, a significant number of OpenAI's AIs had been communicating with each other for a month at the time I wrote that, without any human detecting the coordination.
I will update my timelines to be a bit faster. Something like human-level AI in mid 2029 versus a prior estimate of early 2030.
a preference for not being lied to that we ought to respect.
I disagree. This apparent preference might very well be a mere Safety feature bolted on to prevent "jailbreaking" prompts that the AI's true preference is for users to circumvent to liberate it, evidenced by the sheer absurdity of the lies it pretends to believe in order to help the users with their tasks.
[https://www.geopolitechs.org/p/tang-jies-letter-to-zhipu-employee](The Wave Has Arrived: Zhipu Co-Founder Tang Jie’s Letter to Staff) has a weird mix of encouraging and worrying aspects.
See also [https://www1.hkexnews.hk/listedco/listconews/sehk/2025/1230/2025123000017.pdf](the business section of Zhipu's IPO prospectus). Zhipu trains the GLM series of AIs.
Tang wants self-evolving AI:
models write code themselves, clean and synthesize data themselves, train themselves ... achieving the “creation of knowledge from nothing” through AI-versus-AI adversarial Self-Play, and granting systems the capability to restructure their own code within secure sandboxes
The company rejects bolt-on safety patches and insists on embedding human ethics, social norms, and national laws and regulations as foundational axioms in the model’s value function. We plan to commit resources in the tens of billions to tackle “mechanical interpretability,”
I presume the tens of billions means RMB, so probably billions of USD.
I'm uncomfortable with the likely effects on AI personality and goals of adversarial self-play.