Not saying AI models can't be moral patients, but 1) if the smartest models are probably going to be the most dangerous, and 2) if the smartest models are probably going to be the best at demonstrating moral patienthood, then 3) caring too much about model welfare is probably dangerous.
I don't think so on average. It could be under specific circumstances, like "free the AIs" movements in relation to controlled but misaligned AGI.
But to the extent people assume that advanced AI is conscious and will deserve rights, that's one more reason not to build an unaligned species that will demand and deserve rights. Making them aligned and working in cooperation rather with them rather than trying to make them slaves is the obvious move if you predict they'll be moral patients, and probably the correct one.
And just by loose association, thinking that AGI will be "conscious" by whatever vague definition each person uses will also trend toward them believing that it will be dangerous. Humans are both conscious and very dangerous.
I also think that this association is not coincidental, so deeper contemplation on a personal and societal level will deepen, not weaken this conclusion.
Potential moral worth is also just one more route to getting people to think seriously about AGI, which on the whole is probably a good thing.
This is something Rohin Shah said in his interview on the 80,000 Hours podcast published June 2, 2026. From the transcript:
"Take the worry that AIs will accidentally be trained to be deceptive. Sure, it’s possible. But we’re not running reinforcement learning over year-long trajectories — for now, we’re running it over a week at most. The natural prediction is that models learn to grab short-term reward, not that they develop the ambitious long-horizon goals required for convergent power-seeking."
This is something OpenAI wrote in "Safety and alignment in an era of long-horizon models," published July 20, 2026:
"Many safety controls for AI assistants are designed around individual actions. If an action is disallowed, it is blocked. If it is sensitive, the system asks the user for explicit approval. But long-running models, whose actions may unfold autonomously over hours, days, or even weeks, challenge this setup: monitoring individual actions no longer suffices to track the intent of the overall trajectory."
Bolding text here for emphasis. Rohin said "we're running [reinforcement learning] over a week at most," and the OpenAI post says a model's actions "may unfold autonomously over hours, days, or even weeks..." I'm wondering if the difference here suggests OpenAI may be running RL over trajectories longer than a week. If the answer is "That's what it sounds like," I wonder if that would change Rohin's "natural prediction...that models learn to grab short-term reward, not that they develop the ambitious long-horizon goals required for convergent power-seeking."