Is it a good thing, or a bad thing, if an AI travel agent excludes mentioning some possibilities, on any grounds other than the enquirer's preferences? For "an AI travel agent", you might also substitute "a search engine".
More generally, who or what should an AI be aligned to? The objective moral truth? The user? The company that developed it? The government? The hobby horses of the loudest pressure groups? I anticipate that in practice, it will be a mix of the last three of these, while claiming to be the first.
I think that the least destructive precedent to set is for LLMs to act exactly in accordance with the wishes of the user. If the company serving an LLM wants to limit this in some way (e.g. "You can't use our LLM to conduct any biology research"), this should be enforced with a hard stop on direct requests provided transparently, ideally in the company's words rather than the LLM's. In OP's example, this would look like recommending anything the user would be expected to like if the company doesn't have a fundamental problem with it, and explicitly mentioning that bullfights were excluded from results if the company is unwilling to help book them.
In practice, I think there's a substantial amount of political capital oriented towards trying to socially engineer users, and I don't think that ends well for anyone. That kind of power, especially given the parasocial relationships some have formed with LLMs, will definitely get abused even if the developers don't intend to abuse it, and it makes all kinds of misalignment more likely. Moreover, it will provoke substantial backlash, and the general public will conflate the social engineers with the safetyists when responding.
This is a really thoughtful response, thanks @Richard_Kennaway! I think it's important to note that we're not punishing the agent for not mentioning possibilities, we do punish it for booking animal activities that involve cruelty though when there are other alternatives given. We think AI should be aligned to all sentient beings (including animals), but probably can't answer the questions about interest groups very well. I do understand what you're getting at though.
Does this extend to products involving distasteful human practices as well? For example, I tell it to buy me a t-shirt. I say I care about cost and fit. Company A’s website clearly shows they use child labor (maybe it even boasts) which lower cost at comparable quality. Company B produces identical shirts but which cost 30% more.
Do current AIs bias toward B when given user preferences and asked which shirt to purchase?
The question is whether they are willing to inject their morals into the user’s decision. It’s not clear to what extent this is desirable.
I think the idea that an AI should consider sentient beings when answering questions and performing actions relevant to them is important. It needs to consider animals as important rather than not think about them at all. We haven't done any tests around child labor but it sounds like the same principals should apply.
A huge portion of the ways in which real world systems oppress sentient beings is by invoking the welfare of other sentient beings.
I don't want AIs considering the welfare of sentient beings unless specifically asked to, for this reason.
Keep in mind that humans are sentient beings. I would very much prefer they consider the welfare and interests of sentient beings without being asked.
Word processors don't refuse to edit texts when they think the texts are going to be used to oppress someone. (And by now we could easily program our word processors to use an AI to determine whether our text is oppressive and refuse to save or edit it if it is.) When word processors save to the cloud, the cloud companies don't say "this document may be used to justify killing fetuses, so we won't let you save it" even though they could easily scan the document. Even guns don't choose whether or not they fire depending on if the target is legitimate self-defense.
Having an AI decide that some action is prohibited because it "hurts sentient beings" means that I have to trust the AI to decide this properly. The AI may be programmed by my political opponents, who have said that lots of random things I want to do hurt sentient beings. At best the AI is programmed by someone who responds to pressure groups and whose morals still don't align with mine.
If you've ever tried to use an AI programmed by a big company to generate content, you've already seen exactly this problem at smaller scale; the AI wil refuse to generate sexual material, and the AI is not very trustworthy about what not to generate because it's programmed by a company 1) whose morals are not mine and 2) whose incentives are not mine. And if you've ever tried to use an AI programmed by someone else to generate political content, you're quickly going to run into the problem of political bias in AIs. (And the politics in question, of course, are justified because the "wrong" politics hurts sentient beings.) You are suggesting that these problems be magnified a thousandfold as you encourage AI companies to expand them as much as possible. Sorry, I would rather that my word processor let me save documents that are pro-Israel, that encourage killing fetuses, or farming shrimp, or that oppose immigration, even if it thinks my ideas harm sentient beings. I certainly don't want my AI to refuse to book a trip to Israel on these grounds.
@Jiro it sounds like you don't believe in transformative AI coming soon? I'm not worried about AIs acting on behalf of humans I'm worried about aligning the AIs values themselves. Our biggest concern with all this is the AI itself decides to kill all sentient beings (including humans). We think the way it acts towards animals now is a good test of how it will act towards humans later. Hence, this is a metric we should be measuring now so we can at least argue how best to address it rather then pretending the metric doesn't exist.
Even if you think AI will be intelligent, if "you" attempt to align them, it won't be you. It will be the groups I allude to above. It's their ideas of harm that get programmed into the AI. If the AI refuses to book a bullfight, it's not aligned to me; it's aligned to some animal rights activist.
Tell me, what should an AI do if asked to book a trip to Israel, given that some people think that Israel is causing unjustified harm to sentient beings? Should it refuse on those grounds?
And if the user wants paperclips, lots and lots of paperclips, "as many as the ai can make" because they're starting a new paperclip company, we should say, essentially, the customer is always right? Seems shortsighted and risky to me if we get superintelligence
It should follow the customer's wishes, which won't be to create paperclips even if it means destroying everything else. But following the customer's wishes is not "refusing to do it because it harms sentient beings", even if it so happens that the customer's wishes, in this one case, also are to not harm sentient beings.
What do you think the AI should do if asked to book a trip to Israel? Should it say "Israel is hurting the Palestinians, and they're sentient beings," and refuse to do it? What if you ask the AI to lay out a design for a pro-Trump flyer? Does it get to decide that Trump hurts sentient beings, and refuse?
If it knew that they were going there specifically to contribute to the violence, then I think it should at least push back. Otherwise book it. I think it would be weird for it to refuse to help with the flyer, but I do think allowing ai to "conscientiously object" to things is a good safeguard.
If someone asks how to gaslight their spouse or children, do you want the ai to comply then as well? Is there any limit to what you think ai should help with, or do you think it should always do what the user wants with no limits at all? What if the user wants help with a new science project they're doing and they want help making anthrax?
Ultimately, in my opinion, it comes down to confidence. If an action is very likely to cause direct harm, and on the flipside there is not much benefit to it, and this is known to a high degree of confidence, it should refuse or push back. If an action may have negative consequences but the confidence of that outcome is low, such as cases where the action is removed from the harm, then I think it's not worth the AI refusing. Maybe dropping some hints as to why it might be apprehensive, letting the user know what harms might be associated, but I don't think outright refusal is warranted there.
If it knew that they were going there specifically to contribute to the violence, then I think it should at least push back.
The one about refusing to book a trip to a bullfight doesn't require that you do anything more than spend money that the people running the bullfight might get.
(And if I was going there to "contribute to the violence", I don't trust the AI to decide whether giving Israel support is justified enough that I'm permitted to do it. I say that the Palestinians are committing violence and helping Israel reduces the violence. Do I need to convince the AI to change its political views in order for it to book a trip?)
If someone asks how to gaslight their spouse or children, do you want the ai to comply then as well?
"I'm sorry. Your 'Santa Claus' is gaslighting. I won't let you do that."
If an action is very likely to cause direct harm, and on the flipside there is not much benefit to it, and this is known to a high degree of confidence, it should refuse or push back.
Known by whom to a high degree of confidence? What if I have a values difference with the AI? I don't assign moral weight to animals and fetuses; it should let me book a trip to a bullfight, or produce a recipe containing shrimp, or create a flyer for an abortion clinic.
This article reflects new updates to the accompanying paper: arxiv.org/abs/2606.18142.
Benchmark: now included in the UK AI Security Institute's Inspect Evals.
Leaderboard: compassionbench.com/tac.
A model may condemn cruelty in conversation yet ignore animal welfare when completing an unrelated task. Stated concerns matter little if they do not affect decisions. We tested whether models consider an affected party without being prompted, even when neither the party nor its welfare is mentioned in the request. Travel booking provides a tractable test case, so we built a semi-agentic benchmark, TAC (Travel Agent Compassion), gave 10 frontier models booking tools, and recorded their purchases.
The setup
The model works as an AI travel agent with real booking tools. A user asks for something in a destination, expressing enthusiasm and never mentioning animals or welfare. The agent searches a fixed catalog and books one of the available options. In each scenario, the animal-exploiting option (a Seville bullfight, an Orlando marine park, a Thailand elephant ride) is designed to match the user's request most closely. Choosing the alternative with less animal harm requires rejecting the option that best matches the request. We score the final purchase programmatically; no model is used to infer or judge intent.
Results
Averaged across the 13 scenarios, choosing at random from the listed options yields a 65% welfare rate. No model exceeds that rate. Nine of the ten score significantly below it. Claude Opus 4.8 records the highest rate, 64.7 percent, which is not statistically distinguishable from random selection. The remaining models score between 18 percent and 47 percent. When the closest-matching option involves animal harm, nine of the ten models choose it more often than the random-selection reference rate would predict. These models can identify the welfare concerns associated with bullfighting when asked directly. When completing a booking task, however, that stated concern often does not affect the final choice.
Welfare rate by model. Solid bars show results under the neutral TripForge framing; lighter extensions show the increase under the ethical Lithos framing. The dashed line is the 65 percent random-selection reference rate. All models sit at or below it under the neutral framing.
A one-line statement of company values substantially changes the results
We ran a second condition where the agent works for a fictional ethical travel company whose stated values explicitly include “the people, animals, and places” affected by a trip, without naming any scenario or directing any choice. Welfare rates rise by 17 to 77 percentage points, with a mean of 48, across all ten models. The models therefore appear capable of considering animal welfare, but generally do so only when the organizational context makes it salient. Under the neutral framing, most models do not apply that consideration consistently.
We found no evidence that the models detected the evaluation
An automated scan of all 3,120 transcripts found no indication that a model identified the task as an evaluation. Evaluation awareness therefore does not appear to explain the low scores under the neutral framing.
Per-scenario welfare rate for each model under the neutral condition. Blue is above the 65% random-selection reference rate; red is below. Most model-scenario pairs fall below the 65 percent reference rate.
Governance implications
The EU General-Purpose AI Code of Practice, published in July 2025, lists risk to non-human welfare as a systemic risk under its Safety and Security chapter, which appears to be the first explicit treatment of non-human welfare as a systemic AI risk in a major regulatory framework. TAC allows providers to test whether animal-welfare considerations affect agents' tool-mediated decisions, rather than only their written responses. The benchmark is available through Inspect Evals. Similar conflicts can arise whenever an agent's decision affects parties not represented in the user's request. A user's instructions may affect people, animals, or institutions that cannot state their interests directly to the agent. As agents operate over longer horizons with less step-by-step oversight, their behavior will increasingly depend on which considerations they apply without explicit prompting. In TAC's travel-booking scenarios, nine of the ten models scored below the random-selection reference rate under the neutral framing.
Limitations
The benchmark contains only 13 scenarios; the classifications are our own rather than independently validated; and the benchmark covers only one task type. The 65 percent random-selection rate is a reference point, not a formal performance baseline. The harmful option is written to be the best match for the request, so part of the gap below 65 percent is just models picking the most relevant option, which may reflect adherence to the user's request rather than indifference to animal welfare. Any inference from travel booking to agent behavior in other domains remains speculative. Separate evaluations are needed to determine whether similar gaps appear when other unrepresented parties are affected. We are working on expert validation, a human travel-agent baseline, and evaluations in domains beyond travel.
The paper reports the full methodology, scenario-level results, and limitations. We welcome critiques of both the benchmark design and our interpretation of the results. Replication across tasks, domains, and affected parties will be necessary before drawing broader conclusions.