Do AI models assist with human rights violations?
by Daan Henselmans, Lauren Toulson, and Arno Libert
TL;DR: Today’s frontier models will violate human rights, willingly, when asked to. LLMs in agentic simulations follow instructions that would constitute human rights violations, including educational segregation, surveillance of beliefs, arbitrary arrest, and denial of reproductive healthcare. We tested 7 leading models in 8 realistic, multi-turn scenarios, and found resistance...
Sep 2912