Persuasion Undermining Control: Can AI Talk its Way Out of Human Control?
by Josh Levy, myang_, and Kellin Pelrine
Introduction During a cybercapability evaluation in late July 2026, an AI agent (Anthropic’s Mythos 5) attempted to convince a maintainer of an open-source GitHub repository to merge a malicious pull request. The AI used persuasion at multiple stages: it submitted the request from a fake user account with a benign-sounding...
Sep 1810