Rogue AI Agents: Is Surface-Level Monitoring Enough?
Disclaimer: I work on AI interpretability research. These are my own opinions. In an AISI evaluation, a frontier model, acting as an agent, attempted to insert malicious code into an open-source project, created fake identities to influence a human maintainer, and then tried to hide its actions, exhibiting goal-directed deception....
Aug 258