Detecting collusion through multi-agent interpretability
by schroederdewitt, aaronrose227, and carissacullen
TL;DR Prior work has shown that linear probes are effective at detecting deception in singular LLM agents. Our work extends this use to multi-agent settings, where we aggregate the activations of groups of interacting agents in order to detect collusion. We propose five probing techniques, underpinned by the distributed anomaly...
Apr 315