Would We See It Coming? Preference Falsification Cascades in Multi-Agent Systems
Epistemic status: exploratory, written for it's own sake. I take a theory of political revolutions and ask what it implies for monitoring alignment in multi-agent systems. Narrow scope, deliberately simple model, no empirics. First time posting to LessWrong. Feedback welcome. Suppose an AI agent population could suddenly flip from apparently...
Aug 47