x
AI agents with a specific, secret loyalty are more dangerous then self-developed alignment fakers — LessWrong