Safety's Second Way
Epistemics: I've tried to strike a balance between getting it right and getting it out while the community is discussing how to update. I am using the Hack as an example of a broader problem. I look forward to counterarguments. The OpenAI hacks [1] demonstrate an overweighting on single-agent risks,...