One of your counterarguments is that "the purpose of an AI pause is to do safety work".
A couple points, one that applies even if you think the only acceptable reason to pause AI is to avoid human extinction, the latter that doesn't apply if you believe that:
So, for me with regards to the second top-level bullet-point, "safety"/"alignment" are necessary but far from sufficient conditions to address my objections - the objections that are the reason I have made enormous changes to my life in the past six months to dedicate myself to effectively advocating for a Pause. To me and many others, the purpose of a Pause is to determine what on earth we want society to look like going forward and figure out how to address the new challenges posed by the seemingly-impending appearance of ASI, not just for technical researchers to solve some problems and then unleash a completely unprecedented future world-state on all of humanity absent our collective decision-making.
I've recently heard a surge of support for an old argument: AI safety researchers should not work at frontier AI companies because this reduces the likelihood of non-lethal warning shots, and we need warning shots to build support for an AI pause/slow-down. This argument has several components:
I think this argument has some merit. I expect that RLHF++ will probably be insufficient to align TED-AI circa-2028 and such systems will be terrifically difficult to monitor or control. However, I think there are also significant weaknesses to this argument. Here are some countervailing points to consider:
Overall, I am cautiously optimistic about working on certain types of safety research at frontier AI companies, particularly if a coordinated AI slowdown (e.g., Plan A) occurs. I nevertheless feel highly uncertain about the "Alignment MVP" strategy in light of recent containment failures and alignment results. I think this merits serious consideration.
Disclosure: I'm the CEO of MATS, which trains AI safety researchers, including with mentors at frontier AI companies, and a meaningful fraction of our alumni end up working at those companies. If the argument I'm responding to is right, a good chunk of what MATS has done might be counterproductive, so I have an obvious incentive to find it wrong. I've tried to steelman it anyway and I'd appreciate feedback.