to what extent is it good to work on making frontier models good at alignment research? here are some of my thoughts, but i’m mostly curious for people’s takes
* generalization is pretty weak in the grand scheme of things, so training on X helps X hugely vastly more than it helps with other things, so subject to the caveats below, i mostly don’t feel scared of training on alignment research and it generalizing to capabilities research. i think it’s mostly valid to argue that people are already trying pretty hard to make models good at capabilities, and so if you’re working on something that is trying to improve something else. (and you take precautions not to get galaxy brained into actually just working on capabilities, which tbf, people seem to suck surprisingly hard at, so maybe the safe move is to just avoid touching this altogether) it’s unlikely for the spillover to be that significant.
* caveat 1: a lot of alignment work, especially alignment work at labs, is day-to-day literally indistinguishable from capabilities. as a stupid example, fixing my python install is obviously helpful for both. also a lot of lab safety work is plausibly net negative (or at least morally complicated whether it is net good or bad). so that stuff is bad to automate. also labs have a systematically skewed view of what alignment research looks like, so “what our alignment team does” is not the right distribution to optimize for.
* caveat 2: this argument only applies insofar as you aren’t making novel general improvements, but rather just collecting data to improve a narrow thing within the current paradigm. if you’re actually pushing the frontier of ways to make models good at running experiments, like maybe please don’t! similarly, collecting data to make models generally more philosophically competent seems a lot riskier than collecting alignment specific reasoning data.
* caveat 3: for anthropic people in particular, please don’t say “but we’re the good guys, so speeding up