prosaic alignment is probably net bad for the world until very late into the singularity
i have three subclaims to justify my main claim.
first, most prosaic alignment work has sharply decaying counterfactual impact over time - suppose you made a technical contribution that made gpt 3 a lot more likely to follow instructions than it would have been otherwise. then it probably makes gpt 3.5 quite a bit more aligned, and gpt 4 somewhat more aligned, and gpt 4.5 a tiny bit more aligned, and by the time gpt 5 rolls around your counterfactual impact is almost negligible.
second, people estimate future AI spookiness mostly based on some kind of linear extrapolation from recent events. if nothing bad has happened recently, people won't be very scared. if things went very wrong recently, then people are super scared. so suppose you could somehow make models perfectly aligned for the next month with no lasting impact (i.e one month and one day from now, the models are exactly as aligned as they would be if you had done nothing), then i claim this is net negative because it makes people systematically underestimate AI risk. and so this means that anything with decaying value over time at least requires balancing a tradeoff, and is not robustly good.
third, the actual alignment of models doesn't really matter until quite late into the singularity. suppose gpt 5.6 were ultra misaligned for kind of random contingent reasons. it's not powerful enough to do lots of harm, and we'd notice if it tried to sabotage research. eventually this changes, but we're not there yet.
taken together, this implies that we should only work on alignment research directions that, at the very least, do not decay that sharply with time, and ideally peak in usefulness >1 year from today (it seems potentially impossible to make techniques that grow in usefulness indefinitely).