something I failed to consider on the compute overhang:
~3 years ago, if asked if it was good that frontier labs were pushing out models at their then-current pace I would have said :
it seems good that we are maxing out our compute.
i now think a blind spot of mine was how much frontier labs’ demand for compute accelerated compute supply, therefore shortening timelines.
To recount the compute overhang argument:
if we were to pause / proceed slowly, compute hardware would still improve, so we would get a spiky increase in capabilities. And without intermediate models to study, this would have been much more dangerous.
we did get intermediate models to study. But compute supply itself has also been accelerated.
I think most people now agree that frontier labs’ compute demand has led to the building of more fabs by TSMC + memory makers.
vibe researching this: installed compute capacity was growing at ~2.5× per year before chargpt. Extrapolating that trend to 2027, we would have had roughly 1/4 of the installed compute capacity we now expect in 2027 if growth had remained at pre-chatgpt levels. So if you take pre-chatgpt levels as the counterfactual, in 2027, frontier labs have accelerated installed compute by ~1-2 years.
I’m not sure how to value the extra ~1-2years. and whether that outweighs the value of the intermediate models we got to study.
but it definitely makes me less confident that the argument was sound, and makes me suspicious of similar arguments.
i updated on this take after reflecting on richard ngo's paragraph on compute overhangs.
in general if you are a big enough player (which most impactful people hope to be) , you should expect your actions to shape the environment quite abit.
website to sum up resources / tweet thread/ discussion for our introspection paper
https://modelintrospection.com
Some people (my mentor ethan perez ) said my weekly MATS research update slides were nice. Some rough tips i have: