This post is surprising to me. In the past few years, models have been getting smarter so fast that superintelligence seems very close, probably no further than 2030. That's without looking at spending at all (in dollars, flops or whatever). You seem to be saying, based on spending, that it's actually much further out. How can that be true?
The key premises are the slow model-building cycles of LLM/RL AGI technology, which need to be running for the biggest models, and the EUV machine bound compute slowdown (since transistors no longer significantly shrink). The total number of EUV machines doesn't double every year, and AI compute has already eaten most of TSMC capacity (see the plots under the words "demand is unlikely to be entirely fulfilled"). So without an industrial explosion, there just won't be FLOP/s to spend in the whole world to 1000x the compute compared to what it takes to run a model-building cycle for the kinds of models that can participate in prosaic RSI, because the biggest models will barely qualify for that role just as the compute slowdown starts.
superintelligence seems very close, probably no further than 2030
The slow-learning prosaic RSI model I'm using clearly doesn't predict superintelligence in 2030. The key bottleneck is the inability to autonomously advance cultural accumulation very quickly (come up with new ideas visible from the current standpoint, add them to the collection of deep skills a model has, thus advancing the current standpoint, repeat). Most important inventions crucially rely on what's already understood, and starting from there they take one more step. Smart LLMs might take longer steps (from the current standpoint of what those LLMs already understand) than smart humans, but that doesn't matter if the steps fail to chain and accumulate much faster than humanity's own (technological) culture. Fast reasoning, slow learning.
Being smarter in the moment only matters if a few steps taken by the LLMs suffice to defeat the hobbling of slow model-building cycles. And I'm not even sure LLMs of the early 2030s will be much smarter, based on the effects of scaling so far (this gets clearer in 2028-2029, halfway to pre-slowdown scaling of 2031-2032). I think it's very likely they will be solidly in the very smart human range (in pretraining-induced general awareness, not just RLVR-targetable narrower skills), and thus ready to be trained to fully automate the next-model building loops.
You seem to be saying, based on spending, that it's actually much further out.
As I estimate in the model size scaling post, the 1.4 quadrillion total params model of 2031 is pretrained for 2e29 FLOPs using 10 GW of Feynman chips. You need 10 TW of Feynman chips to 1000x that, so it's not even close to being a matter of spending, at least when we are talking 2030s and the chips are not yet growing on trees.
It sounds like you think one specific drawback of current LLMs - their inability to have some "sleep" and consolidate their memory - will be enough to slow down progress by a decade or more. To me this seems unlikely. People (and LLMs) will be chipping away at this problem pretty fast.
To estimate an upper bound on time to software-only singularity and ASI (excluding a ban/pause or other disruptions), this post considers the most prosaic nothing-ever-happens scenario (based on my understanding of the current/imminent paradigm). For this scenario, the bound of 2040-2045 turns out to be likely, and the bound of 2050 extremely likely. I'm not claiming it to be likely that this scenario itself persists all the way to 1000x compute. If a breakthrough arrives earlier, then ASI just happens earlier, without invalidating the upper bound on its timing from the slower prosaic scenario.
But also, breakthroughs are fickle, and strong continual learning (online full weight updates that instill RLVR-grade learnings without trashing the existing model capabilities or sanity) is already a very old problem, so going unsolved by humanity for another 15 years doesn't seem very unlikely (other ways of triggering ASI are probably even less promising in this timeframe). There's a boost from the remaining scaling before the 2032+ slowdown (that might unlock things that are too hard to pinpoint at a lower scale), and possibly low-hanging fruit that the 2028-2032 LLMs uncover (things that don't require accumulation of novel deep technological culture, but would otherwise take humanity longer to notice), which is why I'm giving it 50% by 2032. But then what's left is human ingenuity and the next-model building cycle speedup that I'm estimating to only start kicking in as a result of industrial explosion, which gets 10x strong by 2038-2043 and 100x strong by 2042-2047.
"it doesn't seem too unlikely that nothing substantively new gets invented until 2040-2050"
This seems like it would be the longest single drought in major industrial / civilizational advances since we invented flight, at the vey latest, and the slowest in computational methods and advances since before the vacuum tube - why would you expect everything to slow down so much?
I'm not saying the lack of the relevant paradigm-breaking innovations is likely, merely that it's "not too unlikely", so it's still worth considering. It did in some sense take more than 60 years to invent transformers. This post explores an upper bound on how long it takes to set off a software-only singularity. It might still happen at any moment, as I mentioned in the post, and the current compute already seems sufficient to scale it well past a takeover.
I'd give 50% for takeoff happening by 2032 (before the compute slowdown that precedes industrial explosion), the substantive observation in the post is that it won't be happening after 2060 without a ban/pause or another disruption, and this estimate doesn't depend on how long it takes to do the basic research (because by 2045-2050 even the "slow-learning" LLM/RL AGIs will have completed centuries worth of research).
"It did in some sense take more than 60 years to invent transformers."
I think this is wrong in a meaningful sense; transformers weren't useful until we had enough compute. A bare-minimum transformer model would be tens of millions of parameters, and require tens of millions of FLOPs to do inference, and a large multiple of that to train. So it's like saying we didn't have Minecraft back then; true, but not useful; it couldn't have been written until computers were fast enough and had good enough graphics.
A key part of the answer is that algorithmic progress is scale-dependent to an extent that is often not realized, and >90% of algorithmic progress is basically downstream of scaling compute, so any slow down in the scaling of compute automatically slows down all other progress as a side effect:
Here's a useful article below:
...but then this is an argument that more compute won't be available, and the chip roadmap for NVIDIA goes through the Feynman Architecture in 2028, then panel level packaging and HBM5, along with increased production volumes - so it seems implausible we'd see a compute availability slowdown?
Also, the 'algorithmic progress is actually compute progress' argument is far older and more general than the version they present for LLMs, and is convincing to some extent, but also weaker than I think it appears if you look at the data.
The slowdown is relative to the 2022-2027 exponential. After 2028-2029, there will be increasingly serious problems with increasing incremental volume (the amount of the new buildout per year). Then there are a few years of accumulation at a slowly increasing incremental volume, and around 2032 decomissioned (or no-longer-relevant) old compute starts catching up with the new compute (see the plot under the words "how silicon capacity holds back AI chip deployment").
Thanks; that makes sense, but I think it doesn't make the case I think you're expecting. The claim linked in that article is that TSMC won't be able to get cleanroom space in 2026/2027; unless they are blind, they'll avoid repeating that mistake for 2028/2029 - https://newsletter.semianalysis.com/p/the-great-ai-silicon-shortage But even if all of that happens, the new chips are faster, and will be made available; the increase of NVIDIA's chips just won't be quite as exponentially large as historically.
But even then, it's not like Google's TPUs and others are so far behind that a multi-year fumble wouldn't allow other firms to catch up, and China is certainly going all in. So even if the analysis is correct, it won't eliminate the exponential trend.
Chips will get faster, but chips are made out of logic dies, and transistors no longer shrink that much. So FLOP/s per GW won't get a lot better, and silicon capacity is binding for FLOP/s (which is more about EUV machines than TSMC fabs). Future process nodes consume more EUV cycles per wafer, so global incremental FLOP/s manufactured per year don't obviously grow (notably) faster than the number of available EUV machines.
In my queue of posts to write is 'Industrial explosion before intelligence explosion?' - I think you probably nailed that better here than I would! Though I might write my own spin on it anyway at some point.
Perhaps my best case against:
Slow learning is about the reasoning-learning cycles, inability to advance their serial depth quickly and thus accumulate technological culture from distant future. Better research taste doesn't help if the results it generates can only be learned at the low model-building speed, rather than at the high token-generation speed. If an extremely smart (but not acutely superintelligent) model still stands on the shoulders of humanity in its initial understanding when working on a problem, it won't be able to leap very far. It also needs to be able to keep advancing the state of the art, and to keep learning what it invents as fast as it invents it (rather than much slower), so that it can put the new learnings to work immediately and keep iterating.
Taste seems to be something like generalization in the classical sense, doing well in increasingly unfamiliar situations (since research taste deals with things that have never happened before, both recent experimental results and possible future experiments). Prosaic RSI of this post defeats generalization by automatically designing RL tasks/environments/graders that develop relevant and more verifiable proxy skills to patch the capability gaps for the situations actually encountered in practice. Humans also need robust feedback loops (easier relevant proxies even for the tasks that are verifiable themselves, just too hard to directly solve), so this is not an unusual requirement for a unit of learning, or an unusual relationship between a unit of efficient learning and the more nebulous purpose of that learning; fields like philosophy suffer greatly from inability to find relevant verifiable skills.
I don't think automated generation of RL tasks has bad sample efficiency with respect to the data describing the situation that witnesses a capability gap and inspires the design of new proxy skill learning feedback loops. Proxy skills that are more verifiable (or easier/faster to practice) also inherently have some modularity (compared to the situations that inspire them), so some generalization in the classical sense (which is taste, the way I understand it) naturally comes along. Pretraining scale should also help, and the 2028-2029 models will be close enough to the pre-slowdown 2031-2032 models to see how far that goes.
Recent results and this kind of thinking has lead me to believe that RSI is just not so useful a concept anymore. When AGI/ASI was further away it did, but now the specific details are what matters.
Specifically cracking the "neural code" is what will matter, rather than if our AGI can act slightly better than humans at coordinating large training runs of inferior architectures.
Its likely true that the transformer architecture in combination with our compute, data constraints etc wont lead to the kind of take-off we have in mind. Sample efficiency to me is the last and very important issue to solve. That is humans etc are about 3-6 OOM faster are learning from small amounts of data. Whatever approach cracks that inevitably gets to at least mild ASI. For example, lets say we get this code from biology experiments and are unsure how to apply it to silicon. Our current human knowledge is more than sufficient to do this and likely lead to a large improvement in capabilities, faster than if we had actual RSI on the transformer architecture.
To me there are cruxes, rather than RSI.
RLVR training doesn't have a sample efficiency problem with respect to the inspiration data for automatically written RL tasks/environments/graders, that's the key feature of the current paradigm that makes general learning possible at all (even if the more nebulous tasks have to be refactored into multiple proxy subskills that are more verifiable first). And it's only slow if the learning loops can't be applied online, without many such loops destroying the model over time.
What makes prosaic RSI surprising is the disparity between fast reasoning and slow learning. The classical argument for why RSI very quickly leads to a takeoff is that chips are faster than brains, therefore once the thing works at all, it also goes much faster than humanity at accumulation of technological culture. I'm going along with the erosion of the term "RSI" to also extend to the slow-learning regime because I think this is how it's in fact being used by the LLM companies right now, and the qualifier of "prosaic" is my attempt to distinguish this very different kind of RSI (that the LLM companies fail to distinguish, and that many people expect to also go very fast, just as the classical argument said).
Yet strong RSI with fast learning is still going to happen at some point, whether it's in 2-10 years or in 10-20 years, and I don't see any issues with the classical notion of RSI in describing what it would look like. Even the industrial explosion that makes prosaic RSI go fast is not yet strong RSI, that milestone is only reached once learning is no longer slow with respect to reasoning, once it's efficiently using the available compute without sacrificing serial depth of the reasoning-learning cycles. Prosaic RSI going fast (or some other route of independent invention) is merely a way to get that different regime of strong RSI started (which is just RSI in the classical sense).
So it seems difference between our positions is that I think strong/classical RSI may never happen and not be possible. That is once you apply the lessons from biology to our GPU/computer you get into diminishing returns. The mild superintelligence that results from this then tries and fails to find another architectural jump (or long significant series of minor jumps), or proves that there isn't one.
In that case you need to distinguish three things:
I think strong continual learning quickly leads to architecture-rewriting strong RSI, because it's sufficient to quickly generate/accumulate technological culture and thus discover new more efficient architectures that can convert the same compute into even more intelligence. And that prosaic RSI is likely too slow to quickly discover strong continual learning or any other things in the architecture-rewriting strong RSI attractor. Thus prosaic RSI is also a stable regime that can persist for a while, either until an architecture leading to the attractor of achitecture-rewriting strong RSI is discovered by humanity, or until the industrial explosion produces enough compute for prosaic RSI to itself go much faster than humanity despite its inefficiency, discovering an architecture of strong RSI.
Thank you for sharing! I got a kick out of the term "prosaic RSI" haha. I think all your timelines seem reasonable to me as an upper bound. The main place I’d differ is that I expect robot mass production to begin before AGI makes robots economically profitable. As we’ve seen with the compute buildout, investors and governments are perfectly willing to pour money and resources into an industry if they believe it will be strategically critical in the future. So I’d probably move that milestone back by a few years.
I don't understand even this level of pessimism; robots are already profitable, and they have grown 5x in the past 20 years, to be a $50b industry, and the "coming" wave of investment already started; projections from "optimists" have it growing to $2.5tr by 2035, and they aren't banking on ASI at all, just current methods working out increasingly well.
I presume you meant humanoid robots not industrial ones and not vacuum cleaners, right?
~85% of the humanoid robots are sold in China, and about 3/4 of all the demand in China is from the education and research sectors—including, notably, government-backed training centers which sell the data back to the manufacturers: https://www.ft.com/content/26735a23-315f-47ef-8cf2-6c6ea9713998
Yep, apologies, I meant specifically humanoids or, more broadly, “general-purpose” robots. I take your point that, at least in China, we seem already to have reached an inflection point. My intuition is that robotics today may be roughly where LLMs were in 2023. Not yet broadly economically transformative, but experiencing a rapid surge in investment and technical progress, but with China and the United States effectively swapping places in terms of who is leading the buildout.
Technologies can pass this spot of "rapid surge of investment and technical progress" and then flop commercially, think of all the nuclear-powered (civilian) transport theoretically made available by the development of compact PWR in the 1950s and 1960s, supersonic air travel in the 1970s or GOFAI in the 1980s. For more contemporary examples, consider high-temperature semiconductors, fuel-cell cars, VR or thermonuclear fusion, which honestly hasn't flopped yet but IMHO is going to.
I don't argue that humanoid robots are necessarily going to flop, it's more common that a technology can find its limited niche (think of segways or blockchain), but there's a fundamental hard tradeoff between expensive strain wave reduction gears allowing high torque needed for carrying and moving heavy objects and cheaper planetary gears. Unlike AI, humanoids are essentially a mature technology
what is even the steelman case for humanoid robots? it seems like the argument from fiction written large and expensive. is it just that "our built environment is built largely for humanoid-shaped beings"? or maybe "for fast adoption, people need to be ... comfortable(?) ... familiar(?) with the shape of our robots"? is there some sort of "for a single tool that can do a wide variety of jobs, the humanoid shape is optimal"? (because this is super not true).
I'd say the most plausible way this occurs is if something close to human-like hands is necessary for a lot of robotics value, then it's plausible that the shape of general/special purpose robots will be constrained enough to force it to be humanoid.
Don't expect this to be true, because while I believe the premise is likely true, the implication probably doesn't hold, and I expect non-humanoid shapes to be used during an industrial explosion much more than humanoid shapes.
Yeah, kind of "We have built the environment for humans, robots have to be produced in large quantities in order to become cheap due to Wright's law, humanoid form factor together with AGI provides very large total addressable market because of direct human substitution".
Nesov's conservative timeline provides almost a decade to develop the strain-wave gear industry, which is at least plausible, even though that requires very large investments over the years
AI datacenters also existed before the ChatGPT moment, they just didn't get to a trillion dollars of buildout per year until a few years after the unit-level profitability (at a large TAM) prompted the industrial conversion. Similarly, there are some robots now, and there will be more robots before they become very useful for everything, but getting to the kind of scale that can match human industry's compute buildout as its side effect will still take a lot of time after the robots become profitable at the unit level (with the kind of somewhat-concretely-visible TAM that dwarfs the automotive industry).
It really seems like you're assuming that investors won't notice a widely predicted trend that starts materializing quickly enough to make money on it, which...?
I understand that there are some bottlenecks in robotics equipment, but they aren't as fundamental as the UV lithography constrain for chips - and one of the key bottlenecks is chips, and we're obviously seeing lots of money invested in building out that capacity already.
My claims aren't that different from what you said in the other comment:
projections from "optimists" have it growing to $2.5tr by 2035, and they aren't banking on ASI at all
The timeline in this post puts parity with automotive industry at 2034-2039, which is currently at about 2-3 trillion dollars of revenue per year. I'm also explicitly not banking on ASI at all with the conservative assumptions of this post.
I'm reading this post as an attempt to upper bound the time to ASI (basically, "even if near-term software intelligence explosion fails, the robot doubling will still bring about ASI by 2050 by sheer brute force"), and to me this line of reasoning seems sound, provided that human regulations allow the robot doubling to happen to the extent described in the post, either on Earth or in space, and nothing else like WW3 stops it.
I think if the post is indeed trying to set an upper bound rather than an average estimate, it should be clearer about it, otherwise the post can be misread as providing a takeoff estimate rather an upper bound, which to many is excessively slow because this post rests on the assumption that near-term software intelligence explosion is not possible with current compute.
The opening paragraph has the following words, with a link that goes into more details clarifying:
The invention of ASI ... is still possible at any time (and very quickly scales, given all the compute)
And at the end of the post:
This is the upper bound on the timing of feasibility of superintelligence, which mostly assumes just the current paradigm ... rather than any particular future breakthroughs
You do have a point in that apparently this failed to robustly communicate, given some of the other comments. This probably needed to be part of the framing rather than merely stated outright. Hopefully the multiple comments I wrote in response have put the matter to rest.
Given that we are in the AI age AND Fermi paradox AND Late Filter based on SIA Doomsday, AI must be the Late Filter, but not ASI. Thus there should be a relatively long period between AGI and ASI, during which AI will be used by bad actors for wars. Your post suggests plausible mechanism for such pause.
Industrial explosion is what will make the next-model building loops (and thus learning) with LLMs 1000x faster by about 2050, if indeed the slow-learning prosaic RSI becomes AGI before the big compute buildout slowdown of 2032+ that is already starting. This puts an upper bound on how long it takes to invent ASI that sets off software-only singularity, implementing efficient online learning and fixing all the other hobblings of the likely near-future AGI technology (LLMs/pretraining/RL). The invention of ASI in that sense is still possible at any time (and very quickly scales, given all the compute), but the likely initial state of slow-learning AGIs of 2028-2032 doesn't seem to give them a significant advantage over humanity in getting there faster. And so it doesn't seem too unlikely that nothing substantively new gets invented until 2040-2050, when the LLM/RL AGIs start accelerating because of the industrial explosion they set off.
Fast Reasoning, Slow Learning
The current methods are likely to enable automated general learning (thus AGI) very soon, using automated creation of RL tasks/environments/graders filling the visible gaps in model capability for the topics and situations that happen to be borderline unfamiliar for that model, followed by automated next-model building. This teaches LLMs deep skills, but operates at the speed of next-model building (weeks to months for one iteration of advancing the deep skill frontier) rather than at the speed of next-token generation (100-1000x the human speed). Humans are not the key bottleneck to the speed of next-model building loops, there's still a lot of waiting for the compute to do its thing in training, so achieving prosaic RSI by teaching the near-future LLMs all the skills necessary to perform it automatically doesn't make it go too fast. Using smaller LLMs to make everything faster doesn't work because the current frontier LLMs are probably borderline insufficient for learning the prosaic RSI skills that automate the next-model building loop. The LLMs of 2028-2031 (that are very likely sufficient) will be even bigger, though the cost of training or running them is not as bad as "quadrillion total params" sounds. This cost can't be circumvented by using different hardware that makes LLM inference much faster, because different hardware doesn't reduce the necessary number of FLOPs, which are not terribly wasted even in RL training and inference that involve bandwidth bound decode. Fundamentally, cost is the amount of compute, and the only thing that overcomes it is the scale of the global buildout.
Compute Slowdown, Industrial Explosion
Since LLMs become ready to close the next-model building loop (that enables AGI) just as the human industry runs out of various kinds of fuel for quickly increasing the scale of the compute buildout, there is no opportunity for another near-term 1000x increase in the available compute (and thus the speed of next-model building loops), the way compute was increasing in 2022-2027, and the way it'll keep increasing (a bit slower) in 2028-2032 until the pace of decommissioning old compute somewhat catches up to the 2028+ pace of producing new compute, set by factors like availability of EUV machines and skilled human labor. Thus the big compute slowdown of 2032+, stronger than the end of the current exponential scaling of compute by 2028+.
Without paradigm-breaking algorithmic innovations, the slow-learning AGIs can't quickly invent such innovations, and so the more predictable component of the pace of progress is set by the pace of the compute buildout. But also, the AGIs (likely available since 2028-2032) make the industrial explosion of robot-building robots a predictable medium-term development. It probably doesn't start right away, since the AGIs are not much faster than humanity at taking care of all the novel engineering challenges (requiring many next-model building loops to get good), and before it goes into full swing humanity still needs to handle the industrial side of things (at the human pace).
The automotive industry and the compute buildout acceleration of 2022-2027 seem like good anchors for how this might unfold. The process starts once AGIs unlock an outsized demand for robots (by making them very useful for everything), and the industry starts reshaping itself to increase the supply as fast as it can, similarly to the consequences of the ChatGPT moment. Robot production exhausts the industrial capacity of the supply chains within 3-5 years (similarly to how it took 5-6 years to reshape compute production). At that point, the scale of the robot supply (anchored to the current automotive industry) approaches a significant portion of human labor, so the process continues right past the limits of human industry without another big slowdown. If the doubling time of the robot-building industry (autonomously operated by AGIs using the existing robots) is around 1 year, then 10 years of this process increase the industrial capacity about 1000x. The countdown should probably start from the end of the 3-5 year period of industrial conversion, when the robot industry first matches a sufficient fraction of the human industry to also start producing as much compute (together with all the other precursors), and the slow-learning AGIs probably need the time for the next-model building cycles to figure out how to automate everything.
Prosaic Timeline to Takeoff
The timeline starts with prosaic RSI in 2028-2032. The resulting slow-learning AGIs first make robots very useful generally within 2-3 years, in 2030-2034, setting off the industrial conversion of human industry towards robot production. This lasts another 3-5 years, and by 2034-2039 the automated robot-building industry operated by robots and AGIs first matches the human industry's capacity in terms of the compute buildout it can support. If this industry can quickly reach a doubling time of 1 year, it can 1000x the compute buildout by 2044-2050. At that point, the next-model building cycles take 1000x less time, and so the unpredictable paradigm-breaking algorithmic innovations necessary to set off a software-only singularity happen on the scale of months instead of centuries, and would've happened at some point earlier than that, probably by 2040-2045, but certainly by 2050. This is the upper bound on the timing of feasibility of superintelligence, which mostly assumes just the current paradigm (extremely hobbled in its efficacy at superhuman invention) rather than any particular future breakthroughs.