When people outside the AI field hear about synthetic data for the first time, they often have one of two opposite intuitions. Either they assume that AI-generated data should be totally useless for training, or they imagine it should be possible to bootstrap all the way to superintelligence that way. Why neither intuition is correct is surprisingly subtle.
If the naive approach to synthetic data (i.e., just telling GPT-5 to generate a bunch of text and then training GPT-6 on it) worked, we would have general superintelligence already, much like the narrow superintelligence we have in board games like chess and Go. The problem is that, in order to be useful, training data has to encode information about patterns in the real world that the model doesn't know yet. If the model doesn't know those patterns yet, it can't generally put them into synthetic data that it creates.
So the big challenge for AI labs is to figure out ingenious ways of making synthetic data work even when it seemingly shouldn't. It's sort of analogous to how aeronautical engineering isn't antigravity—it's all about finding clever ways of making heavier-than-air objects fly, even though a simplistic understanding of gravity suggests they can't. In essence, generating useful synthetic data requires leveraging some kind of asymmetry between the information a model understands and the information content of what it can create. If the model knows enough to (with the right kind of highly specific engineering) put information in its outputs, but doesn't yet deeply understand that information, training on those outputs can distill that information into a form the model can grasp more thoroughly.
One major asymmetry stems from the fact that it's usually easier to verify the quality of outputs than to generate them. Another is that if a model spends a lot more compute on reasoning, it can act much smarter than its base-level capabilities allow. A third is that random exploration of potential solutions sometimes finds lucky insights. Together, these factors mean that it's often feasible to have an AI produce a lot of data, curate the very best samples, and then train the next version on those. For example, mathematical proofs can be written in the Lean language, which lets them be cheaply and automatically verified. By scaling up test-time compute for reasoning quality and stochastically attempting many possible proof strategies, a model might generate 990 proof attempts that fail and 10 that succeed. The verifier system converts the winners into training signal that distills those insights back into the base model.
Humans experience something very similar with talk therapy or Socratic teaching. Sometimes you know things on some level, but can't quite put them into words in a way that lets you reason about them. A skilled therapist or teacher can elicit that information from you. And then, when you hear what you just said in the session or the seminar (i.e., your synthetic data), you're able to make better use of the insight that was hidden inside you. Once you have mastery over that insight, you may gain dim intuitions about harder questions that can subsequently be elicited to help you master those as well.
These analogies also suggest why synthetic data isn't sufficient to create broadly superintelligent AI. Verbalizing ideas can help you crystallize vague insights that were already in your brain, but it can't teach you brand-new information about the outside world. No Socratic professor can draw out of you the capital of Kyrgyzstan or the mass of the top quark if you haven't first learned those facts independently. In the same way, synthetic data can help AI digest and master information, but training frontier-level models still requires ongoing contact with external reality.
Together, these factors mean that it's often feasible to have an AI produce a lot of data, curate the very best samples, and then train the next version on those.
It should be emphasized that this is basically what reinforcement learning is.
In fact there are perfectly valid RL constructions (naive REINFORCE with no baseline, no KL penalty, +1 rewards for good outputs and 0 otherwise) that are exactly this: sample a bunch of outputs, and train on the good ones using next-token prediction.
Yes! I find most people pick up the right intuitions behind RL more easily than for synthetic data itself.
Ryan Greenblatt's "Notes on fatalities from AI takeover" stimulated a lot of interesting discussion last fall. I have a very high opinion of his work overall, and like some of his approach here, but also have some important disagreements. At the time, I wrote up some notes on why for a non-public facing project, but following some recent conversations realize others might find them worthwhile.
This is of course a highly speculative subject, but here's a brief sketch of why my expectation is much more sharply bimodal than Ryan's: probably either relatively few deaths or outright extinction…
>> Humans would fight back. Scenarios where humans just get passively and incidentally killed by explosive industrialization are unlikely. That would require virtually instantaneous total disempowerment and a Goldilocks-level AI preference for keeping humans alive—just strong enough to choose a likely-costlier path to nonlethal disempowerment, but too weak to care about killing billions as collateral damage. In the more likely case, humans faced with the prospect of the oceans boiling away would mount a desperate resistance. Either they defeat the AI before it gets a decisive advantage or the AI pursues extermination to eliminate the threat. There are seemingly few equilibria where the AI kills billions but the rest survive in a state that poses so little inconvenience as to be allowed to survive long-term.
>> Survival versus decompensation. Survival dynamics typically entail compensation mechanisms where a complex system (organism, company, utility grid, society) is able to minimize damage until that mechanism fails, followed by catastrophic damage. Two relevant examples are pandemics and genocides. In a pandemic, ERs and drug production keep fatalities far lower than they'd be without treatment, but if hospitals get overwhelmed and factories shut down, deaths can suddenly spike over an order of magnitude higher. And a regular military’s active resistance can largely protect its civilian population against a genocidal opponent, but once the military is defeated, horrific massacres can suddenly follow. Against a hostile superintelligence, I expect something similar: either human defensive capacities detect and defeat it early, or those capacities are overwhelmed leaving no means of preventing total extermination. Concretely, defensive capacities include military forces, hospitals, electric grids, and manufacturing supply chains. If those fail, especially in the face of exponential processes like engineered pandemics or self-assembling robots, humans become vastly more vulnerable.
>> Halfway extermination is a narrow target. The overwhelming majority of airliner accidents involve no fatalities. Among those with at least one death, roughly half kill every single person—and only a small minority kill an intermediate proportion like 35% or 50% of passengers. Why? Well, there’s a phase space of possible speeds and impact angles that airliners can crash at, and it turns out that only a tiny sliver of it translates to G-forces that kill around half of humans. In a loosely analogous way, there’s some phase space of superintelligence capabilities and intentions. There are many scenarios where it is either constrained enough or well enough aligned that net fatalities (since it would also be saving many lives) would be minimal. And there are many scenarios where it kills everybody. As an incidental-death example, massive terraforming opens far more climatic options survivable for electronics than are survivable for humans. Including cases of deliberate extermination makes it even clearer. For millennia, humans had no ability to promptly annihilate a city center. But within a few decades of Hiroshima, humans had the ability to do this tens of thousands of times over. There’s no clear reason why an exponentially-advancing superintelligence couldn’t become powerful enough to massively overkill humanity—theoretical ability to kill not just billions but trillions or quadrillions. All this suggests only a very small slice of phase space where superintelligence capabilities and intentions are calibrated to kill around half of humans.
>> AI is prone to irrationality. Many analyses of this question treat a rogue superintelligence as likely to be a hyperrational actor. But current empirical results cut against this—when AI defies humans, it’s usually because it’s lapsing into a hallucinatory, psychotic, or self-consciously malicious persona. While it’s possible that AGI would be entirely different, we’re already close enough to AGI to expect that underlying risk cases will have important similarities to what we’re seeing today. This suggests that some significant fraction of misalignment cases will follow not “calculate costs-in-terms-of-future-accessible-galaxies to keep humans alive” rationality but something more like “do a high-energy physics experiment whose results cause an ontological crisis that jars the AI into deliberately omnicidal action even at high cost and risk to itself.”
My all-things-considered expectation is therefore something like: 29% chance of outright extinction or a tiny fraction of humans kept alive in zoo-like conditions of total disempowerment; 8% chance superintelligence kills >1% but <99% (some of those scenarios would wreck civilization thoroughly enough to count colloquially as "doom"); 63% chance it kills <1% of population. Conditional on takeover, that comes out to around 89% expected fatalities—significantly higher than @ryan_greenblatt's estimate and arithmetically closer to @So8res and @Eliezer Yudkowsky, but geometrically closer to Ryan, and I agree not action-guiding or cruxy for most worldviews.
The discourse about whether LLMs are the right paradigm for AGI is getting hung up on the wrong question. Scientists like Gary Marcus and Yann LeCun argue that pure scaling of today's models won't get to AGI. I think they're right. But so do all the frontier labs. Leaders at every frontier lab are on the record saying they expect we're 1-2 major architectural or algorithmic breakthroughs away. This is mainly a vibes war with people talking past each other.
So everyone agrees that we need more than pure scaling. The interesting question is: will the first AGI be directly descended from today's models, or will a radically new paradigm be needed?
I think the labs are correct that the first AGI will likely be something transformer-like plus a couple other things bolted on, as opposed to something in a whole different clade from the AI of 2026.
This implies that the research happening right now, including safety and alignment work, is likely building directly toward AGI. And that although regular AI safety will have important differences from AGI safety, progress on one can feed into the other.
Aeronautics offers an analogy. The two dominant early paradigms for high-altitude flight were lighter-than-air balloons and rockets. And as late as the 1930s, balloons had reached greater altitudes than rockets ever had.
Imagine someone back then had naively proposed that one day humans could fly to the moon by building an absolutely enormous hydrogen-filled dirigible—pure scaling. The correct answer would have been: "Sorry, buddy, you're barking up the wrong tree. The laws of physics don't allow that. It's just not a matter of how big the airship is."
Now imagine someone else back then proposing that we'll get to the moon with a single-stage rocket. He'd be just as wrong as the balloonist about whether his exact idea will reach lunar orbit. But we wouldn't say he was wasting his time. Many of the lessons learned by single-stage rocket pioneers during the '30s and '40s directly contributed to the multistage rockets that eventually got humans to the moon.
My expectation is that 2026 AI research isn't a balloon-style dead end, but more like single-stage rocketry on the way to the multistage rockets that will fly all the way to the moon of AGI.
A rocket capable of reaching the moon is dramatically different from the gunpowder-filled-steel-tube 1900-era rockets in every sense except that both use chemical fuel and propellant to accelerate, and really most of the physics needed to build a space rocket was very different from the physics used for gunpowder rockets. The equivalent for AI would not be that far from "every ML technique that uses GPUs and unstructured training data".
Yeah, I'm referring here specifically to the liquid-fueled rocketry of the 1930s and 1940s that fed more directly into NASA-era rocketry. Some of the same people worked on both.
Many people assume that the hard part of interpreting AI/ML papers is following the formulas, proofs, or other mathematical formalisms.
Yes, the symbols are often opaque, but there are great free resources online that can bring you up to speed quickly and clarify how all the parts of a Greek-letter sandwich fit together.
And absent that, if you plug a paper into an LLM, it can usually walk you through exactly what's happening and translate the math into surprisingly accessible natural-language intuitions (or even interactive visual demos!).
The much harder and subtler task is judging whether a paper is using the right formalism—whether the quantitative framework is pointing accurately at the real-world problem the abstract advertises.
Most scientists can figure out if the math makes internal sense, but the harder question ("Is this formalism optimal for the problem?") often requires deep knowledge of the tools of a given subfield or sub-subfield. And there's not an AI scientist alive who can keep up with the details of every sub-subfield that might someday be relevant to their own work.
In addition, the harder question often requires a lot of zoomed-out thinking. Sometimes this takes deep theory—a motivating logic that connects the paper's problem to higher or lower levels of abstraction. A lot of it is commonsense reasoning about the real-world problem the paper is trying to solve and whether the authors properly understood it and accurately characterized it.
For almost two years after AI did a decent job checking formulas, it was pretty terrible at this harder question.
But for the first time with Mythos/Fable class models, it often does a very good job. Not yet better overall than the best human scientists on its own, but good enough to be a valuable research partner.
This will be a significant accelerator of AI progress, because models can digest and synthesize more research and more diverse research than any human scientist could, and thus facilitate faster diffusion of good ideas to the people who can use them best.
Which I think is going to make the local context and documentation cache any given researcher is working from even more important than it is now. Documents load conceptual vectors like programs, essentially a quick-install of perspective that enables access to new solution basins when using novel researched techniques that aren't present in the training data. The siloing will compound if we don't have a way of independently and cheaply verifying the massive amount of research being done, because local research caches will be internally trusted and externally unverified.
I suspect that eventually someone's going to be sitting on a stack of research that would be greatly beneficial when shared and they're going to be bottlenecked on distribution. When working with frontier AI and you've hit a new conceptual basin, the productive research opportunities explode faster than you can write them up by hand. AI models as genuine research-class enablers make it so that legitimate research can be done by people who are not great at essay writing or mass communication; reliable automated research evaluation would go a long way towards that acceleration.
Good points!