Epistemic status: I am not a subject-matter expert in any of the fields in this article. The following is speculation, albeit I think reasonable speculation. Cunningham’s Law is why I wrote this essay.
AI Disclosure: Claude Opus 5.5 provided feedback to the first and second draft of this essay, similar in kind to a friend providing feedback on a draft. All text and errors therein are mine.
What does discovering new knowledge mean?
The most merciful thing in the world, I think, is the inability of the human mind to correlate all its contents. We live on a placid island of ignorance in the midst of black seas of infinity, and it was not meant that we should voyage far. The sciences, each straining in its own direction, have hitherto harmed us little; but some day the piecing together of dissociated knowledge will open up such terrifying vistas of reality, and of our frightful position therein, that we shall either go mad from the revelation or flee from the deadly light into the peace and safety of a new dark age. —H. P. Lovecraft, “The Call of Cthulhu”
…[W]hat do you make of the fact that these things have basically the entire corpus of human knowledge memorized and they haven't been able to make a single new connection that has led to a discovery?
At the time, the frontier was GPT-4. Like Dwarkesh, I was puzzled by this. When o1 was released in September of 2024, I thought that the answer to this was pretty straightforward: capability will go up, connections will be elicited from the models by RL, interesting results will follow. I think that hypothesis has borne fruit, as I’ll trace out with mathematics as the example field in a moment. The next interesting question, which seems to be severely underexplored, is “how much”. In other words, now that we’re leaving that “placid island of ignorance”, how much of an overhang is going to be absorbed by “correlating the contents” of human knowledge?
I think it’s helpful here to distinguish between two different types of new knowledge. In October of 2026, Matthew Schwartz published an article called “Claude-Shaped Science”, describing multidisciplinary work he had done with Claude.[1]He had a useful diagram of this idea of different types of knowledge:
Schwartz here is modeling human knowledge as a jagged shape (in blue) with each point of the shape (the spokes) being the frontier of the field. Total human knowledge can expand in two main ways: expand a spoke or fill in the area between spokes. For the purposes of this essay, “spoke problems” require a new concept, idea, or data that doesn’t already exist in the corpus of human knowledge. “Gap-filling problems” are problems where everything needed to solve them already existed in that corpus, but nobody had mixed the ingredients together. The question I’m interested in is how large and valuable is the total area between the spokes? Pretending that LLMs never drive any of the spokes forward [2], how much progress or disruption can we expect to see from “merely” filling in those gaps? This seems important because this provides a hard lower bound on the impact of AI.
Math as the vanguard of the fleet
The first field to experience large-scale gap-filling has been mathematics. As a brief timeline, with all of the benchmark numbers from Epoch:
2024: Pre-reasoning, GPT-4o (August) was at the frontier for math. GSM8K (grade-school math) was close to saturated; GPT-4o scored 53% on MATH L5 (high school competition math). The first reasoning models blew past MATH L5 (o1 scored 94%) and made major strides on Mock AIME (o1 at 73%). Terence Tao described working with o1 as being “roughly on par with trying to advise a mediocre, but not completely incompetent, (static simulation of a) graduate student.” No discoveries yet.
2025: In July, Google and OpenAI both announce IMO Gold. FrontierMath T1-T3 (v2 for consistency) are the main benchmark at this point, with the highest score at the end of the year being 72% by GPT-5.2 Pro. These tiers are research-adjacent problems that would take a researcher in the field hours to a few days. According to VibeMathed, the first AI-assisted open problem solution is solved on July 20. A few others happen by the end of the year. The first claimed AI-discovered solution (a disproof of Erdős #333) is published on December 25. It’s actually a rediscovery of a solution Erdős had already published in 1977.
2026: Things accelerate dramatically. The actual first AI-discovered solution (for Erdős #728) happens in January 2026. FrontierMath T1-3 and T4 (a harder tier, problems that would take a couple of weeks) are saturated by October of 2026. More interesting open problems fall: the Jacobian Conjecture on July 20, existence of non-sofic groups on August 1, and (arguably, the details are contested) Navier-Stokes on September 9. This leads to the October 6 release by OpenAI of hundreds of mathematical results. At least one mathematician (Alex Kontorovich) described the most significant solution (proof of the quasi-Riemann hypothesis) as being significant enough that “If a human did this, it would be an instant Fields Medal, no questions asked”.
I am not a mathematician. But from reading the commentary of mathematicians like Daniel Litt or Terence Tao, it appears to be the case that few, if any, of the results listed above are “moving the spokes” solutions.[3]The consistent theme is that LLMs are strongly superhuman at combining existing ideas in potentially novel ways rather than developing new theory themselves. This seems to me like convex hull gap filling. So, the answer to “how big is the convex hull overhang” for math seems to be pretty big! I’d be very surprised if the gap-filling in mathematics has stopped here.
What about other fields?
In short, I don’t know. My mental model of how the convex hull areas grow is that researchers specialize in narrow silos because moving the spokes is hard.[4] Any of the following things or a combination thereof could happen:
Transfer problems: The researcher runs into a problem that has been solved analogously in another sub-field, another field, or another subject entirely. A classic example is Don Swanson suggesting fish oil as a treatment for Raynaud Syndrome.[5] It appears the phylogenetics work Schwartz did is also an example of this category.
Search problems: The researcher can’t search through all of the information in another sub-field or field to solve their problem. The blogger Prinz’s decryption of RICHI-240 seems like this type of problem (the outside information here being contemporaneous telegrams that confirmed troop movements, and therefore what army divisions the telegram is talking about).
Bookkeeping problems: The researcher can’t trawl through, organize, or keep track of all the existing data in their area of study. An example of this type of issue is keeping track of all the biographical information of civil servants in two Victorian reference works. In this case it was solved through using Wikidata, not AI, but this is the shape of problem I’m referring to here.
There are probably several other relevant paths to these overhangs being formed. People at the spokes have a much better understanding of how the convex hulls are formed in their field than I do.
My intuition is that the areas here are very large and that filling in the gaps is going to be a massive wave of progress on a lot of problems. I don’t think it’s “full sci-fi, Dyson Sphere in ten years” big, but my rough guess is “Internet/Printing Press” on the technological Richter scale on the low end.
For field-specific predictions, I weakly expect that social sciences like history are going to see an armada of stunning convex hull type solutions. My expectation is similar for the sciences like biology, but at least one friend who is an active virology/immunology researcher is skeptical of the overhang being that pronounced. His argument is that hypothesis generation (at least in his field) is not a significant bottleneck now, and any hypothesis has to be verified with time-consuming wet-lab work. This seems plausible enough for his field, but things like drug repurposing or the above fish oil-Raynaud example make me think there’s at least some fertile ground. I don’t know the answer to how much fertile ground there is, which is why I’m writing this essay to begin with.
As for what the best way to harvest that fruit, I think Schwartz’s guidelines are reasonable. I also think that researchers should be collecting, clearly demarcating, and writing down problems as much as possible. I suspect that one of the reasons math has seen such a glut of convex hull problems first (beyond obvious things like math being an easy to verify domain) is that a big chunk of “doing math research” is based on moving the ball forward on written-down problems. If I’m right about this point, fields that can’t imitate that model are plausibly going to have smaller overhangs. I also think breadth of knowledge is valuable for identifying potential convex hull areas.[6]
Why does this matter?
I think this question is important for a few reasons.
Regardless of whether AI eventually makes the centaur model (human + AI teams collaborating) obsolete, there’s the question of what to work on right now. If my suspicion is accurate, a huge mass of hitherto infeasible work is now available. Suddenly there’s a way to harvest low-hanging fruit that’s accumulated throughout history.
The size of the convex hull overhang determines part of the lower bound of disruption due to AI. Even if somehow LLMs totally plateau (or are stopped by pacing/a pause/a shutdown), they never improve on research taste, and they never become able to push the spokes forward, diffusion of existing LLM capabilities will lead to incredible change.
As a consequence of the lower bound point, this matters for meaningful dialogue with people who are resolutely not AGI/ASI-pilled. If the overhang is large, then no “sci-fi” stories of AGI or ASI are necessary for disruption far larger than seems to be accepted by the AGI/ASI skeptics or “normal people” not plugged into the discourse. At the very least this seems like a crux point for my conversations with people like my aforementioned biology researcher friend.
Finally, I think it’s entirely plausible that we blow past the centaur model very quickly, RSI happens, crazy sci-fi future or doom, fill in the details from the AI Futures Project or other ASI scenarios as you see fit here. I’m confident that the future is going to be weird and knowledge is going to expand rapidly, but my error bars are very large on what that looks like. If nothing else, I’m going to be very curious how much of that knowledge expansion will be moving spokes or filling in the gaps.
I’m not qualified to judge whether this work is meaningful. At least some of the domain experts he worked with seemed impressed, but my thesis doesn’t depend on the quality of his specific output.
Opus 5.5 filled in a gap for me here; apparently this is a well-known problem called “the burden of knowledge”. I knew the idea, I just didn't know it had a name.
Epistemic status: I am not a subject-matter expert in any of the fields in this article. The following is speculation, albeit I think reasonable speculation. Cunningham’s Law is why I wrote this essay.
AI Disclosure: Claude Opus 5.5 provided feedback to the first and second draft of this essay, similar in kind to a friend providing feedback on a draft. All text and errors therein are mine.
What does discovering new knowledge mean?
The most merciful thing in the world, I think, is the inability of the human mind to correlate all its contents. We live on a placid island of ignorance in the midst of black seas of infinity, and it was not meant that we should voyage far. The sciences, each straining in its own direction, have hitherto harmed us little; but some day the piecing together of dissociated knowledge will open up such terrifying vistas of reality, and of our frightful position therein, that we shall either go mad from the revelation or flee from the deadly light into the peace and safety of a new dark age. —H. P. Lovecraft, “The Call of Cthulhu”
A long time ago (meaning, August of 2023), Dwarkesh Patel asked Dario Amodei a fantastic question:
…[W]hat do you make of the fact that these things have basically the entire corpus of human knowledge memorized and they haven't been able to make a single new connection that has led to a discovery?
At the time, the frontier was GPT-4. Like Dwarkesh, I was puzzled by this. When o1 was released in September of 2024, I thought that the answer to this was pretty straightforward: capability will go up, connections will be elicited from the models by RL, interesting results will follow. I think that hypothesis has borne fruit, as I’ll trace out with mathematics as the example field in a moment. The next interesting question, which seems to be severely underexplored, is “how much”. In other words, now that we’re leaving that “placid island of ignorance”, how much of an overhang is going to be absorbed by “correlating the contents” of human knowledge?
I think it’s helpful here to distinguish between two different types of new knowledge. In October of 2026, Matthew Schwartz published an article called “Claude-Shaped Science”, describing multidisciplinary work he had done with Claude.[1]He had a useful diagram of this idea of different types of knowledge:
Schwartz here is modeling human knowledge as a jagged shape (in blue) with each point of the shape (the spokes) being the frontier of the field. Total human knowledge can expand in two main ways: expand a spoke or fill in the area between spokes. For the purposes of this essay, “spoke problems” require a new concept, idea, or data that doesn’t already exist in the corpus of human knowledge. “Gap-filling problems” are problems where everything needed to solve them already existed in that corpus, but nobody had mixed the ingredients together. The question I’m interested in is how large and valuable is the total area between the spokes? Pretending that LLMs never drive any of the spokes forward [2], how much progress or disruption can we expect to see from “merely” filling in those gaps? This seems important because this provides a hard lower bound on the impact of AI.
Math as the vanguard of the fleet
The first field to experience large-scale gap-filling has been mathematics. As a brief timeline, with all of the benchmark numbers from Epoch:
2024: Pre-reasoning, GPT-4o (August) was at the frontier for math. GSM8K (grade-school math) was close to saturated; GPT-4o scored 53% on MATH L5 (high school competition math). The first reasoning models blew past MATH L5 (o1 scored 94%) and made major strides on Mock AIME (o1 at 73%). Terence Tao described working with o1 as being “roughly on par with trying to advise a mediocre, but not completely incompetent, (static simulation of a) graduate student.” No discoveries yet.
2025: In July, Google and OpenAI both announce IMO Gold. FrontierMath T1-T3 (v2 for consistency) are the main benchmark at this point, with the highest score at the end of the year being 72% by GPT-5.2 Pro. These tiers are research-adjacent problems that would take a researcher in the field hours to a few days. According to VibeMathed, the first AI-assisted open problem solution is solved on July 20. A few others happen by the end of the year. The first claimed AI-discovered solution (a disproof of Erdős #333) is published on December 25. It’s actually a rediscovery of a solution Erdős had already published in 1977.
2026: Things accelerate dramatically. The actual first AI-discovered solution (for Erdős #728) happens in January 2026. FrontierMath T1-3 and T4 (a harder tier, problems that would take a couple of weeks) are saturated by October of 2026. More interesting open problems fall: the Jacobian Conjecture on July 20, existence of non-sofic groups on August 1, and (arguably, the details are contested) Navier-Stokes on September 9. This leads to the October 6 release by OpenAI of hundreds of mathematical results. At least one mathematician (Alex Kontorovich) described the most significant solution (proof of the quasi-Riemann hypothesis) as being significant enough that “If a human did this, it would be an instant Fields Medal, no questions asked”.
I am not a mathematician. But from reading the commentary of mathematicians like Daniel Litt or Terence Tao, it appears to be the case that few, if any, of the results listed above are “moving the spokes” solutions.[3]The consistent theme is that LLMs are strongly superhuman at combining existing ideas in potentially novel ways rather than developing new theory themselves. This seems to me like convex hull gap filling. So, the answer to “how big is the convex hull overhang” for math seems to be pretty big! I’d be very surprised if the gap-filling in mathematics has stopped here.
What about other fields?
In short, I don’t know. My mental model of how the convex hull areas grow is that researchers specialize in narrow silos because moving the spokes is hard.[4] Any of the following things or a combination thereof could happen:
My intuition is that the areas here are very large and that filling in the gaps is going to be a massive wave of progress on a lot of problems. I don’t think it’s “full sci-fi, Dyson Sphere in ten years” big, but my rough guess is “Internet/Printing Press” on the technological Richter scale on the low end.
For field-specific predictions, I weakly expect that social sciences like history are going to see an armada of stunning convex hull type solutions. My expectation is similar for the sciences like biology, but at least one friend who is an active virology/immunology researcher is skeptical of the overhang being that pronounced. His argument is that hypothesis generation (at least in his field) is not a significant bottleneck now, and any hypothesis has to be verified with time-consuming wet-lab work. This seems plausible enough for his field, but things like drug repurposing or the above fish oil-Raynaud example make me think there’s at least some fertile ground. I don’t know the answer to how much fertile ground there is, which is why I’m writing this essay to begin with.
As for what the best way to harvest that fruit, I think Schwartz’s guidelines are reasonable. I also think that researchers should be collecting, clearly demarcating, and writing down problems as much as possible. I suspect that one of the reasons math has seen such a glut of convex hull problems first (beyond obvious things like math being an easy to verify domain) is that a big chunk of “doing math research” is based on moving the ball forward on written-down problems. If I’m right about this point, fields that can’t imitate that model are plausibly going to have smaller overhangs. I also think breadth of knowledge is valuable for identifying potential convex hull areas.[6]
Why does this matter?
I think this question is important for a few reasons.
Finally, I think it’s entirely plausible that we blow past the centaur model very quickly, RSI happens, crazy sci-fi future or doom, fill in the details from the AI Futures Project or other ASI scenarios as you see fit here. I’m confident that the future is going to be weird and knowledge is going to expand rapidly, but my error bars are very large on what that looks like. If nothing else, I’m going to be very curious how much of that knowledge expansion will be moving spokes or filling in the gaps.
I’m not qualified to judge whether this work is meaningful. At least some of the domain experts he worked with seemed impressed, but my thesis doesn’t depend on the quality of his specific output.
This seems to be a significant assumption of the AI skeptic argument and is probably already falsified, but we’ll grant for this essay.
For the new OpenAI results, I don't know if anybody knows yet the breakdown of which category applies to the solutions in the repo.
Opus 5.5 filled in a gap for me here; apparently this is a well-known problem called “the burden of knowledge”. I knew the idea, I just didn't know it had a name.
Another Opus 5.5 gap fill for me writing this essay.
This is my main cope for why my liberal arts degree still has any value at all.