After attending ILIAD: Aeneid and talking with @Richard_Ngo, I've been thinking a bit about how to get ideas, particularly by doing mathematics.
In scientific inquiry, the true hypothesis often hasn't occurred to you yet. Worse, the truth might be too complex to hold in mind, so that any hypothesis you can consider must be incomplete. This is the type of situation that I believe Richard likes to think about; he claims that we do not have the right concepts yet to understand agency, and developing them is robustly beneficial for A.I. safety.
(But it's not always about truth. Sometimes you just need better ideas, because all of your options are looking doomed. Agent foundations is about trying to deeply understand agents, but conceptual A.I. safety research can be broader, also including the invention of devices to control agents.)
A.I. safety needs to invent better concepts and better ideas. I think that agent foundations has cultivated a particular way of doing mathematics which aims to inspire such creativity.
Why math?
At ILIAD, Eliezer questioned whether anyone's alignment agenda was actually bottlenecked on solving a math problem. ILIAD attendees do a lot of math nominally related to A.I. safety, but no one really has a story about how even a proof oracle gives you safe A.S.I. So what are we all doing, spending our time on math problems that don't tell us how to save the world?
I think we're mostly "scrying."
Scrying is staring hard at one thing, in order to try to find out about another. In the occult origins, it's often a mirror or glass, viewed under darkened conditions to inspire hallucination.
MIRI was mathematically scrying when they thought about probabilistic truth predicates, Löb's theorem, and logical induction. Their goal was not to directly use these results to build self-improving A.S.I., but to become less confused about embedded agency.
Today, I think that Resolution is basically scrying about the interactive proofs of complexity theory, hoping to get ideas that will work for natural language reasoning, which is called debate.
Scrying is not the only use of math in A.I. safety, it's the indirect use of math for hopefully stumbling upon A.I. safety ideas. Some math for A.I. safety is more direct than scrying, usually because it's modeling something. For example, @Guillaume Corlouer tries to understand how stochastic gradient descent on neural networks works. He might start with some simplified neural nets, but he seems to be trying to build up to proving theorems with airtight implications about real neural networks.
To be a bit more explicit:
Modeling. A model is some simplified partial representation of reality. Modeling is when you use a model to make predictions about reality.
The mathematics for A.I. safety community is particularly interested in building mathematical models for A.I. Within Iliad, these are often high-fidelity models of deep learning.
Modeling can be much more direct and practical than scrying, since fitting a good model let's you make useful predictions without inventing any new ideas.
However, one can certainly scry in order to look for better models, or use models in order to scry.
We do not have a formal theory of how scrying works. That means it is hard to evaluate the promise of a given scrying glass.
Information theory tells us that learning about X cannot teach you anything about Y, if X and Y have no mutual information. In practice, many things are slightly connected, but we can still form reasonable expectations about relative information value. For example, I would not roll dice to help me guess the outcome of the 2028 U.S. presidential election.
I do not know in what sense X must be related to Y, such that thinking about X is a good strategy for coming up with ideas about Y.
However, I'm confident that there are occupational risks of scrying, chiefly the deadly condition known as nerdsnipe.
Nerdsnipe
On an individual level, nerdsnipe is a cute and charming excess of (mis)directed curiosity, often displayed to signal nerdiness.
On an institutional level, nerdsnipe is a blight that stalls intellectual progress within entire disciplines.
The obsession with tabular MDPs, Nash equilibria, VC dimensions and generally i.i.d. learning for their own sake seems to have stalled academic learning theory and diminished its relevance to agent foundations. These are/were all useful glasses to scry on while searching for better ideas, but have also spawned (sub)disciplines with high-status open problems.
I think that nerdsnipe is a serious risk of scrying - if the scrying glass is too interesting in itself, you can forget that you're trying to learn about something else. This failure mode is virulent, an infohazard that repurposes institutional power structures to propogate itself far past the point that the resources of natural curiosity are exhausted.
Indeed, the agent foundations community is highly prone to nerdsnipe, of the sort "Let's take a quick detour to learn category theory, and formalize it all in lean!"
I would guess that mathematical modelers are a bit less prone to nerdsnipe than mathematical scryers. There's a clear external feedback signal on whether a model is successful, and models are not designed to stimulate the mind the way that scrying glasses are. This could be an interesting question to investigate.
A.I. for math
It looks like A.I. is about to automate mathematics in the sense that proofs become very cheap. I think it's an open question how much this will help with A.I. safety, but I will make some guesses about the benefits for mathematical scrying.
A binary provability oracle for mathematical statements might be useful for theory building (for example by checking which definitions have interesting theorems). However, binary yes/no answers to all of my existing conjectures probably would not accelerate my agenda that much, because the proofs and intuitions behind them are what give me ideas for the next round of conjectures.
A proof oracle would be substantially more useful. However, I worry that both myself and my research fellows would get nearly as much practice thinking in terms of algorithmic information theory by (only) reading proofs as we have by writing them.
A math textbook oracle would be a game changer. If I could simply ask for a high-quality textbook about any topic in mathematics, I would quickly develop entire new frames for thinking about alignment. For example, reading Arora and Barak's "Computational Complexity" substantially enriched my concepts for reasoning about intelligence, despite the fact that I (of course) invented none of the results therein and have never done much research in the area.
At the moment, LLMs act a bit like weak proof oracles. In the future, their frontier reasoning may become increasingly inscrutable as they are trained to generate lean proofs to resolve impressive problems, but we can probably also automate the translation back into legible English. So we may have slightly stronger provability oracles than proof oracles.
Eventually, there will be textbook oracles. However, writing a textbook about some question is an open-ended task where the best answer may be quite surprising and general. In some sense, the textbook oracle would need to avoid falling prey to nerdsnipe and internalize its own healthy mathematical scrying process. By the time that A.I.'s can mathematically scry about A.I. safety beyond human level, we are probably already dead.
These considerations are one reason that building a mathematical research lab as mathematics is increasingly automated has been a very weird experience.
At AIXI Labs
My favorite scrying glass is Solomonoff induction, which I use to get ideas about learning and generalization. Solomonoff induction is part of AIXI, which I use as a mathematical model of A.S.I. Sometimes, I apply this model to analyze A.I. safety schemes, which is partially to predict their failures and partially to scry for better ideas.
So, across multiple projects, I think about Solomonoff induction to scry, to study models, and as a model to scry with.
I would be quite interested in a model of scrying. Solomonoff induction seems poorly suited for this, since its "reasoning" is not bounded. Perhaps logical induction is rich enough?
Blue and Green
Scrying is certainly not the only way of looking for ideas. It's a particularly inward-looking method. Another option is looking out into the world, which I associate with original seeing,attunement, and perhaps naturalism. Scrying might be best known today as a magic: the gathering mechanic, and in that game/world's color wheel scrying is blue, while the other options above seem green.
Agent foundations is often criticized for being too abstract and disconnected from the world. However, empiricism as practiced by the frontier labs seems to me rather incremental and uncreative, recycling and refining old ideas rather than producing new ones. If agent foundations is too blue, I would like to see a more fertile and green approach.
Epistemic status: Exploratory thinking.
After attending ILIAD: Aeneid and talking with @Richard_Ngo, I've been thinking a bit about how to get ideas, particularly by doing mathematics.
In scientific inquiry, the true hypothesis often hasn't occurred to you yet. Worse, the truth might be too complex to hold in mind, so that any hypothesis you can consider must be incomplete. This is the type of situation that I believe Richard likes to think about; he claims that we do not have the right concepts yet to understand agency, and developing them is robustly beneficial for A.I. safety.
(But it's not always about truth. Sometimes you just need better ideas, because all of your options are looking doomed. Agent foundations is about trying to deeply understand agents, but conceptual A.I. safety research can be broader, also including the invention of devices to control agents.)
A.I. safety needs to invent better concepts and better ideas. I think that agent foundations has cultivated a particular way of doing mathematics which aims to inspire such creativity.
Why math?
At ILIAD, Eliezer questioned whether anyone's alignment agenda was actually bottlenecked on solving a math problem. ILIAD attendees do a lot of math nominally related to A.I. safety, but no one really has a story about how even a proof oracle gives you safe A.S.I. So what are we all doing, spending our time on math problems that don't tell us how to save the world?
I think we're mostly "scrying."
Scrying is staring hard at one thing, in order to try to find out about another. In the occult origins, it's often a mirror or glass, viewed under darkened conditions to inspire hallucination.
MIRI was mathematically scrying when they thought about probabilistic truth predicates, Löb's theorem, and logical induction. Their goal was not to directly use these results to build self-improving A.S.I., but to become less confused about embedded agency.
Today, I think that Resolution is basically scrying about the interactive proofs of complexity theory, hoping to get ideas that will work for natural language reasoning, which is called debate.
Scrying is not the only use of math in A.I. safety, it's the indirect use of math for hopefully stumbling upon A.I. safety ideas. Some math for A.I. safety is more direct than scrying, usually because it's modeling something. For example, @Guillaume Corlouer tries to understand how stochastic gradient descent on neural networks works. He might start with some simplified neural nets, but he seems to be trying to build up to proving theorems with airtight implications about real neural networks.
To be a bit more explicit:
Modeling. A model is some simplified partial representation of reality. Modeling is when you use a model to make predictions about reality.
The mathematics for A.I. safety community is particularly interested in building mathematical models for A.I. Within Iliad, these are often high-fidelity models of deep learning.
Modeling can be much more direct and practical than scrying, since fitting a good model let's you make useful predictions without inventing any new ideas.
However, one can certainly scry in order to look for better models, or use models in order to scry.
We do not have a formal theory of how scrying works. That means it is hard to evaluate the promise of a given scrying glass.
Information theory tells us that learning about X cannot teach you anything about Y, if X and Y have no mutual information. In practice, many things are slightly connected, but we can still form reasonable expectations about relative information value. For example, I would not roll dice to help me guess the outcome of the 2028 U.S. presidential election.
I do not know in what sense X must be related to Y, such that thinking about X is a good strategy for coming up with ideas about Y.
However, I'm confident that there are occupational risks of scrying, chiefly the deadly condition known as nerdsnipe.
Nerdsnipe
On an individual level, nerdsnipe is a cute and charming excess of (mis)directed curiosity, often displayed to signal nerdiness.
On an institutional level, nerdsnipe is a blight that stalls intellectual progress within entire disciplines.
The obsession with tabular MDPs, Nash equilibria, VC dimensions and generally i.i.d. learning for their own sake seems to have stalled academic learning theory and diminished its relevance to agent foundations. These are/were all useful glasses to scry on while searching for better ideas, but have also spawned (sub)disciplines with high-status open problems.
I think that nerdsnipe is a serious risk of scrying - if the scrying glass is too interesting in itself, you can forget that you're trying to learn about something else. This failure mode is virulent, an infohazard that repurposes institutional power structures to propogate itself far past the point that the resources of natural curiosity are exhausted.
Indeed, the agent foundations community is highly prone to nerdsnipe, of the sort "Let's take a quick detour to learn category theory, and formalize it all in lean!"
I would guess that mathematical modelers are a bit less prone to nerdsnipe than mathematical scryers. There's a clear external feedback signal on whether a model is successful, and models are not designed to stimulate the mind the way that scrying glasses are. This could be an interesting question to investigate.
A.I. for math
It looks like A.I. is about to automate mathematics in the sense that proofs become very cheap. I think it's an open question how much this will help with A.I. safety, but I will make some guesses about the benefits for mathematical scrying.
A binary provability oracle for mathematical statements might be useful for theory building (for example by checking which definitions have interesting theorems). However, binary yes/no answers to all of my existing conjectures probably would not accelerate my agenda that much, because the proofs and intuitions behind them are what give me ideas for the next round of conjectures.
A proof oracle would be substantially more useful. However, I worry that both myself and my research fellows would get nearly as much practice thinking in terms of algorithmic information theory by (only) reading proofs as we have by writing them.
A math textbook oracle would be a game changer. If I could simply ask for a high-quality textbook about any topic in mathematics, I would quickly develop entire new frames for thinking about alignment. For example, reading Arora and Barak's "Computational Complexity" substantially enriched my concepts for reasoning about intelligence, despite the fact that I (of course) invented none of the results therein and have never done much research in the area.
At the moment, LLMs act a bit like weak proof oracles. In the future, their frontier reasoning may become increasingly inscrutable as they are trained to generate lean proofs to resolve impressive problems, but we can probably also automate the translation back into legible English. So we may have slightly stronger provability oracles than proof oracles.
Eventually, there will be textbook oracles. However, writing a textbook about some question is an open-ended task where the best answer may be quite surprising and general. In some sense, the textbook oracle would need to avoid falling prey to nerdsnipe and internalize its own healthy mathematical scrying process. By the time that A.I.'s can mathematically scry about A.I. safety beyond human level, we are probably already dead.
These considerations are one reason that building a mathematical research lab as mathematics is increasingly automated has been a very weird experience.
At AIXI Labs
My favorite scrying glass is Solomonoff induction, which I use to get ideas about learning and generalization. Solomonoff induction is part of AIXI, which I use as a mathematical model of A.S.I. Sometimes, I apply this model to analyze A.I. safety schemes, which is partially to predict their failures and partially to scry for better ideas.
So, across multiple projects, I think about Solomonoff induction to scry, to study models, and as a model to scry with.
I would be quite interested in a model of scrying. Solomonoff induction seems poorly suited for this, since its "reasoning" is not bounded. Perhaps logical induction is rich enough?
Blue and Green
Scrying is certainly not the only way of looking for ideas. It's a particularly inward-looking method. Another option is looking out into the world, which I associate with original seeing, attunement, and perhaps naturalism. Scrying might be best known today as a magic: the gathering mechanic, and in that game/world's color wheel scrying is blue, while the other options above seem green.
Agent foundations is often criticized for being too abstract and disconnected from the world. However, empiricism as practiced by the frontier labs seems to me rather incremental and uncreative, recycling and refining old ideas rather than producing new ones. If agent foundations is too blue, I would like to see a more fertile and green approach.