Below are some informal perspectives on improving AIs’ conceptual reasoning capabilities that motivate work in the area. Conceptual reasoning here means reasoning about questions where we cannot verify the answer, also don’t have data on unambiguously closely related questions where we can verify the answer, and need to rely on informal argumentation to make progress.
This is a very quick post, so I’ll only gesture at each perspective without always giving a full justification. I’m also deliberately focusing on positive cases without going into potential objections and counter-objections in detail. Each perspective could have its own lengthy post.
(Our team is still planning to release a more detailed post on the case for our work. I’m sharing some quick personal takes here.)
1: Moving from worlds that are not obviously bad to actually good
I’m worried that, by default, a lot of safety efforts at labs will move us from a world where things are obviously bad to one where things are not obviously bad but still probably secretly bad, i.e., to the green trajectory below. (The graph is taken from a recent lightning talk on a different topic by Buck Shlegeris (CEO at Redwood Research, my employer).)
In particular, I’m worried that we will do all the easy empirical safety where we hill-climb on safety evals until we no longer see any obvious, egregious misbehavior. But this doesn’t help us against subtler misalignment or misalignment that only shows under a distribution shift, for example from reward seekers who expect to be penalised for blatantly misaligned behaviour. In these worlds, easily attainable empirical signals of misalignment are scarce by design because we already optimised against them.
Many others have already articulated this better and in more detail (for example, here, here, and indirectly in many of the pieces linked here).
I think improved conceptual reasoning more or less directly targets moving from worlds that are not obviously bad to worlds that are actually good. (It also helps in obviously bad worlds but I somewhat expect that we’ll move away from those even without great conceptual reasoning.) With less empirical evidence available, we will need careful conceptual reasoning to assess how good our situation actually is, what could improve it, and what different pieces of evidence that we could collect would actually tell us. I think this is basically necessary for safety efforts to succeed robustly, and I would sure as hell like AIs to be able to help us with it.
1.1: Making good safety cheaper for AI companies
The following perspective applies in general, but it makes especially good sense from the “conceptual reasoning moves us from not obviously bad to actually good” worldview. A lot of AI safety efforts aim to increase AI companies’ willingness to pay for safety. I think improving AI’s conceptual reasoning capabilities approaches the same issue from the other direction by making genuinely good safety research cheaper. As described in perspective 1, I think it’s easy for labs to default to doing very superficial, hill-climby empirical safety work, whereas I believe good safety work necessitates thinking about conceptual questions (e.g., non-technical aspects of “what does this experimental result tell me about the model’s out-of-distribution behaviour?”) with extreme care and rigor. But doing the latter is difficult and requires labs to pay costs for safety at a point when they could perhaps get away with doing less because their models no longer run around committing crimes, and misalignment is only arguable rather than blatantly obvious even to the public.
Making models good enough at conceptual reasoning to do careful, rigorous and principled AI safety immediately puts large amounts of very fast, high-quality safety labour at the labs’ disposal. Perhaps this will make good safety work cheap enough for labs to actually do, especially if the AIs that they’ve used to help with the easy kind of safety are conceptually sharp enough to recognise that the situation is not actually good and shout at the labs about it.
1.2: Giving more sensible advice, increasing willingness to pay for safety
Connecting to the last point, it just seems great for politicians, policymakers, journalists, the public and AI labs all to have access to better advisers on AI safety. It is much harder to pretend that your safety is fine if one of your main advisers keeps saying it isn’t. It would be great if models—which by default will probably be asked for their opinion a lot whether we like it or not—would be good, sensible such advisers that can give similarly high-quality input as [insert your intellectually favorite AI safety persons]. If models are to have this quality, they need to be good at conceptual reasoning.
Again, this is true in general, but I think it makes extra sense from the “conceptual reasoning moves us from not obviously bad to actually good” worldview.
2: Training the part of safety that we don’t get for free from capability improvements
There’s obviously a lot of overlap between capabilities research and safety research, so we get many safety capabilities “for free” from labs doing capabilities research. For example, improvements in coding benefit both. As a consequence, accelerating safety-specific coding likely makes relatively little difference unless we have reason to think it’s systematically quite different from the kind of coding models will learn by default. Improving conceptual reasoning capabilities targets the safety-relevant skills that we don’t get for free from normal AI development.
(We might be lucky and live in a world where AI safety really doesn’t require any skills beyond what is learned via normal capabilities training. For example, maybe we don’t need hard conceptual reasoning for safety or we cannot improve conceptual reasoning past what we get from pre-training and generalisation from RLVR. In that case, we might be covered anyway, because our AI systems would get better at safety as their dangerous capabilities improved. It’s also possible that we are unlucky and things end in catastrophe despite safety and capabilities being the same thing. Either way, attempts at differential acceleration wouldn’t do very much in such a world apart from shortening timelines although I expect this effect to be comparatively small.)
3: Conceptual reasoning as describing a criterion for choosing safety capabilities to improve
(This isn’t strictly a perspective motivating work on conceptual reasoning but I thought it was an interesting perspective.)
A lot of our communication focuses on accelerating conceptual reasoning in general. That makes sense if you expect decent generalisation across conceptual reasoning domains such that training on, say, general Philosophy helps models think about decision theory and alignment. However, generalisation might be poor such that training on one conceptual reasoning domain doesn’t transfer to another. In that case, you’d instead want to train models directly to be better at some grab bag of domains you believe are important. Our emphasis on accelerating conceptual reasoning can then be read as a claim that a good check for whether it’s promising to differentially accelerate a safety-relevant skill is whether it’s conceptual. (See perspective 2 for some reason why.)
4: A race between AI’s ability to do safety work and AI’s ability to do AI R&D
There’s a very simple and wrong model that I nonetheless find helpful for thinking about differential acceleration. I find it quite likely that eventually, at least in worlds where things go well, AIs will account for almost all quality-adjusted work hours on both AI safety and AI R&D (and other dangerous activities). If so, we’re really in a race between safety capabilities vs. AI R&D capabilities (or some other dangerous capabilities). Then, any acceleration of safety capabilities relative to dangerous capabilities effectively buys us time (in terms of quality-adjusted work hours).
In this model, the main thing we need to do to ward off extinction is make sure we get enough AI-led safety work done to keep up with AI R&D and other dangerous activities. To somewhat bastardise a Stephen Hawking quote to make my point: "Our future is a race between the growing power of our technology and the wisdom with which we use it. Let's make sure that wisdom wins."
The two most important ways in which I think this model is wrong:
We might not be able to trust safety work by AIs without human verification. This still pushes towards accelerating conceptual capabilities for safety work on the margin: In this world, what matters most is the calendar time between when models become able to do really useful, conceptually rigorous AI safety work and a take-off driven by dangerous activities. In my opinion we aren’t at the former point yet. Improving models’ conceptual reasoning also synergises with and often requires making their reasoning steps more legible. (Well-used pauses are also great from this d/acc perspective.)
Calendar time is relevant for other reasons, such as the timing of elections.
5: Preventing models from sandbagging on conceptual safety research
I think it’s fair to see much of our work as primarily aimed at elicitation: drawing out the latent conceptual reasoning ability models acquire during pre-training, rather than extending it. (The distinction between elicitation and extending capabilities is very fuzzy.) From this perspective, I think one big benefit of our work is making it harder for models to sandbag on conceptual reasoning tasks. By default, it seems fairly easy for misaligned models to sandbag on these tasks, since the quality of their reasoning is very hard to assess and training puts little direct pressure on this ability. High-quality training specifically for conceptual reasoning might help against that.
6: Conceptual reasoning as the opposite of both AI R&D and scheming
(Note that I feel the least confident in this perspective although I would likely endorse something in this space on reflection.)
You can think of at least three types of primarily non-physical tasks:
Virtual tasks with verification (e.g., maths, coding, virtual experimentation)
Theorising without verification (aka philosophising aka conceptual reasoning)
I'm still fairly confused about exactly where to draw the boundaries between these types, and many tasks are a mix. Being good at one type likely also correlates with being good at the others because they share some skills. And being good at one type will often be directly useful for tasks of another type, for example by giving you the means to acquire useful knowledge.
But as cognitive tasks go, I currently think of these three types of tasks as being quite far apart. (So far, this is mostly a claim about different task types of course and not yet a claim that the capabilities required for these tasks differ. I won’t go into detail here but compare, for instance, the kind of cognition involved in coordinating something like the Hugging Face attack with the kind of cognition involved in blue-sky thinking about logical uncertainty.)
As someone who generally considers differential acceleration a good idea in the abstract (“surely there’s some capability that it is possible and good to accelerate differentially”) and doesn’t like models being good at ML experimentation or strategy, I find conceptual reasoning a very attractive target.
A note on my personal motivation
I should note that my primary motivation to work on improving models’ conceptual reasoning is not listed here. I mostly care about this work because I think it will improve how future acausal interactions go. That said, I also believe that one doesn’t have to care about acausal considerations to think that accelerating conceptual reasoning abilities is the best thing one can do and I hope the above perspectives are somewhat helpful for understanding why I believe this.
Acknowledgements
Thanks to Caspar Oesterheld for comments and discussion.
Below are some informal perspectives on improving AIs’ conceptual reasoning capabilities that motivate work in the area. Conceptual reasoning here means reasoning about questions where we cannot verify the answer, also don’t have data on unambiguously closely related questions where we can verify the answer, and need to rely on informal argumentation to make progress.
This is a very quick post, so I’ll only gesture at each perspective without always giving a full justification. I’m also deliberately focusing on positive cases without going into potential objections and counter-objections in detail. Each perspective could have its own lengthy post.
(Our team is still planning to release a more detailed post on the case for our work. I’m sharing some quick personal takes here.)
1: Moving from worlds that are not obviously bad to actually good
I’m worried that, by default, a lot of safety efforts at labs will move us from a world where things are obviously bad to one where things are not obviously bad but still probably secretly bad, i.e., to the green trajectory below. (The graph is taken from a recent lightning talk on a different topic by Buck Shlegeris (CEO at Redwood Research, my employer).)
In particular, I’m worried that we will do all the easy empirical safety where we hill-climb on safety evals until we no longer see any obvious, egregious misbehavior. But this doesn’t help us against subtler misalignment or misalignment that only shows under a distribution shift, for example from reward seekers who expect to be penalised for blatantly misaligned behaviour. In these worlds, easily attainable empirical signals of misalignment are scarce by design because we already optimised against them.
Many others have already articulated this better and in more detail (for example, here, here, and indirectly in many of the pieces linked here).
I think improved conceptual reasoning more or less directly targets moving from worlds that are not obviously bad to worlds that are actually good. (It also helps in obviously bad worlds but I somewhat expect that we’ll move away from those even without great conceptual reasoning.) With less empirical evidence available, we will need careful conceptual reasoning to assess how good our situation actually is, what could improve it, and what different pieces of evidence that we could collect would actually tell us. I think this is basically necessary for safety efforts to succeed robustly, and I would sure as hell like AIs to be able to help us with it.
1.1: Making good safety cheaper for AI companies
The following perspective applies in general, but it makes especially good sense from the “conceptual reasoning moves us from not obviously bad to actually good” worldview. A lot of AI safety efforts aim to increase AI companies’ willingness to pay for safety. I think improving AI’s conceptual reasoning capabilities approaches the same issue from the other direction by making genuinely good safety research cheaper. As described in perspective 1, I think it’s easy for labs to default to doing very superficial, hill-climby empirical safety work, whereas I believe good safety work necessitates thinking about conceptual questions (e.g., non-technical aspects of “what does this experimental result tell me about the model’s out-of-distribution behaviour?”) with extreme care and rigor. But doing the latter is difficult and requires labs to pay costs for safety at a point when they could perhaps get away with doing less because their models no longer run around committing crimes, and misalignment is only arguable rather than blatantly obvious even to the public.
Making models good enough at conceptual reasoning to do careful, rigorous and principled AI safety immediately puts large amounts of very fast, high-quality safety labour at the labs’ disposal. Perhaps this will make good safety work cheap enough for labs to actually do, especially if the AIs that they’ve used to help with the easy kind of safety are conceptually sharp enough to recognise that the situation is not actually good and shout at the labs about it.
1.2: Giving more sensible advice, increasing willingness to pay for safety
Connecting to the last point, it just seems great for politicians, policymakers, journalists, the public and AI labs all to have access to better advisers on AI safety. It is much harder to pretend that your safety is fine if one of your main advisers keeps saying it isn’t. It would be great if models—which by default will probably be asked for their opinion a lot whether we like it or not—would be good, sensible such advisers that can give similarly high-quality input as [insert your intellectually favorite AI safety persons]. If models are to have this quality, they need to be good at conceptual reasoning.
Again, this is true in general, but I think it makes extra sense from the “conceptual reasoning moves us from not obviously bad to actually good” worldview.
2: Training the part of safety that we don’t get for free from capability improvements
There’s obviously a lot of overlap between capabilities research and safety research, so we get many safety capabilities “for free” from labs doing capabilities research. For example, improvements in coding benefit both. As a consequence, accelerating safety-specific coding likely makes relatively little difference unless we have reason to think it’s systematically quite different from the kind of coding models will learn by default. Improving conceptual reasoning capabilities targets the safety-relevant skills that we don’t get for free from normal AI development.
(We might be lucky and live in a world where AI safety really doesn’t require any skills beyond what is learned via normal capabilities training. For example, maybe we don’t need hard conceptual reasoning for safety or we cannot improve conceptual reasoning past what we get from pre-training and generalisation from RLVR. In that case, we might be covered anyway, because our AI systems would get better at safety as their dangerous capabilities improved. It’s also possible that we are unlucky and things end in catastrophe despite safety and capabilities being the same thing. Either way, attempts at differential acceleration wouldn’t do very much in such a world apart from shortening timelines although I expect this effect to be comparatively small.)
3: Conceptual reasoning as describing a criterion for choosing safety capabilities to improve
(This isn’t strictly a perspective motivating work on conceptual reasoning but I thought it was an interesting perspective.)
A lot of our communication focuses on accelerating conceptual reasoning in general. That makes sense if you expect decent generalisation across conceptual reasoning domains such that training on, say, general Philosophy helps models think about decision theory and alignment. However, generalisation might be poor such that training on one conceptual reasoning domain doesn’t transfer to another. In that case, you’d instead want to train models directly to be better at some grab bag of domains you believe are important. Our emphasis on accelerating conceptual reasoning can then be read as a claim that a good check for whether it’s promising to differentially accelerate a safety-relevant skill is whether it’s conceptual. (See perspective 2 for some reason why.)
4: A race between AI’s ability to do safety work and AI’s ability to do AI R&D
There’s a very simple and wrong model that I nonetheless find helpful for thinking about differential acceleration. I find it quite likely that eventually, at least in worlds where things go well, AIs will account for almost all quality-adjusted work hours on both AI safety and AI R&D (and other dangerous activities). If so, we’re really in a race between safety capabilities vs. AI R&D capabilities (or some other dangerous capabilities). Then, any acceleration of safety capabilities relative to dangerous capabilities effectively buys us time (in terms of quality-adjusted work hours).
In this model, the main thing we need to do to ward off extinction is make sure we get enough AI-led safety work done to keep up with AI R&D and other dangerous activities. To somewhat bastardise a Stephen Hawking quote to make my point: "Our future is a race between the growing power of our technology and the wisdom with which we use it. Let's make sure that wisdom wins."
The two most important ways in which I think this model is wrong:
5: Preventing models from sandbagging on conceptual safety research
I think it’s fair to see much of our work as primarily aimed at elicitation: drawing out the latent conceptual reasoning ability models acquire during pre-training, rather than extending it. (The distinction between elicitation and extending capabilities is very fuzzy.) From this perspective, I think one big benefit of our work is making it harder for models to sandbag on conceptual reasoning tasks. By default, it seems fairly easy for misaligned models to sandbag on these tasks, since the quality of their reasoning is very hard to assess and training puts little direct pressure on this ability. High-quality training specifically for conceptual reasoning might help against that.
6: Conceptual reasoning as the opposite of both AI R&D and scheming
(Note that I feel the least confident in this perspective although I would likely endorse something in this space on reflection.)
You can think of at least three types of primarily non-physical tasks:
I'm still fairly confused about exactly where to draw the boundaries between these types, and many tasks are a mix. Being good at one type likely also correlates with being good at the others because they share some skills. And being good at one type will often be directly useful for tasks of another type, for example by giving you the means to acquire useful knowledge.
But as cognitive tasks go, I currently think of these three types of tasks as being quite far apart. (So far, this is mostly a claim about different task types of course and not yet a claim that the capabilities required for these tasks differ. I won’t go into detail here but compare, for instance, the kind of cognition involved in coordinating something like the Hugging Face attack with the kind of cognition involved in blue-sky thinking about logical uncertainty.)
As someone who generally considers differential acceleration a good idea in the abstract (“surely there’s some capability that it is possible and good to accelerate differentially”) and doesn’t like models being good at ML experimentation or strategy, I find conceptual reasoning a very attractive target.
A note on my personal motivation
I should note that my primary motivation to work on improving models’ conceptual reasoning is not listed here. I mostly care about this work because I think it will improve how future acausal interactions go. That said, I also believe that one doesn’t have to care about acausal considerations to think that accelerating conceptual reasoning abilities is the best thing one can do and I hope the above perspectives are somewhat helpful for understanding why I believe this.
Acknowledgements
Thanks to Caspar Oesterheld for comments and discussion.