I think this is missing the main reason to not devote more resources to the meta-problem.
One of the main patterns in alignment work is that people run into the hard parts of the problem, where it's hard to even think clearly at all or find any hint of traction, and then they mentally "slide off" to something which feels easier. I claim that "go meta" is one of the most common ways that people slide off. Problem is, the meta problem is not actually easier than the object-level problem. It feels easier, because it's further removed from concrete challenges and therefore easier to not-notice the ways in which one's vague abstract ideas don't match reality. But in actuality, if someone is struggling to even keep their attention on the object-level problems, they're not going to be able to make actually-useful progress on the meta problem.
I resonate with this.
Alignment felt hard, so I concentrated on community building, but one thing I've realised over time is that doing community building well doesn't remove the need to place bets on what solving the problem is likely to look like.
Figuring out how to best "place bets one what solving the problem will look like" IS exactly the meta-problem I'm talking about.
Community-building is not part of the meta-problem as I'm defining it, although I certainly think it's valuable work.
If you mean we've got to actually place the bets, yes and we're already doing that, just (to be a little provocative) without thinking about it all that carefully because we're not bothering to pay anyone to do it.
If your point is that individuals may be spending too much time on the meta-problem, despite not being funded, I agree. I think that's separate from whether we should fund more full-time work. I think that would actually address this sliding-off the hard parts issue.
I agree that there's a tendency to slide off the hard parts, including into the meta-problem. But the meta-problem definitely includes accurately identifying the hard parts, so that fewer people keep sliding off them in the future.
But the more common direction of sliding off the hard parts these days is into empirical work.
One of the main points of increased funding for the meta problem is to address exactly this failure to meet the problem at its hard parts. Good, focused work and publication on the meta-problem would make the hard parts more apparent.
You and others think the hard parts are apparent. Many prosaic alignment researchers do not agree. Better work on that disagreement would help resolve it. So it seems like we should fund that work.
So: (edit) would you really object to funding a few more people to work on the meta-problem? I would've thought you'd support funding some aspect of what I mean by the meta-problem, since that's how I read much of your contribution to the field. Your recent post on streelighting, and slop not scheming as the median doom path was two top-notch pieces of work on the meta-problem IMO. They're widely cited and guiding work in better directions, exactly what I mean by work on the meta-problem.
Hmm, I don't know that I agree with this statement:
> I think that we could probably get usefully better answers at relatively low cost.
I think it is easy to convince oneself that some work is better to do than other work, but it is hard to reach actually sound conclusions about "X is obviously the best thing to work on". The way you get the clearest signal is by trying the idea and seeing where reality pushes back on your efforts. Indeed, I'm on a team doing alignment research and we are constantly trying to "understand what work would solve alignment". This is just hard to do! (especially under the time crunch.)
Absolutely we need to guide work by empirically seeing what works!
I see little chance we'll stop doing enough of that. The suggestion here is just that we fund a little more abstract prediction/theory of what would be valuable if it works. A little more seems like a reasonable split, since it's currently a tiny amount relative to funding for empirical work.
What would you guess is the fraction? I'd say spending on the meta-problem is maybe on the order of .1% of total alignment funding?
To clarify: by "usefully better answers at relatively low cost" I didn't mean vastly better at trivial cost, just a good expected return on investment.
Most researchers agree that mech interp, refining and improving alignment training, control, theory, and miscellaneous techniques like confession are useful for solving alignment. Working toward regulation and slowdown/pause is also commonly considered useful in the governance space, and spreading awareness of alignment risks is pretty obviously useful for accelerating and enabling all of this work. But we don't know what variants of this or other work make best use of our limited time and funding. Which of them, or what others, are the most efficient use of resources? I think that we could probably get usefully better answers at relatively low cost.
Currently work on the meta-problem is done either in researchers' spare time, during grant applications, or inside funding orgs and AI developers. All have limited incentives to publish their thinking legibly, and incentives during grant applications and inside dev orgs are subject to Goodharting and motivated reasoning: sounding good vs. being good. More funding directly for analysis and planning seems like an efficient leverage point.
The AI Futures Project is the exception that demonstrates the trend. They are funded to do prediction work, but that spreads into many aspects of the meta-problem. They make Gears-Level Models of how alignment and governance work might succeed or fail, and so identify work that's likely to be more and less useful. They have the time and focus to explicitly include the many cross-dependencies between governance, public opinion, and technical alignment work, in contrast to people who do this in their spare time and mostly focus on their own area of expertise. (Other orgs and people do aspects of this work, but arguably do less direct work on the meta-problem as I'm thinking of it.)
How could we even make progress on the meta-problem? For example, we had Anthropic claim that not even Claude Mythos 5.1 provided a 2x acceleration on their R&D projects, while Kwa demonstrated disbelief back in July, moved to OAI to construct RSI evals and Ngo ended up claiming that "measures and models of RSI are the kinds of thing which are near-optimal for propagating "feeling the AGI" inside OpenAI" and that Anthropic would have gone twice as fast if it propagated feeling the AGI and OAI would have quadrupled the acceleration rate. In circumstances like these, I would expect that the main technical meta-level work is to resolve as many object-level disagreements as we can.
The main governmental meta-level work seems to be either already done (e.g. AI-2040's Plan A and the doorstoppers on strategy, out of which I would throw away the sloppier parts) or consist of finding the best strategy for awakening politicians and/or promoting the awakened ones onto relevant positions (think of Bores' failed campaign to become the congressperson).
Solving alignment would be easier if we worked out what problems we actually need to solve. This could be called the alignment meta-problem. Work on this problem is rarely directly funded. More focused work on it should let us use our limited time and funding more efficiently.
The diagram implies narrowing alignment work, but I expect meta-problem work to also identify high-payoff "fringe" approaches.
If we're driving toward a cliff, maybe we should buy better headlights.
All too often we're doing work that merely sounds or feels good, and optimizing less than we could for work that drives most efficiently toward success. Some of this is inevitable and some of it is useful, but we could do more to light the path ahead.
Most researchers agree that mech interp, refining and improving alignment training, control, theory, and miscellaneous techniques like confession are useful for solving alignment. Working toward regulation and slowdown/pause is also commonly considered useful in the governance space, and spreading awareness of alignment risks is pretty obviously useful for accelerating and enabling all of this work. But we don't know what variants of this or other work make best use of our limited time and funding. Which of them, or what others, are the most efficient use of resources? I think that we could probably get usefully better answers at relatively low cost.
Currently work on the meta-problem is done either in researchers' spare time, during grant applications, or inside funding orgs and AI developers. All have limited incentives to publish their thinking legibly, and incentives during grant applications and inside dev orgs are subject to Goodharting and motivated reasoning: sounding good vs. being good. More funding directly for analysis and planning seems like an efficient leverage point.
The AI Futures Project is the exception that demonstrates the trend. They are funded to do prediction work, but that spreads into many aspects of the meta-problem. They make Gears-Level Models of how alignment and governance work might succeed or fail, and so identify work that's likely to be more and less useful. They have the time and focus to explicitly include the many cross-dependencies between governance, public opinion, and technical alignment work, in contrast to people who do this in their spare time and mostly focus on their own area of expertise. (Other orgs and people do aspects of this work, but arguably do less direct work on the meta-problem as I'm thinking of it.)
Many researchers and orgs do some of this work despite it not being directly funded, because they think it's so valuable. And people at funding orgs do meta-problem work as part of their job. But publications of funders' models tend to be short and limited (although kudos to the funders who do carefully work out and publish their theories-of-change!).
Why not to fund more work on the meta-problem
There are good reasons not to fund this type of work more heavily. In brief:
These are all important points. Each is something to guard against, or a reason to not over-fund the meta-problem.
Arguments in favor, compressed
Each of the arguments against suggests an argument for:
This post is a teaser for a longer, more worked-out draft that's been stuck in my queue. If responses here change my mind or indicate that No One Cares, it will probably stay unfinished. This is pretty telegraphic; I can elaborate, give examples, and clarify in the comments.
All questions, feedback and suggestions are appreciated!
Disclosure: Much of My research has recently been on aspects of the meta-problem. This is a result of increasingly thinking such work is even more neglected than the object-level technical alignment work I started out doing. I feel very lucky to be able to devote time to this type of work.
Disclaimer: I'm not deeply familiar with all the orgs, and I'm sure I've overlooked orgs funding overlapping work. Apologies to those I've missed.
Thanks to Peter Gebauer and others for many conversations touching on this claim, and Peter for comments on an earlier draft. This framing was inspired when a donor friend said "I might fund more alignment work if I could tell what would be useful," causing me to say "Maybe we should be funding figuring that out!"