Seems like the fundamental problem is that humanity's strategic competence is extremely inadequate relative to its technical capabilities, and this (limited) strategic competence was front-loaded into AI risk (the most strategically competent people foresaw the problem and went into the field early, before AI was even a thing for everyone else). As AI risk awareness grows, the average strategic competence goes down as more and more people enter the field. (Consider the gap between someone who "are going to OpenAI without even having considered x-risk for five minutes", and what's actually needed to make a positive contribution or avoid making a negative contribution.)
This seems clearly true and like an interesting frame, but I'd have to think more about it to agree with the full conclusion. It seems to me not clear that average strategic competence is a thing to care about. I think there's a good case that average people with very low agency still have an instinctual respect for/deference to people who have higher agency, and that our natural/institutional hierarchical structures support this stratification. Most problems seem to be solved like this, if imperfectly? Organize people into systems where small numbers of outliers with high strategic competence are deferred to by larger numbers of people with only technical competence.
To me it seems like we have a crazy problem where the so-called highly strategically competent people are not able to capture the highly technically competent nerds from academia.
Note that I never used "highly strategically competent"! If we did have a group of highly strategically competent people, then what you say becomes a lot more plausible, but even the early AI risk people are far from that. Consider for example that a strategic consideration as obvious (in retrospect) as Legible vs. Illegible AI Safety Problems was written down only in late 2025.
During the 2012-2014-ish era, I recall a fairly common statement from MIRI being "you might think AI alignment is a computer programming problem but it's more of a math problem", and there was some attempt to carve it into fun little chunk for Math people.
Then MIRI mostly gave up on solving it's technical agenda in time. I'm not sure what the state of things are in terms of "what problems are still open and seem to matter?".
Seems probably worth someone putting in legwork to organize the state of the problem into something comprehensible to the math world. I vaguely am thinking of @Max Harms maybe being a good fit for this?
Curious for braindumps from @abramdemski @Scott Garrabrant, @TsviBT, @johnswentworth and @Alex_Altair on the state of thing heres.
(I've been "out of the game" for a couple years.) I would append to
you might think AI alignment is a computer programming problem but it's more of a math problem
a follow-on:
you might think AI alignment is a math problem but it's more of a philosophical / conceptual / mental-phenomenology problem
(I hesitate to use the word "phenomenology" because it will definitely be misunderstood by ~everyone, but there just isn't a better word; an internal phrase that was floated was "core theory of mind", and I mean to gesture at "mathematico-introspective investigation of core theory of mind".) This is elaborated on (cryptically / elliptically, but you could read slowly & think) here: https://www.lesswrong.com/posts/TNQKFoWhAkLCB4Kt7/a-hermeneutic-net-for-agency
In this context, the long and short of it is, I don't think you can feasibly point people at the parts of the problem that matter. I'm just such a special genius that only I could possibly understand the real alignment problem For some reason people don't seem to be inclined to plant difficult questions and let them grow over time, sit with unresolved questions, look at big things and keep their bigness firmly in mind while also somehow making cumulative progress, overhaul key elements or at least keep them firmly provisional, keep staring outside of the streelight's glow for years, and so forth. (Something something Grothendieck something something Peter Scholze something something staring at conceptual foundations.) Or maybe it's the thing about security mindset (Yudkowsky), or overconfidence (Dai), etc. Or maybe we're just not high g enough. Or all of the above or something else, IDK. But again, it just doesn't seem to work to point people at the problem. Or maybe I'm deluded, who could know.
If you take all the alignment research, insofar as I'm aware of it, in the past 2 decades, and multiply that by 3, it still doesn't even come close to solving alignment. This is very rough and grim, but it's what I think.
I'm not sure I buy that most mathematicians will be out of a job. Maybe some of the number theorists, analysts, and combinatorialists? Or something, IDK, I'm making that up. (Based on what fields vaguely seem to have some niches for people doing mostly high-algebraicness work.) Or maybe they will be.
What should they do instead? IDK. Human intelligence amplification is still insanely underresearched, and can definitely use more brainpower. I'd be very happy to be connected with smart motivated scientists / mathematicians who may want to work on empowering humans!! My gmail for that would be: tsvibtcontact
To be clear, when writing this post, recruiting any mathematicians to alignment work was much more of a stretch goal.
The low-hanging fruit I'm trying to point to is that there are droves of academics currently considering going to work at frontier AI labs without having even heard the basic arguments for x-risk. I think it would be much much easier to convince a substantial fraction of them to do NOT THAT, than to direct them to do any active technical alignment work.
Having mulled this slightly more... I don't actually have a very clear idea on what would work here (in the sense of "be memetically successful enough to reach a meaningful number of people.")
I think there's a moderately effortful job here of "figure out how to communicate about this in terms the math community will be motivated by and engage with" which can only really be done by someone in the math community.
I'm curious, when you imagine doing "the obvious thing", what ends up being hard or not working? (presumably there's a reason you wrote this post targeted at LW users than some other post targeted at math people?)
The "obvious thing" that I have done is state my views on doom to people when it comes up and point them to IABIED, and I think it has largely been successful in causing ~20 people I have personal relationships with to arrive at p(doom) > 5%.
The next "obvious thing" is to scale this up by a factor of 10 to people I'm only vaguely acquainted with, and to write more publicly in a mathematician-facing way as you suggested. These feel both costly, risky, and less effective in ways that the previous thing was not. I am still considering doing it, but need to sit down with my aversions for a while. Some examples: these conversations sometimes require me to be forceful/activist to keep people in my frame instead of sliding off into more comfortable topics, and they plausibly only work because I have significant personal rapport already with a colleague or student.
I would certainly prefer to live in a world where 1 in 20 mathematicians are already awake and I can just create common knowledge that the private conversation thing works, and then not do anything more effortful. As far as I know the only other academic mathematician who writes about Doom is Jacob Tsimerman, who just won the Fields Medal and left to do safety at OpenAI. He would certainly be a person who could do the broadcasting thing (and is doing it already, but in a less forceful way than I'm imagining).
(This is probably not helpful, but just want to note that I'm interested, in a "hobbyist" sense, in the problem of the general kind of confrontation you're mentioning, under the moniker "confrontation-worthy empathy"; if you happen to want to chat about it I'm interested, e.g. to consider different ideas for making the process more wholesome and similar.)
I would be interested in having this conversation! It seems to me that the kind of confrontation I'm mentioning is a much easier version of the thing you're pointing to? I tend to just point amateur levels of unconditional positive regard at people and under pre-conditions of existing trust conversations mostly go well. I imagine scaling up the operation requires more finesse, but even so it's hard to imagine needing the amount of skill you describe there. I'll message you.
Yeah, I think it should be much less hard for people not already embedded in working on AI; just seems kinda related.
Yeah fair.
Partly, I am kinda assuming people have a harder time wrapping their brain around "what not to do" vs "what to do", and I think you'll be more successful with that goal if you present a positive vision of how to orient to x-risk that speaks their language.
But, yeah that is a pretty different framing that would be suspicious if it didn't look different at all from where I was pointing.
Partly, I am kinda assuming people have a harder time wrapping their brain around "what not to do" vs "what to do"
IDK, there being problems that are hard to fix but easy to make worse seems like a sort of common sense situation? Though I'm struggling to think of any good examples off the top of my head. Also just "we aren't ready for the machine god, don't help them build it faster" is super simple?
I guess math very much isn't like this. You can generally make at least some progress on a problem, but you can't really make a problem worse except by giving bad advice to others working on it?
I think there's something especially bad about AI, where people disagree a lot or are confused about what exactly the problems are, and have a lot of motivated cognition towards thinking "if they're working on it, it'll go better instead of worse."
(i.e. if people were right that the main problem is misuse, going to work at a lab is more reasonable)
There's something additionally hard about the existential stakes, where it's hard to admit the problem is real if you don't feel like you have some way of engaging with it.
E.g. Musk seems like mega-biased toward taking action, with bad results (responsible for like a quarter of frontier AI lab stuff, in some sense). The politician's syllogism bites real hard when accelerating bad stuff is 1. convergent and 2. much easier than accelerating good stuff.
"Please read IABIED" is my first thought, but sounds like you've probably tried that? Maybe there should be a TLDR aimed at mathematicians? We could ask Yudkowsky to write one?
I think recommending a book just doesn't work without a long preamble or some unusual social leverage. I've had better success recommending online articles, my current go-to is Scott's book review of IABIED, which is still tl;dr and has a bunch of missing context.
The main thing a good first-intro has to do, honestly, is to be as short as possible while explaining instrumental convergence, orthogonality, and countering some representative examples of crazy moon logic that people come up with initially.
Have you tried AISafety.info's intro sequence? There's also a shorter stand-alone article, but that one is written for the average person.
I'd be curious if you (or anybody else) want to test to see if my intro is superior[1]. I weakly think it would be, but I'm pretty unconfident and there are good reasons for me to share my guide with people I know over other people's guides regardless of objective merit.
I deliberately wrote the intro to avoid many mistakes I saw in other intros (eg too many extraneous details, appealing to overly complex analogies, shadowboxing insider objections, or "going meta" in an intro post).
eg by randomizing my intro vs Scott's review to send to different people.
My sense is many of these problems are either ill-defined or too hard to be tractable.
In certain fields like computability theory most problems are intractable, just because programs are very complicated and diverse objects that are hard to prove things about. Progress in such fields is made by working in the areas of the field where there is enough structure. Unfortunately proofs over programs have featured in agent foundations since the beginning (eg tiling agents).
As for ill-defined problems, ontology identification, embedded agency, and decision theory are full of them. Eg finding something that behaves like counterlogicals, which are nonexistent objects. Because they are nonexistent objects it requires philosophical progress to make a list of properties they need to satisfy. This doesn't mean it's impossible to make progress, but the problems need to be formalized in a way that are tractable and don't lose all their relevance to AI.
A specific thing I'm curious about: there was a strategy of "turn some open alignment problems into nice little math problem to nerd-snipe people with." Did that strategy accomplish anything? (i.e did it successfully get someone who would have bounced off the general conceptual arguments to engage enough to seem to eventually get the conceptual argument? And/or just produce obviously useful work)
The MAIS (Math for AI Safety) repo is my new effort to draw mathematicians into AI safety.
It started as a survey paper Math for AI Safety: An Invitation for Mathematicians. During the writing of that paper, Claude and GPT generated 100+ pages of open problems organized in eight research agendas. I'm releasing that material as a public hub for mathematicians to collaborate on AI safety problems.
The repo just went live today, and it's very much a work in progress! I'd appreciate any feedback on how to make it better.
Hi Lionel! I remember your name from doing some research on chip-firing and rotor-routers in high school, many lifetimes ago. I didn't know this existed and am glad it's a thing!
We've been shouting about approaches that could use mathematical research talent for several years, which don't have huge downsides such as those pointed out by Wei Dai. If mathematicians are interested, they should look at, e.g., the work ARIA has been doing on mathematical foundations, at work and the research agenda by Vanessa Kosoy, and at the new work being launched by Resolution.
So far, the big breakthroughs are coming from strong professionals asking LLMs to prove truly important results.
While the amateurs do get empowered quite a bit, right now the situation rewards high competence.
I will say that Fable + Codex & on-demand VMs honestly feels superhuman, lacking only the "spark of genius" heuristic (heuristic set?). The problem search spaces are nevertheless very large, so brute-force simply isn't an option without some clever perspective.
I think also that people underestimate the utility of partial results and we need a way to verify and document these to avoid wasted and duplicated work.
I get that change is scary, but this is honestly the most exciting time to be alive as a mathematician. even if all LLM use were banned tomorrow, each of these discoveries would still be talked about in 100 years. And, capabilities-wise, this is the worst it is ever going to be.
given that AI-associated x-risk largely depends on our level of mathematical competence, this is also arguably the most important time in history for the social role of mathematicians.
AI-associated x-risk largely depends on our level of mathematical competence
So far, mathematics have been used mainly to make the global situation worse, i.e., by enabling AI development to proceed in a way that maximizes capabilities with almost no regard for how easy it will prove to maintain control or alignment of the capabilities.
It would take a lot to persuade me that that trend will reverse any time soon.
I will concede that doing mathematics is probably good training for alignment research because it teaches mental moves that are useful outside mathematics, but I wouldn't recommend mathematics as the majority of the training time.
I say this as someone who really enjoys doing mathematics.
given that AI-associated x-risk largely depends on our level of mathematical competence, this is also arguably the most important time in history for the social role of mathematicians.
Why do you think this?
I think controlling AI x-risk requires insight into the behavior of complex mathematical objects. To be a little less fuzzy, we need a working theory of mind. It is our mathematical competence which will determine how much control we have here.
I think RSI (a.k.a. self-improving AI autoresearchers) is inevitable for a wide variety of commercial, military, and/or geopolitical reasons, no matter what treaties are signed, no matter what promises are made. At least the NSA will do it. This is a curse in that it speeds the timeline. This is a blessing in that it gives us a chance. Again, though, the key to bounding the alignment delta between versions is going to be a theory of mind and its associated complex mathematics. If the delta is unbounded or even just too large, we lose and die. If the delta is small, we have a chance.
I think just waiting, or rather just researching, allows successive model versions to drift apart in mind space and makes it more difficult to bound the alignment delta. An all round catastrophe. We need to cognitively enhance each mathematician we have and put them to work developing (the mathematics behind) a theory of mind and/or on their favorite alignment subproblem.
given that AI-associated x-risk largely depends on our level of mathematical competence, this is also arguably the most important time in history for the social role of mathematicians.
Strongly doubt that.
My view is that we probably aren't anywhere remotely close to knowing how to align ASI, and also that the current people don't seem to be capable of making progress, so we should be working on augmenting human intelligence instead while halting AGI research. So the two things I'd want to suggest they work on are:
(I believe this is essentially the position of IABIED)
Both of these endeavors could absorb a lot of brain power, maybe especially the augmentation. I'd expect mathematicians to have an appreciation for how much humans vary in cognitive ability, and have benefited from high cognitive ability themselves; maybe this helps with pitching them on how valuable it would be to augment human intelligence.
Obvious followup to this: would you be interesting/willing to work on human cognitive enhancement? If not, why?
On Copenhagen Interpretation of Ethics grounds I precommitted to not responding to queries like this, or else it would look like anyone who tried to solve any problem had an obligation to consider every related problem.
Fair enough. I guess I want to get a sense of how hopeless an endeavor what I've proposed above is.
My intuition is that “the smart shot“ (whatever that is) plus a neural lace in an exponentially expanding number of people looks very similar to an RSI, except that a datacenter is unneeded, the nodes of the ASI have human rights, and it will never actually happen because it requires human experimentation, which would be banned.
As a mathematician currently in the middle of this transition (leaving an academic math-phys postdoc for AI safety work), two data points that might be useful here:
easy thing to do: commit to voting against ai datacenters and convince at least one other person to also do so - on the argument of extinction
extra: commit to voting against ai datacenters and convince at least one other person to also do so and convince them such that they also want to convince at least one other person - on the argument of extinction
I think there's a critical opportunity for someone here.
Mathematicians are feeling the doom (mostly in the "lose our jobs" sense).
Academics are freaking out about the daily news that amateurs are asking GPT "prove career-defining theorem, make no mistakes" and it's just working.
Senior researchers are leaving for frontier AI labs. Many are mentally spiralling or flailing about their life's work not mattering anymore.
Last year, when I tried explaining IABIED to colleagues, I would be met with incredulous stares. This year, I'm met with incredulous stares and "So what should I do now?" (I don't have a good answer for them, which is part of why I'm posting this.)
I've quarantined AI discussions on my research discord because otherwise it would overwhelm everything else.
Top mathematicians, people on par in mathematical ability with Critch, Christiano, and Steinhardt, people who have been running leading-edge research groups for decades, people with enormous soft power in academic circles, are going to OpenAI without even having considered x-risk for five minutes.
Many more will be leaving soon.
If you want these folks to hear something at all, to consider some other option in the rest of their life, now is the time.