What does it mean to teach AI Safety?
On teaching AI Safety
Quick history of some of my AI safety involvement and takes:
- 2020 : read LW & sequences, yes this seems important, will get to it at some time (for the classic EA reasons of it's the most important good), but was looking for my first job in France and took some software engineer position in a startup. I continue reading and upskilling on AIS during those 2 years.
- 2022: quit my SWE job to go into AI safety. reasoning: yeah this is not a problem for the future, this is coming soon. I meet some people who have <5 year AGI timelines. I don't know that it's true, but it's worth considering since it's the very start of scaling and it could have been possible intelligence scaled much faster with scaling than it ended up in practice scaling. I work on AI Safety field building and infrastructure. There are then still <100
- 2023 : Still doing AIS field building, but looks like bottleneck is governance, I mostly keep pitching people around me to do comparatively more governance. (over time, this is one of the contributions that leads to the formation of CeSIA, the French Centre for AI Safety)
- 2024 : Involvement with a wider range of ideas from the field has generally updated me more towards the OpenAI and Anthropic and Deepmind positions that prosaic alignment for human level AGI is technically feasible and rather tractable, and that this is indeed a very important input to how we most likely will progress on ASI alignment. I systematically distinguish between AGI risks and ASI risks and clarify the cruxes and assumptions for different threat models and am annoyed that MIRI threat models often skip to the end-game without consideration that certain trajectories towards that end-game falsify important assumptions (see Superintelligence of the Gaps for some elaboration there). I still believe governance is the bottleneck though, and continue contributing ideas that feed into https://www.lesswrong.com/posts/EexsebbYhbe2gXkPP/the-current-bottleneck-is-political-will-not-research
- 2025 : start the year with a burnout after some intense governance work. Ideas wise not too much change with 2024, still doing field building and teaching AI safety theory occasionally, but also take a long break and learn more widely from other wisdom traditions (eg. buddhism, tpot and postrat stuff).
- early 2026 : man does getting good US AI governance fast look intractable given the current administration, at this point it seems better to accelerate good safety work within the AI companies and in the surrounding ecosystem. I broadly wanna contribute to us surviving in the world where we don't get much of a pause, where governments aren't that competent and coordinated. I would work for any of the AGI companies on safety if found an adequate role. I think short timelines to human level AGI (eg. 2028) is plausible and preparing for if algorithmic improvement doesn't asymptote too shallowly is important. Governance might not affect this in time. Being in the room where it happens seems likes the highest leverage way to increase the probability of good ai futures. A bunch of my theory of change is just helping the people in the room where it happens be more wise.
I would worry about the "be in the room" strategy. It seems like most people who justify their career decisions that way wind up getting captured by the groupthink of the org. They think they'll be the ones who can resist it, but they won't.
To the extent that an org does have a "groupthink" and its own theory&value system, then whether a given person in the room adheres to it seems to depend on :
1) the selection effect of who wanted to be there, in that particular org vs others
2) the discussions with other people in the org changing that person's views
3) systematic pressures, eg. greedy/selfish parts of them optimizing to continue getting revenue
If I joined Anthropic and 1 year from now people thought I had surprisingly Anthropic-like views, I'd guess it's mostly because of 1) and 2). 2) happens a lot but is broadly good. 3) is the one that's mostly bad, from the outside/civilizational point of view, and the prior should be most people are susceptible to this, but this can be updated away from seeing particular life accomplishments. In my case, I have enough history of independence, selflessness and moral upstandingness that I don't think 3) will influence me substantially, but I don't recommend this path to those without that history.
Here’s an example of 3) happening to someone, and they noticed it: https://forum.effectivealtruism.org/posts/rHyAmvXiqrC9iAR9T/jay-bailey-s-shortform
I would really suggest reflecting on it if you haven’t already. “I am special and can resist the groupthink” is often a false belief.
As for 2), I don’t think it is necessarily good? If most people there have 1) and 3) influencing them, that will filter the kinds of opinions they have which then get transmitted to you in conversation, and now even if you are stalwartly resisting the direct pull of 3), it’ll still reach through others to pull at you.
Not to mention, workers at frontier labs seem to be doing a fourth thing, delegating increasingly large amounts of trust and thinking to their AIs, in ways which might be troublesome; you would be signing up for this. It can happen indirectly, even if you don’t do it yourself, because others will launder AIs’ beliefs as their own.
Separately, there’s the issue that leadership at the labs simply have their own beliefs about various important issues, and don’t care for the opinions of the rank and file. Anthropic defanged the RSPs it arguably drew in many researchers with; OpenAI let two alignment teams wither; DeepMind sold out to the military over its employees’ objections. The explicit goal of these labs is RSI, and the first workers they want to unemploy are their own, especially their juniors. The remaining employees at late stage AI labs will mostly be a core of leadership and senior researchers whose research taste is still required. What sorts of impactful decisions would you be able to meaningfully influence in the window between signing on and obsolescence, if any?
Thanks for your comments. I don't expect doing an analysis of my situation in particular is best use of our time but I do think these are helpful questions to consider for people in my situation or similar.
Re your last paragraph, I'd happily bet that Anthropic has not reduced their workforce 2 years from now. Yes relative employee disempowerment is an important factor I care about, but it is precisely in worlds where alignment is not that good that having humans in the loop is important (in an obvious seen-by-leadership way). It is only reasonable to automate everything with very very high trust in both the competence and alignment of AI systems, and Anthropic as a company is not that unreasonable, they definitely do find and classify many Claude behaviors as undesirable, and will continue doing so.
There's a usual back and forth about how much to distrust leadership of AGI companies which is hard to ground in material fact. Some people take the lack of safety actions now to mean lack of care for when it will matter, but conversely the fact that it never mattered yet is a good reason for them not to have cared for these inconsequential things. The explanation for defanging the RSP is a good one, I don't think people should tie themselves to masts and go blind into the unknown unknowns of AGI development. They should build capacity to remain aware, capacity to pause, have institutions that can do independent audit and have real power to stop them, but not fixed RSP-like stuff.
Finally, still on last paragraph, the "window between signing on and obsolescence" is very dependant on people's rates of growth, but also where they can work immediately. I am generally glad that Joe Carlsmith joined Anthropic to help with the Claude Constitution, I think he immediately is having very significant impact. There is much object level work to make the chances of better futures to be done. Even if one later gets automated, having made alignment that much better before full automation could be a significant difference.
I would like if LessWrong provided an optional newsletter like the EA forum digest, for people who want the chance to catch non-curated posts without having to open LessWrong and sift through it directly every few days.
Here's what the EA Forum digest looks like : a list of titles + author + time to read.
I don't know exactly how much manual curation goes into it and I'm not asking for that. I'd find a simple karma threshold and this format valuable.
I am also writing up this quick take notably because I've had discussions with other people who'd like this, and because recently costs of development and maintenance of software like this have gone down.
I don't find the existing RSS feed a preferable alternative.
- I have never setup an RSS feed reader or similar process, I don't think I want to and guess most LW readers are similar.
- It seems it would on top of that would take extra work to get the format I want out of it
This Feb 2026 survey of some AI safety leaders found median timelines of 2033 for the following definition of AGI
An AI system (or collection of systems) that can fully automate the vast majority (>90%) of roles in the 2025 economy. A job is fully automatable when machines could be built to carry out the job better and more cheaply than human workers. Think feasibility, not adoption.
It featured the following comment
“I think >10% of roles in the 2025 economy are either manual or otherwise require human-like bodies: construction, barbers, restaurant server, etc. If we restrict to knowledge workers (roughly, jobs that can be done on a laptop), these dates move even closer.”
On the current paradigm, AI capabilities progress on niche tasks and diffusion will be linked[1] and diffusion can go rather slowly even when tools are incredibly productivity enhancing, thus there could be an intuitively surprisingly large gap between automation of 50% human tasks[2] and 90% and 99%, true even if we restricted the prediction to computer work tasks.[3]
I'm 80%+ confident we get automated expert+ level coding and ml research by 2030, and that there will be a significant amount of low hanging fruit in software/algorithmic space to allow fast progress on all tasks for which we have data, but I believe generalisation will stay somewhat limited (very very far from "figure out gravity from a picture of a bent blade of grass, more like "when speaking to a human expert in a niche field, knows how to interview them over 10 to 100 hours to extract most important info and then be mostly autonomous on known tasks, but still needs feedback from reality to learn more"), aka ~human level generalisation at best up to 2031.
The combination of "need feedback from reality" and slow diffusion makes slower timelines to "superintelligence" (eg. better than all humans at 99.99%+ of 2026 tasks) surprisingly plausible (eg. 5 to 10 years between AGI and ASI, thus ASI by 2040). I guess without a pause/significant politically influenced slowdown, we'd 80%+ have ASI by 2040. I'd set my 50% for ASI around 2036.[4]
I think technical alignement for human level AGI is solvable and not even off track, thus the world will look fine/good in 2030 (few to zero severe power seeking and deceptive misalignment problems in deployment from Anthropic AI systems) but have high uncertainty about the "use ai to do ai safety work" plan allowing us to successfully know how to train aligned ASI within five years of that. Overall I place myself at 10% or less p(doom) from sharp left turn risks, but around 40% all things considered p(doom) by including gradual disempowerment/value drift and societal response.
We need people to be deploying the technology to gather the relevant data to train/learn from, because generalisation is limited and because lots of expert knowledge only exists in human minds and structures of human relationships right now.
Note I'm weighing by "meaningfully different task" rather than "frequency of task". Given power law distributions most tasks might be "read email/slack, respond", which computer use will know how to operate, but not be able to respond to intricacies of different work situations.
Because computer work often involves using domain expert knowledge to do the right things on the computer.
I haven't researched robotics enough to know how fast we could produce and deploy 100 million humanoid robots worldwide which seems like an appropriate level of effort required to gather the required data.