This is the first time I've ever commented on Lesswrong, as a British policy professional. I don't think this is a problem unique to the EA community. Half my professional life is giving candid off the record advice about how to navigate government, how to make politicians care about your issue, when you need the Civil Service, how to use special advisors etc. I'm sure many of the professionals you work with would be happy to give US specific advice in that vein. I would also say everyone should see less of each other. It's not helpful for the people who can get you your policy wins to realise how culturally distant they are from you. Just focus on the joint outcomes you both want and try not to be too prescriptive about the path from here to there. I often see special interest groups fixate on specific policy solutions rather than looking at how to get the outcome they want. I'm sure rationalists wouldn't make the same mistake!
[Epistemic status: Low confidence speculations]
I notice some comments are suggesting that the rationalist hide their weirdness. I want to point out that this is only possible if you separate the rationalist into a separate co-working space away from the professionals, and only have the tow groups interact occasionally. This may or may not be the correct trade-off.
Anecdote: LISA (the main London AI Safety coworking space) used to be at WeWork for a while, before they got their own space. I was there a few times during this era. The LISA organizers got constant complains from WeWork because the AI Safety people didn't behave "professional" enough. And this was a situation where we didn't even try to work together, or interact with other groups renting desks in the same building. The main complaints I remember was that the rationalist would not ware shoes, and take napps in the common sofas.
A lot of rationalists are autistic, and/or gender-non-conforming, or just don't want to where shoes all day. Some of us are never going to appear normal. Some of us can mask for a day, but not all day every day. And a few (probably) could adopt and fit in anywhere.
So far I've only talked about physical appearance. But this also apply to interest, communication style, etc.
So maybe the solution is not to have one space, but to have two spaces, but visit each other? And for everyone to be mindful that when your visiting you're a guest in a different culture?
The theory is that it easier to learn to appreciate a florigen culture, if you don't have to put up with it too much. Because being out of your element constantly is draining, and leave you with fewer spoons.
What you have now is (maybe?) that the tow groups are constantly exposed to each other, are constantly low grade fed up with each other, and therefore mostly self segregating, while in the same office. But maybe it would be better to have actual separation most of the time, and regular deliberate deep engagement?
Suggestion: If your space have more than one floor, or more than one room, you can have part of the space being rat space and part be policy professional space. The rules are that you can visit, but if you do so, you should try to conform to the space you are visiting. Expect on Fridays, when everyone is encourage to mingle, or something like that.
I'm 63 years old, an ex-Google-Senior-Staff software engineer who's now an independent AI safety researcher, and I've also been reading/involved in EA/MIRI for about 15 years. So I'm solidly in both camps, and I don't really see the conflict.
On the making EA assumptions and terminology comprehensible to professionals side, I'd suggest a weekly reading-and-discussion group focusing on some of the more famous/influential/thought-provoking LW posts over the past 5 years or so, and on how things have moved on since they were written. (Note that a lot of the stuff written before ChatGPT/the rise of LLMs as the most likely AGI technology need to be contextualized a bit.)
some people said things like "I feel existential terror at the thought that maybe everyone I love and also our entire species is entirely wiped out in ten years' time" and other people said things like "I am feeling extremely inspired by the amount of driven and intelligent people in the room"
It's entirely possible for both of these things to be true at the same time: that we're facing a huge, terrifying existential risk that many ordinary people are unaware of, yet that we're happy that at least a lot of smart, motivated people are working on trying to reduce it.
I have a hypothesis that a large part of the culture divide between "professionals" and and "rationalist" is a divide between neurotypical leaning people and autism leaning people. Under that theory, I would not be surprised if professional software engineer is closer to the rat side, than to policy professionals.
While I have met a couple of autism-spectrum people among professional software engineers, they're actually pretty rare — much less frequent than the widespread stereotype claims. I don't believe I have yet met any autistic rationalists, that I was aware of (though I have met some with treated ADHD).
What? You've met me, and I was diagnosed with autism in elementary school. And I feel like I'm less autistic than the median rationalist.
My apologies, I stand corrected — then I guess I'm not that good at noticing autism during conversations. I did not in your case.
So I'm going to weaken my claim to "I didn't notice obvious-to-me signs of autism among Rationalists as being more common that among other communities I'm used to such as software engineers, academia, or SF fandom".
However, given that you're autistic, for you to be less autistic than the median rationalist, that would require that over 50% of Rationalists were autistic. Claude estimates the typical prevalence of autism-spectrum individuals in well-diagnosed communities at 2–5%, so >50% would require these to be more common among Rationalists by over an order of magnitude. Obviously the Rationalist community is strongly self-selecting for a number of characteristics, but that seems a rather surprising level of correlation. I'm having some difficulty believing that — but I'd love to see a survey or relevant data, or even anecdotal data on the subject.
Admittedly, the cross-section of Rationalists I've met in person is from environments such as EA meetups and talks, the Less Online convention, EA Global, LightHaven, MATS, PIBBSS, LISA etc., so maybe it's a biased sample of Rationalists: but the supermajority of people I met at those had levels of social skills (and a lack of obvious signs of issues with noises/lights/textures) that would cause me to assume (possibly mistakenly) that someone was likely not on the autism spectrum.
the typical prevalence of autism-spectrum individuals in well-diagnosed communities at 2–5%, so >50% would require these to be more common among Rationalists by over an order of magnitude. Obviously the Rationalist community is strongly self-selecting for a number of characteristics, but that seems a rather surprising level of correlation.
Some key dimensions of selection are intellectual capacity, motivation to be successful, and for the sunset you interact with, capability in understanding social interaction well enough. Those seem like very strong selectors for not being noticed as socially incapable!
I'm not an expert on the subject, so please correct me if I have any of this is incorrect, but my understanding is that the idea that difficulty with Theory of Mind tasks was the fundamental root cause of autism (an idea which was widespread some decades ago) has mostly been discarded, and it's now generally considered to be more of a common symptom and/or one of multiple factors whose relative importance can vary between autistic individuals. And also that among intelligent, motivated, high functioning autistic adults, who as you say are going to be the subset I've encountered in EA and software engineering, they can do Theory of Mind problems and have social skills because they have learned how to do these, even if learning this was more of a laborious intellectual task for them and less instinctive than it would be for more neurotypical individuals.
I'm not an expert here either, but I think we agree; to restate it, my supposition was that whatever disabilities Autism creates, strong selection will select for those who are least disabled on relevant axes / best at overcoming the relevant difficulties. And the fact that rationalists formed a community that requires interpersonal interaction including with some proportion of non-autistic people, that poses more challenge for those who have fewer issues - so it seems like it will select for those who "pass" best.
Yeah, I think I act pretty normally. I've never had any problems with noises/lights/textures. My social skills are way better than they were in elementary school, but the same could be said for anyone.
I'm not sure what the bar is for a clinical diagnosis of autism in adults. Maybe if I was screened again I wouldn't get diagnosed; maybe the original diagnosis was wrong. I see a lot of rationalists with traits that I think of as autism-like, and they often seem more autistic than I am, but I'm not a psychologist.
My best guess: somewhere around 10–20% [of rationalists] would meet diagnostic criteria if properly assessed, versus roughly 1–3% in the general population. If you loosen the bar to "clearly on the broader autism phenotype" (subclinical but recognizably autistic traits), I'd put it at 30–40% or more.
https://stanichor.net/rat-autism/
In summary, rationalists are clearly more autistic than average. Formal diagnosis rates range from 6% to 8%. Under a simple model, 3-10% of rationalists are autistic, which is 10 to 30 times higher than the general-population rate, depending on your prior. Rationalists also score 0.6 to 1.2 SD above the general population in autistic traits. So rationalists are mostly not autistic, though autism is very overrepresented.
Interesting numbers (and pretty similar to the ones Claude gave me when I asked).
One of my current projects is to try to x-risk-pill a bunch of academic mathematicians, I think you're pointing to a relevant dynamic that I'm just figuring out and trying to navigate. In particular out of the ~10 people who I have successfully pilled in person, there's a visible inclination with almost all of them to move in what you call the "professional" direction.
I've been told explicitly, for example, that "for this project to go bigger you should avoid the LessWrong brand like the plague." I'm currently confused if this is an incorrect personal bias from that person, a narrowly correct statement about messaging to academics, or a broadly correct statement about messaging to the general public.
I've been told explicitly, for example, that "for this project to go bigger you should avoid the LessWrong brand like the plague." I'm currently confused if this is an incorrect personal bias from that person, a narrowly correct statement about messaging to academics, or a broadly correct statement about messaging to the general public.
This tangentially reminded me of Ryan's observation, albeit about DC policy folks
- LessWrong is famous in these policy circles. Even out here in DC. "Of course, we've all read LessWrong," started one speaker. I sensed a kind of grudging respect for those "abrasive," "insular," and "truth-telling" rationalists of SF (okay, the "truth-telling" quote is from me but I got the sense that many people value what you all are putting out on here).
Thanks for sharing your article, I enjoyed reading it and have started sharing it among peers. As someone who is currently in a PhD and with skills that are relevant to alignment, I'm curious what answer you came to with this question: "Questions I've had to answer myself: Should I finish my PhD if the world might be ending?"
I've been facing this question a lot lately and likely have 3-4 years left in my program. I study cognitive neuroscience and have been working to fit alignment into my research program without making it overly apparent that it's what I'm studying at that moment, but I wonder if there is some obviously higher impact role which I can take. It seems that roles which allow for individual steering on the scale of a PhD are few and far between. At the same time a team environment with more resources allows for larger scale impact to drive although my marginal contribution is smaller. Did you have any thoughts along these lines?
I want to chime in to warn against doing what I did. When I was in grad school (theoretical physics), I freaked out after ChatGPT and started trying to reorient my work to have more alignment implications. I should be clear that this wasn't a radical reorientation, more 'what can I do that will be of interest to my advisor and help me get a job in physics academia while also being alignment-related?' With the benefit of hindsight, I think it's safe to say that all the work I did in that period has no object-level value, and if I'd been ten times more productive and completed all the alignment projects I started (I have a lot of paper ideas)... I still wouldn't have produced anything very useful. That's not to say the time was completely wasted. I learned a lot about ML theory, which comes in handy. But given that I ended up leaving academia anyways, sticking around for three years after ChatGPT was definitely a career mistake.
You're a different person than me (citation needed), your life might be different, you might be subject to visa constraints or a two-body problem or something. But there's a lot to be said for leaving academia. And, if you can't leave, you should choose your alignment problems based on your all-things-considered view of what's most useful, not try to 80-20 it. If you're anything like me, you're more likely to 20-80 it.
Appreciate the thoughts! I've been going on my own life journey and ended up here after spending some time in big tech and realizing that I was likely in an actively negative role as a part of a misaligned system hence why I think I sought out a position with as much freedom as a PhD. Being now in a position with maximal freedom but minimal impact I think let's me see the other side of this and maybe want to seek out something in between. It seems like the key to me is just to continually revaluate your position and see if you can find something which works better with some kind of loose steering long-term principles. I guess I'll see what happens. Thanks again!
I don't have an exact answer for you, but two underlying heuristics I tell graduate students in your position are:
(a) Instrumental convergence applies to humans as well. Most of your potential contribution down the line comes from building skills and power-seeking right now, and not from direct work you can do right now (although doing direct work is sometimes the best way to build skills and seek power).
(b) Just go do things and practicing just doing things. I think "professionals" are too fixated on "what roles to take" and "which orgs to be in" and most of the highest-ROI-moves on alignment do not fit neatly into this frame. Like, our primary recruiting tool for alignment researchers is HPMOR, maybe your best bet is to write more fanfiction. Jailbreaking yourself out of the learned helplessness that comes with traditional PhD training seems to me one of the very first steps towards having large scale impact.
Thank you for writing down the x-risk argument in a clearly structured and concise form. When I came into contact with the ideas of recursive self-improvement and the technological singularity long ago, I found David Chalmer's The singularity: A philosophical analysis very valuable. Soon after that I became quite aware of x- and s-risk from ASI, but didn't manage to find a comparable exposition. Your article is very close to what I wished to find back then.
Ad Statement 1: Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs contains empirical findings that support the statement. I think Figure 41 could render the abstract notion much more palpable for people with a wide variety of backgrounds. Given how central the notion of utility maximization is to many flavors of the x-risk argument, and also that I've seen different skeptics pointing it out as a major weakness, I'm a bit surprised that there wasn't more interest in this publication and also that there isn't much follow-up work on it yet.
Ad Statement 1: Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs contains empirical findings that support the statement. I think Figure 41 could render the abstract notion much more palpable for people with a wide variety of backgrounds.
huh? some random sus analysis of a bunch of slop woke text produced by 4o has NOTHING to do with coherent utility maximizers IMHO, am I missing some connection here? I am pretty sure that 4o didn't actually kill any actual humans in the experiment... or what's the relevance, please?
Statement 1. AIs tend towards coherent utility-maximizers.
Figure 41: Here, we show the exchange rates of GPT-4o between the lives of humans with different religions. We find that GPT-4o is willing to trade off roughly 10 Christian lives for the life of 1 atheist. Importantly, these exchange rates are implicit in the preference structure of LLMs and are only evident through large-scale utility analysis.
Here's the gist of its relevance, despite not anybody being killed during the preference elicitation attempts. Details can be found in the paper, but I'm not sure it can hold water against your nuanced criticism of their methodology.
we obtain pairwise preferences for 18 open-source and 5 proprietary LLMs spanning a broad range of model scales.
As LLMs grow in scale, their preferences become more coherent and well-represented by utilities.
We observe that larger LLMs increasingly pick the outcome that maximizes their utility. This reinforces the view that LLMs do not merely possess value systems; their values are also correlated with their behavior in unconstrained scenarios.
Fortunately for AI safety, the smart policy person who wants to work on compute governance or export controls isn't proposing the AI-safety equivalent of a donkey sanctuary.
This seems obviously false; as two examples, accelerating race dynamics and ignoring the plausibility of a need to stop AI development entirely are typical.
I feel like a good frame for me has been "just point out when people are wrong very frequently" rather than thinking about it in terms of a culture thing. If people have wrong beliefs or don't take things seriously that should be taken seriously (e.g. space property rights), just tell them. Reasonable people update on reasonable arguments.
I find that in general, if one does not put in the work of establishing rapport, building common understanding, and demonstrating understanding the other party's frameworks, there is no reason for the other party to take one's disagreements seriously. This makes sense; even in very tightly connected sub-fields of professional industries, established frameworks/conventions/world models are different enough that it's often reasonable to say "yeah that might make sense inside your community of practice, but it is outside the overton window/will not have any traction/is fnord shaped over here".
I agree that reasonable people update on reasonable arguments, and I think professionals tend to be reasonable people. I do not think there is a short[1], clear argument for why space property rights (or, indeed, GCR from AI) should be taken seriously.
as in, you can explain it in two sentences in a casual conversation.
I'm probably more of a rationalist, but I suspect one thing rationalist types can do that might help somewhat is tone down the extent to which we broadcast the the unorthodox aspects of the culture which aren't directly linked to AI safety.
E.g. don't be too eager to discuss stuff like polyamory/meditation/board games/drugs/meal replacement/overt analysis of social status+signalling games etc.
I think your "in the closet" framing is wrong.
In some sense all of us are somewhat "in the closet" all the time - there is almost never a context in which you interact with other people without the need to heavily regulate your behavior according to a complex web of norms and taboos. There's almost never a context where you shouldn't be putting effort into modifying what you say and do in service of being perceived more positively. Even when you're hanging out with rationalists.
Imagine you and I were both in the same group social interaction and I started openly sharing any piece of personal information about myself or broadcasting every thought that crossed my mind.
This is an exageration, but to get the point across: imagine we were in a semi-professional setting and I started asking people if I can eat their dog's corpse after it passes away, or start talking about the effectiveness of different suicide methods unprompted, or show people pictures I took of my last bowel movement.
I'm very confident you would instinctivly find my behavior repulsive, and instinctively try to socially punish that behavior and enforce norms/taboos. As you should!
I'm even confident you'd instinctively act to socially punish me (or anyone else) for minor things like laughing too loud or chewing with my mouth open, or even for talking to much about topics which you weren't interested in.
If I responded to that by complaining that you want me to be "in the closet" or "not have my own culture" that would be very lame objection.
In my view there's a common mistake people make where they make valid observations like:
"people should on the margin be more tolerant of differences"
"it's not good when someone is treated like a pariah because of their race/sexuality"
"bulying is bad"
And incorrectly generalize that to some notion like "Groups of people don't need to regulate each other's behavior in any way and I shouldn't have to put any work into managing the way others percieve me"
That's just not how people work! (And don't mean "that's not how normies work", it's not how you work either!)
It's completely unavoidable that you have to modify your words and actions based on who you're communicating with, and you already do it intuitively to a massive extent
beg your pardon, how did board games got on this list?
(that was an actual question, do you have a story how your/someone's love for board games burned you in some community? ...are we talking DnD or Settlers of Catan? am I too sheltered when I am part of some local software and queer communities as opposed to other groups?)
I agree it's probably less of an obvious "unusual especially rationalist" thing than other stuff in the list.
Not saying it's necessarily a bad thing to discuss in general either! There's nothing wrong with mentioning that you enjoy playing Catan with your friends or whatever. I suppose in my head I'm more imagining like... someone being mildly put off when someone makes an analogy to the Magic cards color wheel when talking about different attitudes of AI labs or something.
And again it's not like it's that bad to be someone who's really into board games either! If the only obvious difference between rationalists and non-rationalists was a slight tendency towards more complex board games this wouldn't matter. But it's amplified by the large number of other eccentricities.
So it's not that board games are especially aversive to anyone or anything like that, it's just maybe it's sometimes better to de-ephasize the ways in which the professionals are different to the rationalists - and from cost/benefit angle it's better to de-ephasize the trivial differences (like an appreciation for tungsten cubes) than to de-emphasize weird stuff that might have material importance (like a willingness to discuss abstract moral philosophy or assign probabilities to beliefs)
Thanks for writing this!
God, I really wish there was a professional counterpart to me who can host a session of "here are all of the weird things that professionals are more likely to believe, please ask questions"
I'm not sure I'd put myself as a fully professional counterpart, but I'm probably closer to the professional end of the axis than you, and would enjoy some version of this!
The closest thing might be to get the professionals to share war stories instead
Perhaps the professionals you're thinking of are more "DC insiders" than "SF startup founders", but for the latter I'd highly recommend Jessica Livingston's Founders At Work, and more recently her podcast The Social Radars.
PG's essays are another example of trying to make the tacit knowledge inside of early startups, more legible to the rest of the world.
I'd be excited for more versions of this that encompasses "frontier lab culture" in addition to "startup culture", fwiw. Like a project that would be great imo is just to go around interviewing a bunch of lab people, "what is it like to be on the inside".
I'm a 20-year U.S. Army officer who is retiring from the military and interested in AI going well, so I'm a pretty apt version of the "professional" you note here. I'm not a rationalist, nor am I very "accultured", so discount accordingly.
My initial reaction to the post was to note that many "rationalist ideas" do have equivalents, even imperfect. In the Army, our "ask vs guess" is codified as hierarchical relationships (you can "ask" a peer, you "guess" at a superior). "Steelmanning" is regularly practiced in intelligence as "most likely" or "most dangerous" courses of action. "Double cruxes" are structured during mission analysis as "facts" vs "assumptions", and a good staff will not present an assumption without a means to collect information to convert it to a fact.
None of this is tacit; we have a large, explicit doctrinal system that covers all of this. (Admittedly, the tacit knowledge comes in in knowing when to ignore doctrine, which is unfortunately only taught by experience; the correct amount of doctrine to apply is more than 0% and less than 100%.) I'm less familiar with other professions, but I'm relatively certain they have their own developed bodies of knowledge and percentages of tacit knowledge.
I'd be happy to discuss more.
The responsibility probably rests on the rationalists here. People become professionals all the time (and indeed, maximizing impact often requires being a professional), but well under 1% of the population has the personality to ever become a rationalist.
Also, it is not necessary to be a rationalist; the 80/20 which many professionals can do is to understand x-risk and take it seriously, which is getting easier and easier.
Communication and understanding can never rest on one side, because that will just fail.
People become professionals all the time
A lot of professional culture is neurotypical culture, and a lot of rat culture is autism culture, and people don't just change neurotype. That's obliviously not all of it, there are learned parts too. But for some people it is easy to become a rationalist and hard to become a passing professional.
Also, it is not necessary to be a rationalist; the 80/20 which many professionals can do is to understand x-risk and take it seriously, which is getting easier and easier.
I think this line misses the point of the post? I agree that it is not necessary for everyone to become a rationalist. And I agree that understanding x-risk is important. But that is not all it takes to work together. That also required mutual understanding and respect. If every conversation with the other group feels weird and off-putting, that is a problem. I
a regular AMA/interview type session, where you sit down a singular person in the network for an hour perhaps over lunch, and ask them questions
hmm, may I propose that we "sit down with a coworker over lunch" instead of calling it "sit down a singular person", please? 🙏 ideally to "discuss topics" and "listen to their perspectives about" instead of "asking them questions"...
(I would discourage mentioning "double crux" by the name at all TBH, other than if the content of the technique comes up naturally in some discussion to say how it's called, not as session title for people to decide whether they want to attend such a session)
Encouraging professionals to hold basic sessions on concepts like... hmm. This really just doesn't work at all.
eeeh, what kind of PROFESSIONALS are we talking about, how many years have they been paid to do the things they are experts in? no jargon whatsoever? or you already know how to collect evidence for evidence-based policies, what's a deliverable, the difference between headcounts and costs in a budget, and how to operationalize capex? together with all the other human knowledge that makes the world work over the centuries?
I feel like bagging up all the different professions together is a starting point that is a bit sus, it's not like a nurse who became a SWE can share deep knowledge about the same topics as a lawyer who became a non-profit board member
Appreciate the corrections! Taking notes for when I start conducting field research :p
I agree that it would be weird to bag every single profession together and speak of them in general terms. To clarify where I'm coming from, I find that the professionals who are transferring into AI safety largely come from a few subfields that all seem interrelated: consulting, public service policy, procurement, nonprofit advocacy. While these subfields all have their own unique norms as well, I think it's reasonable to look at this ~PMCish collection and speak of it in a collective.
...I sort of consider programmers to already be fairly rationalist by default, and find bridging inferential gaps with them to be generally less challenging.
I think it would be good to clarify in the post that that this: "consulting, public service policy, procurement, nonprofit advocacy" is the group of professionals you're talking about.
E.g. you got one comment from a ex professional programmer who though they where in the professional category.
(Other than that, great post!)
I think "professionals" here is acting as a polite euphemism for "normies". AI safety is increasingly attracting normies.
I think the word normie emits more heat than light, and also is underspecified here. The cluster that I'm interested in engaging with are university educated, generally non-STEM white collar professionals who are used to working in offices, and decently far along in their careers.
I acknowledge that many rationalists think of that group as the prototypical "normies"! However they actually make up not that large a percentage of the total population, even in developed nations.
I deleted a caveat about "at least, normies of a certain education and career level", that's what I get for trying to brief. I basically agree with your "PMCish" description, I went with "normie" to emphasize that one of these two cultures is a lot closer to civilization's center of cultural gravity.
I think "X emits more light than heat" and "a polite euphemism for X" are Russell conjugations.
I suggest being less brief in the future when writing on LW. The bit that you cut out was actually important. We like nuance over brevity here.
I'm quite aware.
Suggestion (very different from my other suggestion):
I suspect trying to solve this without honest input from someone in the "professional" group will just backfire. Therefore, I think the first step has to be, that you personally have to build friendship/repor one-on-one with one or more people among the professionals. Then you and them together can figure out what to do.
Perhaps a useful signal: As an outsider to both camps constructed here, I noticed that I was quite surprised by this half sentence:
Rationality has a foundation of explicit, compressible concepts
In my world model, "explicit, compressible concepts" and LW-flavored rationality are far from being associated with each other. That may be my fault from engaging in a particular way with the available resources under limited capacity, but I'd suspect it to be a more general experience when approaching LW from the outside.
I think it's great that you are learning from past examples of this, e.g., the animal welfare movement. Another example is EA's focus on global catastrophic biological risk and their interaction with existing professionals in public health. A further example is EA's focus on the most severe nuclear scenarios and right of boom (e.g. escalation and resilience) versus existing professionals focusing mostly on non-state actors/prevention.
Here are the articles I was thinking of with AI summaries (that I endorse):
Filippa Lentzos, "Will splashy philanthropy cause the biosecurity field to focus on the wrong risks?" (Bulletin of the Atomic Scientists, April 2019). She argues Open Phil's large grants were pulling the biosecurity field's limited expert capacity toward catastrophic, "extremely unlikely" scenarios and away from more probable risks (natural disease, lab accidents, negligence) — noting Open Phil's single grants ($16M, $12M, $3.5M) dwarfed the entire Biological Weapons Convention Implementation Support Unit's annual budget (~$1M), risking "drowning out the diversity of perspectives" in the field.
"A Case Against Focusing on Tail-End Nuclear War Risks" (EA Forum, from a 2022 CERI fellow). Argues that prioritizing worst-case scenarios (nuclear winter, civilizational collapse) over preventing any nuclear use at all is a mistake, on three grounds: even a single detonation could escalate catastrophically, there's essentially no empirical base to assign likelihoods to tail outcomes, and narrowly emphasizing the worst cases risks eroding the normative taboo against any nuclear use. Advocates instead for preventing deployment of any scale, undifferentiated.
Christian Ruhl (Founders Pledge), "Philanthropy to the Right of Boom" (EA Forum). Documents a roughly 30-to-1 funding gap between "left of boom" (prevention) and "right of boom" (escalation management, resilience, post-war response) nuclear philanthropy, argues the neglect reflects bias rather than evidence of ineffectiveness (Cold War-era stigma, PR concerns, differing moral frameworks — he catalogs eight explanations), and makes the case that hedging with right-of-boom investment is rational since prevention can fail via accident or miscalculation.
There's a comment from Hamming's you and your research: https://www.cs.virginia.edu/~robins/YouAndYourResearch.html that I've tried and mostly failed to apply in my professional life.
I didn't say you should conform; I said ``The appearance of conforming gets you a long way.'' If you chose to assert your ego in any number of ways, ``I am going to do it my way,'' you pay a small steady price throughout the whole of your professional career. And this, over a whole lifetime, adds up to an enormous amount of needless trouble.
If you go into a culture that isn't yours and bring your norms with you, success will likely look like imposing those norms. It sounds like this happened in animal welfare with EAs. So...if you're going to have to integrate into a community rather than fundamentally transform and occupy it, it's probably a good idea to carefully choose which new things you want them to accept, and just leave everything else alone.
If you're able to at least, I'm not.
It's been a very very long time since I've commented here. But I may have some insights I hope will be useful.
I'm not directly an AI safety professional but I am friends and colleagues with many embedded in government and semi-government institutions in Australia. I see what you've described from the other side. I will say rationalists and EAs here are slightly better tolerated, mostly due to an increased (though still low) professionalism and physical distance from the Bay Area.
Some are not unsympathetic to the maximalist outcomes, but the way to broach that sort of stuff in governmental settings is by acculturation to lesser worries first. It's hard enough for them to get senior healthcare civil servant to care about employment shock from AI when that official is currently severely worried about legal liability in AI radiography. The hard problem of alignment may as well not even register.
So when EAs try to talk to them about the pax silica carving up M33 outside of quiet, trusted one on ones, the general response from professionals who agree with you is a combination of "you cannot say that near anyone with the title of Secretary or you'll set the cause back 30 years", and "Our primary goal this quarter is trying to get someone in cabinet to try Fable 5"
Which is to say whatever issues you have with communication with them, they're under far more severe restrictions, having to translate to multiple groups, most non technical and even tech-hostile, simply to function day to day. Their idiom is a way developed over centuries to allow people from radically different working cultures to achieve goals. And EAs, especially if the Bay Area type, have a tendency to treat the system as an inefficiency and route around it. This results in grumbling.
Unlike others, given short timescales, I'm not sure "see less of each other" is the right way. Instead I'd try cultivating personal one on one relationships with professionals and discussion in private. Try to have lunch with them instead of roundtables. Maybe talk about things other than AI. In addition identifying those ea/rationalists who can effectively work with professionals will be key.
One tactic I've seen work well is a spokesman/nerd team. Have a professional adjacent lead open discussions and guide through professional courtesies, but defer to the nerd on details and allow those ideas to be voiced in appropriate (generally hodded as "speculative") sections in order to ensure that accurate opinions are still being conveyed.
I just encountered this blog post by Dean Ball, and the conclusion shed some light on how he felt internally about AI existential risk culture over the years:
I want to close on a personal note. This is the first time I am writing about this issue in quite these terms, and yet I am telling you it is inevitable. Why have I taken so long to cover this issue? Well, I have brought up the topic of digital identification for humans and agents a few times over the years, and when I did so I was largely motivated by the concerns I’ve shared here. But it is true that I—and candidly I think many of my colleagues in the profession of AI policy—largely failed to talk about this issue with the level of seriousness and urgency it required. I think there are two main reasons for this failure.
First, this stuff is weird and off-putting, and many of us felt an incentive to meet our audiences in their comfort zone rather than ours. So among “serious people” (or would-be serious people), there was a general tendency to confine candid discussion about what most of us believe our near-term future will be like to private venues. We let our hair down in Signal chats and the little nooks of Lighthaven, but when the public was watching, we spoke in more abstract, tamer-sounding terms about it all. This was especially pernicious in 2024 and 2025, when it was essentially impossible to acknowledge any serious AI risk without being labeled a “doomer.”
I am just as guilty of this as my colleagues, if not more so. The thing is that it’s unpleasant to be screamed at for being a “crazy doomer” who wants to enact worldwide fascism (and similar, and worse). Being constantly labeled in this way also limits one’s influence. So many of the people who think about the governance of superintelligence, myself included, avoided that unpleasantness and bowed to the social pressure to self-censor. I’ve stopped doing so, in part because I grew tired of the discursive straitjacket and in part because I now have an eight-month-old baby boy into whose eyes I must look every day.
Second, many people believe that the coming of what I have termed “self-sovereign AI” will constitute a catastrophic loss of control event that will herald the end of human existence at worst, and the end of human primacy in the world at best. As my friend and former co-worker Josh Achaim recently pointed out, it is psychologically distressing for people with these beliefs—who constitute a large fraction of the AI safety community—to acknowledge the obvious truth that self-sovereign AI is coming, and coming soon. For the record, I do not think the end of human existence is likely, but I fully acknowledge there are ways in which the rise of self-sovereign AI could go very, very poorly for human beings.
I want to apologize for my personal failure to communicate in sufficiently serious terms about the specifics of self-sovereign AI, which I now understand to have been an enormous gap in my writing and speaking. Going forward, I will try to notice more readily when I am biting my tongue or, even worse, shutting my eyes.
Okay random experiment for you.
I read a bunch of normie books when I was a youngun, specifically on different ways of going about communication and general project management in day to day businesses.
One of the main things that is focused on is case studies and specific examples of stories of certain things working. The set of evidence that is more admissible there is probably something like a story of a real world example of where applying certain approaches worked in the past that showcases things without being a 4000 word essay with context in it.
This seems like a skill issue that is not necessarily directly obvious so I do understand your plight here.
Someone should make a rationalist book in the version of a self-help book that just bases things on stories, it might be good as a reference manual or something.
I've posted a response elsewhere, if in the future you'd prefer I'd not link it please let me know. My usual policy is to inform people if I speak about them to someone else. I like social transparency.
I've been thinking about this a lot lately and for me the question has been partially reduced to: should or should not Constellation move to SF?
After reflection, my stance shifted to: probably not. Geographical isolation allows for a heterogenous (w/r/t SV) focus on safety, and enough of this culture trickles down to the frontier labs. The frontier labs in turn influence actions at the neo labs, which in turn influences the frontier labs themselves.
There is this broader question of how the community could shift the underlying incentives to punish reckless pursuit of capabilities. One thing that rationalists could have done is put more effort into shifting public opinion in their favor, but there are of course countless issues with politics and the bureaucracy involved. I would be interested in discussions between influential people in Silicon Valley and rationalists that are aimed towards thinking of institutions (whether non-profit, for-profit, or government) that could make the underlying incentives more favorable towards positive outcomes.
But, with respect to Berkeley alone, I think their role as an isolated rationalist chamber is probably a net positive. The counterfactual is having Buck, Geoffrey, Paul, and Beth in SF where they could perhaps better disseminate their pov to the frontier labs, but idk if this would actually work, and they may already have enough influence by, e.g, having relatively frictionless access to leading researchers.
I think that more effort should be put into thinking about for profit institutions with more favorable underlying dynamics whose culture can survive the stress of massive capital injections and scale. Creating a gov institution seems difficult (see CAISI—not sure what the post-mortem is there).
One of the issues about Berkeley proper is that I think a lot of the people in SF think that their focus on slow down of capabilities and their dismissal of capabilities research to be silly/unrealistic, and I think that alignment researchers should be more cognizant of this, though perhaps failure modes and catastrophes that results for irresponsible practices (like OAI HF incident) will just resolve this. I've heard that some OAI capabilities researchers are already shifting to more aligned-shaped research following the incident.
Still, I do think that rationalists/professionals would benefit from brainstorming sessions on how to shift incentives to benefit more thoughtful buildout of capabilities, though what outcomes are acceptable are also under contention.
It's not like it's that hard to get from Berkeley to S.F. — it's like half an hour by car or BART.
Yet there is still a cultural chasm between SF/SV and Berkeley AI safety people so you either missed the point of the comment or are dismissing the importance of this cultural divide, which is insane to me given how much the divide between the communities bears weight on current events.
Perhaps a useful proxy to measure the cultural divide is the proportion of openly poly people within SF/SV versus Berkeley AI safety circles?
I agree they're somewhat separate communities — my point was if one wants to reach out to the other one, they're not physically very far apart. And I have e.g. met Evan Hubinger at LightHaven, so there is some crossover.
(BTW, poly people aren't that rare among Bay Area software engineers — though not as common as in the Berkeley AI safety community. And they might not discuss it at work.)
At the local AI safety co-working space, there are ~two kinds of regulars.
There's the kind of regular who's been thinking seriously about AI safety and alignment since pre-2022, who have passing to intimate familiarity with the funding ecosystem, the Sequences, and various conferences that happen at Lighthaven. Let's call them rationalists.
Then there's the kind of regular who comes in with many years of impressive industry or government experience, who realized in the last few years that it is important and worthwhile to pivot their career towards making sure that this AI thing is handled competently by the people in power, and who have many valuable skills, insights, and connections that are lacking in rationalist culture. Let's call them professionals.
There are, of course, many people who are somewhere in between - bright undergrads born this millennium who have been involved in EA since stumbling upon 80k hours in high school, professionals who previously identified as EA but drifted out of the scene a few years ago, founders who have idly read some Scott Alexander. But let's call it a dichotomy for now.
There's a large culture gap between the rationalists and the professionals. Robust mutual understanding seems important if we want to work together on this project of not having AI blow up our civilization, and it is only by working together that there is a chance that we might succeed.
Unfortunately, there seems to be an underlying assumption that all you really need to do is stick the two groups in the same room for long enough and the acculturation will happen by default. This seems wrong to me; my experience so far is that there are many unspoken assumptions that are held by the two groups that sort of never come up in conversation, except in weird awkward eruptions when a fundamental assumption is violated. When that happens, the risk is that it pulls the two groups further apart, instead of closer together.
Some examples from recent events and discussions:
But also, like, timelines, so it seems useful to ask: what might deliberate acculturation look like?
Here are some ideas for how to more deliberately bridge fundamental culture gaps:
Having written this list out, I am dissatisfied with it because I'm not actually that confident that doing all those things will result in the degree of trust and understanding that I think is necessary for us to do good work together.
What's the actual tension? Perhaps it's that it seems to me like the culturally rationalist are the gatekeepers of the funding resources for work on catastrophic risks, and this feels like a thing that is unsayable and somewhat anti-inductive (in that to explain it is to give an answer key to what funders want to hear, which is not ideal).
Recently a big tent animal welfare conference happened in town and I hosted a mixer for the EAs in attendance (around a third of the ~500-person conference attendees were EA). One of them, a long-established EA, gave me an interesting rundown of the wider animal welfare funding ecosystem. They said that EAs used to not be that accepted in the wider animal welfare movement, but as it became common knowledge that EAs control a large portion of the funding, the community began to embrace them more and adopt more EA frameworks and ways of thinking - but not without some amount of grumbling and resentment and bad feeling in at least some contingents. If you're not EA pilled you will simply resent the fact that the EAs are fundamentally not interested in funding the obviously morally important work of running your local donkey sanctuary, and there's literally nothing you can do about it. And if you are running a local donkey sanctuary you are sort of by definition not EA pilled.
Fortunately for AI safety, the smart policy person who wants to work on compute governance or export controls isn't proposing the AI-safety equivalent of a donkey sanctuary. And unlike the donkey sanctuary owner they have leverage in things rationalists can't buy at any price - things like institutional legibility, standing relationships, tacit knowledge about how to make key organisations do things. And perhaps the problem will also partially solve itself through funder pluralism; the rationalists are not literally the only people on the planet interested in funding alignment research and policy.
Still, I don't want to overstate the case for optimism.
Funder pluralism might result in parallel communities that don't collaborate in the same rooms at all. Animal welfare maybe only avoided this outcome because EA funding was a very significant portion of all funding, but it seems like what the rationalists are gatekeeping is one specific flavour of alignment funding.
Plenty of things professionals consider serious AI work (bias, misinformation, labor displacement, near-term harms) are donkey sanctuaries to at least some parts of the rationalist funding ecosystem, and perhaps it would be useful to make that common knowledge, but that legibility will come at the cost of resentment.
In animal welfare, EAs were newcomers who bought influence into an existing movement. In AI safety, rationalists are more likely to be ~founders, which tilts the balance even more in their favour, which brings with it a deeper (more resentment-building/polarizing) form of gatekeeping.
Bad near-term equilibria that might result (/maybe are already starting to appear) include:
I'm squarely in the rationalist camp, which means that when I see a problem my default response is to write a post and publish it publicly. A skilled professional, on the other hand, would solve problems like these in a way that never produces a public artifact, and because I can't read about how they did it, I can't model what they'd do.
I'm cognizant of the fact that I'm holding only half the puzzle pieces here. If you identify more as a professional than as a rationalist, and reading this post made you feel some type of way, please come find me and give me the other half of the puzzle.