Sam Altman and Dario Amodei are the public faces of AI safety and that's poisoning the well
Normies just smell a rat whenever those two bring up AI safety. I think when the AI safety community gets a word in edge-wise (like you're arguing with your friends or maybe a tweet goes viral), we usually say something like
"it could kill everyone. but China could also kill everyone, so now more than ever we have to supercharge OpenAI and Anthropic."
and we just sound like we're "one of them."
I feel like I should be hearing "shut it down" about as often as I hear "broke containment" and that's not really happening, instead it's more like "let OpenAI cook." So a number of people are pretty soured on AI safety because of this.
When have AI safety people said that OpenAI and Anthropic should be supercharged? Maybe you hang out with a very different set of AI safety people, because this has not been my experience at all.
Less_raichu seems to be saying that those two say that thing and that they're the loudest spokespersons for ai safety.
A lot of AI safety people say things about Dario Amodei's decision to move from a contract that forbids the military from allowing to use Claude to run disinformation campaigns to sway European elections to offering one that allows the military to do so as Dario standing up for his principles.
Given that both the first Trump and Biden administration ran their antivax disinformation campaign in the Philippines, thinking that no harm will come from allowing the present Trump administration to run Claude supported disinformation campaigns is bad.
Anthropic did suffer consequences for not rolling over on autonomous killing decisions and domestic surveillance but the decision was still giving up a good portion of the principles that are actually important.
I think this concretely comes up with "do you support a datacenter moratorium?" If you say no, you've taken a confusing stance with respect to the "it might kill everyone" stance.
I'll say you have a point, I can revise like this: the public does sometimes get exposure to a genuine AIS x-risk argument, but they reject it because they mostly hear about x-risk from AI CEOs who conspicuously don't say "shut it down." (Which is not the same as "the AIS person said don't shut it down.")
And not everyone rejects it, and the Overton window is probably moving to include mass unemployment and x-risk.
You're painting a very binary, "us vs. them" picture of the option space here. There are many policy positions one might take between the extremes "datacenter moratorium" and "cheering the leading labs"; rejecting the idea of a moratorium does not mean you want the labs to be supercharged.
For some examples of such policies, see e.g. the Radical Optionality suggestions.
Consensus processes have to reach agreement among a lot of people, some of whom are in different belief states. It's not surprising to me that we'd see GPT2 level reasoning like "if you don't vocally support a datacenter ban, you don't believe in xrisk", which I read as being because to the weak model a crowd behaves as, datacenter ban is a lesser version of banning it all, and therefore if you support banning it all, you support a datacenter ban. The shallow model the averaged crowd behaves as, in the part of the individual's brains running the Keynesian beauty contest, sees a low dimensional subspace where there's a success policy vertex at the extrema "ban it all" and the consensus process is to push towards that with lesser policies incrementally.
So then someone says, for example, "banning new datacenters is irrelevant because we need to shut down all computers capable of running a takeoff at all", and it sounds like not wanting to hike through the intermediate policies to get the one you want.
I make no claim any of this is the slightest bit reasonable. I am reading off my intuitive model of crowds.
I think the simple belief "it might kill everyone => shut it down" is a pretty good belief, and people rightly resist attempts to move them off it, and they are rightly skeptical of people who believe the first but not the second.
The crowd effect has more to do with the outgroup and less to do with matching what one's ingroup thinks. If a stranger say "x-risk => build data centers", they are not showing ingroup harmony, so they get evaluated for outgroup, and the answer is something like "I can easily see Sam and Dario lying about believing in x-risk as part of building hype, this stranger is a tech bro." I wish those two were not the most visible examples of people who believe in x-risk.
I was thinking of nate soares or so in my previous comment - people who say things like "sure, whatever, don't build datacenters, but that won't really help very much". If you're really just thinking of sam/dario types, then yeah agreed that their positions are barely consistent in the terrain. I do think if they can end up right, then everything will be fine. I am as worried as I am because I think their chance of ending up right - "build the safe superintelligence first" - is extremely slim.
According to many thinkpieces: people are reading less, students are cheating more, professors are angry, English professors have to treat their class like it's remedial, high school teachers are crying students can't write a full paragraph.
According to the chess community: everyone is better, across all levels. GMs train against computers, beginners have multimedia learning resources and can have their games auto-critiqued, intermediates get exposed to chess opening theory sooner, and everyone is playing more games with all sorts of people.
The "chess" experience is pretty techno-optimist, and since techno-pessimism is in vogue, you have to almost pay attention more carefully to notice the pocket of techno-optimism.
Has anyone tried squaring these two viewpoints? They seem pretty directionally opposed.
A simple hypothesis to check would be "things people do of their own free volition are getting better; other things are getting worse". Because for other things, you want to just cheat / take the easy route / whatever. (Probably "of their own free volition" is not the right category, curious what would be the right category.)
This is a continuation of a trendline that predates LLMs I think. Intellectuals have been becoming less literate for decades. At the same time, we're getting generally better at stuff, especially stuff that's easy to measure.
Why intellectuals (at least) becoming less literate? My guesses: Maybe because social media captures our attention, so we read less. Maybe because literacy is ephemeral and hard to measure, and our current society in general optimizes things that are easy to measure. Written exams are subjective compared to Scantrons, so writing skill is deprioritized too. The collapse of the expert class means nobody is in a good position to judge good writing from bad.
What's some terminology or lit to discuss the takeover scenario where "AI just sort of gradually takes over and no one notices"?
Sam Altman brought it up in a talk a year or two ago. It almost seems the extreme murkiness of it is part of the danger. It's not clear at what point AI integration has become somehow uniquely bad for humanity's quest to explore the universe, and it's also not clear at what point it's all that different from oligarchy/feudalism government, if you happen to share that cynical political view, which is trendy nowadays. (I guess I share it enough to care about it but not enough to reduce my worldview to it, and not enough to stop thinking takeover is worse.)
I was thinking of this in the context of open world evals which seem to point to this takeover case. The thing is, some open world evals are simply AI grading other AI. For example submitting an app to review by Apple involves an opaque process, which is good for un-gameability by AI training on public data, but bad if Apple is simply using similar AI in the process.
What you don't want to happen is the best benchmark to start scraping the ceiling of humanity's collective ability to reject AI slop, followed by us ceding the benchmarks to AI.
(Maybe these are two contrary points -- open world evals wouldn't be trusted if they're flawed in their reliance on AI? What if an open world eval achieves trust on human data but that eval itself is later ceded to AI as those humans automate?)
epistemic status = probably unoriginal :) but maybe it's fine to have an unoriginal quick take, I'm learning too
I don't know how motivated rationalists are to work on "defending democracy" causes. I am becoming convinced authoritarians seem willing to do what unaligned AI wants, something I thought was kind of a joke or exaggeration, but when you have data centers competing for basic human rights like water and electric and winning sometimes, I'm less sure. This does suggest pro-democracy causes are pro-alignment.
When you're reading and you feel yourself learning something, what do you do next?
Example: I started "Manufacturing Consensus", a book among other things about how bots might upvote content on social media as part of a propaganda campaign. On the first page I got to the sentence, "From his chair, he recruits people across multiple social media sites to essentially rent out their profiles for money."
This isn't too groundbreaking for me. But the word "rent" is new. I hadn't pictured that part of the market, that a social media profile might simply rent itself out, as opposed to being hacked.
So what I do next is... something else?
This is not great and I feel a tension here. Reading forward is simply wrong. I need to actually absorb what I just learned. Take a walk for instance. But taking such a break defeats my rote ideas of productivity and I might pick up a video game or music video or something, and the bigger problem is I lack an internal clock telling me when the break has sufficiently happened, and to go back to the next task of some sort.
Just looking for opinions on how people manage these moments, both giving themselves space to think and slowdown but also minding the overall focus level.
Oh also -- when I'm under work time pressure, yeah I'll bulldoze through these things. I don't know if this means I'm artificially slow when not under time pressure, or artificially shallow when under time pressure.
Reading forward is simply wrong. I need to actually absorb what I just learned.
Here's advice to the contrary from Ravi Vakil for potential PhD students:
Here's a phenomenon I was surprised to find: you'll go to talks, and hear various words, whose definitions you're not so sure about. At some point you'll be able to make a sentence using those words; you won't know what the words mean, but you'll know the sentence is correct. You'll also be able to ask a question using those words. You still won't know what the words mean, but you'll know the question is interesting, and you'll want to know the answer. Then later on, you'll learn what the words mean more precisely, and your sense of how they fit together will make that learning much easier.
The reason for this phenomenon is that mathematics is so rich and infinite that it is impossible to learn it systematically, and if you wait to master one topic before moving on to the next, you'll never get anywhere. Instead, you'll have tendrils of knowledge extending far from your comfort zone. Then you can later backfill from these tendrils, and extend your comfort zone; this is much easier to do than learning "forwards". (Caution: this backfilling is necessary. There can be a temptation to learn lots of fancy words and to use them in fancy sentences without being able to say precisely what you mean. You should feel free to do that, but you should always feel a pang of guilt when you do.)