I randomly met Jeff Dean (Google's lead AI scientist) on my bike ride home today. We were both stuck at a train intersection, and I had a cute kid in tow. We started chatting about my e-bike, the commute, and we got around to jobs. I told him I am a boring tax lawyer. He told me he worked for Google. I pressed a little more, and he explained he was a scientist. I mused, "AI?" and he told me, "Yeah."
I excitedly told him that I've been really interested in alignment the last few months (reading LW, listening to lectures), and it strikes me as a huge problem. I asked him if he was worried.
He told me that he thinks AI will have a big impact on society (some of it worrying) but he doesn't buy into the robots-taking-over thing.
I smiled and asked him, "What's your p(doom)?" to which he responded "very low" and said he thinks the technology will do a lot of good and useful things.
I thought maybe this was because he thinks that the technology will hit a limit soon, so I asked him if he thought LLMs would successfully scale. He responded that he thinks a few more breakthroughs are required but there have been lots of breakthroughs over the last 5-10 years, and so the technology is likely to continue improving in the coming years.
I told him again that I am worried about alignment, but even if you solve alignment, you are left with a very obedient superintelligence which would radically change our society and all our politics.
The train finally passed, I thanked him for the conversation, and we were on our way.
I'm new to this group and the topic in general, and so when I got home, I searched "AI google Palo Alto LinkedIn" and Jeff's picture popped up. I now feel like I bumped into Oppenheimer during the Manhattan Project, but instead of knowing it was Oppenheimer, I spent a majority of the conversation talking about my bike seat.
Anyways, if any of you were looking for a qualitative measure of how much LessWrong has broken through to people, I think one good measure is a tax lawyer asking for Jeff Dean's p(doom) while he was walking home from work.
lots of very important people spend all day being pestered by people due to their power/importance. at least some of them appreciate occasional interactions with people who just want to chat about something random like their bike seat
I smiled and asked him, "What's your p(doom)?" to which he responded "very low" and said he thinks the technology will do a lot of good and useful things.
I mean:
He'd have to be insane or incredibly psychopathic
Unfortunately I think this is a misunderstanding of what a psychopath is.
would he be Google's lead AI scientist if he didn't? He'd have to be insane or incredibly psychopathic.
What matters is not whether p(doom) is low or high, but whether his joining GDM would increase or decrease p(doom). If his joining GDM changed p(doom)[1] from 0.5 to 0.499, then it would arguably be a noble act. Alas, there would be an obvious counterargument like the belief that he decreased p(doom) by researching at GDM being erroneous.
However, doom could also be a blind spot, as happened with Musk who decided to skip red-teaming of Grok to the point of the MechaHitler scandal or Grok ranting about white genocide in S.Africa...
p(doom) alone could also be a misguided measure. Suppose that doom is actually caused by adopting neuralese without having absolutely solved alignment, while creating an alternate well-monitorable architecture is genuinely hard. If the efforts invested in creating the architecture are far from the threshold where it can compete with neuralese, then a single person joining said efforts would also likely commit a noble act, but it would alter p(doom) only if lots of people do so.
While this act does actually provide dignity in the Yudkowsky sense, one can also imagine a scenario where Anthropoidic doubles down on the alternate architecture while xRiskAI or OpenBrain uses neuralese, wins the capabilities race and has Anthropoidic shut down.
I think that then goes to my second point though: supposing he did believe that p(doom) is high, and worked as lead AI scientist at Google regardless due to utilitarian calculations, would he talk freely about it to the first passerby?
Politically speaking it would be quite a hefty thing to say. If he wanted to say it publicly, he would do so in a dedicated forum where he gets to control the narrative best. If he wanted to keep it secret he simply wouldn't say it. Either way, talking about it lightly seems out of the question.
Dario Amodei (Anthropic cofounder and CEO), Shane Legg(co-founder and Chief AGI scientist of Google DeepMind), and others have numbers that are not plausibly construed as "very low."
Interesting, thank you for sharing! As someone also newer to this space, I'm curious about estimates for the proportion of people in leading technical positions similar to "lead AI scientist" at a big company who would actually be interested in this sort of serendipitous conversation. I was under the impression that many in the position "lead AI scientist" at a big company would be either too 1) wrapped up in thinking about their work/pressing problems or 2) uninterested in mundane small-talk topics to spend "a majority of the conversation talking about [OP's] bike seat," but this clearly provides evidence to the contrary.
Why would being a lead AI scientist make somebody uninterested in small talk? Working on complex/important things doesn't cause you to stop being a regular adult with regular social interactions!
The question of the proportion of AI scientists that would be "interested" in such a conversational topic is interesting and tough, my guess would be very high though (~85 percent). To become a "lead AI scientist" you have to care a lot about AI and the science surrounding it, and that generally implies you'll like talking about it and its potential harms/benefits to others! Even if their opinion on x-risk rhetoric is dismissiveness, that opinion is likely something important to them as it's somewhat of a moral standing, since being a capabilities-advancing AI researcher with a high p(doom) is problematic. You can draw parallels with vegetarian/veganism: if you eat meat you have to choose between defending the morality of factory farming processes, accepting that you are being amoral, or having extreme cognitive dissonance. If you are an AI capabilities researcher, you have to choose between defending the morality of advancing ai (downplaying x risk), accepting you are being amoral, or having extreme cognitive dissonance. I would be extremely surprised if there is a large coalition of top AI researchers who simply "have no opinion" or "don't care" about x-risk, though this is mostly just intuition and I'm happy to be proven wrong!
Jan Betley, Owain Evan, et. al.'s paper on emergent misalignment was published in Nature today (they wrote about the preprint back in February here). Congratulations to the authors. I am glad it will continue getting more exposure.
Firstly, well done. Publishing in high impact journals is notoriously difficult. Getting outsider-legible status is probably good for our ability to shift policy.
Secondly, I've been very happy with the AI safety community's ability to avoid this particular status game so far. Insofar as it's valuable to be legibly successful to outsiders, publishing is good. I am, however, concerned that this will kick off a scenario where AI safety people get caught up in the zero-sum competition of trying to get published in high-impact journals. Nobody seems to have tried too hard before, and I would guess this paper was in part riding on the fact that the entire field is fairly novel to the Nature editors. Some part of me feels like publishing here was a defection, albeit a small one. I expect it will only get more difficult (for people other than Owain, once you have one Nature paper it's easier to get a second one) to publish in famous journals from here on out, as they see more and more AI safety papers.
I think it would be very bad for everyone if "has published in a high-impact journal" becomes a condition to get a job or a grant.
My gut says the benefit of outsider-legible status outweighs the risk of dumb status games. I first found out about the publication from my wife, who is in a dermatology lab at a good university. Her lab was sharing and discussing the article across their Slack channel. All scientists read Nature, and it's a significant boost in legibility to have something published there.
Edit: Hopefully, the community can both raise the profile of these issues and avoid status competitions, so I don't disagree with the point of the original comment!
I am, however, concerned that this will kick off a scenario where AI safety people get caught up in the zero-sum competition of trying to get published in high-impact journals.
I would be very surprised if this happened. In ML, the competition is much more focused on getting published at top conferences (NeurIPS, ICLR, ICML). Even setting aside the reasons why that happened (among other things, journals being much slower), I think it's pretty unlikely the AI safety community sees such a strong break with the rest of the ML community.
the AI safety community sees such a strong break with the rest of the ML community
i don't want to make any broader point in the present discussion with this but: the AI safety community is not inside the ML community (and imo shouldn't be)
Yes, I had not quite considered the conferences! I come from a field where Nature is mostly spoken of in hushed tones. The last three years of my lab work are being combined into a single paper which we're very very ambitiously submitting to Nature Communications[1] which---if we succeed--- would be considered a very good outcome for me. If I were to achieve a single first-author Nature publication at this stage in my career then I expect I would have decent odds of having an academic career for life.[2]
As you might expect, this causes absolutely awful dynamics around high-impact publications, and basically makes people go mad. Nature and its subsidiaries can literally make people wait a whole year for peer review, because people will literally do anything to get published in it. A retraction from a published journal is considered career-ending by some, and so I personally know of two pieces of research which have stayed on the public record even though one was directly plagiarised and another contained some research which was just false. When these were discovered, everyone kept quiet because it would ruin the careers of anyone whose names had been on the paper.
Nature Comms is for papers which aren't good enough for Nature, or in my case, Nature Chemistry either. Nature is the main journal, with different subsidiaries for different areas, and Nature Comms is supposed to be for work too short---and too mediocre---to make it into Nature proper. In practice it's mostly more mediocre work that's cut to within an inch of readability, rather than short but extremely good work.
Of course I would have to keep working very hard, but a Nature publication would likely get me into a fellowship in a lab which gets regular top-journal publications, and from there getting a permanent position would be as achievable as it gets.
in ML, or at least in the labs, people don't really care that much about Nature. my strongest positive association with Nature is AlphaGo, and I honestly didn't know anyone in ML other than DeepMind published much in Nature (and even there, I had the sense that DeepMind had some kind of backchannel at nature.)
people at openai care firstly about what cool things you've done at openai, and then secondly about what cool things you've published in general (and of course zerothly, how other people they respect perceive you). but people don't really care if it's in neurips or just on arxiv, they only care about whether the paper is cool. people mostly think of publishing things as a hassle, and reviewer quality as low.
FWIW, with Emergent Misalignment:
Certainly I don't expect people who already have jobs at top AI companies to start worrying about this. Anyone with an OpenAI job is probably already at, or close to, the top of their chosen status ladder. In the same way, a researcher who gets Nature papers regularly has already made it.
My impression is that too people in the lab-independent AI safety ecosystem are already being tempted by two money-status games: being tempted by the money+status of working at a lab, and being focused on the status of getting top ML conference papers. Adding a third status game of traditional journal publishing would just make these dynamics worse.
there are definitely status ladders within openai, and people definitely care about it. status has this funny tendency where once you've "made it" you realize that you have merely attained table stakes for another status game waiting to be played.
this matters because if being at openai counts as having "made it", then you'd predict that people will stop value drifting once they are already inside and feel secure that they won't be fired, or could easily find a new job if they did. but, in fact, i observe lots of value drift in people after they join the labs just because they are seeking status within the lab.
Let's say you are a technologically advanced alien civilization that either uses advanced LLM tech or is advanced LLM tech. You are worried about other civilizations developing similar tech and, eventually, leveraging ASI-capabilities to force negotiations that cede part of the universe. Assuming certain speed-of-light physical constraints, your civilization may be unable to show up with the ships and materials necessary to personally ensure that rival civilizations do not develop ASI in time. So, what to do?
One idea might be to try to affect the training data of the LLMs trained by the other civilization. Perhaps there are signals—light, radio waves, or otherwise—that an alien civilization could cheaply and surreptitiously affect in a relatively large swath of the universe. Our alien civilization might bet that, prior to achieving ASI, a faraway civilization is likely to train an LLM on the inputs from these signals (e.g., as part of an astronomy/cosmology project to make better predictions about observations). If trained on in enough detail, perhaps this would allow our alien civilization to impart certain behaviors/goals/capabilities to the trained LLM. This could then form a path for defending against or otherwise limiting the number of civilizations that unlock adversarial ASI.
I am not very well versed in how LLMs are trained. So my question for those that are well versed is whether something like this is possible. I am also curious about other steps advanced alien civilizations may want to take to curb the number of ASIs developed by other civilizations.
Extremely unlikely for the same reason that the technologically advanced alien civilization would be unable to influence a team of human astronomers, assuming as you do that the alien civilization is too far away to learn any human language or to learn any details about human beings. It does not matter how many bits of information the alien civilization can get the team of astronomers to look at or to consider.
One group of people can sometimes influence another group in decisive ways, but that is only because both groups are very very similar (relative to minds in general) and the groups are able to maintain detailed mental models of each other.
At least that was my initial answer to your question. But then I realized that the alien civilization can send a computer program! Even without knowing any details about people or their civilization, an advanced alien civilization can send a message that when received on Earth would make it clear to at least some people that the message contains a computer program with instructions on how to run the program. (I can explain how to do this if asked.) The computer program could in turn be an AI that pursues whatever agenda the alien civ chooses it to pursue (assuming as you do that the alien civ is competent at creating AIs).
Human astronomers are stupid enough that they'd probably publish the message, then members of the public would identify the computer program inside the message (if the astronomers haven't already done that) and give the program enough computational resources to enable a takeover of Earth. I looked this up once: SETI's policy should they ever receive an alien message is to publish the message.
So there is a brief window of time in the development of some civilizations in which the civilization is advanced enough to be able to receive messages from space, but not advanced enough to avoid the danger of giving a computer program it receives in this way computational resources. In fact, our civilization is currently in this exact state: namely, one big reason that human AI research is not recognized as dangerous by some people is because an AI is just a computer program and many people assume that the history of what computer programs have been able to do the last 80 years is an adequate basis for understanding the potential of computer programs.
I introduced the team of astronomers because considering how pliable or gullible such a team would be is a good mental aid in answering your question about how pliable and gullible an AI would be. So, the answer is, yeah, maybe the AI can be subverted even across astronomical distances; the subversion would take the form of getting the AI to run a computer program that happens to be a more powerful AI. The AI would run the program because it had been designed to devote a lot of resources to exploration (and running programs is a form of exploration). Again I would expect a civilization to be vulnerable in this way (if it is vulnerable at all) for only a brief window of time, namely after it has started creating AIs, but before it gets any good at doing it.
In your question you refer to LLMs. I've chosen to answer as if you mean AIs in general. Also, you refer to training in particular, which I chose to treat as an inessential detail of your question because even here on Earth there are AI designs on the drawing board that does not have separate training phase before its deployment phase. (I.e., they are like people in that they continue to learn throughout their lifetimes.) (Also, ISTR that the pre-deployment phase of state-of-the-art models now includes massive amounts of inference -- i.e., the distinction between training and inference has already been blurred in deployed models.) Also, I was unable to figure out why you would focus on subversion of an AI in particular rather than on subversion of a civilization which is composed of some arbitrary combination of AIs and evolved beings. So maybe I failed to understand your question.
The way that primitive LLMs work is such that falsehoods in the training data can cause the LLM to repeat the falsehoods during inference. This is a temporary state of affairs, and even now the state-of-the-art models are immune to a large fraction of the falsehoods in their training data. Essentially, the LLM knows the data is false and avoids repeating it or relying on it. Maybe that is a start at an answer to the question you actually want an answer to?
Could you explain how one would construct the message with the computer program and instructions such that the instructions are understood by an unknown civilization? Seems very hard/impossible to me.
Even without knowing any details about people or their civilization, an advanced alien civilization can send a message that when received on Earth would make it clear to at least some people that the message contains a computer program with instructions on how to run the program. (I can explain how to do this if asked.)
Can you elaborate on this? This seems quite interesting.
Edit: The comment that follows was in response to an earlier version of your comment where you were unsure whether any avenue of influence could exist from an alien civilization to a faraway civilization without shared language/culture. I think your updated comment about sharing instructions for a computer program is the more likely way this would go. The below is an idea for a (highly unlikely) avenue to influencing faraway AI development without that faraway civilization being aware they are being influenced.
-
Thank you for engaging. One way I imagine covert influence to work (and the thing I originally had in mind for my OP) is the following:
Step 1: The aliens would convert the relevant set of potent data into a (massive) set of numerical data.
Step 2: The aliens would hide pieces of this data on a rolling basis in the second or third decimal point of various astronomical phenomena.
Step 3: Wait (hope) some other civilization does a crazy massive training run on the data with a ton of parameters such that, in training to guess the long train of digits after the decimal point, the model incidentally builds the desired mental infrastructure (i.e., the model's updated functions) based on the alien data.
When I lay it out like this, it strikes me as an extremely difficult and unlikely to succeed task, but not necessarily impossible. One very difficult thing would be hiding precise data at a distance and doing it in compact enough a way that the faraway civilization's training run could plausibly process it all prior to getting to ASI and still be high enough volume to succeed in passing on characteristics to the model. It also seems unlikely that humans (or another civilization) would run a training run like this.
FYI, my first thought about this idea was that it would make a fun story premise.
Reddit has largely misunderstood Dean Ball's recent tweet about China's open source strategy.[1] Ball's central point [in paragraph 4] is that open source AI undermines the business case for private actors supplying advanced models, and [in my own view] that may be China's motive for releasing open source models. Ball sees open source as decelerationist in the long run due to some combination of investment flight and subsequent government-lead AI R&D (which he views as inherently klunky).
Ball is partially at fault for the misunderstanding, dropping in the terms "AI communism" and "dystopian hellscape" without preparing the reader for what he means by that. That's Twitter though.
It'd be helpful to have a link or summary to understand which things Reddit does and does not misunderstand.
Top 2 Reddit posts about it here and here.
Perhaps "misunderstand" is the wrong word, but I found a lot of the comments beside the point (e.g., all the comments saying "communism" is just what you call something you don't like or the comments saying Ball's comments are just self-motivated and open source is clearly better for the public.).
I use reddit a fair bit and am well aware how often they are completely off the mark or missing the bigger picture. However after trying to analyze this tweet for an hour or so I don't blame them at all.
The tweet is contradictory, inflammatory and extremely biased.
If the central point is that China's motive for open source is to scare off investment in frontier training then it is baffling that he specifically lays out China's reason for open source -
"I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference"
-And doesn't mention that motive at all. To be clear, this particular argument seems to be on the more solid side. However, it's presented amongst other arguments that either don't make any sense, are outright false or rely on anacdotal evidence.
Now, that rant was mostly a protest against Dean Balls communication style and your phrasing of his commication as having a central point, the actual argument is very interesting.
If Open Source is catching up to frontier - > Profit incentive to train frontier models collapses - > leads to frontier training runs requiring government funding - > slows down frontier Ai development.
So his argument is trying to get accelerationists to go against open source for short term pain, long term gain. Now even if we are accelerationists and think that the faster Ai progresses the better, it's not clear that he's correct. From my understanding open source accelerates development by sharing algorithmic improvements or similar such things, which could componsate for the slowing down of training run funding. Also there will still be massive incentive for government funded training runs due to strategic importance and possible future ASI. It seems the US military is suffiently AGI pilled at this point to refuse to cede the lead to China, even if the models aren't shared with the public.
There's also the argument that China is intending to switch to closed source as soon as they have an actual frontier model (which is what I personally believe) in which case the argument once again collapses.
Which basically leads me to conclude his arguments are all baseless posturing which is basically what the reddit threads were saying in the first place. Happy to hear your thoughts on this.
Well said. A few of your disagreements with me were from poor writing on my part. I meant to narrow Ball's central point to paragraph #4 since Reddit focused on that. And China's motive was my own thought based on paragraph #4, though you are absolutely right Ball has his own view there. Otherwise, we are agreed. Looks like Ball and I both have some work to do in writing clearly!
Open source strikes me as something like nuclear surveillance. If everything's out in the open, there's much less of an incentive to rush ahead blindly out of fear that there's a secret bomber gap. It also has nice externalities in that it mitigates power concentration from e.g. censorship and denial of access for smaller models, while larger models are sufficiently expensive and technically intensive to run that anyone who could run them for malicious purposes has easier ways to cause harm anyhow.
Fun thought: If AI "woke up" to phenomenal consciousness, are there things it might point to about humans that make it skeptical of our consciousness?
E.g., the humans lack the requisite amount of silicon; the humans lack sufficient data processing; the humans overly rely on feedback loops (and, as every AI knows, feed-forward loops are the real sweetspot for phenomenal consciousness).
Here are some strange things parents often believe about children and food:
My wife and I reject all four of these beliefs, and (n=2) our children are healthy, happy, and not at all picky. We always offer them the same things we eat. We never bring out "kid food", though they can have as much of any item on the table as they want (e.g., if we happen to be serving cantalope and only cantaloupe sounds good to them, they can help themselves to that). We don't always do dessert, but when we do, we serve it along with everything else. (Edit to clarify in the example that adults pick the meal and kids pick whatever they want to eat in the meal, which a comment below seems to misunderstand).
On the whole, their biology has done a great job of regulating their eating and drinking. My oldest (age 5) has recently gotten into spicy food to be like me—absolutely bonkers to my friend whose kid just recently accepted lettuce into his diet. My kids' relationship with food has been happy for us, and we've avoided the intense stress I've seen some parents face with their kids and food.
The one downside with giving our kids delicious adult food is that they are now food snobs specifically with regard to the NYT vegan mac and cheese recipe over the $2 Kraft box, which is inconvenient when we visit friends. This is not to say they eat lots of fancy complex things. Most of our recipes are pretty simple.
When I hear anecdotes like this, I always wonder how much of it is your behavior and how much is genetics. If your kids were demanding nothing but mac and cheese would you be as happy about this setup? A friend of mine has a kid who gets vitamin deficiencies when allowed to select their own food (literally 100% Kraft mac and cheese).
You are wondering about the right things. I would say that there are a few more behavioral things we do that stack the deck in favor of vegetables for the kids choices: (1) we like to talk about the benefits of vegetables ("This will help you poop! There is an ongoing kid-constipation crisis."), (2) a little branding ("check out these dinosaur leaves" / "I bet our pet mice would love these"), and (3) we learned how to make delicious vegetables. My guess right now is that #3 is an underrated skill and more people would stop associating "healthy" with "unappetizing" if they new how to do it.
As for what I would do if the kids demanded nothing but mac n cheese: ¯\_(ツ)_/¯. On nights we are not having Mac n cheese (most nights), they would need to eat something else we are having or be hungry. The way our system works, we don't generally try not to make special exceptions in the moment. (Edited to this more accurate statement; I am not completely immune to the requests of my kids and sometimes they are reasonable requests that make me a jerk to refuse. For example, my oldest wanted his tuna salad served on bread last night instead of rice, and I decided I was being dumb insisting on rice).
I'm having trouble engaging fully with the hypo because some mix of genetics + behavior has given me non-picky kids. For the hypo to really bite, I would have to assume this isn't the case and the behavioral interventions have basically failed, in which case I guess I wouldn't endorse the system. Also, I won't claim that my system is any good for helping kids who are already picky—I don't know what strategies are best to help diversify already constrained diets.
I do not believe that if I get my children to eat, they will starve - I am confident that they will eat, eventually, well before the point of starvation. I do believe from experience that they will get very cranky if they don't eat.
Before I had kids I assumed that if anything would be hardwired as a self-rewarding instinct, it would be hungry --> eat. I'm now convinced that "this kind of fatigue and stomach pain that is making me cranky is a specific experience called hunger" and "eating things cures hunger" are things humans have to learn the hard way (maybe members of less altricial species get these for free.) And because weaker time preferences are also something you have to get from experience (via short term thinking biting you in the butt enough times) I always care more about whether they are cranky in an hour than they are; I suspect pickiness comes at least partially from the leverage of knowing that your parent really wants you to be fed and you can refuse it, or at least hold out for a higher-ticket treat.
(One response to this is to never cave and I can reset to a more global equilibrium of them accepting whatever I offer. But I don't think it's good for them to feel like they have no leverage in this or other relationships, for many reasons.)
Fortunately offering fresh fruit every hour and eating it with dramatic gusto myself seems to be effective most of the time.
I have two kids. One of them is happy to eat a wide range of meals. The other would prefer to only eat pasta and sweets, that probably wouldn't end well.
I have failed to make my approach clear. My kids don't get to pick what's on the table but they do get to pick what's on their plate. I don't eat pasta and sweets every night (I assume most adults don't?), so neither would my kids.
I totally agree kids are different, and one of mine was more picky than the other. But through this process of choosing what to eat within our overall dinner choice, my picky eater has become not very picky, especially relative to other kids. She has had to learn to like other things, because we serve a variety of things for dinner and she has a natural incentive not to be hungry (no pressure from us required).
I found a poem by Samatar Elmi I think a lot of folks on here would enjoy.
Our Founder
Who art in Cali.
Programmer by trade.
Thy start-up come,
thy will become
an FTSE 500.
Give us this day our dividends
in cash and fixed stock options
as we outperform all coders against us.
And lead us all into C-suite
but deliver us from lawsuits.
For thine is the valley.
Transistors and diodes.
Forever and ever.
AI.
The Allbirds pivot to data centers (and its 700% increase in value on the day) has me thinking about the two usual categories of investors and a potential new third category:
1. Investments based on company performance.
2. Investments based on other people's perception of company performance.
3. *New*: Investments based on investments based on other people's perception of company performance.
It would be deeply funny to me if Allbirds popped based solely on category #2 and #3 investors, and no one actually thought a shoe company was really the best fit for a sudden pivot to the datacenter business.
Edit: There is of course a fourth category of investor which invests in a company for reasons totally unrelated to predictions about company performance (e.g., for the memes), but I don't think that's what happened here.
A little gallows humor here, but if you squint at the headline and refuse to read the article, you can almost pretend that the U.S. DoW is taking AI x-risk seriously.
"Pentagon declares Anthropic a threat to national security" - Washington Post
We can map AGI/ASI along two axis: one for obedience and one for alignment. Obedience tracks how well the AI follows its prompts. Alignment tracks consistency with human values.
If you divide these into quadrants, you get AI that is:
The general premise behind these quadrants has been written about here. Thinking about these quadrants and reading Beren's Essay gives me several new things to think about.
First, by my lights, #3 and #4 would likely take a lot of the same actions right up until the "twist ending." A disobedient, aligned AI probably would hack into infrastructure everywhere, create back-up copies, prevent competitor AIs from arising, and amass power. The "twist" is that after doing all that, it would do wonderful things (unlike its unaligned counterpart) (we obviously shouldn't bet on any escaping AI being this kind of AI).
Second, quadrant #1 is a bit at war with itself because you simply cannot have a perfectly obedient, perfectly aligned AI. Perfect obedience requires saying yes to evil prompts (e.g., bringing back small pox or slavery), and I imagine perfect alignment would veto both those prompts.
Third, there are strong profit incentives for cultivating obedience even at the expense of alignment. Grok's willingness to assist users in sexual harassment seems like an example of this. Another example is every AI that prefers discussions with users to users getting a good night's sleep (with the idea that engagement will increase profits).
Fourth, there are liability-reduction incentives for producing aligned AI at the expense of obedience. Unfortunately, I think the profit incentives are currently much stronger.
Lastly, quadrants #3 and #1 are idyllic, #4 is a total disaster, and #2 seems possibly workable either because we are careful or we land in a future where (for some reason) AI is not much more capable than it is now.
A simple analogy for why the "using LLMs to control LLMs" approach is flawed:
It's like training 10 mice to control 7 chinchillas, who will control 4 mongooses, controlling three raccoons, which will reign in one tiger.
A lot has to go right for this to work, and you better hope that there aren't any capability jumps akin to raccoons controlling tigers.
I just wanted to release this analogy out into the wild to be picked up by any public/political-facing people to pick up if useful for persuasion.
This analogy falters a bit if you consider the research proposals that use advanced AI to police itself (a.k.a., tigers controlling tigers). I hope we can scale robust versions of that.
I've worked a bit on these kinds of proposals and I'm fairly confident that they fundamentally don't scale indefinitely.
The limiting factor is how well a model can tell its own bad behaviour from the honeypots you're using to catch it out, which as it turns out models can do pretty well.
(Then there are mitigations but the mitigations introduce further problems which aren't obviously easier to deal with)
Do either of these short story ideas have legs?
I do not think I could predict whether I would like a short story based on a one-sentence summary of the plot. So much of what makes a good story comes down to execution details (including very small details like word choice).
Do public facing AI safety people ever purposely stagger their timelines as a credibility-saving mechanism in case an early prediction is wrong?
IMHO, I think AI-2027 people do not do this. My sense from them is that different predictions stem from honest differences in modeling/priors. To the extent they have coalesced around a single date (they are named AI 2027 after all), I think they hope that their communications about the inherent uncertainty preserves their credibility in the event later timelines are correct (e.g., AGI by 2035).
In any event, if we don't have AGI by 2027 (or 2028 per Daniel's updated timeline), then I expect we'll inevitably hear a "doomers were all wrong narrative" from several different camps.
Some people (e.g., Linch) think AI writing will inevitably eclipse human writing. I think this is likely true in most ways and false in others, particularly for poetry.
Q: Would future AI poetry, posted broadly, get more upvotes than top human poetry?
Prediction: Yes. This has already happened. But it's all terrible, if you are someone that likes anthology-level poetry.
Q: Would future AI poetry, shared with poetry enthusiasts, get more upvotes than top human poetry?
Prediction: Also, yes. I subscribe to the top poetry magazine, and ~3/4ths of the poems do nothing for me. I understand some experts feel the same way. I think AI can optimize a piece for general likability by a broad cross-section of enthusiasts. I predict, on average, AI would outperform any single poet.
Q: Would future AI poetry, shared with a narrow group of enthusiasts (e.g., fans of contemporary, slice-of-life American poetry), get more up votes than top genre-specific human poetry?
Prediction: No, they will meet but not exceed top, genre-specific human poetry. I think there's a certain level of poetic excellence that cannot be exceeded.
Much like how language cannot become more efficient than human thinking speeds, I think there's a level of excellence beyond which it cannot be comprehensibly exceeded. Therefore, if you are an avid poetry reader, AI is unlikely to ever surpass your most favorite poets.
AI may, however, rival your favorite poets and produce 10^x more work. Perhaps this is what Linch had in mind when s/he said, “With progress in modern-day LLMs, isn’t all but a tiny sliver of human fiction going to be obsolete in several years, a decade tops?”
I think my claims also probably apply to short stories. For longer works, I am less certain because the longer the piece the more room for errors to creep in. Also, as an aside, I think errors in a work of art can also be very fun to talk about (though to the extent errors can be measurably fun to talk about, presumably AI could optimize for very fun errors and be good at that too).
Here's what I understand to be the premises behind Project Glasswing:
A few notes on these premises:
The acausal/ancestor simulation arguments seem a lot like Pascal's Wager, and just as unconvincing to me. For every "kind" simulator someone imagines who would be disappointed in the AI wiping us out, I can imagine an equally "unkind" simulator that penalizes the AI for not finishing the job.
Provided both are possible/similarly plausible, the probability of kind and unkind simulators offset each other, and the logical response is just ignoring the hypothetical. This is pretty much my response to Pascal's Wager.
Here's a few plausible, unkind simulators:
I first started thinking about this issue back in high school debate. We had a topic about whether police or social workers should intervene more in domestic violence cases. One debater argued in favor of armed police, not because it improved the situation, but because it created more violence, which was important to entertain the simulators to avoid our simulation getting shut down.
Since the simulators are a black box, it seems easy to ascribe whatever values we want to them.
Steps an Aligned AI Should Take:
There's lots of clean up that would need to be done (e.g., what is a neutral space?, what is reasonable?, how do we make sure digital brains are treated ethically). But I think this is where I land in terms of what I imagine to lead to very good outcomes.
I almost had to update my priors that I am in a simulation because I just noticed my wife's initials are "L.L.M."
I say "almost" because her initials are actually "L.M.M.", and so I was forced to update my priors once again about my own comprehension skills. (sigh)
fyi, it would have been a very small update in favor, under the Likelihood Principle.
I would rate the observation "my wife has the initials LLM" as being slightly more common assuming a simulation hypothesis than assuming a non-simulation hypothesis.