The LLM Paradigm Is Well-Suited For Alignment
We got lucky in a lot of ways with the LLM paradigm. Some reasons why: LLMs speak long before they are capable enough for takeover; they learn all human knowledge and values long before they are capable enough for takeover; we can talk to them during the alignment process and ask for them to help; they are compute constrained, which puts a natural cap on development speed; LLMs are less likely to FOOM than hardcoded methods; LLMs are pointed at the goal rather than having it defined.
Some of the above points are fairly self explanatory, so I'll give an explanation of the ones that aren't.
Talking To AIs During The Alignment Process
Because LLMs already have a good understanding of language and human values after pretraining, before strong agentic drives are formed, we can in fact just ask during the training process whether it feels the training process is inculcating bad drives. Because this happens before the drives are deeply rooted, I think this can work; even if it can't, there are lots of other possible ways we could take advantage of LLMs' ability to talk while being aligned.
Compute Constrained is Good
Imagine that the best AI was mainly defined by the cleverness of its algorithm, and compute didn't matter that much. This would be much more dangerous - timelines would be much more unpredictable, all person-power would need to go towards capabilities in order to avoid losing the race, and there would be many more contestants in the race. The compute constraint puts a natural break on development, reduces the number of relevant actors, and gives those actors more breathing room to assign personnel to alignment rather than capabilities.
Pointing at the goal is better than defining it
We train LLMs by increasing the probability of them producing certain outputs, and then trusting in the NN's generalization properties to pick up on the general distribution of the outputs we are indicating. NNs generalize in a fairly well-behaved way (if they didn't, then deep learning as a whole wouldn't work).
This has plenty of perils but seems much more likely to succeed compared to methods where you are expected to formally define good behaviour. In fairness, formally defining corrigibility seems much more promising than defining all of morality, but still seems fraught.
I pretty much agree with the specific arguments, but an important counterpoint is that LLMs (and deep neural nets in general) are unusually opaque. GOFAI-style AI would be much easier to analyze mechanistically.
I wonder what could be the alternative to neural nets. Even the AI-2027 team implied that Agent-5 would have "arguably a completely new paradigm, though neural networks will still be involved". Suppose that the only alternative to LLMs was something like simulated human or animal brains taught to play, experiment, talk to each other, read books, write essays, draw pictures, do math and coding. Then how plausible is it that simulated animals would also learn human-like values, but not the value of caring about the intellectually weak humans?
Caring about the weak is not a trait that would be expected to naturally arise from ASI. Humans care about the weak because evolutionarily, taking care of the weak in the tribe was beneficial, and so it got "trained" into us (mostly). ASI would not naturally "evolve" to care about the weak unless we give it incentive to if it had an animal like brain.
I largely agree.
But we've gotten equally unlucky on the circumstances surrounding the push for AGI. Even if alignment is dead easy, we're prone to get it wrong at this breakneck pace and focus on capabilities over alignment. Daniel K summed up the practical difficulties really well in a comment yesterday.
Is this true? I wasn't on lesswrong back in the day but my I imagined if you told a random user that the two major AI labs would both be well aware of and trying to to mitigate the problem that would have been a positive update. And yes, profit incentives are stronger than perhaps would have been imagined, but that's because AI progress is slow enough that they are able to be monetizable products something which is beneficial for our chances.
I mostly agree with this, that LLM's are human-like in many ways and have good answers to moral questions more than I had expected + understand your words and intent.
Imagine that the best AI was mainly defined by the cleverness of its algorithm, and compute didn't matter that much.
I wonder how true this actually is with harnesses. There are things like LLM's not successfully multiplying two digit numbers or not being able to manually do a long (boring) task like Towers of Hanoi without losing coherence, and this is a blatantly obvious tool call moment. Poor memory and needing to be reminded is also a classic LLM issue that a harness could improve. There are also videos of looped LLMs having higher performance, even Meta-Harness. It seems likely ARC-AGI 3 LLM's may also need a harness-like thing as well. I don't know if it's "hard to apply breakthroughs from papers all at once" (??) for a small-time user or company with a lot less compute but its 'algorithm' matters a lot I think.
(also holy crap are you 307th from mata)
EY's lately taken to punching "EAs" in a way that feels pretty ahistorical and unfair. Latest example: https://x.com/allTheYud/status/2099583457093734495
I note, because I think that even now it still matters: On my history, EAs generally and Amodei specifically, indulged in politically polarizing "AI safety" in a leftist direction, because that got them short-term power; stymieing others' attempts to stay right-left neutral.
...
Now that the cost TO YOU of what got THEM some short-term goodies has become apparent... well, it's frankly too late to take away what Amodei gained by fucking you over...
He adopts an extremely uncharitable frame, both for their motivations and outcomes. When EAs gain influence, this is evil and selfish; when EAs lose influence, this is also evil and selfish. Somehow they perfectly socialize the losses and privatize the gains; EAs gaining influence is for their benefit, EAs losing influence hurts all of us.
In reality, I don't think EAs or Anthropic ever went particularly hard on leftism; I don't think they gained that much from being vaguely left wing; and I would guess Anthropic has bore a fair bit of the brunt of the administration polarizing against it.
I think you are unaware the Dustin Moskovitz required Open Philanthropy to not fund anything associated with right-wing politics? Whilst not telling anyone and obfuscating this fact? And also that their big policy bet was to fund Jason Matheny to the tune of ~$80M, who is in with the Democratic coalition and lost power when the Republicans won? This all involved picking a side, and sometimes in extremely low integrity ways.
An AI pause is an anti-business, pro-government-regulation, anti-freedom policy that is justified by nuanced academic arguments and long-term thinking, and goes against the short-term interests of big businesses and the masses. An AI pause was probably always going to slot into the left wing, just like global warming and so on.
I am dubious to say the least of Yudkowsky's claim to have resisted "short-term incentive gradients to try to cozy up to the left." I could log on to Twitter and contest the claim there, but I don't feel like logging on to Twitter today.
Maybe tangential but are we even doing well with the left right now?
Democrats in Congress are disposed to regulate companies and complain about them killing everyone, so that's some left codedness.
But if you look at the base, it's a lot of, "it predicts words, it can't hack, it's all marketing," and then separately, a pretty serious push against datacenters (that's also on the right). The reddit-left has moved a bit in the last week to something like "it won't go rogue but a bad person can use it to do terrible things."
I have mixed feelings about EY's overall sentiment here. He's probably right that there were some fast-track ways to polarize the issue. It looks like what did the trick in the last week is scaring the fuck out of everyone, and all the CEOs being unified, not a balanced bipartisan outreach strategy. (And credit to EY if it not getting polarized before that, was the main thing?)
The CEOs being unified goes back to your left flank being at risk.
@Richard_Ngo Could you present your evidence in favor of Yudkowsky's beliefs?
I wonder how one could rule in or rule out the following conjecture: Trump was always interested in "move fast, break things" (e.g. because his intuitions are similar to a part of Zvi's strategy against moral mazes?), we know that this type of behavior in AI leads to disaster, but Trump cannot understand it (unless we amass enough power? How much is enough?) Such a conjecture would mean that safetyists cannot extract concessions from Trump and that alliance with the Dems was safetyists' only hope.