Post from DJT a few hours after the one quoted in this post (bold is mine):
https://truthsocial.com/@realDonaldTrump/posts/117270591511950591
Concerning AI, when, in the History of Business, did anyone see the Leaders of an Industry call for Regulation that, if strongly implemented, will drive them into oblivion and bankruptcy? AI taking over the World, destroying Humanity, and all other things bad, is a HOAX, no different from RUSSIA, RUSSIA, RUSSIA — UKRAINE, UKRAINE, UKRAINE — IMPEACHMENT HOAX #1 — IMPEACHMENT HOAX #2 — and all of the other HOAXES and SCAMS that America was forced to endure through the Destructionists’ and Deviants’ foul play and illegal conduct. President Xi, of China, just announced that China will be doing absolutely nothing to stand in the way of AI, or its future. Google has recently stated that they want to build a massive Plant in Finland, all because they are finding permitting too difficult in the United States. I am not happy about this, and want them to change their thinking. AI, and Data Centers, will be the Greatest Economic Development Engine in History — Bigger than Oil, Gold, Diamonds, or even the Internet. It will not be stopped by brilliantly run Destructive Forces during the Term of President DONALD J. TRUMP!
Even interpreting this charitably, the plain claim that "AI taking over the World, destroying Humanity, and all other things bad, is a HOAX" provides permission for his followers to dismiss AI risk entirely. When taken in the gestalt of his stated positions, it is about the worst possible take one could see from the President of the United States from an x-risk perspective.
His narrow focus on building the "Greatest Economic Development Engine" is the paperclip maximizer wearing a suit and tie.
We should just win, gotta Beat China. The winner would then be the AIs, not us. That is not to say that the China issue is not real, hence the export controls and attempted international agreement.
This is Trump's position. God help Trump become more sane or set aside...
Dario Amodei has a new essay that finally says the thing: We Must Pace the Frontier, naming his call after the Pacing the Frontier letter lab employees signed in July.
As in, we need to slow the rate at which AIs increase their capabilities, to allow for the necessary alignment and safety work.
He explained that, without pacing, he expects things to escalate quickly. He offered three proposals, and unilaterally committed to the first one. OpenAI followed, and both Elon Musk and Demis Hassabis endorsed the overall proposal.
There is still a long way to go. The odds are still against us. The situation remains grim. The hard part lies ahead. We do not agree on what ‘Pacing the Frontier’ will mean in practice. But this is Actual Progress. The work can begin.
Table of Contents
Pacing Does Not Mean Pausing
AI capabilities are improving very fast. Even I cannot keep up.
Think about what models were like even one year ago.
Inside Anthropic and OpenAI, internal models are improving even faster. Over the summer there was a step change, as Mythos and Astra started kicking off the early stages of recursive self-improvement (RSI).
Dario worries that we by default are 6-12 months away from a rogue swarm of AI agents, similarly misaligned to the ones in the OpenAI-HuggingFace incident, being able to use a persistent botnet to take over the internet, or worse.
That is how fast he expects default progress to be. If we went a lot less fast than that, it would still be extremely fast.
Thus (bold his):
Pacing does not mean pausing. As Dario Amodei says, progress will still seem fast.
I strongly encourage you to read the original essay in its entirety.
Dario’s First Proposal: Embedded Evaluators
This is an excellent proposal. I am very happy that Anthropic and OpenAI will do it.
If we do not know what is going on inside the labs, we cannot do anything about it, and the labs have to worry that their competitors are speeding ahead.
Dario frames the benefits as:
Thus the first proposal enables the second and third proposals. It is also a good idea anyway. We need more visibility into the labs. It would have been very good to have such evaluators during recent incidents. So I call upon the other major labs, that have not yet done so, to also commit to this first step.
I agree with Dario that this should be made mandatory in its full form. If you are well-resourced enough to pursue plausibly frontier models, then you can afford to do this, especially if you are our biggest open model advocates, as in Meta, Google or Nvidia.
If you think that your operation could not survive if there were embedded evaluators checking its safety, then ask yourself why you think this.
Anthropic is proposing to empower the evaluators quite a bit:
It is easy to imagine a mostly fake version of embedded evaluators. This is promising to very much not be that, and to unusually empower the evaluators to report, if it was fake, that the arrangement was indeed fake. These details, if followed, answer the ‘oh you can just fake this’ objection.
The good objections to this are about implementation. We need enough evaluators, they need to be qualified, and they need to be independent and trustworthy.
METR is great, but METR cannot do this alone, and we have one hell of a set of incentive problems to solve. We also need to both have them be competent, and also not too linked to the existing ecosystems and labs, and the funding will have to come from somewhere.
You can get all three, but not if you define ‘independent’ as ‘no one ever pays for it’ and ‘no one who ever worked for the major labs.’ For METR in particular, the rule is the labs don’t pay, lab staff do not direct payments, and funders that are seen as biased, such as Coefficient Giving (formerly OpenPhil) also do not pay. Exceptions are made for free tokens. But of course someone, somewhere, has to pay. And similarly, you need to get your expertise from somewhere.
The White House made a similar mistake when they demanded that CAISI not be led by someone with experience at a major lab. That rules out everyone qualified.
If we choose superficially ‘credible’ or distributed sources I expect them to have no idea what they are doing in important ways. Cherry picking becomes a threat, especially for the labs whose hearts aren’t in it. Yes, it is important that in the real sense not all the evaluators be based in Berkeley. It will be tricky.
My true preference is something like ‘you need both METR-style people and also full outsiders as evaluators,’ and both need to be at each lab. You need both the people who can provide an independent inside view, and also the fully outside independent fresh eyes perspective, to avoid blind spots.
Part of this debate will be a coordinated effort to discredit METR and Redwood Research, and anyone else who understands the problem. As in: We get claims that ‘METR is not independent.’
Dario’s Second Proposal: Democratic Coordination
This is also an excellent proposal. I very much want to pursue it.
‘Government support’ here does not have to mean mandating or restricting anything. What Dario Amodei is requesting here is an antitrust waiver or active facilitation by the government, so that he can form agreements to not rush ahead without this running afoul of the Sherman Act.
In practice I would be unsurprised by threats or calls for antitrust actions, including from the White House, and I expect them from the likes of former AI Czar David Sacks. But I would be very surprised if our government would be so unhinged as to actually try to pursue a case under the Sherman Act. If they did, I would expect it to drag on for years, and ultimately cost a highly affordable amount at most. But in practice the companies are going to be very reluctant to risk this.
Such waivers are common. There are DOJ business review letters. There is the NCRPA safe harbor for joint research. This is not an extraordinary request.
It is deeply, deeply standard for industry to get together and agree on safety standards. If you oppose this, the reasonable alternative is that the government needs to directly impose its own safety standards instead. That is a reasonable position.
‘We should not have safety standards for building minds smarter than humans, that are getting radically smarter very quickly’ is not a reasonable position.
Dario says yes, government regulation on the frontier labs in particular (which, as always, would exclude all but something like at most five AI companies) would be first best, because they could be mandatory. No one wants to have to get buy-in from Mark Zuckerberg, for overdetermined reasons.
But realistically, by the time we could get actual government regulation, it would be too late, so at most the government will be facilitating discussions.
Again, if you think that others who are ahead of you meeting to agree on safety standards and to slow down their rates of capabilities research would be a threat to your competing business, rather than a boon, you might want to ask yourself why.
Dario proposes limits based on some combination of system capabilities and also inputs like compute, training run details or the nature of internal AI use on AI. Dario’s preference is to base this on capabilities and evals that relate to potential threat models, which would trigger required safety certifications.
Both kinds of restrictions can be gamed, but contra Dario here I worry more about evals. Another problem is that the evals won’t meaningfully measure what you want, as Anthropic’s Evan Hubinger warned us about recently, and as I’ve observed quite a lot. If you are relying on anything remotely like ‘automated alignment evals’ against potential superintelligences, I predict that you are rather cooked.
He also wants to combine this with strong export controls on chips to maintain our lead, and strong security on model weights and to protect against distillation attacks. Yes, obviously. That can then keep us in a position of strength and be a bargaining, well, you know.
Dario’s Third Proposal: Global Coordination
Ultimately, global coordination is The Way.
We can and should do things without it, to enable coordination, and to show good faith, and because as things are we would be better off acting unilaterally, including because our Chinese competition is fast following. If we slowed down they would be slowed down as well, especially if we enforced strong chip controls as Dario proposes.
There are four broad categories of objection.
Maybe it won’t work. We still have to try.
Dario lists four levels we can try.
Remember that ‘they will probably say no’ is not a good argument for not trying, even if you are right that the odds are not so great.
Why Pace Now?
One reason is because the alternative to pacing would escalate rather quickly. The other is that we can now do a lot more with the time.
Dario explains that if we had paced before, it would have been too early to do alignment or interpretability work that would have been meaningful. It would have, he says, been like ‘studying the psychology of humans by performing experiments on bacteria.’ We would not have had the tools to work with. Now we do, and there are infinite things to do. I think that wildly overstates the case, and not all work needs to take the form of such experiments, but directionally it is a strong point.
I agree that there are currently infinite things to do. If I was in charge of an alignment or interpretability team I would never run out of experiments to run and things to try, including having the time to properly talk to the models in the first place. The only reason I’m not working on that now is I think what I am doing instead is more impactful.
The same goes for testing and evaluation, which he also lists. So much to do.
Then there is his other suggestion of operational excellence. Right now, everything the labs do is rushed to the breaking point. Prosaic work is done sloppily. The training and testing environments are often broken, and Dario admits this extends to Anthropic and caused their recent issues. Best practices are not followed. We all saw what was going on inside OpenAI. Anthropic seems not to be failing quite that badly, but the situation is not great. Time would at least fix such basic mistakes.
I would also include that time gives us a chance to process what is happening, and to better pursue alternative pathways. I don’t love going down the LLM route to superintelligence.
The question was, now that Dario said it, would anyone answer the call?
Yes.
Sam Altman Agrees and Commits to Embedded Evaluators
Perfection. OpenAI will join Anthropic and commit to embedded evaluators.
The devil remains in the details. Altman did not commit, that I have seen, to the details as outlined in Dario’s essay. Without those it would be easy to do the fake version of this. And of course this is necessary as a first step but insufficient. Feet must continue to be held to the fire.
Still, progress. Sam Altman has changed his tune quite a bit, in a very good way.
This is a one page document that then must be fleshed out, negotiated and implemented over many hard steps, but yes. The core agreement could be very easy.
Here is his statement from Sunday night, where he cites the Narrow Path rhetoric and fears about concentration of power in order to try and calm the flip side:
Also this, which sums up OpenAI’s position taken over the last few weeks:
The proper capitalization lets you know he is serious. So is this:
OpenAI Will Not IPO This Year
File under costly signals:
From the same interview:
I mean, obviously, yes. That’s the premise of the question. If that was necessary, you do it. It doesn’t mean you should ask ‘what did Altman see?’ more than you should already have been asking that.
Altman also says he expects that there will be multiple points in which safety and alignment will demand a pause in capabilities development. He also says we likely cannot ‘push much further on capabilities without making more progress on monitorability, alignment, the ability to understand what a model are doing.’
Elon Musk Agrees
Elon Musk did not sign the July Pacing the Frontier letter that was signed by 1,386 frontier AI company employees, and has made statements recently that amount to ‘humanity will lose control over AI, so we here at SpaceX have to make sure we build it first.’
So it was a highly welcome surprise that Elon Musk responded the ideal way, except that to my knowledge he has not yet committed SpaceX to embedded evaluators:
The usual suspects responded to him with dismay. Elon Musk made clear he means it.
I am 47 years old, so one trick I use is I remember myself at 27 years old, and how I was less wise but in many ways I used to be smarter and faster.
When you see the usual suspects and their vibe warriors turn against Elon Musk and call him all the same names they call everyone else, it is a tell.
It is not a coincidence that all the top labs are founded by people who believe that AI is an existential threat to humanity, and that those VCs who do not believe this mostly missed AI and kept on dismissing it until remarkably late in the game.
Demis Hassabis Agrees
I am very sad that Google has managed to push him aside at DeepMind.
Dario acknowledges Demis’s proposal in the essay.
Google DeepMind’s new overlords Pichai and Kavukcuoglu have so far said nothing.
Microsoft CEO Satya Nadella Agrees And Talks His Book
I will quote his statement in full. This is not a full ‘Dario is right.’
It does welcome ‘deliberate pacing needed to get alignment right as the design goal.’
Satya welcomes ideas like ‘embedded evaluators’ but does not himself commit to them, and he calls for ‘broad representation across the ecosystem, countries and fields, including academia.’ Implementing that would require at least a waiver, or operating under government facilitation.
He then tries to position Microsoft as taking this style of approach. Okie dokie.
Anthropic’s Long-Term Benefit Trust Is On Board
The purpose of the Long-Term Benefit Trust is to safeguard Anthropic’s mission. They are supporting Dario’s call, and indeed helped inform the recommendations.
General Online Reactions
I will not quote from the chorus, but general online reception was about as positive as you could hope for, given this is a safety proposal from Anthropic.
There were some that said ‘this will not be enough,’ and yes Neel Nanda is right that you have to actually do things once you have your verification mechanisms, but almost everyone in that camp agreed this was an excellent start.
Most of those who are worried about AI killing everyone were quite happy.
Many of those who are not as worried about AI killing everyone were still happy.
Andrej Karpathy thinks embedded evaluators are a great idea.
There were those who raised practical objections or were skeptical of buy-in. Fair.
Then there were those who objected, including loudly.
All the prominent names were exactly the ones you would expect, and all the arguments were exactly the ones you would expect, focused on their mantras of ‘regulatory capture’ and ‘ban open source.’ The essay does not mention open source in any way, and the arguments involved had little to do with the actual contents of the essay beyond ‘Anthropic proposed acting responsibly.’
The comments on Elon Musk and Sam Altman’s agreement statements were of course flooded with generic ‘regulatory capture’ and ‘ban open source’ memes that every such Tweet always gets. The alliance of vibe warriors is still there. The rest of us have just realized that this is not real life, those people are not persuadable, and we do not have to care.
This is very similar to the reactions to Jacob Coxon’s resignation. The usual suspects who assume everything is a conspiracy made their usual accusations, a few raised good objections, and everyone else was happy.
Mainstream Press Coverage
Washington Post’s Ted Hesson, Ian Duncan and Gerrit De Vynck: Anthropic CEO calls for the AI industry to slow down. Good quick coverage of the basic developments.
Wall Street Journal’s Robert McMillan: Biggest AI Rivals Agree They Need to Slow It Down, focusing on the consensus between Dario Amodei, Sam Altman and Elon Musk, and the commitments from Anthropic and OpenAI to provide third-party evaluators with ‘permanent, employee-level access to our systems.’ Demis Hassabis has since agreed as well, although he no longer leads DeepMind.
Wall Street Journal’s Tim Higgins: Anthropic’s Moral Conflict Is Playing Out in Real Time. He calls this a ‘classic prisoner’s dilemma.’ People forget that, while the single-shot true prisoner’s dilemma is hard, the iterated prisoner’s dilemma is relatively easy.
Higgins is right to point to the IPO: If you really believed that Anthropic needs to pace the frontier, Anthropic would be wise to consider following OpenAI’s lead and pulling the IPO, as much as it would be good to unlock quite a lot of philanthropic dollars tied up in Anthropic stock.
The biggest mistake here is thinking Anthropic was founded to avoid this scenario. It was quite the opposite. Anthropic was founded so that, when the time came, there would be at least one responsible voice in the chorus or horse in the race, who could do things responsibly. They expected to end up here, in the end.
Dario Amodei was the lead on Face the Nation. There is an extended interview here. From that interview:
This is not a new Anthropic position.
In the extended interview, Dario Amodei also said that the AI industry ‘lied to people about the fact that this technology had risks.’ Yes.
Gavin Baker offers a summary of the first 24 hours. I disagree with some of his characterizations, explanations and predictions. If you think this is about things like Section 230 you are not pondering what the labs are pondering. But the facts are right.
OpenAI Researcher Explains What The Labs See And It’s a Rocket Ship
I do not agree with the first line under even cursory familiarity with the situation, but the rest of this is a very good explanation. I will quote it in full.
Consider the Alternative
What happens if we do not proactively pace the frontier?
There are a number of possibilities.
My baseline would be the same as that of many lab employees. We would race to superintelligence without knowing how to align, understand or control it, and all die.
Another possibility is that we would get a bigger warning shot, and a bigger reaction.
Or that we simply see the arguments for a full pause carry the day, which is looking increasingly plausible.
It’s Totalitarianism, Joe
Opposing driver’s licenses is reasonable, likely even correct. The mistake is equating them to totalitarianism.
There are four arguments that will reliably be brought out, by the exact same people as always, against any attempt to not die, or indeed any attempt to make AI safer.
They also try to label anyone proposing any such action, any at all, as a ‘doomer.’
It doesn’t matter if it is even a government action.
It doesn’t matter if it would apply to open models.
It doesn’t matter if it would explicitly be a competitive advantage for open models.
It doesn’t matter how light touch it is.
It doesn’t matter what it is, at all.
These people would see you literally shoot yourself in the foot, and say that is a plot against little tech because they can’t afford good health insurance. Regulatory capture. Plot to ban open source. Totalitarianism. You will lose to China.
Or, when you said you agreed not to shoot yourself in the foot? Open models can’t be made to not shoot people in the foot. Thus, this is regulatory capture. Plot to ban open source. Totalitarianism. You will lose to China.
A regulation that says closed frontier labs have to follow light touch safety procedures they themselves select, while open models, academics and anyone below an explicit high size threshold for money or compute spent are explicitly immune from all requirements? Plot to ban open source. Regulatory capture. Totalitarianism.
OpenAI announces plans to put pineapple on pizza? Totalitarianism, plot to ban open source, regulatory capture. You will lose to China.
OpenAI announces plans to not put pineapple on pizza? Totalitarianism, plot to ban open source, regulatory capture. You will lose to China.
Yes, it is always the exact same people, spouting the same Obvious Nonsense, no matter what you say. You will lose to China.
In this case: The top labs voluntarily commit to embedded third party inspectors? Plot to ban open source. Regulatory capture. Also totalitarianism, presumably, and you’ll lose to China.
The only alternative would be to give open model creators money, and exempt them from all liability or responsibility for what they do, and sell our best chips to China. That’s freedom.
At some point, you need to stop taking these objections as mapping onto reality.
Yes, such competitors have the resources to train frontier models but not the resources to allow for a few embedded evaluators, so voluntarily having embedded evaluators to check your power is instead a plot to concentrate power.
That does not mean that these concerns are never real. Some actions would indeed cause some combination of these four things to become more likely, or move us in such directions. And yes, we should take that into account.
But please, stop entertaining these same people saying this same thing every damn time anyone tries to do anything helpful. It is Obvious Nonsense, and Content-Free.
That also means not trying to placate or calm such folks, or address their concerns. You can’t. Their concern is that you want to take costly action to not die and perhaps not give them the maximally cool toys maximally fast. They are against this, and they are against it the same fixed amount no matter what.
I would not even dignify it by calling it a ‘conspiracy theory.’ There is no theory.
If your response is ‘such folks are reacting to Dario’s second and third proposals, not his first one’ then flat out no, that is not what is happening. You are incorrect.
You do want to address the underlying real concerns, to the extent they are legitimate. But do not make the mistake of trying to convince them you are doing this, or taking their reactions into account. They are a rock with these lines written on it. Period.
Yes We’re The Baddies How Did You Know?
The usual suspects from the previous section do have one redeeming feature.
They wear metaphorical skulls on their metaphorical uniforms.
If you must take the role of a cartoon villain, this is very good form.
As in:
No, seriously, this did not get deleted and he doubled down that he is not kidding:
Sometimes People On the Internet Just Lie
I am done pretending ‘they don’t know the facts.’ They lie. And not well.
There are plenty of understandable confusions about the HuggingFace attack. Accusing OpenAI of instigating the attack explicitly and on purpose? Yeah, no. That is not a misunderstanding Jason could possibly have reached, other than on purpose.
Dwarkesh Patel dutifully smacked Jason down while acting as if Jason is misinformed.
In case you were wondering if they are principled libertarians who don’t want the government on their side, well, no. As an example, Jason suggested responding to ‘the new models might kill everyone’ with requiring models be open sourced after a year. His solution is ‘we get to take your stuff.’
David Sacks Groks The Situation
Here’s a fun interaction, and no, as usual, you did not need Pangram to notice that the post was in the style of an AI, most likely Grok.
Okie dokie, sir. We learned something new about you today.
I strongly believe that it is indeed a major faux pas to post AI text as your own. It is fine to post it if it is clear that the writing is AI.
David Sacks Says Go Ahead
On substance, Sacks’s message boils down to: You have the advantage right now, and you’re the ones who say you’re doing something so dangerous, so you two (Anthropic and OpenAI) should slow down first, without worrying about antitrust laws and commercial consequences, and let the rest of us catch up to you. It would be good business, and then it would buy you goodwill.
And yeah, okay, Sacks is overstating in many places including that one, but at the center of it is a pretty good point. Not that Sacks or his ilk would ever listen, they will read such a concession as weakness and foolishness and propaganda, and attack even more, maybe even call for antitrust action against them.
Indeed, Sacks himself says the companies should ‘agree’ to do this as a duopoly, which is a big no no without a waiver, and exactly the thing we have to dance around until the government does (less than) the absolute minimum and gives us the freaking waiver. What happened instead was that Anthropic did it on its own, and OpenAI followed.
You cannot simultaneously say ‘go ahead and do [X] in the illegal way’ and also say ‘without us giving you a waiver that permits [X].’ Well, I mean you can, Sacks did it, but it is not a good faith move.
That doesn’t matter. It is not about the mustache-twirling villains. It is not about fair.
It is about the fact that if you are about to do something existentially risky, maybe the first thing you should do is Stop It.
Lies and Confusions About Who Previously Claimed What
This keeps happening, and it will continue to happen, as part of the campaign of association to call anyone who warns about any downsides of tech as ‘doomers’ and then to associate them all with each other, to then dismiss all concerns.
I do not know of anyone prominently warning about AI existential risk who previously predicted extinction or other unrealized major downsides from climate change. Not one case. Those worried about AI acknowledge that climate change is a real problem and try to be helpful in finding real solutions, without catastrophizing.
This is the one that blows my mind. There is almost no overlap between NFT evangelists and those worried about AI killing everyone. It is quite the opposite, as Yglesias and Gross say below.
The counterargument is ‘well they vibe the same to me, Jack’:
There is actually a huge difference, but yes, a lot of people simply file everyone under ‘outgroup’ and then link them accordingly.
An especially fun one is to claim that all of this concern over existential risk is ‘new’ or that it needs to be ‘explained’ by some cynical reason, or even that it (lol) shows that they aren’t making progress on capabilities, are you kidding me.
House Speaker Mike Johnson Wants To Lock Everyone In a Room
The reporting on this has often been absolutely terrible. Yes, Mike Johnson has correctly pointed out that the AI companies have an obligation to make their products safe, the same as every other product. In no way has Mike Johnson said the safety of AI is not also the responsibility of Congress or the White House.
Congress has not already instituted meaningful AI guardrails. It is bizarre to be using that as a talking point. But yes, there is an obvious corporate responsibility to ensure that your products are safe, even in the sense of mundane safety. Certainly you also are obligated to make sure your product does not (checks notes) kill everyone on Earth.
At no point did he then say ‘and this is not my problem.’
One core objection of his is that while all the AI leaders say we need guardrails, different leaders all want different guardrails. This is a reasonable objection that can be fixed.
He actually said:
Alas, Johnson falls back on Lose to China, but he does so in a newly balanced way.
Words like ‘pause’ and ‘moratorium’ are scary to those with this mindset. I get it. And yes, if you waited long enough and China proceeded as per normal then China would potentially match and then pass us, although it would take longer than people think because the Chinese are mostly fast following.
Mike Johnson has no intention of having the House of Representatives in session debating things and passing laws at this time, including on AI, due to midterm elections. That is different from saying Congress has nothing to do with any of this. His reluctance to ‘rush in and pass a piece of legislation’ has multiple causes.
Well said.
Here is Mike Rogers, the Republican running for Senate in Michigan:
Here is the Minority Leader:
This is exactly the right priority. You want to ensure the ability to intervene, including the ability to slow down.
Donald Trump is Not Tired of Winning
Donald Trump has already ‘paced the frontier’ somewhat, slowing down releases of Fable and Sol, and even yanking Fable from the market. Donald Trump’s administration was willing to act when the threat was concrete and in front of them.
He is willing to put guardrails on AI. But he’s all about keeping the vibes good, except where he’s all about keeping the vibes bad. Gotta have the right vibes. And Trump’s not yet buying this ‘existential risk’ thing, that’s ‘negative forces.’
How dare ‘negative forces’ ‘bring it up’ when it ‘won’t happen.’ Why won’t it happen? It won’t happen. You’re bad for the vibes, dude. Trump is willing to do what it takes, so long as you play nice and have the right vibes, which can in some contexts mean alarming or bad vibes, big scary vibes, the worst vibes, but not quite here and now.
The misleading headline chosen for this was ‘Donald Trump rejects calls from tech bosses for an AI slowdown’ and that he ‘denounces demands for regulation’ but that is terrible reporting. It is not what he said. I would caution against reading too much into Trump’s statement. You can find the full clip here and I encourage you to watch it from about 7:30 when the question gets asked.
Trump is asked ‘should AI be regulated or slow down?’ From his perspective those are trigger words that mean something very different than ‘embedded evaluators’ and ‘not build Dyson Spheres in 2029,’ also he is unlikely to have been properly briefed. After the above passage, including agreeing that ‘we can put guardrails,’ Trump pivots to ‘look at all the factories we are building.’ What Trump is trying to do here is head off the data center protestors and full pause advocates and negative vibes. Then he talks about which reporters look better.
Trump then reiterated this on Truth Social, which is easy to misread but you have to actually read the words:
Trump is mad about the anti-data center thing and the potential conflation of the two, and about the implication that things might go wrong on his watch, and taking a potshot at Dario Amodei for once again bringing the bad vibes. But obviously you should not interpret this as saying ‘the guardrail is literally that I am the President and therefore we do not need to, what’s the word, actually ever do anything.’ The guardrail is that he will choose the guardrails, and so on. Keep those vibes good, folks.
This is also standard muscle flexing, which may or may not have anything to do with threatening antitrust actions. I doubt it, but such strategies thrive on strategic ambiguity. It’s hilarious how much Fable takes these ‘criminal threats’ literally.
Trump is indeed strongly opposed to anything like a full pause, and is a big fan of the data centers, but none of that is news. It also could change, and fast.
Thus, after a brief period of hopefulness, the chance of a federal AI safety bill this year is back down to 18%, which I interpret as correctly saying ‘not without another incident, and maybe not even then.’
We Must Avoid Polarization on AI at (Almost) All Costs
I endorse this.
How Will We Know If They Actually Paced?
I disagree with Daniel in that I am very confident this is not ‘regulatory capture.’
I agree with him that we should be very concerned that it is kabuki, or that the labs do not meaningfully slow down, or they might only slow down in the places where prosaic safety would have stopped them anyway for prosaic reasons.
We don’t know what the counterfactual pace of progress is, and it is hard to know what progress is supposed to look like when cashed out into practical and real-world capabilities and applications.
This can also be viewed as a question of what pacing means. Does it mean ‘this is costly but worthwhile’ or does it mean ‘this is necessary and we would have to do it anyway, and thus it is worthwhile’?
It can be both, or that can become a real disagreement. What is the counterfactual? If you couldn’t proceed because your product was too misaligned to use internally, did you actually ‘pace’? Kinda yes, kinda no, right?
The embedded evaluators are designed to be able to address this question, by reporting on what steps have been taken and what the counterfactual might have been. This has to become a ‘trust but verify’ situation.
If we can keep things on the current trend lines, certainly that is way better than breaking the trendlines upwards into imminent recursive self-improvement. We would go from remarkably little chance to at least some chance. But I agree with Daniel that if we see ECI staying on-trend, and that is an accurate representation, we are still moving way too fast. The last year was too fast, and we are paying even the prosaic bills now in terms of the crisis of misalignment.
The Real Frontier Is Internal Models At Top Labs
Roon is correct that misuse risks are a problem but internal deployments going haywire are now the primary concern. We are launching swarms of 10,000 Astra-variant agents at plausibly impossible tasks like Millennium problems a week after training starts. It’s not a great leap to imagine how that might go horribly wrong.
This is what must be paced. This is where the battle will be won or lost.
We can finally get to work, and have our chance. Don’t waste it.