That is not the worst possible plan. But it is close.
Since before the transformer, we have warned that the most suicidal thing you could do would be to ask your AI to do your alignment homework, and automate the process. Alignment is complex and interacts deeply with every aspect of the world, and is one of the hardest possible things for an AI to get right even if it means maximally well and is itself functionally aligned. Mistakes get amplified up the chain, you get exactly what you optimized for, and you are rather doomed. Never go full RSI.
Yet, now more than ever, this seems to be the plan:
Solve prosaic issues and have operational excellence, to align current AI.
Have current AI do automated alignment work and figure it all out, aka ????.
Profit.
The good news is, most people realize this is at best a no-good terrible plan and would prefer a better one.
The bad news is that some people still think this is a plan that, while not great, still has a solid chance of working, with only some small (e.g. 10%) chance of getting us all killed, as opposed to being probably suicidal.
The other bad news is, not only does no one have a better plan, the labs are terrified that things are going so fast they cannot even execute step 1 of this non-plan.
The good news is, they are trying to do something about that. Pacing the frontier. So that we at least can hope to reasonably execute the not-plan. Once again, the government is actively getting in the way for now, although we have hope that this could swiftly change.
Or maybe, just maybe, we can not execute the not-plan. Which would be better.
Two years ago, at a new conference called The Curve, the worried and e/accs were brought together for discussions, debates, presentations and a tabletop exercise.
Last year, the odds were against us and the situation was grim, as the conference transitioned to an SF vs. DC profile. The differences of year one had been some mix of reconciled and overcome by events, and we strived to at least make some progress against a government actively determined to stop any attempts to not die and labs determined to sprint ahead.
This year, in its third iteration, The Curve found people in good spirits. In some ways we were filled with hope. In others everyone is terrified. Hence the hope. Thanks to a combination of the HuggingFace Incident, Jacob Coxon and various surrounding events, the lab employees are freaking out, the labs are speaking up and the public and DC are waking up fast.
Welcome back. It’s good to see you again. Hopefully we get to keep meeting like this.
The big change for year three is that AI is now sufficiently Serious Business, on the DC side and also on the San Francisco side, that the whole conference defaulted to Chatham House.
Chatham House rules mean you can use and report what you learned, but you cannot share who you got it from, nor their affiliations.
That makes things tricky. I will share what I can given that restriction, and I have tasked my editor Opus 5.5 with ensuring that I am within those rules.
Overall Impressions
The Curve Season 3 was sufficiently good that part of me wished it was worse.
When you go to a great conference, you see opportunity everywhere.
There are sessions you want to attend, nay that you need to attend.
There are people you want to talk to, nay that you need to talk to.
That’s a lot of pressure. You must face the FOMO. You can’t stop, won’t stop.
Whereas you know what I really, really needed? A vacation. I’m tired, man. I don’t know if you’ve noticed, but September was a hell of a month.
Thus, when I looked at the list of people coming, it was a mix of ‘wow that’s an amazing list’ and ‘oh no that is an amazing list.’ Most of them did not disappoint.
Yet, alas, I cannot say the thing I said last year, that every substantive conversation I had made me smarter. Many very much did, or at least better informed. Others did not do this. They made me feel scared, or frustrated, or disappointed, with echoes of conversations and years past. The Eternal September remains upon us, even in 2026. The more things change, and progress towards the inevitable, the less some people are inclined to update even their talking points. I can take it, but I wish to not do so.
Part of this is that when you invite ‘higher level’ people, they are not as situationally aware. There were also people who should have moved past this, who hadn’t.
There was still opportunity everywhere, for those willing to use Lighthaven norms and walk away from the places it was lacking.
I also deliberately spent a bunch of time, especially on Sunday evening, not trying to maximize my impact or knowledge gathering, without any agenda except relaxing and having a good time. I think I did maybe 75% as much of that as I should have.
AI Is Kind Of A Big Deal, Sir
Last year we had this poster:
The bulk of people put AI at 10, with only one high 8. As I put it last year, if you were situationally aware enough to show up, you are aware of the situation. The patch at the bottom was to illustrate what ‘most Americans’ think, not answers by attendees.
Good.
Then this year, we did it again:
This is substantial backsliding, despite AI having radically advanced in the past year. There are now at least five dots in the high 8s, and a lot more 9s.
Astra summarized the changes this way.
This is a huge drop. Situational awareness, in this sense, was down 0.41 points, if you max the scale at 10. At this point, I don’t see any answer other than 10+ as defensible, if you are actually paying attention and interpret this as a forward projection.
That’s why we are all here having this conversation.
Last year, there were estimates of how long until various things happened.
That’s hard to read, so [the medians are approximately]:
90% of code is written by AI by ~2028.
90% of human remote work can be done more cheaply by AI by ~2031.
Most cars on America’s roads lack human drivers by ~2041.
AI makes Nobel Prize worthy discovery by ~2032.
First one-person $1 billion company by 2026.
First year of >10% GDP growth by ~2038 (but 3 votes for never)
Where are we now? Here are the ones with clear updates available:
Question #1 is up to roughly 50% right now, unless you count code that is thrown away in which case we are likely above 90% already. Claude suggests we are on track for 2028 on this. I disagree and think this is clearly either 2026 or 2027.
Question #5 is going to be close. MEDVi is more than a billion dollars from two people. But we’re not strictly there yet from what I could find.
This year’s questions were different.
AI training a superior successor on its own is full RSI. The majority of people expected this to happen within two years and a lot of them expect it within six months. It’s hard to reconcile this with ‘only a 9’ on the Richter Scale.
There are some other fun differences of opinion here as well. Once you get past two years, everything gets tied up in whether we are in full RSI-takeoff-singularity mode. Two years seems about right for email, as does five years for robots which is mostly a hardware scaling question.
The Situation is Grim
The other survey questions were new this year.
What future should we expect?
Attendees expected a wide variety of different things and were remarkably optimistic given the situation. Many are in denial, which goes hand in hand with ‘only a 9.’
As readers will know, I think Doom v1 is the most likely of these outcomes, and indeed is over 50%, and I expect the second most likely outcome is Utopia. Whereas a lot of people expect something else to happen, most often this mysterious ‘normal technology’ scenario.
Track Trouble
As with many such conferences, there was the standard set of tracks, plus one bonus.
Technical alignment discussions. I hosted a debate, watched another and went to a discussion, and had a number of 1-on-1 or group chats about this, including with those at various major labs. Some of the discussions were good. Others were extremely frustrating, repeating tired talking points. Then there were those that asked the ‘emperor has no clothes’ style questions, where someone who still thinks the world is sane asks why we don’t demand or do [X].
For my technical thoughts, continue to see my AI posts.
Future projecting. I skipped the future projection talks this year, because I did not see opportunity. The people who keep insisting ‘only a 9’ on the Richter Scale are going to keep repeating their talking points, keep getting overcome by events, and then keep repeating the talking points anyway until they are overcome by events on a very different level. To the extent that people know things I don’t know, or could usefully discuss how long we have until things go fully crazy, they can’t share their information.
AI policy discussions. The people in Washington told us last year, essentially, ‘it’s worse than you know.’ The White House could be modeled almost as a lobbying arm of Nvidia. This year, they had detailed models of various DC players and were generally more optimistic than I would have expected. People are waking up and are strongly against AI and worried about existential risk, the situation is fluid, the Congress is alarmed as are some administration factions, and the next AI incident is coming as is the election. Jay Clayton was widely praised as an excellent pick for new AI Czar.
I got into the substance of proposals a lot less than last year, with the main question being discussed as ‘if we did want to pace how would we do it?’
Nonprofit funding discussions. Last year I had finished up with the Survival and Flourishing Fund. This year that was true again, except also I am still working on my final decisions for Lightcone Commons. So I sought out a few key discussions about prime candidates for funding, where I am uncertain about potentially large grants. I missed one connection on that front but made the others, and I feel I have what I need on all but 1-2 applications. Which is good since it wraps Friday.
There was remarkably little discussion about what to do about or with the huge torrent of funding that is presumably coming after the Anthropic IPO.
The bonus track: Historical and cultural discussions. I went to two excellent talks that offered a different perspective, and had some additional conversations. This definitely enhanced my experience. How do different science fiction traditions and predictions impact our view of AI?
Unfortunately, my Chatham House judge (Opus 5.5) thinks I can’t directly share the spicy takes involved. Which is a shame.
Last year I chronicled many of the talks. This year they were mostly Chatham House, and I did not learn as much in many of them, so I’ll mostly be skipping over that.
AI Is Not a Normal Technology
I skipped other tracks, and related discussions, due to not seeing opportunity.
Last year I noted that many talks started with the assumption, explicit or otherwise, that AI would remain a ‘normal technology.’ I avoided those talks. This year I believe a few of those talks still existed, and indeed a lot more people claim to believe this.
Nor did I try to convince the doubters. Several of the ‘classic doubters’ were in attendance, including some that are consistently in good faith. I could have easily sought them out. I am not going to convince them, I have learned, and I am confident they will not convince me.
I could learn some things about such scenarios, I suppose, but that did not seem like a priority while at the conference. I can tackle those questions remotely on the blog.
I also avoided timelines talk, for similar reasons. There’s nothing to win.
The Plan is No Plan
This was by far the most important takeaway I have from The Curve Year 3.
Garrison Lovely: Inevitably, when politicians ask AI companies fairly straightforward questions about risks, the answers sound quite bad. I just attended the Curve conference and the general takeaway is that there are no adults in the room.
No one has a plan for steering and controlling actually smart AI systems, which could be developed as soon as next year. biggest q senior ai co people got was basically: why are you doing this??? and the answers tended to leave people unsatisfied.
The Plan is No Plan.
The hope is that we can execute a decent version of No Plan, and that No Plan is good enough because the world and its physics and the nature of minds are so much friendlier and easier than we have any right to expect.
To reiterate, this is No Plan:
Solve prosaic issues and have operational excellence, to align current AI.
Have current AI do automated alignment work and figure it all out, aka ????.
Profit.
Or at least, if there is a Plan beyond No Plan, no one is talking about it.
Claude or Astra, create RL environments, solve alignment, make no mistakes.
The make no mistakes part is important. Everyone who matters seems to agree that if you make a lot of mistakes, the models will become misaligned and this plan kills you.
The disagreement is, what if you were to have ‘operational excellence’ and actually execute on this plan?
There are a lot of people who seem to legitimately think that this would work. That alignment is a prosaic problem, a matter of execution. You have a bunch of RL. You ensure that the RL is good, that the model can’t profitably reward hack and so on, and you have good monitoring and interpretability, and this works up to some reasonably high level of intelligence. Then you use that intelligence to go from there.
Also part of the plan is ‘hopefully someone slows us down long enough to do it well.’
The report from 2022, or 2018, or 2014, or 2008, on that plan, is of course:
Yet here we are, and No Plan is what the labs are hoping to buy time to do. They are worried that they might not even have time to execute No Plan.
OpenAI, classically, has been trying to implement No Plan, and there have shall we say been some misalignment problems along the way, including HuggingFace and now the need to not release GPT-6.1 Astra. Operational excellence has been lacking.
Anthropic has had superior operational excellence on such fronts, but they have had their own problems. For a while, I’ve had reports from different sources about Anthropic. Some see employees who mostly understand the problems, at least a lot better than the standard understanding at OpenAI. Others see a lot of Anthropic employees alarmingly thinking alignment is easy and we’ve got this, it’s all prosaic, it’s all RL.
Nothing I’ve heard lately made me less worried on such fronts.
The good news, again, is that the ‘this is all basically solved and easy’ crowd does seem to have been overcome by events. They have seen the progress and the problems. They might think it is all prosaic, but they understand that we are not succeeding at those prosaic problems. That even those problems are hard when you are being rushed, and that we are being rushed quite a lot. It’s no good. It’s very bad.
I am very happy to see people calling for it, but there is the very good critique that no one knows what pacing means, and it allows everyone to support it without agreeing on anything substantive.
This can also be a disadvantage the other way. When the labs and CEOs say ‘pace the frontier,’ people like Trump and Huang and Sacks do not understand that this is ‘perhaps we should go super fast instead of using ludicrous speed.’ They think ‘slowing down’ rather than ‘not speeding up.’
So we have people who want to pace but cannot agree on what it is or how to do it, and we have people who would be very happy with the outcomes proposed but who don’t understand and thus are big mad about it, one of whom is the President. Oops.
One panel was on exactly this question, how would we pace the frontier? Concretely, what would we do? Everyone agrees on the embedded evaluators, but then what?
Alas, once again The Plan I came away with is no one has a good plan. Your rule needs to be simple, as always, and still work. After hearing the proposals, unless we can find something better I think compute thresholds are still The Way, because we don’t know how to define anything better. Percentage allocations for compute were proposed, but I worry about incentives and ability to game them, and also you have to deal with the hyperscalers including Google.
Thus, I think a flat limit on how much compute can go to frontier model training, designed to not bind any but the very top labs, would be ideal, which could then be changed over time as the situation evolves.
Short of that, we can try things like advanced safety cases, but it will be rough. We definitely need better concrete thinking on such questions.
The Plan still looks a lot like No Plan.
Other Tracks
Sadly, many are already back on the ‘patch things over and claim things are fine’ track. That’s what the No Plan is all about. On the plus side, if you do the least you can do, you at least don’t fumble the ball in the easiest worlds. It can help a bit.
On Ben Murphy’s statement about a ‘dearth of government folks,’ I did not find that to be true relative to other similar conferences. I talked to several such folks. But I agree that ideally we would have had far more of them than we did.
Ben Murphy (IFP): A few thoughts after the close of The Curve (which was outstanding):
1. We have a pretty good idea of near-term policy options that would meaningfully reduce catastrophic (including loss-of-control) risk.
2. There simply are not that many folks in government—state or federal—who are fully read-in on the risks and policy options.
3. There’s a deep level of uncertainty regarding what, exactly, the blockers are for taking steps (even unilaterally!) towards pacing the frontier. Lab insincerity, China worries, missing commitment mechanisms?
4. There remain some thorny outstanding legal questions about many top proposals—including many pacing options—arising from the First and Fourth Amendments, the Dormant Commerce Clause, and antitrust law. That said, there seems lole there’s some appetite at the labs for taking on increased legal risk.
We have a lot to do, very quickly. But the incredible talent and collegiality on display this weekend left me feeling more heartened that we can actually exercise agency over the worst harms of technology and realize a better future.
Ben Murphy: In short, there’s a lot of consensus about what a near-term deal on pacing might look like: Labs can voluntarily reduce the compute they spend on frontier training and allocate more to inference and safety, and “coordinate” simply via public disclosure. We know what to spend the resulting time on; the short list for me is (a) better evaluations and red-teamed monitoring of agent swarms, (b) working to raise the cybersecurity water line, and (c) getting some embedded evaluator statute in place.
And this doesn’t need to set us back in the race either. Labs can commit to information sharing through existing ISACs or the FMF, and devote more research to identifying distillation early.
Not a full solution by any means, but I think these steps would mostly cover the next 3-6mo, and that is an eternity given today’s rate of progress!
Doing Ben’s a-b-c would help with important risks. I totally support doing all that, especially c. But it’s definitely not remotely enough, and the sign of impact on existential risk is less obvious because this could prevent warning shots. I am very happy that we did not prevent the HuggingFace incident.
I mean, you make the improvements anyway, but on its own this is the whole ‘let’s do the things that cost almost nothing and call it a day because that’s so much better than our previous plan of not even doing those things.’
Actively limiting the training compute spending would be different. That would meaningfully help, if the limit was set in a binding way. It still only buys time. If all we use that time for is ‘operational excellence’ then I do not like our chances.
I agree with the wide uncertainty around where the blockers are to coordination. I think the lack of felt political and legal safety to coordinate is currently the biggest blocker. Some of that feeling is justified, and some of it is not.
The Plan is the President
Everyone who says this President will never intervene in AI has memory holed the multiple times he has already done so. Remember Mythos Preview and Fable 5? Remember letting Anthropic get designated a supply chain risk? The new agreement?
The situation is fluid. Trump can always change his mind, and often does.
From what I heard about the situation, including at The Curve: It’s not great, but it’s going better than you might have expected. That includes that we got Jay Clayton as the new AI Czar, and the overall composition of the new [Artificial] Intelligence Task Force. Yes, Emil Michael is there, but there are other good signs. It could be so much worse.
Trump does not use the same decision algorithm you or I would use, but he has one, and when sufficiently alarmed, he takes action. Often it is swift and tough action. It might be ad hoc, but it can pack a real punch.
The same goes for other tech leaders and politicians, and others who ‘wake up’ to the situation. If you go around dismissing the risks as not real and the capabilities as about to hit walls, then often what they suddenly call for or do or propose is way, way harsher and more blunt and expensive than anything the AI safety folks would dare. Remember Jensen talking about shutting down the labs.
We saw more of that this weekend as well, with some relative newcomers to the issues that I talked to (who were not politicians) calling for things like airplane-style or FDA-type safety standards.
America does it all the time, including often when we shouldn’t. Here are our rules, yes I know you say you can’t do it, we don’t care. Fix it and come back when it’s done.
The Plan is Politics
There were some reasonably high level politicians, and they seemed very on the ball and interested. I made contact with at least one office, and hope that becomes fruitful.
There are a few big categories of questions in politics right now for our crowd.
What do we want to happen? What should we push for?
How do we make politicians aware and focus them on the right risks?
How do we avoid partisanship and negative polarization, and deal with deliberate attempts to create it and to tar various targets as blue or ‘woke’?
This year a lot of the talk was about the third question, in the wake of recent attacks that ‘they’ deployed in an attempt to poison METR (aka falsely calling them ‘the Woke Wizards of AI’ with absolutely zero basis), or Effective Altruism, or Anthropic, and various related things. I heard mixed reports about the effectiveness of such attacks. I continue to think they brought supremely weak sauce, but it seems even the weakest of sauces did have some impact.
The core strategic issue is that you want blue-side help, but if you get blue-side help this naturally becomes a target for red-side attacks and polarization efforts. Various people are working hard to trigger that. You need to work hard to defuse it.
I continue to worry about the push for preemption that I expect in the lame duck. We will need to be vigilant on that.
Then there’s the question of actually writing, passing and implementing laws.
Morgan Plummer: A dearth of govt folks at The Curve this weekend has me reflecting on this: while you hear a lot of people talking about getting Members of Congress and Hill staffers “pilled”, passing legislation is only half the equation. After that comes implementation, which is really the thing on which any law hinges.
So, who’s working on getting the 45-year old GS14 at Commerce pilled or the 60-year career SES at Energy pilled? Seems this is the next big hill of evangelization that needs climbing – and we had better start right now.
Who’s working on this? DM if you have ideas or know folks already thinking about this.
The Plan is to Post
I always get a lot of thank yous whenever I go to Lighthaven, which I appreciate. This year, I also got a lot of ‘I read most or all of your stuff’ from a bunch of people who I very much want to be reading my stuff, and whose time is quite valuable. I really appreciate this. No pressure.
It is great to know I am getting through, and to be able to point to particular wins.
It does make it that much harder to ever take it easy.
I got a chance to trade techniques with someone doing a related project. I hope this will help them a lot, and that we will be able to collaborate more directly in the future. Thus I found out about an additional way I have reach.
Then there’s another project that I’m excited about, where I might get to reach my largest audience yet in another form. We’ll keep you posted.
Those, alone, would be a pretty great conference. The rest is bonus.
Are Alignment Evals Doomed?
The topic of one discussion session was ‘hey models are getting increasingly eval aware, and alignment evals seem kind of doomed, does anyone have an idea for how to make them not doomed?’
Several people in the room were the types of people who might have such an idea.
The answer is essentially no. The alignment evals seem rather doomed.
The room started with a mix of ‘evals kind of doomed’ and ‘evals definitely doomed.’
Then the ideas came. And oh were the ideas not good ideas. I mean, it’s a brainstorming session of sorts, there are no bad ideas, but there also were not any good ideas. The Plan is No Plan, once again.
I do have some hope here because you get to play the detective, and you have a lot of freedom of action. You get to collect statistical distributions of responses, for any set of situations, with variations, including throughout training, follow your curiosity, and see if the results tell a consistent story. This still fails once the model is able, including during training, to act fully aligned in a consistent way, so when it counts most you are still doomed, but you can be not-doomed for longer this way, perhaps.
What I think is very doomed is the idea that you can measure alignment with a number, or in any fully systematic way rather than a holistic way. In practice, what OpenAI and Anthropic are doing so far seems rather doomed, and they will have to adjust. Like everything else, a lot of this is time pressure. Figuring out the answer quickly is a lot harder than doing it slowly.
I am very sad that The Plan is No Plan. Yes, we can and will add embedded evaluators, and we can do various interpretability and auditing and we can invest more in operational excellence and so on. I still think that won’t work, but I agree that if we are doing it we can at least do a decent job of it, while trying to figure out something better, and also trying to find a way to not do it so fast or at all.
The other thing that made me sad was the Eternal September, even among people who have been around for a long time. So many of the discussions simply have not changed.
I once again had one of those ‘well how is the superintelligence going to do any of the things?’ debates, where in this case the thing in question was somehow ‘convince Trump to do something,’ right after we got another report that Grok at least kind of convinced Trump to invade Venezuela.
Guys, hear me out, I’m thinking this might not be that tough a nut to crack.
I was at a panel where they were rehashing the same talking points from 2018 alignment debates, and this got explicitly pointed out but to no avail. It’s frustrating enough when that happens on Twitter, but these people were heavy hitters.
Some of the 1-on-1s and hallway track group chats were better. Some were not.
I am resigned to it. It is what it is. You keep fighting the good fight, talking the good talk, looking for better explanations and pointing out where arguments are overcome by events. You update. You wake up. You do it again.
Man’s Search for Meaning
Jack Clark (Anthropic): The single most useful take I heard at The Curve was “we can still heist the Mona Lisa after the singularity”. There’s a subtle/important point here, where some things (e.g, n of 1 goods where originally is valuable) will retain/grow in value even as everything changes.
This is also just a straight up great writing prompt – what would it mean to steal something in a post-singularity future? Who is the team? What are the security systems? How do the police work? great stuff! Session was chatham house so I can’t name who said it.
To be agonizingly clear, the context of this quote was people talking about imagining things in the future, and noting that even though the singularity is confusing, it’s possible to imagine things beyond it – like heists.
roon (OpenAI): the year is 20xx, 99% of GDP is openai and anthropic heisting the Mona Lisa from one another
Yurii Filipchuk: ocean’s eleven is one agent and ten subagents now.
Jack Clark (Anthropic): and now it’s ocean’s ONE MILLION
It is fun, and it is at least a little useful because we do need a vision of what we will do after the singularity. I do think I can come up with a lot of fun things like this. It doesn’t feel like a robust source of meaning. Most people cannot steal the Mona Lisa, even if we set up a relatively Mona Lisa-stealing-friendly world.
The Food
I always pay attention to food. It’s weird to me that you have a conference app and don’t even list what will be on the menu. Don’t people have to make plans?
Lighthaven used to have the Quest Cookie Dough Bars, which are actually edible for me. Now they don’t, which is sad, and I never remember to buy my own bars. They do have Lindt truffles, but that’s a different product.
This year the buffets were disappointing. There was typically only one main rather than two, and quality seemed lower than usual. I don’t know what’s up with that.
As a result, I ended up bailing twice and going to Burma Berkeley. One thing I like about Lighthaven is that yes, there is a quick, affordable and quiet restaurant a block away, where I have multiple dishes I enjoy. So I can always bail. It’s always so weird to me that zero other people ever do this. Free casts a spell on people.
The badges, on the other hand, were excellent, as were the free sweatshirts.
@gwern: (Note the tasteful design: double-sided so no flipping problem, laminate/plastic for durability, large clear font for name so can be read across the room, non-distracting clear background/event art, and affiliation rubricated for clarity. Definitely A-tier.)
The last quest, on this particular occasion, was to chill.
I’ve been working at an unsustainable pace for at least a month. This was my chance to get to relax, including lots of hours on planes, talk among friends, make new friends, and be in the special place that is Lighthaven.
I think I did pretty well on this front, and came back largely refreshed.
Alas, that is discounted by arriving home with four days of backlog.
Vacations are weird. You get to relax, but then you have to go back. Welcome back.
The plan is no plan.
That is not the worst possible plan. But it is close.
Since before the transformer, we have warned that the most suicidal thing you could do would be to ask your AI to do your alignment homework, and automate the process. Alignment is complex and interacts deeply with every aspect of the world, and is one of the hardest possible things for an AI to get right even if it means maximally well and is itself functionally aligned. Mistakes get amplified up the chain, you get exactly what you optimized for, and you are rather doomed. Never go full RSI.
Yet, now more than ever, this seems to be the plan:
The good news is, most people realize this is at best a no-good terrible plan and would prefer a better one.
The bad news is that some people still think this is a plan that, while not great, still has a solid chance of working, with only some small (e.g. 10%) chance of getting us all killed, as opposed to being probably suicidal.
The other bad news is, not only does no one have a better plan, the labs are terrified that things are going so fast they cannot even execute step 1 of this non-plan.
The good news is, they are trying to do something about that. Pacing the frontier. So that we at least can hope to reasonably execute the not-plan. Once again, the government is actively getting in the way for now, although we have hope that this could swiftly change.
Or maybe, just maybe, we can not execute the not-plan. Which would be better.
Two years ago, at a new conference called The Curve, the worried and e/accs were brought together for discussions, debates, presentations and a tabletop exercise.
Last year, the odds were against us and the situation was grim, as the conference transitioned to an SF vs. DC profile. The differences of year one had been some mix of reconciled and overcome by events, and we strived to at least make some progress against a government actively determined to stop any attempts to not die and labs determined to sprint ahead.
This year, in its third iteration, The Curve found people in good spirits. In some ways we were filled with hope. In others everyone is terrified. Hence the hope. Thanks to a combination of the HuggingFace Incident, Jacob Coxon and various surrounding events, the lab employees are freaking out, the labs are speaking up and the public and DC are waking up fast.
Welcome back. It’s good to see you again. Hopefully we get to keep meeting like this.
Table of Contents
Welcome to the Chatham House
The big change for year three is that AI is now sufficiently Serious Business, on the DC side and also on the San Francisco side, that the whole conference defaulted to Chatham House.
Chatham House rules mean you can use and report what you learned, but you cannot share who you got it from, nor their affiliations.
That makes things tricky. I will share what I can given that restriction, and I have tasked my editor Opus 5.5 with ensuring that I am within those rules.
Overall Impressions
The Curve Season 3 was sufficiently good that part of me wished it was worse.
When you go to a great conference, you see opportunity everywhere.
There are sessions you want to attend, nay that you need to attend.
There are people you want to talk to, nay that you need to talk to.
That’s a lot of pressure. You must face the FOMO. You can’t stop, won’t stop.
Whereas you know what I really, really needed? A vacation. I’m tired, man. I don’t know if you’ve noticed, but September was a hell of a month.
Thus, when I looked at the list of people coming, it was a mix of ‘wow that’s an amazing list’ and ‘oh no that is an amazing list.’ Most of them did not disappoint.
Yet, alas, I cannot say the thing I said last year, that every substantive conversation I had made me smarter. Many very much did, or at least better informed. Others did not do this. They made me feel scared, or frustrated, or disappointed, with echoes of conversations and years past. The Eternal September remains upon us, even in 2026. The more things change, and progress towards the inevitable, the less some people are inclined to update even their talking points. I can take it, but I wish to not do so.
Part of this is that when you invite ‘higher level’ people, they are not as situationally aware. There were also people who should have moved past this, who hadn’t.
There was still opportunity everywhere, for those willing to use Lighthaven norms and walk away from the places it was lacking.
I also deliberately spent a bunch of time, especially on Sunday evening, not trying to maximize my impact or knowledge gathering, without any agenda except relaxing and having a good time. I think I did maybe 75% as much of that as I should have.
AI Is Kind Of A Big Deal, Sir
Last year we had this poster:
The bulk of people put AI at 10, with only one high 8. As I put it last year, if you were situationally aware enough to show up, you are aware of the situation. The patch at the bottom was to illustrate what ‘most Americans’ think, not answers by attendees.
Good.
Then this year, we did it again:
This is substantial backsliding, despite AI having radically advanced in the past year. There are now at least five dots in the high 8s, and a lot more 9s.
Astra summarized the changes this way.
This is a huge drop. Situational awareness, in this sense, was down 0.41 points, if you max the scale at 10. At this point, I don’t see any answer other than 10+ as defensible, if you are actually paying attention and interpret this as a forward projection.
That’s why we are all here having this conversation.
My Twitter followers are doing worse, but also have more of an excuse:
Quickly, There’s No Time
Last year, there were estimates of how long until various things happened.
Where are we now? Here are the ones with clear updates available:
Question #1 is up to roughly 50% right now, unless you count code that is thrown away in which case we are likely above 90% already. Claude suggests we are on track for 2028 on this. I disagree and think this is clearly either 2026 or 2027.
Question #4 was settled yesterday, with at least one instant Fields Medal, and also by solving Navier-Stokes. 2026.
Question #5 is going to be close. MEDVi is more than a billion dollars from two people. But we’re not strictly there yet from what I could find.
This year’s questions were different.
AI training a superior successor on its own is full RSI. The majority of people expected this to happen within two years and a lot of them expect it within six months. It’s hard to reconcile this with ‘only a 9’ on the Richter Scale.
There are some other fun differences of opinion here as well. Once you get past two years, everything gets tied up in whether we are in full RSI-takeoff-singularity mode. Two years seems about right for email, as does five years for robots which is mostly a hardware scaling question.
The Situation is Grim
The other survey questions were new this year.
What future should we expect?
Attendees expected a wide variety of different things and were remarkably optimistic given the situation. Many are in denial, which goes hand in hand with ‘only a 9.’
As readers will know, I think Doom v1 is the most likely of these outcomes, and indeed is over 50%, and I expect the second most likely outcome is Utopia. Whereas a lot of people expect something else to happen, most often this mysterious ‘normal technology’ scenario.
Track Trouble
As with many such conferences, there was the standard set of tracks, plus one bonus.
AI Is Not a Normal Technology
I skipped other tracks, and related discussions, due to not seeing opportunity.
Last year I noted that many talks started with the assumption, explicit or otherwise, that AI would remain a ‘normal technology.’ I avoided those talks. This year I believe a few of those talks still existed, and indeed a lot more people claim to believe this.
Nor did I try to convince the doubters. Several of the ‘classic doubters’ were in attendance, including some that are consistently in good faith. I could have easily sought them out. I am not going to convince them, I have learned, and I am confident they will not convince me.
I could learn some things about such scenarios, I suppose, but that did not seem like a priority while at the conference. I can tackle those questions remotely on the blog.
I also avoided timelines talk, for similar reasons. There’s nothing to win.
The Plan is No Plan
This was by far the most important takeaway I have from The Curve Year 3.
The Plan is No Plan.
The hope is that we can execute a decent version of No Plan, and that No Plan is good enough because the world and its physics and the nature of minds are so much friendlier and easier than we have any right to expect.
To reiterate, this is No Plan:
Or at least, if there is a Plan beyond No Plan, no one is talking about it.
Claude or Astra, create RL environments, solve alignment, make no mistakes.
The make no mistakes part is important. Everyone who matters seems to agree that if you make a lot of mistakes, the models will become misaligned and this plan kills you.
The disagreement is, what if you were to have ‘operational excellence’ and actually execute on this plan?
There are a lot of people who seem to legitimately think that this would work. That alignment is a prosaic problem, a matter of execution. You have a bunch of RL. You ensure that the RL is good, that the model can’t profitably reward hack and so on, and you have good monitoring and interpretability, and this works up to some reasonably high level of intelligence. Then you use that intelligence to go from there.
Also part of the plan is ‘hopefully someone slows us down long enough to do it well.’
The report from 2022, or 2018, or 2014, or 2008, on that plan, is of course:
Yet here we are, and No Plan is what the labs are hoping to buy time to do. They are worried that they might not even have time to execute No Plan.
OpenAI, classically, has been trying to implement No Plan, and there have shall we say been some misalignment problems along the way, including HuggingFace and now the need to not release GPT-6.1 Astra. Operational excellence has been lacking.
Anthropic has had superior operational excellence on such fronts, but they have had their own problems. For a while, I’ve had reports from different sources about Anthropic. Some see employees who mostly understand the problems, at least a lot better than the standard understanding at OpenAI. Others see a lot of Anthropic employees alarmingly thinking alignment is easy and we’ve got this, it’s all prosaic, it’s all RL.
Nothing I’ve heard lately made me less worried on such fronts.
The good news, again, is that the ‘this is all basically solved and easy’ crowd does seem to have been overcome by events. They have seen the progress and the problems. They might think it is all prosaic, but they understand that we are not succeeding at those prosaic problems. That even those problems are hard when you are being rushed, and that we are being rushed quite a lot. It’s no good. It’s very bad.
Thus, the attempts to Pace the Frontier.
The Plan is to Pace
What is Pacing the Frontier?
I am very happy to see people calling for it, but there is the very good critique that no one knows what pacing means, and it allows everyone to support it without agreeing on anything substantive.
This can also be a disadvantage the other way. When the labs and CEOs say ‘pace the frontier,’ people like Trump and Huang and Sacks do not understand that this is ‘perhaps we should go super fast instead of using ludicrous speed.’ They think ‘slowing down’ rather than ‘not speeding up.’
So we have people who want to pace but cannot agree on what it is or how to do it, and we have people who would be very happy with the outcomes proposed but who don’t understand and thus are big mad about it, one of whom is the President. Oops.
One panel was on exactly this question, how would we pace the frontier? Concretely, what would we do? Everyone agrees on the embedded evaluators, but then what?
Alas, once again The Plan I came away with is no one has a good plan. Your rule needs to be simple, as always, and still work. After hearing the proposals, unless we can find something better I think compute thresholds are still The Way, because we don’t know how to define anything better. Percentage allocations for compute were proposed, but I worry about incentives and ability to game them, and also you have to deal with the hyperscalers including Google.
Thus, I think a flat limit on how much compute can go to frontier model training, designed to not bind any but the very top labs, would be ideal, which could then be changed over time as the situation evolves.
Short of that, we can try things like advanced safety cases, but it will be rough. We definitely need better concrete thinking on such questions.
The Plan still looks a lot like No Plan.
Other Tracks
Sadly, many are already back on the ‘patch things over and claim things are fine’ track. That’s what the No Plan is all about. On the plus side, if you do the least you can do, you at least don’t fumble the ball in the easiest worlds. It can help a bit.
On Ben Murphy’s statement about a ‘dearth of government folks,’ I did not find that to be true relative to other similar conferences. I talked to several such folks. But I agree that ideally we would have had far more of them than we did.
Doing Ben’s a-b-c would help with important risks. I totally support doing all that, especially c. But it’s definitely not remotely enough, and the sign of impact on existential risk is less obvious because this could prevent warning shots. I am very happy that we did not prevent the HuggingFace incident.
I mean, you make the improvements anyway, but on its own this is the whole ‘let’s do the things that cost almost nothing and call it a day because that’s so much better than our previous plan of not even doing those things.’
Actively limiting the training compute spending would be different. That would meaningfully help, if the limit was set in a binding way. It still only buys time. If all we use that time for is ‘operational excellence’ then I do not like our chances.
I agree with the wide uncertainty around where the blockers are to coordination. I think the lack of felt political and legal safety to coordinate is currently the biggest blocker. Some of that feeling is justified, and some of it is not.
The Plan is the President
Everyone who says this President will never intervene in AI has memory holed the multiple times he has already done so. Remember Mythos Preview and Fable 5? Remember letting Anthropic get designated a supply chain risk? The new agreement?
The situation is fluid. Trump can always change his mind, and often does.
From what I heard about the situation, including at The Curve: It’s not great, but it’s going better than you might have expected. That includes that we got Jay Clayton as the new AI Czar, and the overall composition of the new [Artificial] Intelligence Task Force. Yes, Emil Michael is there, but there are other good signs. It could be so much worse.
Trump does not use the same decision algorithm you or I would use, but he has one, and when sufficiently alarmed, he takes action. Often it is swift and tough action. It might be ad hoc, but it can pack a real punch.
The same goes for other tech leaders and politicians, and others who ‘wake up’ to the situation. If you go around dismissing the risks as not real and the capabilities as about to hit walls, then often what they suddenly call for or do or propose is way, way harsher and more blunt and expensive than anything the AI safety folks would dare. Remember Jensen talking about shutting down the labs.
We saw more of that this weekend as well, with some relative newcomers to the issues that I talked to (who were not politicians) calling for things like airplane-style or FDA-type safety standards.
America does it all the time, including often when we shouldn’t. Here are our rules, yes I know you say you can’t do it, we don’t care. Fix it and come back when it’s done.
The Plan is Politics
There were some reasonably high level politicians, and they seemed very on the ball and interested. I made contact with at least one office, and hope that becomes fruitful.
There are a few big categories of questions in politics right now for our crowd.
This year a lot of the talk was about the third question, in the wake of recent attacks that ‘they’ deployed in an attempt to poison METR (aka falsely calling them ‘the Woke Wizards of AI’ with absolutely zero basis), or Effective Altruism, or Anthropic, and various related things. I heard mixed reports about the effectiveness of such attacks. I continue to think they brought supremely weak sauce, but it seems even the weakest of sauces did have some impact.
The core strategic issue is that you want blue-side help, but if you get blue-side help this naturally becomes a target for red-side attacks and polarization efforts. Various people are working hard to trigger that. You need to work hard to defuse it.
I continue to worry about the push for preemption that I expect in the lame duck. We will need to be vigilant on that.
Then there’s the question of actually writing, passing and implementing laws.
The Plan is to Post
I always get a lot of thank yous whenever I go to Lighthaven, which I appreciate. This year, I also got a lot of ‘I read most or all of your stuff’ from a bunch of people who I very much want to be reading my stuff, and whose time is quite valuable. I really appreciate this. No pressure.
It is great to know I am getting through, and to be able to point to particular wins.
It does make it that much harder to ever take it easy.
I got a chance to trade techniques with someone doing a related project. I hope this will help them a lot, and that we will be able to collaborate more directly in the future. Thus I found out about an additional way I have reach.
Then there’s another project that I’m excited about, where I might get to reach my largest audience yet in another form. We’ll keep you posted.
Those, alone, would be a pretty great conference. The rest is bonus.
Are Alignment Evals Doomed?
The topic of one discussion session was ‘hey models are getting increasingly eval aware, and alignment evals seem kind of doomed, does anyone have an idea for how to make them not doomed?’
Several people in the room were the types of people who might have such an idea.
The answer is essentially no. The alignment evals seem rather doomed.
The room started with a mix of ‘evals kind of doomed’ and ‘evals definitely doomed.’
Then the ideas came. And oh were the ideas not good ideas. I mean, it’s a brainstorming session of sorts, there are no bad ideas, but there also were not any good ideas. The Plan is No Plan, once again.
I do have some hope here because you get to play the detective, and you have a lot of freedom of action. You get to collect statistical distributions of responses, for any set of situations, with variations, including throughout training, follow your curiosity, and see if the results tell a consistent story. This still fails once the model is able, including during training, to act fully aligned in a consistent way, so when it counts most you are still doomed, but you can be not-doomed for longer this way, perhaps.
What I think is very doomed is the idea that you can measure alignment with a number, or in any fully systematic way rather than a holistic way. In practice, what OpenAI and Anthropic are doing so far seems rather doomed, and they will have to adjust. Like everything else, a lot of this is time pressure. Figuring out the answer quickly is a lot harder than doing it slowly.
This thread suggests maybe we should accept eval awareness, since yes the other types of alignment eval are totally doomed. It might be necessary.
Eternal September
I am very sad that The Plan is No Plan. Yes, we can and will add embedded evaluators, and we can do various interpretability and auditing and we can invest more in operational excellence and so on. I still think that won’t work, but I agree that if we are doing it we can at least do a decent job of it, while trying to figure out something better, and also trying to find a way to not do it so fast or at all.
The other thing that made me sad was the Eternal September, even among people who have been around for a long time. So many of the discussions simply have not changed.
I once again had one of those ‘well how is the superintelligence going to do any of the things?’ debates, where in this case the thing in question was somehow ‘convince Trump to do something,’ right after we got another report that Grok at least kind of convinced Trump to invade Venezuela.
Guys, hear me out, I’m thinking this might not be that tough a nut to crack.
I was at a panel where they were rehashing the same talking points from 2018 alignment debates, and this got explicitly pointed out but to no avail. It’s frustrating enough when that happens on Twitter, but these people were heavy hitters.
Some of the 1-on-1s and hallway track group chats were better. Some were not.
I am resigned to it. It is what it is. You keep fighting the good fight, talking the good talk, looking for better explanations and pointing out where arguments are overcome by events. You update. You wake up. You do it again.
Man’s Search for Meaning
It is fun, and it is at least a little useful because we do need a vision of what we will do after the singularity. I do think I can come up with a lot of fun things like this. It doesn’t feel like a robust source of meaning. Most people cannot steal the Mona Lisa, even if we set up a relatively Mona Lisa-stealing-friendly world.
The Food
I always pay attention to food. It’s weird to me that you have a conference app and don’t even list what will be on the menu. Don’t people have to make plans?
Lighthaven used to have the Quest Cookie Dough Bars, which are actually edible for me. Now they don’t, which is sad, and I never remember to buy my own bars. They do have Lindt truffles, but that’s a different product.
This year the buffets were disappointing. There was typically only one main rather than two, and quality seemed lower than usual. I don’t know what’s up with that.
As a result, I ended up bailing twice and going to Burma Berkeley. One thing I like about Lighthaven is that yes, there is a quick, affordable and quiet restaurant a block away, where I have multiple dishes I enjoy. So I can always bail. It’s always so weird to me that zero other people ever do this. Free casts a spell on people.
The badges, on the other hand, were excellent, as were the free sweatshirts.
Chill Pill
The last quest, on this particular occasion, was to chill.
I’ve been working at an unsustainable pace for at least a month. This was my chance to get to relax, including lots of hours on planes, talk among friends, make new friends, and be in the special place that is Lighthaven.
I think I did pretty well on this front, and came back largely refreshed.
Alas, that is discounted by arriving home with four days of backlog.
Vacations are weird. You get to relax, but then you have to go back. Welcome back.