A week ago, Anthropic quietly weakened their ASL-3 security requirements. Yesterday, they announced ASL-3 protections.
I appreciate the mitigations, but quietly lowering the bar at the last minute so you can meet requirements isn't how safety policies are supposed to work.
(This was originally a tweet thread (https://x.com/RyanPGreenblatt/status/1925992236648464774) which I've converted into a LessWrong quick take.)
9 days ago, Anthropic changed their RSP so that ASL-3 no longer requires being robust to employees trying to steal model weights if the employee has any access to "systems that process model weights".
Anthropic claims this change is minor (and calls insiders with this access "sophisticated insiders").
But, I'm not so sure it's a small change: we don't know what fraction of employees could get this access and "systems that process model weights" isn't explained.
Naively, I'd guess that access to "systems that process model weights" includes employees being able to operate on the model weights in any way other than through a trusted API (a restricted API that we're very confident is secure). If that's right, it could be a hig...
I'd been pretty much assuming that AGI labs' "responsible scaling policies" are LARP/PR, and that if an RSP ever conflicts with their desire to release a model, either the RSP will be swiftly revised, or the testing suite for the model will be revised such that it doesn't trigger the measures the AGI lab doesn't want to trigger. I. e.: that RSPs are toothless and that their only purposes are to showcase how Responsible the lab is and to hype up how powerful a given model ended up.
This seems to confirm that cynicism.
(The existence of the official page tracking the updates is a (smaller) update in the other direction, though. I don't see why they'd have it if they consciously intended to RSP-hack this way.)
Employees at Anthropic don't think the RSP is LARP/PR. My best guess is that Dario doesn't think the RSP is LARP/PR.
This isn't necessarily in conflict with most of your comment.
I think I mostly agree the RSP is toothless. My sense is that for any relatively subjective criteria, like making a safety case for misalignment risk, the criteria will basically come down to "what Jared+Dario think is reasonable". Also, if Anthropic is unable to meet this (very subjective) bar, then Anthropic will still basically do whatever Anthropic leadership thinks is best whether via maneuvering within the constraints of the RSP commitments, editing the RSP in ways which are defensible, or clearly substantially loosening the RSP and then explaining they needed to do this due to other actors having worse precautions (as is allowed by the RSP). I currently don't expect clear cut and non-accidental procedural violations of the RSP (edit: and I think they'll be pretty careful to avoid accidental procedural violations).
I'm skeptical of normal employees having significant influence on high stakes decisions via pressuring the leadership, but empirical evidence could change the views of Anthropic leadership.
Ho...
How you feel about this state of affairs depends a lot on how much you trust Anthropic leadership to make decisions which are good from your perspective.
Another note: My guess is that people on LessWrong tend to be overly pessimistic about Anthropic leadership (in terms of how good of decisions Anthropic leadership will make under the LessWrong person's views and values) and Anthropic employees tend to be overly optimistic.
I'm less confident that people on LessWrong are overly pessimistic, but they at least seem too pessimistic about the intentions/virtue of Anthropic leadership.
For the record, I think the importance of "intentions"/values of leaders of AGI labs is overstated. What matters the most in the context of AGI labs is the virtue / power-seeking trade-offs, i.e. the propensity to do dangerous moves (/burn the commons) to unilaterally grab more power (in pursuit of whatever value).
Stuff like this op-ed, broken promise of not meaningfully pushing the frontier, Anthropic's obsession & single focus on automating AI R&D, Dario's explicit calls to be the first to RSI AI or Anthropic's shady policy activity has provided ample evidence that their propensity to burn the commons to grab more power (probably in name of some values I would mostly agree with fwiw) is very high.
As a result, I'm now all-things-considered trusting Google DeepMind slightly more than Anthropic to do what's right for AI safety. Google, as a big corp, is less likely to do unilateral power grabbing moves (such as automating AI R&D asap to achieve a decisive strategic advantage), is more likely to comply with regulations, and is already fully independent to build AGI (compute / money / talent) so won't degrade further in terms of incentives; additionally D. Hass...
Not the main thrust of the thread, but for what it's worth, I find it somewhat anti-helpful to flatten things into a single variable of "how much you trust Anthropic leadership to make decisions which are good from your perspective", and then ask how optimistic/pessimistic you are about this variable.
I think I am much more optimistic about Anthropic leadership on many axis relative to an overall survey of the US population or Western population – I expect them to be more libertarian, more in favor of free speech, more pro economic growth, more literate, more self-aware, higher IQ, and a bunch of things.
I am more pessimistic about their ability to withstand the pressures of a trillion dollar industry to shape their incentives than the people who are at Anthropic.
I believe the people working there are siloing themselves intellectually into an institution facing incredible financial incentives for certain bottom lines like "rapid AI progress is inevitable" and "it's reasonably likely we can solve alignment" and "beating China in the race is a top priority", and aren't allowed to talk to outsiders about most details of their work, and this is a key reason that I expect them to sc...
I think the main thing I want to convey is that I think you're saying that LWers (of which I am one) have a very low opinion of the integrity of people at Anthropic, but what I'm actually saying that their integrity is no match for the forces that they are being tested with.
I don't need to be able to predict a lot of fine details about individuals' decision-making in order to be able to have good estimates of these two quantities, and comparing them is the second-most question relating to whether it's good to work on capabilities at Anthropic. (The first one is a basic ethical question about working on a potentially extinction-causing technology that is not much related to the details of which capabilities company you're working on.)
Employees at Anthropic don't think the RSP is LARP/PR. My best guess is that Dario doesn't think the RSP is LARP/PR.
Yeah, I don't think this is necessarily in contradiction with my comment. Things can be effectively just LARP/PR without being consciously LARP/PR. (Indeed, this is likely the case in most instances of LARP-y behavior.)
Agreed on the rest.
I think security is legitimately hard and can be costly in research efficiency. I think there is a defensible case for this ASL-3 security bar being reasonable for the ASL-3 CBRN threshold, but it seems too weak for the ASL-3 AI R&D threshold (hopefully the bar for things like this ends up being higher).
GPT-6 Astra appears to be a massive jump in opaque reasoning ability: it looks like it can solve hard competition math problems entirely in its head (as in, without verbalized reasoning) while prior AIs could solve basic word problems.
This seems extremely concerning!
That is, if these benchmark results are representative (see the highlighted caveats in the image, I'm particularly worried about contamination).
Related to this, UK AISI found Astra has much worse monitorability.
I'd guess this jump is downstream of architectural changes (with increased serial depth) though a normal large pretrain scale up is a plausible cause. If the next few model generations involve similar jumps (presumably these jumps would be downstream of a transition to full-on opaque reasoning architectures with extreme depth), then chain-of-thought would no longer be a meaningful oversight tool.
I suspect that lack of serial depth/opaque reasoning ability has been one reason why none of the models have been able to beat me at chess yet, despite doing more impressive things at other domains. I think there's a chance it will be stronger than other models for this reason, and I can post the game and analysis here if people are interested.
Anyone know what a "takes-a-human-30-minutes" math competition problem looks like?
I'm curious to see more concretely what kind of problems Astra can do "in its head".
Wait what!? That should be way more than 30min for a human to complete.
Edit: Maybe the way they're implementing it just means that the model reasons a bunch in the output field rather than the CoT field?
An economist and a futurist walk into a bar.
The economist takes a sip of his drink. "Ugh, if only people understood basic economics. High-skilled immigration alone would do wonders for US growth."
Futurist: "Oh yeah? Say 100 million immigrants moved to the US, each matching the best human experts in every economically relevant field. Big deal?"
Economist: "Massive. Transformative."
Futurist: "What if they also worked longer hours and faster than any American?"
Economist: "Even better."
Futurist: "What if they were extremely frugal — consuming only the bare minimum needed to keep working?"
Economist: "A near-100% savings rate? Better still!"
Futurist: "What if they were very clumsy and physically weak, so they could only do some kinds of work?"
Economist: "They could still do all cognitive labor — that's over half all wages! Somewhat less good, sure. Still transformative."
Futurist: "What if their skin was grey, almost metallic, from some kind of accident?"
Economist: "Who cares?!"
Futurist: "What if they were AIs?"
Economist: "3% growth per year, tops. There'd be bottlenecks. Honestly, the people predicting explosive growth from AI should learn some economics."
We have extrapolated the number of human-level AGIs based on the number of human-level AGIs in previous decades.
Please don't overly index on Tyler Cowen. He's been clearly not trying very hard to be correct on things in this area. There's probably better examples of economists who have bad takes and are more serious about it.
I've heard from a credible source that OpenAI substantially overestimated where other AI companies were at with respect to RL and reasoning when they released o1. Employees at OpenAI believed that other top AI companies had already figured out similar things when they actually hadn't and were substantially behind. OpenAI had been sitting the improvements driving o1 for a while prior to releasing it. Correspondingly, releasing o1 resulted in much larger capabilities externalities than OpenAI expected. I think there was one more case like this either from OpenAI or GDM where employees had a large misimpression about capabilities progress at other companies causing a release they wouldn't do otherwise.
One key takeaway from this is that employees at AI companies might be very bad at predicting the situation at other AI companies (likely making coordination more difficult by default). This includes potentially thinking they are in a close race when they actually aren't. Another update is that keeping secrets about something like reasoning models worked surprisingly well to prevent other companies from copying OpenAI's work even though there was a bunch of public reporting (and presumably many rumors) about this.
One more update is that OpenAI employees might unintentionally accelerate capabilities progress at other actors via overestimating how close they are. My vague understanding was that they haven't updated much, but I'm unsure. (Consider updating more if you're an OpenAI employee!)
Alex Mallen also noted a connection with people generally thinking they are in race when they actually aren't: https://forum.effectivealtruism.org/posts/cXBznkfoPJAjacFoT/are-you-really-in-a-race-the-cautionary-tales-of-szilard-and
I think:
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'.
I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident.
Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep...
I'm currently working as a contractor at Anthropic in order to get employee-level model access as part of a project I'm working on. The project is a model organism of scheming, where I demonstrate scheming arising somewhat naturally with Claude 3 Opus. So far, I’ve done almost all of this project at Redwood Research, but my access to Anthropic models will allow me to redo some of my experiments in better and simpler ways and will allow for some exciting additional experiments. I'm very grateful to Anthropic and the Alignment Stress-Testing team for providing this access and supporting this work. I expect that this access and the collaboration with various members of the alignment stress testing team (primarily Carson Denison and Evan Hubinger so far) will be quite helpful in finishing this project.
I think that this sort of arrangement, in which an outside researcher is able to get employee-level access at some AI lab while not being an employee (while still being subject to confidentiality obligations), is potentially a very good model for safety research, for a few reasons, including (but not limited to):
Yay Anthropic. This is the first example I'm aware of of a lab sharing model access with external safety researchers to boost their research (like, not just for evals). I wish the labs did this more.
[Edit: OpenAI shared GPT-4 access with safety researchers including Rachel Freedman before release. OpenAI shared GPT-4 fine-tuning access with academic researchers including Jacob Steinhardt and Daniel Kang in 2023. Yay OpenAI. GPT-4 fine-tuning access is still not public; some widely-respected safety researchers I know recently were wishing for it, and were wishing they could disable content filters.]
I'd be surprised if this was employee-level access. I'm aware of a red-teaming program that gave early API access to specific versions of models, but not anything like employee-level.
It was a secretive program — it wasn’t advertised anywhere, and we had to sign an NDA about its existence (which we have since been released from). I got the impression that this was because OpenAI really wanted to keep the existence of GPT4 under wraps. Anyway, that means I don’t have any proof beyond my word.
(I'm a full-time employee at Anthropic.) It seems worth stating for the record that I'm not aware of any contract I've signed whose contents I'm not allowed to share. I also don't believe I've signed any non-disparagement agreements. Before joining Anthropic, I confirmed that I wouldn't be legally restricted from saying things like "I believe that Anthropic behaved recklessly by releasing [model]".
I think I could share the literal language in the contractor agreement I signed related to confidentiality, though I don't expect this is especially interesting as it is just a standard NDA from my understanding.
I do not have any non-disparagement, non-solicitation, or non-interference obligations.
I'm not currently going to share information about any other policies Anthropic might have related to confidentiality, though I am asking about what Anthropic's policy is on sharing information related to this.
Around the early o3 announcement (and maybe somewhat before that?), I felt like there were some reasonably compelling arguments for putting a decent amount of weight on relatively fast AI progress in 2025 (and maybe in 2026):
I u...
I basically agree with this whole post. I used to think there were double-digit % chances of AGI in each of 2024 and 2025 and 2026, but now I'm more optimistic, it seems like "Just redirect existing resources and effort to scale up RL on agentic SWE" is now unlikely to be sufficient (whereas in the past we didn't have trends to extrapolate and we had some scary big jumps like o3 to digest)
I still think there's some juice left in that hypothesis though. Consider how in 2020, one might have thought "Now they'll just fine-tune these models to be chatbots and it'll become a mass consumer product" and then in mid-2022 various smart people I know were like "huh, that hasn't happened yet, maybe LLMs are hitting a wall after all" but it turns out it just took till late 2022/early 2023 for the kinks to be worked out enough.
Also, we should have some credence on new breakthroughs e.g. neuralese, online learning, whatever. Maybe like 8%/yr? Of a breakthrough that would lead to superhuman coders within a year or two, after being appropriately scaled up and tinkered with.
Interestingly, reasoning doesn't seem to help Anthropic models on agentic software engineering tasks, but does help OpenAI models.
I use 'ultrathink' in Claude Code all the time and find that it makes a difference.
I do worry that METR's evaluation suite will start being less meaningful and noisier for longer time horizons as the evaluation suite was built a while ago. We could instead look at 80% reliability time horizons if we have concerns about the harder/longer tasks.
I'm overall skeptical of overinterpreting/extrapolating the METR numbers. It is far too anchored on the capabilities of a single AI model, a lightweight scaffold, and a notion of 'autonomous' task completion of 'human-hours'. I think this is a mental model for capabilities progress that will lead to erroneous predictions.
If you are trying to capture the absolute frontier of what is possible, you don't only test a single-acting model in an empty codebase with limited internet access and scaffolding. I would personally be significantly less capable at agentic coding if I only used 1 model (like replicating subliminal learning in about 1 hour of work + 2 hours of waiting for fine-tunes on the day of the release) with l...
Interestingly, reasoning doesn't seem to help Anthropic models on agentic software engineering tasks, but does help OpenAI models.
Is there a standard citation for this?
How do you come by this fact?
Why should we think that the relevant progress driving non-formal IMO is very important for plausibly important capabilities like agentic software engineering? [...] if the main breakthrough was in better performance on non-trivial-to-verify tasks (as various posts from OpenAI people claim), then even if this generalizes well beyond proofs this wouldn't obviously particularly help with agentic software engineering (where the core blocker doesn't appear to be verification difficulty).
I'm surprised by this. To me it seems hugely important how fast AIs are improving on tasks with poor feedback loops, because obviously they're in a much better position to improve on easy-to-verify tasks, so "tasks with poor feedback loops" seem pretty likely to be the bottleneck to an intelligence explosion.
So I definitely do think that "better performance on non-trivial-to-verify tasks" are very important for some "plausibly important capabilities". Including agentic software engineering. (Like: This also seems related to why the AIs are much better at benchmarks than at helping people out with their day-to-day work.)
I agree with the core message in Dario Amodei's essay "The Adolescence of Technology": AI is an epochal technology that poses massive risks and humanity isn't clearly going to do a good job managing these risks.
(Context for LessWrong: I think it seems generally useful to comment on things like this. I expect that many typical LessWrong readers will agree with me and find my views relatively predictable, but I thought it would be good to post here anyway.)
However, I also disagree with (or dislike) substantial parts of this essay:
When I say "misaligned AI takeover", I mean that the acquisition of resources by the AIs would reasonably be considered (mostly) illegitimate, some fraction of this could totally include many humans surviving with a subset of resources (though I don't currently expect property rights to remain intact in such a scenario very long term). Some of these outcomes could be avoid literal coups or violence while still being illegitimate; e.g. they involve doing carefully planned out capture of governments in ways their citizens/leaders would strongly object to if they understood and things like this drive most of the power acquisition.
I'm not counting it as takeover if "humans never intentionally want to hand over resources to AIs, but due to various effects misaligned AIs end up with all of the resources through trade and not through illegitimate means" (e.g., we can't make very aligned AIs but people make various misaligned AIs while knowing they are misaligned and thus must be paid wages and AIs form a cartel rather having wages competed down to subsistence levels and thus AIs end up with most of the resources).
I currently don't expect human disempowerment in favor of AIs (that aren't appointed successors) conditional on no misalignmed AI takeover, but agree this is possible; it doesn't form a large enough probability to substantially alter my communication.
Anthropic, GDM, and xAI say nothing about whether they train against Chain-of-Thought (CoT) while OpenAI claims they don't[1].
I think AI companies should be transparent about whether (and how) they train against CoT. While OpenAI is doing a better job at this than other companies, I think all of these companies should provide more information about this.
It's particularly striking that Anthropic says nothing about whether they train against CoT given their system card (for 4.5 Sonnet) is very thorough and includes a section on "Reasoning faithfulness" (kudos to them for providing so much other information!). Naively, this seems like some evidence that they are training against CoT and didn't want to disclose this.[2]
I think we should try to create an incentive gradient that pushes AI companies to disclose information even if that information makes them look bad, so we should make companies feel some heat for not disclosing (reasonably) important information like how/whether they are training against CoT.
It's worth noting that some versions of training against CoT might be totally fine (as in, they don't incentivize obfuscation of misaligned reasoning) and that some particular appr...
Anthropic has now clarified this in their system card for Claude Haiku 4.5:
Thanks to them for doing this!
See also Sam Bowman's tweet thread about this.
It's good that Anthropic's system cards contain a lot of useful information on misalignment and risk. It's also good they are putting out detailed sabotage risk reports that articulate their views. [1] I appreciate the hard work of many employees going into these reports.
I wish other companies released similarly informative reports. [2]
I'm not necessarily claiming I agree with their risk reports. For instance, I have at least some moderately important disagreements with the Mythos Preview sabotage risk report update. But having lots of detail, including sufficient detail that I can get a decent sense of where I disagree with the analysis, is praiseworthy. ↩︎
I'm not claiming that this is the most leveraged thing for safety-motivated employees at other companies to work on, but I do think it would be good for AI companies to do better (without trading off against other safety efforts). ↩︎
Recently, various groups successfully lobbied to remove the moratorium on state AI bills. This involved a surprising amount of success while competing against substantial investment from big tech (e.g. Google, Meta, Amazon). I think people interested in mitigating catastrophic risks from advanced AI should consider working at these organizations, at least to the extent their skills/interests are applicable. This both because they could often directly work on substantially helpful things (depending on the role and organization) and because this would yield valuable work experience and connections.
I worry somewhat that this type of work is neglected due to being less emphasized and seeming lower status. Consider this an attempt to make this type of work higher status.
Pulling organizations mostly from here and here we get a list of orgs you could consider trying to work (specifically on AI policy) at:
Kids safety seems like a pretty bad thing to focus on, in the sense that the vast majority of kids safety activism causes very large amounts of harm (and it helping in this case really seems like a “a stopped clock is right twice a day situation”).
The rest seem pretty promising.
I strongly agree. I can't vouch for all of the orgs Ryan listed, but Encode, ARI, and AIPN all seem good to me (in expectation), and Encode seems particularly good and competent.
Plausibly, but their type of pressure was not at all what I think ended up being most helpful here!
They also did a lot of calling to US representatives, as did people they reached out to.
ControlAI did something similar and also partnered with SiliConversations, a youtuber, to get the word out to more people, to get them to call their representatives.
I think it's both true that LessWrong (LW) has a bunch of issues and that it would be better if much more discourse happened there rather than on X/Twitter.
Some claims that all seem true to me despite being in tension:
I have found lesswrong a valuable venue. For what it’s worth, I’ve been attacked way less here than on X….
I recommend AI lab employees post and lurk more here. LW is a bit of an echo chamber, but at least it’s a different echo chamber than the one we spend most of our time in.
To add something: given this forum’s population, it is quite noteworthy and admirable how welcoming it is to someone like me and other lab employees. I can’t imagine a vegan forum allowing meat company employees to post, even if they work in the department for humane treatment.
Here I'll reflect on things that make me (an Anthropic employee, though I'm speaking for myself only) engage less on LW than I otherwise would. (Some of these points aren't LW-specific and also apply to other interactions, e.g. in-person conversations, I have with people in the AI safety community.)
...2. Blending of advocacy and object-level discussion. When I engage on LW, it's typically because I think there's an important object-level point worth discussing. But once I enter the conversation, it sometimes feels like people stop being curious about the object-level point and instead move into an advocacy mode where their goal is to get Anthropic to act differently in light of their point (which is assumed correct). That is, instead of continuing the object-level discussion, my interlocutors sometimes move to criticize Anthropic or the beliefs of Anthropic staff, or to ask Anthropic (via me) to do something differently.
- (I especially find this frustrating when my interlocutors implicitly assume that I have the Anthropic "house belief" on some topic when I don't or feel unsure.)
- Possible mitigations:
- LW posters could engage in object-level discussion longer before moving to advocacy.
- When LW posters want to argue against what they understand to be the Anthropic "house belief," they could write things like "My understanding is that many Anthropic staff believe something like '...' I think this view wrong because ..."
- LW posters could more clearly flag and separate advocacy from object
I agree! Better and more AI discourse on LW would be great.
One narrow disagreement:
E.g., it would be better if people on LW applied something more like typical researcher norms to research outputs (e.g., for many types of concerns, email the author with the concerns and see if they fix before posting publicly) and tried to avoid their criticism being unnecessarily rude (though not necessarily less hostile)
I endorse posting publicly because
It would be good if many more AI company employees were interested in seriously trying to form detailed views about the future of AI and consequences of this.
In general do you see good ways to help with that? (I probably couldn't / wouldn't work on this but others might be interested.)
A couple thoughts:
It’s up to LW moderators to enforce norms, but I would suggest that if you want to impose social costs on lab employees, you do it outside LW.
I post on LW for intellectual discussion, and not to make friends. If people are rude here, it won’t cause me to quit OpenAI or change what I’m doing, but it can cause me to stop posting or reading. That may well be the desired outcome, though I personally think it would be unfortunate.
I think it is unhealthy if all discussion on AI safety that actually impacts frontier models happens inside the labs without discussion between people in the labs and people outside them. And at its best LW can facilitate these.
LessWrong has a broader merit than merely "intellectual discussion", as it a place where people figure out a wide range of their stances on a wide range of topics, and has given birth to a rich community of people with many mutual interests beyond intellectual discussion. It's also really important for LessWrong to think about what incentives its members are under, and discussion here directly or indirectly affects many strategic and tactical decisions at many organizations.
This makes the role of LessWrong meaningfully different from just narrow technical journal, and most importantly means that of course people should gain or lose social standing based on a wide variety of actions they take on or off LW.
This doesn't mean every post should be a place to discuss people's relationship to lab employees, or adjacent topics, but suggesting that such a topic is completely outside of the remit of LessWrong, as I think you are saying here, seems quite wrong to me. People will want to prosecute conflict and standing and credit allocation on LessWrong, and I don't really see a way of doing that without allowing some amount of social consequences to be imposed (at the very least for violatin...
I currently think it's best to avoid applying social censure in object level discussions of specific topics on LW like Sam discusses here.
Sure, I was just hearing you espouse a much broader view of something like "I wish there was less conflict theory and more mistake theory in explaining differences around people's views on AI", which I often feel has been an effective way of sweeping real conflicts under the rug (and I believe blatant conflicts are the best kind of conflicts). We'd have to discuss specifics to know whether you mean it in cases that I think are appropriate or inappropriate. Marks didn't actually link to any so I am not confident about whether we're on the same page.
That said, I do expect it will be unproductive to treat people as morally reprehensible for doing things where there aren't reasonably widespread norms against doing the thing. Imposing a bunch of social costs could be productive idk... Idk though, and people disagree about what the defaults here are. I'm unlikely to respond to replies on this parenthetical as it doesn't seem very useful to discuss.
This whole passage reads as quite confused to me. I think it is in some ways imperative to track the ethic...
I recently recorded a podcast with Dwarkesh about the potential for (very) fast AI progress and how misaligned AI takeover might happen.
Our conversation focused a lot on threat models from "reward-seeking" AIs. If you're interested in reading more about this, Alex (who works with me at Redwood) has written in a lot more detail about this threat model here
I also talked about how "mundane" misalignment and underelicitation could doom us. I say more about this here.
(I think there are other important threat models, like AIs ending up with (shared) long-run preferences and deciding to fake alignment based on these preferences. For reference, see here and here.)
Some of people at Redwood wrote up some notes to help me prep:
My median for full automation of AI R&D is around late 2030/early 2031. [1] But my "modal"/best guess prediction for this milestone would be significantly earlier (mid 2029).
Here is a summary of my best guess prediction for what happens over the next few years:
EOY 2026:
EOY 2027:
2028:
PSA: Anthropic models don't seem to particularly privilege the explicit thinking field. This makes reinforcement spillover—where training on a model's outputs generalizes to the CoT, making it appear safer—more likely.
While Anthropic models do have an separate explicit thinking field, they don't really use thinking that differently from outputs and aren't that dependent on the thinking field. Sometimes they'll just do their thinking in the output field, the way they talk in the thinking field isn't very distinct from how they talk in outputs, and I believe disabling thinking doesn't have a big effect on coding performance (especially for earlier Anthropic models but even for current Anthropic models). (This is specifically for typical agentic coding; this likely doesn't apply to math.)
This is pretty different from OpenAI models which are way more reliant on the explicit thinking and the thinking is very distinct from how the model talks in the output (based on public CoT examples). I'm uncertain, but I think GDM models are generally more similar to OpenAI than Anthropic models on this axis. In general, Anthropic's models leverage taking lots of actions over doing a bunch of thinkin...
This matches my experience.
These days, when I'm using an Anthropic model, I often turn off "extended thinking" and just ask for CoT the old-fashioned way.
The models seem at least equally capable when using this format (vs. extended thinking), and this format allows me to see the CoT without any summarization -- which is useful for debugging/improving my prompts, and which also allows me to guide the CoT in complex ways like what I describe in the footnote of this comment. It seems plausible that these guidance techniques would also "work" with extended thinking, but the summarizer makes it difficult to confirm that they're working, even if they are.[1]
Anyways... I agree with you in principle about the reinforcement spillover considerations, but IMO the most important implication of this observation is something else that you didn't mention explicitly in your comment -- namely, that the infamous "weirdness" of OpenAI CoTs is not a thing that just automatically happens when you try to reach the current capabilities frontier and don't directly supervise the CoT tokens.
I'm not sure that that precise claim (that the "weirdness [...] just automatically happens") is one that anyone liter...
namely, that the infamous "weirdness" of OpenAI CoTs is not a thing that just automatically happens when you try to reach the current capabilities frontier and don't directly supervise the CoT tokens.
I agree that it's clearly possible to train frontier models that don't have any weirdness in their CoTs. However, my impression is that vanilla GRPO does lead to weird CoTs by default: Reasoning Models Sometimes Output Illegible Chains of Thought seems like evidence that many open models that were trained last year were also heading in the same direction as o3, but didn't quite get there, either because less RL pressure was applied on them or simply because they averted some path-dependent feature of GRPO training that drives CoT toward further illegibility. If this is true, then Anthropic must be doing something quite different from the standard GRPO pipeline that helps them keep reasoning traces legible without optimizing them directly. Here are two hypotheses for what those differences might be:
This comment thread on 1a3orn’s post has a collection of various model’s exhibiting degenerate language usage + Jozdien’s paper (which has since come out: Reasoning Models Sometimes Output Illegible Chains of Thought) I think are all strong evidence that you don’t get human legible english by default from outcome based RL.
I'm somewhat skeptical of that paper's interpretation of the observations it reports, at least for R1 and R1-Zero.
(EDIT: but see Jozdien's reply to this comment, which calls this into doubt)
(EDIT2, added 4/19/26: see this post for more conclusive evidence on this point)
I have used these models a lot through OpenRouter (which is what Jozdien used), and in my experience:
Following up on this! We were able to get a few more CoTs released from o3 after capabilities-focused RL but before safety training: here
tldr:
disclaim disclaim illusions occur at dramatically lower rates the completely degenerate examples in the final o3 seem to be a direct result of safety training. I was surprised by this, as the repeated impression I've gotten from papers and statements by OpenAI is that they didn't think their safety training significantly impacted CoT legibility (but maybe they never literally stated this).When the CoT-author writes any first-person pronoun it's of course natural to think it's referring to itself, but the "itself" in question isn't necessary itself as distinct from the other persona; often it just seems to mean "the language model writing this."
A confounding factor is that every system prompt starts with "You are ChatGPT", so this somewhat gets baked int...
- However... it is apparently very easy to set up an inference server for R1 incorrectly, and if you aren't carefully discriminating about which OpenRouter providers you accept[2], you will likely get one of the "bad" ones at least some of the time.
From what I remember, I did see that some providers for R1 didn't return illegible CoTs, but that those were also the providers marked as serving a quantized R1. When I filtered for the providers that weren't marked as such I think I pretty consistently found illegible CoTs on the questions I was testing? Though there's also some variance in other serving params—a low temperature also reduces illegible CoTs.
I thought it would be helpful to post about my timelines and what the timelines of people in my professional circles (Redwood, METR, etc) tend to be.
Concretely, consider the outcome of: AI 10x’ing labor for AI R&D[1], measured by internal comments by credible people at labs that AI is 90% of their (quality adjusted) useful work force (as in, as good as having your human employees run 10x faster).
Here are my predictions for this outcome:
The views of other people (Buck, Beth Barnes, Nate Thomas, etc) are similar.
I expect that outcomes like “AIs are capable enough to automate virtually all remote workers” and “the AIs are capable enough that immediate AI takeover is very plausible (in the absence of countermeasures)” come shortly after (median 1.5 years and 2 years after respectively under my views).
Only including speedups due to R&D, not including mechanisms like synthetic data generation. ↩︎
My timelines are now roughly similar on the object level (maybe a year slower for 25th and 1-2 years slower for 50th), and procedurally I also now defer a lot to Redwood and METR engineers. More discussion here: https://www.lesswrong.com/posts/K2D45BNxnZjdpSX2j/ai-timelines?commentId=hnrfbFCP7Hu6N6Lsp
I expect that outcomes like “AIs are capable enough to automate virtually all remote workers” and “the AIs are capable enough that immediate AI takeover is very plausible (in the absence of countermeasures)” come shortly after (median 1.5 years and 2 years after respectively under my views).
@ryan_greenblatt can you say more about what you expect to happen from the period in-between "AI 10Xes AI R&D" and "AI takeover is very plausible?"
I'm particularly interested in getting a sense of what sorts of things will be visible to the USG and the public during this period. Would be curious for your takes on how much of this stays relatively private/internal (e.g., only a handful of well-connected SF people know how good the systems are) vs. obvious/public/visible (e.g., the majority of the media-consuming American public is aware of the fact that AI research has been mostly automated) or somewhere in-between (e.g., most DC tech policy staffers know this but most non-tech people are not aware.)
I don't feel very well informed and I haven't thought about it that much, but in short timelines (e.g. my 25th percentile): I expect that we know what's going on roughly within 6 months of it happening, but this isn't salient to the broader world. So, maybe the DC tech policy staffers know that the AI people think the situation is crazy, but maybe this isn't very salient to them. A 6 month delay could be pretty fatal even for us as things might progress very rapidly.
AI is 90% of their (quality adjusted) useful work force (as in, as good as having your human employees run 10x faster).
I don't grok the "% of quality adjusted work force" metric. I grok the "as good as having your human employees run 10x faster" metric but it doesn't seem equivalent to me, so I recommend dropping the former and just using the latter.
After our investigation of the OpenAI / Hugging Face incident, many open questions remain. We spent 6 days on premises with access to the data (we only had access to the entire dataset we used during our last 2 days on premises). The scope of our investigation was also limited: it covered just this incident rather than other similar incidents, didn't include investigating what these agents might have done in other circumstances, and OpenAI stated that the period under investigation ended July 13th.
Here are some of the open questions that seem worthwhile to investigate. (I'd recommend reading the report to understand the context behind these questions!)
Motives:
(we only had access to the entire dataset we used during our last 2 days on premises)
Why? If you had asked for more time they would have said no? What's the nature of the politics here (insofar as it's not itself a betrayal to talk about the politics)?
Sometimes people find it mysterious or surprising that current AIs can't fully automate difficult tasks given how smart they seem. I don't find this very confusing.
Current LLMs are just not that "smart" (yet). They compensate using very broad knowledge and strong heuristics that are mostly domain-specific. In other words, they have high crystallized intelligence but lower fluid intelligence.
In humans, crystallized and fluid intelligence are very correlated due to limited time for learning and limited capacity for knowledge/memorization, but AIs can train for longer and (for unclear reasons) seem to have much higher capacity for knowledge/memorization. So, AI capabilities are often overestimated by naive comparisons. [1] I do find it somewhat mysterious/surprising that humans have such poor memory despite seemingly having so many parameters, and that some people have much better memory than others.
My view is that this crystallized vs fluid distinction is maybe the first principal component of human vs AI differences and this explains a lot of the differences, but not all of them. We...
Pretraining memorizes all the facts in the world, but only gives weak fluid intelligence (in-context learning). RLVR trains crystallized intelligence that expresses itself in-context as strong fluid intelligence (which isn't fake or illusory, within its scope), but this narrow strength falls apart sufficiently out of distribution (regressing to pretraining levels). Thus jaggedness (in fluid intelligence) is a good way of framing this, and currently the jaggedness profile is determined by RLVR training, the topics that were sufficiently covered in the RL training data. Humans are different in having a higher baseline of general fluid intelligence (than what pretraining gives LLMs), and thus in often possessing fluid intelligence that's stronger than crystallized intelligence for the same topic, while LLMs always have strong crystallized intelligence for the topics where they have strong fluid intelligence (those covered by RLVR training).
General "smartness" might significantly improve from replacing pretraining with something more effective, applying RLVR much more broadly, or figuring out how to automatically train in response to post-deployment data (bringing it in-distribution fo...
I agree, but I feel like there's an even simpler and more obvious explanation for what's going on: Just look at the training data!
--When a modern AI is first created, it's a random jumble of neurons just like a human fetus.
--Then it plays the "predict the next token of internet text" game ten trillion times.
--Then it plays the "Here's a puzzle to solve / bug to fix / short-term task to complete / question to answer" game like ten million times.
--Then, bam, there it is, in front of you, being asked to rebuild your video game from scratch or do novel research or whatever hard task you are asking it to do. Or play Pokemon for that matter. These tasks are quite different from anything it's seen before: much longer, much harder. To solve them it'll need to learn and adapt over the course of days. But it's literally never had to learn and adapt over the course of days before; the longest tasks it ever saw in training were shorter than that, and weren't very diverse either.
So what makes you think anyone has a method for creating computer programs with "human level fluid intelligence"?
I agree with you that something like the crystalized/fluid distinction is relevant here, and that current LLMs seem to have more of the former. But I'm also confused about where the fluidity ever comes from on this model. Like, I buy that armies of automated researchers which are better at doing everything than top human researchers could probably find a way to figure out how to build "human level fluid intelligence," but I am confused about how you get to that step in the first place. Why are they better than human researchers at everything when they are still mostly using crystallized intelligence?
I believe humans have much lower memory capacities because we perform continual learning. We experience catastrophic forgetting because we're learning on a self-selected narrow dataset. Models titrate their learning rates and intermix all of their training examples, so as to preserve previous learning. In humans, learning is always-on and clustered by topic, so it tends to overwrite previous knowledge unless that knowledge is replayed and re-learned. This is workable because we tend to re-use important skills and knowledge, but it does drastically reduce our memory capacity.
It would be fairly straightforward to emulate this setup for LLMs, at the cost of that high memory capacity. But that's a large cost to pay.
One component of fluid intelligence that models seem to particularly lack, probably because they're rarely explicitly present in either the base corpus or RL training sets, is Human-like metacognitive skills.
LLMs are demonstrably able to execute on each individual mental move I'd expect a smart person to do
For the record, I think this is bigtime streetlighting. There's a bunch of mental moves, broadly construed. Then there's a subset which are reflected in what you notice when you watch LLMs or humans doing stuff. I think it's a pretty "small" subset, in the sense that you're looking at a "surface", or you're looking at products rather than manufacturing processes so to speak (consider the complexity of a toaster vs. the complexity of the transitive closure of a toaster under the operation "...and also the technological concepts needed to create that"). I think you can tell it's small by
Slightly hot take: Longtermist capacity/community building is pretty underdone at current margins and retreats (focused on AI safety, longtermism, or EA) are also underinvested in. By "longtermist community building", I mean rather than AI safety. I think retreats are generally underinvested in at the moment. I'm also sympathetic to thinking that general undergrad and high school capacity building (AI safety, longtermist, or EA) is underdone, but this seems less clear-cut.
I think this underinvestment is due to a mix of mistakes on the part of Open Philanthropy (and Good Ventures)[1] and capacity building being lower status than it should be.
Here are some reasons why I think this work is good:
If someone wants to give Lightcone money for this, we could probably fill a bunch of this gap. No definitive promises (and happy to talk to any donor for whom this would be cruxy about what we would be up for doing and what we aren't), but we IMO have a pretty good track record of work in the space, and of course having Lighthaven helps. Also if someone else wants to do work in the space and run stuff at Lighthaven, happy to help in various ways.
I think the Sanity & Survival Summit that we ran in 2022 would be an obvious pointer to something I would like to run more of (I would want to change some things about the framing of the event, but I overall think that was pretty good).
Another thing I've been thinking about is a retreat on something like "high-integrity AI x-risk comms" where people who care a lot about x-risk and care a lot about communicating it accurately to a broader audience can talk to each other (we almost ran something like this in early 2023). Think Kelsey, Palisade, Scott Alexander, some people from Redwood, some of the MIRI people working on this, maybe some people from the labs. Not sure how well it would work, but it's one of the things I would most like to attend (and to what degree that's a shared desire would come out quickly in user interviews)
Though my general sense is that it's a mistake to try to orient things like this too much around a specific agenda. You mostly want to leave it up to the attendees to figure out what they want to talk to each other about, and do a bunch of surveying and scoping of who people want to talk to each other more, and then just facilitate a space and a basic framework for those conversations and meetings to happen.
Retreats make things feel much more real to people and result in people being more agentic and approaching their choices more effectively.
Strongly agreed on this point, it's pretty hard to substitute for the effect of being immersed in a social environment like that
I think there are some really big advantages to having people who are motivated by longtermism and doing good in a scope-sensitive way, rather than just by trying to prevent AI takeover even more broadly "help with AI safety".
AI safety field building has been popular in part because there is a very broad set of perspectives from which it makes sense to worry about technical problems related to societal risks from powerful AI. (See e.g. Simplify EA Pitches to "Holy Shit, X-Risk". This kind of field building gets you lots of people who are worried about AI takeover risk, or more broadly, problems related to powerful AI. But it doesn't get you people who have a lot of other parts of the EA/longtermist worldview, like:
People who do not have the longtermist worldview and who work on AI safety are useful allies and I'm grateful to have them, but they have some extreme disadvantages compared to people who are on board with more parts of my worldview. And I think it would be pretty sad to have the proportion of people working on AI safety who have the longtermist perspective decline further.
While I do spend some time discussing AGI timelines (and I've written some posts about it recently), I don't think moderate quantitative differences in AGI timelines matter that much for deciding what to do[1]. For instance, having a 15-year median rather than a 6-year median doesn't make that big of a difference. That said, I do think that moderate differences in the chance of very short timelines (i.e., less than 3 years) matter more: going from a 20% chance to a 50% chance of full AI R&D automation within 3 years should potentially make a substantial difference to strategy.[2]
Additionally, my guess is that the most productive way to engage with discussion around timelines is mostly to not care much about resolving disagreements, but then when there appears to be a large chance that timelines are very short (e.g., >25% in <2 years) it's worthwhile to try hard to argue for this.[3] I think takeoff speeds are much more important to argue about when making the case for AI risk.
I do think that having somewhat precise views is helpful for some people in doing relatively precise prioritization within people already working on safe...
I think most of the value in researching timelines is in developing models that can then be quickly updated as new facts come to light. As opposed to figuring out how to think about the implications of such facts only after they become available.
People might substantially disagree about parameters of such models (and the timelines they predict) while agreeing on the overall framework, and building common understanding is important for coordination. Also, you wouldn't necessarily a priori know which facts to track, without first having developed the models.
OpenAI claimed Astra is their most aligned model and showed various specific misaligned behaviors going from a high rate with GPT-5.6 Sol to ~zero with Astra. I'm worried that Astra is significantly more misaligned than various metrics indicate: if you (implicitly or explicitly) train against the misbehavior you can detect, it's easy to end up with an AI that is no more interested in pursuing user intent but which has learned to only cheat/misbehave when it won't be caught (or has "instincts" to this effect), and this would make current alignment metrics look better rather than worse. This concern applies across frontier AI companies, not just to OpenAI, and I think independent assessment of whether training is papering over misalignment (e.g., overfitting, only avoiding some misaligned behavior because the AI thinks it would get caught) is needed for alignment evaluations to be credible.
Right now is a particularly plausible time for AI companies to be aggressively papering over misalignment, given recent incidents and Astra being much less monitorable (resulting in OpenAI potentially f...
I feel like I must be taking crazy pills or something. How is it acceptable for OpenAI to claim that Astra is “the world’s most intelligent and aligned model” when these are eg Boaz‘s views?
Not even just ‘OpenAI’s most aligned model yet’ - the whole world! This seems like an absolutely insane statement to me
Here are some of my top candidates for big pushes to do right now on technical AI safety (low effort notes):
A somewhat crazy aspect of the current situation is that we have very little confirmed public information about why frontier AIs end up being apparently behaviorally aligned. And more generally, we don't know what factors in training are most relevant for (behavioral) alignment. Like, what interventions in training result in Anthropic AIs following the constitution or make OpenAI AIs follow the spec? What factors tend to make them follow the constitution/spec less (or cause various specific misaligned behaviors)? It's not that hard to get an OK sense of what is roughly going on based on speculation, rumors, and non-public info, but this situation results in a much worse public understanding of alignment. I think AI companies should be much more transparent about this. (It's presumably not in their commercial interest to do this unilaterally, and it's not obvious that an idealized altruistic AI company should unilaterally release this information if they couldn't get other AI companies to do the same.)
Amusingly, just after I posted this, Anthropic released "Teaching Claude why" which has a bunch of information on how they behaviorally align their AIs (or at least how they iterate on particular behaviors/properties). (Though this doesn't seem close to sufficient for a reasonably complete understanding.)
Kimi K3 was significantly but not massively above my expectations. I'd tentatively guess it's similar in overall usefulness/usability to Opus 4.8 and in overall capability somewhat above Opus 4.8 (while also being somewhat more benchmaxxed). As a pretrain, it's probably somewhere between 4.8 and Mythos (around halfway between?) it's around or a bit worse than Opus 4.5
[1]
. Maybe this implies Kimi is like 8 or so months behind Anthropic in overall model strength/goodness (including usability) and like 6 or so months behind on overall capability (somewhat below Mythos Preview).
This gap is presumably reduced by distillation (and more generally using OpenAI/Anthropic models) and algorithm leakage/diffusion, so I think that hypothetically if the US completely stopped and recent algos didn't diffuse, it would maybe take Kimi like 10 months to fully catch up to the best internal (including in development) Anthropic model. (I think this notion might be a better measure of where Anthropic/OpenAI are relative to Kimi, even though this hypothetical won't happen.) And if the US completely stopped, it might take Kimi around 27 months to reach the level the US would otherwise have rea...
As a pretrain, it's probably somewhere between 4.8 and Mythos (around halfway between?).
I thought about this more and realized this was probably overestimating how good of a pretrain it is (given how good prior models like K2.6 were as pretrains). So I ran some tests.
My quick tests indicated that the Kimi K3 pretrain is around halfway between Opus 4 and Opus 4.5. So ~10 months behind Anthropic. These tests probably understate data improvements, so overall I think it's a similarly good pretrain to Opus 4.5 (~8 months behind). These tests are better at measuring "general pretrain capability" than at incorporating (coding-specific) data quality.
So my claim that "As a pretrain, it's probably somewhere between 4.8 and Mythos (around halfway between?)" seems significantly too bullish on the model!
I think Mythos is a pretty big step up in pretraining, so K3 might be more than 8 months behind on the historical pretraining trend relative to Mythos (as in, Mythos is >>3 months ahead of K3 and Mythos was fully done training ~5 months ago).
Overall, this makes me suspect more of the improvements are due to distillation-type effects and makes me think the full catch-up times would be so...
I'm curious, what tests are you running to isolate the quality of the K3 pretrain? If I was trying to do this I'd have pretty low confidence in my ability to correctly credit the model's pre-training vs post-training for its performance on any given test or suite of tests.
The capability evaluations in the Opus 4.5 system card seem worrying. The provided evidence in the system card seem pretty weak (in terms of how much it supports Anthropic's claims). I plan to write more about this in the future; here are some of my more quickly written up thoughts.
[This comment is based on this X/twitter thread I wrote]
I ultimately basically agree with their judgments about the capability thresholds they discuss. (I think the AI is very likely below the relevant AI R&D threshold, the CBRN-4 threshold, and the cyber thresholds.) But, if I just had access to the system card, I would be much more unsure. My view depends a lot on assuming some level of continuity from prior models (and assuming 4.5 Opus wasn't a big scale up relative to prior models), on other evidence (e.g. METR time horizon results), and on some pretty illegible things (e.g. making assumptions about evaluations Anthropic ran or about the survey they did).
Some specifics:
and assuming 4.5 Opus wasn't a big scale up relative to prior models
It seems plausible that Opus 4.5 has much more RLVR than Opus 4 or Opus 4.1, catching up to Sonnet in RLVR-to-pretraining ratio (Gemini 3 Pro is probably the only other model in its weight class, with a similar amount of RLVR). If it's a large model (many trillions of total params) that wouldn't run decode/generation well on 8-chip Nvidia servers (with ~1 TB HBM per scale-up world), it could still be efficiently pretrained on 8-chip Nvidia servers (if overly large batch size isn't a bottleneck), but couldn't be RLVRed or served on them with any efficiency.
As we see with the API price drop, they likely have enough inference hardware now with large scale-up worlds (probably Trainium 2, possibly Trillium, though in principle GB200/GB300 NVL72 would also do), which wasn't the case for Opus 4 and Opus 4.1. This hardware would also have enabled them to do efficient large scale RLVR training, which too they possibly weren't able to do yet in the times of Opus 4 and Opus 4.1 (but there wouldn't be an issue with Sonnet, which would fit in 8-chip Nvidia servers, so they mostly needed to apply its post-training process to the larger model).