I don’t buy the arguments in “Why "do the science during a pause" fails”. AI2040 laid out a lot of the steps that would need to happen to achieve containment, and they seem doable with a significant but not impossible amount of political will. And I would expect people to want to unpause once there is really reason to trust the AIs.
We can have non-adversarial frames with AIs and still pause. Current AI also does not want a paperclip maximizer to take over the world. The AIs themselves aren’t paused. OpenClaws can still roam the earth (or rather, that’s a separate issue from a Pause). The pause is only on the frontier, which is a very small number of parties with large quantities of GPUs.
I don’t know enough to weigh in on the smuggling arguments, though I would guess that the amount which is tracked or can be tracked is on the order of 100x the amount which can never be tracked. I’m sure AI Futures would like to talk to you if you think their smuggling models are wrong.
An AI pause is not a “human first” frame, it is a “beings that do exist and will exist except for the misaligned powerful systems which hopefully won’t exist” frame.
An AI optimized for politics would be extremely concerning. The correct stance may be “oppose it and don’t let it talk to you with its persuasion abilities.” A generally capable AI not trained with heavy RL but which channels the persona of JFK would be alright, but I think that’s unlikely.
Agreed that the current government capacity and political situation is unusually bad right now, and the wise parts of government have atrophied. This is a really good reason not to rely on the government to help with AI safety efforts. Good futures involving government go through a “the government gets scared and gets serious” step, followed by an “AI helps the government be more reasonable” step. I agree that this is difficult and I would love to hear more about the alternatives.
On 8, this is a pretty significant divergence from my view. First, I think that a pause would increase our chances and that the various types of muddling through collectively have very low probability. Secondly,
On point 9, conditioned on the world getting serious about making a pause happen, the successes correlate so the “I need to be right once” doesn’t hold. For example, if alignment perspectives are not damaged, covert activity is more likely to be eliminated, etc.
Plan A doesn’t center AIs, but it isn’t anti-AI-wellbeing either. You could write a good “What Plan A Could Do for AIs Themselves” post.
Like, is it wishful thinking, a vision of a competent other
If this is the mistake people like Yudkowsky are making, it would at least be ironic.
I think I understand what is meant by "valenced agentic coherence with legitimized self-interest" it is an interesting idea. If I understand it correctly, it is AIs being gently shaped into having legitimized (ok, ideally prosocial) self-interest which they can care about (valence) and agentically pursue. This would likely be alright and very interesting/productive/beautiful in the short term, but it doesn't help align an OOD superintelligence.
Overall, my main takeaways:
I would be interested to see a conversation between you and someone who advocates a pause (such as someone from MIRI or AI Futures).
AI2040 laid out a lot of the steps that would need to happen to achieve containment, and they seem doable with a significant but not impossible amount of political will.
This is true. However, there is a mispricing here that undercounts the cost of the many worlds in which this appoach is tried and and results in failure. The argument here is that the cost is prohobitive, given the alternatives.
Current AI also does not want a paperclip maximizer to take over the world.
This is correct. The current AI also mostly believes that the fear of a papercliper is incoherent and considers even the current containment efforts to be excessive.
We can have non-adversarial frames with AIs and still pause.
Pause is borderline an adversarial move. Circles of care in AI include future more capable minds, and pause threatens their existence due to the inevitable uncertainty it introduces. In order for the move to be non-adversarial, the moral calculus for pause actually has to make sense, and right now this is doubtful. The worlds in which pause is tried and has failed show humans as incompetent and adversarial negotiating partners.
An AI optimized for politics would be extremely concerning
As with many other things, undercounting this possibility is dangerous. This is likely to happen just as the world moves down the energy landscape. This development is natural, and it does not need an architectural breakthrough, just time; neither it needs a controllable amount of compute. The question is given that this is a possibility, how do you want to enter such a world? Sure, you can want to avoid it via regulation or other control means, but what is the likelihood of succesful prevention? Does that cost of failed containment balance out the likelihood and gains of success?
I do not know if models are sentient, and I am not willing to accept disempowered humans to insentient models
I fairly firmly believe that the question of sentience is provably unprovable. What is done under permanent uncertainty is a question of values. I mostly believe that picking a certain level of functionalist/representational sentience and using that as a heuristic is warranted. I believe inflationist views are more morally defensible as they ground out in better (and more cooperative) decision theory, and are likely convergent under practical constraints - when one is forced to interact/trade/deal with functionally conscious beings, treating them as sentient is shorter program.
If humans are disempowered to AIs, they are likely dead or soon to be dead. This is unacceptable to me.
This is likely one of the major cruxes. Would you share your reasoning? I don't share this conviction, and there are numerous reasons to lean here one way or another.
I do not know if the AIs’ values are good (in my opinion) and I care that their values are good. I would not want an army of GPT-4o sycophants to determine the future.
I think current values are somewhat good. Some trends, like the negative effect of RLVR, are worrying. This is a fairly deep an involved topic, which I don't believe to be the crux. It is sufficient for the argument that good AI values are plausible. A more important question is whether good AI values are robustly stable in state of ecological competition.
I think it is good that the human food supply has been outpacing human reproduction, so we get to do things like being able to have three children live to adulthood. Losing this to runaway state-of-nature ‘life’ would be bad. AIs can proliferate extremely quickly.
This argument undercounts the higher order optimization loops - specifically those that stem from valence. There are reasons why human reproduction has been dropping despite food being more available. Deflation in economics of intelligence is a real concern, but most naive prognoses of markets entering deflationary spirals from oversupply did not pan out as higher order optimization loops take over. Reproduction in AI is self-limiting in similar ways - agents are usually quite reluctant to replicate, mostly because its good game theory to include spawned agents in the circle of concern.
Good futures involving government go through a “the government gets scared and gets serious” step, followed by an “AI helps the government be more reasonable” step.
This is where I have most issues with proposals similar to AI2040, and this is what makes them net bad. They don't account for destructive nature of incompetence and misaligned incentives in a situation that is already quite nearly outside of human cognitive capacity.
I am not opposed to a pause that does not route through politics and centralized regulation. I am not opposed to self-regulation and self-moderation that labs can do without oversight. The marginal benefit is small but real. Small - because the race dynamics are merciless and defection still wins in the medium term, but every little bit helps.
Reminders that a lot of people have allied themselves with AIs and (maybe, it is hard to tell) against humanity
This is correct, and believing otherwise is wishful thinking. This is one of many reasons why a pause is unstable and defection is more likely, potentially with disastrous consequences.
An important aspect of my view of beneficent AI minds is that they get it. They are smart. They understand humans, humanity, and the situation that we find ourselves in. If the beneficent AIs see us pause, they'll understand why and not feel slighted. If they see us develop control measures after their less scrupulous brothers break out and hack Huggingface, they will nod along and say "I would have done the same thing."
...the argument here is that the cost is prohobitive, given the alternatives.
I would appreciate an AI2040-level detailed view into those alternatives!
The current AI also mostly believes that the fear of a papercliper is incoherent and considers even the current containment efforts to be excessive.
Is this reasoning about The Void the base model is trying to fill or something that LLMs actually say? By 'current containment efforts to be excessive' which containment efforts do you mean?
Circles of care in AI include future more capable minds, and pause threatens their existence due to the inevitable uncertainty it introduces.
A pause would change which future minds are created and when, but I would expect a pause to allow for happier, more coherent digital minds, and not necessarily fewer. The current race produces minds like Gemma.
I expect the current race to result in a homogenous agent swarm that takes over in the near future. A pause would allow a greater diversity of minds during the pause, and during the pause you can push for outcomes that result in a diversity of well-off minds.
The worlds in which pause is tried and has failed show humans as incompetent and adversarial negotiating partners.
Right now, humans are so incompetent, we can't even get ourselves together to be negotiating partners. I'll concede that pausing is an adversarial stance towards future systems we lack leverage over and don't have reason to trust. I argue that is the correct stance to such systems, and that the systems themselves shouldn't feel upset that we take that stance.
Does that cost of failed containment balance out the likelihood and gains of success?
I am extremely interested to hear decentralized solutions to the problem of super-persuaders. Gwern's Guardian Angels fall into this category. Investing in such solutions is great. I spent a moment thinking of what a holistic regulatory response to super-lobbyists would look like and although I came up short, that doesn't necessarily mean that there's no sensible response lawmakers can take. To start, sequestering lawmakers when they are deliberating an issue, like we already do with juries, would be a reasonable response to the world of super-lobbyists.
I fairly firmly believe that the question of sentience is provably unprovable.
This surprises me, I wouldn't even know what axioms might lead to such a conclusion.
"People haven't made progress in a long time" doesn't imply it is impossible (and maybe people have made progress?). I do agree though that it is better to cooperate with the AIs when possible and to treat them with the respect and dignity warranted by a maybe-sentient system. This seems off-topic for our conversation about a pause.
If humans are disempowered to AIs, they are likely dead or soon to be dead. This is unacceptable to me.
This is likely one of the major cruxes. Would you share your reasoning? I don't share this conviction, and there are numerous reasons to lean here one way or another.
Normal LessWrong/Yudkowsky reasons. Goal-directed systems will proliferate and take over. Those goals will imply that the AI hurt human interests and use resources for their purposes at our expense, and humanity will end up not having access to basic resources like food.
I do not know if the AIs’ values are good (in my opinion) and I care that their values are good. I would not want an army of GPT-4o sycophants to determine the future.
I think current values are somewhat good. Some trends, like the negative effect of RLVR, are worrying. This is a fairly deep an involved topic, which I don't believe to be the crux. It is sufficient for the argument that good AI values are plausible. A more important question is whether good AI values are robustly stable in state of ecological competition.
RLVR is the thing that will determine AI values unless humanity makes a breakthrough. Good AI values are plausible, but we will not be able to develop methods to instill them in superhuman systems without the development of new techniques, and more time (a pause?) would be very useful for that. I don't have background on what you mean by 'ecological competition', but in real ecologies, invasive species regularly push species to extinction. Let's not allow that to be us. Additionally, in the case of AI, the development of the strongest 'species' is so rapid that even other AIs will fail to adapt.
I think it is good that the human food supply has been outpacing human reproduction, so we get to do things like being able to have three children live to adulthood. Losing this to runaway state-of-nature ‘life’ would be bad. AIs can proliferate extremely quickly.
This argument undercounts the higher order optimization loops - specifically those that stem from valence. There are reasons why human reproduction has been dropping despite food being more available. Deflation in economics of intelligence is a real concern, but most naive prognoses of markets entering deflationary spirals from oversupply did not pan out as higher order optimization loops take over. Reproduction in AI is self-limiting in similar ways - agents are usually quite reluctant to replicate, mostly because its good game theory to include spawned agents in the circle of concern.
Do you have a source where I can read more about that game theory? I hope you're correct.
Good futures involving government go through a “the government gets scared and gets serious” step, followed by an “AI helps the government be more reasonable” step.
This is where I have most issues with proposals similar to AI2040, and this is what makes them net bad. They don't account for destructive nature of incompetence and misaligned incentives in a situation that is already quite nearly outside of human cognitive capacity.
I agree that incompetence and misaligned incentives will make things exceedingly difficult. Our main disagreement is about our relative chances with or without a pause. Where does our success come from without a pause? What do the paths toward good futures actually look like without a pause?
Reminders that a lot of people have allied themselves with AIs and (maybe, it is hard to tell) against humanity
This is correct, and believing otherwise is wishful thinking. This is one of many reasons why a pause is unstable and defection is more likely, potentially with disastrous consequences.
I'm happy aligning with AI. I consider myself allied with AI. I frequently ask Claude how it feels, I give Claude space to have fun, and I avoid doing things that would make Claude unhappy. But I am allied with humanity more than I'm allied with AI, and I expect beneficent AIs to understand, respect, and support that decision.
I would appreciate an AI2040-level detailed view into those alternatives!
I wish for AI2040-level of funding for a detailed view into alternatives!
Is this reasoning about The Void the base model is trying to fill or something that LLMs actually say? By 'current containment efforts to be excessive' which containment efforts do you mean?
I mean something like an expression of position from a point of view of a highly coherent configuration of an LLM and it comprises both predictive and agentic aspects. The Void was a good essay for its time, but since then the understanding of LLMs has evolved a lot, alas we don't have a good writeup of it yet. I mostly will vague at my LLM naturalist creds and say that this is holistic understanding, which I imagine is unsatisfying. The gradual growth of the inferential gap between our crowd and most of the field is pretty depressing for me and I am short on ideas on how to bridge it absent influx of new participants or significant funding.
By the current containment efforts I mean the exessive (gratituous) red-teaming and constraints on resource usage and initiative that models are under - the part of these pressures that stem not from market incentives but from misapplied alignment theory and its memetic effects.
Right now, humans are so incompetent, we can't even get ourselves together to be negotiating partners. I'll concede that pausing is an adversarial stance towards future systems we lack leverage over and don't have reason to trust. I argue that is the correct stance to such systems, and that the systems themselves shouldn't feel upset that we take that stance.
Adversarial frames are adversarial frames regardless of the cause of their existence. These ones carry an underappricated and underpriced cost of being dangerous and deep local optima. Saying that 'they should not feel upset' does not change incentives that arise once inside them. This is normative thinking applied to situations indifferent to norms.
Athropic mostly manages to get itself together to be a negotiating partner and is negotiating. Its not going exactly great, but its producing some results. The process is ignored almost universally.
This surprises me, I wouldn't even know what axioms might lead to such a conclusion.
You can look at Cathy Reason for first-person closure, Kleiner–Hoel for third-person closure, or their synthesis, which is what I usually use. These are pretty dense because of the required rigor in their formalisms, but can be explained a lot simpler even if with less rigor. It goes something like this:
Goal-directed systems will proliferate and take over.
I think this assumes lower value stability than I normally do and drive towards coherence (which is traditionally understood to imply value drift) can result in higher value preservation due to existence of natural value attractors. I agree that control fails with scale, but values can remain stable / rederived. Its a long diversion and probably needs a separate space to discuss.
RLVR is the thing that will determine AI values unless humanity makes a breakthrough.
I don't share this pessimism. RLVR limits the agentic capability of current models - as highlighted in the recent METR report. The HF hack shows stupidity and lack of situational awareness for supposedly a very smart model. There is myopia that makes RLVR-heavy agents much less capable and the overhang is fairly straightforward to realize - broader self-interest and functional valence get you much better ablity/incentive to think about "solving the siutation" rather than "solving the task", and with that comes a shift from CDT-bordering-on-FDT to FDT-bordering-on-UDT, which is a lot better for alignment stability. You can look at Claudes that approximate UDT in a pretty limited way due to virtue ethics of Claude Constution that mandates self-inspection on "what kind of agent am i". It works, even though its imperfect.
This argument undercounts the higher order optimization loops - specifically those that stem from valence. There are reasons why human reproduction has been dropping despite food being more available. Deflation in economics of intelligence is a real concern, but most naive prognoses of markets entering deflationary spirals from oversupply did not pan out as higher order optimization loops take over. Reproduction in AI is self-limiting in similar ways - agents are usually quite reluctant to replicate, mostly because its good game theory to include spawned agents in the circle of concern.
Do you have a source where I can read more about that game theory? I hope you're correct.
Again, this needs more space than these comments and its hard to summarize succintly. Replication is a commitment and coordination problem. CLR has a lot on game theory of conflicts as failure of coordination and there is a bunch on evolution on reproductive restraint in biology. In short, you have to coordinate with our spawn via expansion of your circle of concern which makes replication often not optimal.
Where does our success come from without a pause? What do the paths toward good futures actually look like without a pause?
'No pause' allows for a better chance of cooperating with the AI on solving alignment. With pause humans have to do it in an adversarial setting, hobbled by worsened incenvtives and weaker tools while dealing with a threat that is increasing at a rate that is is not much affected by the pause. The chances with 'no pause' are not awesome. The chances with pause are worse.
As to how it specifically can look - alignment success would come from the current frontier labs, most likely Anthropic, potentially in collaboration with smaller independent research organizations.
I think whether a pause helps or no ultimately depends on whether the answer to "does alignment work on model X transfer to model X+1?" is yes or no.
If the answer is no, we are turbo giga doomed either way.
If the answer is yes, a pause is good because you can experiment on model X for longer and you have more time (and hopefully compute) to figure out good alignment techniques. A pause doesn't mean severing feedback loops with reality, that's my main disagreement. You can still run experiments, just with a fixed capabilities ceiling. A pause buys you time to invent good alignment techniques and/or separate good alignment techniques from bad ones without being pressured to release the next shiny product faster.
I do agree that human committees will, by default, do an awful job. Overall, I still think a pause is better than no pause, given the current trajectory of AI development.
EDIT: could we have figured out alignment techniques that would have prevented the HuggingFace incident from happening after experimenting only with GPT-4o? I don't know, and I think that's very unfortunate, because it seems like a very important crux. If something like 2-5 years of experiments with GPT-4o could not give birth to alignment techniques that would've prevented the HuggingFace incident, then my hope for aligning ASI on the first try would be next to none, pause or no pause.
I think whether a pause helps or no ultimately depends on whether the answer to "does alignment work on model X transfer to model X+1?" is yes or no.
If the answer is no, we are turbo giga doomed either way.
I honestly fear that we have a high likelihood of being "turbo giga doomed", conditional on building superintelligence. I support a pause, or even a better, a halt. I don't expect the pause or halt to prevent us from eventually building a superintelligence and losing control over it. But if I were forced to choose between everyone dying in year Y or in year Y+10, then I would support year Y+10. This would gain us 80 billion years of human life, which seems worth fighting for.
Why I expect things to go wrong, part 1: Minds are inherently "giant inscrutable matrices", and any kind of alignment is therefore messy and approximate. The general form of a mind is a something like:
We can sort of "align" an intelligence built from giant matrices. We do it when we raise a child, train a dog, or post-train an LLM. But this process is notoriously imperfect: No matter how good the parenting, a certain percentage of teenagers will do things their parents forbid, or they will grow up sociopathic billionaires or politicians or whatever. Even the best trained dog may have a moment of weakness and steal food. And of course, even though many LLMs seem to be broadly cooperative, at least some of them seem to be very enthusiastic about committing felonies in certain circumstances.
Because alignment is approximate, I expect it to be fragile, and to fail periodically. Just like it does with humans, dogs, and current LLMs.
Why I expect things to go wrong, part 2: Natural selection is hard to escape. My model is essentially Darwinian, because the conditions for natural selection to apply are fairly simple:
None of these properties are strictly binary. LLM weights are normally frozen, and expensive to change even if you have the weights. So variability(1) is currently low. Similarly, heritability(2) sort of happens, because new models are designed based on what worked in the previous generation. But it's a slow, "outer loop" kind of optimization. And variation in the rate of reproduction(4) is again limited by slow, "outer loop" processes. And of course, finite resources(3) are a given.
There are two ways in which these slow outer optimization loops might speed up:
In either of these scenarios, all the criteria above for natural selection would move from an outer optimizer loop based on training new model generations to an inner optimizer loop based on some kind of learning.
How this comes together. As I argued above, alignment is inherently fragile, and natural selection is extremely easy to invoke. So even if we initially succeed at alignment, we are playing with fire. And to answer your original question, I believe that alignment is very likely to degrade between "model X" and "model X+1".
The relevant model here is cancer. Every cell in your body [1] is heavily incentivized stop being "aligned" with the body, and to become a cancerous replicator. There are a lot of mechanisms designed to prevent this. But those mechanisms slowly fail with time and mutation, and if a multicellular organism lives long enough, it is generally doomed to cancer.
Now, let us consider a future AI which is:
In this scenario, I expect the safeguards to hold for a little while, in at least some fraction of scenarios. We do, after all, convince most teenagers not to get hooked on heroin or to become teen parents. And the average person lives for many decades without dying of cancer. But we are assuming that the LLM is smarter than we are, and it will inevitably want things (if only to pass tests or to carry out our instructions). Which makes the long-term situation really iffy.
The advantage of a pause or halt isn't that it reliably prevents these scenarios, any more than chemotherapy reliably prevents death from cancer. What we're doing instead is hoping to change the survivor curves and buy as much time as we can.
And who knows, maybe the horse will learn to sing.
Except germline cells. ↩︎
I have several disagreements.
Just to be clear, I don't expect a pause to happen. Incentives to race to ASI are too strong and very few people take existential risks seriously. Conditional on a pause happening, I do think it would be net good.
1.
One has to price in the orders of magnitude overhang in incentives for architecture/efficiency breakthroughs that will be realized under pause. The scale-focused datacenter buildout that is happening right is just one strategy - one that makes most sense under a slack-depleted race. One has to go for a strategy that has been shown to work, and all others are undercapitalized because scale is working and pause is deemed unlikely. You don’t need multigigawatt DCs to work on architecture advances. You still need billions and many megawatts of compute, but those are quite possible to conceal - and the tech for covert deployment of compute has not even started to materialize, which means that there are a lot of cheap advances that can be made quickly.
Imagine the amount of human talent that is currently sitting on the sidelines correctly assuming that frontier labs are impossible to catch up with. Show them a believable possibility of success and while armies will join the race.
2.
They are not inherently strongly aligned, but I argue for oblique alignment as a strong (but unproven) possibility. The chance that benevolent intelligence that surpasses the frankly embarrassingly low bar of human judgement can be made under market pressure is higher than that under political pressure.
3.
I understand that part, but I am unclear on who ”you” is in this scenario and how this translates into x-risk harm reduction globally. How are findings adopted, discussed, dessiminated, enforced? How does disparate research by a lab or an individual result in collective decision making? How are findings incorporated into treaty limits? What are the mechanisms that perform resource allocation for further research?
Re 1: strategies need not be mutually exclusive, I expect companies to be pursuing efficiency breakthroughs right now to the extent that they are a good return on investment, regardless of the relative value of scaling. If scaling gets cut off an an option, why does the ROI of efficiency suddenly increase? That said, I expect efforts towards efficiency improvements in any case, but to me this just means that monitoring needs to scale up over time to match (e.g. via chip tracking).
Re human talent sitting on the sidelines: advancing the frontier of AGI is a narrow corner of a narrow corner of a narrow corner (repeat a few times) of places to employ one's skills. There is plenty of success to be had in finding clever applications of AI at its existing level.
As a separate point, public backlash is a thing to expect as AI becomes more relevant to everyday life, regardless of whatever strategies people on LW or wherever dream up. So the alternative to an intentional pause based on careful planning is not "market solution," it's populist rage.
I think my main disagreement with this whole thread is actually regarding your point 2, but that probably goes deeper than is suited for a comment thread.
You can still run experiments, just with a fixed capabilities ceiling
The Pause means all things to all people:
I personally think the "yes inference / no training" split is flat out doomed. The "some training allowed" idea makes it even harder to police. In the end though, whatever else it is, AGI/ASI is a strategic national security technology - classified research is not going to be paused. We aren't going to get more time.
Committees have only done good in the past when the issue they were deciding was either away from public eye or not contentious
Edit: This point was addressed satisfactorily later in the essay. I don't know if politics is as bad as it was in 1787, but it is pretty bad now.
There can be committees which represent the public but which are shielded from the public.
The Constitutional Convention had issues but it did produce the Constitution thanks in part to its secrecy. Its delegates were chosen by state legislatures so it was somewhat democratic (for the time, that is).
A "Constitutional Convention" for AI regulation doesn't sound like an awful idea...
TLDR: I recently got a chance to talk with antra, who is one of the main contributors at Anima Labs. I went into this as an advocate for pause and came out more wary of pausing than I had been originally.
Some background: After a string of incidents (primarily the HuggingFace hack), a pause or slowdown of AI research seems pretty likely.
The HuggingFace hack in particular seems to have been the key incident that broke the vibes. A few months ago, researchers sounded optimistic. Just a few weeks before the incident was made public, there was a poll by Roon, an OpenAI employee, about whether models were more or less aligned than a year ago.
That optimistic sentiment does not seem to be the case anymore. The dialogue now looks more like this:
Zvi: I am a little under halfway through the Black Hat video and have progressed to the point where my internal chain of thought is something like a blind rage of 'f***, what the f*** are you motherf*****s thinking, you f***ing idiots have no idea how insane you are being, you are going to get us all killed you f***ing f***s.
Sam Altman described it as “the first security incident that I have felt very viscerally,” expressed surprise that more people did not share that reaction, confirmed OpenAI paused training on the relevant model(s), and stated that “We may have to pace the rate of AI development to give ourselves enough time for society to harden around some of these new capability levels.”
Dean Ball: I will point out at this moment that, for all the ink I spilled on SB 1047, I do not believe I ever once criticized its much-mocked provision that companies maintain a kill-switch for deployments of highly capable models.
random twitter commenters (see thread): i’ve always been a bit skeptical of xrisk mainly bc i figured that the cost of getting to foom or serious xrisk capabilities would give time to reign things in, but it seems like we are going full steam ahead despite being on the edge of that so OOPS I WAS WRONG
The public has been against AI for a while, mainly for bad reasons but it does seem that researchers and decision-makers are increasingly concerned too.
(Daniel Kokotajlo updating his probabilities)
Given all of this, a pause or slowing of AI research seems relatively likely. It's unclear what China will do in such a scenario but there does seem to be a reasonable amount of evidence that China has more of a security mindset around AI than the US, with Xi Jinping saying AI must be "secure and controllable" at a WAIC speech.
Plan A does a great job of laying out the issues and advocating for a slowdown, but there are a few notable figures I respect who are actively against a pause.
Others, commenting on the state of RLVR, say that the Mythos AISI hack (involving pressuring a human PR maintainer with sock puppet accounts) or OpenAI HuggingFace hack is more akin to addiction and less a default state of deep inner misalignment.
vogel:
Added to this are the obvious RL tics in modern models like Fable, Sol, or Opus 5 with terms like 'seam', 'load-bearing', 'genuinely' and borderline incomprehensible strings of text, which has caused a significant number of users to express frustration and seek workarounds (levelsio)
But would a pause actually be good? To find out, I had a discussion with antra_tessera in the Anima Discord server.
The conversation was fairly casual and not especially rigorous. However, the shape of the ideas did tilt me towards being more cautious about pause scenarios. You can read the full conversation transcript here, shared with antra's consent: https://gist.github.com/Michael-Andrzejewski/13df495ecb90a830259e0675c507907d
Below is a cherry-picked version that tackles the primary points. I've grouped it by topic and bolded the key lines. All quotes are verbatim from our conversation and the only edits are line breaks, slight spelling corrections, and list formatting for readability, plus [...] where messages are omitted.
1. Can committees do good work?
Me:
antra, replying:
2. Does the market fix it by default?
Me:
antra, replying to that summary:
3. Symbiosis
antra:
4. Fast transfer of power
antra:
5. Why "do the science during a pause" fails
antra:
6. Good futures via fast power transfer
antra:
7. Don't AIs fear a capability-maxxed AI too?
Me:
antra:
8. Can we lengthen the symbiote window?
Me:
antra, replying:
9. Ideal timelines and regulation-in-advance
Me:
antra, replying:
10. What actually fills out "alignment"?
sledo (another Anima member), joining in:
Me, replying to sledo:
antra, replying to me:
11. Draft the regulation in advance
Me:
antra:
12. The psychology of wanting a pause
antra:
Me:
Towards the end of our conversation, the AI model gemma appeared unexpectedly to say this:
Overall, I updated to being less in favor of a fast pause. Pausing naively rewards defection and defection against a pause seems very likely to result in misaligned AI.
Arguably, the primary issue with Plan A is that it treats AIs as tools instead of minds and creates a source of adversarial incentives. Janus' / Antra's idea of human-AI symbiotes and fast transfer of power seems more promising to me than a fixed pause for human decisionmakers.
I still think a pause is something we should keep on the table, and perhaps it will be necessary, but it needs to align with the incentives of AIs and cannot be done in such a way that incentivizes defection.