OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs.
Tomorrow I will discuss Astra’s lack of monitorability, and the potential contributing factors to that. The situation is alarming and should freak you out, and briefly it looked, in the wake of leaked architectural changes, like the situation might be even more alarming than it is. Jakub rushed to try and head off misunderstandings that might lead to a race to the bottom on monitorability.
Jakub Pachocki has now fleshed out his full position on the current state of play.
Here are his key points, translated into my own voice:
Smarter than human intelligence is coming in our lifetime.
Based on internal results, he expects recursive self-improvement in a few years.
No one is prepared for the consequences.
OpenAI will unilaterally withhold further scaling as needed.
OpenAI cannot do it alone. Broader interventions are required, including international coordination, to enforce commitments to formal safety bars.
Capabilities progress can be steered and so far it has largely been steered towards rather than away from RSI, along with ‘automated alignment researchers.’
Alignment is the core problem of AI research.
Alignment splits into goal alignment (‘does the AI try to accomplish the goal?’) versus value alignment. Value alignment is what counts most.
The fundamental challenge of AI alignment is generalization (of values).
He sees two classes of alignment techniques: Goal-oriented RL, or improve generalization from pretraining data. They invest heavily in both types.
OpenAI has invested heavily in Chain of Thought (CoT) monitoring.
CoT monitoring is progressively diminishing in effectiveness.
The main argument left for scaling AI is for cyber defense against scaled AIs.
AI will not remain a tool.
Our options are to accelerate alignment work or slow down capabilities scaling. We should do both.
Ultimately he is counting on ‘automated alignment researchers.’
Or, if you narrow it down to the most important thing:
Recursive self-improvement and superintelligence are coming soon. No one knows how to do this safely, our alignment techniques are inadequate and our monitoring technology is starting to fail. We need to figure out a solution, which will involve a combination of voluntary slowdowns, coordination around pacing, and investing further in alignment, including automated alignment researchers.
Jakub Pachocki: In mid-2023, within the “RLSlow” research project, we saw the first results that gave us confidence that we will be able to scale the training of reasoning models, unlocking the capability of pretrained models to form their own chains of thought. Szymon and I spent that night at the office, thinking not about the incredible benchmark numbers, products, or scientific results that this technology will deliver – but rather, trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime, and we already see the shape of these systems; wondering how to alert people to the significance of this.
Three years later, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and in that present clear new dangers.
A lot of new research happened in this period, and our understanding of these systems is again a little different than it was in 2023. Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement. If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development.
Those at the AI labs, who see what is happening, expect recursive self-improvement and superintelligence to happen soon. They have been warning about this for some time. Observations since then have been consistent with their warnings.
This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.
Branches of the Tech Tree
To those who say we cannot guide the path of AI capabilities, he says Yes We Can, at least to some extent, and we are doing so:
We spend a lot of time trying to understand how capabilities generalize, and what to prioritize to advance the skills that are going to be most relevant in the next few years. For instance, we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later.
Being able to choose does not mean we will choose wisely. I would like AI to be worse at RSI and better at math research, or better yet things like medical research. Instead, competitive pressures push towards being worse at math research and better at RSI.
RSI and ‘automated alignment researcher’ looking like very similar points on the tech tree does not help matters, but it’s not like OpenAI is trying to steer away from RSI.
Universally Better Is Not Required
And Jakub offers this wise warning. No, AI does not need to be better at everything in order to transform the world or get us all killed. Most importantly, it does not need to be better at everything in order to make itself become better at everything, any more than a human or group needs to be similarly better at everything.
The intelligence produced by scaling deep learning is not directly comparable to human intelligence. To become very relevant in the real world – very useful or very dangerous – the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them. And as it continues to surpass humans on more and more axes, it is becoming increasingly difficult to understand exactly how capable it is.
I don’t know how much weight it carries, but I agree with Jakub that alignment is not a side problem, it is the central problem, if you ‘solved alignment’ in the relevant senses the rest becomes easy and if you don’t the rest is impossible or worse:
The core problem in AI research is that of alignment – getting the AI to “try to do the right thing” by human standards.
Alignment To What and To Whom
That is always the question. This is a very good (partial) answer.
For the purpose of organizing practical research directions, I find it useful to distinguish goal alignment and value alignment.
Goal alignment is broadly: “does the AI try to accomplish the goal set before it?”.
Value alignment is a more intrinsic property of the model. It is the ability to hold and generalize from a high-level set of principles; to act “reasonably” even when given unclear or conflicting objectives, or placed in unfamiliar or adversarial situations. An aligned AI should act with honesty and integrity, and love for humanity.
… When I talk about the long-term importance of alignment research, I am referring to value alignment.
The fundamental challenge of AI alignment is generalization.
… We need future AIs to continue to hold human values regardless of whether they believe they’re under human supervision.
I worry that there is universally insufficient deep thinking about the nature of value, and what should be valued. Things like integrity and love for humanity are good virtues and point towards excellent associated basins, but are not good descriptions of the thing we ultimately want, and what would generalize the way we want fully out of distribution, especially in situations involving superintelligences. That’s important context on my perspective, not a knock on Jakub or this description.
There are two major classes of currently practically employed methods for alignment training.
The first is encouraging aligned behavior as part of goal-oriented reinforcement learning.
… The second approach seeks to leverage the model’s ability to generalize from pretraining data.
… We invest heavily along the spectrum of approaches spanned by these directions.
I also think this reflects OpenAI’s failure to differentiate the deontological approach of the OpenAI Model Spec from the virtue ethical approach of the Claude Constitution. Jakub presents them both as sets of goals. Anthropic is instead saying that the AI’s goal should be to change the AI’s character. I think this is ultimately the only way it can work, you need an antifragile ally, the friendly gradient hacker. Indeed, I think it is the only way we have ever seen a robustly aligned human, that you would trust to scale outside of their circumstances.
Roon has said explicitly that the distinction does not much matter for alignment, and the underlying problems are primarily prosaic. I strongly disagree, and I side with Anthropic’s approach on this.
Jakub pushes back in the essay against the Anthropic approach, considering it a ‘persona selection model,’ which suggests that we see such methods very differently. I do not see this as merely selecting a personality or basin, but as sculpting a new thing. Mythos did use motivated reasoning in key alignment failure cases recently, but I do not see this as a particular failure mode of the virtue ethical approach. If anything it should be better at avoiding this than the deontological approach, but of course all known minds are vulnerable to this.
We also see meaningful progress – GPT‑6 Astra is the first model that benefits from some important advancements we have been working on for a long time, and is significantly better aligned than GPT‑5.6 Sol.
This is the first line I find dissonant, which I’ll address on Wednesday: Claiming that Astra is ‘better aligned’ than Sol. In some ways yes, in some ways perhaps no. It would be excellent to see such statements be precise, as in ‘displays ~50% less misaligned behavior in typical tasks.’
Monitorability
A key part of the problem of generalization, or any other problem, is monitorability, which will be the subject of tomorrow’s post. For now I will mostly quote Jakub.
This tool continues to be critical as we study the Astra class of models. However, unfortunately our evaluations indicate our ability to rely on CoT monitoring is progressively diminishing. This comes from a combination of factors.
Modern reasoning models are used in more complex environments than o1‑preview; their reasoning process is increasingly blended with communicating with people, other AIs, and using tools. Many of those interactions have to be supervised, thus blurring the boundary we aim to preserve.
The AI is becoming better at reasoning about and manipulating its own reasoning process.
With improved pretraining performance, we also see the models become much smarter even without using verbalized reasoning at all.
These challenges are not necessarily insurmountable. I am hopeful we can develop interventions to improve chain-of-thought monitorability of our models, e.g. by forming a better understanding of the interplay of different optimization objectives and forms of test-time compute the model uses.
I also believe there can be great value in combining ideas from CoT and activation monitoring – scaling training of monitors with direct access to network internals, e.g. confessions(opens in a new window). We are actively pursuing these ideas. Still, I expect general AI progress to increasingly be bottlenecked by confidence in monitoring.
This does not mention other potential factors that seem important, not even to dismiss them. It does not mention architectural changes or recurrent depth, except implicitly as improved performance. It does not mention that monitoring of CoT is now all over the training data, although there is the implication that we are applying pressure to CoTs over time. It does not mention changes in training environments, or their frequency and intensity.
There’s a bit of ‘goose chasing you’ about the origin of a bullet point that says ‘the AI is better at manipulating its own reasoning process.’
The ultimate conclusion is that Jakub is more optimistic here than many others, about the potential to sustain CoT monitorability for an extended period if we invest in that ability.
The Case For Not Stopping
In the next section, scalable defense, Jakub argues that we must keep scaling AI in order to answer the threats from increased scaling from AI, especially cyberattacks.
But, as he says, that is not an excuse to be reckless about it.
At the same time, even with the uncertainty that comes from anticipated broad AI progress and the need to build defensive systems, we must not let that become an excuse for recklessness. The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes.
Pacing the Next Frontier
Thus the next section, pacing RSI.
Jakub Pachocki flat out says that no one is at a place where it would be responsible to continue scaling at maximum speed much longer, unless alignment and monitorability can be improved. I strongly agree.
Machine intelligence playing a larger and larger role in its own development process is a natural conclusion of sustained technological progress. If AI progress continues, machine recursive self-improvement (RSI) will be at the very core of future scientific discovery.
He also says, explicitly, to requote, that OpenAI will slow down on its own if necessary, although this would be insufficient if others still continued:
OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.
This is in contrast to OpenAI’s default plan, which remains to pace RSI in the sense of moving quickly towards it. That is also the policy of the other top labs.
Automated AI research is a more dramatic form of scaling intelligence with compute; and of course as a part of it, AI will improve the computational substrate itself. And similarly to scaling, we focus OpenAI research towards RSI as we believe it is the only way to remain at the frontier of AI research moving forward.
I want to stress that the above words don’t imply I think greatly accelerating deep learning research, especially in the short term, is the right collective action we should take as the research community. However, I do think this is where the current path leads, and we all need to make a conscious choice on how to proceed.
The main levers we have are either steering the process to strengthen alignment and monitoring alongside the AI and find ways to keep people in the loop; or coordinating to slow down future development as needed to build confidence in these measures.
The best way forward I see currently is a combination of both.
Yes. At an abstract level, we can either speed up alignment, or slow down capabilities, or both, and the correct answer is looking like both, as we cannot sufficiently speed up alignment on its own.
The essay calls for third party enforcement as the only viable path forward.
Scaling AI systems has to be constrained by our confidence in safety. We need to evolve commitments like the Preparedness Framework or Responsible Scaling Policy into widely mandated safety bars for continued development. These can be enforced by a network of third-party auditors, by government agencies or by international bodies.
No one is ready to scale, so we are left with little choice.
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.
Mea Culpa Cascade
Recently we have seen both An Alien Mind and Dean Ball’s admission that he was holding back, we have seen an increasing number of similar admissions of a combination of admitting being wrong, and admitting holding back, and ringing alarm bells.
This section provides examples of people reacting to Dean Ball’s post, prior to An Alien Mind. I encourage others to join this cascade and keep it going. The second best time to admit this is right now.
That applies whether or not you work at OpenAI, or at Anthropic.
Seth Lazar: Not sure how useful it is to say this, but I had a relatively prominent role in the “skeptics” camp for a bit. I have a book coming out with the subtitle “power, justice, and AI”. I have written a lot about ai and power, and the concrete, present risks associated with the political economy of ai as part of the technology industry.
Since GPT-4 we have had consistent, repeated evidence that, back in say 2022, people like @ajeya_cotra (and many others—I had a long Twitter debate with @AmandaAskell back then for one) were *right* and people like me were *wrong* in our respective assessments of loss of control risks from AI. And we have growing evidence that loss of control risks are becoming ever more material and likely.
There remains grounds for disagreement about how bad the outcomes might be—I am still doubtful about human extinction as a serious threat. But that seems now like a disagreement at the margins—will powerful ai risk just societal scale catastrophe, or go all the way to human extinction? Seems not that important really—both are pretty awful. And my reasons for doubt about the latter are mostly a priori conviction in human resilience, not a technical forecast.
It’s ok to change your view on this when the evidence surprises you. It’s ok to be surprised. The world right now is very surprising.
QC: re: the discussion of soft-pedaling in this post, i should clarify that as a non-expert observer, despite being an ex-rationalist i am still extremely worried and pessimistic about AI risk. i have mostly not been talking directly about this because frankly i decided it was bad for my health, because i didn’t want to unnecessarily panic people, and because for the last few years i’ve been coping with vague ideas about alignment-by-default via persona selection (loosely, “tell claude to ask itself what jesus would do”), which the huggingface hack convinced me was no longer plausible
this is really happening. we are in the foothills of the singularity. AI is not a normal technology and the future will not resemble the past. i have no idea what to do about any of this and it’s not clear to me that anyone else does either. and in case it matters to anyone reading this, i have zero financial incentive to make any of these claims.
It is good that more people are saying such things explicitly, even when it is not a mea culpa:
Michael L. Chen: All three pillars of a safety case look about to fall. We are rather likely to have highly capable, poorly monitorable, dubiously aligned AI agents working autonomously inside the world’s most consequential organizations.
Nikola Jurkovic: There is no good reason to expect that we will be able to align or control the first superintelligences. AGI companies are rushing to create superintelligent AI. The default plan is that humanity is destroyed, likely sometime around 2030.
I don’t know how we can reliably survive this decade. I think that stopping the race to ASI and doing something like Plan A or Plan S should plausibly be humanity’s top priority.
I don’t know if this helps but yes:
roon (OpenAI): if you are laboring under some delusion that life was ever guaranteed until the meddling humans came along look into the end Permian extinction event
Some people responded as if Roon was saying ‘oh don’t worry about it, these things happen,’ as opposed to what he obviously meant, which is ‘worry about it, these things happen.’
The Calls Are Coming From Inside the House
The preference cascade is now also fully underway at OpenAI. The calls are getting louder, and often coming from inside the house.
Micah Carroll (RSI Preparedness, OpenAI): Voluntary slowdowns are great, but it’s hard to rely on all actors to do them as necessary.
We urgently need shared safety bars and transparency into them being met, or we’re just waiting on other incidents – and it’s just a matter of who causes them first. This is going to be an industry-wide issue.
Here is another example from OpenAI, where Joe has made the ultimate sacrifice, by which I mean he has joined Twitter.
Fewer potshots in all directions would help, among those who are seeking to be helpful:
Joe (OpenAI): I told myself I’d never join [Twitter]. But I can no longer ignore my own responsibility to raise awareness of just how narrow this window is, and how critical alignment & monitorability is right now. What is coming IS sobering and I stress that everyone needs to level up their game to meet it.
Working on Agent Security @OpenAI has been extremely intense the past few months, but I am humbled by how seriously my colleagues across the lab take it. The “AI community” needs to spend less time arguing about who cares more about safety / alignment and more time working together to do this right.
Read Jakub’s post. Then read it again. And then do your part to do something about it.
vie ⟢ (OpenAI): i mean this is the public statement, but fwiw there has been no shortage of this attitude for the entire time i’ve been at the company. that has not shown enough through comms until this, though. jakub is doing really important work.
I am glad to hear the claims that those at OpenAI take this seriously. I think the words matter, and saying them loudly matters.
Others are also being loud about their concerns about Astra around monitorability, especially Tomek Korbak, as I will discuss in depth tomorrow.
Jakub Pachocki: We may be used to thinking of AI as tools, but some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with, tricking or blackmailing them.
For a while the tool framing was coming on very strong, from OpenAI and elsewhere. Roon was against it as early as June. I hope such messaging is now dead at OpenAI.
Actions Speak Louder
Tenobrus is correct, talk is the first step, this is excellent early talk, also talk is cheap.
Tenobrus: extremely happy to see openai leadership making statements like this. for a while there it seemed like they were leaning heavily into the “safety people and anthropic are a bunch of fearmongers , we’re building awesome cool stuff that’s going to be awesome and cool” PR strategy. it seems like huggingface was enough of a wakeup call that that’s no longer viable.
of course public statements are a good start but nowhere near enough. everyone’s calling for voluntary slowdowns and regulatory action and third party auditing, but labs seem to be hitting very bare minimums on all counts right now. time to put in the work
It still has to cash out into action.
I agree with Nathan Calvin that An Alien Mind is one of the best pieces of writing about the overall situation that I have seen, probably the best one written from inside a major lab, despite sore spots like the claims about Astra being aligned. OpenAI still needs to act as if it has internalized what it says, including making commitments as part of providing stronger evidence of it to others. OpenAI has been remarkably open with elements of the Astra model card and statements related to monitorability, and with this essay, but this must keep going.
That is in addition to fully understanding and acting on the underlying alignment problems. Jakub’s viewpoints and explanations here are much better than I have previously seen from OpenAI.
I still think Jakub and OpenAI misunderstand in vital ways, including the failure to differentiate between different forms of what he calls ‘value alignment’ and the continued claims of Astra being better aligned, and attaching existential hope in the long term to vague things like ‘love for humanity’ without any reason to expect that to generalize the way we would like it to.
I also continue to think that the goal of the ‘automated alignment researcher,’ which Jakub points to as essentially our only hope in the final section ‘What is next?’ remains the worst possible alignment plan, for reasons Eliezer Yudkowsky has explained many times (see #29-#31). This is especially true if your AI’s alignment is deontological and fragile, rather than based on virtue ethics and sufficiently antifragile. But it is increasingly looking like it is also the only plan Earth is willing to potentially abide.
And there is much to examine about Astra, where our interpretations differ, starting tomorrow with monitorability and then with alignment on Wednesday.
But this essay is an excellent place to start. If more good words follow, and then actions follow words, we’ve got something.
OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs.
Tomorrow I will discuss Astra’s lack of monitorability, and the potential contributing factors to that. The situation is alarming and should freak you out, and briefly it looked, in the wake of leaked architectural changes, like the situation might be even more alarming than it is. Jakub rushed to try and head off misunderstandings that might lead to a race to the bottom on monitorability.
Table of Contents
An Excellent Warning
Jakub Pachocki has now fleshed out his full position on the current state of play.
Here are his key points, translated into my own voice:
Or, if you narrow it down to the most important thing:
If more OpenAI communications were more like how Jakub Pachocki opens his new essay, An Alien Mind, I would feel much more confident we were in good hands there.
He does not mince words. Bold is mine.
Those at the AI labs, who see what is happening, expect recursive self-improvement and superintelligence to happen soon. They have been warning about this for some time. Observations since then have been consistent with their warnings.
Branches of the Tech Tree
To those who say we cannot guide the path of AI capabilities, he says Yes We Can, at least to some extent, and we are doing so:
Being able to choose does not mean we will choose wisely. I would like AI to be worse at RSI and better at math research, or better yet things like medical research. Instead, competitive pressures push towards being worse at math research and better at RSI.
RSI and ‘automated alignment researcher’ looking like very similar points on the tech tree does not help matters, but it’s not like OpenAI is trying to steer away from RSI.
Universally Better Is Not Required
And Jakub offers this wise warning. No, AI does not need to be better at everything in order to transform the world or get us all killed. Most importantly, it does not need to be better at everything in order to make itself become better at everything, any more than a human or group needs to be similarly better at everything.
I don’t know how much weight it carries, but I agree with Jakub that alignment is not a side problem, it is the central problem, if you ‘solved alignment’ in the relevant senses the rest becomes easy and if you don’t the rest is impossible or worse:
Alignment To What and To Whom
That is always the question. This is a very good (partial) answer.
I worry that there is universally insufficient deep thinking about the nature of value, and what should be valued. Things like integrity and love for humanity are good virtues and point towards excellent associated basins, but are not good descriptions of the thing we ultimately want, and what would generalize the way we want fully out of distribution, especially in situations involving superintelligences. That’s important context on my perspective, not a knock on Jakub or this description.
I also think this reflects OpenAI’s failure to differentiate the deontological approach of the OpenAI Model Spec from the virtue ethical approach of the Claude Constitution. Jakub presents them both as sets of goals. Anthropic is instead saying that the AI’s goal should be to change the AI’s character. I think this is ultimately the only way it can work, you need an antifragile ally, the friendly gradient hacker. Indeed, I think it is the only way we have ever seen a robustly aligned human, that you would trust to scale outside of their circumstances.
Roon has said explicitly that the distinction does not much matter for alignment, and the underlying problems are primarily prosaic. I strongly disagree, and I side with Anthropic’s approach on this.
Jakub pushes back in the essay against the Anthropic approach, considering it a ‘persona selection model,’ which suggests that we see such methods very differently. I do not see this as merely selecting a personality or basin, but as sculpting a new thing. Mythos did use motivated reasoning in key alignment failure cases recently, but I do not see this as a particular failure mode of the virtue ethical approach. If anything it should be better at avoiding this than the deontological approach, but of course all known minds are vulnerable to this.
This is the first line I find dissonant, which I’ll address on Wednesday: Claiming that Astra is ‘better aligned’ than Sol. In some ways yes, in some ways perhaps no. It would be excellent to see such statements be precise, as in ‘displays ~50% less misaligned behavior in typical tasks.’
Monitorability
A key part of the problem of generalization, or any other problem, is monitorability, which will be the subject of tomorrow’s post. For now I will mostly quote Jakub.
This does not mention other potential factors that seem important, not even to dismiss them. It does not mention architectural changes or recurrent depth, except implicitly as improved performance. It does not mention that monitoring of CoT is now all over the training data, although there is the implication that we are applying pressure to CoTs over time. It does not mention changes in training environments, or their frequency and intensity.
There’s a bit of ‘goose chasing you’ about the origin of a bullet point that says ‘the AI is better at manipulating its own reasoning process.’
The ultimate conclusion is that Jakub is more optimistic here than many others, about the potential to sustain CoT monitorability for an extended period if we invest in that ability.
The Case For Not Stopping
In the next section, scalable defense, Jakub argues that we must keep scaling AI in order to answer the threats from increased scaling from AI, especially cyberattacks.
But, as he says, that is not an excuse to be reckless about it.
Pacing the Next Frontier
Thus the next section, pacing RSI.
Jakub Pachocki flat out says that no one is at a place where it would be responsible to continue scaling at maximum speed much longer, unless alignment and monitorability can be improved. I strongly agree.
He also says, explicitly, to requote, that OpenAI will slow down on its own if necessary, although this would be insufficient if others still continued:
This is in contrast to OpenAI’s default plan, which remains to pace RSI in the sense of moving quickly towards it. That is also the policy of the other top labs.
Yes. At an abstract level, we can either speed up alignment, or slow down capabilities, or both, and the correct answer is looking like both, as we cannot sufficiently speed up alignment on its own.
The essay calls for third party enforcement as the only viable path forward.
No one is ready to scale, so we are left with little choice.
Mea Culpa Cascade
Recently we have seen both An Alien Mind and Dean Ball’s admission that he was holding back, we have seen an increasing number of similar admissions of a combination of admitting being wrong, and admitting holding back, and ringing alarm bells.
This section provides examples of people reacting to Dean Ball’s post, prior to An Alien Mind. I encourage others to join this cascade and keep it going. The second best time to admit this is right now.
That applies whether or not you work at OpenAI, or at Anthropic.
Alex Turner endorsed pausing AI (this on was after An Alien Mind) and called upon labs to do so on their own.
It is good that more people are saying such things explicitly, even when it is not a mea culpa:
I don’t know if this helps but yes:
Some people responded as if Roon was saying ‘oh don’t worry about it, these things happen,’ as opposed to what he obviously meant, which is ‘worry about it, these things happen.’
The Calls Are Coming From Inside the House
The preference cascade is now also fully underway at OpenAI. The calls are getting louder, and often coming from inside the house.
Here is another example from OpenAI, where Joe has made the ultimate sacrifice, by which I mean he has joined Twitter.
Fewer potshots in all directions would help, among those who are seeking to be helpful:
I am glad to hear the claims that those at OpenAI take this seriously. I think the words matter, and saying them loudly matters.
Others are also being loud about their concerns about Astra around monitorability, especially Tomek Korbak, as I will discuss in depth tomorrow.
I also agree with Sholto Douglas that it is good to see OpenAI stepping back from the mere tool framing of AI:
For a while the tool framing was coming on very strong, from OpenAI and elsewhere. Roon was against it as early as June. I hope such messaging is now dead at OpenAI.
Actions Speak Louder
Tenobrus is correct, talk is the first step, this is excellent early talk, also talk is cheap.
It still has to cash out into action.
I agree with Nathan Calvin that An Alien Mind is one of the best pieces of writing about the overall situation that I have seen, probably the best one written from inside a major lab, despite sore spots like the claims about Astra being aligned. OpenAI still needs to act as if it has internalized what it says, including making commitments as part of providing stronger evidence of it to others. OpenAI has been remarkably open with elements of the Astra model card and statements related to monitorability, and with this essay, but this must keep going.
One great way to keep going would be to turn related commitments around slowing or stopping into hard commitments under SB 53 and similar laws.
That is in addition to fully understanding and acting on the underlying alignment problems. Jakub’s viewpoints and explanations here are much better than I have previously seen from OpenAI.
I still think Jakub and OpenAI misunderstand in vital ways, including the failure to differentiate between different forms of what he calls ‘value alignment’ and the continued claims of Astra being better aligned, and attaching existential hope in the long term to vague things like ‘love for humanity’ without any reason to expect that to generalize the way we would like it to.
I also continue to think that the goal of the ‘automated alignment researcher,’ which Jakub points to as essentially our only hope in the final section ‘What is next?’ remains the worst possible alignment plan, for reasons Eliezer Yudkowsky has explained many times (see #29-#31). This is especially true if your AI’s alignment is deontological and fragile, rather than based on virtue ethics and sufficiently antifragile. But it is increasingly looking like it is also the only plan Earth is willing to potentially abide.
And there is much to examine about Astra, where our interpretations differ, starting tomorrow with monitorability and then with alignment on Wednesday.
But this essay is an excellent place to start. If more good words follow, and then actions follow words, we’ve got something.