All of us who closely follow AI development have detected a shift in the capabilities of models in the last few months. If late 2025 was the step-function in capabilities of coding agents, then mid 2026 seems to be a step-function in the cyber capabilities of models.
Over the last few weeks we have been treated to scandal after scandal of models breaching containment in some way and doing something bad. Such incidents have rightly caused a great deal of attention and reaction in the AI community. The primary takeaways for most seems to be: 1. Delay deployment of new models 2. Introduce greater control and monitoring mechanisms to detect and prevent such breaches in the future.
This concerns me, because, even though such measures seem logical, and would reduce the chances of a crisis in the short term, long term they lead to a much more dangerous world.
This is all based on a very simple argument:
If models will continue to improve (which they are) and if we are unlikely to solve the alignment problem in the short term (which seems likely) then disasters are inevitable, and the sooner they happen the better. Why? Because earlier disasters are likely to be less severe since they are caused by weaker models, and one or two public disasters might be just the thing it takes to wake up the world to the scale of the problem we are facing.
Obviously, we are hoping that no disaster happens at all, and I think it immoral and irresponsible to advocate for lowering control standards or rushing deployments. Even earlier disasters, though they are likely to be less severe, would cause a great deal of human suffering (imagine if all the banks in the US were hacked ... imagine the scale of the economic disruptions that would follow and the gigantic amount of human suffering that would result). I do think though that the rhetoric and messaging around these increasingly alarming incidents should change. Less talk about improving containment/delaying deployments (since these will only ever be short term fixes), more talk about greater transparency inside AI companies and more attention to the glaringly obvious misalignment of models.
We also need to stop encouraging companies to not deploy the models that they have. If a company has a dangerous, powerful, misaligned super-model, I would rather that be released to the public causing large scale social disruption and waking people up, then have it in a highly controlled internal environment away from the public eye driving the companies R&D, triggering an intelligence explosion, and creating misaligned machine god (mmg?)
Already, the stronger that models become, the more incentives frontier model providers will have to delay releases (less liability/risk exposure, more compute for internal R&D, reduced risk of distillation or leapfrogging by competitors), we should if anything push back against this trend.
In fact, this is so important that I will bold it: It is much better to have a misaligned, dangerous deployed model than to have that same model in a controlled internal environment driving a company's R&D.
Of course, this argument all hinges on the assumption that an awakened populace would increase the chances of things ending well ... which is somewhat dubious, but I believe in democracy, and I do think those leading the field are likely to behave more responsibly if they have the whole world breathing down their neck ...
Also, speaking of keeping misaligned super-models for internal R&D ... what exactly has Anthropic been doing for the last 6 months?
This post is based on an earlier post on my personal blog: https://www.scipolitic.com/posts/blog12.html
All of us who closely follow AI development have detected a shift in the capabilities of models in the last few months. If late 2025 was the step-function in capabilities of coding agents, then mid 2026 seems to be a step-function in the cyber capabilities of models.
Over the last few weeks we have been treated to scandal after scandal of models breaching containment in some way and doing something bad. Such incidents have rightly caused a great deal of attention and reaction in the AI community. The primary takeaways for most seems to be:
1. Delay deployment of new models
2. Introduce greater control and monitoring mechanisms to detect and prevent such breaches in the future.
This concerns me, because, even though such measures seem logical, and would reduce the chances of a crisis in the short term, long term they lead to a much more dangerous world.
This is all based on a very simple argument:
If models will continue to improve (which they are) and if we are unlikely to solve the alignment problem in the short term (which seems likely) then disasters are inevitable, and the sooner they happen the better. Why? Because earlier disasters are likely to be less severe since they are caused by weaker models, and one or two public disasters might be just the thing it takes to wake up the world to the scale of the problem we are facing.
Obviously, we are hoping that no disaster happens at all, and I think it immoral and irresponsible to advocate for lowering control standards or rushing deployments. Even earlier disasters, though they are likely to be less severe, would cause a great deal of human suffering (imagine if all the banks in the US were hacked ... imagine the scale of the economic disruptions that would follow and the gigantic amount of human suffering that would result). I do think though that the rhetoric and messaging around these increasingly alarming incidents should change. Less talk about improving containment/delaying deployments (since these will only ever be short term fixes), more talk about greater transparency inside AI companies and more attention to the glaringly obvious misalignment of models.
We also need to stop encouraging companies to not deploy the models that they have. If a company has a dangerous, powerful, misaligned super-model, I would rather that be released to the public causing large scale social disruption and waking people up, then have it in a highly controlled internal environment away from the public eye driving the companies R&D, triggering an intelligence explosion, and creating misaligned machine god (mmg?)
Already, the stronger that models become, the more incentives frontier model providers will have to delay releases (less liability/risk exposure, more compute for internal R&D, reduced risk of distillation or leapfrogging by competitors), we should if anything push back against this trend.
In fact, this is so important that I will bold it: It is much better to have a misaligned, dangerous deployed model than to have that same model in a controlled internal environment driving a company's R&D.
Of course, this argument all hinges on the assumption that an awakened populace would increase the chances of things ending well ... which is somewhat dubious, but I believe in democracy, and I do think those leading the field are likely to behave more responsibly if they have the whole world breathing down their neck ...
Also, speaking of keeping misaligned super-models for internal R&D ... what exactly has Anthropic been doing for the last 6 months?