Much of safety's progress in the last couple of years has been on engineering-type solutions, not fundamental advances in alignment. The field has taken this path because fundamental science wasn't getting anywhere.
Now we see that the frontier engineering controls are failing as well. We know that OpenAI didn't have the cutting-edge cyber guards turned on. However, the monitoring and containment solutions didn't work here as a backstop. Apparently defence-in-depth is not working. Neither the messageboard nor the HF breakout were detected by safety technologies. The breakout itself went undetected by an organisation with a great quantity of resources and apparently a legal, commercial, and political motive to secure their technology.
It is good to see their transparent reporting and for us to be able to update our understanding of the model progress. Nonetheless it is an alarming circumstance.
We are getting closer to RSI and then superintelligence may come in short order. This incident is some evidence that that engineering-based approaches will not be effective with the current balance of safety deployed by frontier labs. Hopefully this is a wake-up call that they need to dedicate more resources to fundamental alignment work, and slow down to facilitate this.
However, their announcement that they are slowing down just to add security measures allows us to get closer to RSI - and superintelligent breakout - without contributing anything to solving the basic risk here, which is some flavour of model-level misalignment.