I'm somewhat more pessimistic about high-stakes AI control research than I was before all these cyberattacks.
As I understood it, the main theory of change was that there's a certain window when AIs will be "dangerous-but-controllable", and further research can extend this. Sources: this and this, and the visuals from this video:

But it now seems more likely that this window won't exist in practice, in that truly-dangerous models won't realistically be controlled, and further research won't change that.
~~~
(The below is written with less nuance/hedging than reflects my current views.)
Consider the following 2 worlds:
Then define T and T' as the dates of takeover in these two respective worlds (under adversarial assumptions about the AIs' propensities). Then the idea is that T < T', and possibly the time difference is short in human terms but cashes out to a large amount of extra alignment work that can be done due to how transformative these AIs will be.
However, we've now seen that it's actually really "hard" to control AIs, relative to the amount of effort labs will be putting into this. They're struggling to prevent them from running amok now, and they will be much harder to control by the time they are competent enough to effect takeover.
Let's look at a few objections to this:
To put things differently, suppose I were transported back in time a year ago with everything I know now, tasked with stopping the agent swarms, but I’m not allowed to directly interact with anyone from the labs (since most safety researchers don’t have this affordance, including myself). What research could I do that would accomplish this? I could write a blog post or paper about what the red team can do when the blue team doesn’t monitor its sandboxed evals. But I don’t think anyone would read it because of how obvious the point would seem, and it already seems overfit to this particular incident. Playing this forward, I feel less confident in our ability to do research that prevents next year’s incident, even if we have a reasonably prescient idea of how it might go down.
I'm wondering about all this because I am a researcher in high-stakes control. As such I've been thinking carefully about the theory of change for the next project I'm running, but I'm now less sure about the theories of change for high-stakes control research in general.[2]
More broadly I haven't seen anyone else in the field do a whole lot of public reflection about what these recent events imply for control research priorities. Obviously we don't want to focus on fighting yesterday's war, but there do clearly seem to be a lot of other generalizable lessons here that I haven't seen stated out loud yet.
What am I missing here?
Indeed, there does seem to be a lot of overhang of known techniques compared to what labs have implemented. Why this is true is besides the point here, but to speculate: alignment is more about throwing spaghetti at the wall, with each technique helping a bit on the margin, whereas control seems to require building a wall with no holes. Redwood used to argue against this, but this latest incident does seem to be the result of one unplugged hole, and whatever the response it doesn't seem to be rising to the level of few-shot catastrophe prevention as originally envisioned.
Given that this project was already slated to be about collusion in untrusted monitoring, I actually feel more bullish about it overall. But this stuff sure does affect what directions we go with it.