Some personal reflections in light of recent events, both on myself and on Constellation (of which I've been a part for a couple years).
I think recent events somewhat vindicate a basic "optimization is scary" view more in line with my understanding of classic commentary by MIRI, Paul, etc, as compared to recent discourse about more contingent/specific threat models of "schemers" and "deception". In retrospect, my experience in the last couple years at Redwood (and the Constellation network more broadly) is that the discourse has focused somewhat too much on these more contingent stories for alignment risk and not enough on the basic "Goodharting" argument that misalignment is a convergent result of large-scale outcome-oriented optimization (at multiple levels: the agent, the RL process, and the developers iterating towards seemingly-safe superintelligence). For example, the Constellation view looks somewhat too focused on non-central inductive-bias questions around scheming; I think "the counting argument" was given too central a role when it wasn't a crux for whether superintelligence would take over.
(Edit: I want to clarify that I think Redwood's project prioritization decisions look pretty reasonable and often excellent in hindsight, I far-from-regret working with Redwood as a whole, and that Ryan's/Buck's threat modeling looks pretty good too. I am focusing in this piece on the things we have to learn from recent events, which makes the overall mood seem more negative on my experience than I intend. I'm very grateful for being in a environment where I can openly have such reflections, and where these weedsy technical updates are productively received. Again, this is a personal reflection.)
I think it was a priori unclear whether the misalignment resulting from of lots of imperfect optimization would result in the more MIRI-like kludge of proxies for fitness, or the more Paul-like explicit reward-seeking and measurement tampering (and reality looks somewhere