Thank you for your research!
One connection that comes to my mind is the concern that, in pursuiing a goal, an AI system could find acquiring additional resources (compute/data/tool access/influence over people) increases its probability of success. This would become more relevant as models can operate over increasingly longer sessions.
But based on your findings, it appears that models can learn to optimize its strategy given resource constraints while preserving performance. Is the implication here that we could at least reduce incentives for unnecessaril... (read more)
Very interesting thought process. If I'm understanding correctly, the key risk comes from a capability leap that is too large, such that we have very low probability of aligning the next model.
Reading your post, I find myself more optimistic about AI alignment than beating Magnus.
First, one distinction between AI alignment and the chess analogy is that we don't have to "beat" AI. Unless the model is actively resisting us (which means we've done something terriblely wrong), we don't need to outthink a much more intelligent opponent at its own game. At lea... (read more)
Very thought provoking question. To prevent what you described from happening, I think the best way is to have open source models that remain close to the frontier, so that ai is not concentrated in a few companies or governments.
I also wonder to what extent the end state you described is the most likely path. We will probably go through a period of "hybrid" workforce where ai needs to work with humans. There will be political parties, firms, governments, alliances, and rival states all competing for "power" (quoted because what counts as power can vary w... (read more)