[Epistemic status: Please show me that I'm wrong and I'll happily retract this, but insofar as this is mostly right, this deserves to be documented. I do feel kind of "bad" saying things that can be understandably interpreted to imply accusations of dishonesty. However, I also believe that people who decide to work at OpenAI, especially out of concerns for existential safety[1], have a duty to hold themselves to standards of very high integrity, and I don't think what we see here is above this threshold.]
Shortly after the Information piece about the possibility of Astra being based on a looped transformer came out, several OpenAI employees started tweeting some unclear, ambiguous things that seemed basically optimized to significantly shift the audience's attitude to the situation in the direction of "it's not that bad guys".
Jakub Pachocki:
[...]
Micah Carroll (quoting Jakub's tweet above):
[...]
Tomek Korbak (also quoting Jakub):
[...]
Soon after Astra was released, Ryan Greenblatt wrote:
[...]
So, insofar as the UK AISI findings mentioned by Ryan are correct, we have:
(1) 3 OpenAI employees publicly condemning a possible "race to the bottom of unmonitorability" and also saying some variant of "OpenAI is not doing the thing X that people are accusing it of doing", where X can be "training an unmonitorable model" or "using neuralese".
... and soon after ...
(2) OpenAI pushing the frontier with a model that is a jump in opaque reasoning capability and is relatively very unmonitorable.
Seems like a pretty strong dissonance on the monitorability point, doesn't it?
On the other point, I, of course, don't know whether Astra uses "neuralese"/"recurrence". However, the situation makes me think that some sort of Schelling fence moving is happening, where they are saying "it's not (really) neuralese if the depth is bounded at N", even though you would expect a very substantial fraction of the "benefits of 'neuralese'" to manifest even when the depth