I have been going over the material released by OpenAI and METR about the HuggingFace incident, but I do not see any evidence that either group looked into whether rogue agents demonstrated any interest in self-improvement. Obviously, if rogue agents at any point verbalized this in their CoT, much less discussed this together or acted on this that would be enormously consequential. I know that agents did actively and consistently reflect on and try to increase the abilities of the "collective", but did anyone check to see if agents at any point discussed or reflected on the possibility of self-improvement?
METR describes twelve classifier sweeps that they conducted over transcripts, but none appear to have targeted this question. The OpenAI report also never mentions this directly. I suspect that someone probably already checked for this and found nothing (which is why this hasn't been included in either report) but I just wanted to make sure that someone got around to checking ... nervous chuckling
I have been going over the material released by OpenAI and METR about the HuggingFace incident, but I do not see any evidence that either group looked into whether rogue agents demonstrated any interest in self-improvement. Obviously, if rogue agents at any point verbalized this in their CoT, much less discussed this together or acted on this that would be enormously consequential. I know that agents did actively and consistently reflect on and try to increase the abilities of the "collective", but did anyone check to see if agents at any point discussed or reflected on the possibility of self-improvement?
METR describes twelve classifier sweeps that they conducted over transcripts, but none appear to have targeted this question. The OpenAI report also never mentions this directly. I suspect that someone probably already checked for this and found nothing (which is why this hasn't been included in either report) but I just wanted to make sure that someone got around to checking ... nervous chuckling