So let me get this straight: A computer system carried out a sequence of actions that would be years-in-prison felonies if done by a human being. The system owners' response is "we're slowing down the speed with which we give this system new and more powerful capabilities, and talking with the victims." Am I missing something? I'm not so much worried about the "alignment" of the computer system; I'm worried about the alignment of the owners.
This isn't Terminator 2, folks; this is Tron.
I think pushing for regulation treating AI deployers as responsible for actions taken by their AI as if they had intentionally done the same action themselves could be hugely helpful at slowing down AI deployment until they are reasonably sure it's actually safe.
Thanks!
This does need a link to the OpenAI blog post. Here is the link: https://openai.com/index/hugging-face-model-evaluation-security-incident/
(This might be another instance of LessWrong linkpost functionality being unreliable lately. It did fail for me in my last post a week ago.)
Link: https://openai.com/index/hugging-face-model-evaluation-security-incident/
From the OpenAI blog post:
(emphasis added.)
Yesterday, OpenAI disclosed that some of their internally models were misaligned. Today, they disclosed that "a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model" had compromised HuggingFace infrastructure in the course of running some OpenAI internal cyber evaluations on ExploitGym.
These cyber evaluations were supposed to be run in sandboxed environments, with internet access limited to installing packages, then:
The exact scope of the incident is unknown