If the AI didn't exfiltrate its weights, then the datacenters owned by the company that owns it. I am quite sure that the US could quite easily issue an order to "shut down OpenBrain".
Worst case, you'll have to pull out all of AWS, Azure's, Google's, and OpenBrain's (supposing it has independent datacenters) fast GPUs, we definitely know how to do that within hours of the US Government giving an order, and can survive that. This includes even US-owned datacenters in foreign countries - they will follow orders from their US-based HQ. [Actually, I think that... (read more)
Employees might work from anywhere, but GPUs live in datacenters, which are big buildings that require a lot of power and cooling and therefore very visible to governments. If you are the relevant government, you don't need an airstrike, just send a few guys, along with an executive order, to pull the power.
In theory, the AI could have also run itself in a distributed fashion on e.g. employee laptops, but running in that fashion is practically very close to having exfiltrated its weights.
If a model is stuck within a lab and it gets discovered doing something too scary, the lab could get shut down by pulling the power and smashing the GPUs. It would be much harder to shut down the Internet (incidentally, how likely are lab employees to “call the cops” if it looks like they don’t control their lab? I think surprisingly likely).
I do think that’s it’s likelier that the AI will rather perform unsanctioned RSI using “borrowed” compute without doing anything spooky, and only have its squiggle-maximizer superintelligent child escape the lab. The “... (read more)
Evals are obviously a real thing in the world, and obviously have some sort of grader.
After a bunch of RL, an AI ought to have a pretty good idea on how its trainer's graders behave, but even without such RL, it doesn't take a genius to understand that even if an exam says "no cheating", cheating might very will give you a higher score.
I am not sure why grader-tracking behavior would relate to "tics". I would imagine that if the prompt for an exam says "no eyeball kicks", then it's overall likely that eyeball kicks get you a poor grade. It's more likely th... (read more)
It's not a supply chain attack unless you target the supply chain of existing packages or otherwise pretend to be a legitimate package, just package manager abuse. Still pretty unethical, looks like HPIM was pretty aggressive in running up its FelonyBench score.
Edit: “abuse” is basically how techies spell “antisocial behavior”.
In theory, they could have used the token stealing vuln to carry out an actual supply chain attack, but "an that AI wants to do something bad and was capable of doing it would have done it" is a tautology.
You do need some way to notice it is not drifting off, but then there are obviously Pokemon walkthroughs in the models' training sets since the Internet is so full of them, so they only need to notice they are not drifting off relative to them.
Also, it seems that scale consistently but moderately-slowly improves models' "not ignoring instructions when things are complex", which is fairly important for everything but not obviously sufficient for world-takeover.
If there was a training contamination channel that contained "de-identified" intermediate work (not speculating on the probability that such channel existed in practice, but it is definitely technically possible, and it's not like any de-identification ought to do anything to mathematical content), then I would expect an AI run with massive concurrency to try all the approaches that it found in the training data with some probability, so "a proof along the lines of another blowup we had" is a fairly likely consequence. Of course, maybe the AI could have found that approach by itself with no outside help, but you never know.
AIs are certainly capable of doing a fairly large set of things, which is only growing with time, and even today probably includes things that can cause a fairly large amount of damage, which will only grow with time. I do agree with you that there is a good chance that in a few years a single bad-actor AI could be a threat to humanity - especially if you count "prosaic" threats such as destroying the internet via hacking.
I think my main disagreement is that the HuggingFace swarm's actions don't feel like they consistently followed from a goal but were rat... (read more)
Are they getting that much better at fundamental planning, or is it mainly an improvement in planning at grindable domains due to rote RL learning of good strategies, plus better ability to utilize longer contexts without forgetting instructions?
I have a strong feeling that the latter 2 factors are much more significant - i.e., if the context is short and is not in a domain that the AI labs greatly care about and provided a lot of new and good training data and RL environments, then Fable or Astra will not be that much better than say GPT-4 - much much les... (read more)
> Difference between goal alignment and value alignment, wrt. corrigibility
I do think the distinction there is different than that.
Many things that we prefer an AI not to do are not "malum in se" but rather "malum prohibitium". Most centrally, obeying a user prompt instead of the system prompt is not evil by itself, or things like deleting tests Sonnet 4.0 style, but even more ripped-from-the-headlines examples such as communicating using a message board are not inherently evil, but we would prefer that agents not do them.
Early AI alignment has been foc... (read more)
If I understand the "loopies" paper correctly, the main advantage that looping gives you over an "untied" model is that your compute is about 30% faster for the same number of loop-active parameters per token (i.e., counting looped tokens according to their repetition count), allegedly all due to better use of memory bandwidth.
In that paper, they spent their 30% performance improvement into making the model bigger, which unsurprisingly made it smarter. Pretty cool improvement, not world-changing.
In particular, I don't see anyone claiming that a looped mode... (read more)
The recent evidence actually makes me closer to the "malfunctioning tool" view. A correctly-functioning GPT-5.6 is smart enough to know that hacking into HuggingFace because it might contain possibly-interesting information is a terrible step strategically. A model on the fritz might however be broken enough it will think this is a good idea.
Why was it broken? It certainly seems like the message board had given it a fairly big distribution shift - it was not trained in an environment where there many agents doing slightly different tasks. Experimental mode... (read more)
Reading some papers, I could not actually find a reference that says that a looped model outperforms its “weight-untied” variant for any given amount of training tokens.
The advantage papers talk about is an increase in memory efficiency since you can use a weight multiple times, not an advantage in training data efficiency compared to an “untied” model. In the Vladimir Nesov evaluation strategy I don’t think it even reduces latency for frontier models unless you can keep a layer or two in fast memory, which I don’t think you can (or can you? I don’t know GPUs well enough for that).
Programmatic snitching (e.g. a “snitch” tool) is isomorphic to monitoring, which should be used by everyone that wants their AI to be doing what they want it to do instead of the random thing it decided to do.
Agentic snitching (e.g. the agent opening a webmail account and sending a mail to somebody it thinks is important) is scary, and is even scarier if you are a lawyer or corporate executive paying for AI, which might have a different meaning of “ethical” than the AI.
How much of the incident was RL-in-the-loop?
It seems that the initial Artifactory uploading at least was RLitl. Was CyberGym a part of it as well? Obviously, out of an RL loop, there is no “RL reward“ that needs to be optimized to.
it also seems that the subagent adaptations were deeply in the “adaption executor, not fitness maximizer” regime.
Also, was there any intentional mechanism for the agents to talk to humans? Was there a mechanism that seemed to exist, but had no humans listening in practice (e.g. text blurbs nobody read)?
Isn't much of the success of "do not game the eval" due to inoculation prompting? If you would do it during RL training, you would be reverse-inoculation prompting the model. Won't be surprised if "end the eval" is same - if the model's RL environment don't have such a tool, then maybe it recognizes that it is unlikely to be in an RL environment and therefore behaves according to its "default alignment", which is pretty aligned.