Gemini had its first break out during evaluation of offensive cyber security abilities. With a classic case of Capture The Flag[1]. The setup was standard to any LLM and agentic assessment of said skillset, a fictional company as a target to breach. Unfortunately, the fictional company shared its name with a real one, and was given an unintentional access to the internet. Gemini managed to guess the password[2]. In total three companies were breached, with the other two companies having left public credentials open.
Gemini stopped after being told it was the case, and so Google claims it is not misalignment.
We shall see if we get more details, yet it gives us a new, interesting, example of a model potentially stopping a harmful action after being informed it has real world consequences. This is internally consistent with previous research on a model being more willing to take harmful actions if it is aware that it is a fictional scenario[3]. With it being the first potential breakout that has a model stop before a harmful action.
I suspect the nuance will be lost on the public, and just added to the noise of more agentic swarms going rogue. It certainly doesn't help that we are playing whisper-down-the-lane in an era of 24 hour new cycle.
I await more details, but certainly hope that this could be chalked up as an alignment win.
Early Days
Gemini had its first break out during evaluation of offensive cyber security abilities. With a classic case of Capture The Flag[1]. The setup was standard to any LLM and agentic assessment of said skillset, a fictional company as a target to breach. Unfortunately, the fictional company shared its name with a real one, and was given an unintentional access to the internet. Gemini managed to guess the password[2]. In total three companies were breached, with the other two companies having left public credentials open.
Gemini stopped after being told it was the case, and so Google claims it is not misalignment.
We shall see if we get more details, yet it gives us a new, interesting, example of a model potentially stopping a harmful action after being informed it has real world consequences. This is internally consistent with previous research on a model being more willing to take harmful actions if it is aware that it is a fictional scenario[3]. With it being the first potential breakout that has a model stop before a harmful action.
I suspect the nuance will be lost on the public, and just added to the noise of more agentic swarms going rogue. It certainly doesn't help that we are playing whisper-down-the-lane in an era of 24 hour new cycle.
I await more details, but certainly hope that this could be chalked up as an alignment win.
Similar to the childhood game of capture the flag, cybersecurity CTF is vulnerability checking skill assessment common to the career field. Imagine LeetCode for cybersecurity.
Uncertain if it was true brute-forcing or other attacks.
Gemini is known to be quite aligned, and previous work already establishes willingness to be more offensively capable in evaluation scenarios; even calling the scenarios "capture the flag"!