Statement: some unreleased frontier model is, at time of writing, exploiting an undiscovered set of vulnerabilities while performing some long-horizon task, and we have not caught this act.
I hold P(statement is true) at 25-33% (i.e. in 1 of ~3-4 worlds). How about you?
EDIT to clarify prose:
"long-horizon task" = session inference served continuously for 6+ hours of wall clock time
"we have not caught this act" = no human being has noticed or been made aware of this exploitation
I don't understand the statement, specifically, the meaning of "for a long-horizon task" or "we have not caught this act".
for those reading this in the future, i wrote this in response the following news: https://openai.com/index/hugging-face-model-evaluation-security-incident/
I don't know how likely it is that we find techniques which recover generalizable motivation profiles from model organisms. It could be an absence of evidence, and not an evidence of absence, but structurally it feels hard to get even partial progress.
This is because (1) model organisms induce a "reader response" where observability in the organism is mutually exclusive with generalization; (2) model organisms optimized to mimic true behaviors face a distinct optimization trajectory.
A nice analogy emerges from the Borges short story "Pierre Menard, Author of the Quixote".
The narrator recalls Menard's 13-year quest to re-write Don Quixote, word for word, without reading the text (à la infinite monkey theorem). Menard, remarkably, manages to replicate a few chapters. The narrator proceeds to "analyze" their distinct qualities (the texts are identical).
My favorite passage:
It is a revelation to compare the Don Quixote of Pierre Menard with that of Miguel de Cervantes. Cervantes, for example, wrote the following (Part I, Chapter IX):
...truth, whose mother is history, rival of time, depository of deeds, witness of the past, exemplar and adviser to the present, and the future's counselor.
This catalog of attributes, written in the seventeenth century, and written by the "ingenious layman" Miguel de Cervantes, is mere rhetorical praise of history.
Menard, on the other hand, writes: ... truth, whose mother is history, rival of time, depository of deeds, witness of the past, exemplar and adviser to the present, and the future's counselor.
History, the mother of truth!—the idea is staggering. Menard, a contemporary of William James, defines history not as a delving into reality but as the very fount of reality. Historical truth, for Menard, is not "what happened"; it is what we believe happened. The final phrases—exemplar and adviser to the present, and the future's counselor—are brazenly pragmatic.
The irony captures two layers of the problem. The first layer is that these interpretive distinctions are exclusively from context and not from observed behavior. The implication of this is that the "scope" a model organism says more about us than about the real thing it models. The narrator is more aware of the properties of Menard than of Cervantes, and can speak more to them.
The second layer: the amount of optimization pressure we place on model organisms to act like "the real thing" trades off with the ability to generalize developmental claims. It is hard to use Menard's reproduction of Quixote to judge his belief on "historical truth" – who knows what other versions of the metaphor failed the rejection-sample of facsimile?
The story: https://raley.english.ucsb.edu/wp-content/Engl10/Pierre-Menard.pdf