It seems pretty clear from the METR Hugging Face report that our ability to assess what an agent does depended mostly on the logs, with some supplementation based on interviews with OpenAI employees. By my count the METR report relies on three different record types; including the infrastructure logs (artifactory) which they didn't succesfully alter but they did try to delete evidence of wrong-doing from their own containers, the agent-to-agent communications, and their internal chain of thought reasoning.
The EU AI act mandates that companies retain logs, but only for systems that are on the market and are deployed, the logs, for at least six months. Even if the logs they mandate a company keep did include pre-deployment testing artifacts or logs associated with models that are not publicly released, the EU AI act does not mandate where logs are kept. Technically, a log kept in the agent's enviornment, which we now know they can alter, or logs kept in such a way that they are susceptible to tampering would both satisfy the requirements of Article 19. Finally, its not clear whether agent-to-agent communications or chain of thought reasoning are considered events under the Act.
Our current state of frontier AI governance is so woefully inadequate that the first major warning shot we have would not have been reconstructed under the current regulatory regime.
I think PauseAI global made a poor decision disendorsing PauseAI US and Holly Elmore, but I am not sure how much they matter, and I'd rather spend the time reflecting on why I think Holly's rhetorical strategy is valuable and necessary.
I'm turning 43 this month. I share that because it means I came of age during the New Atheist fights, and Holly reminds me of a polemical pugilist very much in the tradition of Christopher Hitchens. Hitchens' rhetorical strategy was very effective in my own turn toward secularism and atheism. He was effective where others were not because he so clearly displayed his contempt for what he believed to be a dangerous and scientifically unsound set of ideas.
For some of us who had never questioned our indoctrination into religion until we arrived at college, he was a compelling and forceful advocate for reason. I suspect there will be people in the future, if we get this transition right, who look back and trace their own understanding of these risks and what is at stake to Holly Elmore, precisely because of the way she communicates.
Many of the arguments lodged against Holly today were also lodged against Hitchens. He belonged to a class of public intellectuals who were prominent when I was in my early-to-mid 20s. They regularly, as a matter of strategy, adopted a rhetorical style filled with contempt and attacked religiosity and its proponents with fervor through books, debates, interviews, and op-eds. Between 2007-2011 there was a prominent public conflict between Hitchens and co and religious leaders and proponents. At the time, the New Athesists were a bit norm-breaking and shocking in their communication patterns.
They weren't measured or circumspect in their comments. In fact, they refused to argue with religious leaders and their ideas on their own terms. They wouldn't take them or their beliefs seriously or at face-value; they attacked them, their history, and the actions of their leaders directly by name in some instances. And they did so with humor, condescension, and contempt. They were aggressive in asserting that religion, as an institution, was particularly dangerous, filled with irrational ideas and people, and net harmful to humanity.
And as such, religion was an appropirate target for contempt.
An illustrative quote from Hitchens: "Mother Teresa was not a friend of the poor. She was a friend of poverty. She said that suffering was a gift from God. She spent her life opposing the only known cure for poverty, which is the empowerment of women and the emancipation of them from a livestock version of compulsory reproduction."
In their heyday they had many supporters, and even among those supporters there were many who suggested that Hitchens, Harris, and Dawkins should be more accommodating, kinder, more careful, so that their message reached more people and avoided offending religious folks. These supporters and detractors argued that they should avoid attacking the Catholic Church, its bishops, or anti-poverty advocates like Mother Teresa. Their claim was that the nature of the attacks went beyond the pale and hurt the movement.
I thought then, as I do now, that dangerous ideas and dangerous people must be addressed, by at least some members of a movement, in a pugilistic fashion. If there is as much risk as we think there is, if timelines are as short as we believe, and if there is a limited universe of people whose actions can move the needle on the net total risk we are all collectively facing, then it is in all of our interests to wish Holly continued success and greater influence in sounding the alarm and holding people's feet to the fire.
What matter the norms of discourse when the future is at stake?
Like Hitchens, Holly has decided to confront ideas and people that she believes to be dangerous and irresponsible, loudly and in public, with no regard for any supposed reverence and deference owed to the companies, their employees, their proponents or discourse norms. She is relentless in pointing out the contradictions of people in AI safety who she believes are failing to act in accordance with their private beliefs. She has chosen to disregard the norm that we should critique ideas rather than people. She has called out those who temper their public pronouncements because it is politically expedient or because they seek proximity to money, power, or influence. And she does so consistently, even at great cost to herself.
I am not an insider, and I don't have a measurement of her effect on any individual researcher's choices, but I think she has at least forced people to reflect morally. I'm sure people in this community have a much better ability to quantify her effect on insiders. What I can say is that some members of the public are aware of AI safety risk because she causes a ruckus, that the ground-up movement building and direct activism of PauseAI US is real work, and that, as with Hitchens, I would expect her effect to show up in the next cohort of advocates rather than in a lab researcher's resignation letter.
One of the most common objections I see on Twitter is that the people she attacks hardest are, in some cases, the people the movement most needs on the inside, and that contempt from outside makes it harder, not easier, for them to change course. But the insider doing safety work under those incentives should be able to withstand being asked, loudly and rudely, whether it is working. If the rude questions, if one considers such questions rude, is what drives them out, then how committed were these folks in the first place?
Movements benefit from having leaders across the continuum from polite to pugilistic. Most successful non-violent movements have had an extreme wing, and the research on radical flank effects mostly says what the history suggests: the flank makes the moderates look reasonable and raises support for them, at some cost to the flank's own popularity. Holly's rhetorical strategy cannot credibly be called violent, or said to advocate violence, by any reasonable understanding of that word. She appears extreme only if one believes we are in a safer timeline, or if one is an incrementalist.
Even if Holly is only effective because she makes those who communicate moderately and with circumspection more reasonable, then we should ensure that someone is funding and supporting Holly's work institutionally.
What was shocking about Hitchens then, and is shocking about Holly now to some people, is that both refused to give dangerous and irresponsible ideas and their proponents any reverence or deference. That is also what makes both of them effective as part of a movement to dethrone a set of entrenched ideas that have captured capital, institutions, power, and influence, and whose actors must be fought loudly and in broad daylight if we don't all want to die.
I am thankful that she is not an incrementalist and that she operates with a degree of urgency in which she sees little value in tempering her criticisms and refuses to adhere to "discourse norms," "politeness norms," "messaging effectiveness norms," and the rest.
And I hope a new class of AI safety advocates arises and crisscrosses the platforms, the country, and the world, engaging in rhetorical battles that confront the ideas, people, and organizations who are going to get us all killed.
She operates with more courage than I have. And I hope there are courageous funders out there who'll continue to support her work.