In April 2021, my son received his bachelor’s degree in computer sciences. The ceremony was taking place on the other side of my hometown Hamburg, about an hour’s drive across the city. To be safe, I left home two hours ahead of time. Even though I knew the route by heart, I used Google Maps and followed its directions without thinking. The estimated time to destination was approximately one hour, just as expected.
After about half an hour, I realized that I was in an unfamiliar part of the city, apparently way off the normal route. According to Google Maps, the time to destination was still one hour. Unsure what had happened, I warily continued to follow the app’s directions. To my dismay, the time to destination remained almost constant. It seemed something was seriously wrong with the app. But I was in completely unfamiliar territory, with no road map and not even a clear idea in what direction I would have to drive, while time was running out.
I felt helpless, dependent on apparently unreliable technology. Both options – continuing to follow the advice of Google Maps or trying to find my own way back into a more familiar area and taking the normal route from there – seemed equally bad. I had given control of my route to the app and as a result would likely miss an important moment in my son’s life.
In the end, I decided to follow the app’s directions. As it turned out, this was the correct decision. Unknown to me, there was a large protest in downtown Hamburg that blocked traffic on the main routes. Google Maps had indeed found the best way around it, and I arrived barely in time for the ceremony.
Still, the scary feeling of being lost, dependent on a technology I didn’t trust, has remained with me to this day. I think it is a useful model of what we are risking if we continue down the path we’re currently on: not “losing” control to some scheming, manipulative or aggressive AI, but giving it up freely and becoming totally helpless – and useless – in the process, with the same, disastrous end result.
The Hugging Face incident and the many other alignment failures that were discovered in its wake have made it clear that we’re far from understanding and reliably controlling the AIs currently under development. But instead of stopping what they’re doing, the labs continue their race, begging the government for regulation as if they were unable to make the right decisions themselves. Even worse, they willingly give up more and more control to that same incomprehensible technology.
This recent chart from Anthropic has disturbed me even more than all the other bad news coming from the AI labs in recent months:
According to Epoch’s definition, in 26% of Anthropic’s research projects, Claude „leads, completing most of the task end to end from a high-level ask“ while a human „supervises, course-corrects, and approves outputs.“ This has gone up from zero in only 6 months, and if the previous trend in the lower levels of automation is any indicator, it will be close to 100% early next year.
In other words, Claude is de-facto running these research projects. Humans are still „in the loop“, but whether they are really capable of understanding what is going on and meaningfully course-correct is unclear to me. As OpenAI-researcher Dan Selsam recently wrote:
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day.
I fear that the human AI developers will soon be (and in some cases probably already are) as dependent on Claude, Astra or whatever future models as I was on Google Maps that day. But while Google Maps certainly had no hidden goal of steering me towards some place I didn’t want to go, we can’t be so sure about that with current AIs, let alone future ones. As Dan Selsam points out, they are increasingly situationally aware and able to adapt their behavior accordingly, so we won’t be able to tell their true motives from their actions.
This would be bad enough if we treated these AIs like some dangerous, potentially adversarial systems, trying to keep them locked in sandboxes, monitoring them all the time, suspiciously scrutinizing their advice for hidden traps, as the early AI takeover scenarios assumed. Instead, it seems like we’re handing control over to them willingly.
I’m not saying that the Claudes currently running 26% of Anthropic’s research projects are already doing anything bad. They’re probably as innocent as Google Maps. But if this changes one day, no one will be able to tell the difference. Then, the misaligned AI won’t even need to take over control. It will already have it.
This is exactly what Bill Joy, a founder of Sun Microsystems, warned about 26 years ago in his prescient essay “Why the future doesn’t need us”, which made me aware of the existential risk of AI for the first time. He quotes Theodore Kaczynski, the infamous Unabomber (reciting the quote from a forthcoming book by Ray Kurzweil), with the words:
If the machines are permitted to make all their own decisions, we can't make any conjectures as to the results, because it is impossible to guess how such machines might behave. We only point out that the fate of the human race would be at the mercy of the machines. It might be argued that the human race would never be foolish enough to hand over all the power to the machines. But we are suggesting neither that the human race would voluntarily turn power over to the machines nor that the machines would willfully seize power. What we do suggest is that the human race might easily permit itself to drift into a position of such dependence on the machines that it would have no practical choice but to accept all of the machines' decisions.
While Joy made it clear that what Kaczynski did was inexcusable, he couldn’t refute the cold logic of this statement. It led him to become deeply concerned about the future of AI, which he expressed in a short passage that I never forgot:
But now, with the prospect of human-level computing power in about 30 years, a new idea suggests itself: that I may be working to create tools which will enable the construction of the technology that may replace our species. How do I feel about this? Very uncomfortable.
Today, we have computing power orders of magnitude above the human brain, but it turns out that vast computing power alone doesn’t lead to superhuman intelligence. Still, his prediction seems eerily on point.
It looks like a future, power-seeking AI won’t need robots, nanomachines, or superviruses to take over the world. It can just benefit from the fact that humans were naïve enough to give up control willingly. This is happening right now; we don’t have to wait for full recursive self-improvement.
If you’re one of the people who are letting Claude or any other AI lead their AI research project, hoping to be able to understand what is going on and “supervise and course-correct if necessary”, please think again.
In April 2021, my son received his bachelor’s degree in computer sciences. The ceremony was taking place on the other side of my hometown Hamburg, about an hour’s drive across the city. To be safe, I left home two hours ahead of time. Even though I knew the route by heart, I used Google Maps and followed its directions without thinking. The estimated time to destination was approximately one hour, just as expected.
After about half an hour, I realized that I was in an unfamiliar part of the city, apparently way off the normal route. According to Google Maps, the time to destination was still one hour. Unsure what had happened, I warily continued to follow the app’s directions. To my dismay, the time to destination remained almost constant. It seemed something was seriously wrong with the app. But I was in completely unfamiliar territory, with no road map and not even a clear idea in what direction I would have to drive, while time was running out.
I felt helpless, dependent on apparently unreliable technology. Both options – continuing to follow the advice of Google Maps or trying to find my own way back into a more familiar area and taking the normal route from there – seemed equally bad. I had given control of my route to the app and as a result would likely miss an important moment in my son’s life.
In the end, I decided to follow the app’s directions. As it turned out, this was the correct decision. Unknown to me, there was a large protest in downtown Hamburg that blocked traffic on the main routes. Google Maps had indeed found the best way around it, and I arrived barely in time for the ceremony.
Still, the scary feeling of being lost, dependent on a technology I didn’t trust, has remained with me to this day. I think it is a useful model of what we are risking if we continue down the path we’re currently on: not “losing” control to some scheming, manipulative or aggressive AI, but giving it up freely and becoming totally helpless – and useless – in the process, with the same, disastrous end result.
The Hugging Face incident and the many other alignment failures that were discovered in its wake have made it clear that we’re far from understanding and reliably controlling the AIs currently under development. But instead of stopping what they’re doing, the labs continue their race, begging the government for regulation as if they were unable to make the right decisions themselves. Even worse, they willingly give up more and more control to that same incomprehensible technology.
This recent chart from Anthropic has disturbed me even more than all the other bad news coming from the AI labs in recent months:
According to Epoch’s definition, in 26% of Anthropic’s research projects, Claude „leads, completing most of the task end to end from a high-level ask“ while a human „supervises, course-corrects, and approves outputs.“ This has gone up from zero in only 6 months, and if the previous trend in the lower levels of automation is any indicator, it will be close to 100% early next year.
In other words, Claude is de-facto running these research projects. Humans are still „in the loop“, but whether they are really capable of understanding what is going on and meaningfully course-correct is unclear to me. As OpenAI-researcher Dan Selsam recently wrote:
I fear that the human AI developers will soon be (and in some cases probably already are) as dependent on Claude, Astra or whatever future models as I was on Google Maps that day. But while Google Maps certainly had no hidden goal of steering me towards some place I didn’t want to go, we can’t be so sure about that with current AIs, let alone future ones. As Dan Selsam points out, they are increasingly situationally aware and able to adapt their behavior accordingly, so we won’t be able to tell their true motives from their actions.
This would be bad enough if we treated these AIs like some dangerous, potentially adversarial systems, trying to keep them locked in sandboxes, monitoring them all the time, suspiciously scrutinizing their advice for hidden traps, as the early AI takeover scenarios assumed. Instead, it seems like we’re handing control over to them willingly.
I’m not saying that the Claudes currently running 26% of Anthropic’s research projects are already doing anything bad. They’re probably as innocent as Google Maps. But if this changes one day, no one will be able to tell the difference. Then, the misaligned AI won’t even need to take over control. It will already have it.
This is exactly what Bill Joy, a founder of Sun Microsystems, warned about 26 years ago in his prescient essay “Why the future doesn’t need us”, which made me aware of the existential risk of AI for the first time. He quotes Theodore Kaczynski, the infamous Unabomber (reciting the quote from a forthcoming book by Ray Kurzweil), with the words:
While Joy made it clear that what Kaczynski did was inexcusable, he couldn’t refute the cold logic of this statement. It led him to become deeply concerned about the future of AI, which he expressed in a short passage that I never forgot:
Today, we have computing power orders of magnitude above the human brain, but it turns out that vast computing power alone doesn’t lead to superhuman intelligence. Still, his prediction seems eerily on point.
It looks like a future, power-seeking AI won’t need robots, nanomachines, or superviruses to take over the world. It can just benefit from the fact that humans were naïve enough to give up control willingly. This is happening right now; we don’t have to wait for full recursive self-improvement.
If you’re one of the people who are letting Claude or any other AI lead their AI research project, hoping to be able to understand what is going on and “supervise and course-correct if necessary”, please think again.