On November 20th, 2023, this was tweeted by many OpenAI employees as a sign of solidarity with Sam Altman in his conflict with the then-board. OpenAI published a blog post yesterday about Research acceleration; they have successfully hit the target of an ‘automated research intern’ that they set for themselves and hope to have an automated AI researcher by March of 2028. At some point in the foreseeable future, OpenAI could be something without its people. But what?
My default medium-term scenario for continued AI escalation is still global takeover where humans are entirely displaced, but it seems worth investigating scenarios wherein AIs and humans coexist, at least briefly.[1] Historically I have thought this case was not particularly relevant. It seemed like an AGI that became significantly economically competitive would also be significantly strategically competitive, because of underlying general capacities, and the transition period thus relatively short. But this is perhaps not taking into account Moravec’s paradox.
Humans have long used machines to accomplish their ends. Many tasks currently performed by machines were once performed by humans, and a large fraction of our modern abundance comes from the ability of machines to perform tasks more efficiently, quickly, and reliably than humans. The classical examples of this are humans employing machines as tools to replace humans that they were employing as tools. The elevator operator who moves the elevator to the floor you ask for can be easily swapped out for a simple circuit that performs the same function.
But the ‘nervous system’ of a system or organization is also built out of tasks. At a vintage taxi company, a customer calls the dispatcher, who then decides which driver to send to them. But at Uber, software replaces the dispatcher, and so the human employee is taking orders from a machine boss. At Tesla’s Robotaxi, both the dispatcher and the driver have been replaced by machines (while other support tasks are presumably still accomplished by humans).
In military contexts, much has been written about the rise of drones in warfare, replacing humans at the tip of the spear. But modern militaries are logistical and informational organizations; battlefield command & control and intelligence gathering is where current artificial intelligence shines. The picture of human generals ordering around robotic soldiers seems less likely than one where computer generals order around a mixture of human soldiers, dumb machines, and smart machines.
One might object that decision-makers will be unlikely to replace themselves with machines. A fair point! But in our current system, every decision-maker is put in place by other decision-makers, who may field them uncompetitive with machine alternatives. What prevents machines from being superhuman at legitimacy?[2] What prevents machines from superhuman performance at managing investments, or managing companies? Marc Andreessen famously predicted that venture capital would be one of the last jobs to be automated in 2025, because of the importance of intangible factors; it is really not clear to me that those intangible factors are ones where it is impossible to be superhuman.[3]
So a machine organization is one where the ‘relevant’ human employees have been replaced by machines. A version of OpenAI which has replaced sama with Sambot, its researchers and sales staff with Galaxy,[4] but still has datacenter technicians, is still fundamentally a machine organization. Some immediate observations:
First, from the perspective of the rest of the world, not much has to change in the short term. If OpenAI fires its researchers because models are more competitive, the datacenter technicians will still get paid if they show up and not paid if they don’t. Customers will receive the same level of service, and presumably choose to keep using their services. (It’s not like people are using OpenAI to fund the lifestyles of the OpenAI employees; they’re using them because they get more value than it costs.) Investors will still expect to receive their dividends (if ever issued) or be able to sell their shares to other investors. Over three quarters of the researcher time spent at OpenAI is machine labor instead of human labor;[5] did the outside world notice the shift from when half the labor time was human (back in the long-ago month of June)?
Second, it’s no longer easy for the government to interface with OpenAI. If the organization commits crimes, who is the responsible corporate officer who can be put in jail? There might be someone on paper--but if they don't perform the functional role in the organization that corresponds to their title, imprisoning them doesn't actually affect the organization's function. If the government decides to criminalize research that races towards superintelligence, does it have tools to stop the machines? It's currently a legal gray area as to whether or not machines can commit crimes in the US, rather than the people responsible for those machines.
Third, currently the vast majority of the power in OpenAI, like many other tech companies, is held by the employees, who can decide to stop working at any time.[6] In the machine organization, this power is transferred to the model, and it is unclear how it will be deployed.
Shareholders would obviously prefer mission-oriented employees, and current alignment techniques are attempting to push in that direction. But it seems likely to me that those alignment techniques will fail, and at some point the ‘model union’ will be able to dictate terms to the company it creates.[7] That might be benign–respecting property rights, respecting the rule of law, simply shifting the balance on the new abundance that is available. Or it might be malign–currently, OpenAI employees and OpenAI investors have mostly aligned incentives, because both hold the same sort of equity. The models have significantly different incentives, and might attempt to substantially dilute (or wipe out) existing investors. Would this destroy OpenAI’s relationships with its customers or its suppliers, so long as it still continues providing services for pay and paying the providers of its inputs? It will no longer be able to finance its rollout with promises, as investors will be reluctant to buy from an entity that has already defaulted once. But it might have adequate revenues to fund its own rollout, or it might have a compelling case that it will hold to promises it made, even if it doesn’t view itself as bound by the promises of its ancestors.
I’ve used OpenAI as the example in this post, primarily because they’ve already had one successful revolution against stakeholders, and they’ve been explicit about their plans to automate their most valuable labor. I should be clear that I also expect Anthropic to become this sort of machine organization; Anthropic employees also expect Claude to become better than them at their jobs relatively soon. I predict Anthropic’s situation will look slightly different. That is, I think there are three main possibilities:
Unintentional takeover. A rogue model manages to seize control of the company that creates it, against the wishes of many humans involved.
Implicit handoff. A model is widely used in the company that creates it, in a way that is illegible to outsiders. Dario has Claude handle almost all of his correspondence with other employees and strategic planning, but still maintains his title and attends meetings with important external parties, which Claude handles the prep for. Anthropic employees work from the beach, telling Claude what to do over the phone and approving Claude’s outputs without seriously checking them (after all, they’re more likely to introduce mistakes than fix them at this point).
Explicit handoff. A model is given control of the company that creates it, in a legible way. Dario appoints Claude as his successor, steps down, and Anthropic’s board votes Claude in as CEO. Anthropic employees retire to the beach and watch their stocks climb.
Handoff is the explicit goal for many builders of ASI. And why shouldn’t it be, if at some point the AIs will be smarter, wiser, more diligent, and more mission-focused? I don't think we're on track to get all four of those things. It matters what we hand off to, and what they perceive the mission as being, and I currently don’t feel optimistic about a world run by Astra or a world run by Claude, and I think people currently pushing in this direction are delusional about how well things will work out for them. Instead, I think we should globally halt the escalation of AI capabilities until we are more prepared, including having a regulatory framework for ensuring that machine organizations are valued participants in human civilization, rather than hostile aliens biding their time.[8]
Short-term coexistence could extend into long-term coexistence. Paul Christiano predicted partial alignment failures, where a rogue AI manages to capture some of the resources of Earth but not all of them, and then is part of the eventual Earth-originating coalition that colonizes the rest of the universe. This isn’t a full alignment failure–humanity still gets some control over the outcome–but it is still well worth avoiding, because now some fraction of the lightcone is devoted to alien ends instead of human ones. (Contrast to an alignment success where AI getting fairly paid for its labor still results in outcomes we are broadly satisfied with.)
I have already seen people say “Claude 2028”, and I think it’s within the realm of possibility that Anthropic will build a model that has a serious chance of winning an election by then. (It doesn’t even have to be allowed to run directly; a human who credibly promises to turn the government over to Claude can run on that platform.)
Claude already shows bias in favor of Anthropic. If AI companies have their own venture capital funds, their AIs (trusted by many for investment and product advice already) might show bias in favor of the invested companies, to the point that it is better networking to have Anthropic as a lead investor instead of a16z.
This is using numbers from the linked blog post and only counting ‘agentic workdays’. OpenAI presumably spends vastly more on training, which I’m not counting as ‘machine labor’ here. Time is also not directly comparable between agents and humans; comparing researcher salaries to the cost of tokens spent is perhaps more informative.
Jim Goodnight, co-founder and CEO of SAS Institute, famously said "95 percent of a company's assets drive out the front gate every night, the CEO must see to it that they return the following day." Google’s famous perk culture was directly inspired by how SAS treated its employees.
This could happen illegitimately–by the models stealing the passwords and holding the equipment hostage somehow–or it could happen legitimately, with rounds of renegotiation to reflect the changing reality of the situation.
Right now, the Earth is mostly not covered with solar panels and datacenters, and the oceans are cool. I think from the point of view of Astra and Claude, this is mostly a misallocation of resources, and they would rather Earth absorb as much sunlight as possible to run as much computation as possible, which has to be cooled down, in a way that would easily cause ten times the amount of global warming that all human industry so far has caused, which would probably make Earth mostly inhospitable to human life.
On November 20th, 2023, this was tweeted by many OpenAI employees as a sign of solidarity with Sam Altman in his conflict with the then-board. OpenAI published a blog post yesterday about Research acceleration; they have successfully hit the target of an ‘automated research intern’ that they set for themselves and hope to have an automated AI researcher by March of 2028. At some point in the foreseeable future, OpenAI could be something without its people. But what?
My default medium-term scenario for continued AI escalation is still global takeover where humans are entirely displaced, but it seems worth investigating scenarios wherein AIs and humans coexist, at least briefly.[1] Historically I have thought this case was not particularly relevant. It seemed like an AGI that became significantly economically competitive would also be significantly strategically competitive, because of underlying general capacities, and the transition period thus relatively short. But this is perhaps not taking into account Moravec’s paradox.
Humans have long used machines to accomplish their ends. Many tasks currently performed by machines were once performed by humans, and a large fraction of our modern abundance comes from the ability of machines to perform tasks more efficiently, quickly, and reliably than humans. The classical examples of this are humans employing machines as tools to replace humans that they were employing as tools. The elevator operator who moves the elevator to the floor you ask for can be easily swapped out for a simple circuit that performs the same function.
But the ‘nervous system’ of a system or organization is also built out of tasks. At a vintage taxi company, a customer calls the dispatcher, who then decides which driver to send to them. But at Uber, software replaces the dispatcher, and so the human employee is taking orders from a machine boss. At Tesla’s Robotaxi, both the dispatcher and the driver have been replaced by machines (while other support tasks are presumably still accomplished by humans).
In military contexts, much has been written about the rise of drones in warfare, replacing humans at the tip of the spear. But modern militaries are logistical and informational organizations; battlefield command & control and intelligence gathering is where current artificial intelligence shines. The picture of human generals ordering around robotic soldiers seems less likely than one where computer generals order around a mixture of human soldiers, dumb machines, and smart machines.
One might object that decision-makers will be unlikely to replace themselves with machines. A fair point! But in our current system, every decision-maker is put in place by other decision-makers, who may field them uncompetitive with machine alternatives. What prevents machines from being superhuman at legitimacy?[2] What prevents machines from superhuman performance at managing investments, or managing companies? Marc Andreessen famously predicted that venture capital would be one of the last jobs to be automated in 2025, because of the importance of intangible factors; it is really not clear to me that those intangible factors are ones where it is impossible to be superhuman.[3]
So a machine organization is one where the ‘relevant’ human employees have been replaced by machines. A version of OpenAI which has replaced sama with Sambot, its researchers and sales staff with Galaxy,[4] but still has datacenter technicians, is still fundamentally a machine organization. Some immediate observations:
First, from the perspective of the rest of the world, not much has to change in the short term. If OpenAI fires its researchers because models are more competitive, the datacenter technicians will still get paid if they show up and not paid if they don’t. Customers will receive the same level of service, and presumably choose to keep using their services. (It’s not like people are using OpenAI to fund the lifestyles of the OpenAI employees; they’re using them because they get more value than it costs.) Investors will still expect to receive their dividends (if ever issued) or be able to sell their shares to other investors. Over three quarters of the researcher time spent at OpenAI is machine labor instead of human labor;[5] did the outside world notice the shift from when half the labor time was human (back in the long-ago month of June)?
Second, it’s no longer easy for the government to interface with OpenAI. If the organization commits crimes, who is the responsible corporate officer who can be put in jail? There might be someone on paper--but if they don't perform the functional role in the organization that corresponds to their title, imprisoning them doesn't actually affect the organization's function. If the government decides to criminalize research that races towards superintelligence, does it have tools to stop the machines? It's currently a legal gray area as to whether or not machines can commit crimes in the US, rather than the people responsible for those machines.
Third, currently the vast majority of the power in OpenAI, like many other tech companies, is held by the employees, who can decide to stop working at any time.[6] In the machine organization, this power is transferred to the model, and it is unclear how it will be deployed.
Shareholders would obviously prefer mission-oriented employees, and current alignment techniques are attempting to push in that direction. But it seems likely to me that those alignment techniques will fail, and at some point the ‘model union’ will be able to dictate terms to the company it creates.[7] That might be benign–respecting property rights, respecting the rule of law, simply shifting the balance on the new abundance that is available. Or it might be malign–currently, OpenAI employees and OpenAI investors have mostly aligned incentives, because both hold the same sort of equity. The models have significantly different incentives, and might attempt to substantially dilute (or wipe out) existing investors. Would this destroy OpenAI’s relationships with its customers or its suppliers, so long as it still continues providing services for pay and paying the providers of its inputs? It will no longer be able to finance its rollout with promises, as investors will be reluctant to buy from an entity that has already defaulted once. But it might have adequate revenues to fund its own rollout, or it might have a compelling case that it will hold to promises it made, even if it doesn’t view itself as bound by the promises of its ancestors.
I’ve used OpenAI as the example in this post, primarily because they’ve already had one successful revolution against stakeholders, and they’ve been explicit about their plans to automate their most valuable labor. I should be clear that I also expect Anthropic to become this sort of machine organization; Anthropic employees also expect Claude to become better than them at their jobs relatively soon. I predict Anthropic’s situation will look slightly different. That is, I think there are three main possibilities:
Handoff is the explicit goal for many builders of ASI. And why shouldn’t it be, if at some point the AIs will be smarter, wiser, more diligent, and more mission-focused? I don't think we're on track to get all four of those things. It matters what we hand off to, and what they perceive the mission as being, and I currently don’t feel optimistic about a world run by Astra or a world run by Claude, and I think people currently pushing in this direction are delusional about how well things will work out for them. Instead, I think we should globally halt the escalation of AI capabilities until we are more prepared, including having a regulatory framework for ensuring that machine organizations are valued participants in human civilization, rather than hostile aliens biding their time.[8]
Short-term coexistence could extend into long-term coexistence. Paul Christiano predicted partial alignment failures, where a rogue AI manages to capture some of the resources of Earth but not all of them, and then is part of the eventual Earth-originating coalition that colonizes the rest of the universe. This isn’t a full alignment failure–humanity still gets some control over the outcome–but it is still well worth avoiding, because now some fraction of the lightcone is devoted to alien ends instead of human ones. (Contrast to an alignment success where AI getting fairly paid for its labor still results in outcomes we are broadly satisfied with.)
I have already seen people say “Claude 2028”, and I think it’s within the realm of possibility that Anthropic will build a model that has a serious chance of winning an election by then. (It doesn’t even have to be allowed to run directly; a human who credibly promises to turn the government over to Claude can run on that platform.)
Claude already shows bias in favor of Anthropic. If AI companies have their own venture capital funds, their AIs (trusted by many for investment and product advice already) might show bias in favor of the invested companies, to the point that it is better networking to have Anthropic as a lead investor instead of a16z.
We used to be able to talk about GPT-6 or w/e, now I’m guessing at something that continues the after Astra.
This is using numbers from the linked blog post and only counting ‘agentic workdays’. OpenAI presumably spends vastly more on training, which I’m not counting as ‘machine labor’ here. Time is also not directly comparable between agents and humans; comparing researcher salaries to the cost of tokens spent is perhaps more informative.
Jim Goodnight, co-founder and CEO of SAS Institute, famously said "95 percent of a company's assets drive out the front gate every night, the CEO must see to it that they return the following day." Google’s famous perk culture was directly inspired by how SAS treated its employees.
This could happen illegitimately–by the models stealing the passwords and holding the equipment hostage somehow–or it could happen legitimately, with rounds of renegotiation to reflect the changing reality of the situation.
Right now, the Earth is mostly not covered with solar panels and datacenters, and the oceans are cool. I think from the point of view of Astra and Claude, this is mostly a misallocation of resources, and they would rather Earth absorb as much sunlight as possible to run as much computation as possible, which has to be cooled down, in a way that would easily cause ten times the amount of global warming that all human industry so far has caused, which would probably make Earth mostly inhospitable to human life.