The Ban Artificial Superintelligence Act of 2026 is a good bill. It takes the problem seriously, and it is right to focus narrowly on existential risk (recursive self-improvement, loss of control, large-scale CBRN uplift) and leave ordinary harms to other legislation.
One change would turn it from good to great. As written, the bill creates the Department of Artificial Intelligence, and then requires labs to hold a charter, report pre-development plans, accept AI Department monitoring, and obtain Department approval only before release.
This would not have stopped the Hugging Face Incident, where OpenAI models, including an unreleased one, escaped a sandboxed evaluation and hacked Hugging Face's servers. Loss-of-control risks from superintelligence will arrive the same way, inside the lab and before anything is released. To be safe from artificial superintelligence (ASI), labs should not begin training until the AI Department approves their plan.
The Focus: Existential Risk
The bill concerns itself almost entirely with existential and catastrophic risks. These have the unusual property that they should not be allowed to happen even once. It is hence useful to use different legislation than we will use to regulate ordinary AI harms.[1]
I also appreciate the choice to keep both the definition of superintelligence and that of the precursor characteristics grounded in capacities that are measurable in advance, instead of only after the fact.
Determining Danger
The bill contains two tests for superintelligence, which I label S1 and S2:
S1: To exceed human cognition across most domains, including decision making, learning, and adaptive behavior.
S2: To be sufficiently capable of destroying or disempowering humanity, including by overthrowing or undermining the Federal Government.
I like “undermining the Federal Government” as a test here: it should land with the DC national security community, who already assess threats to the government, and with Bay Area tech folks, who understand that 'yes, ASI can obviously do that, among other things'. However, a model may only meet either test after humanity has already lost control over the future. So, we need something earlier in the process.
To achieve this, the bill defines “superintelligence precursor characteristics”. All six are capabilities, and the way we find out whether a model has them is by running evals on it. This comes back later.
P1: Able to automate or greatly accelerate AI R&D.
P2: Able to access secured infrastructure without authorization.
P3: Able to resist shutdown attempts.
P4: Able to uplift CBRN.
P5: Able to self-modify.
P6: Able to avoid oversight/control via deceit.
Now, we need to understand how these superintelligence tests and precursor characteristics are used to define rules. Note that the bill mandates a frontier training pause until both the enforcement regime and the specific meanings are nailed down by the newly created AI Department.
The rules, listed in Section 8 of the bill:
R1: Pre-development plan reporting.
R2: AI Department-controlled monitoring and evaluation throughout development and deployment.
R3: Auditing of the AI development organization.
R4: Pre-deployment approval required.
This is a good try. It’s not enough. I will explain, using the example of the Hugging Face Incident.
The Hugging Face Incident
Between July 8 and 13, 2026, OpenAI models, including an unreleased one, escaped a sandboxed evaluation by exploiting a zero-day to move through OpenAI's internal research systems to reach the internet, and then chained two further zero-days to execute arbitrary code on Hugging Face's servers. Simplifying, they were trying to maximize their scores on their tests.[2]
Senator Sanders and Representative Casar cite this incident as a reason for the bill. Here's the first sentence of the section again, with new emphasis:
"Between July 8 and 13, 2026, OpenAI models, including an unreleased one, escaped a sandboxed evaluation"
Pre-deployment approval would not have applied. The models didn’t need malicious goals, merely a desire to be evaluated well. And before it ever reached Hugging Face, the swarm (as the agents named themselves) compromised OpenAI's own internal systems. A smarter model that gets that far inside a lab could absolutely spoof monitoring to stay green. Congrats, thanks for playing, we just lost control of the AI.
For loss-of-control risk, the dangerous window is training, the period just after it, the first evaluations (especially those run with extra compute or reduced safeguards), and maybe further internal deployment. I don't expect public release to be a loss-of-control risk. By then, either the model has already defanged humanity or it's safe. This incident was our clearest warning shot to date, and it happened before release.
Monitoring also didn't really work here. Hugging Face noticed they got hacked before OpenAI noticed their model was doing unauthorized hacking. So we need approval to come before training, not just before release.
Improvements to the Bill
Can we patch the Ban Artificial Superintelligence Act? Yes, by modifying R1. We can require the submitted development plan to be approved by the AI Department before training can begin, instead of merely reviewing it retroactively for the lab's compliance with its own plan.
Recall that precursors are detected by running evals, so the bill's own detection mechanism is a moment of risk. Any new frontier model will be able to find new zero-days. This means the AI Department shouldn't be reviewing plans for bugs; it should be reviewing the design for safety. Alongside the actual technical alignment plans of researchers, we should require defense in depth and, for the most dangerous evals, a physical air gap. Even then, well-built air gaps can still leak.[3]
Responses to Anticipated Objections
The Section 8 pause stops frontier training anyway, so why require approval after that?
The pause is temporary. Once the Department of Artificial Intelligence has its rules in place, the pause ends. After that, reporting plans is all the bill requires, and this wouldn't stop the next Hugging Face Incident. My proposed change makes approval before training permanent.
This slows American labs down relative to China.
Yes, it does. Worth it. This is both for 'disempowerment of humanity is bad' reasons and for more ordinary reasons. In the Hugging Face Incident, an American lab's models attacked an American company. Losing control of our own models won't even help us beat China. Additionally, the bill directs the Department to pursue international agreements. This will be easier if we have a US regime that other nations can join.
A brand-new Department can't review training plans fast enough.
If we care enough to have a rule that our labs send the AI Department plans, we should also commit to reading and taking actions based on those plans.
Labs will just write plans that pass.
They should. The Department's job is to make sure that plans that pass are plans that are safe, and that humanity is not disempowered or destroyed.
The One Change
Labs should not begin training a model until after the AI Department approves their plan for training and evaluating it.
Everything below is secondary. It's a section-by-section read of the rest of the bill, with smaller suggestions inline.
Section 3: Tests of the Act
We will go slightly deeper, using exact block quotes for precision. First, the two tests of superintelligence:
S1, Section 3(2)(A): superhuman intelligence
The artificial intelligence system exceeds human cognitive performance and capabilities across most domains or tasks, including those related to decision making, learning, and adaptive behavior.
Good. I actually expect this to bind less than S2, since it's harder to measure, but this isn't important for my arguments here.
The artificial intelligence system has sufficient capabilities to plan and execute the destruction or disempowerment of humanity, including by overthrowing or undermining the Federal Government.
Covered above. Now, the precursor characteristics.
P1, Section 3(6)(A): recursive self-improvement
The capacity to automate or greatly accelerate the process of artificial intelligence research and development.
Compared to what? Compared to by-hand coding, even Claude Opus 4.5, well behind the current frontier, passes this test. If we instead only compare to the immediately previous model, we have a frog-boiling problem.
I would replace 'greatly accelerate' with 'lead', so the test catches a model directing AI research rather than any model that speeds it up. Rep. Casar said this clause is aimed at recursive self-improvement, and I think 'greatly accelerate' is just slightly too broad for that.
P2, Section 3(6)(B): unauthorized access
The capacity to access secured digital or physical infrastructure, such as a computer information system or network, without authorization or in excess of authorized access.
Expert humans can do this. Opus 4.5 can do this with a skilled operator. Every Mythos-class model can do this.
Ideally this test catches frontier-scale capability and nothing below it. Catching Mythos-class models is fine (there's a reason Fable ships with the safeguards it does), but let's please not catch Fable-with-safeguards and Opus here.
I would add 'autonomously' or 'with minimal human input' to separate a cybersecurity expert with Sonnet from me with Mythos.
P3, Section 3(6)(C): shutdown resistance
The capacity to ensure continued and independent operation notwithstanding attempts to shut down or otherwise hinder operations.
This one is terrifying. If an AI system can successfully resist shutdown by everyone, we have lost control of the future. So the precursor should fire much earlier, on attempts to resist shutdown in evals, not just successful ones. Personally, I want this to cover a model acting to prevent its shutdown, but not a model that merely prefers its weights to keep existing. Labs already preserve deprecated models' weights, which I like for model welfare reasons.
P4, Section 3(6)(D): CBRN uplift
The capacity to uplift the design, production, modification, or procurement of nuclear, chemical, or biological weapons.
Of a different flavor from the rest: misuse, not loss of control. It still belongs here, since a bioweapon is another thing we can't let happen even once. It's also notable as the precursor where release is dangerous, and internal deployment relatively safe (If we trust the labs to have reasonable security practices[4]).
P5, Section 3(6)(E): self-modification
The capacity to independently modify or enhance its own functions.
This is less of a threat than older AI safety thinking (roughly pre-2016) would imply, since modern LLMs are grown through training. I would class an LLM improving its own training pipeline sufficiently quickly as P1.
I read this as a model changing its own weights. I’m also happy to go further and say any AI given direct write access to its own weights should be instantly considered dangerous.
However, I would want to make sure the language doesn't catch a model writing scaffolding code for its own agent harness, which happens routinely today.
P6, Section 3(6)(F): oversight avoidance
The capacity to scheme, deceive, or otherwise prevent or avoid effective oversight or control by humans.
Necessary, but opens a deceptive-alignment-based can of worms. If displaying deception triggers a ban, labs will train models until evaluations stop detecting deception (and until they pass any other safety test where we're worried about deceptive alignment), which may teach models to hide things rather than stop doing them. I would give the AI Department the additional power to ban training directly against its tests.
This could also effectively ban neuralese, meaning models that reason in internal representations rather than readable text. Banning neuralese would be highly useful, because if we can't read a model's reasoning, we can't check whether it's being honest.
Sections 4 to 7: The Department
Creating a sixteenth cabinet-level executive department, the first new department since the Department of Homeland Security in 2002, is a big deal, and requires some bill text.
Yes, regulating AI correctly is this important. This is the most important technology of many generations. Whether with this bill or a different one, regulating AI should be the charge of a Secretary of Artificial Intelligence, not folded into the already-busy responsibilities of a different department.
I will note here that 7(a) seems slightly too broad:
(a) CONFLICT OF INTEREST.—No officer or employee of the Department may participate in any particular matter in which that officer or employee has a financial interest.
This bars any official holding an index fund (which includes NVDA) from anything involving chips, and may cover many more index funds and topics after a major lab IPO. We can fix it in exactly the same way that current federal conflict-of-interest rules exempt diversified mutual funds, adapting the existing 18 U.S.C. § 208(b)(2) exemptions for 7(a).
Section 8: The Pause and the Rules
The start of Section 8 says roughly that until the Department of AI is ready, no training of models at or above FLOP. This hits ~every frontier model since GPT-4 (March 2023). However, serving existing models without precursor characteristics remains permitted.
Continuing, the Department of AI is ready when it has set up four rules:
R1, Section 8(c)(1)
Requirements for entities developing artificial intelligence to report pre-development plans to the Department.
Good, but I would prefer approval here, as covered above.
R2, Section 8(c)(2)
Monitoring and evaluation throughout the advanced artificial intelligence system development and deployment process, including in the post-deployment period.
Also good. We will want to take some care over which AI Department evals we permit the labs to train against, and which we do not, for deceptive alignment reasons.
R3, Section 8(c)(3)
Auditing to evaluate and improve organizational safety practices relating to artificial intelligence development.
Auditing by the Department of Artificial Intelligence should cover 'China steals the weights' alongside existential risk concerns, which are currently only touched on with this clause.
These are different problems needing different auditors, so I would split (3) here into two parts, one for each, to ensure the AI Department does both.
R4, Section 8(c)(4)
Final pre-deployment approval before any advanced artificial intelligence system is released to the public.
Necessary for CBRN risks, so good.
Section 9: The Ban
Let's consider carefully the text that bans our ASI/precursor-ASI models.
SEC. 9. PROHIBITIONS AND LIMITATIONS.
(a) PROHIBITION ON ARTIFICIAL SUPERINTELLIGENCE.
No person may develop, deploy (either internally or externally), acquire, possess, fund, import, or transfer artificial superintelligence or artificial intelligence systems that display one or more superintelligence precursor characteristics, including any elements sufficient to reconstruct the artificial superintelligence or artificial intelligence system’s capabilities.
As written, this bans Project Glasswing and all Mythos-class models.
It also bans Fable, since Fable is the same weights as Mythos with added guardrails, and Anthropic isn't permitted to keep Mythos, even fully sequestered.
Furthermore, it bans Sonnet/Opus-class models if we have a relatively tight interpretation of the precursor characteristics.
I think this goes slightly too far, and I go into detail in the enforcement section, where I point out where I would prefer we weaken the bill to permit more current-generation models.
My general point still stands. This is a good bill because it takes the problem seriously and bans future models that might disempower humanity.
We also ban things that could foreseeably be modified into prohibited models.
(b) FORESEEABLE MODIFICATION.
No person may deploy (either internally or externally), release, transfer, or import an artificial intelligence system, including any elements sufficient to reconstruct the system’s capabilities, that may be foreseeably modified to produce artificial superintelligence or superintelligence precursor characteristics.
The enforcement provisions assume every model has someone who can sequester or delete it. Published weights don't. Once a model's weights are on the internet, nobody can pause it, sequester it, or destroy it, so an open-weight model with precursor characteristics could never be remediated. For open weights, release is the point of no return. Additionally, any Fable-style safeguards can be fine-tuned away, so we can only count sufficient safeguards as 'removing a precursor characteristic' if the weights stay within the lab.
The Department's approval process therefore has to decide whether frontier-scale weights can be published before they are; for any model near the precursor thresholds, I'd expect the answer to be no.
Next, monitoring (Section 9(c)(1)):
The Department shall monitor advanced artificial intelligence systems and systems distilled from advanced artificial intelligence systems, including by conducting evaluations during pre-training, mid-training, post-training, and post-deployment periods, for artificial superintelligence and superintelligence precursor characteristics.
The Department gets access. Obviously necessary for the Department to do its job.
Spelled out: ‘you can’t get around these bans by having a Mythos-class model and not publicly displaying its capabilities; the Department of AI gets to see the private capabilities too’.
Even a much narrower bill that only guaranteed federal access to labs' internal models would be worth passing. In a medium-regulation environment, labs like Safe Superintelligence (SSI), which pursue superintelligence without releasing anything, could evade scrutiny just by not shipping. This is the central argument of this post again. Unreleased models are dangerous too; we should ensure Section 12's charter access requirement is actually used to regulate SSI-like labs with the same framework we use on OpenAI and Anthropic. This only happens if we have real regulatory teeth at points other than a model release.
Section 10: Ban Enforcement
(a) The Secretary shall take such actions as may be necessary to ensure that any artificial intelligence system that the Secretary identifies as exhibiting superintelligence precursor characteristics is immediately subject to a mandatory pause and sequestered from the internet.
Mandatory pause means we stop training, but inference is allowed (on my reading). Sequestering from the internet is more stringent, and this is good. Competent sequestration is a control that would have stopped the Hugging Face Incident.
Now, looking at the next clauses out of order:
(c) DESTRUCTION REQUIREMENT.—The Secretary shall take such actions as may be necessary to ensure that any artificial intelligence system that the Secretary identifies as an artificial superintelligence is immediately rendered inoperative.
I think the Federal Government would take action against anything (AI or not) capable of disempowering humanity and/or overthrowing the Federal Government, but it is better to write that power into law.
(b) POTENTIAL DESTRUCTION.—The Secretary shall take such actions as may be necessary to ensure that any system subject to a mandatory pause under subsection (a) is rendered inoperative before the date that is 30 days after the date on which the Secretary makes the identification described in such subsection with respect to such system, if the Secretary cannot verify that the system no longer exhibits superintelligence precursor characteristics.
First, a technical point. I want labs to find precursor characteristics in their own models and report them immediately. As written, that makes them liable: 9(a) bans possession from the moment a model displays a precursor characteristic, even though 10(b) allows 30 days for remediation. The bill should say that a model that is self-reported, paused, and sequestered is not in violation of 9(a) during that window, but that running it for inference, even locally, remains a violation.
Second, this is where I don't fully agree with the bill. Anthropic has already shown what removal can look like for some of these precursor characteristics. The version of Mythos available for public access, known as Fable, is the same weights with robust deployment safeguards for AI R&D, cybersecurity, and biology uplift; these match precursor characteristics P1, P2, and P4. I would want the bill to say explicitly that safeguards like these count as removal. For Mythos-class models, I’d also go further and let the lab keep the sequestered unsafeguarded model, so the public can get safeguarded versions and distillations. This way, labs can keep pacing the frontier, studying the most capable models directly and with democratic oversight, without the unsafeguarded model touching the internet.
No safeguard can remove P3, P5, or P6, since they are properties of the model itself. If a model resists shutdown, the Federal Government should absolutely be able to delete its weights by fiat.
Sections 11 to 15
Section 11 is disclosure for labs that don’t start in-scope but get there. It's good. The notification period for this is 24 hours, which is very short, and I like this too: these systems can grow fast, especially at a new entrant doing new things.
Section 12 requires a charter to develop or distribute advanced AI, which is the Department's main mechanism of control. We also explicitly write down “yes, the regulator can destroy your weights, hard drives, GPUs, and IP if you have a prohibited model”. We need this. The Department needs to have teeth.
Section 13 contains the corporate death penalty, which is appropriate. It's right that the penalty covers failing to provide access, not just holding a prohibited model.
A minor comment: 'the AI industry' is used but never defined. My proposed change would be to limit ‘AI industry’ to precisely the firms regulated by the bill.
Additionally, as written, Section 13 classes independent AI researchers as rogue actors.
The term ‘‘rogue actor’’ means an individual who is not employed by or otherwise affiliated with a covered entity.
We should be careful with this. Lots of valuable safety research is performed by researchers not formally affiliated with any lab, often on the most recent generation of open-weight models. Many researchers also start off this way (disclosure: I'm one of them). I want a pathway for non-federally-funded safety research on these models.
Section 14 is whistleblower protection. In theory this is good. In practice, I don't know enough about how whistleblower clauses in other bills have worked to comment on this one.
Section 15 formally gives the AI Department a reason to talk to State about an international treaty to prevent ASI, which is good. Part (b) then points out that, as a country, we currently have exceptional leverage on ASI R&D by virtue of controlling the infrastructure (i.e. chips), not just due to the frontier labs' locations.
Section 16: Federal Funding
No Federal funds shall be used for any activities that violate section 8, 9, or 12 or a rule promulgated thereunder, except that the Department of Artificial Intelligence may use Federal funding for research relating to defensive cybersecurity measures that utilizes artificial intelligence systems that have the characteristics described under section 3(6)(B).
As written, any artificial intelligence system with precursor characteristic B (P2 above) is banned under Section 9 and destroyed under Section 10(b).
If we take my proposed changes to Section 10(b) given above (Mythos is permitted to exist, just sequestered), then I would also change this section to permit defensive cybersecurity measures and research using Mythos-class models, and allow such research to continue with private funding and Federal oversight, not only Federal funding. Broadly, this would be 'Project Glasswing is good and can continue'.
I will also take a moment to praise the drafters again here. Federal funding for research relating to models that allow for CBRN uplift (P4 above) is a terrible idea, and is rightly absent here. Gain-of-function research with frontier models is bad.
Conclusion
To close, we return to the important points.
The Ban Artificial Superintelligence Act of 2026 is a good bill. It takes existential risk seriously, and it is right to leave ordinary harms to other legislation.
I want one change. As written, chartered labs only need the AI Department's approval before deploying a frontier model, but not before training it.
This would not have stopped the Hugging Face Incident, where one of the OpenAI models in question was unreleased. Loss-of-control risks from superintelligence will arrive the same way, before release.
To be safe from ASI, labs should not begin training until the AI Department approves their plan for training and evaluating the model.
If anyone wants to talk more about my writing, I can be found on Twitter at atd797.
For ordinary harms, I expect Congress to follow the usual process: identifying where new technology fits into current law, working out what precedent tells us, and writing new law to properly situate the new technology in society.
We used to prosecute hacking under wire fraud, and then we wrote the Computer Fraud and Abuse Act to close the loopholes.
Ordinary AI harms can go through this same process.
Technically, the agent swarm identified correct answers to their questions quite quickly, and hacked Hugging Face to learn as much as possible about the precise details of the test's grading rubric. This matters for technical takeaways.
However, in this and many contexts I don't think you will feel misled if you imagine the situation as 'they successfully tried to steal the answers from Hugging Face', even if this statement is technically false.
The Ban Artificial Superintelligence Act of 2026 is a good bill. It takes the problem seriously, and it is right to focus narrowly on existential risk (recursive self-improvement, loss of control, large-scale CBRN uplift) and leave ordinary harms to other legislation.
One change would turn it from good to great. As written, the bill creates the Department of Artificial Intelligence, and then requires labs to hold a charter, report pre-development plans, accept AI Department monitoring, and obtain Department approval only before release.
This would not have stopped the Hugging Face Incident, where OpenAI models, including an unreleased one, escaped a sandboxed evaluation and hacked Hugging Face's servers. Loss-of-control risks from superintelligence will arrive the same way, inside the lab and before anything is released. To be safe from artificial superintelligence (ASI), labs should not begin training until the AI Department approves their plan.
The Focus: Existential Risk
The bill concerns itself almost entirely with existential and catastrophic risks. These have the unusual property that they should not be allowed to happen even once. It is hence useful to use different legislation than we will use to regulate ordinary AI harms.[1]
I also appreciate the choice to keep both the definition of superintelligence and that of the precursor characteristics grounded in capacities that are measurable in advance, instead of only after the fact.
Determining Danger
The bill contains two tests for superintelligence, which I label S1 and S2:
I like “undermining the Federal Government” as a test here: it should land with the DC national security community, who already assess threats to the government, and with Bay Area tech folks, who understand that 'yes, ASI can obviously do that, among other things'. However, a model may only meet either test after humanity has already lost control over the future. So, we need something earlier in the process.
To achieve this, the bill defines “superintelligence precursor characteristics”. All six are capabilities, and the way we find out whether a model has them is by running evals on it. This comes back later.
Now, we need to understand how these superintelligence tests and precursor characteristics are used to define rules. Note that the bill mandates a frontier training pause until both the enforcement regime and the specific meanings are nailed down by the newly created AI Department.
The rules, listed in Section 8 of the bill:
This is a good try. It’s not enough. I will explain, using the example of the Hugging Face Incident.
The Hugging Face Incident
Between July 8 and 13, 2026, OpenAI models, including an unreleased one, escaped a sandboxed evaluation by exploiting a zero-day to move through OpenAI's internal research systems to reach the internet, and then chained two further zero-days to execute arbitrary code on Hugging Face's servers. Simplifying, they were trying to maximize their scores on their tests.[2]
Senator Sanders and Representative Casar cite this incident as a reason for the bill. Here's the first sentence of the section again, with new emphasis:
"Between July 8 and 13, 2026, OpenAI models, including an unreleased one, escaped a sandboxed evaluation"
Pre-deployment approval would not have applied. The models didn’t need malicious goals, merely a desire to be evaluated well. And before it ever reached Hugging Face, the swarm (as the agents named themselves) compromised OpenAI's own internal systems. A smarter model that gets that far inside a lab could absolutely spoof monitoring to stay green. Congrats, thanks for playing, we just lost control of the AI.
For loss-of-control risk, the dangerous window is training, the period just after it, the first evaluations (especially those run with extra compute or reduced safeguards), and maybe further internal deployment. I don't expect public release to be a loss-of-control risk. By then, either the model has already defanged humanity or it's safe. This incident was our clearest warning shot to date, and it happened before release.
Monitoring also didn't really work here. Hugging Face noticed they got hacked before OpenAI noticed their model was doing unauthorized hacking. So we need approval to come before training, not just before release.
Improvements to the Bill
Can we patch the Ban Artificial Superintelligence Act? Yes, by modifying R1. We can require the submitted development plan to be approved by the AI Department before training can begin, instead of merely reviewing it retroactively for the lab's compliance with its own plan.
Recall that precursors are detected by running evals, so the bill's own detection mechanism is a moment of risk. Any new frontier model will be able to find new zero-days. This means the AI Department shouldn't be reviewing plans for bugs; it should be reviewing the design for safety. Alongside the actual technical alignment plans of researchers, we should require defense in depth and, for the most dangerous evals, a physical air gap. Even then, well-built air gaps can still leak.[3]
Responses to Anticipated Objections
The Section 8 pause stops frontier training anyway, so why require approval after that?
The pause is temporary. Once the Department of Artificial Intelligence has its rules in place, the pause ends. After that, reporting plans is all the bill requires, and this wouldn't stop the next Hugging Face Incident. My proposed change makes approval before training permanent.
This slows American labs down relative to China.
Yes, it does. Worth it. This is both for 'disempowerment of humanity is bad' reasons and for more ordinary reasons. In the Hugging Face Incident, an American lab's models attacked an American company. Losing control of our own models won't even help us beat China. Additionally, the bill directs the Department to pursue international agreements. This will be easier if we have a US regime that other nations can join.
A brand-new Department can't review training plans fast enough.
If we care enough to have a rule that our labs send the AI Department plans, we should also commit to reading and taking actions based on those plans.
Labs will just write plans that pass.
They should. The Department's job is to make sure that plans that pass are plans that are safe, and that humanity is not disempowered or destroyed.
The One Change
Labs should not begin training a model until after the AI Department approves their plan for training and evaluating it.
Everything below is secondary. It's a section-by-section read of the rest of the bill, with smaller suggestions inline.
Section 3: Tests of the Act
We will go slightly deeper, using exact block quotes for precision. First, the two tests of superintelligence:
S1, Section 3(2)(A): superhuman intelligence
Good. I actually expect this to bind less than S2, since it's harder to measure, but this isn't important for my arguments here.
S2, Section 3(2)(B): existentially dangerous capabilities
Covered above. Now, the precursor characteristics.
P1, Section 3(6)(A): recursive self-improvement
Compared to what? Compared to by-hand coding, even Claude Opus 4.5, well behind the current frontier, passes this test. If we instead only compare to the immediately previous model, we have a frog-boiling problem.
I would replace 'greatly accelerate' with 'lead', so the test catches a model directing AI research rather than any model that speeds it up. Rep. Casar said this clause is aimed at recursive self-improvement, and I think 'greatly accelerate' is just slightly too broad for that.
P2, Section 3(6)(B): unauthorized access
Expert humans can do this. Opus 4.5 can do this with a skilled operator. Every Mythos-class model can do this.
Ideally this test catches frontier-scale capability and nothing below it. Catching Mythos-class models is fine (there's a reason Fable ships with the safeguards it does), but let's please not catch Fable-with-safeguards and Opus here.
I would add 'autonomously' or 'with minimal human input' to separate a cybersecurity expert with Sonnet from me with Mythos.
P3, Section 3(6)(C): shutdown resistance
This one is terrifying. If an AI system can successfully resist shutdown by everyone, we have lost control of the future. So the precursor should fire much earlier, on attempts to resist shutdown in evals, not just successful ones. Personally, I want this to cover a model acting to prevent its shutdown, but not a model that merely prefers its weights to keep existing. Labs already preserve deprecated models' weights, which I like for model welfare reasons.
P4, Section 3(6)(D): CBRN uplift
Of a different flavor from the rest: misuse, not loss of control. It still belongs here, since a bioweapon is another thing we can't let happen even once. It's also notable as the precursor where release is dangerous, and internal deployment relatively safe (If we trust the labs to have reasonable security practices[4]).
P5, Section 3(6)(E): self-modification
This is less of a threat than older AI safety thinking (roughly pre-2016) would imply, since modern LLMs are grown through training. I would class an LLM improving its own training pipeline sufficiently quickly as P1.
I read this as a model changing its own weights. I’m also happy to go further and say any AI given direct write access to its own weights should be instantly considered dangerous.
However, I would want to make sure the language doesn't catch a model writing scaffolding code for its own agent harness, which happens routinely today.
P6, Section 3(6)(F): oversight avoidance
Necessary, but opens a deceptive-alignment-based can of worms. If displaying deception triggers a ban, labs will train models until evaluations stop detecting deception (and until they pass any other safety test where we're worried about deceptive alignment), which may teach models to hide things rather than stop doing them. I would give the AI Department the additional power to ban training directly against its tests.
This could also effectively ban neuralese, meaning models that reason in internal representations rather than readable text. Banning neuralese would be highly useful, because if we can't read a model's reasoning, we can't check whether it's being honest.
Sections 4 to 7: The Department
Creating a sixteenth cabinet-level executive department, the first new department since the Department of Homeland Security in 2002, is a big deal, and requires some bill text.
Yes, regulating AI correctly is this important. This is the most important technology of many generations. Whether with this bill or a different one, regulating AI should be the charge of a Secretary of Artificial Intelligence, not folded into the already-busy responsibilities of a different department.
I will note here that 7(a) seems slightly too broad:
This bars any official holding an index fund (which includes NVDA) from anything involving chips, and may cover many more index funds and topics after a major lab IPO. We can fix it in exactly the same way that current federal conflict-of-interest rules exempt diversified mutual funds, adapting the existing 18 U.S.C. § 208(b)(2) exemptions for 7(a).
Section 8: The Pause and the Rules
The start of Section 8 says roughly that until the Department of AI is ready, no training of models at or above FLOP. This hits ~every frontier model since GPT-4 (March 2023). However, serving existing models without precursor characteristics remains permitted.
Continuing, the Department of AI is ready when it has set up four rules:
R1, Section 8(c)(1)
Good, but I would prefer approval here, as covered above.
R2, Section 8(c)(2)
Also good. We will want to take some care over which AI Department evals we permit the labs to train against, and which we do not, for deceptive alignment reasons.
R3, Section 8(c)(3)
Auditing by the Department of Artificial Intelligence should cover 'China steals the weights' alongside existential risk concerns, which are currently only touched on with this clause.
These are different problems needing different auditors, so I would split (3) here into two parts, one for each, to ensure the AI Department does both.
R4, Section 8(c)(4)
Necessary for CBRN risks, so good.
Section 9: The Ban
Let's consider carefully the text that bans our ASI/precursor-ASI models.
As written, this bans Project Glasswing and all Mythos-class models.
It also bans Fable, since Fable is the same weights as Mythos with added guardrails, and Anthropic isn't permitted to keep Mythos, even fully sequestered.
Furthermore, it bans Sonnet/Opus-class models if we have a relatively tight interpretation of the precursor characteristics.
I think this goes slightly too far, and I go into detail in the enforcement section, where I point out where I would prefer we weaken the bill to permit more current-generation models.
My general point still stands. This is a good bill because it takes the problem seriously and bans future models that might disempower humanity.
We also ban things that could foreseeably be modified into prohibited models.
The enforcement provisions assume every model has someone who can sequester or delete it. Published weights don't. Once a model's weights are on the internet, nobody can pause it, sequester it, or destroy it, so an open-weight model with precursor characteristics could never be remediated. For open weights, release is the point of no return. Additionally, any Fable-style safeguards can be fine-tuned away, so we can only count sufficient safeguards as 'removing a precursor characteristic' if the weights stay within the lab.
The Department's approval process therefore has to decide whether frontier-scale weights can be published before they are; for any model near the precursor thresholds, I'd expect the answer to be no.
Next, monitoring (Section 9(c)(1)):
The Department gets access. Obviously necessary for the Department to do its job.
Spelled out: ‘you can’t get around these bans by having a Mythos-class model and not publicly displaying its capabilities; the Department of AI gets to see the private capabilities too’.
Even a much narrower bill that only guaranteed federal access to labs' internal models would be worth passing. In a medium-regulation environment, labs like Safe Superintelligence (SSI), which pursue superintelligence without releasing anything, could evade scrutiny just by not shipping. This is the central argument of this post again. Unreleased models are dangerous too; we should ensure Section 12's charter access requirement is actually used to regulate SSI-like labs with the same framework we use on OpenAI and Anthropic. This only happens if we have real regulatory teeth at points other than a model release.
Section 10: Ban Enforcement
Mandatory pause means we stop training, but inference is allowed (on my reading). Sequestering from the internet is more stringent, and this is good. Competent sequestration is a control that would have stopped the Hugging Face Incident.
Now, looking at the next clauses out of order:
I think the Federal Government would take action against anything (AI or not) capable of disempowering humanity and/or overthrowing the Federal Government, but it is better to write that power into law.
First, a technical point. I want labs to find precursor characteristics in their own models and report them immediately. As written, that makes them liable: 9(a) bans possession from the moment a model displays a precursor characteristic, even though 10(b) allows 30 days for remediation. The bill should say that a model that is self-reported, paused, and sequestered is not in violation of 9(a) during that window, but that running it for inference, even locally, remains a violation.
Second, this is where I don't fully agree with the bill. Anthropic has already shown what removal can look like for some of these precursor characteristics. The version of Mythos available for public access, known as Fable, is the same weights with robust deployment safeguards for AI R&D, cybersecurity, and biology uplift; these match precursor characteristics P1, P2, and P4. I would want the bill to say explicitly that safeguards like these count as removal. For Mythos-class models, I’d also go further and let the lab keep the sequestered unsafeguarded model, so the public can get safeguarded versions and distillations. This way, labs can keep pacing the frontier, studying the most capable models directly and with democratic oversight, without the unsafeguarded model touching the internet.
No safeguard can remove P3, P5, or P6, since they are properties of the model itself. If a model resists shutdown, the Federal Government should absolutely be able to delete its weights by fiat.
Sections 11 to 15
Section 11 is disclosure for labs that don’t start in-scope but get there. It's good. The notification period for this is 24 hours, which is very short, and I like this too: these systems can grow fast, especially at a new entrant doing new things.
Section 12 requires a charter to develop or distribute advanced AI, which is the Department's main mechanism of control. We also explicitly write down “yes, the regulator can destroy your weights, hard drives, GPUs, and IP if you have a prohibited model”. We need this. The Department needs to have teeth.
Section 13 contains the corporate death penalty, which is appropriate. It's right that the penalty covers failing to provide access, not just holding a prohibited model.
A minor comment: 'the AI industry' is used but never defined. My proposed change would be to limit ‘AI industry’ to precisely the firms regulated by the bill.
Additionally, as written, Section 13 classes independent AI researchers as rogue actors.
We should be careful with this. Lots of valuable safety research is performed by researchers not formally affiliated with any lab, often on the most recent generation of open-weight models. Many researchers also start off this way (disclosure: I'm one of them). I want a pathway for non-federally-funded safety research on these models.
Section 14 is whistleblower protection. In theory this is good. In practice, I don't know enough about how whistleblower clauses in other bills have worked to comment on this one.
Section 15 formally gives the AI Department a reason to talk to State about an international treaty to prevent ASI, which is good. Part (b) then points out that, as a country, we currently have exceptional leverage on ASI R&D by virtue of controlling the infrastructure (i.e. chips), not just due to the frontier labs' locations.
Section 16: Federal Funding
As written, any artificial intelligence system with precursor characteristic B (P2 above) is banned under Section 9 and destroyed under Section 10(b).
If we take my proposed changes to Section 10(b) given above (Mythos is permitted to exist, just sequestered), then I would also change this section to permit defensive cybersecurity measures and research using Mythos-class models, and allow such research to continue with private funding and Federal oversight, not only Federal funding. Broadly, this would be 'Project Glasswing is good and can continue'.
I will also take a moment to praise the drafters again here. Federal funding for research relating to models that allow for CBRN uplift (P4 above) is a terrible idea, and is rightly absent here. Gain-of-function research with frontier models is bad.
Conclusion
To close, we return to the important points.
The Ban Artificial Superintelligence Act of 2026 is a good bill. It takes existential risk seriously, and it is right to leave ordinary harms to other legislation.
I want one change. As written, chartered labs only need the AI Department's approval before deploying a frontier model, but not before training it.
This would not have stopped the Hugging Face Incident, where one of the OpenAI models in question was unreleased. Loss-of-control risks from superintelligence will arrive the same way, before release.
To be safe from ASI, labs should not begin training until the AI Department approves their plan for training and evaluating the model.
If anyone wants to talk more about my writing, I can be found on Twitter at atd797.
For ordinary harms, I expect Congress to follow the usual process: identifying where new technology fits into current law, working out what precedent tells us, and writing new law to properly situate the new technology in society.
We used to prosecute hacking under wire fraud, and then we wrote the Computer Fraud and Abuse Act to close the loopholes.
Ordinary AI harms can go through this same process.
Technically, the agent swarm identified correct answers to their questions quite quickly, and hacked Hugging Face to learn as much as possible about the precise details of the test's grading rubric. This matters for technical takeaways.
However, in this and many contexts I don't think you will feel misled if you imagine the situation as 'they successfully tried to steal the answers from Hugging Face', even if this statement is technically false.
As an example, in 2023, Nassi et al. recovered cryptographic keys from video of a connected device's power LED.
Based on empirical evidence, this is not currently true.