How would this work technically? . . . Ideally this compliance check would be bundled into the closed-source driver powering NVIDIA GPUs
Even better would be having the technology built right into Nvidia's chips -- like Apple has technology built into its A-series and M-series chips that prevent software not signed by Apple from running on them. This technology is called remote attestation or (more recently) trusted computing or sometimes "confidential computing" (although this latter term assumes a particular application of the tech).
How would you get NVIDIA to implement this? . . .Why would they implement a feature that constrains the sort of models which can run on their GPUs?
Nvidia has implemented it (in hardware) in some of its GPUs, and there are companies that use Nvidia's GPUs to host open-weights models in such a way that (if the technology has been designed and implemented correctly -- i.e., without any security holes) the hosting company cannot eavesdrop on the communications between the customer and the model. (The same basic technology --remote attestation-- can be used for the purpose you want.) tinfoil.sh is one such company. I used Tinfoil just yesterday to have a enhanced-confidentiality chat with Kimi K3.
This post from 2 months ago gives an overview of the technology.
In a recent comment an expert on this technology (remote attestation) states that Nvidia's implementation of the technology probably hasn't been tested much yet, so there's a good chance that there is some flaw in it that could be exploited with enough labor by experts. This is in contrast to Apple's implementation, which is many years old at this point and has been looking quite solid for years. Nobody for example has published a jailbreak for a recent iPhone running a recent version of iOS and even though there is a lot of interest in iPhone jailbreaks and a lot of glory for anyone who manages to publish one.
Thank you for your very informative comment. I took the time to get familiar with the different info/sources you provided.
I think NVIDIA's chip-level Confidential Computing and Attestation tech is interesting. Their whitepaper here gives information on how this is implemented on their H100 Hopper GPU. Pages 14 and 15 outline what threats their tech would and would not protect against - you might find that interesting.
What they have however, is more for their hardware proving its identity as a legit NVIDIA GPU, providing detailed information about the running firmware, microcode etc., and attesting that the GPU is impervious to snooping by e.g the cloud computing company managing it.
It doesn't provide an enforcement layer - the discretion for the GPU to decide whether or not to accept commands/models.
Your comparison of NVIDIA and Apple in this regard is interesting, however after doing some reading, I've realized there are a few inaccuracies:
- Apple prevents unsigned software from running on their chips, but this has no bearing on AI models. For example I could run any model freely on my M-series Mac, as long as it'll fit.
(Also, this restriction is more lax on Mac than on iOS. On Mac you can implement ad-hoc signatures which provide lower-tier privileges, but can still e.g. enable software run locally)
- Apple's vendor lock for software is more similar to NVIDIA's Root-of-Trust enabled Secure Boot process (which their Confidential Computing builds on), which ensures that only signed and authenticated firmware is used to boot the GPU (Page 9 of the whitepaper).
Apple (especially on e.g. iOS) just extends that signing requirement further up to the application layer, while NVIDIA intentionally makes such layers less restrictive. But again, even Apple doesn't restrict what AI models you can run on an M-series GPU.
To implement the core idea I'm outlining in this post, one approach would be for these chip makers to extend that 'ability to refuse' from firmware-level operations to AI model-level operations.
For example, Cloud computing providers could then take up the responsibility of leveraging this GPU's ability to refuse models, to ensure that only verifiably safe models can be run on their platform. This is something of a flip from the context NVIDIA's current Confidential Computing is framed in: Where users distrust the cloud providers managing these GPUs in the first place.
Because there is too much text on LW, I've taken this conversation private. If you (Mayowa or anyone else) want to be part of the conversation or would prefer the conversation to remain public, let me know.
This substack post has a similar idea, and goes into some more technical detail on verifiable hardware, if it's useful.
Thanks. I'll drop a comment on that post, expressing my perspective on it.
I actually spent a good amount of time building a detailed technical proof-of-concept for this GPU-level compliance based on AMD's SEV-SNP trusted execution environment. NVIDIA's CUDA is closed-source, so that's not exactly accessible for a POC. You can check out the GitHub here. I plan to possibly discuss it in detail, in a subsequent post.
However I felt like the biggest unanswered questions were more about economic incentives and industry consensus, than technical possibility. That's why I didn't go so much into technical implementation details here (possibly I'll insert a short reference to this work/Github repo in the post).
Tldr:
AI Agents (e.g. based on models like Claude Opus and Fable) are now powerful enough to be used as autonomous tools for large-scale cyberattacks.
This most powerful class of agents generally tends to be based on closed-weight (closed source) models (like Claude and Fable), which generally have significant safety guardrails and monitoring implemented by their parent companies.
Open weight models are not quite there yet, but they're not far behind (leaderboard, GLM 5.2). These models however, are distributed without the guardrails and monitoring infrastructure present in their closed-weight counterparts.
Given that this infrastructure tends to be the last resort against successful attempts to manipulate models to do harm, open-weight models lacking this safeguard poses a very pronounced cybersecurity risk (amongst other kinds of risk).
Here I'm proposing an additional approach for mitigating the risk from frontier open-weight models - GPU-level model-safety compliance.
How Do We Enforce Model Safety Today?
We do this across multiple layers of the AI LLM/Agent stack:
These are safety preferences baked into model weights. They include preferences from training data, fine-tuning policies, safety-oriented model architectures, etc.
This is usually only possible if you control the server where the model is deployed. If not, you never really get the chance to assess these inputs before they're processed by the model, or screen outputs before they're passed to the user. Examples include Anthropic's Constitutional Classifiers, toxicity filters, etc.
Again, only possible if you control the server where the model is deployed, or the agent harness within which the model operates (this was how Anthropic was able to detect the first reported AI-orchestrated cyber-espionage campaign). Generally does not apply to open-weight models.
This is a recent development where a national government like the White House in the US decides to evaluate new frontier models, before approving their public release. This most notably happened with Anthropic's Fable and Mythos models which were redeployed after White House evaluation, following the initial suspension of public access. So far, this approach has been applied to closed-weight models being developed by frontier AI companies in the USA.
Majority of the most capable open-weight models today however, are being built in China, and aren't subject to such a process.
Enforcing model safety at the GPU level, seems to me like a valuable addition to what we do today.
As with any sort of model safety enforcement approach, it's likely not going to be a catch-all that solves every possible concern in AI Safety - I'm just proposing it as yet another thing that could be done, to hedge against the risks of a massively powerful open-weight AI model being wielded by bad actors to effect great harm.
The Case For GPU-Level Safety Enforcement:
One of the core determinants of AI progress today, is availability of compute. The ecosystem has evolved to depend on GPUs (specifically NVIDIA GPUs) as its core computing infrastructure.
These GPUs are a scarce resource, and will continue to be pivotal to the operation of frontier AI models. Their importance is only set to grow as these models get bigger, more powerful and more compute-hungry.
Given all of this, it makes a lot of sense to me that this scarce, increasingly-valuable resource should be leveraged as a bottleneck where model safety is enforced. As far as we know, the eventual open-weight AI model that reaches AGI is still going to run on GPUs. Preparing a GPU-level safety implementation ahead of time, seems like a wise bet.
What Would This Look Like?
Essentially a GPU would refuse to load a model's tensors unless that model could prove that it met the requirements of some given AI Safety compliance standard.
An analogy which seems relevant here, is HTTPS certificates. Certificate Authorities (like the Electronic Frontier Foundation's Let's Encrypt) serve as a compliance body of sorts, providing certificates which attest that a web resource accessible via http requests to a public IP, is handling user data safely by encrypting it appropriately.
Today, any public-facing website which wants to be taken seriously, obtains certificates from these authorities to put members of the public at ease when visiting their site.
I'm proposing doing the same with GPUs and AI models. In the HTTPS analogy, a web developer wants access to the valuable resource which is public attention of the internet. They get a https certificate to show that they intend to handle that attention in a trustworthy way.
Here, a model wants access to the valuable resource with is GPU compute. Similarly, it requires a certificate to show that it intends to handle access to that compute, in a trustworthy manner.
Questions I'm Thinking About (And Would Appreciate Thoughts On):
A proxy for this however, could be model integrity. If we're to assume AI labs generally release models with at least a decent level of inbuilt guardrailing and safety enforcement mechanisms, then compliance here could just be ensuring that only these original models are given access to compute - denying access to variants which have been fine-tuned for malevolent purposes. This leads to the next question:
A concession could be: Only imposing this requirement on models judged to be dangerously powerful - for example models meeting or exceeding Anthropic's ASL-3 capability level. The more powerful a model is, the more havoc can be wreaked with a variant fine-tuned to undo its inbuilt safety guardrails.
You'd want it to be deep enough in the GPU stack, so it couldn't be easily circumvented. Ideally this compliance check would be bundled into the closed-source driver powering NVIDIA GPUs - CUDA. It could be as simple as adding this compliance check feature in the next CUDA version.
Newer GPUs tend to only be compatible with recent CUDA versions, so you can rest assured that when the massive "world ending" open-weight AI model does get released, it's only going to run efficiently on NVIDIA GPUs compatible with just the recent CUDA versions (which would have this safety compliance feature).
(You can run new massive models on clusters of older GPUs, but that generally tends to be less efficient and more expensive)
Would love thoughts, thanks.