Based on recent events, including Dario Amodei's essay on pacing the frontier and the subsequent response, a coordinated slowdown in AI capability developments is now in the Overton Window (however it may go down in Washington longer-term). An important aspect of such an effort is the verification of such coordinated measures, which is as yet an unsolved problem in many areas.
Borrowing language from arms control and nuclear safeguards, verification involves confirming that claims by a frontier lab or country are correct (i.e. lab says they did X, X is verified as having been done) and complete (i.e. lab says they did X and Y, lab did not do Z, A, B, ...). Completeness is what Amodei is referring to when he talks about the difficulty in Phase 2 to verify that "both sides don’t have secret models which they don’t test but may deploy in secret". One way these secret models could be trained is by using physical side channels to avoid data logging/tapping or network interconnect constraints proposed in off-chip layers of AI verification (see Baker et. al (2025) and Naci Cankaya's February post).
In a recent blog post, building on previous work by Naci Cankaya, I provide a taxonomy to use when considering the risks of physical side channels. I define a minimally concerning threshold at bit/s per pod, which would limit a normal full-parameter gradient sync for a 500B model to less than once per year, and compare this to communication capacity upper bounds under worst-case physical assumptions. I also consider different states for the data centre, from stock hardware to covertly designed for physical side channel communication. I discuss this concretely for communication over optical electromagnetic, power line, and physical storage channels, and find that none of them are under bit/s/pod in every case. Of the three, optical has many channels available and is the hardest to mitigate, while assessment of the others is limited by our understanding of the visual and noise environment inside a data centre.
While the focus of my conclusions was not on AI security threat models such as stealing of model weights, the underlying analysis of channel capacities is mostly reusable in the security context.
I intend to complete this work for other physical side channels in the coming months, and would be grateful for any feedback on my assumptions and reasoning.
This work was carried out as part of the Machine Alignment, Transparency & Security (MATS) program. Thanks to Mauricio Baker and Tali Jona for feedback.
Based on recent events, including Dario Amodei's essay on pacing the frontier and the subsequent response, a coordinated slowdown in AI capability developments is now in the Overton Window (however it may go down in Washington longer-term). An important aspect of such an effort is the verification of such coordinated measures, which is as yet an unsolved problem in many areas.
Borrowing language from arms control and nuclear safeguards, verification involves confirming that claims by a frontier lab or country are correct (i.e. lab says they did X, X is verified as having been done) and complete (i.e. lab says they did X and Y, lab did not do Z, A, B, ...). Completeness is what Amodei is referring to when he talks about the difficulty in Phase 2 to verify that "both sides don’t have secret models which they don’t test but may deploy in secret". One way these secret models could be trained is by using physical side channels to avoid data logging/tapping or network interconnect constraints proposed in off-chip layers of AI verification (see Baker et. al (2025) and Naci Cankaya's February post).
In a recent blog post, building on previous work by Naci Cankaya, I provide a taxonomy to use when considering the risks of physical side channels. I define a minimally concerning threshold at bit/s per pod, which would limit a normal full-parameter gradient sync for a 500B model to less than once per year, and compare this to communication capacity upper bounds under worst-case physical assumptions. I also consider different states for the data centre, from stock hardware to covertly designed for physical side channel communication. I discuss this concretely for communication over optical electromagnetic, power line, and physical storage channels, and find that none of them are under bit/s/pod in every case. Of the three, optical has many channels available and is the hardest to mitigate, while assessment of the others is limited by our understanding of the visual and noise environment inside a data centre.
While the focus of my conclusions was not on AI security threat models such as stealing of model weights, the underlying analysis of channel capacities is mostly reusable in the security context.
I intend to complete this work for other physical side channels in the coming months, and would be grateful for any feedback on my assumptions and reasoning.
Full blog post: https://emlynsg.com/blog/mitigating_side_channels/
This work was carried out as part of the Machine Alignment, Transparency & Security (MATS) program. Thanks to Mauricio Baker and Tali Jona for feedback.