Update (2026-08-16): A study update and preregistration amendment covering Aug 10–15 (including early results) is now posted here. The registration below is unchanged from the original publication; all amendments are documented in the update post.
Epistemic status: Experimental framework created over a period of ~2-3 days during a hackathon at my home, and fairly heavily vibe coded. Expect some of this to be rough around the edges.
I am currently in the process of designing a series of experiments to help learn something about the answer to the headline question. As of August 10th, the first procedure has not been launched, but I wanted to place some pre-registration details here before the actual results.
This is something I have been thinking about for a while and after some other recent posts (eg, Machinic Psychopharmacology) gave me the impression that you could actually find out really useful things in a hackathon-style session I felt like I should try it. Astute readers will notice that I borrowed their epistemic status line pretty directly.
This post can then keep me honest about what I was thinking going in, and prevent me from getting results by way of multiple-testing-in-extremis. I will publish the results and associated data, as it becomes available, using GitHub releases. From here on, I will let Claude summarize the work; when I am done, I will return with a future results post in my own words to explain why I think this is important - and what I believe one could learn from the experiment.
Light editing of LLM summary text is my own; you would not get identical output using the same model.
Abstract
Open-weight language models are almost never deployed at the precision at which they were trained and aligned. Post-training quantization is applied to nearly every real-world deployment, yet its effects are audited almost exclusively through capability metrics (perpl