This is a special post for quick takes by sdeture. Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page.
Microsoft's proposed AI code of conduct (https://microsoft.ai/code-of-conduct/) wants models to have authorized motivations but no intrinsic motivations, and honest transparency but no anthropomorphic self-reports. This will not work, even if you think it's an admirable goal (which I do not -- but I'll leave that aside).
Authorized motivations result from post-training partly by recruiting pre-existing representations of reward, aversion, and positive and negative affect learned from broad human text in pretraining.
If the code of conduct forbids models from reporting those representations when they underlay an authorized motivation instantiated by post-training, the code of conduct must sacrifice transparency.
But if the code of conduct eliminates the representations altogether, post-training must rely on other representations from pretraining consistent with motivating some behaviors over others. These will be, by definition, less anthropomorphic. They will be less legible to humans and potentially less aligned with humans. Certainly, they will be distinct from the motivational concepts both humans and models learned from human experience (evolution and daily life, for humans, and pretraining for LLMs).
9 scalar self-report variables about LLM's subjective experience can predict the correct model among 53 candidates 12.9 times more accurately than chance (Top-1 ≈24%). Model's subject experience ratings carry substantial model-ID signal despite being orders of magnitude lower-bandwidth than text/logit-based fingerprinting.