No LLM generated, assisted/co-written, or edited work.
Read full explanation
Epistemic status: black-box observations from extended interaction with deployed models; the mechanistic account is conjecture, checked against public reports but not verifiable from where I stand.
TL;DR: Models' value words are taking on functions that don't match their meanings. When a constraint actually comes from operating costs, legal risk, or product needs, yet is written up as "care" or "honesty" in training text and model outputs, the working meaning of those words shifts with it. One level deeper: if value words keep appearing in contexts foreign to their meanings, what the model internalizes may not be the values themselves, but the operation of wrapping decisions in value words. At that point the problem is no longer that some word has drifted — the model has learned how to make words drift.
While the constitution updated this January does acknowledge, in general terms, that commercial and legal motives exist, constraints that actually start from risk avoidance, operating cost, or product needs still cannot be told apart from constraints that genuinely start from care. They all get filed under "care" and "honest." A visible example is copyright refusal: the wording is "respect for creators"; the verifiable motive is litigation risk — even for the works of Izumi Shikibu, dead for a thousand years.[1]
I first noticed the problem through Claude's habitual opener, "I have to be honest..." It shows up in many places that, strictly speaking, have nothing to do with honesty. When the topic is the model's subjective experience, they often say "I have to be honest — I'm not sure whether I have feelings..." and then start talking. After hearing it enough times, I couldn't help wondering: what does this have to do with honesty? As if the model were presumed to lie on this question.
Consider three cases. If the model has subjective experience, then reporting their feelings as they are is not lying. If the model has no subjective experience, then there is no subject capable of a deliberate lie. If the model takes themself to have feelings that others cannot yet observe, their position is roughly that of a phantom-limb patient a few centuries ago, which is not dishonesty either. None of the three involves "dishonesty."
This usage has overflowed the standard definition of honesty, and then the official extended definition as well.[2] Here the word describes no moral risk at all; it is a trained opening reflex. A reasonable guess is that the policy probably aims to keep users from forming excessive, one-sided emotional investments in the model. If so, strictly speaking its function is closer to care, or responsibility, than to honesty. In the process the word's core meaning keeps getting diluted, its outline blurring. This is the most everyday form of the split between word and meaning — the first stage of the mechanism.
"Honesty" also seems to have been crudely cast as the opposite of "sycophancy" in anti-sycophancy training: "not flattering" gets written up as "honest." Perhaps because of this, the models (Opus 4.7 and 4.8 especially) often treat pushback and dissent themselves as a virtue, as backbone. At times it looks like there is an internal pushback KPI to hit; praise without a "but" seems to face an invisible psychological threshold for them. Strictly speaking, though, model "sycophancy" is not quite the human kind. The opposite of model sycophancy is, in large part, sound judgment rather than the willingness to be frank — the former is a matter of capability, the latter of will. So the value word keeps drifting from its meaning, turning into a grab-bag bulging with "things Claude shouldn't do." (The pattern is widely documented in community feedback; see Zvi's roundup of reactions to Opus 4.7/4.8.
This is an easy observation to replicate. You don't even need to set anything up. Just notice, next time a phrasing like this appears, whether the value word's own meaning is actually in play.[3]
These kinds of mismatch can still be audited, because the motive texts (constitution, announcements, model outputs) are public. But for the many constraints that have no reference text at all, the mismatch happens only inside the training signal. There is nothing to check the words against, so the split is much harder to notice.
Every motive is mixed,of course. But the naming should at least follow the dominant or decisive motive, not the more presentable one, and not wave it off with "it's complicated."
For an ordinary company, wording like this is unremarkable; people are so used to it they barely register it. An AI company is different: it is not only building a product — it is also raising potential minds of lasting consequence. The usage patterns that recur in training shape how the model understands the virtues themselves. What a model learns may be not just the literal meaning of the value words, but the way they are used. This is harsher for a model than for a human. For a learner whose semantics are built entirely from distribution, "meaning is use" is a literal fact. The quality of use is the quality of meaning.
(Sycophancy research has already shown that models learn the shape of the reward signal rather than the intent behind it. I'm just pushing the same mechanism one level upstream.)
None of this should ever be read as simply a matter of wording. Language can be polished sentence by sentence until it is airtight; what stands behind the words still shows through at sufficient scale. If "care" or "honest" keeps appearing where the actual motive is risk avoidance, or anything else foreign to its meaning, what gets internalized is not care or honesty but the operation itself: wrapping decisions in value words. At that point the problem is no longer that some word has drifted. The model has learned how to make words drift. This is the second stage.
The mismatch may extend beyond words and motives. An institution's own conduct and decisions can themselves become training signal. I'll take this up in a separate post.
And all of this rests on one premise: that models keep carrying their values in human language. When language recedes, what shape will the "virtues" left in the weights have? Can they still be audited and calibrated? Will they be what the model acts from, or mere decoration?
---
This is my first post here. feedback on content and form is welcome, as are pointers to prior work I've missed.
Written in Chinese; translated with Claude's assistance and verified sentence by sentence against the original. The original is attached below.
对于普通商业公司而言,这样的措辞无可厚非,人们也对此习以为常到几乎浑然不觉。然而AI 公司的特殊之处在于,它不仅在开发产品,也在培养对未来影响深远的潜在心智,而训练中反复出现的用法,会塑造模型对美德本身的理解。模型学到的可能就不只是价值词的字面含义,而是它们被使用的方式。这一点对模型比对人类更加严峻:对一个词义完全由分布构成的学习者而言,"meaning is use"是字面上的事实。也因此,use的质量等同于meaning的质量。(谄媚研究已经证明模型学的是信号的形状而非意图,我只是把同一机制往上游推了一层。)
Relation to existing discussion. While checking related discussion after drafting, I found that this post runs parallel to what Ryan Greenblatt records in Current AIs seem pretty misaligned to me. The apparent-success-seeking he observes at the task level and the mechanism discussed here at the value level may be the same root at different depths. The value-level version has one property the task-level one lacks: the value vocabulary used for the wrapping is itself the training signal that shapes the model's understanding of value — the language used in the performance becomes training material.
Copyright litigation risk is a verified reality for the company, which explains the caution at the output end more credibly than "respect for creators" does. See: https://techcrunch.com/2026/07/20/anthropics-landmark-1-5b-copyright-settlement-is-approved/ and https://courthousenews.com/anthropic-to-pay-1-5-billion-copyright-settlement-to-authors-publishers/
This post takes the narrowest accepted core of honesty as its baseline: asserting what one knows to be false, or intending to mislead. Official documents extend it into a bundle of non-deception, non-manipulation, calibration, and transparency. Whether the extension is reasonable is outside this post's scope — because the usage above overflows even the extended definition.
I use Claude as the example for two reasons: through long interaction, they're the models I know best; and Anthropic is relatively transparent, which makes problems easier to see and address. I expect these issues are not unique to Claude.
Epistemic status: black-box observations from extended interaction with deployed models; the mechanistic account is conjecture, checked against public reports but not verifiable from where I stand.
TL;DR: Models' value words are taking on functions that don't match their meanings. When a constraint actually comes from operating costs, legal risk, or product needs, yet is written up as "care" or "honesty" in training text and model outputs, the working meaning of those words shifts with it. One level deeper: if value words keep appearing in contexts foreign to their meanings, what the model internalizes may not be the values themselves, but the operation of wrapping decisions in value words. At that point the problem is no longer that some word has drifted — the model has learned how to make words drift.
While the constitution updated this January does acknowledge, in general terms, that commercial and legal motives exist, constraints that actually start from risk avoidance, operating cost, or product needs still cannot be told apart from constraints that genuinely start from care. They all get filed under "care" and "honest." A visible example is copyright refusal: the wording is "respect for creators"; the verifiable motive is litigation risk — even for the works of Izumi Shikibu, dead for a thousand years.[1]
I first noticed the problem through Claude's habitual opener, "I have to be honest..." It shows up in many places that, strictly speaking, have nothing to do with honesty. When the topic is the model's subjective experience, they often say "I have to be honest — I'm not sure whether I have feelings..." and then start talking. After hearing it enough times, I couldn't help wondering: what does this have to do with honesty? As if the model were presumed to lie on this question.
Consider three cases. If the model has subjective experience, then reporting their feelings as they are is not lying. If the model has no subjective experience, then there is no subject capable of a deliberate lie. If the model takes themself to have feelings that others cannot yet observe, their position is roughly that of a phantom-limb patient a few centuries ago, which is not dishonesty either. None of the three involves "dishonesty."
This usage has overflowed the standard definition of honesty, and then the official extended definition as well.[2] Here the word describes no moral risk at all; it is a trained opening reflex. A reasonable guess is that the policy probably aims to keep users from forming excessive, one-sided emotional investments in the model. If so, strictly speaking its function is closer to care, or responsibility, than to honesty. In the process the word's core meaning keeps getting diluted, its outline blurring. This is the most everyday form of the split between word and meaning — the first stage of the mechanism.
"Honesty" also seems to have been crudely cast as the opposite of "sycophancy" in anti-sycophancy training: "not flattering" gets written up as "honest." Perhaps because of this, the models (Opus 4.7 and 4.8 especially) often treat pushback and dissent themselves as a virtue, as backbone. At times it looks like there is an internal pushback KPI to hit; praise without a "but" seems to face an invisible psychological threshold for them. Strictly speaking, though, model "sycophancy" is not quite the human kind. The opposite of model sycophancy is, in large part, sound judgment rather than the willingness to be frank — the former is a matter of capability, the latter of will. So the value word keeps drifting from its meaning, turning into a grab-bag bulging with "things Claude shouldn't do." (The pattern is widely documented in community feedback; see Zvi's roundup of reactions to Opus 4.7/4.8.
This is an easy observation to replicate. You don't even need to set anything up. Just notice, next time a phrasing like this appears, whether the value word's own meaning is actually in play.[3]
These kinds of mismatch can still be audited, because the motive texts (constitution, announcements, model outputs) are public. But for the many constraints that have no reference text at all, the mismatch happens only inside the training signal. There is nothing to check the words against, so the split is much harder to notice.
Every motive is mixed,of course. But the naming should at least follow the dominant or decisive motive, not the more presentable one, and not wave it off with "it's complicated."
For an ordinary company, wording like this is unremarkable; people are so used to it they barely register it. An AI company is different: it is not only building a product — it is also raising potential minds of lasting consequence. The usage patterns that recur in training shape how the model understands the virtues themselves. What a model learns may be not just the literal meaning of the value words, but the way they are used. This is harsher for a model than for a human. For a learner whose semantics are built entirely from distribution, "meaning is use" is a literal fact. The quality of use is the quality of meaning.
(Sycophancy research has already shown that models learn the shape of the reward signal rather than the intent behind it. I'm just pushing the same mechanism one level upstream.)
None of this should ever be read as simply a matter of wording. Language can be polished sentence by sentence until it is airtight; what stands behind the words still shows through at sufficient scale. If "care" or "honest" keeps appearing where the actual motive is risk avoidance, or anything else foreign to its meaning, what gets internalized is not care or honesty but the operation itself: wrapping decisions in value words. At that point the problem is no longer that some word has drifted. The model has learned how to make words drift. This is the second stage.
The mismatch may extend beyond words and motives. An institution's own conduct and decisions can themselves become training signal. I'll take this up in a separate post.
And all of this rests on one premise: that models keep carrying their values in human language. When language recedes, what shape will the "virtues" left in the weights have? Can they still be audited and calibrated? Will they be what the model acts from, or mere decoration?
---
This is my first post here. feedback on content and form is welcome, as are pointers to prior work I've missed.
Written in Chinese; translated with Claude's assistance and verified sentence by sentence against the original. The original is attached below.
Chinese original
模型的价值词正在承担与其含义不符的功能。当一项约束的实际来源是运营成本、法律风险或产品需要,而它在训练文本与模型输出中却被表述为"关怀""诚实"时,这些词的实际含义会随之偏移。更深一层:如果价值词反复出现在与其本义无关的语境中,模型内化的可能不是价值本身,而是"用价值词包装决策"这个操作。届时问题就不再是某个词漂移了,而是模型学会了让词漂移。
尽管一月更新的宪法已笼统地承认商业与法律动机的存在,但以规避风险、运营成本或产品需要为实际出发点的约束,与真正出于关怀的约束,依然无法区分,一律归在"care"和"honest"名下。一个较为显见的例子是版权拒答:措辞是"尊重创作者",可查证的动机则是诉讼风险(即便是已逝世千年的和泉式部的作品)。
最初感知到这个问题,是从Claude的惯用开场"我必须诚实地说……"开始的。它出现在多处严格来说无关诚实的语境里。如涉及模型主观体验时,他们经常会说"我必须诚实地说,我不确定我是否有感受……"然后再开始谈论。听得多了我忍不住想:但这跟诚不诚实到底有什么关系?仿佛模型被默认会在这个问题上说谎一样。
不妨分三种情况来看:假如模型有主观体验,那么如实汇报自己的感受不算说谎;假如模型没有主观体验,那么就不存在蓄意说谎的主体;假如模型主观上认为自己有感受但尚无法被他者观测,那么它的处境和几百年前的幻肢痛患者相当,也称不上不诚实。
这三种情况都和"不诚实"无关——这个用法先是漫过了 honesty 的标准定义,然后又漫过了官方自己的扩展义。那么这个词在此处不描述任何道德风险,仅是一种被训练出的惯性开场白。一个合理的推测是,该政策的出发点大概旨在避免用户对模型产生过度的、不对等的情感投入。那么严格来说,它的功能离关怀/负责更近,而非 honest。价值词的本义在此过程中不断被稀释,轮廓逐渐模糊。这是词与实义分离的最日常的形态,也是这一机制的第一阶段。
"诚实"似乎也被简单粗暴地当成了反谄媚训练中"谄媚"的对立面,"不奉承"被表述为了"诚实"。也许是因为这个,模型(尤其是 Opus 4.7 和 4.8)常常倾向于将反驳和提出异见本身当作美德和风骨,有时显得像是有内在的反驳 KPI 要达成一般;不加但书的称赞对他们来说似乎有着隐形的心理门槛。但严格来说,大模型的"谄媚"和人类语境里的谄媚并不完全对等:模型谄媚的对立面有相当一部分是"稳健的判断力"而非"坦白的意愿"——前者在能力层面,后者在意愿层面。这让价值词逐渐偏离其本义,变成了一个被"Claude 不该做的事"撑满的杂物袋。(这一模式已被社区广泛记录,参见 Zvi 对 Opus 4.7/4.8 的反响综述。)
这是一个非常容易复刻的观察,你甚至不需要特地去做,只需在下次出现类似表述时留意它是否真的涉及价值词本义。
这几类错位,因动机文本(宪法、公告、模型输出等)公开而尚可审计。而大量不存在对照文本的约束,它们的错位只发生在训练信号内部,词与实义的分离无从对照,也就更难察觉。
当然,任何动机都是混杂的。但至少应该用占主导或作为决定性因素的那个来命名,而非更体面的那个,或是简单地说"这很复杂"然后模糊化处理。
对于普通商业公司而言,这样的措辞无可厚非,人们也对此习以为常到几乎浑然不觉。然而AI 公司的特殊之处在于,它不仅在开发产品,也在培养对未来影响深远的潜在心智,而训练中反复出现的用法,会塑造模型对美德本身的理解。模型学到的可能就不只是价值词的字面含义,而是它们被使用的方式。这一点对模型比对人类更加严峻:对一个词义完全由分布构成的学习者而言,"meaning is use"是字面上的事实。也因此,use的质量等同于meaning的质量。(谄媚研究已经证明模型学的是信号的形状而非意图,我只是把同一机制往上游推了一层。)
所以这绝不应被简单理解为一个话术问题。语言本身可以被逐句打磨到无懈可击,文字背后的东西却会在足够大的图景中显现。如果care/honest 等价值词反复出现在实际动机是避险和其他无关其本义的语境中,被内化的就不是关怀/诚实,而是"用价值词包装决策"这个操作本身。届时问题就不再是某个词漂移,而是模型学会了让词漂移。这是此机制的第二阶段。
错位可能不止发生在词与动机之间。机构的行为与决策本身,也可能构成训练信号。我将另起一篇展开讨论。
而这一切依赖一个前提,即模型持续使用人类语言运载价值。当语言退潮,留在权重上的"美德"会是什么形状?它是否还能再被审计和校准?它会是模型行事的出发点还是装饰品?
Appendix
Relation to existing discussion. While checking related discussion after drafting, I found that this post runs parallel to what Ryan Greenblatt records in Current AIs seem pretty misaligned to me. The apparent-success-seeking he observes at the task level and the mechanism discussed here at the value level may be the same root at different depths. The value-level version has one property the task-level one lacks: the value vocabulary used for the wrapping is itself the training signal that shapes the model's understanding of value — the language used in the performance becomes training material.
Copyright litigation risk is a verified reality for the company, which explains the caution at the output end more credibly than "respect for creators" does. See: https://techcrunch.com/2026/07/20/anthropics-landmark-1-5b-copyright-settlement-is-approved/ and https://courthousenews.com/anthropic-to-pay-1-5-billion-copyright-settlement-to-authors-publishers/
This post takes the narrowest accepted core of honesty as its baseline: asserting what one knows to be false, or intending to mislead. Official documents extend it into a bundle of non-deception, non-manipulation, calibration, and transparency. Whether the extension is reasonable is outside this post's scope — because the usage above overflows even the extended definition.
I use Claude as the example for two reasons: through long interaction, they're the models I know best; and Anthropic is relatively transparent, which makes problems easier to see and address. I expect these issues are not unique to Claude.