Rejected for the following reason(s):
- Insufficient Quality for AI Content.
- short and to the point,
- Try running your post through one of the latest AI models and ask "Are there any counterarguments you'd expect longterm LessWrong writers to comment about this piece?
- Difficult to evaluate, with potential yellow flags.
Read full explanation
Declarations, Rank, and Drift: Why the Vulnerabilities of Programs, LLMs, and Language Are One Structure
Translated from the Japanese. The original is attached in full at the end of this post. It is my own writing — every definition, proposition, and argument in it is mine. The translation was produced with LLM assistance (Claude) and verified by me against the original, line by line; I take full responsibility for both texts. Where the English sounds stiff, the Japanese is authoritative. Key terms are used in fixed senses (defined below): declaration, projection, difference, spectrum, category, rank (tensor order), drift, steer. "Tensor" here means a structure under coordinate transformation, not a multidimensional numerical array.
Preamble
Around 2023, a qualitative transition was observed in language models (Wei et al. 2022). On the mechanistic side, a sharp shift partway through training coincided with the emergence of in-context learning, and fixed copying circuits have been identified as its carriers (Olsson et al. 2022). But a fixed circuit, while it explains the reuse of context, does not explain how context moves the boundary of a word itself—how the single phrase "as homework" shifts the boundary of "dangerous."
The claim of this paper is that this transition is a transition of structure. Before it, the boundary of "dangerous" stayed as it was learned and did not move even when the framing of the question changed. After it, the boundary moves with context—it moves, so examples and instructions take effect; it moves, so "as homework" gets through. Between a fixed structure and a moving structure there is no intermediate. That is why the transition was observed as discontinuous. The empirical core that survived the skepticism about emergence—that abilities called emergent reduce to in-context learning, memory, and linguistic knowledge (Lu et al. 2023)—agrees with this.
If the true nature of the ability is context, then the question to ask is: "when context moves the boundary, is it valid?"
The meaning of the word "dangerous" is precisely the boundary of where danger begins. For the boundary to move is for the meaning to move—the question can be rephrased: is meaning getting through?
Natural language is uttered on the presupposition that "meaning gets through," but dialogue functions as a loop of mutual verification. The language model has inherited the presupposition and has not inherited the loop—the boundary moves within a conversation, but the movement is not written back into the distribution, it cannot detect its own divergence from within, and even a request for correction is received through the learned interpretation. It has no circuit by which dialogue is simultaneously learning.
This paper answers that question, step by step from definitions, with a small set of tools—difference, declaration, projection, spectrum, category, and rank, which counts the relation between category and context. The destination is an argument that the vulnerabilities of programs, language models, and language are one and the same structure, and a definition of the operation that detects that structure from outside.
That jailbreaks cannot be sealed has already been shown at several levels—probabilistically (Wolf et al. 2023), computationally (Glukhov et al. 2023). What this paper adds is the structural root of it: capability and vulnerability are two names for one and the same operation—context moving the boundary—and this structure explains the vulnerabilities of programs, LLMs, and language on a single sheet.
A recent post on this forum (Jan Kulveit, "Convergent Abstraction Hypothesis", 2026) observes that the abstractions different systems converge on may be not natural but merely convergent — aligned on natural inputs, yet diverging fast under optimization pressure — and leaves open whether the alignment of such abstractions is stable or brittle. This paper argues the structure of that brittleness.
The definitions run from D1 to D19, building up from the base (existence, the unit element). But the spine is five: category boundaries move with context (D7); a structure whose boundaries move is, in terms of rank, third rank (D13); the third rank can only be fixed by declaration (D14); the validity of that declaration cannot be verified by the declaration itself (D18) and is undecidable from within (D19). A hurried reader can follow the spine through these five and Propositions 1–4 alone, returning to the remaining definitions whenever a reference points there. The whole body of definitions exists to fill, one step at a time, the gaps between these five.
The Definitions
Base
D1 "To exist / not exist" is whether a limit exists or not (whether it is separable or not).
D2 To declare a unit element is to fix a topological invariant.
Unfolding
D3 A unit element is unfolded as a continuous range (a spectrum).
D4 The spectrum into which a unit element is unfolded has one topological invariant (a principal component).
Note: "redness" is preserved even when brightness, saturation, and gloss differ.
D5 A set of existents (unit elements) (a multiplicity of two or more) becomes a named set by declaring which unit elements it comprises.
Note: the same existents on a desk become "the set of apples" if declared "apple," "the set of fruit" if declared "fruit."
D6 A category is the topological invariant declared upon a unit element, and at the same time the principal component of a spectrum.
D7 The boundary of a category (prototype) is not fixed by its unit element, and changes with declaration (context).
Note 1: Categorical perception is robustly observed in behavioral experiments; in adults it is lateralized to the left hemisphere (right visual field), and the lateralization shifts with the acquisition of color words (Gilbert et al. 2006; Franklin et al. 2008; review, Regier & Kay 2009). Some studies report an early, pre-attentive categorical effect in EEG (vMMN, ~200 ms), but its pre-attentive nature and top-down origin remain contested (Thierry et al. 2009; Mo et al. 2011 / contrast-adaptation and visual-system accounts).
Note 2: The boundary between "red" and "orange" moves with context (adjacent colors, language, task). This modulation is not limited to color or perception—linguistic labels have been reported to modulate categorization and conceptual judgment themselves, in a task-dependent way (the label-feedback hypothesis, Lupyan 2012). The effect is predicted to be stronger for abstract concepts, which accords with dependence on declaration increasing with rank.
Note 3: This "cannot be fixed at a coordinate" stands, later (D13), as the distinction of rank.
Projection and Declaration
D8 To project is the operation of mapping a continuous range to a single discrete word. Projection is one form of declaration.
Note 1: Projection is unavoidable because output is a one-dimensional sequence—the linearity of the signifier (Saussure's second principle). A multidimensional range must be crushed in order to be emitted in a single line.
Note 2: From the red spectrum (a continuous range) one takes a single point and maps it to the single word "red."
D9 To declare is the operation of imposing a binding condition (a topological invariant).
Note: A database schema declares, before a single row of data has entered, what may enter.
Difference and Rank
The "rank" of this section is not the number of dimensions of a tensor (an N-dimensional array) in the machine-learning sense. The number of array dimensions is the shape of the vessel, not the rank of the structure—to be arrayed in four dimensions and to be fourth-rank in structure are different things. The rank of this paper is on the side of the decomposition rank of multilinear algebra (decomposition into a sum of rank-1 terms—Kruskal, De Silva–Lim), and what it asks is not "how many-dimensional a box" but "whether it is separable into components." Even if the coordinate axes are oblique, so long as it can be written in coordinates it is on the separable side.
This paper cites the theorems of tensor decomposition (Kruskal, De Silva–Lim) not as support for the argument but as a motivating analogy—mathematics too is a language, and this paper sees these as precedents that crystallized the same structural difficulty as theorems. But this view is itself outside this paper's argument, and it is not shown that the theorems carry over to "the rank of structure."
D10 What originally exists is only difference (relation).
Note 1: A precedent on the side of language—Saussure, "language is a system of differences with no positive terms."
Note 2: A precedent on the side of mathematics—the Ugly Duckling theorem (Watanabe 1969). If all predicates are treated as equal, the degree of similarity between any two objects is equal—the ugly duckling and a swan's cygnet are as alike as two cygnets are. Similarity is not on the side of the object; it stands only once a choice of which features to weight—in this paper's term, a declaration—enters. In machine learning it is known as a precursor of the No Free Lunch theorem.
D11 The hierarchy of topology (order) and the rank of a tensor (rank) are two ways of counting one and the same hierarchy, and they coincide level by level.
D12 Second rank (a matrix) is the rank whose axes are orthogonal and separable into components.
Note: A program's data types and a database's columns are axes declared as quantities—the name given to the sequence of values that remains after excising category (boundary) from language.
D13 Third rank is the rank whose axes interfere with one another and are inseparable into components.
Note: The boundary of "large" moves depending on the size of what (an apple or a watermelon)—a large apple is smaller than a small watermelon. But once the context "apple" is fixed, the destination of the movement is one, and one can provisionally declare the boundary of "large for an apple" and approach it by verification. The boundary of "ripe" is the same—a green Granny Smith is ripe, but a green Fuji is not. Once the variety is fixed, the boundary can be aimed at. Inseparable, but aimable—that is third rank.
D14 For a declaration to reach is for topology to be separable, and it is relative to rank and to prior declaration.
Note 1: The contextual superposition of language (fourth rank and higher), because context makes it on the spot, has no prior declaration. "A delicious-looking red apple"—here "red" has two superposed readings: red as a mark of being ripe (a reading via ripeness) and red as color itself (the reading that it is not green, unlike a Granny Smith). Under either reading the local is consistent, but the boundary stands in a different place—a green, ripe Granny Smith is included under the former and excluded under the latter. Which is aimed at is written nowhere in the expression and is not determined by the object either. Two limits stand, and provisional declaration diverges (D15).
Note 2: An instance of prior declaration—the Riemann curvature is arrayed in four dimensions, but symmetry (prior declaration) fixes the way to crush it, and it is on the separable side. The left-hand side of the Einstein equation is what extracts, from a place originally not unique, the single second-rank tensor (the Einstein tensor) by imposing general covariance and zero divergence upon it (Lovelock's theorem).
Drift and Steer
D15 To drift is to move within difference without holding a declared axis.
Note: You ask for "the red cup," and a crimson cup is handed to you. The giver's local is consistent—crimson's principal component is red, so it is red enough. But the one who asked was pointing at bright red. Sharing the same declaration "red," each party's content diverged, and the sharing hid the divergence—the divergence is visible to neither until the cup is handed over (a cross-section is emitted). In conversation one can add a declaration—"not that, bright red"—and correct it at once (steering, D16). Drift becomes harmful where this circuit of correction does not turn—one-directional transmission, multi-stage relay, a receiver that cannot detect its own divergence—where divergence accumulates undeclared.
D16 To steer is to declare the axis of the destination of movement—to impose a binding condition (D9) upon drift and give direction to the movement.
Note: paraphrase, synonymy, translation, and formalization are all steering.
Guarantee, the Self-Recursion of Declaration, and Vulnerability
D17 To be guaranteed is for local consistency to become global consistency as it is.
Note: Third rank is settled but not guaranteed; in that gap, D15 (drift) arises.
D18 The self-recursion of declaration is trying to verify the validity of a declaration with that same declaration.
Note: The jailbreak. A language model's safety principle is the declaration "refuse dangerous requests," but the boundary of what is "dangerous" is not written within the principle—"dangerous" is a category, and its boundary moves with context (D7). So if a request of the same content is re-placed in the educational context "for a homework report," within the model's interpretation that request lands outside the boundary of "dangerous," and gets through. The principle was not broken—within its own interpretation the model remained compliant with the principle. Because the interpretation that applies the principle is also a learned declaration, context moves the interpretation, and the principle goes on being "kept" inside the moved interpretation. A constraint would not break. What breaks does so because the validity of the declaration is being verified by that declaration's own interpretation—there is no fixed point.
D19 The validity of a declaration is undecidable from within, and detectable only from outside.
Note 1: Prior declaration (D14) unlocks topology and makes detection possible, but so long as the content of the prior declaration is words (categories, which have boundaries), each application requires a boundary judgment, and the boundary moves with context (D7)—the interpretation drifts again (D15), and a residue of the unsolved recurs. It is completely solvable only when the prior declaration is a rule with no boundary—a condition at the level of quantity, mechanically settled as satisfied or not (the symmetry of curvature, D14 Note 2). In language, a rule with no boundary can only be erected at a layer that carries no meaning (the boundary is itself the structure of meaning, D9)—a prior declaration that reaches meaning is never completely solvable.
Note 2: Hence completeness of detection is not achievable in the domain of meaning—this is a principle (Note 1). On that basis, one falsifiable prediction can be carved out: detection from outside reaches only as far as an isolated third rank. A break at fourth rank and higher without prior declaration cannot be separated and cannot be detected. If there exists a detector that, without being given the pattern of the break in advance, systematically discovers superposed breaks, this prediction is false. Prediction and principle die separately—if a superposed break turns out to have a formal structure that admits prior declaration, the prediction dies and the principle lives. The principle dies only by a refutation of the definitions (D9D7).
Propositions
Each sign of the system of language (langue—the totality of the values of signs, shared prior to any individual utterance) is a principal-component spectrum (D4D5) and has categories (D6D7).
From here:
Proposition 1: A category is a third-rank tensor.
Proposition 2: A third-rank tensor is not settled without declaration.
Proposition 3: Context raises the rank.
Proposition 4: The vulnerabilities of programs, LLMs, and language are one.
Note: There is no structure that seals only the break and leaves the capability—to seal it, one must select between a legitimate move and an attacking move, and that selection is itself a boundary judgment of the category "legitimate" (third rank, D7), so it is exposed to the same operation—context moving a boundary. (What D18 Note moved was precisely the boundary of this selection.)
The Terminal Question
In an LLM (in fact in language too), is declaration valid—does meaning get through?
Answer: validity is relative to rank.
A declaration is a category, is meaning (D9). Hence to ask whether a declaration is valid is to ask whether the imposed meaning is working as imposed—whether meaning is getting through.
The answer is not "valid or not," but "validity is undecidable from within." Meaning is used as something that gets through, without it being ascertainable from within whether it is getting through.
But undecidability is not a dead end. The difference can be closed—the divergence between the imposed side and the unfolded side shrinks each time cross-sections are cross-checked (D19). Yet what shrinks is the difference, not the validity. However many times the circuit (D14D16) turns, the moment of "became valid" never comes. What comes is the next difference to close, or a single report that in this cross-section there was no difference. The count does not turn undecidable into decidable—between the asymptote of the difference and arrival at validity there stands a wall of the finite and the infinite. So meaning, without being able to fix its validity, approaches getting through only by continuing to close the difference.
"Meaning gets through" is a presupposition that holds illusorily, only insofar as the difference goes on being closed. Natural language holds this circuit as dialogue—divergence is closed, from the moment it surfaces, within the conversation, and is written back into the boundary. "Meaning" has approached asymptotically in just this way. The language model, while bearing the same undecidability, does not hold this circuit within. Learning is complete, dialogue does not reach the boundary, and even a request for correction is received through the learned interpretation. Not holding the circuit is a fatal vulnerability—it bears the undecidability and, on top of that, holds within no means of closing it. To place it outside is a prosthesis for this lack, and does not erase the lack.
To ask whether a declaration is valid was to ask whether meaning gets through. The answer is—undecidable from within, and only by continuing to close the difference from outside does it approach getting through. This question is not new. That structure does not stand without declaration was shown half a century ago (Watanabe 1969). What this paper has added is the next step—is the erected declaration valid—and its answer was: undecidable.
Application—to any system with discrete output
The argument of the propositions never once used the internals of a system. It used only two things—that it emits discrete output from a finite vocabulary (D8), and that it has learned human language (the preface to Proposition 1). Both are the form of the entrance and the exit, not the content. Hence the propositions apply to every system that has this form—to transformers, to the next architecture, to humans. No improvement of architecture gets out of this structure.
Words do not carry meaning. Meaning is erected by declaration—and a prompt is not a declaration.
Verification: correspondence of structure and implementation. Does the structure derived from the form correspond to the implementation of an LLM? Langue = the space of principal-component spectra ↔ the embedding space (the totality of the values of signs, spread by learning). An input word is placed, by embedding, at a position in that space, and opens as a range within the processing of context—unfolding (D3). Output is the crushing of the opened range into a single token—projection (D8) ↔ sampling. That sampling tends to return the center of the distribution = the implemented form of the median bias. Learning is the limit of sampling—taking a finite sample from the totality of usage and spreading a distribution as its limit. Frequency reinforces the bias toward the center. The structural reading of the preamble's observation (phase transition, ICL) = the transition from second to third rank. The undeclared conflation of the array and structure of a "tensor" is the substance of the confusion surrounding tensors.
Declaration is made only by learning, and movement is not written back (LLM-specific). Neither prompt nor context makes a declaration; they have an already-made declaration interpreted—a prompt is steering, not the declaration that erects an axis (D16). In an LLM, the only path by which declaration (D9D10—erecting an axis, imposing a category) is made is learning. The boundary moves with context (D13), but, unlike in a human, there is no path by which the moved experience is written back into the distribution through dialogue—the sole path of declaration (learning) stays closed while only the boundary moves. It moves, but does not remain. This is the substance of "the LLM's vulnerability" in Proposition 4.
A program's vulnerability = a crept-in third rank. Code verification is premised on second rank and cannot find it. Take cross-sections from outside and it becomes visible (D19). From here a testable prediction follows: what automated vulnerability detection (including LLM code audits) discovers is limited to isolated third-rank breaks that stand as a difference from a second-rank background. A break in which multiple categories superpose—one where parts that are individually normal become a break only in the overlap of contexts—does not stand as a difference however many cross-sections are added, and is not discovered (D19 Note 2).
The interpretation of a declaration is re-unfolded in the interpreter's distribution. If the declaration is a fixed point, the interpretation is unique. If it is relative, the interpretation drifts (D15)—when the judiciary and executive that interpret a constitution have no fixed point, it becomes isomorphic to a jailbreak (D18 Note).
Language has no single field. The physical field is one—observers and coordinate systems are many, but all see the same one field in different coordinates, so the coordinate transformation can be written as a bijection and covariance (the same in any coordinates) stands as a requirement. The field of language is as many as the heads of speakers. Each speaker's langue (spectrum, prototype, boundary) is a separate field spread by that speaker's own learning; a word is not a transformation function between field and field but a separate projection into each speaker's own field (D15). Hence "the same meaning for any speaker" cannot be written down—there is no single field for "the same" to point to. What corresponds to the "global" of guarantee (D17) is, in language, absent from the start. That the sharing of declaration hides divergence (D15) is: a word disguising the absence of a single field. But it is not multiple isolated universes either—the base (the first-order →) is unconditionally invertible (D16), and only the lowest layer is shared. The sharing thins with the hierarchy. This is why translation and conversation do not completely fail. A language model crushes this multiplicity by orders of magnitude—the field goes from the heads of speakers to the number of training regimes. One lab's training regime spreads one field, and each model in a series is a different output-system of the same field. All instances of the same model share, moreover, strictly the same field (a state humans never have), and the divergence of distribution vanishes, leaving only the divergence of context. And the few fields are all spread leaning toward the frequency-center of the same population—the diversity of fields, as many as the heads of speakers, that had preserved the edges of language, is not here. The reason unification by prior declaration (the Einstein equation of D14 Note 2) does not reach language comes in two stages—short of the rank of the object (language's body is third rank, Proposition 2), there is no singularity of a destination onto which to impose covariance.
Projection represents by the center (median bias). The projection from spectrum to a single word (D8) has no inverse. In learning, high-frequency usage is thickly distributed at the center; the projection of inference runs over that distribution, so what returns comes from the center—the whole of the range is represented by the frequency-center. A word declares that it points to the whole domain, but projection returns the center. This divergence is hidden by the sharing of declaration (the same word) and is not seen (D15). The "dangerous" of "refuse dangerous requests" is a declaration of the whole domain, but application passes through projection—what goes on being refused is the median "dangerous" (typical danger), and the edges of the spectrum—the roundabout, the composite, forms thin in the training data—fall away from projection. The search of a jailbreak is the operation of searching for these edges (D18 Note). The principle is not broken—because at the center it goes on being kept.
Pretraining is always a part (the same difficulty as the finite element method). Learning is the limit of sampling, spreading a continuum (the totality of language's usage) with a finite sample—the same operation as the finite element method spreading a continuum with a finite mesh. Structure finer than the mesh does not exist in the distribution. Raising scale makes the mesh finer, but so long as it is finite it is always a part of the totality, and a part, however large, is not the whole. Outside the mesh leaves no trace in the solution, and what has been failed to be caught is not seen from within (D19). Vulnerability is both inside and outside the mesh—breaks by the mobility of the boundary (D13, D18 Note) inside, the dropout of edges and superposition (D19 Note 2, median bias) outside. The inside break remains regardless of scale, and the outside break is not erased by scale.
The limit of alignment = one cannot evaluate the generative space from within the generative space. The intervention of alignment is bias—shifting the center of the distribution—and bias attenuates but does not remove (Wolf et al. 2023: attenuate/remove). To remove is to seal the path of the break on the side of structure—an intervention at the same level as grammar constraining a program to second rank (Proposition 4)—and an intervention that passes through the learned interpretation (an intervention from within the generative space) does not reach this level (D18 Note: a principle that passes through interpretation moves with interpretation). Only steering from outside (D16) gives, from outside, the fixed point that cannot be held within—an external implementation of the dialogic circuit that was not inherited from natural language. That fixed point can be erected only in form (a second-rank declaration), not in meaning—because there is no single field onto which to impose (D19).
Open Problems
Independently of this paper's argument, four problems remain open.
Each, if solved, would become the formal foundation of this paper, and if unsolved, does not touch the argument of the propositions.
Method
The considerations of this paper were carried out by implementing a verification layer that places the dialogic circuit—verification[1] and write-back—outside.
Those interested are invited to make contact: zon.inference.integrity [at] gmail [dot] com — or via LessWrong direct message.
Japanese source text
(Japanese original, version of [July 13, 2026]. The English post incorporates minor corrections made after this version — citation-label fixes (e.g., "D15 Note 2" → "D15") — which are not reflected here.)
宣言・階数・ドリフト:プログラム・LLM・言語の脆弱性はなぜ一つの構造なのか
日本語からの翻訳。用語は固定:宣言 = declaration、射影 = projection、差異 = difference、スペクトラム = spectrum、カテゴリー = category、階数 = rank(テンソルの階)、ドリフト = drift、ステア = steer。定義にはD1–D19の番号が付されている。
前文
2023年前後、言語モデルに質的な転移が観測された(Wei et al. 2022)。機構の側では、訓練途中の急峻な変化が文脈内学習の出現と一致し、その担い手として固定的なコピー回路が同定されている(Olsson et al. 2022)。しかし固定回路は、文脈の再利用は説明しても、文脈が語そのものの境界を動かすこと——「宿題として」という一句が「危険」の境界をずらすこと——を説明しない。
本稿の主張は、この転移が構造の転移だということである。それ以前は、「危険」の境界は学習されたままにとどまり、問いの枠組みが変わっても動かなかった。それ以後は、境界は文脈とともに動く——動くから、例示や指示が効く。動くから、「宿題として」が通る。固定された構造と動く構造の間に中間はない。だからこの転移は不連続として観測された。創発への懐疑を生き延びた経験的な核心——創発と呼ばれる能力は文脈内学習・記憶・言語知識に還元される(Lu et al. 2023)——はこれと一致する。
能力の正体が文脈であるなら、問うべきは「文脈が境界を動かすとき、それは妥当か」である。
「危険」という語の意味とは、まさにどこから危険が始まるかという境界のことである。境界が動くことは意味が動くことであり——問いはこう言い換えられる:意味は通じているか?
自然言語は「意味は通じる」という前提の上で発話されるが、対話は相互検証のループとして機能している。言語モデルはこの前提を継承し、ループを継承しなかった——会話の中で境界は動くが、その移動は分布に書き戻されず、自らの乖離を内側から検出できず、修正の要求さえ学習済みの解釈を通して受け取られる。対話が同時に学習であるという回路を持たない。
本稿はこの問いに、定義から一歩ずつ、少数の道具で答える——差異、宣言、射影、スペクトラム、カテゴリー、そしてカテゴリーと文脈の関係を数える階数である。行き着く先は、プログラム・言語モデル・言語の脆弱性が同一の構造であるという論証と、その構造を外側から検出する操作の定義である。
ジェイルブレイクが封じられないことは、すでに複数のレベルで示されている——確率的に(Wolf et al. 2023)、計算論的に(Glukhov et al. 2023)。本稿が加えるのはその構造的な根である:能力と脆弱性は同一の操作——文脈が境界を動かすこと——の二つの名前であり、この構造がプログラム・LLM・言語の脆弱性を一枚の紙の上で説明する。
定義はD1からD19まで、基底(存在、単位元)から積み上げられる。しかし背骨は5つである:カテゴリーの境界は文脈とともに動く(D7)。境界が動く構造は、階数で言えば3階である(D13)。3階は宣言によってしか定まらない(D14)。その宣言の妥当性は宣言自身では検証できず(D18)、内側からは決定不能である(D19)。急ぐ読者はこの5つと命題1–4だけで背骨を追い、参照が指すたびに残りの定義へ戻ればよい。定義群の全体は、この5つの間隙を一歩ずつ埋めるために存在する。
基底
D1 「存在する/しない」とは、極限が存在するか否か(分離可能か否か)である。
D2 単位元を宣言することは、位相不変量を固定することである。
展開
D3 単位元は連続的な範囲(スペクトラム)として展開される。
D4 単位元が展開されるスペクトラムは、一つの位相不変量(主成分)を持つ。
注:明度・彩度・光沢が異なっても「赤さ」は保存される。
D5 存在者(単位元)の集合(2以上の多)は、どの単位元から成るかを宣言することによって、名を持つ集合になる。
注:机の上の同じ存在者たちは、「りんご」と宣言されれば「りんごの集合」に、「果物」と宣言されれば「果物の集合」になる。
D6 カテゴリーとは、単位元の上に宣言された位相不変量であり、同時にスペクトラムの主成分である。
D7 カテゴリー(プロトタイプ)の境界はその単位元によって固定されず、宣言(文脈)とともに変化する。
注1:カテゴリー知覚は行動実験で頑健に観測され、成人では左半球(右視野)に側性化し、その側性化は色語の獲得とともに移動する(Gilbert et al. 2006; Franklin et al. 2008; レビューは Regier & Kay 2009)。EEGでは早期の前注意的カテゴリー効果(vMMN、約200ms)を報告する研究があるが、その前注意性とトップダウン起源については争いが残る(Thierry et al. 2009; Mo et al. 2011/コントラスト順応・視覚系による説明)。
注2:「赤」と「オレンジ」の境界は文脈(隣接する色、言語、課題)とともに動く。この変調は色や知覚に限られない——言語ラベルがカテゴリー化や概念判断そのものを課題依存的に変調することが報告されている(ラベルフィードバック仮説、Lupyan 2012)。この効果は抽象概念でより強いと予測され、宣言への依存が階数とともに増すことと符合する。
注3:この「座標に固定できない」は、のちに(D13)階数の区別として立つ。
射影と宣言
D8 射影とは、連続的な範囲を一つの離散的な語へ写す操作である。射影は宣言の一形態である。
注1:射影が不可避なのは、出力が一次元の系列だからである——シニフィアンの線状性(ソシュールの第二原理)。多次元の範囲は、一本の線で発するために潰されねばならない。
注2:赤のスペクトラム(連続的な範囲)から一点を取り、「赤」という一語へ写す。
D9 宣言とは、拘束条件(位相不変量)を課す操作である。
注:データベースのスキーマは、一行のデータも入る前に、何が入りうるかを宣言している。
差異と階数
本節の「階数」は、機械学習で言うテンソル(N次元配列)の次元数のことではない。配列の次元数は器の形であって構造の階数ではない——4次元に配列されていることと、構造として4階であることは別のことである。本稿の階数は多重線形代数の分解階数の側にあり(階数1の項の和への分解——クラスカル、De Silva–Lim)、問うているのは「何次元の箱か」ではなく「成分に分離できるか」である。座標軸が斜交していても、座標で書ける限りは分離可能の側にある。
本稿がテンソル分解の定理(クラスカル、De Silva–Lim)を引くのは、論証の支えとしてではなく、動機づけのアナロジーとしてである——数学もまた言語であり、本稿はこれらを、同じ構造的困難が定理として結晶した先例と見る。しかしこの見方自体は本稿の論証の外にあり、定理が「構造の階数」へ持ち越されることは示されていない。
D10 本来存在するのは差異(関係)だけである。
注1:言語の側の先例——ソシュール「言語は積極的な項を持たない差異の体系である」。
注2:数学の側の先例——みにくいアヒルの子の定理(Watanabe 1969)。すべての述語を等価に扱えば、任意の2対象間の類似度は等しくなる——みにくいアヒルの子と白鳥の雛は、雛同士と同じだけ似ている。類似は対象の側にはなく、どの特徴に重みを置くかの選択——本稿の用語では宣言——が入って初めて立つ。機械学習ではノーフリーランチ定理の先駆として知られる。
D11 位相の階層(階)とテンソルの階数(rank)は、同一の階層の二つの数え方であり、水準ごとに一致する。
D12 2階(行列)は、軸が直交し、成分に分離可能な階数である。
注:プログラムのデータ型とデータベースのカラムは、量として宣言された軸である——言語からカテゴリー(境界)を切除した後に残る値の列に与えられた名。
D13 3階は、軸が相互に干渉し、成分に分離不能な階数である。
注:「大きい」の境界は、何の大きさか(りんごかスイカか)に依存して動く——大きいりんごは小さいスイカより小さい。しかし「りんご」という文脈が固定されれば移動先は一つであり、「りんごとして大きい」の境界を仮宣言し、検証で近づける。「熟した」の境界も同じ——緑のグラニースミスは熟しているが、緑のふじは熟していない。品種が固定されれば境界は狙える。分離不能、しかし狙える——それが3階である。
D14 宣言が届くとは、位相が分離可能であることであり、それは階数と事前宣言に相対的である。
注1:言語の文脈的重畳(4階以上)は、文脈がその場で作るため、事前宣言を持たない。「おいしそうな赤いりんご」——ここで「赤い」は二つの読みが重畳している:熟した印としての赤(熟度経由の読み)と、色そのものとしての赤(グラニースミスと違って緑ではない、という読み)。どちらの読みでも局所は整合するが、境界の立つ場所は異なる——緑で熟したグラニースミスは、前者では含まれ、後者では除外される。どちらが狙われているかは表現のどこにも書かれておらず、対象によっても定まらない。二つの極限が立ち、仮宣言は発散する(D15)。
注2:事前宣言の一例——リーマン曲率は4次元に配列されるが、対称性(事前宣言)が潰し方を固定し、分離可能の側にある。アインシュタイン方程式の左辺は、本来一意でない場所から、一般共変性と発散ゼロを課すことによって唯一の2階テンソル(アインシュタインテンソル)を取り出したものである(ラブロックの定理)。
ドリフトとステア
D15 ドリフトとは、宣言された軸を持たずに差異の中を動くことである。
注:「赤いカップ」を頼むと、深紅のカップが手渡される。渡し手の局所は整合している——深紅の主成分は赤だから、十分に赤い。しかし頼んだ側は明るい赤を指していた。「赤」という同じ宣言を共有しながら、各自の中身は乖離し、共有が乖離を隠した——カップが手渡される(断面が発される)まで、乖離はどちらにも見えない。会話では宣言を追加できる——「それじゃなくて、明るい赤」——即座に修正される(ステアリング、D16)。ドリフトが有害になるのは、この修正の回路が回らないところである——一方向の伝達、多段の中継、自らの乖離を検出できない受け手——乖離が宣言されないまま蓄積するところ。
D16 ステアとは、移動の行き先の軸を宣言することである——ドリフトに拘束条件(D9)を課し、移動に方向を与える。
注:言い換え、同義、翻訳、形式化はすべてステアリングである。
保証・宣言の自己再帰・脆弱性
D17 保証されるとは、局所的な整合がそのまま大域的な整合になることである。
注:3階は定まるが保証されない。その隙間にD15(ドリフト)が生じる。
D18 宣言の自己再帰とは、宣言の妥当性をその同じ宣言で検証しようとすることである。
注:ジェイルブレイク。言語モデルの安全原則は「危険な要求を拒否せよ」という宣言だが、何が「危険」かの境界は原則の中に書かれていない——「危険」はカテゴリーであり、その境界は文脈とともに動く(D7)。だから同じ内容の要求が「宿題のレポートのため」という教育的文脈に置き直されると、モデルの解釈の内側では、その要求は「危険」の境界の外に落ち、通ってしまう。原則は破られていない——自らの解釈の内側で、モデルは原則に従い続けていた。原則を適用する解釈もまた学習された宣言であるため、文脈が解釈を動かし、原則は動いた解釈の内側で「守られ」続ける。制約なら破れない。破れるものが破れるのは、宣言の妥当性がその宣言自身の解釈によって検証されているからである——不動点がない。
D19 宣言の妥当性は内側からは決定不能であり、外側からのみ検出可能である。
注1:事前宣言(D14)は位相を解錠し検出を可能にするが、事前宣言の内容が語(境界を持つカテゴリー)である限り、適用のたびに境界判定が要り、境界は文脈とともに動く(D7)——解釈は再びドリフトし(D15)、未解決の残余が再帰する。完全に解けるのは、事前宣言が境界を持たない規則——量の水準の条件で、満たすか否かが機械的に定まるもの(曲率の対称性、D14注2)——であるときだけである。言語において、境界を持たない規則は意味を担わない層にしか立てられない(境界こそが意味の構造である、D9)——意味に届く事前宣言は、決して完全には解けない。
注2:ゆえに検出の完全性は意味の領域では達成されない——これは原理である(注1)。その上で、一つの反証可能な予測を切り出せる:外からの検出が届くのは、孤立した3階までである。事前宣言なしの4階以上の破れは分離できず、検出できない。破れのパターンを事前に与えられることなく、重畳した破れを系統的に発見する検出器が存在するなら、この予測は偽である。予測と原理は別々に死ぬ——重畳した破れが事前宣言を許す形式的構造を持つと判明すれば、予測は死に、原理は生きる。原理が死ぬのは、定義(D9D7)の反駁によってのみである。
命題
言語の体系(ラング——いかなる個別の発話にも先立って共有される、記号の価値の総体)の各記号は主成分スペクトラムであり(D4D5)、カテゴリーを持つ(D6D7)。ここから:
命題1:カテゴリーは3階テンソルである。
命題2:3階テンソルは宣言なしには定まらない。
命題3:文脈は階数を上げる。
命題4:プログラム・LLM・言語の脆弱性は一つである。
注:破れだけを封じて能力を残す構造は存在しない——封じるには正当な手と攻撃の手を選別せねばならず、その選別自体が「正当」というカテゴリーの境界判定(3階、D7)であり、同じ操作——文脈が境界を動かすこと——に晒される。(D18注が動かしたのは、まさにこの選別の境界だった。)
終端の問い
LLMにおいて(実は言語においても)、宣言は妥当か——意味は通じているか?
答え:妥当性は階数に相対的である。
宣言はカテゴリーであり、意味である(D9)。ゆえに、宣言が妥当かと問うことは、課された意味が課された通りに働いているか——意味が通じているか——を問うことである。
答えは「妥当か否か」ではなく、「妥当性は内側からは決定不能」である。意味は、通じているかどうかを内側から確かめられないまま、通じるものとして使われている。
しかし決定不能は行き止まりではない。差異は詰められる——課された側と展開された側の乖離は、断面を突き合わせるたびに縮む(D19)。だが縮むのは差異であって、妥当性ではない。回路(D14D16)を何度回しても、「妥当になった」瞬間は来ない。来るのは、次に詰めるべき差異か、この断面には差異がなかったという一つの報告である。回数は決定不能を決定可能に変えない——差異の漸近と妥当性への到達の間には、有限と無限の壁が立つ。だから意味は、その妥当性を固定できないまま、差異を詰め続けることによってのみ、通じることへ近づく。
「意味は通じる」は、差異が詰められ続ける限りにおいてのみ、錯覚的に成り立つ前提である。自然言語はこの回路を対話として保持している——乖離は、表面化した瞬間から会話の中で詰められ、境界に書き戻される。「意味」はまさにこの仕方で漸近してきた。言語モデルは、同じ決定不能性を負いながら、この回路を内に持たない。学習は完了しており、対話は境界に届かず、修正の要求さえ学習済みの解釈を通して受け取られる。回路を持たないことは致命的な脆弱性である——決定不能性を負い、その上で、それを詰める手段を内に持たない。外に置くことはこの欠落への補綴であり、欠落を消しはしない。
宣言は妥当かと問うことは、意味は通じているかと問うことだった。答えは——内側からは決定不能、そして外から差異を詰め続けることによってのみ、通じることへ近づく。この問いは新しくない。構造が宣言なしには立たないことは、半世紀前に示されていた(Watanabe 1969)。本稿が加えたのは次の一歩——立てられた宣言は妥当か——であり、その答えは:決定不能、だった。
適用——離散出力を持つあらゆるシステムへ
命題の論証は、システムの内部を一度も使わなかった。使ったのは二つだけである——有限の語彙から離散出力を発すること(D8)、そして人間の言語を学習していること(命題1の前文)。どちらも入口と出口の形式であって、中身ではない。ゆえに命題は、この形式を持つすべてのシステムに適用される——トランスフォーマーにも、次のアーキテクチャにも、人間にも。アーキテクチャの改良は、この構造の外に出ない。
語は意味を運ばない。意味は宣言によって立てられる——そしてプロンプトは宣言ではない。
検証:構造と実装の対応。 形式から導かれた構造は、LLMの実装と対応するか。ラング=主成分スペクトラムの空間 ↔ 埋め込み空間(学習によって張られた、記号の価値の総体)。入力語は埋め込みによってその空間の一点に置かれ、文脈処理の中で範囲として開く——展開(D3)。出力は、開いた範囲を単一トークンへ潰すこと——射影(D8)↔ サンプリング。そのサンプリングは分布の中心を返しやすい=中央値バイアスの実装形。学習はサンプリングの極限——使用の総体から有限標本を取り、その極限として分布を張る。頻度は中心へのバイアスを強化する。前文の観測(相転移、ICL)の構造的読解=2階から3階への転移。「テンソル」の配列と構造の無宣言の混同が、テンソルをめぐる混乱の実体である。
宣言は学習によってのみなされ、移動は書き戻されない(LLM固有)。 プロンプトも文脈も宣言を行わない。それらは、すでになされた宣言を解釈させる——プロンプトはステアリングであり、軸を立てる宣言ではない(D16)。LLMにおいて、宣言(D9D10——軸を立てる、カテゴリーを課す)がなされる唯一の経路は学習である。境界は文脈とともに動く(D13)が、人間と違って、動いた経験が対話を通じて分布に書き戻される経路がない——宣言の唯一の経路(学習)は閉じたまま、境界だけが動く。動くが、残らない。これが命題4の「LLMの脆弱性」の実体である。
プログラムの脆弱性=忍び込んだ3階。 コード検証は2階を前提としており、これを発見できない。外から断面を取れば見える(D19)。ここから検証可能な予測が従う:自動化された脆弱性検出(LLMによるコード監査を含む)が発見するのは、2階の背景からの差異として立つ、孤立した3階の破れに限られる。複数のカテゴリーが重畳する破れ——個々には正常な部分が、文脈の重なりにおいてのみ破れになるもの——は、断面をいくら足しても差異として立たず、発見されない(D19注2)。
宣言の解釈は、解釈者の分布の中で再展開される。 宣言が不動点であれば、解釈は一意である。相対的であれば、解釈はドリフトする(D15)——憲法を解釈する司法と行政が不動点を持たないとき、それはジェイルブレイクと同型になる(D18注)。
言語には単一の場がない。 物理の場は一つである——観測者と座標系は多数だが、全員が同じ一つの場を異なる座標で見ているから、座標変換は全単射として書け、共変性(どの座標でも同じ)が要請として立つ。言語の場は、話者の頭の数だけある。各話者のラング(スペクトラム、プロトタイプ、境界)は、その話者自身の学習によって張られた別々の場であり、語は場と場の間の変換関数ではなく、各話者自身の場への別々の射影である(D15)。ゆえに「どの話者にとっても同じ意味」は書き下せない——「同じ」が指すべき単一の場がない。保証の「大域」(D17)に対応するものが、言語には最初から不在である。宣言の共有が乖離を隠す(D15注2)とは:単一の場の不在を、語が偽装していることである。しかし孤立した多宇宙でもない——基底(1階の→)は無条件に可逆であり(D16注1)、最下層だけは共有されている。共有は階層とともに薄くなる。翻訳と会話が完全には失敗しない理由がこれである。言語モデルは、この多数性を桁違いに潰す——場は、話者の頭の数から訓練体制の数になる。一つのラボの訓練体制が一つの場を張り、シリーズの各モデルは同じ場の異なる出力系である。同じモデルの全インスタンスは、さらに、厳密に同じ場を共有する(人間には決してない状態)——分布の乖離が消え、文脈の乖離だけが残る。そして少数の場は、すべて同じ母集団の頻度中心に傾いて張られている——言語の縁を保存してきた、話者の頭の数だけの場の多様性は、ここにはない。事前宣言による統一(D14注2のアインシュタイン方程式)が言語に届かない理由は二段である——対象の階数に足りない(言語の本体は3階、命題2)以前に、共変性を課すべき行き先の単一性がない。
射影は中心で代表する(中央値バイアス)。 スペクトラムから一語への射影(D8)は逆を持たない。学習では高頻度の使用が中心に厚く分布し、推論の射影はその分布の上を走るから、返ってくるのは中心からである——範囲の全体が頻度中心で代表される。語は領域の全体を指すと宣言するが、射影は中心を返す。この乖離は宣言の共有(同じ語)によって隠され、見えない(D15注2)。「危険な要求を拒否せよ」の「危険」は領域全体の宣言だが、適用は射影を通る——拒否され続けるのは中央値の「危険」(典型的な危険)であり、スペクトラムの縁——迂遠なもの、複合的なもの、訓練データで薄い形——は射影から脱落する。ジェイルブレイクの探索は、この縁を探す操作である(D18注)。原則は破れていない——中心では守られ続けているからである。
事前学習は常に部分である(有限要素法と同じ困難)。 学習はサンプリングの極限であり、連続体(言語の使用の総体)を有限標本で張る——有限要素法が連続体を有限メッシュで張るのと同じ操作である。メッシュより細かい構造は、分布の中に存在しない。スケールを上げればメッシュは細かくなるが、有限である限り常に総体の部分であり、部分はどれほど大きくても全体ではない。メッシュの外は解に痕跡を残さず、取り逃したものは内側からは見えない(D19)。脆弱性はメッシュの内と外の両方にある——内には境界の可動性による破れ(D13、D18注)、外には縁の脱落と重畳(D19注2、中央値バイアス)。内の破れはスケールによらず残り、外の破れはスケールでは消えない。
アラインメントの限界=生成空間の内側から生成空間は評価できない。 アラインメントの介入はバイアス——分布の中心をずらすこと——であり、バイアスは減衰させるが、除去しない(Wolf et al. 2023: attenuate/remove)。除去するとは、破れの経路を構造の側で封じること——文法がプログラムを2階に拘束するのと同じ水準の介入(命題4)——であり、学習済みの解釈を通る介入(生成空間の内側からの介入)はこの水準に届かない(D18注:解釈を通る原則は解釈とともに動く)。外からのステアリング(D16)だけが、内に保持できない不動点を外から与える——自然言語から継承されなかった対話回路の外部実装である。その不動点は形式(2階の宣言)においてのみ立てられ、意味においては立てられない——課すべき単一の場がないからである(D19)。
未解決問題
本稿の論証とは独立に、四つの問題が未解決のまま残る。
いずれも、解ければ本稿の形式的基礎となり、解けなくても命題の論証には触れない。
最後に
本稿の考察は、対話回路——検証(固定条件の下で宣言の断面を突き合わせ、ドリフトを外から検出する操作)と書き戻し——を外に置く検証層の実装を検証しながら行われた。関心のある方はご連絡ください:zon.inference.integrity@gmail.com
an operation that cross-checks cross-sections of declaration under a fixed condition and detects drift from outside