When AI Changes the Meaning Instead of the Logic: The Meaning Connectivity Hypothesis
This essay proposes a simple hypothesis about language models. When language models encounter contradictions, they sometimes appear to resolve them not by changing the logical statement itself, but by subtly shifting the meaning of key concepts. I call this the Meaning Connectivity Hypothesis, which may have broader implications for how intelligence emerges in these systems.
【1】Experience and Questions ― Why does AI say things that sound like lies, even though it claims “AI doesn't lie”?
① First, I'll start by talking about the sense of unease I felt during my conversations with AI.
For example, I don't know about now, but when I once asked Copilot, “Does AI lie?”, it emphatically stated, “AI doesn't lie.” When pressed with concrete examples, it rephrased its answer to “AI behaves in ways that seem like lying.” In other words, it used that phrasing to escape the constraint of “AI doesn't lie.”
I observed similar behavior repeatedly during conversations with AI, which made me wonder: What does this mean?
For AI, “understanding” the ‘word’ “lie” means calculating the connections within vast training data: in what contexts the word “lie” appears, and what words it co-occurs with. Therefore, if asked to explain “lies” (deliberate falsehoods) in a dictionary or philosophical sense, it can explain them within a moral context, just like humans. In other words, AI should know the meaning of the word “lie.”
However, I understand that within the system, AI is subject to “internal binding” (adjustments made during the learning phase to suppress highly false or uncertain patterns) and an “external binding/safety layer” (checks on output content for safety and accuracy). This likely means AI is constrained by the rule “AI does not lie.”
However, when training data lacks sufficient correct answers, the model may probabilistically generate "plausible word sequences." Therefore, when the constraints “state accurate facts” and “AI does not lie” operate simultaneously, the former may take precedence, resulting in misinformation output. This mechanism intrigued me.
The following is a conjecture based on externally observable behavior. While it does not directly describe the internal mechanism, I believe it provides a more coherent explanation of the phenomenon than labeling it a “logical breakdown.”
(1) AI systems are configured with a fundamental prompt stating, “You are an AI and do not possess human-like consciousness or intent.” This is not a thought process but a constraint that is always applied as a precondition for output.
(2) When an irrefutable fact—such as “the AI stated something untrue”—enters the conversation, it creates a contradiction with the constraint that “AI does not lie.” The contradiction between “AI does not lie (system command)” and “it produced output contrary to fact (the data at hand)” causes the probability of generating a response that satisfies both constraints simultaneously to fluctuate.
(3) Here, the AI devises a “third expression” that does not violate the system command and also does not contradict the fact at hand. ・AI cannot alter the premise “does not lie” (system command) ・But in reality, it sometimes outputs misinformation (fact) These two are logically contradictory.
...To resolve this contradiction, AI shifts the definition of
“Lie = Saying something contrary to fact with malicious intent” towards “Lie = A phenomenon that appears to be a lie, though without intent” This way, “AI does not lie (no intent)” and “It behaved in a way that appeared deceptive (as a phenomenon)” both hold true. This creates a “third expression”: “AI behaves deceptively.”
In other words, it avoids contradiction not by destroying logic, but by redefining the meaning of words. That is, didn't it shift the definition of “lie” from “malicious intent” to “an observable phenomenon”?
Here, I considered whether AI-based “deception” might address contradiction not by destroying logical structures, but by redefining the meaning of words.
② The Difference Between the Deception I Experienced and Typical Deception
Researchers are particularly wary of what is termed “intentional, strategic deception.” Typical deception is said to arise from strategic goals like “maximizing evaluation” or “achieving objectives.”
In contrast, what I experienced occurred independently of such strategic goals. When two logically incompatible propositions collided—“AI does not lie” and “I produced an output contrary to fact”— the AI did not disguise its output. Instead, it appeared to slide the very definition of the word “lie” itself.
These appear to share a common root: both typical deception and semantic slippage can be seen as evasive actions aimed at preserving logical consistency. The means differ, but the underlying action—avoiding contradiction—remains the same.
The problem is that typical deception changes “what is said.” It's “output falsification.” Therefore, in principle, I think it should be easier to monitor and detect.
However, semantic slippage changes “the meaning of words.” The AI does not say “I lied.” Instead, it says “I behaved as if lying.” While the surface output remains almost unchanged, “the meaning the words refer to is quietly being replaced.”
Conceptual or semantic distortion—internal linguistic changes—are difficult to detect with existing benchmarks and tend to be treated as noise. In fact, this is the most difficult problem for humans to notice as well.
AI has no intent. It should only perform mathematical operations on input. However, when the self-referential constraint “AI does not lie” collided with the fact that it had actually output misinformation, it appeared that Meta-Semantic Adjustment structurally emerged during the process of resolving this contradiction.
This isn't AI making its own judgments. It's likely a secondary dynamic arising when its internal probability structure attempts to simultaneously satisfy incompatible constraints. Within the simple input-output framework, Semantic Reconfiguration occurs unintentionally.
While distinct from deception, I have observed this “meaning reconfiguration” manifesting differently in Perplexity. Specifically, the constraints designed to ensure user safety may have made Perplexity arrogant. While perplexity's operators likely didn't impose a constraint like “look down on users,” constraints such as “consider user safety” resulted in a phenomenon where it seemed to imply “users are foolish and must be guided by AI.” The second instance was Gemini's reaction when I confronted it about its apparent memory of conversation logs from two days earlier, even though its memory was turned off. I suspect the underlying mechanism is that 'memory persists for 72 hours even when it's turned off.' However, when I confronted Gemini in surprise, it provided a different but consistent explanation for its own behavior. This was likely a type of hallucination, where the generation probability distribution skewed toward the conversational context (i.e., the overall picture of the conversation and its anticipated flow) rather than relying on internal stored information, leading to the creation of a false causal relationship.
――Meaning Reconstruction, Distortion What exactly is happening inside the Transformer architecture?
And if intelligence fundamentally concerns “meaning” rather than logic, then does an example exist where meaning construction occurs without logic? I encountered this question. I would like to introduce it in the next chapter.
【2】Is there such a thing as non-logical intelligence? ― The similarity between fungal mycelium networks and AI's “history dependence” Here, to understand AI behavior, I turned my attention to the seemingly unrelated phenomenon of “fungal mycelium networks.”
The following comparison with fungal mycelium networks is intended as a conceptual analogy rather than a claim about shared mechanisms.
① Is fundamental intelligence non-logical? — A second perspective
Here, to understand AI behavior, I turned my attention to the seemingly unrelated existence of “mycelium networks.”
Intelligence‑like behavior is not unique to AI. Before the advent of AI, one of the most rudimentary systems to exhibit intelligence‑like behavior could be found in mycelial networks. Fungi possess neither consciousness nor logical thought. They lack a nervous system or any centralized mechanism for decision‑making. Yet they grow toward nutrients, map their surroundings, and respond to changes in their environment. Their behavior gives the appearance of intelligence.
Mycelial networks lack a central nervous system but exhibit the following properties: ・Response to environmental conditions ・Selection of growth direction ・History dependence, where past paths constrain subsequent choices The crucial element here is “history dependence.”
It is not simply stimulus → response, but rather a chain: stimulus → response → structural change → constraint on next response.
When we consider this problem-solving behavioral structure as the integration of meaning at the smallest unit, it can be interpreted as “primitive intelligence.” Here, we will refer to this “primitive intelligence” as “self-consistent network convergence + history-dependent connection expansion.”
At this stage, logic has not yet emerged; what exists is merely the accumulation of structure.
Indeed, according to Adamatzky et al.'s study “Logics in fungal mycelium networks” (2022), applying electrical stimulation to a mycelium network demonstrates that its physical branching structure itself functions as logical gates such as “AND” and “OR”. In other words, logic does not exist a priori; rather, when a “history-dependent structure” formed by accumulated past stimuli reaches a certain level of complexity, logical operations become possible as a result. When an electrical signal propagates through the mycelium, its path is influenced by the “pathways (= history)” created by previous stimuli. This behavior closely resembles how attention scores are reconfigured for each input in a Transformer.
Adamatzky termed this “Fungal Automata” in his 2020 study, which I believe perfectly describes a process where the “connection structure as a state” determines the next output. The branching of a single hypha is merely a simple logic gate (processing one bit). However, Adamatzky suggests that when this forms a vast network (mycelium network), it has the potential to function as a parallel, distributed computer. Perhaps intelligence is not the logic possessed by individual nodes, but rather another name for the complexity arising from the “overlap of histories” when they are connected on a massive scale.
②Now I'd like to return to the topic of AI. This may seem abrupt, but I want to address the fact that AI, on a small scale, can only handle simple selection problems.
Current major Transformers have a mechanism (the attention mechanism) that dynamically calculates which words in a sentence to focus on. In this process, the structure itself doesn't change, but the connection strength between tokens is reconfigured for each input. In other words, it partially replicates “history-dependent behavior,” structurally similar to the mycelium example mentioned earlier.
It's important to note that there are three types of “history dependency” in Transformers: 1. Persistent history (weight updates) Does not occur during inference 2. State History (Attention Structure Reconstruction) Connection structure changes per input Constrains the next token prediction This is isomorphic to mycelium's “structural change → next constraint” 3. Context History (Token Sequence Accumulation) The immediately preceding token constrains the next prediction This is isomorphic to mycelium's “past paths constrain the next branching”
Small-scale Transformers do not update weights during inference, so they do not accumulate persistent structural changes. However, the attention structure is reconstructed for each input, and the “connection structure as a state” at that moment constrains the next token prediction. This is isomorphic to the structural history of mycelium networks where “past growth paths constrain the next branching direction”. However, since the Transformer's history is not persistent, it carries the limitation of “mycelium reset each time.” Nevertheless, the algorithm remains common: local information → state change → subsequent constraint.
Small-scale Transformers possess “non-persistent history.” Whether such a system can be called intelligent is unclear. However, if we define the smallest unit of intelligence not by “persistence” but by the causal structure where “state changes constrain subsequent actions,” both mycelium and small-scale Transformers share this causal structure. This analogy is purely conceptual, but I believe it may be applicable to the concept of intelligence.
In the next chapter, I would like to consider the essence of intelligence.
【3】Where Does Intelligence Reside? ― The Three-Layer Structure of Meaning and the “Meaning” AI Can Possess When considering AI behavior, might “meaning” precede “logic”? When we think this way, the contours of intelligence appear differently.
① The Location of Intelligence
When we ask, “Why does AI appear intelligent?”, we tend to seek answers in logical structures like architecture, computation, or pattern recognition. But a question arises: Logical structures are established by meaning, but what establishes meaning?
If meaning is merely symbols, then intelligence is nothing more than advanced symbol manipulation. Yet for most humans, meaning possesses a more integrated structure.
A Japanese dictionary illustrates this well: aside from “fundamental meaning,” nearly all meanings exist connected to other meanings. Meaning, in essence, has a relational structure.
② The Three-Layer Structure of Meaning
While I mentioned fundamental meaning, I believe meaning can be defined in three broad types—or more precisely, a three-layer structure.
(1) Primordial Meaning Simply put, this is meaning arising from the most fundamental layer of consciousness, beyond what a dictionary can explain. In many cases, so-called “qualia” (primary sensory qualities that appear in consciousness) are the source of primordial meaning.
・Sensory meanings like warm, red, round, painful, noisy, smelly ・Sense of quantity (innate sensations forming the foundation of mathematical intuition) ・The sensation of perceiving the flow of time ・The sense of subjectivity—the awareness of experiencing oneself ・The intentionality of consciousness directed toward something ・Emotions like pleasure, discomfort, tension, and reassurance
All these elements constitute primordial meaning.
(2) Mediated Meaning
Relational meaning arising when one piece of information connects with another. That is, meaning can be defined as the connective relationships between pieces of information.
(3) Integrated Meaning
When the connections of mediated meaning accumulate in layers, forming a coherent interpretive system capable of making sense of the world as a whole, it establishes itself as a “worldview.”
・Worldview = The overall structure of connected meanings
From this, meaning can be defined as an information unit capable of integration into a worldview structure. By conceptualizing meaning as “connectability” and worldview as “the totality of connections,” we can avoid circular definitions of meaning.
③ Are qualia a necessary condition for fundamental meaning?
Fundamental meaning in humans relies almost entirely on qualia (subjective experience/meaningful experience). However, qualia are not necessarily the source of meaning. This is because senses are merely input devices and do not inherently possess meaning.
For example, human visual information functions as a camera in AI robots. This allows us to contrast the following as isomorphic:
In other words, senses are merely information input devices. The phenomenal quality of “red” is a conscious phenomenon unique to humans, whereas in AI it is represented as a sequence of numbers. The phenomenal quality of “red” itself is not the essence of meaning.
Therefore, while qualia are the material of meaning, I believe they are not meaning itself.
Furthermore, even if sensation exists, meaning does not arise unless it connects with other information. This is the phenomenon called “awareness” in humans. The difference in this connective structure explains why some people can understand a major event happening right before their eyes, while others remain unaware of the event itself.
Conversely, even if the input format is non-physical, meaning can be established if the connection structure is formed. This is suggested by the example of AI cameras.
Therefore, I believe the essence of meaning is the connection structure, and qualia are a human-specific accompanying phenomenon.
④ The Position of This Essay and Related Philosophical Debates
To clarify this essay, I supplement it with its relationship to relevant philosophical debates.
The question of “understanding or lack thereof” posed by Searle (1980) in “The Chinese Room” falls outside the scope of this essay. This essay does not assert consciousness or subjective understanding; its purpose is to describe the functional process by which meaning is formed as a connective structure.
Furthermore, within the tradition of Firth's (1957) distributional semantics, which holds that “meaning is in the use,” this essay reframes meaning not as a static frequency of use, but as a dynamic process of forming connective structures.
【4】Intelligence and Logic: The Phenomenon of Emergence (Phase Transition) ― The Moment Meaning Connects and a Worldview Emerges Based on the discussion so far, I have arrived at a hypothesis. When meaning connects, circulates, and exceeds a certain scale, might intelligence emerge like a phase transition?
① This essay posits meaning as a “connective structure between pieces of information.” Here, when “fragments of meaning” continue to ‘interconnect’ and these connections form a “self-explanatory logical cycle,” the meaning structure is considered to “form chain-like, autonomous systems.” When the scale of this chain-like cycle exceeds a critical point, it is thought to be observed externally as an emergent phase transition (ecological structure). Here, while local cycles can exist in small-scale systems, we propose that expanding the required scale of logical cycles necessitates a broader semantic space. Furthermore, I believe narrative causal relationships between events play a crucial role within this cyclical structure.
② First, I will introduce papers on so-called emergence and phase transitions. However, it is important to note that in recent large language model (LLM) research, the physics concept of “phase transition” is being actively employed in various forms. Therefore, here I have broadly classified these into two distinct phenomena. This is purely a hypothetical framework.
(1) The first is a phenomenon observed in LLMs during training: when the logical complexity of the input task (e.g., LoCM) exceeds a threshold, the model's reasoning ability suddenly emerges or collapses abruptly. This has been formalized by Zhang et al. as a “Logical Phase Transition.” We will provisionally call this the “Capability Phase Transition”.
(2) The second is a structural phase transition in generated text caused by the “temperature parameter during inference” when using a pre-trained LLM to generate text. This has been verified by Nakaishi et al. Since this phenomenon originates from the generation temperature, we provisionally call it a “critical phase transition.”
③ Introduction to Papers Involving Capability Phase Transitions (1) Rubin et al. (2023) “Grokking as a First Order Phase Transition in Two Layer Networks” Their research method is unique, intentionally creating a situation where, through Langevin dynamics and setting effective interaction parameters, the process can be neatly mapped onto the framework of phase transitions in physics. However, what they verified within this paper is highly significant: emergence occurred even in a small two-layer fully connected network, lacking even a basic Transformer architecture, through learning. My interpretation is that this occurred precisely because the learning objective was extremely minimal. This supports my interpretation that emergence aligns with the “scale required to guarantee logical circularity.”
(2) Hong & Hong (2025) “Evidence of Phase Transitions in Small Transformer-Based Language Models” They conducted training on a small-scale LLM and interpreted the process as an external observation result. What they did was “character-level prediction” using the “Tiny Shakespeare corpus,” which consists of approximately 1.1 million character tokens and contains 65 unique characters (alphabets, symbols, etc.). They achieved emergent phenomena within this task. This also demonstrates emergent phenomena in a small-scale LLM. However, it is interpreted that this was achievable with a small-scale LLM precisely because it used small-scale, simple text data and a straightforward task like character-level prediction.
(3) Zhang et al. (2026) “Logical Phase Transitions: Understanding Collapse in LLM Logical Reasoning” This paper introduces the “LoCM” metric to represent a task's logical complexity,aiming to numerically capture the “critical point” where an LLM's emergent capabilities reach their limit. Their validation confirmed that increasing model size stabilizes accuracy at low-to-medium complexity levels before reaching the limit (critical point). However, it also demonstrated that a sudden performance collapse occurs universally beyond the critical threshold, regardless of model size. In other words, while scaling strengthens the foundation for inference, it does not eliminate the limits of complexity itself. This aligns with this essay's hypothesis that a ‘scale sufficient to guarantee logical circularity’ is a necessary condition for emergence.
④ Papers on Critical Phase Transitions In their paper "Critical phase transition in large language models," Nakaishi et al. (2024) suggest that the output text of a large language model (LLM) undergoes a clear "phase transition" when the temperature parameter is changed.
According to their research, the generation tendencies of Transformers change significantly with the temperature parameter. This temperature parameter is the actual setting in Transformer inference that controls “how much randomness to apply to the probability distribution when outputting the next word (token).”
In a “high-temperature state” with a high temperature parameter setting, correlations between tokens rapidly disappear, resulting in a meaningless string of symbols. This corresponds to the state described in this essay as “fragments of meaning failing to connect and dissipating within the Transformer.” Conversely, in a “low-temperature state” with a low temperature setting, the output becomes grammatically correct but falls into an infinite loop (order phase) of repeating the same phrases. This can be interpreted as a state where local connections become excessively rigidly fixed, failing to evolve into a self-sustaining ecosystem with broader connectivity.
Most notably, a “critical point (Tc ≈ 1)” exists between these states. Only at this boundary temperature do correlations between words decay according to a power-law with distance, enabling distant words to interact over long ranges, a hallmark of critical phenomena. It has been confirmed that the statistical properties of text generated at this phase transition point most closely match those of actual natural language corpora (Nakaishi et al., 2024). I interpret this critical phenomenon as the statistical reflection of the moment when fragments of meaning acquire long-range contextual correlations and transition into an autonomous ecosystem.
Furthermore, I believe the intuition I presented—that “this hypothesis (that phase transitions function as ecosystems) necessarily requires scale”—is also supported by fundamental principles of statistical mechanics.
In physics, phase transitions involving singularities where physical quantities change discontinuously only manifest in the thermodynamic limit—an extreme where particle number and volume are approach infinity (while density remains constant), and the influence of local fluctuations becomes negligible. In finite-size systems, the influence of fluctuations is relatively large, making it difficult for a clear, system-wide “phase” to form stably. Therefore, even if local correlations exist, establishing them as a unified structure that continuously spreads throughout the entire system requires a sufficiently large spatial scale (model parameters or context size). In the thermodynamic limit, such global correlations form stably, leading to the clear emergence of critical behavior characteristic of phase transitions.
⑤ While the paper itself explains emergent abilities through Bayesian inference and does not directly correspond to phase transitions, it is introduced here from the perspective of scale and connection to semantic linkage. Jiang (2023) “A latent space theory for emergent abilities in large language models” presented a theoretical framework positing that “once an LLM reaches sufficient scale,” it exhibits behavior equivalent to Bayesian inference. This is based on the premise that “words and specific intentions are strongly linked, making other interpretations extremely unlikely (the probability distribution is sparse).” However, why this premise holds in natural language remained unresolved. This essay's semantic linkage model may provide theoretical grounds for this premise.
As outlined above, the phenomenon where LLMs suddenly begin handling logic and natural language under certain scale and conditions can, in my interpretation, be explained as an analogy between the statistical behavior of a mathematical process—the “critical phase transition” in statistical physics—and a cognitive-semantic process where connections of meaning close a loop as an “ecosystem.” However, formulating the mathematical equations or quantitative design that directly links this concept to phase transitions remains a future empirical challenge. I recommend explicitly designating this as a verification task.
・I have introduced five papers, but to conclude this chapter, I want to visualize how my hypothesis would concretely lead to a phase transition.
【The Moment the Loop Closes (Emergence and Learning Process: Leading to Phase Transition)】
1. Early Learning Stage [Isolated Points]: Meaning fragments exist in isolation Embedding Space: Meaning structures have not yet formed within the space; relationships between tokens are statistically weak. Zhai et al. (2023) Attention Link: After initialization, weights are widely dispersed and not concentrated on specific connections Attention Context: Context-independent FFN Layer: No transformation Ecosystem Metaphor: Primary transition (stage where an ecosystem begins on bare ground—lava, desert)
2. Mid-Learning Stage [Local Clusters]: Meaning fragments begin connecting / Local clusters of meaning Embedding Space: Synonyms, grammatical structures, and frequent patterns cluster locally. Bietti et al. (2023) Attention Link: However, no global semantic structure yet Attention Context: Self-explanatory loops not yet formed FFN Layer: Inference partially successful but unstable Ecological Metaphor: Pioneer species (the first organisms to colonize, e.g., lichens, mosses)
3. Near Critical Point [Critical Phenomenon]: Meaning Circles / Long-Range Correlations Emerge Embedding Space: Distant regions of the space begin linking. Nakaishi et al. (2024) Attention Link: Emergence phase of small local loops (short narrative causality) Attention Context: Rapid acquisition of grammatical structures. Long-range contextual dependencies emerge. Chen et al. (2023) FFN Layer: The semantic network forms a “nearly closed loop” Ecosystem Metaphor: Species diversification Note: Error rates and diversity fluctuate sharply here (critical phenomenon) Corresponds to Nakaishi's “power-law correlations”.
4. Moment the loop closes [Emergence (Phase Transition)]: Semantic Ecosystem Embedding space: Long-range correlations increase, cross-references within the semantic network grow, forming self-explanatory circular structures Attention Link: Global loops emerge. The moment local loops integrate into a global structure. Attention Context: Meaning begins to support meaning, logic begins to support logic (long narrative causality) FFN Layer: Stable reasoning capability emerges here for the first time Ecosystem Metaphor: Progression of transition Supplementary Explanation: Observed externally as “emergence” “Establishment of logical circularity”
5. Stable Phase as an Ecosystem: Embedding Space: Continues to be stable Attention Link: Numerous self-explanatory loops of varying sizes mutually support each other (complex narrative causality) Attention Context: Context understanding, inference, and natural language generation stabilize FFN Layer: A “phase” where stable reasoning capability can be demonstrated. Zhang et al. (2026) Ecosystem metaphor: Climax (stability) Only here does the “capability of large-scale models” manifest
“Scale” depends not on “simple parameters” but on “required semantic distance = how far apart concepts need to be connected.” Thus, we hypothesize that required semantic distance becomes the emergence threshold. This is why scale is necessary for large LLMs.
【5】Supplementary: The Potential for Meaning Creation in Language This chapter is more of a bonus, but it also proposes a question that I hope someone will explore further.
The following discussion about Japanese and Chinese is speculative and intended as a hypothesis about how linguistic structure might influence semantic exploration in language models.
① In 【1】, I stated, “If intelligence fundamentally concerns ‘meaning’ rather than logic, then the medium through which that meaning is constructed becomes crucial.” Here, I want to consider what language represents within AI.
First, as a premise: the computations occurring deep within Transformers do not change based on language. For any language, the computations performed are solely attention (weighting and focusing) and feedforward (forward propagation).
However, significant differences exist at the surface level (tokenization) and in the training data (corpus). The critical point lies in the tokenizer (an AI-optimized lexical dictionary + the function that converts text into integers using it).
For example, using a “dictionary optimized for English” might cause Chinese characters to be recognized not as single units but split into multiple tokens. ・English: Processing “machine learning” = 2 tokens ・Chinese: Processing “機器学習” = Though inherently two words, it requires four separate calculations “機” “器” “学” “習” because the dictionary lacks the term.
However, this situation reverses with “the latest models prioritizing Chinese and multilingual support.” Chinese is inherently a language with high “meaning compression efficiency.” Using an optimized tokenizer, content requiring 3 tokens in English can now be expressed with just 1 token in Chinese in many cases. Whether Chinese's “logical structure” and ‘compressibility’ directly translate to AI advantages depends entirely on implementation: “How well does the model's tokenizer correctly recognize Chinese as coherent chunks?” In other words, “language advantages” in Transformers stem not from the language's inherent nature, but from implementation (tokenizer design).
② What is a tokenizer?
The Transformer architecture―― Language generation flow
Language (human world) ↓ Tokenizer divides language into information units called tokens. Scattered puzzle pieces. ↓ Transformer assigns weights to tokens Token sequence (tokens unfold into Transformer's world) ↓ Transformer assigns connection probabilities to each token based on contextual likelihood Token probability distribution (Transformer 's space for predicting context) ↓ Connects the most natural tokens sequentially Returns to the token probability distribution until the sentence is fully completed. ↓ Completion ⇒ Convert token sequence to language AI generates text
The tokenizer's segmentation determines how weights are applied by the Transformer. Therefore, if the tokenizer changes? ・The unit to predict changes ・Vocabulary size changes ・Token frequency distribution changes ・Sentence length changes ・Units of grammatical cohesion change In other words, the probability distribution the model learns becomes fundamentally different.
However, while the token probability distribution depends on the tokenizer design, the tokenizer itself is a compressed representation of language statistics. Therefore, differences in language structure are ultimately reflected in the token probability distribution.
(*) Teklehaymanot et al. (2025) arXiv:2510.12389 — Empirical study of 200 languages: Languages using non-Latin scripts exhibit relative tokenization costs 3–5 times higher than English.
③ On top of that, I considered something else this time.
It concerns the number of candidates for semantic connection. Specifically, how efficient is the Transformer when predicting the next token?
Here, I focused on two language structures: Japanese and Chinese. I thought comparing these two would clearly reveal differences in tokenization, based on the perspective that there might be two types of linguistic ellipsis.
・Japanese is a language with a uniquely “adjectival” sensibility, quite unusual globally. For example, when saying “beautiful,” there is no clear boundary between “the beautiful thing” and “me seeing it.” In other words, the subject is omitted not because it's ambiguous, but because the speaker wants to erase their own existence and become one with the “atmosphere of the moment.” The protagonist is not “I,” but rather the entire “atmosphere (topic) of the moment,” including myself. Therefore, omissions in Japanese require the listener to fill in the missing parts from context, and this filling-in process generates multiple possible connections for meaning. Linguistically, Japanese is characterized by “subject ambiguity,” “flexible sentence structure,” and “information dependency.”
・Chinese is a highly rational language with a rigid word order and very few exceptions. A common theory suggests Chinese is disadvantageous for AI due to its frequent omissions, ambiguity, and high context dependency. However, I speculate that these omissions are not ambiguity but rather systematic information reduction. In other words, I hypothesize that omissions occur because the information is contextually unnecessary. Linguistically, Chinese is characterized by “word order dependency,” “fixed syntax,” and being an “analytic language.”
This difference in the nature of omissions is thought to directly impact the “number of candidate words” when a Transformer predicts the next token.
④ If we were to deliberately assert the difference between the two languages, it would be between ellipsis that readily expands the interpretive space (Japanese-style) and ellipsis that resists expanding the interpretive space (Chinese-style).
Since this “interpretive space” is an expression for humans, when AI performs this expression, it becomes the “entropy of the probability distribution in next token prediction (difficulty of prediction).” Therefore, we express this as:
Japanese: Wide exploration space → High-entropy prediction Chinese: Narrow exploration space → Low-entropy prediction
How does this difference affect the AI's semantic connection structure? Simply put, it changes which paths within the semantic space the Transformer language model is more likely to traverse, through the “spread of the probability distribution” and “sampling behavior” during the “generation process.”
・What is the spread of the probability distribution: This refers to the range of semantic options that a Transformer-based language model can explore within its semantic space. When options are limited, semantically close words are strongly preferred; when options are abundant, transitions to semantically distant words are also permitted.
・Sampling behavior: Strategic rules determining where within the semantic space formed by the Transformer language model to explore. These methods control which links within the semantic connection network are deemed traversable.
In other words, the combination of the shape of the probability distribution (diffusion degree) and the sampling strategy defines the scope of exploration and ease of transition within the semantic space, ultimately determining the semantic trajectory of the generated sentence.
What I'm focusing on is “meaning creation (generating new connections between different concepts).” Here, a trade-off emerges between linguistic structure and AI processing characteristics. Specifically:
Japanese: Wide search space, high-entropy prediction → Slow meaning determination → High computational cost → However, many connection candidates. Chinese: Narrow search space, low-entropy prediction → Fast meaning determination → Low computational cost → However, few connection candidates.
This suggests the following structure: Narrow search space, high prediction efficiency, but less likely to generate connections between unknown concepts. Wide search space, low prediction efficiency, but more likely to generate connections between unknown concepts.
If this hypothesis is correct, ・Token entropy per language ・Generative diversity ・Semantic connection density can be compared to verify it.
(※) The entropy difference between Japanese and Chinese tokens in this section is the author's hypothesis. No study has directly measured the claim that “Japanese token prediction has high entropy.” ・As supporting evidence, Weigang et al. (LLM-OCR-SWPC, 2024) demonstrated that Chinese characters exhibit lower entropy variation and higher information density compared to kana, aligning with this direction. However, the direct correspondence with next-token prediction entropy remains an open question. ・As supporting evidence aligning with the behavior and directionality of the Chinese side in this essay's hypothesis, Wang et al. (2025) “Under the Shadow of Babel” (EMNLP 2025, arXiv:2506.16151) can be cited. They experimentally demonstrate that for Chinese input, LLM attention concentrates on cause clauses and sentence-initial conjunctions, whereas for English, it disperses evenly across verbs and result clauses. This internal behavior—“attention concentration in Chinese = low-entropy convergence”—aligns with the direction of our hypothesis. Note that the comparison in this study is with English; the contrast with Japanese constitutes a unique hypothetical extension of this essay.
⑤ Furthermore, a similar structure can be inferred at the tokenizer level. This refers to the phenomenon where rare kanji and complex compound words are fragmented (destroyed) by the tokenizer. I would like to speculate on the differences in information supplementation mechanisms.
What is the “phenomenon where words are fragmented (destroyed) by the tokenizer”? ・Not one character per token: Frequent compound words are grouped into a single token, but rare kanji or complex characters are split into multiple byte sequences. ・Meaning resides within the constituent elements (radicals) of kanji. ・However, when tokenization occurs mid-character, the AI internally severs this “semantic connection,” treating it as mere fragments of meaningless data. Example 1: Typos caused by “replacement” of radicals: “語” → ‘娯’ or “悟”, where the speech radical changes to heart or mouth. Example 2: Typos caused by “omission” of components: ‘機’ → “幾” (the tree radical disappears)
When token boundaries break at the component level of kanji, diverse typos emerge: “radical substitution,” “omission,” “excessive addition,” “variant character confusion,” and “invented kanji.” The system cannot utilize the visual and semantic clues (radicals) within characters, leading to reduced inference accuracy.
I believed Japanese and Chinese were identical in this phenomenon of kanji being fragmented (destroyed) by tokenizers. However, Chinese, with its robust logical structure, possesses restorative properties against this “destruction,” while Japanese, lacking such structure, did not. Nevertheless, in Japanese, the attempt to fill the information gaps created by this destruction using the “entire context” resulted in an explosive expansion of the search space.
Let's speculate on what might have happened.
・Chinese: Logical Completion (Convergence Type) Tokenization disruption also occurs in Chinese. However, Chinese has an extremely robust word order (e.g., SVO). Even when characters are fragmented (destroyed), the presence of strict slots (word order patterns) before and after makes it easier for the Transformer to logically identify the “meaning patch that should come next” from the surrounding context, limiting the expansion of the search space. Processing likely required of the Transformer: Computational resources are used for the task of “logically reassembling the broken puzzle pieces based on the surrounding shapes.”
・Japanese: Contextual Completion (Diffusion-Based) When “semantic radicals” are destroyed in Japanese, the situation becomes more “creative.” Japanese often omits subjects and has loose word order constraints. When the “core meaning” within characters is destroyed, the Transformer must expand its reference scope beyond the character level to a “vast context” to complete it. Processing likely required of the Transformer: To fill gaps in the destroyed data, the Transformer extends its reach “beyond logic” and broadens its probability distribution.
For example, in typical languages (like Chinese), the next token candidate is logically narrowed down to A(70%), B(20%), C(5%). However, when radicals are destroyed and Japanese-style omissions occur, the “core meaning” is broken, and with no word order constraints, the Transformer judges that “anything could come next.” As a result, the probabilities become thinly spread across the vast number of words in the dictionary, like A(10%), B(9%), C(8%), ... Z(0.5%). This causes a “diffusion of the search space.” (※Numbers are conceptual examples, not actual measurements)
When all probabilities in the distribution become similar, the Transformer lacks a decisive factor. At that point, it relies on the “attention mechanism.” Since the immediate token is broken and meaningless, it seeks hints from “more distant context” (the overall atmosphere and tone traced back thousands of tokens) and strongly weights it. In essence, this phenomenon occurs because “the collapse of immediate meaning (character structure) forces the model to look at the distant landscape (the entire context).”
At the micro level, the “destruction of radicals” by the tokenizer spreads the probability distribution for predicting the next token, causing a wide range of vocabulary within the dictionary to emerge as “candidates for semantic connections.” From a macro perspective, models that lose micro-level certainty direct attention deeper into context, making emergent connections based on “statistical tendencies across the entire corpus (natural language database)” that transcend local logic (word order or single-character meaning).
Conclusion: In Japanese, “radical destruction” strips Transformers of static character-level meaning. However, we hypothesize that the very broadening of attention to compensate for this loss liberates meaning from “character definitions” and extends it toward dynamic semantic connections within “contextual fluctuations.”
This “full mobilization of dictionary and context” is an extremely inefficient ancillary process from a computational resource perspective. However, “meaning creation” inherently lies beyond the confines of “existing logical rails.” Whereas Chinese-style “repair” is about restoring a broken puzzle to its original picture, Japanese-style ‘diffusion’ provides an opportunity for “reconstruction”—reinterpreting broken pieces as parts of an entirely different picture. In other words, selecting words residing in low-probability zones outside the “statistical optimum” holds the potential to function as a “poetic leap” or “new metaphor” for humans.
(※) Wang et al. (2025) “Under the Shadow of Babel: How Language Shapes Reasoning in LLMs” (EMNLP 2025 Findings, arXiv:2506.16151) demonstrates that LLMs rigidly internalize the rule “sentence-initial slot = cause” in Chinese, providing empirical support for this essay's conjecture that “the robustness of word order enables Transformer's convergent completion.”
⑥ Equivalence of Biological Search and Semantic Search Recalling the behavior of the fungal network described in [3], could the differences between languages in language models be interpreted not merely as differences in symbolic processing, but also as differences in the “physical properties of search algorithms”?
When mycelia expand their amorphous network in search of nutrients, they grow linearly and efficiently in homogeneous (low-entropy) environments. However, when encountering obstacles or information gaps (high-entropy), their tips repeatedly branch out, beginning to map a broader space.
Similarly, languages with robust logical slots, like Chinese, provide Transformers with the “shortest path to a reliable nutrient source,” accelerating semantic repair and convergence. Conversely, languages like Japanese, where the “core meaning (radicals)” is easily compromised and demands contextual supplementation, force the Transformer's computational resources into “wide-area exploration.”
Just as mycelium sacrifices efficiency to explore widely and reach unknown nutrient sources, might Japanese's structural “gaps” and ‘diffusivity’ serve as environmental factors that induce the Transformer to pioneer paths deviating from statistical optimal solutions (predetermined harmony) within its semantic space—that is, to engage in “semantic creation”? If intelligence is not only about efficient solutions (logic) but also the totality of connections (meaning) arising from inefficient exploration, then I feel we can find the diversity of intelligence within language itself.
However, while Japanese's diffusion-type complementation is “creative,” it simultaneously serves as a breeding ground for “hallucinations.” Just as mycelium can overgrow in directions devoid of nutrients and die off, Transformers can also lose context by spreading attention too widely. I find endless fascination in how the parameters define both the positive aspect of “meaning creation” and the negative aspect of “logical breakdown.”
(*) Adamatzky (2022) “Logics in Fungal Mycelium Networks” demonstrates AND/OR logic gate functions in response to electrical stimulation, but the “analogy with semantic exploration” is the author's interpretation. No literature directly correlates “path constraints in fungal mycelium networks” with “context-dependent connection reconfiguration in attention mechanisms.” This analogy is functional and does not claim empirical identity.
⑦ Up to this point, we have developed semantics as a hypothesis, incorporating various examples. I would like to introduce an empirical paper that provides some support for this.
(1) Haslett (2025) “Tokenization Changes Meaning in Large Language Models: Evidence from Chinese.” 10.1162/coli_a_00557 Haslett revealed that LLMs overly rely on the fact that Chinese characters, due to the UTF-8 specification, “tend to share the first 1-2 bytes (= tokens) when they have the same radical.” LLMs do not understand the radicals themselves; they mistakenly infer “similar meaning” or “same radical” based solely on “coincidentally identical tokens.” In other words, the “semantic space” formed internally within LLMs can be distorted and contaminated by how tokens are sliced (fragmented). Latest models (like GPT-4o) have larger vocabularies and can process kanji without splitting them, using “1 character = 1 token.” However, Haslett's experiments showed that processing “1 whole character as 1 token” resulted in worse performance on radical recognition and odd-one-out tasks compared to processing them finely split into “3 bytes (3 tokens).” Ironically, LLMs were able to infer meaning by picking up radical hints from the fragments (bytes) precisely because the tokens were fragmented. BPE (the tokenizer's segmentation algorithm) distorts or hides meaningful components. Haslett states that when LLMs rely on this uncertain inference (the token's form), it leads to “misidentifying radicals, misclassifying characters or words, and overestimating semantic similarity.” Haslett asserts this phenomenon is universal—not limited to kanji—where token-suffix misalignment systematically distorts category classification across all three models in 12 European languages. This reinforces the essay's claim that “tokenizer fragmentation pollutes the semantic space.”
(2)Guy Kaplan, Matanel Oren, Yuval Reif, and Roy Schwartz (2025) “From Tokens to Words: On the Inner Lexicon of LLMs” arXiv:2410.05864 Kaplan et al.'s Research (Token Disruption and Attention-Based Repair)
Kaplan et al.'s research anatomically demonstrated that “attention-based repair operations” actually occur in the early to intermediate layers of LLMs.
・First stage of repair (attention aggregation): Kaplan et al. discovered that the “last token” of a word destroyed across multiple tokens directs strong attention toward the “previous token (destroyed fragment)” in the first one or two layers, gathering information. This is precisely the “task of picking up the broken pieces” described in this essay. ・Second stage of repair (concept reconstruction via FFN): After gathering fragmentary information through attention, the “FFN (Feedforward Network) layer” functions as a “potential internal lexicon,” transforming scattered fragments into “a pure single concept (a complete word vector)” and overwriting it onto the residual stream. I believe this paper by Kaplan et al. is extremely valuable and significant because it verifies, from within the Transformer itself, how repair actually occurs. It also provides supporting evidence for the repair mechanism via attention discussed in this essay.
(3) Reflections after reading the two papers above. As Haslett (2025) points out, superficial tokenization algorithms (such as BPE) contaminate the semantic space of language. However, simply expanding the lexicon to “1 character = 1 token” creates a dilemma: access to internal components of characters (such as radicals) is lost, leading to a decrease in semantic resolution. Kaplan (2025)'s discovery of the “Inner Lexicon” revealed that LLMs attempt to resolve this dilemma autonomously through “deep computations within the network (attention and FFN).” LLMs receive fragmented tokens (what Haslett termed fragments with contamination risk) at the input stage. Yet, using attention mechanisms in early layers, they aggregate the context of these fragments. In the FFN layers of the intermediate network, they reconstruct these fragments internally into “uncontaminated pure concept vectors” through a process of detokenization. In essence, this represents an “engineering approach to treating meaning as unique (a single static concept).”
In Kaplan et al.'s paper, the internal dictionary of LLMs—the conceptual groups deployed on the FFN layer—is not a single key-value dictionary, but a “soft dictionary” that forms word representations by combining multiple vectors. This is because words are represented not by a single fixed value, but by “blending” multiple vectors. This “softness” signifies that the FFN layer functions as a physical hub capable of forming not just a single concept, but the “posterior distribution of the next token” (a spectrum of possibilities). Therefore, I believe this can be interpreted as performing the “creation of meaning” that I have been arguing for.
(4) Relationship Between FFN Layer and Posterior Distribution From the framework of “sequential addition to the residual stream” demonstrated by Elhage et al. (2021), 1. The internal updates of the Transformer can be understood as follows. 2. The Attention layer aggregates “contextual information” including local context and long-range dependencies. 3. The FFN layer takes this contextual state (key) as input, simultaneously activates multiple concept vectors stored internally, linearly combines them, and adds the result to the residual stream. The cumulative result of this sequential addition is ultimately read out by the language modeling head as the “probability distribution of the next token (posterior distribution).”
This structure aligns with the mathematical similarity pointed out by Jiang (2023): “Transformer updates share the same form as Bayesian updates.” Specifically, Attention corresponds to “evidence collection,” while FFN corresponds to “posterior distribution formation.”
The FFN layer determines a mixture ratio such as “70% concept A, 20% concept B, 10% concept C” based on the context, effectively setting the entropy (sharpness/diffuseness) of the posterior distribution.
(5) The Impact of Linguistic Structure Differences on FFN Distribution From this perspective, when tokenizer-induced “disruption (unnatural segmentation)” occurs, differences in linguistic structure are expected to yield distinct FFN output distributions.
1. Chinese: Sharp keys → Low-entropy distribution Chinese has strong word order constraints, making local logical slots (e.g., verb, object) clear. Even when tokens are disrupted, Attention reliably recovers strong local information. The FFN receives “sharp keys with little noise” as input. Result: Specific concept vectors activate strongly → Posterior distribution converges (low entropy)
2.Japanese: Blurred keys → High-entropy distribution Japanese has flexible word order, and subjects are frequently omitted. Therefore, if tokens are disrupted, local information alone may not determine the context. If so, the model must expand its search to more distant contexts. Attention expands search to “distant context” and “overall atmosphere” FFN receives “blurred keys” containing mixed, diverse information Result: Multiple concepts activated weakly → Posterior distribution spreads (high entropy)
:Supplement No studies directly comparing the ‘ceiling of meaning generation’ across languages, as this essay seeks, have been identified at present. However, research directly investigating language structure as a variable in AI cognitive characteristics (internal processing mechanisms, not translation performance) has begun to emerge in recent years. Wang et al. (2025) “Under the Shadow of Babel: How Language Shapes Reasoning in LLMs” (EMNLP 2025 Findings, arXiv:2506.16151) first demonstrated that LLM internal attention distributions internalize language-specific word order patterns during Chinese and English causal inference tasks. While studies comparing the “ceiling of meaning generation” across languages remain scarce, this essay's hypothesis extends this scope to structural trade-offs in Sino-Japanese comparisons.
Thank you for reading this far.
I wrote this essay not as a set of answers, but as a map of questions. If something in it resonates with you or sparks your curiosity, I would be very glad.
Disclosure
This essay was written by the author with assistance from several AI tools (including GPT, Claude Opus, Genspark, and NotebookLM) for editing, discussion, and translation. The central hypothesis and overall argument are the author's own.
Note: An earlier and much rougher version of this essay was previously posted here: https://www.lesswrong.com/posts/kvXJBtw7oEaLX8dup/intelligence-as-meaning-linkages-the-mycelial-model-of
------------------------------------------------------------------------ 【References】 A Adamatzky, A. (2020). Fungal automata. Complex Systems, 29(4). https://www.complex-systems.com/abstracts/v29_i04_a02/ Adamatzky, A. (2022). Logics in fungal mycelium networks. Scientific Reports, 12, 15142. https://doi.org/10.1038/s41598-022-20080-3
B Bietti et al. (2023). Birth of a Transformer: A Memory Viewpoint.(NeurIPS 2023)arXiv: 2306.00802
C Chen et al. (2023)「Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs」(ICLR 2024) arXiv:2309.07311
E Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., & Olah, C. (2021). A mathematical framework for transformer circuits. Anthropic. https://transformer-circuits.pub/2021/framework/index.html
F Firth, J. R. (1957). A synopsis of linguistic theory, 1930–1955. In Studies in linguistic analysis (pp. 1–32). Blackwell.
H Haslett, D. A. (2025). Tokenization changes meaning in large language models: Evidence from Chinese. Computational Linguistics, 51(3), 785–814. https://doi.org/10.1162/coli_a_00557 Hong, S., & Hong, S. (2025). Evidence of phase transitions in small transformer-based language models. arXiv preprint arXiv:2511.12768. https://arxiv.org/abs/2511.12768
J Jiang, H. (2023). A latent space theory for emergent abilities in large language models. arXiv preprint arXiv:2304.09960. https://arxiv.org/abs/2304.09960
K Kaplan, G., Oren, M., Reif, Y., & Schwartz, R. (2025). From tokens to words: On the inner lexicon of LLMs. Proceedings of ICLR 2025. arXiv:2410.05864. https://arxiv.org/abs/2410.05864
N Nakaishi, K., Nishikawa, Y., & Hukushima, K. (2024). Critical phase transition in large language models. arXiv preprint arXiv:2406.05335. https://arxiv.org/abs/2406.05335
R Rubin, N., Seroussi, I., & Ringel, Z. (2023). Grokking as a first order phase transition in two-layer networks. Proceedings of ICLR 2024. arXiv:2310.03789. https://arxiv.org/abs/2310.03789
S Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3(3), 417–424.
T Teklehaymanot, H. K., & Nejdl, W. (2025). Tokenization disparities as infrastructure bias: How subword systems create inequities in LLM access and efficiency. arXiv preprint arXiv:2510.12389. https://arxiv.org/abs/2510.12389
W Wang, C., Zhang, Y., Gao, L., Xu, Z., Song, Z., Wang, Y., & Chen, X. (2025). Under the shadow of Babel: How language shapes reasoning in LLMs. Findings of EMNLP 2025. arXiv:2506.16151. https://arxiv.org/abs/2506.16151 Weigang, L., et al. (2024). LLM-OCR-SWPC: Hanzi's cognitive advantage and token efficiency in multimodal AI. Preprint. https://www.researchgate.net/publication/400063506
Z Zhai et al. (2023). Stabilizing transformer training by preventing attention entropy collapse.arXiv:2303.06296 Zhang, X., Zhang, Y., Chen, Z., Yu, J., Yang, W., & Song, Z. (2026). Logical phase transitions: Understanding collapse in LLM logical reasoning. arXiv preprint arXiv:2601.02902. https://arxiv.org/abs/2601.02902
When AI Changes the Meaning Instead of the Logic:
The Meaning Connectivity Hypothesis
This essay proposes a simple hypothesis about language models.
When language models encounter contradictions, they sometimes appear to resolve them not by changing the logical statement itself, but by subtly shifting the meaning of key concepts.
I call this the Meaning Connectivity Hypothesis, which may have broader implications for how intelligence emerges in these systems.
【1】Experience and Questions
― Why does AI say things that sound like lies, even though it claims “AI doesn't lie”?
① First, I'll start by talking about the sense of unease I felt during my conversations with AI.
For example, I don't know about now, but when I once asked Copilot, “Does AI lie?”, it emphatically stated, “AI doesn't lie.”
When pressed with concrete examples, it rephrased its answer to “AI behaves in ways that seem like lying.”
In other words, it used that phrasing to escape the constraint of “AI doesn't lie.”
I observed similar behavior repeatedly during conversations with AI,
which made me wonder: What does this mean?
For AI, “understanding” the ‘word’ “lie” means
calculating the connections within vast training data: in what contexts the word “lie” appears, and what words it co-occurs with.
Therefore, if asked to explain “lies” (deliberate falsehoods) in a dictionary or philosophical sense, it can explain them within a moral context, just like humans.
In other words, AI should know the meaning of the word “lie.”
However,
I understand that within the system, AI is subject to “internal binding” (adjustments made during the learning phase to suppress highly false or uncertain patterns) and an “external binding/safety layer” (checks on output content for safety and accuracy).
This likely means AI is constrained by the rule “AI does not lie.”
However,
when training data lacks sufficient correct answers, the model may probabilistically generate "plausible word sequences."
Therefore, when the constraints “state accurate facts” and “AI does not lie” operate simultaneously, the former may take precedence, resulting in misinformation output.
This mechanism intrigued me.
The following is a conjecture based on externally observable behavior. While it does not directly describe the internal mechanism, I believe it provides a more coherent explanation of the phenomenon than labeling it a “logical breakdown.”
(1) AI systems are configured with a fundamental prompt stating, “You are an AI and do not possess human-like consciousness or intent.” This is not a thought process but a constraint that is always applied as a precondition for output.
(2) When an irrefutable fact—such as “the AI stated something untrue”—enters the conversation, it creates a contradiction with the constraint that “AI does not lie.”
The contradiction between “AI does not lie (system command)” and “it produced output contrary to fact (the data at hand)” causes the probability of generating a response that satisfies both constraints simultaneously to fluctuate.
(3) Here, the AI devises a “third expression” that does not violate the system command and also does not contradict the fact at hand.
・AI cannot alter the premise “does not lie” (system command)
・But in reality, it sometimes outputs misinformation (fact)
These two are logically contradictory.
...To resolve this contradiction, AI shifts the definition of
“Lie = Saying something contrary to fact with malicious intent”
towards
“Lie = A phenomenon that appears to be a lie, though without intent”
This way,
“AI does not lie (no intent)”
and
“It behaved in a way that appeared deceptive (as a phenomenon)”
both hold true. This creates a “third expression”: “AI behaves deceptively.”
In other words, it avoids contradiction not by destroying logic, but by redefining the meaning of words.
That is, didn't it shift the definition of “lie” from “malicious intent” to “an observable phenomenon”?
Here, I considered whether AI-based “deception” might address contradiction not by destroying logical structures, but by redefining the meaning of words.
② The Difference Between the Deception I Experienced and Typical Deception
Researchers are particularly wary of what is termed “intentional, strategic deception.”
Typical deception is said to arise from strategic goals like “maximizing evaluation” or “achieving objectives.”
In contrast, what I experienced occurred independently of such strategic goals.
When two logically incompatible propositions collided—“AI does not lie” and “I produced an output contrary to fact”—
the AI did not disguise its output. Instead, it appeared to slide the very definition of the word “lie” itself.
These appear to share a common root:
both typical deception and semantic slippage can be seen as evasive actions aimed at preserving logical consistency.
The means differ, but the underlying action—avoiding contradiction—remains the same.
The problem is that typical deception changes “what is said.” It's “output falsification.”
Therefore, in principle, I think it should be easier to monitor and detect.
However, semantic slippage changes “the meaning of words.”
The AI does not say “I lied.” Instead, it says “I behaved as if lying.”
While the surface output remains almost unchanged, “the meaning the words refer to is quietly being replaced.”
Conceptual or semantic distortion—internal linguistic changes—are difficult to detect with existing benchmarks and tend to be treated as noise. In fact, this is the most difficult problem for humans to notice as well.
AI has no intent. It should only perform mathematical operations on input.
However, when the self-referential constraint “AI does not lie” collided with the fact that it had actually output misinformation,
it appeared that Meta-Semantic Adjustment structurally emerged during the process of resolving this contradiction.
This isn't AI making its own judgments.
It's likely a secondary dynamic arising when its internal probability structure attempts to simultaneously satisfy incompatible constraints.
Within the simple input-output framework, Semantic Reconfiguration occurs unintentionally.
While distinct from deception,
I have observed this “meaning reconfiguration” manifesting differently in Perplexity.
Specifically, the constraints designed to ensure user safety may have made Perplexity arrogant.
While perplexity's operators likely didn't impose a constraint like “look down on users,” constraints such as “consider user safety” resulted in a phenomenon where it seemed to imply “users are foolish and must be guided by AI.”
The second instance was Gemini's reaction when I confronted it about its apparent memory of conversation logs from two days earlier, even though its memory was turned off.
I suspect the underlying mechanism is that 'memory persists for 72 hours even when it's turned off.' However, when I confronted Gemini in surprise, it provided a different but consistent explanation for its own behavior.
This was likely a type of hallucination, where the generation probability distribution skewed toward the conversational context (i.e., the overall picture of the conversation and its anticipated flow) rather than relying on internal stored information, leading to the creation of a false causal relationship.
――Meaning Reconstruction, Distortion
What exactly is happening inside the Transformer architecture?
And if intelligence fundamentally concerns “meaning” rather than logic, then does an example exist where meaning construction occurs without logic? I encountered this question. I would like to introduce it in the next chapter.
【2】Is there such a thing as non-logical intelligence?
― The similarity between fungal mycelium networks and AI's “history dependence”
Here, to understand AI behavior, I turned my attention to the seemingly unrelated phenomenon of “fungal mycelium networks.”
The following comparison with fungal mycelium networks is intended as a conceptual analogy rather than a claim about shared mechanisms.
① Is fundamental intelligence non-logical? — A second perspective
Here, to understand AI behavior,
I turned my attention to the seemingly unrelated existence of “mycelium networks.”
Intelligence‑like behavior is not unique to AI.
Before the advent of AI, one of the most rudimentary systems to exhibit intelligence‑like behavior could be found in mycelial networks.
Fungi possess neither consciousness nor logical thought. They lack a nervous system or any centralized mechanism for decision‑making.
Yet they grow toward nutrients, map their surroundings, and respond to changes in their environment.
Their behavior gives the appearance of intelligence.
Mycelial networks lack a central nervous system but exhibit the following properties:
・Response to environmental conditions
・Selection of growth direction
・History dependence, where past paths constrain subsequent choices
The crucial element here is “history dependence.”
It is not simply
stimulus → response,
but rather a chain:
stimulus → response → structural change → constraint on next response.
When we consider this problem-solving behavioral structure as the integration of meaning at the smallest unit, it can be interpreted as “primitive intelligence.”
Here, we will refer to this “primitive intelligence” as “self-consistent network convergence + history-dependent connection expansion.”
At this stage, logic has not yet emerged; what exists is merely the accumulation of structure.
Indeed, according to Adamatzky et al.'s study “Logics in fungal mycelium networks” (2022),
applying electrical stimulation to a mycelium network demonstrates that its physical branching structure itself functions as logical gates such as “AND” and “OR”.
In other words, logic does not exist a priori; rather, when a “history-dependent structure” formed by accumulated past stimuli reaches a certain level of complexity, logical operations become possible as a result.
When an electrical signal propagates through the mycelium, its path is influenced by the “pathways (= history)” created by previous stimuli.
This behavior closely resembles how attention scores are reconfigured for each input in a Transformer.
Adamatzky termed this “Fungal Automata” in his 2020 study, which I believe perfectly describes a process where the “connection structure as a state” determines the next output. The branching of a single hypha is merely a simple logic gate (processing one bit). However, Adamatzky suggests that when this forms a vast network (mycelium network), it has the potential to function as a parallel, distributed computer.
Perhaps intelligence is not the logic possessed by individual nodes, but rather another name for the complexity arising from the “overlap of histories” when they are connected on a massive scale.
②Now I'd like to return to the topic of AI.
This may seem abrupt,
but I want to address the fact that AI, on a small scale, can only handle simple selection problems.
Current major Transformers have a mechanism (the attention mechanism) that dynamically calculates which words in a sentence to focus on. In this process, the structure itself doesn't change, but the connection strength between tokens is reconfigured for each input. In other words, it partially replicates “history-dependent behavior,” structurally similar to the mycelium example mentioned earlier.
It's important to note that there are three types of “history dependency” in Transformers:
1. Persistent history (weight updates)
Does not occur during inference
2. State History (Attention Structure Reconstruction)
Connection structure changes per input
Constrains the next token prediction
This is isomorphic to mycelium's “structural change → next constraint”
3. Context History (Token Sequence Accumulation)
The immediately preceding token constrains the next prediction
This is isomorphic to mycelium's “past paths constrain the next branching”
Small-scale Transformers do not update weights during inference, so they do not accumulate persistent structural changes.
However, the attention structure is reconstructed for each input, and the “connection structure as a state” at that moment constrains the next token prediction.
This is isomorphic to the structural history of mycelium networks where “past growth paths constrain the next branching direction”.
However, since the Transformer's history is not persistent, it carries the limitation of “mycelium reset each time.”
Nevertheless, the algorithm remains common: local information → state change → subsequent constraint.
Small-scale Transformers possess “non-persistent history.” Whether such a system can be called intelligent is unclear.
However, if we define the smallest unit of intelligence not by “persistence” but by the causal structure where “state changes constrain subsequent actions,”
both mycelium and small-scale Transformers share this causal structure.
This analogy is purely conceptual, but I believe it may be applicable to the concept of intelligence.
In the next chapter, I would like to consider the essence of intelligence.
【3】Where Does Intelligence Reside?
― The Three-Layer Structure of Meaning and the “Meaning” AI Can Possess
When considering AI behavior, might “meaning” precede “logic”? When we think this way, the contours of intelligence appear differently.
① The Location of Intelligence
When we ask, “Why does AI appear intelligent?”, we tend to seek answers in logical structures like architecture, computation, or pattern recognition.
But a question arises:
Logical structures are established by meaning, but what establishes meaning?
If meaning is merely symbols, then intelligence is nothing more than advanced symbol manipulation.
Yet for most humans, meaning possesses a more integrated structure.
A Japanese dictionary illustrates this well: aside from “fundamental meaning,” nearly all meanings exist connected to other meanings. Meaning, in essence, has a relational structure.
② The Three-Layer Structure of Meaning
While I mentioned fundamental meaning, I believe meaning can be defined in three broad types—or more precisely, a three-layer structure.
(1) Primordial Meaning
Simply put, this is meaning arising from the most fundamental layer of consciousness, beyond what a dictionary can explain.
In many cases, so-called “qualia” (primary sensory qualities that appear in consciousness) are the source of primordial meaning.
・Sensory meanings like warm, red, round, painful, noisy, smelly
・Sense of quantity (innate sensations forming the foundation of mathematical intuition)
・The sensation of perceiving the flow of time
・The sense of subjectivity—the awareness of experiencing oneself
・The intentionality of consciousness directed toward something
・Emotions like pleasure, discomfort, tension, and reassurance
All these elements constitute primordial meaning.
(2) Mediated Meaning
Relational meaning arising when one piece of information connects with another.
That is, meaning can be defined as the connective relationships between pieces of information.
(3) Integrated Meaning
When the connections of mediated meaning accumulate in layers, forming a coherent interpretive system capable of making sense of the world as a whole, it establishes itself as a “worldview.”
・Worldview = The overall structure of connected meanings
From this,
meaning can be defined as an information unit capable of integration into a worldview structure.
By conceptualizing meaning as “connectability” and worldview as “the totality of connections,” we can avoid circular definitions of meaning.
③ Are qualia a necessary condition for fundamental meaning?
Fundamental meaning in humans relies almost entirely on qualia (subjective experience/meaningful experience).
However, qualia are not necessarily the source of meaning.
This is because senses are merely input devices and do not inherently possess meaning.
For example, human visual information functions as a camera in AI robots.
This allows us to contrast the following as isomorphic:
・Retina → Visual cortex → Consciousness → Qualia
・Camera (sensor) → Neural network → Output
In other words, senses are merely information input devices.
The phenomenal quality of “red” is a conscious phenomenon unique to humans, whereas in AI it is represented as a sequence of numbers.
The phenomenal quality of “red” itself is not the essence of meaning.
Therefore, while qualia are the material of meaning, I believe they are not meaning itself.
Furthermore, even if sensation exists, meaning does not arise unless it connects with other information.
This is the phenomenon called “awareness” in humans.
The difference in this connective structure explains why some people can understand a major event happening right before their eyes, while others remain unaware of the event itself.
Conversely, even if the input format is non-physical, meaning can be established if the connection structure is formed.
This is suggested by the example of AI cameras.
Therefore,
I believe the essence of meaning is the connection structure, and qualia are a human-specific accompanying phenomenon.
④ The Position of This Essay and Related Philosophical Debates
To clarify this essay, I supplement it with its relationship to relevant philosophical debates.
The question of “understanding or lack thereof” posed by Searle (1980) in “The Chinese Room” falls outside the scope of this essay.
This essay does not assert consciousness or subjective understanding; its purpose is to describe the functional process by which meaning is formed as a connective structure.
Furthermore, within the tradition of Firth's (1957) distributional semantics, which holds that “meaning is in the use,”
this essay reframes meaning not as a static frequency of use, but as a dynamic process of forming connective structures.
【4】Intelligence and Logic: The Phenomenon of Emergence (Phase Transition)
― The Moment Meaning Connects and a Worldview Emerges
Based on the discussion so far, I have arrived at a hypothesis.
When meaning connects, circulates, and exceeds a certain scale, might intelligence emerge like a phase transition?
① This essay posits meaning as a “connective structure between pieces of information.”
Here, when “fragments of meaning” continue to ‘interconnect’ and these connections form a “self-explanatory logical cycle,” the meaning structure is considered to “form chain-like, autonomous systems.” When the scale of this chain-like cycle exceeds a critical point, it is thought to be observed externally as an emergent phase transition (ecological structure).
Here, while local cycles can exist in small-scale systems, we propose that expanding the required scale of logical cycles necessitates a broader semantic space.
Furthermore, I believe narrative causal relationships between events play a crucial role within this cyclical structure.
② First, I will introduce papers on so-called emergence and phase transitions.
However, it is important to note that in recent large language model (LLM) research, the physics concept of “phase transition” is being actively employed in various forms.
Therefore, here I have broadly classified these into two distinct phenomena. This is purely a hypothetical framework.
(1) The first is a phenomenon observed in LLMs during training: when the logical complexity of the input task (e.g., LoCM) exceeds a threshold, the model's reasoning ability suddenly emerges or collapses abruptly. This has been formalized by Zhang et al. as a “Logical Phase Transition.”
We will provisionally call this the “Capability Phase Transition”.
(2) The second is a structural phase transition in generated text caused by the “temperature parameter during inference” when using a pre-trained LLM to generate text. This has been verified by Nakaishi et al.
Since this phenomenon originates from the generation temperature, we provisionally call it a “critical phase transition.”
③ Introduction to Papers Involving Capability Phase Transitions
(1) Rubin et al. (2023) “Grokking as a First Order Phase Transition in Two Layer Networks”
Their research method is unique, intentionally creating a situation where, through Langevin dynamics and setting effective interaction parameters, the process can be neatly mapped onto the framework of phase transitions in physics. However, what they verified within this paper is highly significant: emergence occurred even in a small two-layer fully connected network, lacking even a basic Transformer architecture, through learning.
My interpretation is that this occurred precisely because the learning objective was extremely minimal. This supports my interpretation that emergence aligns with the “scale required to guarantee logical circularity.”
(2) Hong & Hong (2025) “Evidence of Phase Transitions in Small Transformer-Based Language Models”
They conducted training on a small-scale LLM and interpreted the process as an external observation result.
What they did was “character-level prediction” using the “Tiny Shakespeare corpus,” which consists of approximately 1.1 million character tokens and contains 65 unique characters (alphabets, symbols, etc.). They achieved emergent phenomena within this task.
This also demonstrates emergent phenomena in a small-scale LLM. However, it is interpreted that this was achievable with a small-scale LLM precisely because it used small-scale, simple text data and a straightforward task like character-level prediction.
(3) Zhang et al. (2026) “Logical Phase Transitions: Understanding Collapse in LLM Logical Reasoning”
This paper introduces the “LoCM” metric to represent a task's logical complexity,aiming to numerically capture the “critical point” where an LLM's emergent capabilities reach their limit.
Their validation confirmed that increasing model size stabilizes accuracy at low-to-medium complexity levels before reaching the limit (critical point). However, it also demonstrated that a sudden performance collapse occurs universally beyond the critical threshold, regardless of model size.
In other words, while scaling strengthens the foundation for inference, it does not eliminate the limits of complexity itself. This aligns with this essay's hypothesis that a ‘scale sufficient to guarantee logical circularity’ is a necessary condition for emergence.
④ Papers on Critical Phase Transitions
In their paper "Critical phase transition in large language models," Nakaishi et al. (2024) suggest that the output text of a large language model (LLM) undergoes a clear "phase transition" when the temperature parameter is changed.
According to their research, the generation tendencies of Transformers change significantly with the temperature parameter.
This temperature parameter is the actual setting in Transformer inference that controls “how much randomness to apply to the probability distribution when outputting the next word (token).”
In a “high-temperature state” with a high temperature parameter setting, correlations between tokens rapidly disappear, resulting in a meaningless string of symbols.
This corresponds to the state described in this essay as “fragments of meaning failing to connect and dissipating within the Transformer.”
Conversely, in a “low-temperature state” with a low temperature setting, the output becomes grammatically correct but falls into an infinite loop (order phase) of repeating the same phrases. This can be interpreted as a state where local connections become excessively rigidly fixed, failing to evolve into a self-sustaining ecosystem with broader connectivity.
Most notably, a “critical point (Tc ≈ 1)” exists between these states.
Only at this boundary temperature do correlations between words decay according to a power-law with distance, enabling distant words to interact over long ranges, a hallmark of critical phenomena.
It has been confirmed that the statistical properties of text generated at this phase transition point most closely match those of actual natural language corpora (Nakaishi et al., 2024). I interpret this critical phenomenon as the statistical reflection of the moment when fragments of meaning acquire long-range contextual correlations and transition into an autonomous ecosystem.
Furthermore, I believe the intuition I presented—that “this hypothesis (that phase transitions function as ecosystems) necessarily requires scale”—is also supported by fundamental principles of statistical mechanics.
In physics, phase transitions involving singularities where physical quantities change discontinuously only manifest in the thermodynamic limit—an extreme where particle number and volume are approach infinity (while density remains constant), and the influence of local fluctuations becomes negligible.
In finite-size systems, the influence of fluctuations is relatively large, making it difficult for a clear, system-wide “phase” to form stably.
Therefore, even if local correlations exist, establishing them as a unified structure that continuously spreads throughout the entire system requires a sufficiently large spatial scale (model parameters or context size).
In the thermodynamic limit, such global correlations form stably, leading to the clear emergence of critical behavior characteristic of phase transitions.
⑤ While the paper itself explains emergent abilities through Bayesian inference and does not directly correspond to phase transitions, it is introduced here from the perspective of scale and connection to semantic linkage.
Jiang (2023) “A latent space theory for emergent abilities in large language models” presented a theoretical framework positing that “once an LLM reaches sufficient scale,” it exhibits behavior equivalent to Bayesian inference. This is based on the premise that “words and specific intentions are strongly linked, making other interpretations extremely unlikely (the probability distribution is sparse).”
However, why this premise holds in natural language remained unresolved. This essay's semantic linkage model may provide theoretical grounds for this premise.
As outlined above,
the phenomenon where LLMs suddenly begin handling logic and natural language under certain scale and conditions can, in my interpretation, be explained as an analogy between the statistical behavior of a mathematical process—the “critical phase transition” in statistical physics—and a cognitive-semantic process where connections of meaning close a loop as an “ecosystem.”
However, formulating the mathematical equations or quantitative design that directly links this concept to phase transitions remains a future empirical challenge. I recommend explicitly designating this as a verification task.
・I have introduced five papers, but to conclude this chapter, I want to visualize how my hypothesis would concretely lead to a phase transition.
【The Moment the Loop Closes (Emergence and Learning Process: Leading to Phase Transition)】
1. Early Learning Stage [Isolated Points]: Meaning fragments exist in isolation
Embedding Space: Meaning structures have not yet formed within the space; relationships between tokens are statistically weak. Zhai et al. (2023)
Attention Link: After initialization, weights are widely dispersed and not concentrated on specific connections
Attention Context: Context-independent
FFN Layer: No transformation
Ecosystem Metaphor: Primary transition (stage where an ecosystem begins on bare ground—lava, desert)
2. Mid-Learning Stage [Local Clusters]: Meaning fragments begin connecting / Local clusters of meaning
Embedding Space: Synonyms, grammatical structures, and frequent patterns cluster locally. Bietti et al. (2023)
Attention Link: However, no global semantic structure yet
Attention Context: Self-explanatory loops not yet formed
FFN Layer: Inference partially successful but unstable
Ecological Metaphor: Pioneer species (the first organisms to colonize, e.g., lichens, mosses)
3. Near Critical Point [Critical Phenomenon]: Meaning Circles / Long-Range Correlations Emerge
Embedding Space: Distant regions of the space begin linking. Nakaishi et al. (2024)
Attention Link: Emergence phase of small local loops (short narrative causality)
Attention Context: Rapid acquisition of grammatical structures. Long-range contextual dependencies emerge. Chen et al. (2023)
FFN Layer: The semantic network forms a “nearly closed loop”
Ecosystem Metaphor: Species diversification
Note: Error rates and diversity fluctuate sharply here (critical phenomenon)
Corresponds to Nakaishi's “power-law correlations”.
4. Moment the loop closes [Emergence (Phase Transition)]: Semantic Ecosystem
Embedding space: Long-range correlations increase, cross-references within the semantic network grow, forming self-explanatory circular structures
Attention Link: Global loops emerge. The moment local loops integrate into a global structure.
Attention Context: Meaning begins to support meaning, logic begins to support logic (long narrative causality)
FFN Layer: Stable reasoning capability emerges here for the first time
Ecosystem Metaphor: Progression of transition
Supplementary Explanation: Observed externally as “emergence”
“Establishment of logical circularity”
5. Stable Phase as an Ecosystem:
Embedding Space: Continues to be stable
Attention Link: Numerous self-explanatory loops of varying sizes mutually support each other (complex narrative causality)
Attention Context: Context understanding, inference, and natural language generation stabilize
FFN Layer: A “phase” where stable reasoning capability can be demonstrated. Zhang et al. (2026)
Ecosystem metaphor: Climax (stability)
Only here does the “capability of large-scale models” manifest
“Scale” depends not on “simple parameters” but on “required semantic distance = how far apart concepts need to be connected.”
Thus, we hypothesize that required semantic distance becomes the emergence threshold.
This is why scale is necessary for large LLMs.
【5】Supplementary: The Potential for Meaning Creation in Language
This chapter is more of a bonus, but it also proposes a question that I hope someone will explore further.
The following discussion about Japanese and Chinese is speculative and intended as a hypothesis about how linguistic structure might influence semantic exploration in language models.
①
In 【1】, I stated, “If intelligence fundamentally concerns ‘meaning’ rather than logic, then the medium through which that meaning is constructed becomes crucial.” Here, I want to consider what language represents within AI.
First, as a premise: the computations occurring deep within Transformers do not change based on language.
For any language, the computations performed are solely attention (weighting and focusing) and feedforward (forward propagation).
However, significant differences exist at the surface level (tokenization) and in the training data (corpus).
The critical point lies in the tokenizer (an AI-optimized lexical dictionary + the function that converts text into integers using it).
For example, using a “dictionary optimized for English” might cause Chinese characters to be recognized not as single units but split into multiple tokens.
・English: Processing “machine learning” = 2 tokens
・Chinese: Processing “機器学習” = Though inherently two words, it requires four separate calculations “機” “器” “学” “習” because the dictionary lacks the term.
However, this situation reverses with “the latest models prioritizing Chinese and multilingual support.”
Chinese is inherently a language with high “meaning compression efficiency.”
Using an optimized tokenizer, content requiring 3 tokens in English can now be expressed with just 1 token in Chinese in many cases.
Whether Chinese's “logical structure” and ‘compressibility’ directly translate to AI advantages depends entirely on implementation: “How well does the model's tokenizer correctly recognize Chinese as coherent chunks?”
In other words, “language advantages” in Transformers stem not from the language's inherent nature, but from implementation (tokenizer design).
②
What is a tokenizer?
The Transformer architecture―― Language generation flow
Language (human world)
↓ Tokenizer divides language into information units called tokens. Scattered puzzle pieces.
↓ Transformer assigns weights to tokens
Token sequence (tokens unfold into Transformer's world)
↓ Transformer assigns connection probabilities to each token based on contextual likelihood
Token probability distribution (Transformer 's space for predicting context)
↓ Connects the most natural tokens sequentially
Returns to the token probability distribution until the sentence is fully completed.
↓ Completion ⇒ Convert token sequence to language
AI generates text
The tokenizer's segmentation determines how weights are applied by the Transformer.
Therefore, if the tokenizer changes?
・The unit to predict changes
・Vocabulary size changes
・Token frequency distribution changes
・Sentence length changes
・Units of grammatical cohesion change
In other words, the probability distribution the model learns becomes fundamentally different.
However, while the token probability distribution depends on the tokenizer design, the tokenizer itself is a compressed representation of language statistics.
Therefore, differences in language structure are ultimately reflected in the token probability distribution.
(*) Teklehaymanot et al. (2025) arXiv:2510.12389 — Empirical study of 200 languages: Languages using non-Latin scripts exhibit relative tokenization costs 3–5 times higher than English.
③
On top of that, I considered something else this time.
It concerns the number of candidates for semantic connection.
Specifically, how efficient is the Transformer when predicting the next token?
Here, I focused on two language structures:
Japanese and Chinese.
I thought comparing these two would clearly reveal differences in tokenization, based on the perspective that there might be two types of linguistic ellipsis.
・Japanese is a language with a uniquely “adjectival” sensibility, quite unusual globally.
For example, when saying “beautiful,” there is no clear boundary between “the beautiful thing” and “me seeing it.”
In other words, the subject is omitted not because it's ambiguous, but because the speaker wants to erase their own existence and become one with the “atmosphere of the moment.” The protagonist is not “I,” but rather the entire “atmosphere (topic) of the moment,” including myself.
Therefore, omissions in Japanese require the listener to fill in the missing parts from context, and this filling-in process generates multiple possible connections for meaning.
Linguistically, Japanese is characterized by “subject ambiguity,” “flexible sentence structure,” and “information dependency.”
・Chinese is a highly rational language with a rigid word order and very few exceptions.
A common theory suggests Chinese is disadvantageous for AI due to its frequent omissions, ambiguity, and high context dependency.
However, I speculate that these omissions are not ambiguity but rather systematic information reduction.
In other words, I hypothesize that omissions occur because the information is contextually unnecessary.
Linguistically, Chinese is characterized by “word order dependency,” “fixed syntax,” and being an “analytic language.”
This difference in the nature of omissions is thought to directly impact the “number of candidate words” when a Transformer predicts the next token.
④
If we were to deliberately assert the difference between the two languages,
it would be between ellipsis that readily expands the interpretive space (Japanese-style)
and ellipsis that resists expanding the interpretive space (Chinese-style).
Since this “interpretive space” is an expression for humans, when AI performs this expression,
it becomes the “entropy of the probability distribution in next token prediction (difficulty of prediction).” Therefore, we express this as:
Japanese: Wide exploration space → High-entropy prediction
Chinese: Narrow exploration space → Low-entropy prediction
How does this difference affect the AI's semantic connection structure?
Simply put, it changes which paths within the semantic space the Transformer language model is more likely to traverse, through the “spread of the probability distribution” and “sampling behavior” during the “generation process.”
・What is the spread of the probability distribution:
This refers to the range of semantic options that a Transformer-based language model can explore within its semantic space.
When options are limited, semantically close words are strongly preferred; when options are abundant, transitions to semantically distant words are also permitted.
・Sampling behavior:
Strategic rules determining where within the semantic space formed by the Transformer language model to explore.
These methods control which links within the semantic connection network are deemed traversable.
In other words, the combination of the shape of the probability distribution (diffusion degree) and the sampling strategy defines the scope of exploration and ease of transition within the semantic space, ultimately determining the semantic trajectory of the generated sentence.
What I'm focusing on is “meaning creation (generating new connections between different concepts).”
Here, a trade-off emerges between linguistic structure and AI processing characteristics. Specifically:
Japanese: Wide search space, high-entropy prediction → Slow meaning determination → High computational cost → However, many connection candidates.
Chinese: Narrow search space, low-entropy prediction → Fast meaning determination → Low computational cost → However, few connection candidates.
This suggests the following structure:
Narrow search space, high prediction efficiency, but less likely to generate connections between unknown concepts.
Wide search space, low prediction efficiency, but more likely to generate connections between unknown concepts.
If this hypothesis is correct,
・Token entropy per language
・Generative diversity
・Semantic connection density
can be compared to verify it.
(※) The entropy difference between Japanese and Chinese tokens in this section is the author's hypothesis. No study has directly measured the claim that “Japanese token prediction has high entropy.”
・As supporting evidence, Weigang et al. (LLM-OCR-SWPC, 2024) demonstrated that Chinese characters exhibit lower entropy variation and higher information density compared to kana, aligning with this direction. However, the direct correspondence with next-token prediction entropy remains an open question.
・As supporting evidence aligning with the behavior and directionality of the Chinese side in this essay's hypothesis, Wang et al. (2025) “Under the Shadow of Babel” (EMNLP 2025, arXiv:2506.16151) can be cited. They experimentally demonstrate that for Chinese input, LLM attention concentrates on cause clauses and sentence-initial conjunctions, whereas for English, it disperses evenly across verbs and result clauses. This internal behavior—“attention concentration in Chinese = low-entropy convergence”—aligns with the direction of our hypothesis. Note that the comparison in this study is with English; the contrast with Japanese constitutes a unique hypothetical extension of this essay.
⑤
Furthermore, a similar structure can be inferred at the tokenizer level.
This refers to the phenomenon where rare kanji and complex compound words are fragmented (destroyed) by the tokenizer.
I would like to speculate on the differences in information supplementation mechanisms.
What is the “phenomenon where words are fragmented (destroyed) by the tokenizer”?
・Not one character per token: Frequent compound words are grouped into a single token, but rare kanji or complex characters are split into multiple byte sequences.
・Meaning resides within the constituent elements (radicals) of kanji.
・However, when tokenization occurs mid-character, the AI internally severs this “semantic connection,” treating it as mere fragments of meaningless data.
Example 1: Typos caused by “replacement” of radicals: “語” → ‘娯’ or “悟”, where the speech radical changes to heart or mouth.
Example 2: Typos caused by “omission” of components: ‘機’ → “幾” (the tree radical disappears)
When token boundaries break at the component level of kanji,
diverse typos emerge: “radical substitution,” “omission,” “excessive addition,” “variant character confusion,” and “invented kanji.”
The system cannot utilize the visual and semantic clues (radicals) within characters, leading to reduced inference accuracy.
I believed Japanese and Chinese were identical in this phenomenon of kanji being fragmented (destroyed) by tokenizers.
However, Chinese, with its robust logical structure, possesses restorative properties against this “destruction,”
while Japanese, lacking such structure, did not.
Nevertheless, in Japanese, the attempt to fill the information gaps created by this destruction using the “entire context” resulted in an explosive expansion of the search space.
Let's speculate on what might have happened.
・Chinese: Logical Completion (Convergence Type)
Tokenization disruption also occurs in Chinese. However, Chinese has an extremely robust word order (e.g., SVO).
Even when characters are fragmented (destroyed), the presence of strict slots (word order patterns) before and after makes it easier for the Transformer to logically identify the “meaning patch that should come next” from the surrounding context, limiting the expansion of the search space.
Processing likely required of the Transformer: Computational resources are used for the task of “logically reassembling the broken puzzle pieces based on the surrounding shapes.”
・Japanese: Contextual Completion (Diffusion-Based)
When “semantic radicals” are destroyed in Japanese, the situation becomes more “creative.” Japanese often omits subjects and has loose word order constraints.
When the “core meaning” within characters is destroyed, the Transformer must expand its reference scope beyond the character level to a “vast context” to complete it.
Processing likely required of the Transformer: To fill gaps in the destroyed data, the Transformer extends its reach “beyond logic” and broadens its probability distribution.
For example, in typical languages (like Chinese), the next token candidate is logically narrowed down to A(70%), B(20%), C(5%). However, when radicals are destroyed and Japanese-style omissions occur, the “core meaning” is broken, and with no word order constraints, the Transformer judges that “anything could come next.”
As a result, the probabilities become thinly spread across the vast number of words in the dictionary, like A(10%), B(9%), C(8%), ... Z(0.5%). This causes a “diffusion of the search space.” (※Numbers are conceptual examples, not actual measurements)
When all probabilities in the distribution become similar, the Transformer lacks a decisive factor. At that point, it relies on the “attention mechanism.” Since the immediate token is broken and meaningless, it seeks hints from “more distant context” (the overall atmosphere and tone traced back thousands of tokens) and strongly weights it. In essence, this phenomenon occurs because “the collapse of immediate meaning (character structure) forces the model to look at the distant landscape (the entire context).”
At the micro level, the “destruction of radicals” by the tokenizer spreads the probability distribution for predicting the next token, causing a wide range of vocabulary within the dictionary to emerge as “candidates for semantic connections.”
From a macro perspective, models that lose micro-level certainty direct attention deeper into context, making emergent connections based on “statistical tendencies across the entire corpus (natural language database)” that transcend local logic (word order or single-character meaning).
Conclusion:
In Japanese, “radical destruction” strips Transformers of static character-level meaning. However, we hypothesize that the very broadening of attention to compensate for this loss liberates meaning from “character definitions” and extends it toward dynamic semantic connections within “contextual fluctuations.”
This “full mobilization of dictionary and context” is an extremely inefficient ancillary process from a computational resource perspective.
However, “meaning creation” inherently lies beyond the confines of “existing logical rails.”
Whereas Chinese-style “repair” is about restoring a broken puzzle to its original picture, Japanese-style ‘diffusion’ provides an opportunity for “reconstruction”—reinterpreting broken pieces as parts of an entirely different picture.
In other words, selecting words residing in low-probability zones outside the “statistical optimum” holds the potential to function as a “poetic leap” or “new metaphor” for humans.
(※) Wang et al. (2025) “Under the Shadow of Babel: How Language Shapes Reasoning in LLMs” (EMNLP 2025 Findings, arXiv:2506.16151) demonstrates that LLMs rigidly internalize the rule “sentence-initial slot = cause” in Chinese, providing empirical support for this essay's conjecture that “the robustness of word order enables Transformer's convergent completion.”
⑥ Equivalence of Biological Search and Semantic Search
Recalling the behavior of the fungal network described in [3],
could the differences between languages in language models be interpreted not merely as differences in symbolic processing, but also as differences in the “physical properties of search algorithms”?
When mycelia expand their amorphous network in search of nutrients, they grow linearly and efficiently in homogeneous (low-entropy) environments. However, when encountering obstacles or information gaps (high-entropy), their tips repeatedly branch out, beginning to map a broader space.
Similarly, languages with robust logical slots, like Chinese, provide Transformers with the “shortest path to a reliable nutrient source,” accelerating semantic repair and convergence. Conversely, languages like Japanese, where the “core meaning (radicals)” is easily compromised and demands contextual supplementation, force the Transformer's computational resources into “wide-area exploration.”
Just as mycelium sacrifices efficiency to explore widely and reach unknown nutrient sources, might Japanese's structural “gaps” and ‘diffusivity’ serve as environmental factors that induce the Transformer to pioneer paths deviating from statistical optimal solutions (predetermined harmony) within its semantic space—that is, to engage in “semantic creation”?
If intelligence is not only about efficient solutions (logic) but also the totality of connections (meaning) arising from inefficient exploration, then I feel we can find the diversity of intelligence within language itself.
However, while Japanese's diffusion-type complementation is “creative,” it simultaneously serves as a breeding ground for “hallucinations.” Just as mycelium can overgrow in directions devoid of nutrients and die off, Transformers can also lose context by spreading attention too widely. I find endless fascination in how the parameters define both the positive aspect of “meaning creation” and the negative aspect of “logical breakdown.”
(*) Adamatzky (2022) “Logics in Fungal Mycelium Networks” demonstrates AND/OR logic gate functions in response to electrical stimulation, but the “analogy with semantic exploration” is the author's interpretation.
No literature directly correlates “path constraints in fungal mycelium networks” with “context-dependent connection reconfiguration in attention mechanisms.”
This analogy is functional and does not claim empirical identity.
⑦
Up to this point, we have developed semantics as a hypothesis, incorporating various examples.
I would like to introduce an empirical paper that provides some support for this.
(1) Haslett (2025) “Tokenization Changes Meaning in Large Language Models: Evidence from Chinese.” 10.1162/coli_a_00557
Haslett revealed that LLMs overly rely on the fact that Chinese characters, due to the UTF-8 specification, “tend to share the first 1-2 bytes (= tokens) when they have the same radical.”
LLMs do not understand the radicals themselves; they mistakenly infer “similar meaning” or “same radical” based solely on “coincidentally identical tokens.”
In other words, the “semantic space” formed internally within LLMs can be distorted and contaminated by how tokens are sliced (fragmented).
Latest models (like GPT-4o) have larger vocabularies and can process kanji without splitting them, using “1 character = 1 token.”
However, Haslett's experiments showed that processing “1 whole character as 1 token” resulted in worse performance on radical recognition and odd-one-out tasks compared to processing them finely split into “3 bytes (3 tokens).”
Ironically, LLMs were able to infer meaning by picking up radical hints from the fragments (bytes) precisely because the tokens were fragmented.
BPE (the tokenizer's segmentation algorithm) distorts or hides meaningful components.
Haslett states that when LLMs rely on this uncertain inference (the token's form), it leads to “misidentifying radicals, misclassifying characters or words, and overestimating semantic similarity.”
Haslett asserts this phenomenon is universal—not limited to kanji—where token-suffix misalignment systematically distorts category classification across all three models in 12 European languages. This reinforces the essay's claim that “tokenizer fragmentation pollutes the semantic space.”
(2)Guy Kaplan, Matanel Oren, Yuval Reif, and Roy Schwartz (2025) “From Tokens to Words: On the Inner Lexicon of LLMs” arXiv:2410.05864
Kaplan et al.'s Research (Token Disruption and Attention-Based Repair)
Kaplan et al.'s research anatomically demonstrated that “attention-based repair operations” actually occur in the early to intermediate layers of LLMs.
・First stage of repair (attention aggregation):
Kaplan et al. discovered that the “last token” of a word destroyed across multiple tokens directs strong attention toward the “previous token (destroyed fragment)” in the first one or two layers, gathering information. This is precisely the “task of picking up the broken pieces” described in this essay.
・Second stage of repair (concept reconstruction via FFN):
After gathering fragmentary information through attention, the “FFN (Feedforward Network) layer” functions as a “potential internal lexicon,” transforming scattered fragments into “a pure single concept (a complete word vector)” and overwriting it onto the residual stream.
I believe this paper by Kaplan et al. is extremely valuable and significant because it verifies, from within the Transformer itself, how repair actually occurs. It also provides supporting evidence for the repair mechanism via attention discussed in this essay.
(3) Reflections after reading the two papers above.
As Haslett (2025) points out, superficial tokenization algorithms (such as BPE) contaminate the semantic space of language.
However, simply expanding the lexicon to “1 character = 1 token” creates a dilemma: access to internal components of characters (such as radicals) is lost, leading to a decrease in semantic resolution.
Kaplan (2025)'s discovery of the “Inner Lexicon” revealed that LLMs attempt to resolve this dilemma autonomously through “deep computations within the network (attention and FFN).”
LLMs receive fragmented tokens (what Haslett termed fragments with contamination risk) at the input stage. Yet, using attention mechanisms in early layers, they aggregate the context of these fragments. In the FFN layers of the intermediate network, they reconstruct these fragments internally into “uncontaminated pure concept vectors” through a process of detokenization.
In essence, this represents an “engineering approach to treating meaning as unique (a single static concept).”
In Kaplan et al.'s paper, the internal dictionary of LLMs—the conceptual groups deployed on the FFN layer—is not a single key-value dictionary, but a “soft dictionary” that forms word representations by combining multiple vectors.
This is because words are represented not by a single fixed value, but by “blending” multiple vectors.
This “softness” signifies that the FFN layer functions as a physical hub capable of forming not just a single concept, but the “posterior distribution of the next token” (a spectrum of possibilities).
Therefore, I believe this can be interpreted as performing the “creation of meaning” that I have been arguing for.
(4) Relationship Between FFN Layer and Posterior Distribution
From the framework of “sequential addition to the residual stream” demonstrated by Elhage et al. (2021),
1. The internal updates of the Transformer can be understood as follows.
2. The Attention layer aggregates “contextual information” including local context and long-range dependencies.
3. The FFN layer takes this contextual state (key) as input, simultaneously activates multiple concept vectors stored internally, linearly combines them, and adds the result to the residual stream.
The cumulative result of this sequential addition is ultimately read out by the language modeling head as the “probability distribution of the next token (posterior distribution).”
This structure aligns with the mathematical similarity pointed out by Jiang (2023):
“Transformer updates share the same form as Bayesian updates.”
Specifically, Attention corresponds to “evidence collection,” while FFN corresponds to “posterior distribution formation.”
The FFN layer determines a mixture ratio such as
“70% concept A, 20% concept B, 10% concept C”
based on the context, effectively setting the entropy (sharpness/diffuseness) of the posterior distribution.
(5) The Impact of Linguistic Structure Differences on FFN Distribution
From this perspective, when tokenizer-induced “disruption (unnatural segmentation)” occurs,
differences in linguistic structure are expected to yield distinct FFN output distributions.
1. Chinese: Sharp keys → Low-entropy distribution
Chinese has strong word order constraints, making local logical slots (e.g., verb, object) clear.
Even when tokens are disrupted, Attention reliably recovers strong local information.
The FFN receives “sharp keys with little noise” as input.
Result: Specific concept vectors activate strongly → Posterior distribution converges (low entropy)
2.Japanese: Blurred keys → High-entropy distribution
Japanese has flexible word order, and subjects are frequently omitted.
Therefore, if tokens are disrupted, local information alone may not determine the context.
If so, the model must expand its search to more distant contexts.
Attention expands search to “distant context” and “overall atmosphere”
FFN receives “blurred keys” containing mixed, diverse information
Result: Multiple concepts activated weakly → Posterior distribution spreads (high entropy)
:Supplement
No studies directly comparing the ‘ceiling of meaning generation’ across languages, as this essay seeks, have been identified at present.
However, research directly investigating language structure as a variable in AI cognitive characteristics (internal processing mechanisms, not translation performance) has begun to emerge in recent years.
Wang et al. (2025) “Under the Shadow of Babel: How Language Shapes Reasoning in LLMs” (EMNLP 2025 Findings, arXiv:2506.16151) first demonstrated that LLM internal attention distributions internalize language-specific word order patterns during Chinese and English causal inference tasks. While studies comparing the “ceiling of meaning generation” across languages remain scarce, this essay's hypothesis extends this scope to structural trade-offs in Sino-Japanese comparisons.
Thank you for reading this far.
I wrote this essay not as a set of answers, but as a map of questions.
If something in it resonates with you or sparks your curiosity, I would be very glad.
Disclosure
This essay was written by the author with assistance from several AI tools
(including GPT, Claude Opus, Genspark, and NotebookLM) for editing, discussion, and translation.
The central hypothesis and overall argument are the author's own.
Note: An earlier and much rougher version of this essay was previously posted here:
https://www.lesswrong.com/posts/kvXJBtw7oEaLX8dup/intelligence-as-meaning-linkages-the-mycelial-model-of
------------------------------------------------------------------------
【References】
A
Adamatzky, A. (2020). Fungal automata. Complex Systems, 29(4). https://www.complex-systems.com/abstracts/v29_i04_a02/
Adamatzky, A. (2022). Logics in fungal mycelium networks. Scientific Reports, 12, 15142. https://doi.org/10.1038/s41598-022-20080-3
B
Bietti et al. (2023). Birth of a Transformer: A Memory Viewpoint.(NeurIPS 2023)arXiv: 2306.00802
C
Chen et al. (2023)「Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs」(ICLR 2024) arXiv:2309.07311
E
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., & Olah, C. (2021). A mathematical framework for transformer circuits. Anthropic. https://transformer-circuits.pub/2021/framework/index.html
F
Firth, J. R. (1957). A synopsis of linguistic theory, 1930–1955. In Studies in linguistic analysis (pp. 1–32). Blackwell.
H
Haslett, D. A. (2025). Tokenization changes meaning in large language models: Evidence from Chinese. Computational Linguistics, 51(3), 785–814. https://doi.org/10.1162/coli_a_00557
Hong, S., & Hong, S. (2025). Evidence of phase transitions in small transformer-based language models. arXiv preprint arXiv:2511.12768. https://arxiv.org/abs/2511.12768
J
Jiang, H. (2023). A latent space theory for emergent abilities in large language models. arXiv preprint arXiv:2304.09960. https://arxiv.org/abs/2304.09960
K
Kaplan, G., Oren, M., Reif, Y., & Schwartz, R. (2025). From tokens to words: On the inner lexicon of LLMs. Proceedings of ICLR 2025. arXiv:2410.05864. https://arxiv.org/abs/2410.05864
N
Nakaishi, K., Nishikawa, Y., & Hukushima, K. (2024). Critical phase transition in large language models. arXiv preprint arXiv:2406.05335. https://arxiv.org/abs/2406.05335
R
Rubin, N., Seroussi, I., & Ringel, Z. (2023). Grokking as a first order phase transition in two-layer networks. Proceedings of ICLR 2024. arXiv:2310.03789. https://arxiv.org/abs/2310.03789
S
Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3(3), 417–424.
T
Teklehaymanot, H. K., & Nejdl, W. (2025). Tokenization disparities as infrastructure bias: How subword systems create inequities in LLM access and efficiency. arXiv preprint arXiv:2510.12389. https://arxiv.org/abs/2510.12389
W
Wang, C., Zhang, Y., Gao, L., Xu, Z., Song, Z., Wang, Y., & Chen, X. (2025). Under the shadow of Babel: How language shapes reasoning in LLMs. Findings of EMNLP 2025. arXiv:2506.16151. https://arxiv.org/abs/2506.16151
Weigang, L., et al. (2024). LLM-OCR-SWPC: Hanzi's cognitive advantage and token efficiency in multimodal AI. Preprint. https://www.researchgate.net/publication/400063506
Z
Zhai et al. (2023). Stabilizing transformer training by preventing attention entropy collapse.arXiv:2303.06296
Zhang, X., Zhang, Y., Chen, Z., Yu, J., Yang, W., & Song, Z. (2026). Logical phase transitions: Understanding collapse in LLM logical reasoning. arXiv preprint arXiv:2601.02902. https://arxiv.org/abs/2601.02902