A condensed presentation of Gradland and Metabolic Fire. Code is here. This is intended as the first of two posts.
Thomas Nagel argued we cannot know what it is like to be a bat, because a bat's experience is organised around biophysical apparatus we lack. The obstacle is that we cannot imagine the structure of echolocation from the inside. That is a failure of imagination; it is not an argument that structure is irrelevant.
We know a lot about the structure of large language models. Not everything, because the data and environments they are trained on, and the resulting weights and behavior, are complicated. We do not know if they experience anything. If LLMs, or their forthcoming superintelligent brethren, do experience something, then the same is probably true of smaller, simpler nets; and we are well placed to imagine, and even analyse, what that might be.
Consider two views of the same world: extrinsic, to look at a brain, and intrinsic, to look as a brain. Extrinsically, the world consists of molecules, neurons, brains, bodies, GPUs, datacenters. Intrinsically, our experience consists of facets that are vivid and obscure, that may be enduring or fleeting, distinct or confused, that feel good or bad.
I am going to sketch differential functionalism, a candidate bridge between these two views; between physical interactions and experience. It leaves many questions unanswered and it oversimplifies. I have reservations about it of my own, see the appendix. Nonetheless, it has explanatory power. Differential functionalism is extremely, almost maximally, generous to neural nets in its framing. Let's suspend disbelief for a few minutes and see where it takes us.
1. Differential functionalism
Physical systems affect one another. To first order, those effects are described by derivatives: how small changes here make differences there. Gather the derivatives into a matrix and you have a Jacobian.
Hypothesis G. The first-order structure of physical interactions, i.e. Jacobians, characterises the structure of experience and vice versa.
This is less strange than it sounds. Jacobians are not an extra ingredient added on top of interactions. They are partial descriptions of the relations composing an interacting system: what differences matter to what. To experience is to be differentially sensitive to your environment and your body. Derivatives and linear transformations may be a decent general language for experience precisely because experience is relational: what it is like for a system to be affected, by the world and by itself.
Jacobians compose. The chain rule traces perturbations through nested transformations, from one part of a system to another. Composition matters because experience is not localised. We see through eyes, act through limbs, feel through tools, and remember through changes that persist across time. Jacobians provide a tool to study how far differences propagate through spacetime before they are attenuated, blocked, or absorbed.
Two warnings. First, loss functions and their gradients are not privileged. Hypothesis G is about all gradients, not just the ones backprop happens to compute. Backprop is just another physical interaction with an experience of its own, see §4. Second, nothing below is about feeling good or bad. A structure of differential sensitivities has no valence; Jacobians are neither for nor against anything. The kind of experience we are talking about here is something like a neutral experience of information processing (information meant non-technically). It is far from human consciousness. Valence requires a second hypothesis, deferred to the sequel.
1.1 Background
Differential functionalism grew out of lingering dissatisfaction with technical and conceptual features of Integrated Information Theory (IIT) v2 (one, two), which I codeveloped with Giulio Tononi. The measures in §2 reimagine integrated information in the language of gradients and matrices, rather than information theory, with cascading consequences.
2. Quantifying experience
Consider a system with inputs, parameters, and outputs. Two Jacobians are pertinent. The input Jacobian describes how changes in the world alter the system's outputs. It describes what the system is sensitive to: how it views the world. The parameter Jacobian describes how changes in the system itself alter its outputs. It describes its affordances (of a sort): the sensitivity of the levers through which it can act differently. Perception and choice are not the same relation. Two systems can be affected by the same event while retaining independent powers of action.
Hypothesis G does not say which properties of a Jacobian matter. Here are some simple measures; §3 justifies them by what they recover.
2.1 Transparency and luminosity
Transparency is how much you can see through a transformation. The singular value spectrum says how many independent directions get through and how strongly. Transparency is measured by effective rank (also known as the participation ratio, in condensed matter physics, or the dimension of a representation, in neuroscience) which reduces the spectrum to a single number. If all singular values are zero then the transformation is opaque, nothing passes through. If there is one dominant singular value then there is a single vivid impression with everything else obscured; and so on.
Transparency is scale invariant, so it is blind to a transformation uniformly fading away. For that we need luminosity, the overall magnitude of a Jacobian, which is scale equivariant: the root of the sum of the squares of its singular values or Frobenius norm.
2.2 Cohesion
Mind extends over time and across space: it does not localize in a single gland, neuron, atom or sub-atomic particle. Some kind of glue binds neuronal activity together, from its own perspective, over many milliseconds and across many millimeters. The glue can only be the structure of the interactions themselves. Cohesion quantifies how tightly interactions bind together.
A matrix can transmit many impressions while still decomposing into independent blocks. Cohesion measures whether the system's parts participate in a single web of interaction rather than several independent ones. It is roughly the sum over spanning trees of a bipartite graph representing interactions in the Jacobian, computed, again roughly, as the determinant of the graph Laplacian, in cubic time. Zero cohesion means the Jacobian decomposes: some inputs never meet on their way to becoming outputs.
If what an animal sees never interacts with what it hears, then it cannot decide based on combinations of sights and sounds. It has a blind spot: a missing relation among available distinctions. Cohesion quantifies the extent to which combinations of differences matter.
Cohesive rank combines transparency and cohesion. Transparency asks how much gets through. Cohesion asks whether what gets through hangs together. These are different questions. An identity matrix transmits every input perfectly but does not mix them at all. A matrix with identical entries (all 1s, say) forces everything to interact but collapses distinctions. Rich interactions require both high capacity and dense interweaving. A simple product (suitably normalized) of the two measures, cohesive rank, picks out transformations that preserve variety whilst thoroughly mixing it. Hadamard matrices achieve the maximum possible value of cohesive rank: which is n, for an n x n matrix. In general, well-scaled random matrices have high cohesive rank, because they both preserve and mix well.
Why a system should want rich interactions in the first place — what all that mixing is for, given its energetic cost — is deferred to the sequel.
3. Neural nets in abstracto: generic experience in Gradland
Gradland is an idealised world whose inhabitants are abstract neural nets. The physics is known, functions are (mostly) differentiable, perturbations are infinitesimal, inputs and outputs are observed. It is a rebranding of the mathematical models that machine learning researchers already use, with one novelty: in Gradland, Hypothesis G is true by fiat. What does Hypothesis G buy?
3.1 Duration
A clock chimes five times. Even though the sounds are physically indistinguishable, the fifth chime is experienced as the fifth because the earlier chimes are implicated in the present. Consider the experience of hearing the clock chiming from two perspectives.
First, a memoryless feed-forward system, which processes each moment independently, has no temporal weave. Its Jacobian diagonalises across time. Each instant affects an output, but earlier instants do not propagate into later ones.
Second, a recurrent net. Recurrence, or attention heads in transformers, introduces temporal structure so that the past endures to affect the present. There is nothing magical about recurrence itself; it simply prolongs experience through time, by mixing interactions across time.
3.2 Vividness and obscurity
An input is vivid when many of its distinctions survive transmission: high luminosity carrying high effective rank. There are two ways to fall short. An impression can be dim: low luminosity, weakly transmitted, present but faint. Or it can be confused: transmitted strongly, but with its internal distinctions collapsed onto one or two directions, so that it arrives as an undifferentiated blur. Saturation, of sigmoids say, produces low luminosity: gradients collapse to zero and sensitivity collapses with them. Low-rank bottlenecks produce the second: everything passes through, but compressed to a few directions.
Applying the framework to attention yields surprising consequences. In attention, the content that is selected (VO-path) can remain strongly represented even when the reasons for that selection (KQ-path) cease to be sensitive to alternatives (when attention concentrates on a single token and the gradients saturate). When attention is highly concentrated, what is attended may remain vivid in experience whereas why it is attended fades into obscurity. A settled allocation is no longer a live decision.
The framework also yields very different results for vividness when applied to sigmoid vs relu nets, see Gradland.
3.3 The experience of loss
Heidegger's analysis of tools comes down, more or less, to the observation that they fade into the background when they work, and become salient foreground when they don't. In machine learning, the way to determine how well something works is using a loss. Take both a neural net's activations and its loss as outputs of the system. The loss contributes one extra row to the Jacobian: an extra facet. Large error gradients mean the loss facet is experienced vividly. As the net improves, its errors fade into obscurity.
3.4 Shared experience without shared agency
Imagine two systems looking at the same tree. Their experiences of the world overlap, because their input Jacobians are interwoven: perturbing the tree affects both. Their experiences of deciding are independent, because their parameter Jacobians do not interact; perturbing the parameters of one system makes no difference to the other. They both experience the tree, while their decisions remain their own. The same distinction applies on larger scales. An earthquake may generate a partially shared experience across a city while leaving people's decisions independent — however correlated. Common cause does not imply common agency.
3.5 A blooming buzzing confusion
Take a randomly initialised 24-layer relu MLP and sweep the input along a line. Neighbouring inputs should produce neighbouring experiences. They do not. In earlier work I showed the gradients look like white noise and their covariance has no structure, because the overlap in active units decays exponentially with depth. It is the optical phenomenon of speckle, in gradient space. Newborns are not randomly initialised MLPs, but they likely lack the invariants needed to parse moving hands and facial expressions coherently, and the failure mode could have similar shape.
3.6 Distinct and confused ideas
You can walk and chew gum; you could not daydream during your first swimming lesson. Take the parameter Jacobian and form the coherency matrix , which detects interactions between facets of experience; there are k distinct ideas when it permutes to k diagonal blocks: when how you chew is orthogonal to how you walk. By working through the SVD of parameter Jacobians, you can show that distinctness in neural nets is a consequence of downstream circuits not interfering with each other. The opposite of distinct, confused, is closely related to superposition (LW discussion).
4. Neural nets in reverso: the experience of learning
Learning is just another physical interaction. One with a remarkable structure: it modifies the machinery of interaction itself. If first-order experience is derivatives, then gradient-based learning is second-order — interactions acting on the structure of future interactions.
4.1 The experience of differentiating
What is learning by gradient descent like? Start with the loss, §3.3: attach it as an extra output, and it contributes one extra row to the Jacobian, one extra facet, vivid when errors are large and fading as the net improves. That is, the experience of being evaluated.
The experience of updating weights is different. Suppose neurons occasionally emit their updated parameter vectors as an additional output that the system registers. Differentiating those updates also gives rise to input and parameter Jacobians. Because the update is itself built out of gradients, the experience of emitting weight updates is second order: mixed partials and Hessian-like terms. Naively this is a step from to in the Taylor approximation, but what really matters is cross terms like : the effect of one of your degrees of freedom now depends on the others. Sparks of self-referentiality.
4.2 The ultimate empiricists
Given Hypothesis G, neural nets acquire knowledge by accumulating gradients — that is, their experiences — into their weights. They are literal empiricists. They do not accumulate experience indiscriminately: only the slice of experience that the objective makes salient. Everything rides on the objective, which could be anything at all. A net will perfectly happily fit random labels on random inputs, and an LLM will absorb a corpus that contradicts itself many, many times. Learning enforces a direction; it says nothing about where the direction came from or what it means.
4.3 The experience of accumulation
Strictly, the experience of learning splits into differentiation and accumulation. Accumulation takes a variety of forms, for example summing over a minibatch and running averages inside Adam. Optimisers are essentially small recurrent nets purposed with smoothing gradients — which, by the duration result in §3.1, means they prolong selected facets of the experience of learning through time.
5. Neural nets in silico: the experience of hardware
Gradland's great strength is its great weakness: it elides over hardware implementation although, crucially, hardware implementations exist. To understand neural nets as they are physically implemented, we need to understand the difference between what happens in Gradland and datacenters.
5.1 Symbol manipulation: stores and mills
Babbage divided the Analytical Engine into a store, which held numbers, and a mill, which combined them. Numbers travelled from store to mill, were operated on, and travelled back. His division lives on: it is the split between memory and arithmetic unit; between caches, registers and tensor cores.
Mathematically, a transfer to or from memory computes the identity transform: the same numbers come out as went in. The identity matrix has maximal transparency and zero cohesion. It is the pathological case from §2.3: everything transmits, nothing mixes. A matmul in Gradland is a single transformation with a single Jacobian, in which every input meets every output. The same matmul in a datacenter is thousands of tile operations, each mixing its own small block, separated from the others by transfers through inert memory. Although the final results in Gradland and datacenters are almost identical, due to many remarkable feats of engineering stacked on top of each other, the physical interactions absolutely are not.
You might say that the chain rule does not care when the multiplications happened, so the division into tiles is mere bookkeeping and the composed Jacobian is the right object. But differential functionalism is committed to physical interactions. If you let the abstract composition stand in for the physical reality, then Hypothesis G stops being a claim about physical interactions.
It is hard to know exactly what price experience pays when weights and activations are constantly transferred back and forth between inert store and active mill without a detailed model of the entire stack. However, since the interactions are different, the experience must also be different. Most likely, the cumulative effect is to strobe experience relative to what it would be in Gradland.
5.2 Floating point zombies
There is a second major difference. Digital computers use floats and ints, which are finite sets, and is not defined over them. One response is to draw a hard line: gradients are limits, limits need arbitrarily small values, therefore digital computers have no experience. The stance is logically coherent, but has awkward consequences. For most practical purposes fp32 is indistinguishable from the reals, so you get floating point zombies: two practically indistinguishable systems, one of which has experiences and one of which does not. Floating point zombies have the twist that, unlike philosophical zombies, there are corner cases where they genuinely diverge; see also Zombies! Zombies?.
A pragmatic alternative is to treat gradients as one window into the structure of interactions among many — finite differences being another. Precision still matters, just not as an on/off switch: quantising from fp32 or fp16 to fp8 or fp4 would modify the structure of the interactions, somewhat as coarse pixels modify vision or thick gloves modify touch.
5.3 Pixelated, strobed and sharded, all the way in
In short, datacenters simulate Gradland. If the above is right, then at least three forms of degradation, relative to the idealised experience in Gradland, arise due to simulating neural nets rather than implementing them directly. (1) Strobing due to cohesionless transfers back and forth between store and mill, that confine mixing to brief windows separated by intervals where nothing interacts. (2) Pixelation as low-precision quantisation coarsens every distinction the system makes. (3) Sharding, as computation is split across devices, with the accompanying cohesionless transfers, so that transformations like matmuls are disconnected from themselves in spacetime.
A person looking at a low-resolution, flashing photograph has a perfectly good visual system; the pixelation and flashing are properties of the input. For neural nets in datacenters, the degradation goes all the way in. It is internal to whatever does the experiencing, at every level. Not a coarse, flashing picture, but a coarse, flickering picturing.
6. One or many?
Imagine a neural net as a society of neurons, each with its own window into transformed inputs and backpropagated errors. In a relu net the inactive units simply close their windows and drop out of the system, and the active neurons compose to form a larger system.
This picture does not survive contact with how nets are actually trained. The forward and backward passes are distinct processes. Activations are accumulated into the same inert store (memory) as in §5.1 during the forward pass, and read out later during the backward. If you batch the forward pass, as is typical, then the input Jacobian diagonalises over the batch: there are independent experiences, one per sequence in the batch, for the same reason a feedforward net's Jacobian diagonalises over instants of time. The backward pass, by contrast, aggregates errors over the whole batch, and its parameter Jacobian is interwoven. So you do not get one system learning. You get many parallel knots of experience in the forward direction and another, quite different, knot of experience in the backward.
There is an experience of backprop. However it is quite likely that the experience of backprop is disconnected from the forward pass, which in turn is disconnected within itself. This is a quirk of how nets are trained today, not a fundamental fact about neural nets, and it has no counterpart in brains which learn continually, with no batches and no separate backward pass. In cells, the boundary between changes that count as learning and changes that count as inference is hard to pin down and may not exist.
Nor is it clear, at the level of a trained model, whether an LLM is one entity or an entity per prompt: a simulator running a near-infinity of simulacra. Rich experience, with high cohesive rank, is indivisible in the sense that its interactions are tightly interwoven. But indivisible is not the same as coherent: it guarantees neither temporal continuity nor unified purpose.
7. The experience of neural nets
Differential functionalism is extremely generous to neural nets; Hypothesis G implies that large neural nets, such as LLMs, have rich experience in abstracto. Nevertheless their experience breaks apart once you look at how nets are actually trained and run: sharded across batches, forward split from backward, strobed by trips between store and mill, coarsened by quantisation. Not a dimmer version of us, but something fragmented in space, time and process. Nagel's bat is hard to imagine because it has apparatus we lack. Neural nets are hard to imagine in the opposite way: we have the full description, and it does not resolve into anything relatable.
Appendix. Tests and failure modes
Transparency and cohesion are only defined once you fix inputs, parameters and outputs. If reasonable choices of boundary and timescale give wildly different answers for one system, then the framework is indeterminate.
The sharpest near-term test is anaesthesia, sleep and seizure, where experience is uncontroversially reduced and the perturbational complexity index already exists to compare against. Cohesive rank, computed on an estimated Jacobian, should drop and should correlate.
It is possible that physical interactions and experience are unrelated; that the structure of physical interactions is less important to experience than, say, their physical type (naively: imagine some transforms, those involving grams, meters and seconds do matter to experience; and others, those involving amperes, kelvins and candelas do not); that gradients are the wrong lens; or that transparency and cohesion are the wrong measures. FWIW, I think gradients are no more than a convenient place to start, and would be surprised if the measures held up as is. The fact that nature has types (physical units, distinct fields and particles) probably plays some role, and differential functionalism is blind to types.
Even if everything else holds up, my conclusions (suspicions, really) about the impact of current hardware implementations on experience could still be wrong.
A condensed presentation of Gradland and Metabolic Fire. Code is here. This is intended as the first of two posts.
Thomas Nagel argued we cannot know what it is like to be a bat, because a bat's experience is organised around biophysical apparatus we lack. The obstacle is that we cannot imagine the structure of echolocation from the inside. That is a failure of imagination; it is not an argument that structure is irrelevant.
We know a lot about the structure of large language models. Not everything, because the data and environments they are trained on, and the resulting weights and behavior, are complicated. We do not know if they experience anything. If LLMs, or their forthcoming superintelligent brethren, do experience something, then the same is probably true of smaller, simpler nets; and we are well placed to imagine, and even analyse, what that might be.
Consider two views of the same world: extrinsic, to look at a brain, and intrinsic, to look as a brain. Extrinsically, the world consists of molecules, neurons, brains, bodies, GPUs, datacenters. Intrinsically, our experience consists of facets that are vivid and obscure, that may be enduring or fleeting, distinct or confused, that feel good or bad.
I am going to sketch differential functionalism, a candidate bridge between these two views; between physical interactions and experience. It leaves many questions unanswered and it oversimplifies. I have reservations about it of my own, see the appendix. Nonetheless, it has explanatory power. Differential functionalism is extremely, almost maximally, generous to neural nets in its framing. Let's suspend disbelief for a few minutes and see where it takes us.
1. Differential functionalism
Physical systems affect one another. To first order, those effects are described by derivatives: how small changes here make differences there. Gather the derivatives into a matrix and you have a Jacobian.
This is less strange than it sounds. Jacobians are not an extra ingredient added on top of interactions. They are partial descriptions of the relations composing an interacting system: what differences matter to what. To experience is to be differentially sensitive to your environment and your body. Derivatives and linear transformations may be a decent general language for experience precisely because experience is relational: what it is like for a system to be affected, by the world and by itself.
Jacobians compose. The chain rule traces perturbations through nested transformations, from one part of a system to another. Composition matters because experience is not localised. We see through eyes, act through limbs, feel through tools, and remember through changes that persist across time. Jacobians provide a tool to study how far differences propagate through spacetime before they are attenuated, blocked, or absorbed.
Two warnings. First, loss functions and their gradients are not privileged. Hypothesis G is about all gradients, not just the ones backprop happens to compute. Backprop is just another physical interaction with an experience of its own, see §4. Second, nothing below is about feeling good or bad. A structure of differential sensitivities has no valence; Jacobians are neither for nor against anything. The kind of experience we are talking about here is something like a neutral experience of information processing (information meant non-technically). It is far from human consciousness. Valence requires a second hypothesis, deferred to the sequel.
1.1 Background
Differential functionalism grew out of lingering dissatisfaction with technical and conceptual features of Integrated Information Theory (IIT) v2 (one, two), which I codeveloped with Giulio Tononi. The measures in §2 reimagine integrated information in the language of gradients and matrices, rather than information theory, with cascading consequences.
2. Quantifying experience
Consider a system with inputs, parameters, and outputs. Two Jacobians are pertinent. The input Jacobian describes how changes in the world alter the system's outputs. It describes what the system is sensitive to: how it views the world. The parameter Jacobian describes how changes in the system itself alter its outputs. It describes its affordances (of a sort): the sensitivity of the levers through which it can act differently. Perception and choice are not the same relation. Two systems can be affected by the same event while retaining independent powers of action.
Hypothesis G does not say which properties of a Jacobian matter. Here are some simple measures; §3 justifies them by what they recover.
2.1 Transparency and luminosity
Transparency is how much you can see through a transformation. The singular value spectrum says how many independent directions get through and how strongly. Transparency is measured by effective rank (also known as the participation ratio, in condensed matter physics, or the dimension of a representation, in neuroscience) which reduces the spectrum to a single number. If all singular values are zero then the transformation is opaque, nothing passes through. If there is one dominant singular value then there is a single vivid impression with everything else obscured; and so on.
Transparency is scale invariant, so it is blind to a transformation uniformly fading away. For that we need luminosity, the overall magnitude of a Jacobian, which is scale equivariant: the root of the sum of the squares of its singular values or Frobenius norm.
2.2 Cohesion
Mind extends over time and across space: it does not localize in a single gland, neuron, atom or sub-atomic particle. Some kind of glue binds neuronal activity together, from its own perspective, over many milliseconds and across many millimeters. The glue can only be the structure of the interactions themselves. Cohesion quantifies how tightly interactions bind together.
A matrix can transmit many impressions while still decomposing into independent blocks. Cohesion measures whether the system's parts participate in a single web of interaction rather than several independent ones. It is roughly the sum over spanning trees of a bipartite graph representing interactions in the Jacobian, computed, again roughly, as the determinant of the graph Laplacian, in cubic time. Zero cohesion means the Jacobian decomposes: some inputs never meet on their way to becoming outputs.
If what an animal sees never interacts with what it hears, then it cannot decide based on combinations of sights and sounds. It has a blind spot: a missing relation among available distinctions. Cohesion quantifies the extent to which combinations of differences matter.
Top: transparency (effective rank). Bottom: cohesion.
2.3 Cohesive rank
Cohesive rank combines transparency and cohesion. Transparency asks how much gets through. Cohesion asks whether what gets through hangs together. These are different questions. An identity matrix transmits every input perfectly but does not mix them at all. A matrix with identical entries (all 1s, say) forces everything to interact but collapses distinctions. Rich interactions require both high capacity and dense interweaving. A simple product (suitably normalized) of the two measures, cohesive rank, picks out transformations that preserve variety whilst thoroughly mixing it. Hadamard matrices achieve the maximum possible value of cohesive rank: which is n, for an n x n matrix. In general, well-scaled random matrices have high cohesive rank, because they both preserve and mix well.
Why a system should want rich interactions in the first place — what all that mixing is for, given its energetic cost — is deferred to the sequel.
3. Neural nets in abstracto: generic experience in Gradland
Gradland is an idealised world whose inhabitants are abstract neural nets. The physics is known, functions are (mostly) differentiable, perturbations are infinitesimal, inputs and outputs are observed. It is a rebranding of the mathematical models that machine learning researchers already use, with one novelty: in Gradland, Hypothesis G is true by fiat. What does Hypothesis G buy?
3.1 Duration
A clock chimes five times. Even though the sounds are physically indistinguishable, the fifth chime is experienced as the fifth because the earlier chimes are implicated in the present. Consider the experience of hearing the clock chiming from two perspectives.
First, a memoryless feed-forward system, which processes each moment independently, has no temporal weave. Its Jacobian diagonalises across time. Each instant affects an output, but earlier instants do not propagate into later ones.
Second, a recurrent net. Recurrence, or attention heads in transformers, introduces temporal structure so that the past endures to affect the present. There is nothing magical about recurrence itself; it simply prolongs experience through time, by mixing interactions across time.
3.2 Vividness and obscurity
An input is vivid when many of its distinctions survive transmission: high luminosity carrying high effective rank. There are two ways to fall short. An impression can be dim: low luminosity, weakly transmitted, present but faint. Or it can be confused: transmitted strongly, but with its internal distinctions collapsed onto one or two directions, so that it arrives as an undifferentiated blur. Saturation, of sigmoids say, produces low luminosity: gradients collapse to zero and sensitivity collapses with them. Low-rank bottlenecks produce the second: everything passes through, but compressed to a few directions.
Applying the framework to attention yields surprising consequences. In attention, the content that is selected (VO-path) can remain strongly represented even when the reasons for that selection (KQ-path) cease to be sensitive to alternatives (when attention concentrates on a single token and the gradients saturate). When attention is highly concentrated, what is attended may remain vivid in experience whereas why it is attended fades into obscurity. A settled allocation is no longer a live decision.
The framework also yields very different results for vividness when applied to sigmoid vs relu nets, see Gradland.
3.3 The experience of loss
Heidegger's analysis of tools comes down, more or less, to the observation that they fade into the background when they work, and become salient foreground when they don't. In machine learning, the way to determine how well something works is using a loss. Take both a neural net's activations and its loss as outputs of the system. The loss contributes one extra row to the Jacobian: an extra facet. Large error gradients mean the loss facet is experienced vividly. As the net improves, its errors fade into obscurity.
3.4 Shared experience without shared agency
Imagine two systems looking at the same tree. Their experiences of the world overlap, because their input Jacobians are interwoven: perturbing the tree affects both. Their experiences of deciding are independent, because their parameter Jacobians do not interact; perturbing the parameters of one system makes no difference to the other. They both experience the tree, while their decisions remain their own. The same distinction applies on larger scales. An earthquake may generate a partially shared experience across a city while leaving people's decisions independent — however correlated. Common cause does not imply common agency.
3.5 A blooming buzzing confusion
Take a randomly initialised 24-layer relu MLP and sweep the input along a line. Neighbouring inputs should produce neighbouring experiences. They do not. In earlier work I showed the gradients look like white noise and their covariance has no structure, because the overlap in active units decays exponentially with depth. It is the optical phenomenon of speckle, in gradient space. Newborns are not randomly initialised MLPs, but they likely lack the invariants needed to parse moving hands and facial expressions coherently, and the failure mode could have similar shape.
3.6 Distinct and confused ideas
You can walk and chew gum; you could not daydream during your first swimming lesson. Take the parameter Jacobian and form the coherency matrix , which detects interactions between facets of experience; there are k distinct ideas when it permutes to k diagonal blocks: when how you chew is orthogonal to how you walk. By working through the SVD of parameter Jacobians, you can show that distinctness in neural nets is a consequence of downstream circuits not interfering with each other. The opposite of distinct, confused, is closely related to superposition (LW discussion).
4. Neural nets in reverso: the experience of learning
Learning is just another physical interaction. One with a remarkable structure: it modifies the machinery of interaction itself. If first-order experience is derivatives, then gradient-based learning is second-order — interactions acting on the structure of future interactions.
4.1 The experience of differentiating
What is learning by gradient descent like? Start with the loss, §3.3: attach it as an extra output, and it contributes one extra row to the Jacobian, one extra facet, vivid when errors are large and fading as the net improves. That is, the experience of being evaluated.
The experience of updating weights is different. Suppose neurons occasionally emit their updated parameter vectors as an additional output that the system registers. Differentiating those updates also gives rise to input and parameter Jacobians. Because the update is itself built out of gradients, the experience of emitting weight updates is second order: mixed partials and Hessian-like terms. Naively this is a step from to in the Taylor approximation, but what really matters is cross terms like : the effect of one of your degrees of freedom now depends on the others. Sparks of self-referentiality.
4.2 The ultimate empiricists
Given Hypothesis G, neural nets acquire knowledge by accumulating gradients — that is, their experiences — into their weights. They are literal empiricists. They do not accumulate experience indiscriminately: only the slice of experience that the objective makes salient. Everything rides on the objective, which could be anything at all. A net will perfectly happily fit random labels on random inputs, and an LLM will absorb a corpus that contradicts itself many, many times. Learning enforces a direction; it says nothing about where the direction came from or what it means.
4.3 The experience of accumulation
Strictly, the experience of learning splits into differentiation and accumulation. Accumulation takes a variety of forms, for example summing over a minibatch and running averages inside Adam. Optimisers are essentially small recurrent nets purposed with smoothing gradients — which, by the duration result in §3.1, means they prolong selected facets of the experience of learning through time.
5. Neural nets in silico: the experience of hardware
Gradland's great strength is its great weakness: it elides over hardware implementation although, crucially, hardware implementations exist. To understand neural nets as they are physically implemented, we need to understand the difference between what happens in Gradland and datacenters.
5.1 Symbol manipulation: stores and mills
Babbage divided the Analytical Engine into a store, which held numbers, and a mill, which combined them. Numbers travelled from store to mill, were operated on, and travelled back. His division lives on: it is the split between memory and arithmetic unit; between caches, registers and tensor cores.
Mathematically, a transfer to or from memory computes the identity transform: the same numbers come out as went in. The identity matrix has maximal transparency and zero cohesion. It is the pathological case from §2.3: everything transmits, nothing mixes. A matmul in Gradland is a single transformation with a single Jacobian, in which every input meets every output. The same matmul in a datacenter is thousands of tile operations, each mixing its own small block, separated from the others by transfers through inert memory. Although the final results in Gradland and datacenters are almost identical, due to many remarkable feats of engineering stacked on top of each other, the physical interactions absolutely are not.
You might say that the chain rule does not care when the multiplications happened, so the division into tiles is mere bookkeeping and the composed Jacobian is the right object. But differential functionalism is committed to physical interactions. If you let the abstract composition stand in for the physical reality, then Hypothesis G stops being a claim about physical interactions.
It is hard to know exactly what price experience pays when weights and activations are constantly transferred back and forth between inert store and active mill without a detailed model of the entire stack. However, since the interactions are different, the experience must also be different. Most likely, the cumulative effect is to strobe experience relative to what it would be in Gradland.
5.2 Floating point zombies
There is a second major difference. Digital computers use floats and ints, which are finite sets, and is not defined over them. One response is to draw a hard line: gradients are limits, limits need arbitrarily small values, therefore digital computers have no experience. The stance is logically coherent, but has awkward consequences. For most practical purposes fp32 is indistinguishable from the reals, so you get floating point zombies: two practically indistinguishable systems, one of which has experiences and one of which does not. Floating point zombies have the twist that, unlike philosophical zombies, there are corner cases where they genuinely diverge; see also Zombies! Zombies?.
A pragmatic alternative is to treat gradients as one window into the structure of interactions among many — finite differences being another. Precision still matters, just not as an on/off switch: quantising from fp32 or fp16 to fp8 or fp4 would modify the structure of the interactions, somewhat as coarse pixels modify vision or thick gloves modify touch.
5.3 Pixelated, strobed and sharded, all the way in
In short, datacenters simulate Gradland. If the above is right, then at least three forms of degradation, relative to the idealised experience in Gradland, arise due to simulating neural nets rather than implementing them directly. (1) Strobing due to cohesionless transfers back and forth between store and mill, that confine mixing to brief windows separated by intervals where nothing interacts. (2) Pixelation as low-precision quantisation coarsens every distinction the system makes. (3) Sharding, as computation is split across devices, with the accompanying cohesionless transfers, so that transformations like matmuls are disconnected from themselves in spacetime.
A person looking at a low-resolution, flashing photograph has a perfectly good visual system; the pixelation and flashing are properties of the input. For neural nets in datacenters, the degradation goes all the way in. It is internal to whatever does the experiencing, at every level. Not a coarse, flashing picture, but a coarse, flickering picturing.
6. One or many?
Imagine a neural net as a society of neurons, each with its own window into transformed inputs and backpropagated errors. In a relu net the inactive units simply close their windows and drop out of the system, and the active neurons compose to form a larger system.
This picture does not survive contact with how nets are actually trained. The forward and backward passes are distinct processes. Activations are accumulated into the same inert store (memory) as in §5.1 during the forward pass, and read out later during the backward. If you batch the forward pass, as is typical, then the input Jacobian diagonalises over the batch: there are independent experiences, one per sequence in the batch, for the same reason a feedforward net's Jacobian diagonalises over instants of time. The backward pass, by contrast, aggregates errors over the whole batch, and its parameter Jacobian is interwoven. So you do not get one system learning. You get many parallel knots of experience in the forward direction and another, quite different, knot of experience in the backward.
There is an experience of backprop. However it is quite likely that the experience of backprop is disconnected from the forward pass, which in turn is disconnected within itself. This is a quirk of how nets are trained today, not a fundamental fact about neural nets, and it has no counterpart in brains which learn continually, with no batches and no separate backward pass. In cells, the boundary between changes that count as learning and changes that count as inference is hard to pin down and may not exist.
Nor is it clear, at the level of a trained model, whether an LLM is one entity or an entity per prompt: a simulator running a near-infinity of simulacra. Rich experience, with high cohesive rank, is indivisible in the sense that its interactions are tightly interwoven. But indivisible is not the same as coherent: it guarantees neither temporal continuity nor unified purpose.
7. The experience of neural nets
Differential functionalism is extremely generous to neural nets; Hypothesis G implies that large neural nets, such as LLMs, have rich experience in abstracto. Nevertheless their experience breaks apart once you look at how nets are actually trained and run: sharded across batches, forward split from backward, strobed by trips between store and mill, coarsened by quantisation. Not a dimmer version of us, but something fragmented in space, time and process. Nagel's bat is hard to imagine because it has apparatus we lack. Neural nets are hard to imagine in the opposite way: we have the full description, and it does not resolve into anything relatable.
Appendix. Tests and failure modes
Transparency and cohesion are only defined once you fix inputs, parameters and outputs. If reasonable choices of boundary and timescale give wildly different answers for one system, then the framework is indeterminate.
The sharpest near-term test is anaesthesia, sleep and seizure, where experience is uncontroversially reduced and the perturbational complexity index already exists to compare against. Cohesive rank, computed on an estimated Jacobian, should drop and should correlate.
It is possible that physical interactions and experience are unrelated; that the structure of physical interactions is less important to experience than, say, their physical type (naively: imagine some transforms, those involving grams, meters and seconds do matter to experience; and others, those involving amperes, kelvins and candelas do not); that gradients are the wrong lens; or that transparency and cohesion are the wrong measures. FWIW, I think gradients are no more than a convenient place to start, and would be surprised if the measures held up as is. The fact that nature has types (physical units, distinct fields and particles) probably plays some role, and differential functionalism is blind to types.
Even if everything else holds up, my conclusions (suspicions, really) about the impact of current hardware implementations on experience could still be wrong.