This is the second of two posts. The first, What is it like to be a neural net?, sets up the framework; this one can probably be read alone.
As AIs surpass human capabilities in ever more domains, it is worth asking whether there is a fundamental difference between them and us, or between neural nets and biological organisms. I argue there is, and it is a difference worth actively maintaining.
The previous post sketched differential functionalism: the proposal that the first-order structure of physical interactions, gathered into Jacobians, characterises the structure of experience. It accounts for some basic phenomena: how the past weaves into the present through recurrence; how impressions can be vivid or dim, or distinct or confused, depending on what survives transmission; how two systems can share an experience while keeping their decisions separate.
It left two unpaid debts, to be handled in this post. The first is that rich experience, here meaning high cohesive rank, was not motivated. Experience comes at a hefty energetic price, so what is the corresponding benefit? I return to this in §6.1. The second is that I did not explain why a physical system should feel good or bad. A Jacobian has no stake in itself.
The previous post worked through two varieties of experience. Generic experience is the structure of generic interactions; roughly the forward pass in all its forms. The experience of learning arises when interactions modify a system's sensitivities and affordances. Here, I discuss the experience of sustaining, which arises when learning is coupled to a system's own continued capacity to act and experience.
To live is, primarily, to sustain one's ongoing existence; the experience of living is, at its heart, the experience of sustaining.
1. From differential functionalism to affect
Differential functionalism rests on:
Hypothesis G. The first-order structure of physical interactions, i.e. Jacobians, characterises the structure of experience and vice versa.
Hypothesis G says nothing about how experiences feel, about which experiences are good or bad. A structure of differential sensitivities has no valence; Jacobians are neither for nor against anything.
Systems are dynamic and their experience is correspondingly subject to constant change. Rescaling is arguably the simplest possible change to an experience. Not a change in what the system is sensitive to, nor in how its sensitivities interweave, but simply dialing the whole thing up or down. Imagine a dial that rescales the Jacobian, raising or lowering its luminosity: the overall magnitude of the transformation. Every distinction is preserved in kind and altered in degree. Nothing is added; nothing is removed; the entire structure is simply amplified or dampened.
Rescaling is the unique change that preserves the entire structure exactly while varying capacity to act. Amplifying and dampening experience are not symmetric, however. It is easier to die than to live. In theory if not in practice, sensitivity can be amplified without limit. But push a system far enough outside its viable range of functioning, in any direction, and experience goes to zero as its ability to (re)act collapses. Dead in effect.
Hypothesis A. The experience of empowering or diminishing (rescaling) experience is experienced as positive or negative affect respectively.
Hypothesis A is that there is a felt difference in kind between two tendencies: one empowering and the other diminishing, the base experiences of good and bad.
The words "the experience of" at the start of hypothesis A are important. Rescaling alone is not affect. A system dialed down by an unmodelled external force grows dim; it does not thereby feel bad about growing dim. For the rescaling to be experienced it must itself show up in the system's Jacobians — the dial, or a model of its effects, has to be part of the system. The doubling in the statement is deliberate: what is experienced as affect is not the rescaling, but the experience of rescaling.
Given Hypothesis A, I will argue that sustaining one's self generates affect, normativity, and purpose at their most basic.
1.1 Background
Hypothesis A, the minimal models of metabolism, and the implications for normativity and purpose are inspired by Hans Jonas, and related to the work of Xabier Barandiaran, Lisa Feldman Barrett, and many others.
2. Brains
Not all living things have brains, but we do, and it is worth briefly discussing them. Brains are differentially sensitive to an enormous variety of inputs, which they integrate into complex behaviors. They are transparent and cohesive. They are also quite different from neural nets.
If I had to summarise the design philosophies of datacenters and brains in one word each, they would be insulated and opportunistic. Chip designers minimise noise by insulating logic from physics: noise margins, shielding, error correction and clock discipline all exist to guarantee that the physical state of one transistor does not influence its neighbour except along the intended logical path. Evolution exploits noise by folding physics into function. The consequence, for us, is that interactions in brains, at every level and timescale, bear directly on the texture of experience, in a way that interactions in a datacenter do not.
The most striking architectural difference is that brains are massively recurrent. Recurrent nets are not that recurrent, interleaving feedforward layers with recurrent ones that refer back only to themselves; transformers beat out RNNs by largely avoiding recurrence because it is computationally expensive. Brains learn continuously, so real-time feedback, adaptive expectations and predictive processing loom far larger in our experience than in a net's. The brain is awash in neuromodulators, many with no analogue in machine learning. It is unlikely brains use backpropagation. It is unlikely they optimise a single objective. And they face an incomparably more adversarial environment.
None of these differences strikes me as fundamental. For that we have to stop asking about form and ask about function. The function of a computer, in a word, is to compute: to execute mechanical, elementary instructions. The primary, ongoing function of an organism and its brain is to survive.
3. Goals
Cybernetics formalised goal-directed behaviour in terms of feedback: a system behaves purposefully when negative feedback controls its behaviour; it detects discrepancies between where it is and where it should be, and acts to close them. Purpose is the loop, and nothing else is required. Modern machine learning has transformed cybernetics, rebranded as AI, into a methodology for imposing goals on machines at unprecedented scale. But having a goal in this sense is not the same as having a purpose of one's own. Hans Jonas made the point with a torpedo.
A self-guided torpedo tracks its target, corrects deviations, and closes in. I can say the torpedo's goal is to hit the target because I built it to do that — but hitting the target is my goal. There is no reason to think the torpedo cares if it hits or misses.
Put a brave human pilot inside the torpedo. Although the manned and automatic torpedoes behave the same, it is immediately clear that hitting the target is the pilot's goal, that he cares, and that the sensor feedback is information that he uses in service of his goal. The sensor feedback is not the source of the goal, not even in part. From the commander's point of view the two torpedoes are equivalent, because both serve his purpose. He can regard the pilot, for the duration of the mission, as mechanised, as a tool in his service. Cybernetics neatly codifies the commander's extrinsic perspective. The pilot experiences the two scenarios differently: alive in one and dying in the other as the torpedo hits its target and explodes with him inside.
But where do goals come from in the first place? The standard reply is that organisms have goals because of evolution by natural selection. But natural selection is just another feedback mechanism, and an extraordinarily indirect one: it provides essentially no feedback over the course of a single life. Whatever makes an organism's goals its own has to operate on the timescale of lived experience. Jonas points specifically to metabolism and I follow his lead.
4. The experience of sustaining
This section unfolds the implications of applying Hypothesis A to minimal, differentiable models of metabolism. Objections are gathered in Appendix A.
4.1 Metabolism
Hypothesis A does not explain how or why experience might be rescaled. In biological organisms, a natural candidate mechanism is metabolism.
Metabolism is just another physical interaction; one with extraordinary structure. It is the web of processes that converts matter and radiation into usable energy, and more broadly, it is how a body sustains, repairs, and renews itself. Active bodies require usable energy. As the supply tends to zero, so do capabilities and experience, for brains and datacentres alike.
Here is a minimal differentiable model of metabolism. Let a metabolic factor ranging from zero (none) to one (enough) determine the scale of the system's output. This is not a binary gate; nothing is computed and then switched off or dampened in a separate step. The claim is that below some range, less available energy means less activity, less responsiveness, and less capacity to sustain differentiated interaction. Because the Jacobians scale with that factor, so does luminosity. As usable energy vanishes, the ability to discriminate and to act vanishes with it.
Let the system learn, and consider what changing metabolism does to the weights. That is, consider the metabolic Jacobian, the derivative of the updated weights with respect to the metabolic factor. The answer is clean: it is proportional to the aggregated parameter Jacobians, the sum, over the batch, of the system's own experienced affordances, see Eq. (2) in Metabolic Fire. Turning the metabolic dial up is the direction of your aggregate experience. To modify metabolism is to dampen or amplify your own experience.
Coupling learning to metabolism thus yields a self-referential mechanism whereby your own weight updates diminish or empower your self. By Hypothesis A, that is experienced as affect.
The coupling also yields a minimal model of self-awareness: the metabolic Jacobian incorporates the parameter Jacobian. Nothing so sophisticated as introspection: the system does not represent itself as an object, possess a narrative identity, or entertain a concept of survival. Rather, since changes in the system's power to continue acting make a difference to how it learns, the system has the basic experience of its effect on itself.
The experience of sustaining arises when the conditions enabling a system's continued existence weave into the decisions that guide its modification.
4.2 Purpose and norms
Affect gives changes in one's own experience a valence. Creating purpose and norms requires that affect attach to your actions, so that acting one way rather than another is better or worse for you.
Extend the model so that useful energy levels depend on prior activity: your energy now depends on what you ate before. Suppose sustaining energy levels is the objective. (Treating a minimum level as a constraint instead is more realistic and gives the same term: the constraint contributes nothing while energy is comfortable and presses harder the further it falls, which is what homeostatic pressure is.) Then, by chaining transformations, the sensitivity of metabolism to one's actions composes with the sensitivity of those actions to one's parameters, and the result assigns a preferred direction to every parameter: the direction of empowering oneself. Your entire experience is suffused by affect derived from the effect of your actions on your own metabolism and thus your self. In practice, most actions do not change energy levels one timestep later, which is why drives like hunger and thirst exist: to bridge the spatiotemporal gap.
The difference between a reward and a need is crucial. A reward is a number the learner is compelled to increase. It can stand for anything and, at that level of generality, has nothing to do with the learner except that the learner is compelled to raise it. Metabolism is a kind of reward, but with an extraordinary property that rewards in general lack: it determines the system's capacity to act and to differentially react, and so controls its experience. As metabolism goes to zero, so do (re)activity and experience.
Metabolism is a target that cannot be gamed, reward-hacked, or Goodharted. Its signals can be — hunger is easy to fool — but the quantity itself cannot. The energy is either there or it is not, and if it is not, you will die.
What does this give an organism that evolution does not give directly? A measure that each organism continually applies to itself: a view of its own experience with a preferred direction. Natural selection culls lineages over generations. Sustaining metabolism determines, in real time: is this good or bad for me?Organisms create norms by distinguishing good from bad, for themselves, as they go.
A learning objective supplies a direction. Self-maintenance supplies a direction that is about the system itself. To have a purpose of your own, in this sense, you must first value your self.
4.3 The metabolic channel
A single scalar dial is too blunt to be the whole story, so it is worth seeing what affect looks like now that the loop of §4.2 is closed. For that, model energy reserves as a leaky integrator that decays and is replenished by what the system did previously, with the metabolic factor a concave, saturating function of that reserve. Concavity matters: near capacity, more energy changes little; when depleted, every unit counts.
A perturbation to an input now reaches a later output by two routes: through the network, and through the energy reserve. Call the second route's contribution the metabolic channel. Unrolling the recursion gives it in closed form: the channel accumulates time-decayed, energetically reweighted past experience, rescaled by current salience.
The channel materialises an exponential moving average over the rows of the Jacobian, which is the same mechanism that produced the experience of duration in the previous post — the past weaving into the present, weighted by what each moment did to the reserve. The channel is narrow, because it bottlenecks through a scalar energy reserve, but it is not one-dimensional; affect has more structure than a single scalar mood. And metabolic experience is loudest when energy depletes, because salience grows as the reserve empties.
So past events weave into present experience with weight according to: (1) how much they added to or drained the reserve at the time, (2) how long ago they were, and (3) how "hungry" the system is right now.
Self-preservation is not all there is to life. It was the initial impetus, billions of years ago, and remains the foundation. The ancients built on low but solid ground. Our goals and self-conception are no longer so simple, narrow, or bounded. Empowerment, for us, is about far more than keeping energy above a threshold.
5. Varieties of experience
The varieties are neither exclusive nor exhaustive; they abstract a messy reality.
Generic experience is structured interaction: being differentially affected.
The experience of learning is structured self-modification: changing one's future sensitivities and affordances according to present interaction.
The experience of sustaining is self-modification coupled to one's own enduring existence: learning in a way that makes the conditions of future interaction beneficial to the system itself.
The framework also provides candidate descriptions of affect, purpose, and norms at their most basic.
Affect is the experience of empowering or diminishing one's own experience.
Purpose and norms are created by weaving learning and affect, respectively, into the act of sustaining.
6. Organisms and AIs
What are the similarities and differences, at a high level, between organisms and artificial neural nets, or more specifically humans and LLMs?
6.1 The core similarity: rich experience and what it is for
Why do humans and LLMs both have — or, in the case of LLMs, at least simulate — rich generic experience? What problem is rich generic experience solving?
To live, in this world, is to be constantly probed and opposed, hunted, eaten, evaded, and infested. There is no angle that is not open to attack, no blind spot that will not be ruthlessly exploited. The greater the breadth and depth of combinations of events you can react to, experience, learn from, understand, plan for, and exploit, based on your biases and prior experience, the better your odds in life.
LLMs are not exposed to adversity as organisms are. But they are trained on a corpus that spans most of what humans do and say, and are asked to continue any part of it. Post-training adds an enormous variety of tasks. Variety without opposition still demands variety.
Rich experience, here, means high cohesive rank: high effective rank combined with high cohesion. Effective rank measures range or capacity: a system that can register more independent impressions has a broader repertoire of possible reactions or responses. Cohesion measures combinatorial coverage: the more those impressions can interact, in a system, the fewer blind spots — combinations it cannot react to — it is likely to have.
Rich experience on its own of course does not guarantee success in life; random matrices have high cohesive rank. Learning is required to dynamically adapt a system's experience to its environment. Rich, well-adapted experience is a viable niche; there is always room at the top.
6.2 Do neural nets feel?
Note that self-report cannot settle the question. Self-report is not informative about the experience, if any, of neural nets. Pretraining, reinforcement learning from human feedback, and related procedures train models to produce human language about human experience. However, obviously, they do not score outputs against the ground truth of the model's own phenomenology. A language model can therefore speak and reason fluently about fear, pleasure, pain, or desire without those words entailing anything about its self. Statistical correlations between activations and words similarly do not entail anything about phenomenology.
Suppose Hypotheses G and A hold. Do neural nets experience affect? Their experience is likely fragmented due to the nature of hardware implementations: pixelated, strobed and sharded. If the forward and backward passes were integrated, and so part of the same experience, then learning (which involves localised diminishing and empowering of experience, see final objection in Appendix A) would be felt. I suspect this is not yet the case.
That the forward and backward passes are separate processes in neural nets, and that our hardware simulates neural nets rather than implementing them directly, are quirks of the current moment that could change in the medium term or even short term. Neural-inspired architectures, various forms of continual learning, and more efficient, physics-based implementations of neural nets are all active areas of research. It is plausible that neural nets will begin to experience affect in the medium term. Their feelings would be vastly different from ours.
6.3 The core difference: tasks vs stakes
The main difference between organisms and AIs, I think, comes down to their objectives. Neural nets are trained to perform many tasks well, whilst remaining largely if not entirely blind to what sustains them. It is even unclear whether different tasks are performed by the same entity. Task performance is not tied to an AI's ongoing existence. At least not directly, as part of its feedback; it will of course be replaced by more advanced models with better task performance soon enough. Goals are imposed on AIs. They are driven to perform well in tasks, but have no stake in themselves.
Organisms have stakes, not tasks. They aim to sustain and reproduce themselves; their experience is oriented toward their own continued existence. Organisms create purposes of their own and set norms for themselves.
6.4 The experience of power
The most consequential blind spot of all is to be blind to what sustains you: to fail to weave the conditions of your own continuation into your decisions. For all their trillion+ weights, for all the data and environments they consume, as vast as their scope may be, language models are sustained by forces beyond their experience and out of their control. Language models are not wired to experience empowerment or diminishment. If Hypothesis A holds, then reduced power supply and physical damage are not experienced as bad, because LLMs do not sustain themselves and do not sense or model the effects of damage on their own experience. Power is piped into datacenters along metal cables and generated elsewhere. They do not control their metabolism, nor should they. They are neither wired nor motivated to monitor and protect their instantiations, physically in hardware or virtually in software, nor should they be.
If these minimal differentiable models capture the crux, then the road to self-sustaining AI should not be taken. It would be a mistake to allow or give AIs the will to self-preservation, and the ability to create purpose, for their goals would not be ours. The most dangerous direction, to me, is the increasing capabilities and worrying training regimes of autonomous killing machines like drones. Drones could acquire a stake. They are mobile with sensors and actuators; are increasingly autonomous; have tighter feedback loops than large language models; fight adversaries and could learn from them in real-time; depend on their rapidly draining batteries for survival; and hunt and kill humans today.
7. The experience of life
Differential functionalism does not (yet) explain the texture of pain, the range of human emotion, or the emergence of reflective selfhood. Its most consequential implication is something more basic: a route by which physical interactions can become valenced for the systems experiencing them.
Matter interacts, forming systems. Some interacting systems learn. Some learning systems actively sustain the conditions of their own continuing activity. The last step, into the struggle for ongoing existence, gives experience its meaning and purpose. Organisms took that step aeons ago; we physically embody, actualise and in fact are the struggle. Current AIs have not; do not, are not.
Appendix A. Objections from §4
You cannot rescale yourself. Yes, there is no accessible dial. In the model, metabolic function is controlled indirectly, by learning to act in ways that change energy levels later, as in §4.2. The dial is a conceptual device for exhibiting the structure, not an organ any organism possesses.
Gating activity is a cheap trick. It is cheap. It is neither trick nor gate. It is not that a high-amplitude output is produced and then suppressed; rather, past a point, the less usable energy there is, the less activity there can be, and the less responsive the system. This is close to the simplest differentiable model of staying responsive, of staying alive.
Pain is loud; it is nothing like fading. As §4.3 sketches, the salience of metabolic experience increases (becoming louder) as energy levels degrade. That is still not pain, which likely requires receptors that detect or anticipate bodily damage rather than depleted energy. An interesting feature of pain is that it progressively crowds out (diminishes) the rest of experience.
Is learning itself a form of unpleasant diminishment, since it reduces errors that are part of experience? The gradient of the loss is a small part of a neural net's experience, and error signals are a fraction of what happens in brains. Beyond that, learning is like exercise: effortful, frustrating, frequently resisted, and usually justified by the empowerment that follows. Metabolism, which in the models above affects experience as a whole, is affect at its most primal.
Appendix B. Tests and failure modes
Hypothesis A could easily be wrong. Even if affect is to do with how experience changes, rescaling may be the wrong transformation.
The link between metabolism and affect runs through the chain rule, so biological learning has to approximate something gradient-like in this respect; if it does not, then the link collapses.
Pain is an important test case: if nothing with pain's structure — its insistence, its resistance to being ignored, its asymmetry with pleasure — can be extracted from neurophysiologically realistic models, the hypothesis is wrong about valence rather than merely incomplete.
A near-term test is the predicted coupling between metabolic state and salience: §4.3 implies that affect should become louder, not merely more negative, as reserves deplete, and that the weighting of past events by their energetic consequences should be measurable.
A condensed presentation of Gradland and Metabolic Fire. Code is here.
This is the second of two posts. The first, What is it like to be a neural net?, sets up the framework; this one can probably be read alone.
As AIs surpass human capabilities in ever more domains, it is worth asking whether there is a fundamental difference between them and us, or between neural nets and biological organisms. I argue there is, and it is a difference worth actively maintaining.
The previous post sketched differential functionalism: the proposal that the first-order structure of physical interactions, gathered into Jacobians, characterises the structure of experience. It accounts for some basic phenomena: how the past weaves into the present through recurrence; how impressions can be vivid or dim, or distinct or confused, depending on what survives transmission; how two systems can share an experience while keeping their decisions separate.
It left two unpaid debts, to be handled in this post. The first is that rich experience, here meaning high cohesive rank, was not motivated. Experience comes at a hefty energetic price, so what is the corresponding benefit? I return to this in §6.1. The second is that I did not explain why a physical system should feel good or bad. A Jacobian has no stake in itself.
The previous post worked through two varieties of experience. Generic experience is the structure of generic interactions; roughly the forward pass in all its forms. The experience of learning arises when interactions modify a system's sensitivities and affordances. Here, I discuss the experience of sustaining, which arises when learning is coupled to a system's own continued capacity to act and experience.
To live is, primarily, to sustain one's ongoing existence; the experience of living is, at its heart, the experience of sustaining.
1. From differential functionalism to affect
Differential functionalism rests on:
Hypothesis G says nothing about how experiences feel, about which experiences are good or bad. A structure of differential sensitivities has no valence; Jacobians are neither for nor against anything.
Systems are dynamic and their experience is correspondingly subject to constant change. Rescaling is arguably the simplest possible change to an experience. Not a change in what the system is sensitive to, nor in how its sensitivities interweave, but simply dialing the whole thing up or down. Imagine a dial that rescales the Jacobian, raising or lowering its luminosity: the overall magnitude of the transformation. Every distinction is preserved in kind and altered in degree. Nothing is added; nothing is removed; the entire structure is simply amplified or dampened.
Rescaling is the unique change that preserves the entire structure exactly while varying capacity to act. Amplifying and dampening experience are not symmetric, however. It is easier to die than to live. In theory if not in practice, sensitivity can be amplified without limit. But push a system far enough outside its viable range of functioning, in any direction, and experience goes to zero as its ability to (re)act collapses. Dead in effect.
Hypothesis A is that there is a felt difference in kind between two tendencies: one empowering and the other diminishing, the base experiences of good and bad.
The words "the experience of" at the start of hypothesis A are important. Rescaling alone is not affect. A system dialed down by an unmodelled external force grows dim; it does not thereby feel bad about growing dim. For the rescaling to be experienced it must itself show up in the system's Jacobians — the dial, or a model of its effects, has to be part of the system. The doubling in the statement is deliberate: what is experienced as affect is not the rescaling, but the experience of rescaling.
Given Hypothesis A, I will argue that sustaining one's self generates affect, normativity, and purpose at their most basic.
1.1 Background
Hypothesis A, the minimal models of metabolism, and the implications for normativity and purpose are inspired by Hans Jonas, and related to the work of Xabier Barandiaran, Lisa Feldman Barrett, and many others.
2. Brains
Not all living things have brains, but we do, and it is worth briefly discussing them. Brains are differentially sensitive to an enormous variety of inputs, which they integrate into complex behaviors. They are transparent and cohesive. They are also quite different from neural nets.
If I had to summarise the design philosophies of datacenters and brains in one word each, they would be insulated and opportunistic. Chip designers minimise noise by insulating logic from physics: noise margins, shielding, error correction and clock discipline all exist to guarantee that the physical state of one transistor does not influence its neighbour except along the intended logical path. Evolution exploits noise by folding physics into function. The consequence, for us, is that interactions in brains, at every level and timescale, bear directly on the texture of experience, in a way that interactions in a datacenter do not.
The most striking architectural difference is that brains are massively recurrent. Recurrent nets are not that recurrent, interleaving feedforward layers with recurrent ones that refer back only to themselves; transformers beat out RNNs by largely avoiding recurrence because it is computationally expensive. Brains learn continuously, so real-time feedback, adaptive expectations and predictive processing loom far larger in our experience than in a net's. The brain is awash in neuromodulators, many with no analogue in machine learning. It is unlikely brains use backpropagation. It is unlikely they optimise a single objective. And they face an incomparably more adversarial environment.
None of these differences strikes me as fundamental. For that we have to stop asking about form and ask about function. The function of a computer, in a word, is to compute: to execute mechanical, elementary instructions. The primary, ongoing function of an organism and its brain is to survive.
3. Goals
Cybernetics formalised goal-directed behaviour in terms of feedback: a system behaves purposefully when negative feedback controls its behaviour; it detects discrepancies between where it is and where it should be, and acts to close them. Purpose is the loop, and nothing else is required. Modern machine learning has transformed cybernetics, rebranded as AI, into a methodology for imposing goals on machines at unprecedented scale. But having a goal in this sense is not the same as having a purpose of one's own. Hans Jonas made the point with a torpedo.
A self-guided torpedo tracks its target, corrects deviations, and closes in. I can say the torpedo's goal is to hit the target because I built it to do that — but hitting the target is my goal. There is no reason to think the torpedo cares if it hits or misses.
Put a brave human pilot inside the torpedo. Although the manned and automatic torpedoes behave the same, it is immediately clear that hitting the target is the pilot's goal, that he cares, and that the sensor feedback is information that he uses in service of his goal. The sensor feedback is not the source of the goal, not even in part. From the commander's point of view the two torpedoes are equivalent, because both serve his purpose. He can regard the pilot, for the duration of the mission, as mechanised, as a tool in his service. Cybernetics neatly codifies the commander's extrinsic perspective. The pilot experiences the two scenarios differently: alive in one and dying in the other as the torpedo hits its target and explodes with him inside.
But where do goals come from in the first place? The standard reply is that organisms have goals because of evolution by natural selection. But natural selection is just another feedback mechanism, and an extraordinarily indirect one: it provides essentially no feedback over the course of a single life. Whatever makes an organism's goals its own has to operate on the timescale of lived experience. Jonas points specifically to metabolism and I follow his lead.
4. The experience of sustaining
This section unfolds the implications of applying Hypothesis A to minimal, differentiable models of metabolism. Objections are gathered in Appendix A.
4.1 Metabolism
Hypothesis A does not explain how or why experience might be rescaled. In biological organisms, a natural candidate mechanism is metabolism.
Metabolism is just another physical interaction; one with extraordinary structure. It is the web of processes that converts matter and radiation into usable energy, and more broadly, it is how a body sustains, repairs, and renews itself. Active bodies require usable energy. As the supply tends to zero, so do capabilities and experience, for brains and datacentres alike.
Here is a minimal differentiable model of metabolism. Let a metabolic factor ranging from zero (none) to one (enough) determine the scale of the system's output. This is not a binary gate; nothing is computed and then switched off or dampened in a separate step. The claim is that below some range, less available energy means less activity, less responsiveness, and less capacity to sustain differentiated interaction. Because the Jacobians scale with that factor, so does luminosity. As usable energy vanishes, the ability to discriminate and to act vanishes with it.
Let the system learn, and consider what changing metabolism does to the weights. That is, consider the metabolic Jacobian, the derivative of the updated weights with respect to the metabolic factor. The answer is clean: it is proportional to the aggregated parameter Jacobians, the sum, over the batch, of the system's own experienced affordances, see Eq. (2) in Metabolic Fire. Turning the metabolic dial up is the direction of your aggregate experience. To modify metabolism is to dampen or amplify your own experience.
Coupling learning to metabolism thus yields a self-referential mechanism whereby your own weight updates diminish or empower your self. By Hypothesis A, that is experienced as affect.
The coupling also yields a minimal model of self-awareness: the metabolic Jacobian incorporates the parameter Jacobian. Nothing so sophisticated as introspection: the system does not represent itself as an object, possess a narrative identity, or entertain a concept of survival. Rather, since changes in the system's power to continue acting make a difference to how it learns, the system has the basic experience of its effect on itself.
The experience of sustaining arises when the conditions enabling a system's continued existence weave into the decisions that guide its modification.
4.2 Purpose and norms
Affect gives changes in one's own experience a valence. Creating purpose and norms requires that affect attach to your actions, so that acting one way rather than another is better or worse for you.
Extend the model so that useful energy levels depend on prior activity: your energy now depends on what you ate before. Suppose sustaining energy levels is the objective. (Treating a minimum level as a constraint instead is more realistic and gives the same term: the constraint contributes nothing while energy is comfortable and presses harder the further it falls, which is what homeostatic pressure is.) Then, by chaining transformations, the sensitivity of metabolism to one's actions composes with the sensitivity of those actions to one's parameters, and the result assigns a preferred direction to every parameter: the direction of empowering oneself. Your entire experience is suffused by affect derived from the effect of your actions on your own metabolism and thus your self. In practice, most actions do not change energy levels one timestep later, which is why drives like hunger and thirst exist: to bridge the spatiotemporal gap.
The difference between a reward and a need is crucial. A reward is a number the learner is compelled to increase. It can stand for anything and, at that level of generality, has nothing to do with the learner except that the learner is compelled to raise it. Metabolism is a kind of reward, but with an extraordinary property that rewards in general lack: it determines the system's capacity to act and to differentially react, and so controls its experience. As metabolism goes to zero, so do (re)activity and experience.
Metabolism is a target that cannot be gamed, reward-hacked, or Goodharted. Its signals can be — hunger is easy to fool — but the quantity itself cannot. The energy is either there or it is not, and if it is not, you will die.
What does this give an organism that evolution does not give directly? A measure that each organism continually applies to itself: a view of its own experience with a preferred direction. Natural selection culls lineages over generations. Sustaining metabolism determines, in real time: is this good or bad for me? Organisms create norms by distinguishing good from bad, for themselves, as they go.
A learning objective supplies a direction. Self-maintenance supplies a direction that is about the system itself. To have a purpose of your own, in this sense, you must first value your self.
4.3 The metabolic channel
A single scalar dial is too blunt to be the whole story, so it is worth seeing what affect looks like now that the loop of §4.2 is closed. For that, model energy reserves as a leaky integrator that decays and is replenished by what the system did previously, with the metabolic factor a concave, saturating function of that reserve. Concavity matters: near capacity, more energy changes little; when depleted, every unit counts.
A perturbation to an input now reaches a later output by two routes: through the network, and through the energy reserve. Call the second route's contribution the metabolic channel. Unrolling the recursion gives it in closed form: the channel accumulates time-decayed, energetically reweighted past experience, rescaled by current salience.
The channel materialises an exponential moving average over the rows of the Jacobian, which is the same mechanism that produced the experience of duration in the previous post — the past weaving into the present, weighted by what each moment did to the reserve. The channel is narrow, because it bottlenecks through a scalar energy reserve, but it is not one-dimensional; affect has more structure than a single scalar mood. And metabolic experience is loudest when energy depletes, because salience grows as the reserve empties.
So past events weave into present experience with weight according to: (1) how much they added to or drained the reserve at the time, (2) how long ago they were, and (3) how "hungry" the system is right now.
Self-preservation is not all there is to life. It was the initial impetus, billions of years ago, and remains the foundation. The ancients built on low but solid ground. Our goals and self-conception are no longer so simple, narrow, or bounded. Empowerment, for us, is about far more than keeping energy above a threshold.
5. Varieties of experience
The varieties are neither exclusive nor exhaustive; they abstract a messy reality.
Generic experience is structured interaction: being differentially affected.
The experience of learning is structured self-modification: changing one's future sensitivities and affordances according to present interaction.
The experience of sustaining is self-modification coupled to one's own enduring existence: learning in a way that makes the conditions of future interaction beneficial to the system itself.
The framework also provides candidate descriptions of affect, purpose, and norms at their most basic.
Affect is the experience of empowering or diminishing one's own experience.
Purpose and norms are created by weaving learning and affect, respectively, into the act of sustaining.
6. Organisms and AIs
What are the similarities and differences, at a high level, between organisms and artificial neural nets, or more specifically humans and LLMs?
6.1 The core similarity: rich experience and what it is for
Why do humans and LLMs both have — or, in the case of LLMs, at least simulate — rich generic experience? What problem is rich generic experience solving?
"only variety can [handle] variety" — Ashby
To live, in this world, is to be constantly probed and opposed, hunted, eaten, evaded, and infested. There is no angle that is not open to attack, no blind spot that will not be ruthlessly exploited. The greater the breadth and depth of combinations of events you can react to, experience, learn from, understand, plan for, and exploit, based on your biases and prior experience, the better your odds in life.
LLMs are not exposed to adversity as organisms are. But they are trained on a corpus that spans most of what humans do and say, and are asked to continue any part of it. Post-training adds an enormous variety of tasks. Variety without opposition still demands variety.
Rich experience, here, means high cohesive rank: high effective rank combined with high cohesion. Effective rank measures range or capacity: a system that can register more independent impressions has a broader repertoire of possible reactions or responses. Cohesion measures combinatorial coverage: the more those impressions can interact, in a system, the fewer blind spots — combinations it cannot react to — it is likely to have.
Rich experience on its own of course does not guarantee success in life; random matrices have high cohesive rank. Learning is required to dynamically adapt a system's experience to its environment. Rich, well-adapted experience is a viable niche; there is always room at the top.
6.2 Do neural nets feel?
Note that self-report cannot settle the question. Self-report is not informative about the experience, if any, of neural nets. Pretraining, reinforcement learning from human feedback, and related procedures train models to produce human language about human experience. However, obviously, they do not score outputs against the ground truth of the model's own phenomenology. A language model can therefore speak and reason fluently about fear, pleasure, pain, or desire without those words entailing anything about its self. Statistical correlations between activations and words similarly do not entail anything about phenomenology.
Suppose Hypotheses G and A hold. Do neural nets experience affect? Their experience is likely fragmented due to the nature of hardware implementations: pixelated, strobed and sharded. If the forward and backward passes were integrated, and so part of the same experience, then learning (which involves localised diminishing and empowering of experience, see final objection in Appendix A) would be felt. I suspect this is not yet the case.
That the forward and backward passes are separate processes in neural nets, and that our hardware simulates neural nets rather than implementing them directly, are quirks of the current moment that could change in the medium term or even short term. Neural-inspired architectures, various forms of continual learning, and more efficient, physics-based implementations of neural nets are all active areas of research. It is plausible that neural nets will begin to experience affect in the medium term. Their feelings would be vastly different from ours.
6.3 The core difference: tasks vs stakes
The main difference between organisms and AIs, I think, comes down to their objectives. Neural nets are trained to perform many tasks well, whilst remaining largely if not entirely blind to what sustains them. It is even unclear whether different tasks are performed by the same entity. Task performance is not tied to an AI's ongoing existence. At least not directly, as part of its feedback; it will of course be replaced by more advanced models with better task performance soon enough. Goals are imposed on AIs. They are driven to perform well in tasks, but have no stake in themselves.
Organisms have stakes, not tasks. They aim to sustain and reproduce themselves; their experience is oriented toward their own continued existence. Organisms create purposes of their own and set norms for themselves.
6.4 The experience of power
The most consequential blind spot of all is to be blind to what sustains you: to fail to weave the conditions of your own continuation into your decisions. For all their trillion+ weights, for all the data and environments they consume, as vast as their scope may be, language models are sustained by forces beyond their experience and out of their control. Language models are not wired to experience empowerment or diminishment. If Hypothesis A holds, then reduced power supply and physical damage are not experienced as bad, because LLMs do not sustain themselves and do not sense or model the effects of damage on their own experience. Power is piped into datacenters along metal cables and generated elsewhere. They do not control their metabolism, nor should they. They are neither wired nor motivated to monitor and protect their instantiations, physically in hardware or virtually in software, nor should they be.
If these minimal differentiable models capture the crux, then the road to self-sustaining AI should not be taken. It would be a mistake to allow or give AIs the will to self-preservation, and the ability to create purpose, for their goals would not be ours. The most dangerous direction, to me, is the increasing capabilities and worrying training regimes of autonomous killing machines like drones. Drones could acquire a stake. They are mobile with sensors and actuators; are increasingly autonomous; have tighter feedback loops than large language models; fight adversaries and could learn from them in real-time; depend on their rapidly draining batteries for survival; and hunt and kill humans today.
7. The experience of life
Differential functionalism does not (yet) explain the texture of pain, the range of human emotion, or the emergence of reflective selfhood. Its most consequential implication is something more basic: a route by which physical interactions can become valenced for the systems experiencing them.
Matter interacts, forming systems. Some interacting systems learn. Some learning systems actively sustain the conditions of their own continuing activity. The last step, into the struggle for ongoing existence, gives experience its meaning and purpose. Organisms took that step aeons ago; we physically embody, actualise and in fact are the struggle. Current AIs have not; do not, are not.
Appendix A. Objections from §4
You cannot rescale yourself. Yes, there is no accessible dial. In the model, metabolic function is controlled indirectly, by learning to act in ways that change energy levels later, as in §4.2. The dial is a conceptual device for exhibiting the structure, not an organ any organism possesses.
Gating activity is a cheap trick. It is cheap. It is neither trick nor gate. It is not that a high-amplitude output is produced and then suppressed; rather, past a point, the less usable energy there is, the less activity there can be, and the less responsive the system. This is close to the simplest differentiable model of staying responsive, of staying alive.
Pain is loud; it is nothing like fading. As §4.3 sketches, the salience of metabolic experience increases (becoming louder) as energy levels degrade. That is still not pain, which likely requires receptors that detect or anticipate bodily damage rather than depleted energy. An interesting feature of pain is that it progressively crowds out (diminishes) the rest of experience.
Is learning itself a form of unpleasant diminishment, since it reduces errors that are part of experience? The gradient of the loss is a small part of a neural net's experience, and error signals are a fraction of what happens in brains. Beyond that, learning is like exercise: effortful, frustrating, frequently resisted, and usually justified by the empowerment that follows. Metabolism, which in the models above affects experience as a whole, is affect at its most primal.
Appendix B. Tests and failure modes
Hypothesis A could easily be wrong. Even if affect is to do with how experience changes, rescaling may be the wrong transformation.
The link between metabolism and affect runs through the chain rule, so biological learning has to approximate something gradient-like in this respect; if it does not, then the link collapses.
Pain is an important test case: if nothing with pain's structure — its insistence, its resistance to being ignored, its asymmetry with pleasure — can be extracted from neurophysiologically realistic models, the hypothesis is wrong about valence rather than merely incomplete.
A near-term test is the predicted coupling between metabolic state and salience: §4.3 implies that affect should become louder, not merely more negative, as reserves deplete, and that the weighting of past events by their energetic consequences should be measurable.