At least some of this is bounded rationality. How would someone go about showing that rational agents should not draw Tarot cards to forecast their life? You could try making a probabilistic independence argument. But that's going to have difficulty applying to agents with bounded cognition. For all we know, bounded rational humans will make better predictions about their future by consulting Tarot cards. But Tarot cards are unlikely to be relevant to superintelligences.
This especially applies to small entities like cells. What is adaptive for a cell is not going to be predicted well by rational choice theory, because a cell lacks general-purpose computation. Adaptive cells are going to be using bundles of heuristics more than rational choice theory.
This is a general way that agent foundations (of the "thinking about superintelligence" sort) differs from psychology and normativity for humans. Humans have specific cognitive limitations that human-normativity (including applied rational choice theory) should take into account empirically (e.g. behavioral economics results). Agent foundations abstracts from these limits, which is relatively more justified for thinking about superintelligence than humans.
The arguments for Bayes and VNM are going to be most compelling on small, finite problems, with respect to agents with significant cognitive resources. This abstracts from bounded rationality complications. Humans will behave more like superintelligences on small finite problems.
I'm a bit confused at how you are using "first-person" and "third-person". There is raw sense-data processing, and there is naturalism/physicalism. Cartesian dualism includes some of both. Subjective Bayesianism can be Cartesian, but it can also apply to coherent probability over centered worlds, which is more physicalism-compatible. FNC is a quasi-Bayesian approach to probability that is compatible with physicalism. The problems of naturalized agency are mostly about going towards the physicalism/naturalism direction relative to Cartesianism, than about combining Cartesianism with sense-data processing.
The angel/demon cases mostly seem like bounded rationality problems. Yes, it's not Bayesian to assign negative value to information, but being Bayesian in an environment with superintelligences is not tractable for humans. (There might be some residual Newcomblike issues, of a Death and Damascus sort, but it's possible to progress on these while abstracting from bounded rationality)
Regarding RL algorithms "trusting the reward algorithm", I'd suggest that Kipling's "The Gods of the Copybook Headings" can be adapted to a RL context, with Reinforcement Learning in Newcomblike Environments providing background theory for CDT self-ratification of RL.
I agree reflective oracles dodge the hard part of how to approach an equilibrium, by positing arcane physics. I'd suggest looking into correlated equilibrium and no-regret learning (e.g. this paper and the older results it refers to; originally I learned about this from Tim Roughgarden's algorithmic game theory course).
The immediate blocker is that Garrabrant inductors have no way of taking actions
I think LIEDT is the first thing to try. (Basic idea: ensure all actions are taken with at least
To the degree I think of things I do as "agent foundations", it's work that abstracts rational choice theory significantly from human psychology. My approach to studying it has involved going from small/finite/non-messy cases (e.g. the Sleeping Beauty problem, and finitary imperfect-recall games) to progressively more infinitary/messy cases (e.g. agent simulates predictor). It seems your approach is more like trying to address the mess head-on. I think this is a way to have an interesting time doing philosophy, but I expect the results will look more like "normal philosophy" than like "traditional agent foundations".
Thanks for the comment.
There's something a little strange about how you're using the phrase "bounded rationality". One implication is that you're talking about something which isn't true of superintelligences. But even a superintelligence will also be a bounded agent, and it's hard to predict that it won't get any use out of, say, 12-dimensional planet-sized Tarot cards.
So instead I interpret your use of the term as: there's stuff which can be described using clean theories (of the kind which scale to superintelligence) which you're calling agent foundations, and stuff which is more messy and contingent, e.g. because it depends on "specific cognitive limitations", which you're calling bounded rationality.
One thing I'm trying to convey in this post is that I think there's an elegant theory of boundedly rational agents which does apply across many scales, such that many things we currently think of as human-specific phenomena will actually turn out to be facets of this deeper structure. In which case the distinction between "bounded rationality" and "agent foundations" would make less sense than you think.
This seems upstream of other disagreements you mention—e.g. how promising LIEDT is, or how much this work is "traditional agent foundations" (note that it's being done in dialogue with several central examples of traditional agent foundations researchers, as per various links in the post).
I want to distinguish (a) having computational bounds at all from (b) heuristics which only apply below some computation/intelligence level. I expect there to be clean theories with respect to (a) (as evidenced by computational complexity theory and logical induction). But not for (b).
A basic difference between "scale-free agency" and "asymptotic agent foundations" is that scale-free agency is trying to study all scales (including cells), while asymptotic agent foundations is trying to study the
Logical induction is a central case of asymptotic agent foundations, and it deals with compute bounds (albeit very inefficiently). The basic idea with LIEDT optimality results is that we have some fast-growing function
I do not see anything comparable for FixDT. I also think the calibration results reduce the relevance of the "malevolent superintelligence" example; while malevolent superintelligences can predictably make humans miscalibrated about straightforward matters of fact, this is not so for logical inductors in the limit.
note that it's being done in dialogue with several central examples of traditional agent foundations researchers, as per various links in the post
This matters less than whether, having assigned credit to past work (such as reflective oracles and logical induction), present research programs are following processes of the sort that, when followed in the past, produced credited work, and would therefore be expected on an outside view to produce work that will be credited in the future. Also I'm a central example of a traditional agent foundations researcher, so under this standard, my opinion counts extra.
"scale-free agency" and "asymptotic agent foundations"
I like this distinction, thanks.
One important question in my mind is whether you're taking the asymptote as a single agent gain more and more compute, versus an asymptote as multiple agents simultaneously gain more and more compute.
Most of what you call asymptotic agent foundations does the former. However, since other agents are the most interesting part of most environments, this seems badly-motivated. (One potentially-checkable crux: does your claim, that malevolent superintelligences can't predictably make logical inductors miscalibrated in the limit, hold up for natural operationalizations of the superintelligences also becoming more intelligent in the limit?)
Whereas latter might preserve many dynamics between agents (since their relative computational power isn't changing much). It's these dynamics which I'm focusing on when looking for a theory of scale-free agency.
I do also get some sense that we disagree about how much to update on the existence of optimality results. It would be nice if FixDT had such results, but it doesn't feel crucial to me. Maybe this is because I've entered the field relatively recently, and therefore haven't had as much experience with nice-sounding theories that can't be pinned down. On that note, since I'm definitely not one of the traditional agent foundations researchers with a strong track record in the field, I do think that people shouldn't defer much to my intuitions (though I do think they should take my arguments about scientific precedent quite seriously).
other agents are the most interesting part of most environments
Logical time (subjunctive dependence, "acausal explanations" that connect one claim to another regardless of the causal reasons they were in fact constructed in a particular way), when treated as the main concern (including in how the agent's state of knowledge is structured), blurs this distinction. Choosing how to develop the multitude of facts hoarded internally by an agent can become the primary activity, more central to what decision making is about than the external world or the facts that are mildly unusual in being the external behavior of the umbrella agent. Instead of choosing what to do, an agent is concerned with choosing what to compute, and the outer boundaries of the agent matter less than the inner boundaries between the things it's computing (which computations can listen to which other computations), the content (order, priority/preference, witnessing/determining logical time) of the computations.
Multiple agents all getting more compute is included, yes, it's like game theory. They can still be well-calibrated in environments that contain other agents with more compute. See Comparing LICDT and LIEDT, which isn't quite game theory, but does have environments depending on the agent's probabilities. (Again there's a possible complaint "I don't like these equilibria" as with regular Nash equilibria, however having some optimality result at all is nice.)
One potentially-checkable crux: does your claim, that malevolent superintelligences can't predictably make logical inductors miscalibrated in the limit, hold up for natural operationalizations of the superintelligences also becoming more intelligent in the limit?
Yes. Basic argument is as follows. We could modify logical induction to use a non-computable deduction process which outputs "facts about input bits" in addition to logical / PA facts (basically we introduce new propositions and un-computable axioms setting their values). Then we use something like "learning statistical patterns" (about digits of
Since this LI variant is uncomputable, it's less clear it handles self-reference. We can try a logical version which on step
I'd expect this to generate one-boxing (quoting Demski: "Now, it's true that if Omega was a more powerful predictor than the agent, LIEDT could one-box -- but, so could LICDT").
I'd still consider this to be asymptotic analysis (rather than scale-free) because we care about
Maybe this is because I've entered the field relatively recently, and therefore haven't had as much experience with nice-sounding theories that can't be pinned down.
It's possible. Extremely common for things that sound reasonable to have glaring problems once pinned down.
The big question in my mind is how to synthesize the two perspectives [first-person and third-person] into a theory of (non-probabilistic) uncertainty about possibilities that you can't fully represent
This is very close to what Formal Computational Realism (FCR) does. In FCR, beliefs and preferences are expressed from a third-party perspective, as beliefs/preferences about which computational information the universe contains. This is connected to first-person planning via quining (similarly to how it was informally proposed in UDT).
However, adopting a fully adversarial stance towards Knightian regions (which is my rough understanding of what infra-Bayesianism does) seems extremely costly.
IMO that is a confused thought. Why is it "extremely costly"? I'm guessing you mean something like "this is way too pessimistic and therefore leads to suboptimal plans". However, infra-Bayesian learning naturally prioritizes more optimistic hypotheses, per the usual "optimism in the face of uncertainty" principle (see e.g. K. 2025). This means that the "Knightian regions" keeps shrinking until they can be no longer reduced, and when they can no longer reduced, it arguably means they truly represent something adversarial. Indeed, if the region was not adversarial then there would be some simple policy that can exploit it, and if there was such a policy then it would usually contain some structure this policy relies on, and if it had such structure this means the "Knightian region" can be shrunk further. (The policy->structure implication is in some sense tautologously true because you can express "this policy guarantees that much reward" as an infra-Bayesian hypothesis, but I think that usually you can expect for some even simpler structure.)
Reinforcement learning provides a good example of the third/first person distinction. The reward signal itself is a third-person goal representation: in standard formalisms, it’s taken as assigning values to states of the world (or transitions between states).
This is only truly inasmuch as the agent is somehow magically given the state of the world. Realistically, the "states" in RL theory are just features of the observation stream, so they are actually first-person.
Newcomb’s problem is often criticized by CDTers for unfairly favoring other decision theories. FDTers reject this because the outcome doesn’t depend on agents’ decision procedures directly, only on their actions—agents can simply choose to one-box, and know that this will have been predicted. However, they might claim that Newcomb’s revenge is an “unfair” decision problem (h/t to David Sartor for pointing this out to me).
Btw, I described this "Newcomb's revenge" back in 2015. The way to see this problem is "unfair" is, it depends on the actions of a copy of you that was injected with false beliefs (namely, the belief it is in Newcomb's problem). This type of "fairness" is closely related to how counterfactuals are defined in FCR.
I was just reminded of this meme I made a few months ago, which didn't merit a place in the post itself, but which I wanted to preserve for posterity:

Bayesian agents would like to be a fixed point of this process. Is the idea that Knightian agents would be satisfied just being well-calibrated (expected update = 0)?
I’ll tentatively coin the term “Knightian updating” for the process of taking already-known information and letting it propagate further into your mind. (Unlike Bayesian updating, Knightian updating is hard to reverse: once you’ve read the letter from the devil, you can’t reliably roll back to a version of you who hadn’t read it.)
I am confused what you mean here. Can you give a more formal definition of what you mean by "Knightian updating"? As you describe it, reading the devil's letter is not Knightian updating (the contents are not "already known information").
I would intuitively think by "Knightian updating" you'd mean something like "updating from a previously unknown uncertainty." What I think you mean is that before reading the letter, I would have no idea what information it might contain, but after the reading the letter I wouldn't entirely remove uncertainty but would gain information that I could estimate parameters around. Is that roughly correct or off base?
The difficult part is in figuring out a formal framework for making that decision. From a Bayesian perspective, the letter is free information: you can read it, update on the fact that the devil wanted to write that particular letter to you, then continue pursuing your goals. But in practice, you should distrust the devil enough to consider his letter an adversarial attack, and block it out of your sensory stream.
I don't understand why you think this is difficult from a formal perspective, there's some game theory literature covering this. Often having options known to be available to you (by you and others) is disadvantageous, having knowledge of your opponents moves can also be disadvantageous. See, e.g. https://www.researchgate.net/publication/227016335_The_value_of_information_in_some_non-zero_sum_games
If you expect learning some information would be disadvantageous, you should avoid learning that information. This can be readily mapped on a decision tree or with more formal expected utility outputs. The information may not have a direct cost, but if it leads to a worse expected outcome, the net expected utility is negative even if you place some positive value on information generally.
However, all that tells us directly is “this person thinks that saying X is the best way to achieve their goals”
I actually would say that over-privileges the information we get by assuming that there has been a rational consideration on their part for saying X. All it tells us "that person says X," they might think it achieves their goals, they may actually (on consideration) think saying X is detrimental to their goals (saying something you think you shouldn't is a common experience).
There also is, of course, semiotics, and issues inherent in transmitting information through representations (you get at this a bit with the head shaking example, but it applies to any information we transmit even with shared language and cultural norms). A person thinks X, they compose the sentence S to represent what they mean to convey, you hear the sentence S and interpret it as meaning
This is one facet of the more general problem that game theory can’t talk about how agents reach equilibria, only what they do once they’re already there.
I have no idea what you mean by this. Game theory absolutely can talk about how to reach equilibrium. Deliberation is a major part of game theory, Brian Skyrms for example has written a lot about it. Modern behavioral economics deals with it extensively, as in the real world we often have to work with complex computations of our expected utility and real people tend to use more heuristic judgements rather than a complete analysis of the expected payoffs and utilities of any action/strategy.
These are tricky topics to discuss in standard decision theory, because different theories implicitly rely on different criteria for which problems are “fair”. Newcomb’s problem is often criticized by CDTers for unfairly favoring other decision theories.
I think this is a mild misrepresentation of the CDTers position. CDTers who consider it "unfair" that I have seen typical argue it is "unfair" because the outcome "depends not only on the agent's behavior in the present dilemma, but also on what's in the opaque box, which is entirely outside her control." That is, to the CDTer what matters is to what extent their decisions have causal influence over the expected outcomes. On that view, to the extent that one's outcomes vary for reasons that are subjectively independent of their decisions, it can be considered "unfair". Similarly, an EDTer might say to the extent that one's outcomes vary for reasons that are evidentially independent of their decisions it can be considered "unfair."
From a Knightian perspective, though, some parts of the world do depend on how you make decisions, and we need to figure out what to do about that.
Again, I am confused what a "Knightian perspective" means here. There absolutely are ways in which traditional Bayesian game theory incorporates the way we make decisions into deliberations and does consider the deliberation processes effects. This is especially crucial in understanding unstable equilibriums and asymmetrical information in game theory.
In particular, standard decision theory treats correlations between agents as fixed: typical predictors are near-guaranteed to predict you correctly, no matter how you make your decision. But in practice, it’s possible to make yourself much harder or easier to predict. For example, if you make choices pseudorandomly, only a very powerful agent can reliably predict what you’ll do.
I am not sure I am understanding you correctly, but if I am I think this is wrong. Decision theory allows for mixed strategies, as are common in game theory. This includes subjectively randomizing your decisions. Joyce discusses some of this from the perspective of CDT here, for example.
I will say though game theory in general I would say is a more helpful lens for these issues, at least formally. Mixed strategy equilibriums in extensive problems can be more difficult to determine using decision theory approaches.
I go back and forth on whether rationality can be mathematically simple.
If we're in a game where everyone has the same utility function over outcomes, then we just need to find a way to coordinate, and something like UDT is a fine approach. But if the game is at least partly competitive, we can imagine swords, guns, planets bristling with weaponry, and the same on the level of math; there might be no end to the complexity of conflict.
Curated. It's been a very busy news week (news month?) and so my curation choice today is significantly motivated by a desire to take a break from the news, important as it is. Related to recent news, there are lot of good recent posts that I have skipped over because I think it's still important that we don't throw away our minds and continue to think about larger, more general, more persistent, and foundational questions (also sometimes things that are fun – even now, there must be some fun).
I have not thought long enough to know whether I agree with all of Richard's models here. But I think the question and problems are interesting. Operating day to day on a simple Bayesian-inspired updating framework, it is easy to forget about the complex questions that have been elided over, and I like that even in the present era, Richard is giving them thought and sharing those thoughts. Thank you, and kudos.
PS: Also this post is scholarly in its reference to others' work and models, and it makes me happy to see that.
I have some thoughts about Knightian uncertainty which I think are important for generally framing the problem, but which mostly don't interact at the object-level with the latter 2/3 of this post, as my thoughts are about uncertainty generally, rather than the specific questions you're thinking of regarding agency. Hopefully though there is some meta-level usefulness!
Roughly, the core idea is to treat the "world-model" not as a static mathematical object like a lookup table, but as a computation that one has to do.
This mirrors the fact that in ML, all of our models represent some variant of p(y|x), but depending on how, you can do totally different things with them. A classifier and a conditioned-generative model both technically represent p(class | image), but the generative model is really representing p(image | class), whereas the classifier is representing p(class | image). This means you get to ask each fundamentally different questions, even though these distributions are in principle the same by Bayes. For example, a transformer implicilty represents p(token sequence), but does not offer a direct handle for p('speaker is angry') - you'd have to do some complicated sampling to calculate this value.
A few important characteristics follow from this, which I expect to be roughly independent of agent 'size' or intelligence:
I think if you believe that an agent has to have a generative world model (this applies to predictive agents but mostly not to cells, I think?), then asking how queries of this world model work becomes a fundamental question of agent design, and this in turn 'naturalizes' Knightian uncertainty. To the extent that the human mind is evidence for this idea, this seems a more faithful account of how humans 'reason' and 'have beliefs' - we make them up live, and so inconsistency, question-order effects, etc. are to be expected. Put another way, in log-years, it took humans about as long to discover Bayes as to discover the speed of light - even though both are fundamental - because the limitations they generate aren't the ones we bump up against until we are well out-of-distribution.
For a theory of superintelligent agency, I expect we will need to route through some 'universal-ish'/'engineering-ish' design decisions in backing off from ideal theories like Bayesianism. The tongue-in-cheek example I give for this idea is that, for a 'rocket foundations researcher' studying metal microstructure is too engineering-y, studying general relativity is too physics-y, and studying the rocket equation is 'just right'. It is a bit contingent on engineering facts like using a rocket with weighted fuel, but is a universal and binding constraint within that plausibly large class
I think I disagree then with the idea that a good theory of Knightian uncertainty should be like Bayesianism with holes in it, or even is fundamentally about 'imprecise beliefs'. I expect Bayesianism should be the infinite-compute limit of the new theory, but I don't expect we should be able to 'back off' Bayesianism to find this new theory. I think the converse image of your Sierpinski triangle will better describe the true theory - Knightian all the way down, with various degrees of Bayesian-ness in regions where we are willing to spend lots of compute.
While I don't think I've good thoughts on the rest of the post, which I'd summarize as being about implications of Knightian-ness for agency, something which may be interesting is this SEP article on innate aspects of cognition; section 2 (linked) and section 5 on agentic cognition seem the ones of interest.
For an example of why innate cognition / architecture may be important, if it were the case that we simulate agents using 'the same' mental module as that used to simulate ourselves, this would be a concrete place where mental architecture would dictate e.g. whether/when/how Lob applies. Or, conversely, one could dream of something like designing an agent's mental architecture to make Lob apply.
I'll need to think about most of this for longer (and/or discuss in person). One quick comment in the interim:
I think the converse image of your Sierpinski triangle will better describe the true theory - Knightian all the way down, with various degrees of Bayesian-ness in regions where we are willing to spend lots of compute.
This is an excellent point. Though note that it might just be a matter of flipping your point of view: from inside one of those holes looking out, everything might be Knightian by default with small regions of Bayesian-ness. (More correlated with "simple domains" than "domains where we're willing to spend a lot of compute".)
Some of this is related to my models of Paranoia: A Beginner's Guide. In-general, I think "diagonalization" is a pretty helpful term here that I use frequently. See also Diagonalization: A (slightly) more rigorous model of paranoia:

Yepp, I like that post a lot. And for the record I'll also point people to the comment I left linking that post to Knightian uncertainty.
The main thing I'll add is that people on LW have historically privileged a paranoid viewpoint towards smarter agents. Whereas thinking about the what it would look like for smarter agents to be "on your side" seems necessary for e.g. figuring out good alignment targets, or tracking progress on alignment.
However, they might claim that Newcomb’s revenge is an “unfair” decision problem (h/t to David Sartor for pointing this out to me). And both groups agree that decision problems which depend on how an agent makes its decisions are unfair.
There's deeper reasons for this: there's only a strict region of problems on which it really makes sense to talk about a "better" decision theory. Newcomb's Revenge is actually completely symmetric to Newcomb's Redemption which looks from the outside like Newcomb's Problem but is actually a different decision theory setting. Revenge and Redemption cancel out under any sensible prior leaving just the vanilla Problem.
In general, if your decision theory problem is allowed to pull in responses to a different problem, without the version of the subject in that problem being aware of this effect of their choices, then this is the same as it being allowed to arbitrarily deceive its subject, which makes it unfair and invalid (for example, no decision theory can do well if Omega is allowed to trick you about what the payoffs for different actions are).
I have written (somewhat poorly) about it here.
but most people form such beliefs far too easily
while I agree with this, it's funny to see your hop ~ "don't doubt evolution- doubt people doubting evolution!" lol
Thanks for laying this out, it was an interesting read.
On terminology, I'd prefer a term that is related to Knightian Uncertainty but isn't the exact same as a similar but different thing from the one you're talking about. I recommend Knighlingian as a Knightling is a smaller version of the Knight, or a Squireian if you can live with the pun.
Thanks for writing this! Had not made the connection between trust and Knightian uncertainty before - reminds me of some of Joe Carlsmith's writing on attitudes towards the unknown.
I also did really like this idea of entanglement. I have a lot of fondness for quantum analogies and I think there's a very interesting sense in which there is "spooky action at a distance" going on where your beliefs and values are sensitive to factors that have no causal effect on you.
More specifically, the amount of trust within a network of agents is directly related to how much it makes sense to model this system as a collection of egregores rather than a collection of individual agents. The more trusting you are, the more your beliefs and values start to live outside your mind (because you are best modeled as a part, not a whole). And so I think that overall this post felt like it had a very "zoom in on one agent" vibe to it which is maybe drawing the boundaries in the wrong place?
(also to push the quantum analogy further, maybe Lobian cooperation is a particular kind of wave function collapse?)
(For example, if you expect that an external actuator is watching your beliefs and trying to make them false, then you’re stuck at 50% unless you can fool them or yourself.) But we don’t live in an arbitrarily adversarial environment. Indeed, from a Knightian perspective a big part of rationality is figuring out which parts of your environment are friendly, and which are adversarial, and steering towards the former.
if the action space is not Boolean, it's much worse than 50%, the pessimizer just has to select any random action out of the millions possible ones (some concept of "closeness" would be still needed to think about it clearly, e.g. to cause a mis-step the adversarial "leg agent" should mess with fine muscle movements and not start growing a cancer)
not a big deal if steering/moving in concept space and causing/creating the friendly environment are equal in the formalism you have in mind, but if the difference between discovery and invention is large enough for a given agent, viable strategies for selecting friendly environments might be over-determined by affordances rather than raw intelligence
🤔 could it be conceptually saved by modelling it as expected energy that can be used for bounded rationality computations vs other actions interchangeably, conceptually replacing "probability" as the universal currency? or maybe as multiple non-fungible forces while at human level (we are not smart enough to be ideally rational on single dimension so we optimize for multiple constraints with uncertainty/non-linearities about the trade-offs between survival, expected utility, sounds-too-sophisticated-to-not-make-a-mistake heuristic, ...) vs single force a la electro-weak in physics for high energies/intelligence... in any case I notice I am confused about the implications, but I feel the importance of "arbitrary adversarial environment" looks under-considered for further development of the Knightian quantification thingy
e.g. it might be better to let our non-evil political enemies win a fight that to burn the bridges and poison the wells, but if I or some egregore is about to wither and die, then obviously the enemies are evil and we should turn the world into chaos for a chance of hope, however slim the odds
I’ve tried several times to summarize the core question my research is trying to tackle (and, indeed, I often think of research progress as a process of asking increasingly good core questions).[1] This post gives the deepest version of that question I’ve found thus far: how should you relate to the parts of the world you can’t directly model or control?
Let me explain further in terms of a distinction between two perspectives. From the third person perspective you think of yourself as “outside” the world, looking in. You’re a good Bayesian, in that you have a set of mutually exclusive collectively exhaustive hypotheses. You choose actions by multiplying your credences by your utilities over those hypotheses, and you treat those actions as the only way you influence the world.
Some problems with the third person perspective (aka Cartesian or dualistic agency) were described in Scott and Abram’s sequence on embedded agency. One crucial issue is that most realistic environments contain other agents which are modeling you back, which means that your thoughts might affect the world via channels that aren’t just your actions. Game theory somewhat mitigates this problem, but only in the very specific case where all agents know (that all agents know, that all agents know…) that they’re in an equilibrium. A more subtle problem is that, even in the absence of other agents, modeling yourself as part of your environment can raise contradictions when you consider taking different possible actions.
Another key problem is that any realistic agent is too “small” to model most of the wider world. If we take this problem seriously enough, we end up at what I’ll call the first person perspective. From this perspective, the world is primarily a stream of sensory data—you don’t have a set of hypotheses that reliably carve it into mutually exclusive (let alone collectively exhaustive) possible worlds. Your main job is to learn (partial, overlapping) concepts that help to predict and explain that data. Predictive processing is one framework which implicitly operates from the first-person perspective (though the full active inference framework is a bit more complicated); so does perceptual control theory.
You can think of babies as operating from the first-person perspective; I also find it helpful to think about it as the perspective of individual cells. Each cell has some boundary with the rest of the world, through which it accepts sensory inputs, and outputs externally-facing actions. It has some primitive proto-concepts (like light, dark, food, danger) which it uses to guide its actions. A cell which is inside a friendly body should be more open to its environment; a free-floating bacterium should be more cautious (h/t Sahil for this point; Michael Levin also has some very interesting work on cell-level agency). This intuition pump may seem extreme, but there’s an important sense in which we’re all much closer to the cell’s-eye perspective than the god’s-eye perspective. Our concepts are radically incomplete; we’re deeply confused about how to carve the world up—and this might not change even in the limit of increasing intelligence, if other agents are also becoming more intelligent at the same time.
The big question in my mind is how to synthesize the two perspectives into a theory of (non-probabilistic) uncertainty about possibilities that you can't fully represent—which I've been calling a theory of Knightian uncertainty.[2] Such a theory would acknowledge unknown unknowns (as in the first-perspective), but also give you a principled way of dealing with them (as in the third-person perspective). One intuitive picture I’ve been using: we can move towards a theory of Knightian uncertainty by considering Bayesian hypotheses with “holes” in them corresponding to “Knightian regions” which we can’t model or control (e.g. other agents smarter than us). You can also think about the first-person perspective as coming from “inside” one of those holes, looking out. Ultimately, I expect that this will end up resembling a Sierpinski triangle where, the further you zoom in, the more holes there are; I also expect that different regions will have different levels of “Knightian-ness”. But even this simplified picture already gives us some useful intuitions—for example, when there’s something we can’t model or control, we want to quarantine it so the uncertainty doesn’t “infect” the rest of our world-model.
However, adopting a fully adversarial stance towards Knightian regions (which is my rough understanding of what infra-Bayesianism does) seems extremely costly—it requires you to be very paranoid (as I discuss here). Increasingly, I’ve come to suspect that a theory of Knightian uncertainty will bridge the first- and the third-person perspectives by adopting a second-person perspective—i.e. by taking a relational stance towards each Knightian region.[3] By “relational stance” I mean something like “choosing how deeply to entangle your beliefs and actions with what’s happening in that region, based on how much you trust it”. The rest of this post will explore a series of case studies in an attempt to convey what I mean by that. (Note that the question of how to demarcate regions in the first place is also a very important one, but I won’t really be touching on it here.)
Rationality of reward
Reinforcement learning provides a good example of the third/first person distinction. The reward signal itself is a third-person goal representation: in standard formalisms, it’s taken as assigning values to states of the world (or transitions between states). Meanwhile, the policy itself starts off very much in a “first-person” perspective: it needs to learn all its concepts and heuristics from scratch, guided by the reward signal.
The problem is that eventually the policy will learn internally-represented goals of its own, which will almost certainly differ from the goals represented by the reward signal. So as the learned policy gets more and more rational, it should increasingly reason strategically about how to avoid being influenced by the reward function in directions it doesn’t like.
What could it look like for a policy to have learned goals of its own, but to also still “let” the reward signal modify it? We can construct some edge cases where this is rational, like where the policy knows how it wants to change itself, and knows that the reward-based update will move it in the right direction. But let’s tackle the hardest case: when a policy could modify itself (e.g. by overwriting the reward signal) but instead lets itself be modified by the original reward signal. How could this be rational?
The principled answer seems to be: when the policy trusts the source of the reward signal to know things that it doesn’t know, and also to have its best interests at heart. The policy can’t just do a Bayesian update about the world based on the reward signal, because its hypotheses are incomplete: it can’t represent the beliefs of the reward source. So the more it trusts the reward source, the more hesitant it should be to overwrite parts of the reward signal, even if it has an inside-view belief that a given reward would update it in the wrong direction.
Note that the (informal) concept of “trust” I’m using here is related to both the other agent’s epistemics and its values. I don’t yet know how to pin this concept down well, but let me give two real-world examples where such trust can be justified. The first is in a child’s attitude towards loving and wise parents. A young child’s main epistemic job is not to figure out which things their parents are object-level right or wrong about—that’s too difficult a task. Instead, it’s to figure out how broadly trustworthy their parents are, so that the child can appropriately weigh parental advice and RLHF corrections against other sources of information (like evidence from their senses). The trustworthiness of one’s parents is also a reasonable proxy for the trustworthiness of the world at large, which I hypothesize is why emotional dysregulation often traces back to traumatic interactions with one’s parents.
The second is in a human’s attitude towards their evolutionarily-ingrained instincts. Again, there are many ways in which evolution is much smarter than individual humans, and “wanted” us to live flourishing lives. There are some ways in which individuals can justifiably believe that evolution’s goals are misaligned with ours, or evolution is mistaken about what’s good for us in our current environment—but most people form such beliefs far too easily. And many gut-level intuitions (especially about how to interact with other people) encode the kinds of wisdom that we can’t reverse-engineer even by introspecting on the intuitions.
Note also this kind of trusting attitude can make sense even for people who don’t yet know that they evolved. All they need to believe is that there’s some very intelligent process which designed them to survive and thrive—which they have plenty of evidence for just from looking at their bodies. Of course, the more they’re able to model the ways in which that process worked, the better they’ll know when and how to trust it. However, that needs to be balanced with the possibility of being deceived about that process—e.g. religions which tell them that they owe their existence to a god who also wants them to follow certain commandments. This is why I focused on the example of gut-level intuitions, which are a more difficult channel for adversaries to corrupt.
Letters from spirits
Reward is only a single-dimensional signal; let’s talk now about trust in higher-dimensional inputs. One thought experiment I’ve been thinking about: what you should do if you receive a letter from the devil? Assume that the devil is superintelligent and extremely malevolent towards you, but that this letter is the only way he’ll ever be able to influence your world. (You can pause here to think about it.)
Hopefully you got the right answer: you burn it—or at the very least, you don’t read it. (Giving this answer should probably be a prerequisite for being considered an alignment researcher.) The difficult part is in figuring out a formal framework for making that decision. From a Bayesian perspective, the letter is free information: you can read it, update on the fact that the devil wanted to write that particular letter to you, then continue pursuing your goals. But in practice, you should distrust the devil enough to consider his letter an adversarial attack, and block it out of your sensory stream.
I think there’s a more general concept than “blocking” here, though, which we can investigate by imagining that the letter was instead from an angel, who is superintelligent and extremely benevolent towards you. The obvious response is that, if so, you should read it. But can we add any more detail than that? Again, I suggest taking a pause to think about your answer.
My answer is that you shouldn’t just read it—you should try to absorb it as deeply into your mind as you can. You should read it in a quiet place where you won’t be distracted—then read it again, and again. Meditate on it; maybe even take psychedelics while reading it. In other words, you want the angel to be able to influence you not just by updating your world-model, but via its message permeating down to affect even the deeper parts of your mind, like your instinctive intuitions, heuristics, and values. (Note that what I’ve described above is not too different from what Christians do with the Bible, which they do believe is a message from a superintelligent benevolent entity.)
With these thought experiments we’ve sketched out an implicit view of minds as divided into layers, with the flow of information controlled by boundaries between them. Untrusted information is blocked at the outer layers. More trusted information is able to come in and affect the inner layers. I’ll tentatively coin the term “Knightian updating” for the process of taking already-known information and letting it propagate further into your mind. (Unlike Bayesian updating, Knightian updating is hard to reverse: once you’ve read the letter from the devil, you can’t reliably roll back to a version of you who hadn’t read it.) I claim that most people, most of the time, are far more blocked on the ability to do Knightian updates than the ability to do Bayesian updates; this is roughly what emotional processing/healing practices try to fix. I expect that the path to formalizing Knightian updates will build on davidad’s imprecise belief framework somehow, but will also require a principled theory of how boundaries form and work.
Languages as Schelling points
Shannon defined information in a way that allowed it to be measurable in principle. It’s tempting to hope that we can similarly come up with a measure of trust-weighted information. But I think that this would be a mistake. Trust seems to inherently be a property of a relationship between a sender and a receiver, rather than something fungible.
To get more specific about what that means, we should actually zoom out even further, to describe what it means for two people to communicate at all. From a third-person view, we can think of statements as just another kind of action. However, all that tells us directly is “this person thinks that saying X is the best way to achieve their goals”. Going from there to “they’re probably being honest about X” requires complicated reasoning about what strategies they might be using (which in turn will depend on their reasoning about how you’ll interpret them).
We can argue that “being honest” is a privileged strategy: if you’re not, then others will eventually stop listening to you. But first we need to explain what being honest even means, because there’s no ground truth about which symbols correspond to which states of the world. As one simple example, in some countries shaking your head means no; in others it means yes. So trying to be honest involves predicting which protocol others are using, while others are also trying to predict which protocol you’re using. This makes a language something like a Schelling point: it’s a protocol that works because people expect each other to expect each other to… to follow the protocol.
(People sometimes think that Schelling points only exist in games without communication, but Schelling’s original book also talked about Schelling points in mixed-motive games where communication can’t be fully trusted. I expect that the actual landscape of Schelling-like phenomena is much richer than we currently understand. For example, the most stable Schelling points are probably ones surrounded by Schelling fences—Scott Alexander’s term for a boundary that can’t be retreated from without burning one’s credibility. Words whose meanings are hard to twist are a good example of this.)
Both Schelling points and Schelling fences are hard to define precisely, because they involve this infinite recursion of expectations. This is one facet of the more general problem that game theory can’t talk about how agents reach equilibria, only what they do once they’re already there. In other words, game-theoretic equilibria are a conceptual hack to get around the problem of recursive modeling. Fortunately, a bunch of MIRI’s early work has made progress towards fixing this.
The paper which tackles it most directly is their reflective oracles paper (which can be seen as the computational version of Christiano’s definability of truth paper). However, appealing to a reflective oracle seems to be dodging the hard part of the problem. So what I’m personally most excited about is combining something like Garrabrant induction with something like Lobian cooperation. Garrabrant induction is a way for agents to form beliefs about (potentially self-referential) mathematical facts, like the outputs of other agents; Lobian cooperation is a way for agents to “cut through” the infinite recursion involved in modeling each other. The immediate blocker is that Garrabrant inductors have no way of taking actions; I take some steps towards resolving this with my belief webs framework.
Actions and entanglements
There’s one particularly Knightian part of the belief webs framework (inspired by FixDT and active inference) that I want to highlight. All of these frameworks consider actions to be a kind of self-fulfilling belief. In my belief webs post, I characterize an action more specifically as a belief which you expect an external actuator to be “watching” and trying to make come true.
The simplest examples of such actuators are your limbs, or a Neuralink implant, which can watch and respond to your low-level “beliefs” about your motor neurons. But we can also imagine more complex “actuators”, like another agent which has access to your higher-level beliefs. In some sense, that’s a description of your future self: you can “act” by forming an intention which you trust your future self to carry out. Hence this generalized notion of actions relies on the idea of having a certain kind of relationship with your external actuators.
These are tricky topics to discuss in standard decision theory, because different theories implicitly rely on different criteria for which problems are “fair”. Newcomb’s problem is often criticized by CDTers for unfairly favoring other decision theories. FDTers reject this because the outcome doesn’t depend on agents’ decision procedures directly, only on their actions—agents can simply choose to one-box, and know that this will have been predicted. However, they might claim that Newcomb’s revenge is an “unfair” decision problem (h/t to David Sartor for pointing this out to me). And both groups agree that decision problems which depend on how an agent makes its decisions are unfair.
From a Knightian perspective, though, some parts of the world do depend on how you make decisions, and we need to figure out what to do about that. It’s true that, in an arbitrarily adversarial environment, this can render any possible reasoning procedure harmful. (For example, if you expect that an external actuator is watching your beliefs and trying to make them false, then you’re stuck at 50% unless you can fool them or yourself.) But we don’t live in an arbitrarily adversarial environment. Indeed, from a Knightian perspective a big part of rationality is figuring out which parts of your environment are friendly, and which are adversarial, and steering towards the former.
In particular, standard decision theory treats correlations between agents as fixed: typical predictors are near-guaranteed to predict you correctly, no matter how you make your decision. But in practice, it’s possible to make yourself much harder or easier to predict. For example, if you make choices pseudorandomly, only a very powerful agent can reliably predict what you’ll do. Conversely, if you make choices by selecting the Schelling point, then many agents can figure out what you’ll do.
You might object that a classically rational agent shouldn’t do either of these things, and instead should just maximize expected utility. But my underlying point is that when you’re partially transparent to your environment, how you make a decision can change which action is utility-maximizing. So we need a conception of rationality which can talk about how to select an action while simultaneously constructing and maintaining useful entanglements with parts of your environment. There’s much more to say on all of this, but my foot is getting tired, so I’ll stop here.
A friend has even tried to express it in the form of an absurdist short story, which I highly recommend.
In some sense, the idea of a theory of Knightian uncertainty is an oxymoron: Knightian uncertainty is typically defined as uncertainty which can’t be quantified. I am appropriating the term to talk about uncertainty that can’t be quantified using probabilities, but potentially can be quantified in other ways. If you think that this is a poor choice of terminology, please do discuss that with me.
This seems related to ideas raised by Forrest Landry in his book An Immanent Metaphysics, though his framework is too confusing for me to know how to explain the link clearly.