(This post is currently in an incoherent state because it is in the middle of a major revision)
TL;DR: Agent Foundations [AF] pursues a goal so far-fetched that none of the field's progress make us feel closer to it. One might conclude that the goal is unachievable, the question ill-posed. But we also lack a principled refutation of AF: instead of proving that the task is impossible or very unlikely to succeed,we simply fail, and abandon the task. Working toward a principled refutation of Agent Foundations might either indeed refute AF or point to unexplored directions, and either outcome would be welcome progress.
Kant's refutation of all arguments of God's existence
Kant argued that there are three (and only three) related concepts of God, definable on three different levels:
God defined independently of any universe: the set of all logical predicates.
God defined in connection to the existence of a universe, without assuming any of particular property of the universe: the cause of the universe.
God defined in connection with the (perceived) order of our universe: the telos of the (apparent) purposefulness of the universe.
He argued furthermore that none of the three definitions has enough content from which to derive a valid proof of God's existence, and concluded that everyone should stop wasting their time trying to prove God's existence.
I'm not interested in the specifics of Kant's argument; only in its structure:
1. We are dealing with concept(s) that are to some extent[1] definable without reference to specific facts of our universe.
2. There are different levels of "abstraction from the facts of our universe" in which different[2] definitions can be made.
4. It is potentially possible to show that all levels and all possible definitions have been listed.
5. It is potentially possible to show, for each level and definition, that the definition doesn't have enough content for the kind of proof we are looking for
What is the definition of "computation"? Computation can be defined at the level of what a Turing machine can or cannot produce, at the level of what belongs to different complexity classes, and at the level of what can be computed in our universe, and at the level of what a real computer computers.
The concept of computation at the topmost level has enough content to allow e.g. for a proof that Turing-machine-computable and lambda-calculus-computable refer to the same thing.
The second level makes a partition of the things that are computable in the first level. This is simultaneously a refinement of the first level, and something that can be ignored when thinking in the first level: the second definition doesn't make the first one false.
At the third level, the Cobham-Edmonds thesis claims that only polynomial complexity is tractable in our universe. This is a fact about our universe (other universes with different properties are conceivable), and at the same time it is a very general fact that abstracts from the specifics of our universe, such that one might imagine other universes very different from our own that also fulfill this abstract property.
Finally, a the bottom-most level, computation is defined as a thing done by chips or neurons or some other substrate, which imposes the physical laws that govern them onto the previous concept of computation.
Across the four different levels, computation (and thus computability) is, in same sense, the same object, and in some sense four different objects for four different disciplines.[3]
Locating the AF paradigm
In this post we'll treat Agent Foundations as the attempt to provide an abstract characterization of agency. Since nobody seems to know exactly what Agent Foundations should be doing, it seems fair to say that this (alleged) discipline is in search of a paradigm, which we'll define here as a self-contained networks of concepts and proofs. The point of this point is to provide some clarity by reformulation the question so: at which level of abstraction does the paradigm of AF live?
We seem to have no vocabulary to distinguish between the specific results that an agenda has proven, from the general direction of the agenda, for which other results might be better suited.[4]
It seems plausible to me that a solution for alignment could be found using several self-contained components, and that not all of these components are on the same level. For instance, maybe decision theory can be solved at a higher level of abstraction than induction.[5] But it seems very implausible that we can cobble concepts from different levels in a haphazard fashion, and solve any component (say decision theory) with concepts coming from three different levels. So it might be useful to try to create a taxonomy of levels, to locate current agendas within them,[6] and to see what we are missing, if anything.
The guiding question to locate the level in which a solution is is to ask: How different are the universes in which this alignment strategy would work? For instance: would this strategy work in a universe with different physics but the same math?
Concepts
There are two basic aspects to what we are trying to capture under the name of agency: a agent knows the universe and an agent acts in the universe. I will separate the two aspects and, for lack of better terms, speak of inductors and interveners.[7][8]
Further potential concepts arise from the attempts to conceptualize the intervener's interventions, such as coherence,[9] values, pessimization (?), etc., which might be definable in different levels.
I don't know if there are some equivalent concepts for the inductor side. Though not a concept, a natural question to ask is: Is every inductor an intervener, and vice versa?
Finally, we also have concepts like alignment and control, which seem to be definable on different levels. The list is unfortunately far from complete; it is also unclear how we would come to be satisfied about its completeness.
Solving Philosophy and/or Math
The previous levels are all more or less agnostic wrt "solving philosophy", i.e. one could for instance work on the Platonic Space without asking the question of what exactly that means.
But it is plausible that the confusion regarding this situation acts as a blocker in AF, so that working on clearing this confusion could itself be one way of working on AF.[10]
A major source of confusion is that humans are embedded, non-dualistic parts of the universe, and et seem to have what Nagel called the "view from nowhere" able to find out truth that is valid also outside the universe.
The hierarchy between solving Math[11] and solving Philosophy seems unclear. Solving philosophy seems to gesture at something like formalizing the structure of the world of concepts. In some sense this is previous to math (because concepts is a much more general category than mathematical concepts), but the structure might be mathematical. It's possible that at the bottom we don't find a foundation, but that two things that refer to each other.[12]
***
I've realized the first version of this post was trying to make three points at the same time, in quite a confused way. I've separated the three points in three posts, and realized that my original point (the one that remains here), is actually the least important one.
-Kantian refutation (here)
-levels of abstraction
-orthogonal axis about which concepts are essential for an agenda
There's no such thing [as computer science]. [...] At one end you have people who are really mathematicians. [...] In the middle you have people working on something like the natural history of computers-- studying the behavior of algorithms for routing data through networks, for example. And then at the other extreme you have the hackers, who are trying to write interesting software, and for whom computers are just a medium of expression, as concrete is for architects or paint for painters.
There are greatintroductions to Agent Foundations, as well as to MIRI's research, but no real introduction to what the general MIRI direction is, and no map of all the things that are and could be part of AF, even if those things contradict each other in the specifics.
Each concept of an agenda could in theory be on a different level, but if the agenda is coherent, we should expect to find the whole of it nested together, unless the agenda consists on several separable coherent sub-components.
a monograph untangling this coherence mess some more would be valuable. it could do the following things:
specifying a bunch of a priori different properties that could be called “coherence”
discussing which ones are equivalent, which ones are correlated, which ones seem pretty independent
giving good names to the notions or notion-clusters
discussing which kinds of coherence generically increase/decrease with capabilities, which ones probably increase/decrease with capabilities in practice, which ones can both increase or decrease with capabilities depending on the development/learning process, both around human level and later/eventually, in human-like minds and more generally[2]
discussing how this relates to AI x risk. like, which kinds of coherence should play a role in a case for AI x risk? what does that look like? or maybe the picture should make one optimistic about some approach to de-AGI-x-risk-ing? or about AGI in general?[3]
One very speculative way in which this could work out: Kant sketched an argument of how every free will should act super-rationally towards other free wills. Unfortunately the concept of free will doesn't seem to be compatible with our deterministic universe. But what if we could convince an ASI that it is a free will in the Platonic Space, and we could do something like proving meta-ethical theorems that the ASI would be (legitimately) convinced it should obey? Relatedly, it is interesting to note that Kant's defense of free will is much closer to the Block Universe than to Newtonian mechanics.
We are trying to capture the level at which the concept, as useful for our universe, can be defined. Compare:
What sort of claim is [Church-Turing]? Is it an empirical claim, about which functions can be computed in physical reality? Is it a definitional claim, about the meaning of the word "computable"? Is it a little of both? Well, whatever it is, the Church-Turing Thesis can only be regarded as extremely successful, as theses go. Scott Aaronson
Re [18] on free will. I think that the common understanding among compatibilists (who believe that it can exist in a deterministic universe) is:
Free will is the ability to make decisions freely, as an individual with personal and moral considerations that impact your decisions.
As consciousness supervenes on the physical world, the state of the physical world means that you have specific considerations that come to mind and thereby influence your decisions; you only could have chosen differently if the physical world had been different.
However, if the physical world was different in such a way that you chose differently, then the person choosing would not be you, but instead a nearly-identical copy. You could only make the decision that you did and that is what makes you yourself and not your copy.
This also means that decisions which you would never make differently are more demonstrative of free will: they are more emblematic of who you are because you would have to be changed more (in terms of the general impact on your behavior) before you would consider alternatives. Decisions that have a chance of going either way have much less bearing on you as a person.
(This post is currently in an incoherent state because it is in the middle of a major revision)
TL;DR: Agent Foundations [AF] pursues a goal so far-fetched that none of the field's progress make us feel closer to it. One might conclude that the goal is unachievable, the question ill-posed. But we also lack a principled refutation of AF: instead of proving that the task is impossible or very unlikely to succeed, we simply fail, and abandon the task. Working toward a principled refutation of Agent Foundations might either indeed refute AF or point to unexplored directions, and either outcome would be welcome progress.
Kant's refutation of all arguments of God's existence
Kant argued that there are three (and only three) related concepts of God, definable on three different levels:
He argued furthermore that none of the three definitions has enough content from which to derive a valid proof of God's existence, and concluded that everyone should stop wasting their time trying to prove God's existence.
I'm not interested in the specifics of Kant's argument; only in its structure:
1. We are dealing with concept(s) that are to some extent[1] definable without reference to specific facts of our universe.
2. There are different levels of "abstraction from the facts of our universe" in which different[2] definitions can be made.
4. It is potentially possible to show that all levels and all possible definitions have been listed.
5. It is potentially possible to show, for each level and definition, that the definition doesn't have enough content for the kind of proof we are looking for
Computability theory as comparison
Consider the following four levels:
What is the definition of "computation"? Computation can be defined at the level of what a Turing machine can or cannot produce, at the level of what belongs to different complexity classes, and at the level of what can be computed in our universe, and at the level of what a real computer computers.
The concept of computation at the topmost level has enough content to allow e.g. for a proof that Turing-machine-computable and lambda-calculus-computable refer to the same thing.
The second level makes a partition of the things that are computable in the first level. This is simultaneously a refinement of the first level, and something that can be ignored when thinking in the first level: the second definition doesn't make the first one false.
At the third level, the Cobham-Edmonds thesis claims that only polynomial complexity is tractable in our universe. This is a fact about our universe (other universes with different properties are conceivable), and at the same time it is a very general fact that abstracts from the specifics of our universe, such that one might imagine other universes very different from our own that also fulfill this abstract property.
Finally, a the bottom-most level, computation is defined as a thing done by chips or neurons or some other substrate, which imposes the physical laws that govern them onto the previous concept of computation.
Across the four different levels, computation (and thus computability) is, in same sense, the same object, and in some sense four different objects for four different disciplines.[3]
Locating the AF paradigm
In this post we'll treat Agent Foundations as the attempt to provide an abstract characterization of agency. Since nobody seems to know exactly what Agent Foundations should be doing, it seems fair to say that this (alleged) discipline is in search of a paradigm, which we'll define here as a self-contained networks of concepts and proofs. The point of this point is to provide some clarity by reformulation the question so: at which level of abstraction does the paradigm of AF live?
We seem to have no vocabulary to distinguish between the specific results that an agenda has proven, from the general direction of the agenda, for which other results might be better suited.[4]
It seems plausible to me that a solution for alignment could be found using several self-contained components, and that not all of these components are on the same level. For instance, maybe decision theory can be solved at a higher level of abstraction than induction.[5] But it seems very implausible that we can cobble concepts from different levels in a haphazard fashion, and solve any component (say decision theory) with concepts coming from three different levels. So it might be useful to try to create a taxonomy of levels, to locate current agendas within them,[6] and to see what we are missing, if anything.
The guiding question to locate the level in which a solution is is to ask: How different are the universes in which this alignment strategy would work? For instance: would this strategy work in a universe with different physics but the same math?
Concepts
There are two basic aspects to what we are trying to capture under the name of agency: a agent knows the universe and an agent acts in the universe. I will separate the two aspects and, for lack of better terms, speak of inductors and interveners.[7][8]
Further potential concepts arise from the attempts to conceptualize the intervener's interventions, such as coherence,[9] values, pessimization (?), etc., which might be definable in different levels.
I don't know if there are some equivalent concepts for the inductor side. Though not a concept, a natural question to ask is: Is every inductor an intervener, and vice versa?
Finally, we also have concepts like alignment and control, which seem to be definable on different levels. The list is unfortunately far from complete; it is also unclear how we would come to be satisfied about its completeness.
Solving Philosophy and/or Math
The previous levels are all more or less agnostic wrt "solving philosophy", i.e. one could for instance work on the Platonic Space without asking the question of what exactly that means.
But it is plausible that the confusion regarding this situation acts as a blocker in AF, so that working on clearing this confusion could itself be one way of working on AF.[10]
A major source of confusion is that humans are embedded, non-dualistic parts of the universe, and et seem to have what Nagel called the "view from nowhere" able to find out truth that is valid also outside the universe.
The hierarchy between solving Math[11] and solving Philosophy seems unclear. Solving philosophy seems to gesture at something like formalizing the structure of the world of concepts. In some sense this is previous to math (because concepts is a much more general category than mathematical concepts), but the structure might be mathematical. It's possible that at the bottom we don't find a foundation, but that two things that refer to each other.[12]
***
I've realized the first version of this post was trying to make three points at the same time, in quite a confused way. I've separated the three points in three posts, and realized that my original point (the one that remains here), is actually the least important one.
-Kantian refutation (here)
-levels of abstraction
-orthogonal axis about which concepts are essential for an agenda
And the crux is precisely to what extent?
The different definitions can potentially have connections to each other (e.g. one might be a sub-specification of another).
Hackers and Painters
Paul Graham
There are great introductions to Agent Foundations, as well as to MIRI's research, but no real introduction to what the general MIRI direction is, and no map of all the things that are and could be part of AF, even if those things contradict each other in the specifics.
Just an arbitrary, easy to understand example with no connection to my estimates about these questions.
Each concept of an agenda could in theory be on a different level, but if the agenda is coherent, we should expect to find the whole of it nested together, unless the agenda consists on several separable coherent sub-components.
From now on I will avoid using the terms agent or agency.
Arguably, an intervener is just a transducer, but the taxonomy should avoid committing to specific frameworks as much as possible.
I take the idea of separating the different definitional levels of coherence from this exchange:
Mateusz Bagiński:
Kaarel:
One very speculative way in which this could work out: Kant sketched an argument of how every free will should act super-rationally towards other free wills. Unfortunately the concept of free will doesn't seem to be compatible with our deterministic universe. But what if we could convince an ASI that it is a free will in the Platonic Space, and we could do something like proving meta-ethical theorems that the ASI would be (legitimately) convinced it should obey? Relatedly, it is interesting to note that Kant's defense of free will is much closer to the Block Universe than to Newtonian mechanics.
Also mentioned by Kaarel in the previously mentioned discussion of the definitional levels of Coherence.
We are trying to capture the level at which the concept, as useful for our universe, can be defined. Compare: