Note: I wrote this post in 2023, along with the rest of my meta-rationality sequence. I never got around to uploading it, but a recent comment from Vladimir Nesov inspired me come back to it and fill in the few remaining gaps. I still broadly stand by the ideas in this post (some of which were also covered in my posts on Why I'm Not a Bayesian and Towards a Formal Scientific Epistemology), though I don't necessarily still endorse all the specific claims. My recent thinking along these lines tries to be much more precise, and no longer uses the "frame" terminology very much.
Previously in this sequence I described frames as clusters of mental traits. But what does that look like in practice? In this post I’ll explore in more detail how frames are structured, and how that affects the ways we learn and improve them.
The structure of frames
One reason for using the “frame” terminology is that frames can be seen as an attempt to address the frame problem in AI and philosophy: the problem of how we identify which knowledge is relevant to a particular situation. In order to orient ourselves to a given context, we need to identify what to pay attention to, how to interpret it, and what affordances we have in response. This often occurs not via explicit reasoning, but rather via a type of perceptual shift, as in the famous duck-rabbit illusion. Once we snap into seeing the image as a duck or a rabbit, we can identify the most salient features—there’s the eye, there’s the mouth, and so on—and interpret the rest in light of those.
One way of making sense of such shifts is in terms of the predictive processing framework. According to this framework, our perceptions aren’t just determined from the bottom up by sensory inputs, but also from the top down by our expectations. More specifically, predictive processing postulates a hierarchy of generative models which both pass signals up to the models above them, and make predictions about the models below them. If strong enough, these predictions can even override the signals coming upwards. A top-down prediction of “rabbit” or “duck” can therefore force a reinterpretation of features identified at lower levels.
Frames do the same thing at a more abstract level. (Indeed, it seems plausible that frames should be formalized as hierarchical generative models; I haven’t attempted this primarily because I’ve lacked the time to flesh out the details.) While the duck-rabbit image is an example of the same low-level sensory inputs being interpreted in different ways, a more abstract version of this involves the same high-level concepts being interpreted in different ways. Some examples:
"Electricity" means very different things to a quantum physicist, an electrical engineer, a biologist, a meteorologist, and a layperson—while in some sense they're all thinking about the same concept, their frames focus on very different aspects of it, and there may be few or no specific claims about electricity which are relevant to all their frames.
Consider recounting the history of the 20th century from the perspective of different ideologies: each one will tell an overarching story referring to many of the same events, but with the overall narratives being very different.
People used to classify animals and plants based on their physical features; in the biological frame, though, these same features are used for categorization primarily via being evidence about their location on the evolutionary tree.
Consider looking at the same house from the frame of a builder vs an architect vs a real estate agent vs an economist vs an artist vs a city planner. Each of them will have the concepts of windows and doors and many others in common, but they will focus on different features, and assign them different significance.
Here’s an updated version of the diagram from the previous post, which includes concepts being reused across frames.
However, while I’ve focused on concepts for the sake of simplicity, I previously characterized frames as being composed of a much wider range of mental phenomena, which are also reused across different frames. For example, habits and heuristics like seeking out more information or being cooperative might apply in different ways in different scenarios; A scientist might have a set of deep-rooted intuitions about how to set up, carry out and interpret experiments, which they’re unable to communicate explicitly. (And indeed, over the last few decades philosophers of science have shifted towards characterizing even science as a practice where the empirical components of theories need to be understood in the context of a range of background assumptions, which should be studied from an ethnological or anthropological lens.) The most general version of meta-rationality would treat each of these concepts as frames in their own right, for example as portrayed in the diagram below (with arrows omitted for simplicity).
Learning and improving frames
Aside from the ones we learn during early infancy, almost all of our frames are learned by watching other people or absorbing existing knowledge—i.e. from cultural learning. Human children are strongly hardwired for this, and have sophisticated adaptations for inferring who to learn from. Once we learn to talk, and then to read, we can pick up frames much more quickly (as you’re doing now). This isn’t a fully passive process, since we need to seek out new cultural knowledge, and evaluate its trustworthiness; but it relies relatively little on personal experiences. As I’ll discuss later on, it’s also fairly similar to how cutting-edge neural networks are primarily trained: on large corpuses of data generated by humans or other networks.
We seldom learn frames all at once; instead, we gradually build them up out of simpler concepts, skills, and frames. So the more extensive our existing frames, the more concepts we have available to use in constructing new frames. For example, the more we know about history, the more easily we can grasp frames about modern society. However, since different frames interpret and use concepts differently (as described in the previous section), new frames often seem confused or nonsensical at first. In order to acquire a new frame, you often need to be able to “step outside” existing frames to a sufficient extent that you can grasp the new frame on its own terms. I think of this as the core skill behind scout mindset: the ability to engage with other frames on their own terms, without solely interpreting them through the lens of existing frames.[1] Instead, though, we often feel a strong urge to defend our existing frames, and interpret evidence in ways that are consistent with those frames.
In many cases, the new frames turn out to be straightforwardly compatible. In other cases, you need to learn rules for when to apply one or the other. A key aspect of improving frames comes in making the boundaries between them, and the concepts used by them, less nebulous. Nebulosity is the property of not being precisely defined or bounded, described by Chapman by analogy to clouds:
Boundaries: Clouds do not have sharp edges; they thin out gradually at the margin. As you approach a cloud (in an airplane, or on a mountain hike), you cannot say quite when you have entered it.
Identity: It may be impossible to say where one cloud ends and another begins; whether two bits of cloud are part of the same whole or not; or to count the number of clouds in a section of the sky.
Categories: Cirrocumulus shades into cirrus and into altocumulus; clouds of intermediate form cannot meaningfully be assigned to one or another.
Properties: Depending on temperature and density, clouds may be white, gray, blue, or iridescent. There are no specific dividing lines between these colors. Clouds have diverse, highly structured shapes, which cannot be precisely described. First, because the edges are indistinct; and second because the shape is so complex that a full description would be overwhelmingly gigantic even were it possible. Yet meteorologists find useful phrases like “ragged sheets,” “wavy filaments,” “bubbling protuberances,” and “castle-like turrets.”
In order to make frames less nebulous, passively learning from others’ experience is seldom sufficient. Instead, we need to actively focus our attention towards the areas that the frame considers most important to explore. These might be topics which the frame is uncertain about, or areas which fall near the boundary of the frame. In the scientific context, this corresponds to what Kuhn calls "normal science", in which scientists run experiments or solve puzzles in a well-understood domain. Improvements like these will make frames more internally consistent, allowing some of them to assign coherent credences to their claims.
However, it’s important to be clear about the limitations of this type of work, because the strong desire to draw precise boundaries around the “essential nature” of nebulous categories is behind a large proportion of the mistakes made in academic philosophy, especially in the conceptual analysis paradigm.[2] Questions like “is X really Y?” or “what does it truly mean for X to be Y?” are telling signs of this mistake. Why can’t all frames be made fully precise? The short answer: because there’s a tradeoff between precision and practical usefulness, with the categories that are most practically useful often being fuzzier. And even when concepts are well-defined in some frames, it’s hard to unify those with the versions of the concepts used by other frames. Consider for instance the toy dialogue from which David Chapman takes the name of his book on meta-rationality:[3]
A: Is there any water in the refrigerator?
B: Yes.
A: Where? I don’t see it.
B: In the cells of the eggplant.
“Water” is an unusually precise concept in the frame of chemistry, and B is not wrong, but they’re using the concept of water in a way that’s irrelevant to A’s implicit frame. One frame’s version of a concept is often a noncentral example of another frame’s version: technically correct but so irrelevant as to be actively misleading. Since different frames have different central examples of a concept, trying to fully pin down a concept requires unifying all the frames in which it’s used, which is very hard.
Constructing frames
Almost all complex frames applied by almost all people almost all of the time are ones they’ve learned from others around them, albeit perhaps with incremental improvements they’ve added themselves. However, progress depends on constructing new frames—typically via reinterpretingfoundational concepts in an existing frame, in a kind of ontological shift. I’ll talk about two types of ontological shifts. The first is where you have two well-developed frames which seem contradictory, but which can each be viewed as special cases of a more general principle. I’ll call this case-based merging; one intuition for why it’s so useful is that in high-dimensional spaces, like the space of possible strategies, very few conflicts are unavoidable.
Should you be a mistake theorist or a conflict theorist? Yes: they’re each the appropriate strategy in different contexts. Should you be a decoupler or a contextualizer? Yes: the former in philosophy debates, the latter in policy debates. Should you be an extrovert or an introvert? Yes: you should develop both the skill of getting value out of time by yourself and the skill of getting value out of time spent with others. Should you be an empiricist or a rationalist? Yes: you should develop both the skill of patiently observing the world, and the skill of abstractly reasoning about the world. Should you be a deontologist or a consequentialist? Yes and yes. Should you develop the skills described by bayesian rationalism for judging which hypothesis is more correct, or the skills described by meta-rationalism for merging hypotheses together? Yes: the former for when you’re working within a frame, the latter for when you’re dealing with multiple frames. Of course it can be difficult to figure out when to apply one frame versus another, but in many cases the realization that two frames aren’t inherently opposed is by itself a large proportion of the work required in figuring out how to apply each appropriately.
However, the most important intellectual progress doesn’t just come from case-based merging of existing frames, but instead from ontological shifts which reinterpret the foundational concepts used by existing frames. The biggest scientific breakthroughs reshape the categories we use to understand the world; the biggest ideological shifts involve describing familiar features of society in a novel light; the biggest emotional breakthroughs involve realizing that your motivations were totally different from what you thought they were. (Note, however, that this doesn’t necessarily render the old ontology obsolete—e.g. we often still use the ontology of Newtonian mechanics when it’s not worthwhile to use the ontology of relativity. And sometimes a new ontology is mainly useful for making existing knowledge easier to think about—e.g. I believe Feynman diagrams fall into this category.)
I wish I had a recipe for making such breakthroughs. Philosophers of science long searched for a “scientific method” which would, if followed, straightforwardly lead to scientific progress. However, while the process of “normal science” within any given field may be fairly legible, the biggest leaps come from what Kuhn calls “revolutionary science”—the key step of which doesn’t follow any consistent methodology. Kuhn describes how, as normal science progresses, anomalies pile up which cannot be explained within an existing paradigm, until they become too pressing to ignore. During the subsequent period of crisis, scientists question their previous assumptions, generate an alternative to the old paradigm (the key mysterious step), then continue developing it until it’s widely-accepted and they enter a new phase of normal science.
What can be said about the process of inventing a new paradigm? It seems similar in some ways to doing good analytic philosophy—e.g. grappling with what concepts “really mean” and how they might apply in novel contexts, and doing the conceptual engineering required to design new versions of those concepts which span previously-separate frames. But that process is far from straightforward. While being developed, new frames often have gaping holes in them, making them seem absurd to many observers. Choosing to continue developing them despite such seemingly-insurmountable obstacles is often more like a leap of faith than a carefully-weighed decision.
Perhaps the best frame we have for thinking about inventing new frames comes from George Lakoff, in particular his books on the importance of metaphor in human cognition. The omnipresence of metaphor in even our most foundational terminology suggests that a crucial step in constructing a new frame is identifying some kind of abstract similarity with an existing frame.[4] (Just look at the metaphorical terminology from the last sentence: “foundational”, “step”, “constructing”, “frame”!) Identifying those abstractions often involves acts of creativity that toe the line between inspiration and derangement. Stories abound of scientists grasping key insights via wild metaphysical speculation, or by accident when trying to do something totally different, or by drugging themselves to the gills, or by revelation in a dream. As per Koestler:
The history of cosmic theories may without exaggeration be called a history of collective obsessions and controlled schizophrenias; and the manner in which some of the most fundamental discoveries were arrived at reminds one more of a sleepwalker’s performance than an electronic brain’s.
Visions! omens! hallucinations! miracles! ecstasies! gone down the American river!
Dreams! adorations! illuminations! religions! the whole boatload of sensitive bullshit!
Breakthroughs!
How is this manic energy harnessed to actually make intellectual progress? By forcing scientists to engage with reality, especially by making predictions, in accordance with Strevens’ iron rule: resolve disagreements via empirical tests (tenet 5). Before I explore this, however, I want to contrast the meta-rationalist view of science with the bayesian rationalist view, which I think is flawed in instructive ways.
Bayesianism treats hypotheses as pre-existing; you just need to search through the space of all possible hypotheses (formulated in some suitable language) to find the ones that best fit the available evidence.[5] But this “search” metaphor is deeply misleading. Poets don't search for poems; they write them. Architects don’t search for buildings; they design them. Inventors don’t search for better engines; they build them. More specifically, searching has connotations of traversing within a search space of fully-formed answers, and therefore mischaracterizes design processes in which you construct the final answer gradually throughout the process, only landing in the “search space” at the end. Indeed, the whole point of new ontologies is that they supersede your old ontologies in ways you can’t yet imagine, and so trying to define the search space in a meaningful way is most of the work of actually constructing the new ontology.
You can see how misleading the search metaphor is by looking at the absurdities which result from it. Yudkowsky argues that “if your brain doesn’t have at least 10 bits of genuinely entangled valid Bayesian evidence to chew on, your brain cannot single out a correct 10-bit hypothesis for your attention”—a claim which is technically correct but laughably irrelevant, especially in the scientific domains where he tries to apply it. Every single human who ever formulated any non-trivial scientific hypothesis already had billions of times more “bayesian evidence” than they needed from the trillions of photons that entered their eyes throughout their life. Yudkowsky’s argument is analogous to citing the relativistic energy stored in a human body as an explanation of how we’re able to walk up stairs. In other words, the difference between being justifiably confident a scientific hypothesis is true, and justifiably confident it’s false, has nothing to do with how many bits of “genuinely entangled valid bayesian evidence” you have. (Yudkowsky makes a related point here: he very much groks the scale of the gap between humans and Solomonoff induction, just not why it invalidates most attempts to cross-apply conclusions about the latter to the former.)
To be fair, even once a frame has been constructed, it’s non-trivial to recognize it as correct: many scientists rejected relativity (and evolution, and the periodic table, and indeed almost all other big scientific breakthroughs). But the gap between inventing relativity or evolution and recognizing its correctness is a huge one (analogous to solving problems in NP vs problems in P), and bayesian rationalism has very little to say about the former. Even when it comes to the latter, while it does provide some helpful intuitions, it’s also incorrect in important ways, as I’ll discuss in the next post.
Unfortunately, the scout mindset doesn’t distinguish between cases where you’re just trying to fill in uncertainties (e.g. searching for information online), and cases where you actually need to step back from an existing frame in order to be open to a new one. So when I’m trying to be precise I prefer a slightly different term: the “multiple frames mindset”, in which you do your best to always keep track of at least two different frames during disagreements. I’ve adapted this slightly from the “multiple hypotheses mindset” discussed here by Fernandez. Another related term is Sabien’s ‘split and commit’.
A more extreme version of this mistake led logical positivists to attempt to ground all concepts in observational claims, which in turn were intended to be grounded in formal logic. A more recent version of this comes from bayesian rationalism, which claims that we don’t need to know how to translate between ordinary statements and predictions about sense data, since we can instead try every hypothesis and discard the ones which contradict our observations. David Chapman eloquently critiques “try more possibilities in parallel” as a way of getting around rationalism’s problems.
You can see some examples of this in the sections on mental models from physics and chemistry here, which take precise technical concepts and apply them loosely and metaphorically to try to construct useful new frames; this is a very different type of activity from using (or even incrementally improving) those concepts in their original setting.
Similarly, Garrabrant induction aims to find the traders which make the most profitable trades—but doesn’t give us any insight into how to generate those traders in the first place.
Note: I wrote this post in 2023, along with the rest of my meta-rationality sequence. I never got around to uploading it, but a recent comment from Vladimir Nesov inspired me come back to it and fill in the few remaining gaps. I still broadly stand by the ideas in this post (some of which were also covered in my posts on Why I'm Not a Bayesian and Towards a Formal Scientific Epistemology), though I don't necessarily still endorse all the specific claims. My recent thinking along these lines tries to be much more precise, and no longer uses the "frame" terminology very much.
Previously in this sequence I described frames as clusters of mental traits. But what does that look like in practice? In this post I’ll explore in more detail how frames are structured, and how that affects the ways we learn and improve them.
The structure of frames
One reason for using the “frame” terminology is that frames can be seen as an attempt to address the frame problem in AI and philosophy: the problem of how we identify which knowledge is relevant to a particular situation. In order to orient ourselves to a given context, we need to identify what to pay attention to, how to interpret it, and what affordances we have in response. This often occurs not via explicit reasoning, but rather via a type of perceptual shift, as in the famous duck-rabbit illusion. Once we snap into seeing the image as a duck or a rabbit, we can identify the most salient features—there’s the eye, there’s the mouth, and so on—and interpret the rest in light of those.
One way of making sense of such shifts is in terms of the predictive processing framework. According to this framework, our perceptions aren’t just determined from the bottom up by sensory inputs, but also from the top down by our expectations. More specifically, predictive processing postulates a hierarchy of generative models which both pass signals up to the models above them, and make predictions about the models below them. If strong enough, these predictions can even override the signals coming upwards. A top-down prediction of “rabbit” or “duck” can therefore force a reinterpretation of features identified at lower levels.
Frames do the same thing at a more abstract level. (Indeed, it seems plausible that frames should be formalized as hierarchical generative models; I haven’t attempted this primarily because I’ve lacked the time to flesh out the details.) While the duck-rabbit image is an example of the same low-level sensory inputs being interpreted in different ways, a more abstract version of this involves the same high-level concepts being interpreted in different ways. Some examples:
Here’s an updated version of the diagram from the previous post, which includes concepts being reused across frames.
However, while I’ve focused on concepts for the sake of simplicity, I previously characterized frames as being composed of a much wider range of mental phenomena, which are also reused across different frames. For example, habits and heuristics like seeking out more information or being cooperative might apply in different ways in different scenarios; A scientist might have a set of deep-rooted intuitions about how to set up, carry out and interpret experiments, which they’re unable to communicate explicitly. (And indeed, over the last few decades philosophers of science have shifted towards characterizing even science as a practice where the empirical components of theories need to be understood in the context of a range of background assumptions, which should be studied from an ethnological or anthropological lens.) The most general version of meta-rationality would treat each of these concepts as frames in their own right, for example as portrayed in the diagram below (with arrows omitted for simplicity).
Learning and improving frames
Aside from the ones we learn during early infancy, almost all of our frames are learned by watching other people or absorbing existing knowledge—i.e. from cultural learning. Human children are strongly hardwired for this, and have sophisticated adaptations for inferring who to learn from. Once we learn to talk, and then to read, we can pick up frames much more quickly (as you’re doing now). This isn’t a fully passive process, since we need to seek out new cultural knowledge, and evaluate its trustworthiness; but it relies relatively little on personal experiences. As I’ll discuss later on, it’s also fairly similar to how cutting-edge neural networks are primarily trained: on large corpuses of data generated by humans or other networks.
We seldom learn frames all at once; instead, we gradually build them up out of simpler concepts, skills, and frames. So the more extensive our existing frames, the more concepts we have available to use in constructing new frames. For example, the more we know about history, the more easily we can grasp frames about modern society. However, since different frames interpret and use concepts differently (as described in the previous section), new frames often seem confused or nonsensical at first. In order to acquire a new frame, you often need to be able to “step outside” existing frames to a sufficient extent that you can grasp the new frame on its own terms. I think of this as the core skill behind scout mindset: the ability to engage with other frames on their own terms, without solely interpreting them through the lens of existing frames.[1] Instead, though, we often feel a strong urge to defend our existing frames, and interpret evidence in ways that are consistent with those frames.
In many cases, the new frames turn out to be straightforwardly compatible. In other cases, you need to learn rules for when to apply one or the other. A key aspect of improving frames comes in making the boundaries between them, and the concepts used by them, less nebulous. Nebulosity is the property of not being precisely defined or bounded, described by Chapman by analogy to clouds:
In order to make frames less nebulous, passively learning from others’ experience is seldom sufficient. Instead, we need to actively focus our attention towards the areas that the frame considers most important to explore. These might be topics which the frame is uncertain about, or areas which fall near the boundary of the frame. In the scientific context, this corresponds to what Kuhn calls "normal science", in which scientists run experiments or solve puzzles in a well-understood domain. Improvements like these will make frames more internally consistent, allowing some of them to assign coherent credences to their claims.
However, it’s important to be clear about the limitations of this type of work, because the strong desire to draw precise boundaries around the “essential nature” of nebulous categories is behind a large proportion of the mistakes made in academic philosophy, especially in the conceptual analysis paradigm.[2] Questions like “is X really Y?” or “what does it truly mean for X to be Y?” are telling signs of this mistake. Why can’t all frames be made fully precise? The short answer: because there’s a tradeoff between precision and practical usefulness, with the categories that are most practically useful often being fuzzier. And even when concepts are well-defined in some frames, it’s hard to unify those with the versions of the concepts used by other frames. Consider for instance the toy dialogue from which David Chapman takes the name of his book on meta-rationality:[3]
“Water” is an unusually precise concept in the frame of chemistry, and B is not wrong, but they’re using the concept of water in a way that’s irrelevant to A’s implicit frame. One frame’s version of a concept is often a noncentral example of another frame’s version: technically correct but so irrelevant as to be actively misleading. Since different frames have different central examples of a concept, trying to fully pin down a concept requires unifying all the frames in which it’s used, which is very hard.
Constructing frames
Almost all complex frames applied by almost all people almost all of the time are ones they’ve learned from others around them, albeit perhaps with incremental improvements they’ve added themselves. However, progress depends on constructing new frames—typically via reinterpreting foundational concepts in an existing frame, in a kind of ontological shift. I’ll talk about two types of ontological shifts. The first is where you have two well-developed frames which seem contradictory, but which can each be viewed as special cases of a more general principle. I’ll call this case-based merging; one intuition for why it’s so useful is that in high-dimensional spaces, like the space of possible strategies, very few conflicts are unavoidable.
Should you be a mistake theorist or a conflict theorist? Yes: they’re each the appropriate strategy in different contexts. Should you be a decoupler or a contextualizer? Yes: the former in philosophy debates, the latter in policy debates. Should you be an extrovert or an introvert? Yes: you should develop both the skill of getting value out of time by yourself and the skill of getting value out of time spent with others. Should you be an empiricist or a rationalist? Yes: you should develop both the skill of patiently observing the world, and the skill of abstractly reasoning about the world. Should you be a deontologist or a consequentialist? Yes and yes. Should you develop the skills described by bayesian rationalism for judging which hypothesis is more correct, or the skills described by meta-rationalism for merging hypotheses together? Yes: the former for when you’re working within a frame, the latter for when you’re dealing with multiple frames. Of course it can be difficult to figure out when to apply one frame versus another, but in many cases the realization that two frames aren’t inherently opposed is by itself a large proportion of the work required in figuring out how to apply each appropriately.
However, the most important intellectual progress doesn’t just come from case-based merging of existing frames, but instead from ontological shifts which reinterpret the foundational concepts used by existing frames. The biggest scientific breakthroughs reshape the categories we use to understand the world; the biggest ideological shifts involve describing familiar features of society in a novel light; the biggest emotional breakthroughs involve realizing that your motivations were totally different from what you thought they were. (Note, however, that this doesn’t necessarily render the old ontology obsolete—e.g. we often still use the ontology of Newtonian mechanics when it’s not worthwhile to use the ontology of relativity. And sometimes a new ontology is mainly useful for making existing knowledge easier to think about—e.g. I believe Feynman diagrams fall into this category.)
I wish I had a recipe for making such breakthroughs. Philosophers of science long searched for a “scientific method” which would, if followed, straightforwardly lead to scientific progress. However, while the process of “normal science” within any given field may be fairly legible, the biggest leaps come from what Kuhn calls “revolutionary science”—the key step of which doesn’t follow any consistent methodology. Kuhn describes how, as normal science progresses, anomalies pile up which cannot be explained within an existing paradigm, until they become too pressing to ignore. During the subsequent period of crisis, scientists question their previous assumptions, generate an alternative to the old paradigm (the key mysterious step), then continue developing it until it’s widely-accepted and they enter a new phase of normal science.
What can be said about the process of inventing a new paradigm? It seems similar in some ways to doing good analytic philosophy—e.g. grappling with what concepts “really mean” and how they might apply in novel contexts, and doing the conceptual engineering required to design new versions of those concepts which span previously-separate frames. But that process is far from straightforward. While being developed, new frames often have gaping holes in them, making them seem absurd to many observers. Choosing to continue developing them despite such seemingly-insurmountable obstacles is often more like a leap of faith than a carefully-weighed decision.
Perhaps the best frame we have for thinking about inventing new frames comes from George Lakoff, in particular his books on the importance of metaphor in human cognition. The omnipresence of metaphor in even our most foundational terminology suggests that a crucial step in constructing a new frame is identifying some kind of abstract similarity with an existing frame.[4] (Just look at the metaphorical terminology from the last sentence: “foundational”, “step”, “constructing”, “frame”!) Identifying those abstractions often involves acts of creativity that toe the line between inspiration and derangement. Stories abound of scientists grasping key insights via wild metaphysical speculation, or by accident when trying to do something totally different, or by drugging themselves to the gills, or by revelation in a dream. As per Koestler:
The history of cosmic theories may without exaggeration be called a history of collective obsessions and controlled schizophrenias; and the manner in which some of the most fundamental discoveries were arrived at reminds one more of a sleepwalker’s performance than an electronic brain’s.
Or as Ginsberg put it:
How is this manic energy harnessed to actually make intellectual progress? By forcing scientists to engage with reality, especially by making predictions, in accordance with Strevens’ iron rule: resolve disagreements via empirical tests (tenet 5). Before I explore this, however, I want to contrast the meta-rationalist view of science with the bayesian rationalist view, which I think is flawed in instructive ways.
Bayesianism treats hypotheses as pre-existing; you just need to search through the space of all possible hypotheses (formulated in some suitable language) to find the ones that best fit the available evidence.[5] But this “search” metaphor is deeply misleading. Poets don't search for poems; they write them. Architects don’t search for buildings; they design them. Inventors don’t search for better engines; they build them. More specifically, searching has connotations of traversing within a search space of fully-formed answers, and therefore mischaracterizes design processes in which you construct the final answer gradually throughout the process, only landing in the “search space” at the end. Indeed, the whole point of new ontologies is that they supersede your old ontologies in ways you can’t yet imagine, and so trying to define the search space in a meaningful way is most of the work of actually constructing the new ontology.
You can see how misleading the search metaphor is by looking at the absurdities which result from it. Yudkowsky argues that “if your brain doesn’t have at least 10 bits of genuinely entangled valid Bayesian evidence to chew on, your brain cannot single out a correct 10-bit hypothesis for your attention”—a claim which is technically correct but laughably irrelevant, especially in the scientific domains where he tries to apply it. Every single human who ever formulated any non-trivial scientific hypothesis already had billions of times more “bayesian evidence” than they needed from the trillions of photons that entered their eyes throughout their life. Yudkowsky’s argument is analogous to citing the relativistic energy stored in a human body as an explanation of how we’re able to walk up stairs. In other words, the difference between being justifiably confident a scientific hypothesis is true, and justifiably confident it’s false, has nothing to do with how many bits of “genuinely entangled valid bayesian evidence” you have. (Yudkowsky makes a related point here: he very much groks the scale of the gap between humans and Solomonoff induction, just not why it invalidates most attempts to cross-apply conclusions about the latter to the former.)
To be fair, even once a frame has been constructed, it’s non-trivial to recognize it as correct: many scientists rejected relativity (and evolution, and the periodic table, and indeed almost all other big scientific breakthroughs). But the gap between inventing relativity or evolution and recognizing its correctness is a huge one (analogous to solving problems in NP vs problems in P), and bayesian rationalism has very little to say about the former. Even when it comes to the latter, while it does provide some helpful intuitions, it’s also incorrect in important ways, as I’ll discuss in the next post.
Unfortunately, the scout mindset doesn’t distinguish between cases where you’re just trying to fill in uncertainties (e.g. searching for information online), and cases where you actually need to step back from an existing frame in order to be open to a new one. So when I’m trying to be precise I prefer a slightly different term: the “multiple frames mindset”, in which you do your best to always keep track of at least two different frames during disagreements. I’ve adapted this slightly from the “multiple hypotheses mindset” discussed here by Fernandez. Another related term is Sabien’s ‘split and commit’.
A more extreme version of this mistake led logical positivists to attempt to ground all concepts in observational claims, which in turn were intended to be grounded in formal logic. A more recent version of this comes from bayesian rationalism, which claims that we don’t need to know how to translate between ordinary statements and predictions about sense data, since we can instead try every hypothesis and discard the ones which contradict our observations. David Chapman eloquently critiques “try more possibilities in parallel” as a way of getting around rationalism’s problems.
Originally from Winograd & Flores' Understanding Computers and Cognition.
You can see some examples of this in the sections on mental models from physics and chemistry here, which take precise technical concepts and apply them loosely and metaphorically to try to construct useful new frames; this is a very different type of activity from using (or even incrementally improving) those concepts in their original setting.
Similarly, Garrabrant induction aims to find the traders which make the most profitable trades—but doesn’t give us any insight into how to generate those traders in the first place.