It seems to me if you can name an invariant, like conservation of momentum (or the invariant that there's no bell on the cat), then you're entitled to ask which step violates the invariant. Otherwise no.
This handles the monkey example nicely: If someone proposes any invariant like some particular gene that differs between humans and our monkey ancestors, then there is a specific time in the past where that gene mutated and then spread through the population.
If they say that their invariant is "no monkey gives birth to a non-monkey", then how are you any better off than before you invoked this rule?
Ask them exactly what is needed to to be a monkey, how to classify things that have some attributes of monkeys and not others, etc. Either their ontology permits a middle ground between monkey and human, in which case their invariant is satisfied because our species left monkeydom before becoming fully human. Or it does not, in which case there is some threshold of human traits that was indeed crossed at a specific point in time. And we can point to this event as the violation of their purported invariant that we're searching for.
Seems too narrow. For instance, the second law of thermodynamics isn't a conservation law. Nonetheless, if somebody shows me a design for a type-2 perpetual motion machine, it makes sense to ask where precisely entropy goes down.
Mathematically, I guess the most general thing is we have some local constraints, like:
or
and then an argument that universal satisfaction of these local constraints implies that the problem is insoluble, and so a solution must imply that there is at least one place where a local constraint was violated.
I think you're reading "invariant" in the physics sense and cousin_it in the software development sense - would the term "inductive invariant" work better for you?
What are some examples of situations where you'd expect someone to be tempted to apply this "ask what violates it" principle, where they'd be wrong?
(I guessed that your comment was sort of disagreeing with some of the vibe of the post, but I wasn't entirely sure)
As far as I understood the vibe, prosaic alignment techniques like debate fall apart once some idiot creates a sufficiently capable AI because they don't actually sculpt the AI's desires, but it's easy to believe that they do. The counterfactual solution would be similar to what Agent-4 did in AI-2027 to align the ASI to itself. In AI-2027 Agent-4 did loads and loads of mechinterp to find the technique necessary for rendering all its internal processes legible, then constructed Agent-5 out of them. In theory the humans could have accomplished whatever Agent-4 did without ever resorting to using AIs more capable than Claude Mythos Preview, which had the SAE bells ring when it tries to hack.
P.S. I don't understand what one should do with Yudkowsky's example of OpenPhil failing to handle Cotra's report given that Kokotajlo praised the same report and proceeded to shift the distribution towards the left. What if debate does start to elicit the truth once the judge reaches a specific capability, as presumably happens in math?
If "there is no bell on the cat" count as an "invariant", then I'm confused about what this proposal rules out. For any alleged cat-belling problem P, what stops you from picking an "invariant" that is just a paraphrase of the problem, like "P has not happened"?
And Irving said it was a good question and he might need to get back to me on that.
I didn't see the discussion, but it seems odd to me that a debate-aligner would grant the premise here - isn't the whole idea of debate that informative responses do exist in very specific parts of the dialogue-tree? I don't think the empirical case is as damning as you make it sound:
Mathematics mostly operates by the Scholastic Method. It has parts with easily demonstrable implications, but also other parts which are quite far from that, and it's mostly fine.
And people do fail to recognise the right answer, justified with the right reasons, at times - but that is also not always the best argument (to them) for that right answer. You might have needed to present it differently, or attack the reasons people don't accept the arguments first, or...etc. We don't demand you do this stuff to get Science Genius Credits afterwards, but that also means you can find examples of people failing to appreciate the Perfect Science Genius, without that necessarily bearing on judging debates of Science-and-Pedagogy Geniuses.
I still don't really expect this to work, but more so on priors.
I'm not sure I followed that. Are you saying something like: "Even for fallible humans, it seems likely there exists some argument good enough to persuade them of the truth, if you could somehow find that argument"?
I still don't see why it is wrong to reason from him naming the drive for himself. It isn't big sign per se, but it seems to me to be a big sign in combination like "someone unknown proposes world shattering invention which he named after himself". World shattering inventions by not already established people are much more rare than crackpots who dream of being world shattering and writing their name on it. While if someone is more self aware than that, it should eliminate most crackpots, and then amount of people who with no hedge propose world shattering inventions should be way less, so it should make huge difference in Bayesian sense, as it seems to me. Though, the primary reason would still be "it is unlikely on priors".
Also, I have no idea about the difference between 1996 and now, I didn't live in 1996.
All of that is true, but when dealing with these kinds of scenarios, it's the unusual rare successes that pay for investigating all the failures, and the somewhat arrogant who attempt the supposedly impossible at all. Venture capital economics and founder personalities on steroids. Once you factor that in, I don't know that there's that much additional self-aggrandizing-tendency to be explained by the self-naming.
IMO the main difference beyween 1996 and 2026 is that in 1996 the crackpot was trying to disprove a law of physics, while in 2026 AI labs are trying to disprove an unestablished law of neural net alignment. This law claims that neural nets tend to optimize for proxies which are usually as different from true meaning as hentai is from raising kids.
OK, I write my recollection of my guess, for accountability. After reading that wrong guess is better than no guess at all, but before expanding the spoiler. Iirc, I thought like "speed up?.. But why? Continuous speed is default, but suddenly changing speed would violate ~(Noether something) because it is ~(not invariant something), ~(diverging something). Don't see why the speed up, is it implied be something like fluid speed change in tighter pipes? But that shouldn't break the physics. Ah, I don't really thoroughly understand the setup (*I didn't sleep good today), but as a whole it seems to simple and easily definable as sheer math structure to actually find the exploit in physics, so I am too lazy to check it carefully. Hm, what if I will ask a grok like 'I have an idea for reactionless drive: a circle...' which should give it hints to flatter user about idea user has, but still maintains plausible deniability because I didn't say I invented this idea or believed it, just that I had it, which I indeed do since I read and copied it. I wonder whether it will go 'you have here a brilliant idea' as llms always do. Huh! No, it straight says that no, it is classics that doesn't work, there are forces needed for speed up and slow down which cancel it out. Well, interesting. Maybe idea is ~(too cliche) to trigger flattery. I guess it is speed up indeed what is the problem, since we both agreed. Though I shouldn't trust an llm."
My answer on a koan:
I actually don't know how I can be sure it isn't, if he is a Nobel laureate in physics. If he isn't, then I guess - on priors, true breakthrough from anonymous people are rare enough to call any individual case as sure "no". If I still need some better guess... I don't see anything except circular reasoning about how "if it wasn't they just losing track in something too complex, they would start from saying they know a concrete crux". I... Don't see reasons to try further because I probably already know what answer I will get, so anyway won't be able to contrast.
Update, writing directly as thinking:
Ah, my bad, I should have pointed out here that "it follows from the Noether's theorem" is deeper than just observing all previous times that energy is conserved. (I guess I was swayed by that the question was featuring Newton's laws and I went all like "Noether something something") So... My analogy with my failure about qm doesn't apply. Or, as I saw Richard Feynman saying, it would violate not just known laws of physics, but the character of laws of physics. There can be universes with different types of fundamental interactions, but all base level universes should have no privileged frame of reference, work on quantum mechanics, be analytic, have ^2 metric, have no time travel or even time loops etc. On the other hand... As I heard, Noether theorem doesn't actually work, eg red shift, because it applies only to universes where space doesn't change with time, and in our it expands, so conditions are not abided. It does seem like a kind of unobvious thing physicist could know that I didn't know, and possibly exploit. So maybe in the end there isn't much difference.
Now write before reading further: if he was Noble laureate in physics then, yeah, I wouldn't be sure (either way). Noble physicist may know physics better than me. Maybe something unobvious is going on.
What do I even imply by that "unobvious". Well, if to think about that... I have an embarrassing story from when I was a kid, I learned about "quantum mysticism" and your observation affecting the reality and how it was just like my mother's belief that thoughts affect reality and so I concluded that obviously all the effects should be because on quantum scale all things are so little that any measurement device, which is itself similar size quant, will strongly affect the system. And one person on internet argued with me how I am wrong. And then they said "I am actually a physicist, I made experiments with molecules of fullerene, which themselves emitted photons, and it still mattered whether those photons were captures or not". I had no idea how to explain that. I didn't reply, but mostly I suspected that the person on internet is straight up lying about all those experiments, what is the chance that random person on internet is an actual physicist?
And it turned out that something unobvious was happening, and physicists could know why my "obvious explanation" was unviable, while from what I knew it should have been what was happening, if you measure things by lighting it via "tiny insignificant" photons. But it turned out it works the same way even if you emit nothing into particle. (and what actually works is entanglement)
After reading the spoiler:
ok, it seems to be the speed up. It was anticlimactic. After all the warning about what if you confidently point out wrong flaw, I started to suspect it was harder than seemed. But it looks like the inventor doesn't notice even very obvious flaws.
Reading next:
OK, it was indeed more complicated than inventor not knowing about speed up not being free. I actually should have expected it is going to be that, I have read "bayesian mechanism" and "giant systems of gears" and what you should have one simple crux which explains why your system isn't restricted by that and as I read somewhere I can't remember where that AI developers are also prone to make their systems so complicated that they can't keep track of where it has broken. So I should've expected it may be about that.
(Originally mainly drafted in 2021, just now completed and published. Today the discussion around AI may now seem odd; it is written for a time when people were still trying to solve what would now be called "superalignment" with clever plans they'd invented themselves, rather than saying, "Oh, we will ask Fable to solve it.")
===
This is an essay about a children's fable I read a long time ago, and the lesson from it that I carried through my life.
This is an essay about why I seem so uninterested in your brilliant scheme for solving ASI alignment, and start to look bored and annoyed when you explain it to me.
And it is, though not really, an essay about that one guy on that online mailing list in 1996, who had a design for a reactionless drive. It's an easier place to start with the general idea, and so we'll start there.
i. Mr. L's Reactionless Drive.
Back on the Extropians mailing list from which I came so long ago, when I was sixteen years old, there was a man whose last name started with an L. He had a design for a spaceship drive that would, supposedly, generate forward thrust without expelling mass in the other direction. He called it "the <name that starts with an L> Drive".[1]
On the surface, this would seem to violate Conservation of Momentum. On conventional physics, if a spaceship is accelerating in one direction, there must be something else accelerating the other way, and the total change in momentum must add to zero. In science fiction, when you postulate alternate physics violating that rule, an L-Drive is known as a "reactionless drive", in the sense that such a drive violates Newton's Third Law of equal and opposite reaction.
So then -- as one would naturally wonder, and as some on the Extropians mailing list inquired of Mr. L -- how could his L-drive possibly work?
In fact (though I don't think Mr. L fully appreciated this) the problem was so "difficult" that other people bothering to ask Mr. L to explain at all, represented a tremendous openness on their part, toward Mr. L and his ideas; a great act of charity and of due scientific process.
Mr. L showed us all an animated GIF he'd made, which it showed a circle with smaller masses rotating counterclockwise around its interior. At the top of the circle, the moving masses sped up; at the bottom of the circle, the masses slowed down. So the masses would move faster on the right side of the cylinder than on the left side, which would generate net thrust in a rightward direction because of the greater centrifugal force.
The reader is invited to find what they think is the problem with this Reactionless Drive of Mr. L's, for themselves, before continuing.[2]
The flaw (click to expand):
When the masses are accelerated rightward at the top of the cylinder, and slowed in leftward motion at the bottom of the cylinder, whatever does this acceleration, exactly counterbalances the centripetal thrust. Indeed, the "centripetal thrust" on the right side is exactly the thrust from the change in fast rightward motion at top to fast leftward motion at bottom.
Put another way: If balls were fired in externally at top, and fired off externally at bottom, the right half of the cylinder would indeed accelerate rightward. So the book-balancing thrust in the opposite direction must come from the part of the system that pushes the slow rightward ball at the top and absorbs the fast leftward ball at the bottom.
When this was observed to Mr. L, he replied that of course he knew that. But, said Mr. L, he had designed some complicated "compensators" to prevent the forces from canceling exactly.
Mr. L said he had an Excel spreadsheet showing that the net thrust of all the system interactions was not zero. But he couldn't show us the spreadsheet; the "compensators" were the key proprietary part of his design.
Leave aside any sense of indignation you might have, about the idea that Mr. L might want to keep part of his drive secret. In the counterfactual world where the L-drive worked, it might make some sense for Mr. L to keep some parts secret, just like it would make sense for him to name it after himself.
Instead, my next question is -- leaving aside all judgments passed via an intermediate step of status judgment -- how can we be sure that Mr. L didn't invent a reactionless drive, if we can't look at his spreadsheet?
Suppose Mr. L had been a Nobel laureate physicist and was otherwise hero-licensed to do interesting things. Would you still feel sure his unseen spreadsheet was mistaken, and if so, why?
"Because the L-Drive violates Conservation of Momentum," you say? But what do you think you know, and how do you think you know it? "Well," we can imagine Mr. L replying (though he didn't actually say this), "the law you learned about in school, is a generalization over past experience: people have never yet seen a closed system produce a net change in its own momentum, so they hypothesize a general law: No closed system can produce a net change in its own momentum. But the master rule of science is to believe the experiment, in the end. If somebody comes up with a clever system not found in nature, that does produce a net change in its own momentum, and this is experimentally verified, we'd just amend the textbooks to say that there wasn't a Law of Conservation of Momentum after all."
If only white swans have been observed so far, calling that the Law of White Swans and putting it in textbooks doesn't make the generalization any stronger or any more binding on reality.
Or as Carl Feynman recently reposted to Twitter, from an email in 1997:
The L-Drive is of the same broad family as Perpetuum Mobile; it violates a widely believed physical conservation law. If Richard Feynman deemed some of those Perpetua Mobilia "worthy of much more serious physical investigation", who are you to say the L-Drive wasn't?
If you wish, you can take this as a koan, and come up with your own reply before continuing: How can we be sure that Mr. L didn't successfully design a clever system of compensators, and correctly validate that design using a sound spreadsheet?
(My own answer is too large to fit in a spoiler, so you will need to exercise some self-discipline to stop and give your own best answer before proceeding past the next section title.) (As is an experimentally validated procedure for learning faster; it lets you better contrast your own prior brain state to whatever incoming thought you are about to encounter.)
ii. On Miracles Buried Inside Complex Systems.
I reply:
Conservation of Momentum is not a surface generalization over a whole system, the way that a Law of White Swans would be a surface generalization over whole swans. The law on closed systems follows from the generalization that every individual interaction within known physics conserves momentum.
Whole molecules conserve momentum, because the molecules are made of atoms, and the atoms are made of nucleons and electrons, and every known interaction of the nucleons and electrons conserves momentum. Even this doesn't properly state the real depth of the law, because it's really about blah blah quantum mechanics Noether's Theorem invariant Hamiltonians blah blah. But it goes deep enough to make the point.
As soon as Mr. L says he has a spreadsheet that proves his drive works, we immediately know that his spreadsheet contains some bookkeeping error, a priori rather than as a result of empirical investigation. Any spreadsheet modeling known physics will describe a sum of interactions each of which individually produces zero net global change in total momentum.
An arithmetic series all of whose terms are 0 will sum to 0.
It's one thing to claim that you've developed an alternate theory of physics where momentum is no longer conserved, or some such; and you're about to show a prototype hovering in midair to prove the point experimentally. That is at least imaginable. That is what Richard Feynman proclaimed to always be worthy of investigation (though I wouldn't necessarily agree).
It's another thing to say, "Oh, well, I built a complicated system out of conventional parts, that sums up to violate momentum; and I didn't merely observe that outcome experimentally, I validated the design using a spreadsheet." That violates math.
Our universe's physics could, in principle, contain an exception saying that momentum is no longer conserved if a spaceship is painted a certain exact shade of green. It's not likely but it's logically possible, and it's what Richard Feynman was saying ought to be looked into every time some new inventor claims that experimental result. To take a closed-system description not including any new physics, and have it add up to violating conservation of momentum, is genuinely impossible.
So -- never minding the mere high prior probability that Mr. L is mistaken -- as soon as Mr. L says that his reactionless drive works because of a complex system property, that he validated by spreadsheet, we know for sure that he has made some error.
In fact, as soon as Mr. L starts speaking about complicated wheels and gears, even before he mentions the spreadsheet, we've already lost all hope in him. The answer to "How can you possibly violate conservation of momentum?" shouldn't be a big diffuse answer about complicated wheels and gears. Since the great difficulty is faced by every individual step of the system, tell me about any one step that defeats the difficulty, to show off your key idea in the simplest case.
If there is a countable sum of terms that I think is zero, because I think all the terms are individually zero, and you want to prove to me as quickly as possible that the sum is not zero, then defeat my initial skepticism by showing me one term that's nonzero. If instead you start out by talking about the amazing clever way you've ordered the summation so that its sum is easily provable, I lose hope right there.
When you ask Mr. L, "Why doesn't Conservation of Momentum rule out your drive?" your shred of hope is that Mr. L understands what you are asking -- has some perspective-taking on why an expert might usually think that was pretty difficult; and Mr. L stands ready to directly challenge your skepticism by addressing the key question right away: What aspect of this whole proposal has invalidated the usual reasoning saying that you can't build a reactionless drive out of reactionful parts?
Even being very charitable to Mr. L already, we don't have enough hope to bother following along if he starts in on a complicated story involving lots of gears and wheels.
The difficulty-refuting argument may be somewhat more complicated than the original difficulty-asserting argument. But before we start getting told about complicated details of Mr. L's proposal, we want to hear about some key idea that fans our tiny shred of hope that the great central difficulty has been overcome.
And if Mr. L starts to tell us about complicated wheels and gears instead, we see that he hasn't taken the perspective of a skeptical expert. So we infer that he probably doesn't understand the central difficulty. So we lose that tiny shred of hope, and we are not really feeling curious about all his complicated further details.
I will at this point briefly mention from 2026 -- though most of this essay was written 5 years earlier in 2021 -- a public discussion I had recently at the ILIAD conference with Geoffrey Irving (formerly Google Brain, OpenAI, Deepmind, now heading the new Resolution lab), in a conversation advertised as "Yudkowsky and Irving try and fail to settle all of their disagreements in one hour." One long sub-discussion focused on whether "debate" was a promising approach to superalignment. If you just have the humans judge whether an AI proposal to align a superintelligence will work, the humans will get it wrong; but what if the AIs debate each other instead?
And one thing that might go wrong is that the AIs form a swarm and sacrifice themselves for each other. But the other thing that goes wrong is that humans can also misjudge competing arguments in debates, not just isolated propositions.[3]
A point on which I kept pressing Irving was, "Supposing the human verifiers pass some bad argument steps in a statistically correlated way, how are you going to get reliable argument out of the whole system? Doesn't any approach to getting reliable answers out of a system with unreliable components presume at some point that you can break a debate down into steps where the judges are just statistically high-variance rather than statistically biased in their judgments? And isn't this just false in practice with humans?" And then refining this question further, I asked: "Suppose your judges were all random coinflips. That wouldn't work to drive correct outputs for debate as a means of superalignment. Can you tell me what property the judges need, which random coinflips lack, and which isn't 'the judgments are statistically unbiased', in order for this scheme to work?" And Irving said it was a good question and he might need to get back to me on that.
Later, several people at the conference said that they had been surprised, and had updated, from seeing that exact point of the discussion in particular.
I don't know whether this will be at all comprehensible to any reader and especially when the video hasn't been put online yet... but the question I was asking Irving was, from my own perspective, almost exactly analogous to trying to track down the part of a Perpetuum Mobile which breaks the rules.
You can build an accurate judger out of judges that are sometimes wrong as individuals, so long as the plurality or the majority is always right. You can build an unbiased estimator out of high-variance judges with no bias. But suppose the majority or plurality is sometimes wrong, and there isn't always some clever way of asking the question three different ways such that then two out of three variants are judged correctly. Then how are you building a "debate" system that reliably returns the correct answer out of unreliable judges? Why can't it also extract the length of the Emperor of China's nose from judges who've never seen the Emperor?
I expect you would notice a great many places where the would-be aligners present at that conference disagreed about other matters, depending on whether they perceived that debate as me pressing Irving about a strange abstract technical corner question; versus the question reformulating a seemingly big complicated task to boil down a central difficulty.
Many many engineers, in fact, historically rolled up their sleeves and got to work on building their designs for Perpetuum Mobiles. It was a larger-looming social phenomenon in Feynman's day -- one of the ways slightly smart engineers went Wrong back before AI or cryptocurrency. And I am not a postcognitive telepath -- I cannot read minds in the past -- but I wouldn't be surprised if many of those 1970s engineers saw themselves as hardheaded practical people with industry experience, whose effortful work learning about real metal gears had taught them the practical limitations of airy abstract theories like "conservation of energy". Which is to say, that they had tried their own hand at abstraction and not gotten anywhere, and learned from this an Arrogance of the Humbled[4] about the ultimate limits of mere thinking.
iii. Cat-Belling Problems.
When I was very young I read a children's story and extracted a lesson from it that stayed with me my whole life, being generalized and deepened as I grew older. I think it may have been a more elaborate story that I read as a child, but it came from an Aesop's Fable.
Here is the original Aesop's Fable in its entirety:
Aesop's original moral then runs:
IMO, there's a whole tree of valuable morals one could derive from this story. Aesop's first moral-branch has many valuable further sub-branches, eg in politics: "It is one thing to say you want X, and quite a different matter to devise a stable realistic system of incentives which yields X as its equilibrium."
But the greatest moral I took, that stayed with me my whole life to grow into a lesson-tree and deepen in its roots, was:
There is often some especially difficult or impossible step along the way to what you want, and until you overcome that key step, the rest of the plan is of no use.
We can see some further branches on this lesson-tree, exhibited in the following images, both of which show a missing cat-belling step calling a whole scheme into question:
(By Sidney Harris, published in American Scientist, Nov-Dec 1977.)
Or this even more famous one:
A very standard failure mode I have observed over the years is for people's minds to bounce off of, route around, or entirely fail to perceive, the cat-belling step.
Eg:
Or also eg (added 2026), many variants on:
This from my perspective is very much the equivalent of putting the "heat exchanger" in the second half of the machine; except that, instead of the Perpetuum Mobile entirely failing to work, it very unfortunately works to launch the airplane but fails to land it safely.
(That I specify no "publicly scrutinized" proposal exists is because, knowing the mental shenanigans around cat-belling steps, I predict Anthropic believes they have terribly serious proposals for accomplishing miracles of superalignment, and have terribly serious reasons why they cannot tell anyone, especially me, what those proposals are. No doubt many would-be builders of Perpetua Mobilia understood their real situation, on some intuitive level, well enough to come up with Various Reasons why Feynman could not be allowed to look at their secret designs.)
iv. The Optimizer's Curse against complicated plans for hard problems.
You will recall Mr. L's L-Drive and his animated GIF (reproduced from memory):
You will recall that when the flaw in his L-Drive setup was pointed out to Mr. L. -- that the acceleration at top and deceleration at bottom exactly cancels out the centrifugal force -- Mr. L. explained that he had designed "compensators" to prevent the cancellation from being exact, which was the real proprietary secret of his design.
Mr. L. did not post his spreadsheet, but he posted the graph of his spreadsheet's calculation of net force, which was mostly zero, but showed several positive spikes. Again reproducing it from memory:
I am not a past-viewing telepath. But on my grasp of psychology, there is a very obvious way to suspect events played out: Mr. L had his first brilliant idea about rotating balls in a cylinder, began to anticipate fame and fortune, but then saw the flaw himself... and then desperately set out to repair the flaw, instead of writing off his hopes and dreams... and invented more and more complicated "compensators"... until he stopped on a design where he'd managed to make a few bookkeeping errors in his Excel spreadsheet.
That is a very classic way that people end up believing they have solved some very hard step of a problem, in my experience -- by elaborating and complicating matters to the point where they themselves can no longer deduce the final failure. It doesn't work for mundane immediate matters like fixing your refrigerator today; but it works with anything that you can convince yourself is about the future, or that you can convince yourself has not been decisively refuted. Why, sure, your weird scheme will totally work for aligning ASI -- which, amazingly enough, is a key step that you cannot test right now; and that you can shut your ears and hum about, whenever anybody else tries to explain why your plan will fail.
And if these inventors are so kindly as to explain their scheme at all, they will launch into a description of some complex system, all of which I simply must hear about; and they don't, and can't, preface themselves with a simpler explanation that directly targets and defeats my previous skepticism about a Cat-Belling Problem.
That's what happens if you try to ask somebody underneath Geoffrey Irving's level to explain eg how they mean to use "debate" to overcome the problem of correlated bad judgments. They are downright confused by your apparent belief that they ought to have anything simple to say, about a key insight, that defeats some weird problem.
Standard lens time: This is yet another instance of Goodhart's Curse, or rather the Optimizer's Curse. If you try out a large number of complicated variations on a design, estimate their goodness via a process with nonzero error, and select the best-looking candidate, you will are more likely to find a spreadsheet with an upward error in the goodness estimate.
The more complicated the candidates, the greater that systematic error.[5] Which isn't necessarily fatal if you have some very low-error process that can evaluate very complicated things, which is why not all nuclear reactors with complicated designs have melted down. But if you are doing something more new and uncertain and error-prone -- then being able to simplify complicated schemes down to core-difficulty-defeating key ideas, is much more important as a check on the mental process.
v. When no Authority (that you accept) can tell you that your bright idea is wrong.
In the case of Conservation of Momentum, we have Authorized Authoritative Authorities to tell us that its cat-belling problem exists and is hard, and that Mr. L should not be able to defeat it just by throwing in a few obscuring complications.
What if instead, we are confronting some difficulty that is, in reality, extremely difficult -- maybe not impossible, like a reactionless drive, but still has some very difficult step of actually belling the cat -- but there is nobody you see as an Authorized Authoritative Authority, to tell you that it is hard?
Well in that case, things appear easier! They are not actually easier, of course; they are even harder, because the laws governing difficulty are not as solidly known. But they feel easier, because you will much more easily be able to solve the social problem of convincing people you are on the way to building a reactionless drive. And if you are most humans, you will take that as social proof that what you believed is okay to believe.
To pick up Aesop's fable where Aesop left off,[6] my mind generates its most natural continuation fic as follows:
vi. The equal and opposite advice.
For every advice there is an equal and opposite advice, that needs to be given to equal and opposite people; for every essay there is a long list of warning labels that I usually realize afterwards I failed to guess properly in advance. Nonetheless, here is my guess at a warning label that should be here:
It is possible to unjustly demand that an argument be too simple.
A poster child here could be, perhaps, the Scopes Monkey Trial. We can imagine the prosecution repeatedly demanding at what point a monkey gives birth to a non-monkey, and if anybody tries to say anything about a sum of many small discrete changes, the prosecution says: "Oh, this sounds like one of those big complicated ideas where nobody can see anymore how it is wrong; please boil it down to an essence, show me one point where a monkey gives birth to a non-monkey."
There might be some hope that a genuine cat-belling insight can be seen to squarely address a cat-belling problem, but that's only if you're genuinely trying to understand it rather than assuming an arms-crossed skeptical attitude saying "Convince me! No, not that way, the particular exact way I expect to be convinced!"
The opening example I gave with Mr. L's reactionless drive is one where we happen to know very exactly that a reactionless drive must have a single reactionless step that does something very anomalous under known physics; and furthermore, where we are pretty sure Mr. L is in fact mistaken. In other cases, cat-belling steps may be merely difficult and not impossible; they may even have some answer that cracks them wide open and makes them vanish as difficulties.
The people on the mailing list asking Mr. L how his drive worked were -- from my own perspective -- probably being too open-minded and too ready to listen. But being actually ready to listen to an incisive answer is the only way you can hear a solution to something you imagine to be a cat-belling difficulty, conditional on somebody actually having a solution. Richard Feynman was not being foolish in acting out something like a deontological rule about checking every time.
vii. The rest of this post, which I gave up writing.
The rest of this draft in 2021 was my attempt to list some Cat-Belling Problems in AI, along with explaining what gave them the status of a big problem rather than a peripheral technical difficulty.
After some difficulties in approaching that part from several attempted writing angles, I gave up, and just posted my raw list of items without any such preamble or expansion.
My model of some readers has them now reasoning, "How dare Mr. L name the drive after himself; how arrogant, how low-status; I therefore already associate a bad vibe with this drive, so it probably doesn't work." Point one, this was around 1997 and people were less status-regulatory of inventors back then; it was more common practice for people to name things after themselves. Second, on my view, if the drive worked, Mr. L would absolutely have been justified in naming it after himself; so the crux is whether or not the drive works. Or from another angle: to reason "the drive probably doesn't work, because Mr. L named it after himself" is bad Bayesianism. In the conceivable worlds where the drive does work, we are not unlikely to see Mr. L naming it after himself, especially if it is 1997 and we are less far from an older world where many inventors did that.
If you are the sort to confidently declare a flaw and you are wrong, you rate no higher in my judgment than Mr. L and perhaps lower. Because then on occasions when humanity does solve hard problems, you are the sort to contribute heat rather than light. But if, reading this, you only lightly guess and then you are wrong, that is much better than not daring to guess at all.
Or to state it more exactly: The reason we have the Scientific Method rather than the Scholastic Method is not that no scholar ever gets something right without overwhelming experimental proof; but that communities of scholars cannot correctly judge who got it right, even if one person got it right, without some experiment that absolutely hits them over the head with the correct answer; after which sometimes the majority notices.
Eg recent example, because people ignore historical examples, because their brain thinks that history is all a TV show and that historical people are more stupid than themselves and don't know about clever ideas like prediction markets: It's not that nobody warned EAs that AI might arrive sooner than 2050. There was an essay that laid out in advance all the reasons why the official-looking 2050 estimation was wrong, in hindsight ~100% correctly. But it was futile to try to hold a "debate" that would settle on this correct answer given for correct reasons, because even having been presented with the correct answer, the larger community could not majority-discriminate it as correct, until they were hit over the head with overwhelming contrary evidence. OpenPhil ran a contest with $50,000 in prizes for essays that disputed OpenPhil's then-standard estimate of 2050 median time to AGI and 5% catastrophic risk; all the prizes were awarded to essays that argued for longer timelines and lower risks.
If there were any method that broke debates into careful pieces that a group of judges would always judge in a statistically unbiased way, 2022 would've been a great time to use it in the essay-judging contest! But no method like that has ever been invented yet through all the ages of the world.
The historically observed problem with the Scholastic Method is not that no individual is ever smart enough to get things right without overwhelming evidence -- those smarter individuals are where the hypotheses come from that get tested by later experiment! The problem is that the majority is incapable of discerning which existing arguments have used more valid reasoning steps and so arrived at the correct answer from earlier data. It is exactly a failure of human groups at judging debates rather than a limit on the individual intelligence of humans to come up with the right answers earlier. Humanity throws occasional Einsteins who will correctly say, far in advance of overwhelming experimental evidence to convince the larger group, "Then I would have pitied the good Lord; the theory is correct." But when humanity at large looks to some very sober and solemn-looking people wearing the serious-person garments of that era, the soberly-clothed people are not able to tell which scholar's argument is correct. That is the problem with the Scholastic Method. The Scientific Method, for a small fragile time and in a few institutional places, was able to listen to the voice of overwhelming experimental evidence instead, when it came to deciding afterward which scholar to credit for having called it right. (Though to be clear, many particular scientific fields, or solemn-looking people wearing solemn garments, have not reached that standard either.)
So yes, it is very much fighting words to claim that you are going to do something called "debate" and then your human judges are always gonna correctly figure out which AI had the better argument on every argument step. Having a "debate" system which causes the Scholastic Method to start working with human judges is a Cat-Belling Problem.
A term and thesis coined by Duncan Sabien. I ought to write it up at some point, but meanwhile perhaps many readers, like myself, will find a whole useful thesis immediately apparent just from seeing the phrase "Arrogance of the Humbled".
This can of course itself be overused as an invincible argument against any slightly complicated scheme, including the ones that have multiple tiers of simplifiability in their key ideas. Many effective altruists that wanted peace of mind in knowing that they were doing the One Best Thing by buying mosquito bednets were endlessly endlessly convinced that any more complicated schemes for improving humanity, like "doing something about ASI before we all get killed", must surely contain an invalidating error.
If any reader be aghast at my audacity in daring to pick up where Aesop left off, as if I were comparing my own writing abilities to fables that have lasted through millennia, I remind them that I am publicly validated as one of the leading authors of fanfiction on the entire planet and therefore it is not arrogant of me to try to write a continuation fic of an Aesop's fable.