# Comments: Cat-Belling Problems Post URL (Markdown): [/api/post/cat-belling-problems](/api/post/cat-belling-problems) Post URL (HTML): [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems) Showing 98 of 95 comments (sort=top). For reaction user names, use `?includeReactionUsers=1`. ### Comment by [Steven Byrnes](/users/steve2152) * 2026-09-04 03:23:24Z * Karma: 56 * Voting system: namesAttachedReactions * Approval votes: 27 * Total votes: 27 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/P35gDozbTyZ8BMhLi](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/P35gDozbTyZ8BMhLi) * Markdown permalink: [/api/post/cat-belling-problems/comments/P35gDozbTyZ8BMhLi](/api/post/cat-belling-problems/comments/P35gDozbTyZ8BMhLi) Reactions by quoted text: * "for example, if Eliezer considers his own hidden-stat good enough, then we can note that Eliezer has in fact sometimes learned things from reading debates among humans in the liter..." * important: 1 I too am very skeptical of debate overall, but I think this particular objection can be answered. If the debate judge is a population-median person, then we’re screwed. But that’s not a specific problem with debate. Rather, it’s a universal, unavoidable aspect of the challenge of surviving ASI. For example, if humans do alignment research instead of AIs, then the humans will still argue with each other. (Universal consensus on ASI alignment plans is a pipe dream.) If we ask a population-median person to adjudicate those arguments in order to figure out how to approach alignment, then, again, we’re screwed. Thus, this is a universal problem. So it’s misleading to treat it as an argument against debate specifically. (See my cynical discussion of “law of conservation of wisdom” [here](/api/post/pZhEQieM9otKXhxmd?commentId=Eb5HgqyqT5uAye484).) Instead, we should break the problem of surviving ASI into a conjunction of a number of (IMO very unlikely) hopes. One of the hopes is that we have a technical plan that can work in the hands of people with exceptional conscientiousness, discernment, understanding, security-mindset, etc. A separate hope is that such people do in fact wind up being the decisionmakers, and following that technical plan. Both of these are real problems, and indeed there’s no solution to the first one that’s so good that it removes the second problem. So that’s why, when we’re talking about debate, we should *not* assume a population-median judge, but rather assume a judge with unusual discernment, patience, conscientiousness etc. Not because such a judge is guaranteed, or even likely, but because winding up with such a judge is a *separate* part of the challenge we face. And we don’t want to double-count. Now, the question from the OP was: “what property \[do\] the judges need, which random coinflips lack, and which isn't 'the judgments are statistically unbiased', in order for this scheme to work?” Alas, we don’t have a *great* answer to this question, because nobody has a way to quantify “the quality of being able to tell a good alignment plan from a bad one” among humans. Presumably this hidden stat has some relation to patience, security-mindset, intelligence, etc., but the details are unknown. A hope for debate might be that, if a person has a good-enough hidden stat to construct an adequate alignment plan in the traditional way (reading, thinking, corresponding with other humans, etc.) *given enough time*, then that same person will also be good enough to judge a debate between AIs about alignment plans. However, the latter would enable them to make much faster progress. And maybe also, the latter would be more forgiving about what constitutes a good-enough hidden stat for the scheme to work. Not so forgiving that the population-median hidden-stat is OK! That’s unrealistic. But it might be a bit more forgiving on the margin. So that’s a hope. And it’s not a crazy hope—for example, if Eliezer considers his own hidden-stat good enough, then we can note that Eliezer has in fact sometimes learned things from reading debates among humans in the literature. And maybe he could have figured those things out himself eventually, but it would have taken him longer. Again, I’m not defending debate as a good plan. I don’t think the debate story hangs together. But my cruxiest complaints are pretty different (mainly something like: the AIs will be either so incompetent that none of them have good ideas about ASI alignment, or sufficiently competent that they will subvert the setup). I think there are also *possible* problems more similar to what OP is talking about, I’m just pushing back on the notion that this is a guaranteed blocker. ### Comment by [cousin_it](/users/cousin_it) * 2026-09-03 23:17:12Z * Karma: 54 * Voting system: namesAttachedReactions * Approval votes: 21 * Total votes: 21 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/AfThXr7gLvFE3nuqX](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/AfThXr7gLvFE3nuqX) * Markdown permalink: [/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX](/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX) Reactions (whole comment): * thinking: 1 It seems to me if you can name an invariant, like conservation of momentum (or the invariant that there's no bell on the cat), then you're entitled to ask which step violates the invariant. Otherwise no. ### Comment by [Eli Tyre](/users/elityre) * 2026-09-05 07:26:52Z * Karma: 47 * Voting system: namesAttachedReactions * Approval votes: 23 * Total votes: 23 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/KepDpgXbCtYbmYBhs](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/KepDpgXbCtYbmYBhs) * Markdown permalink: [/api/post/cat-belling-problems/comments/KepDpgXbCtYbmYBhs](/api/post/cat-belling-problems/comments/KepDpgXbCtYbmYBhs) Reactions by quoted text: * "throw their mind away" * typo: 1 > This can of course itself be overused as an invincible argument against any slightly complicated scheme, *including* the ones that have multiple tiers of simplifiability in their key ideas. Many effective altruists that wanted peace of mind in knowing that they were doing the One Best Thing by buying mosquito bednets were endlessly endlessly convinced that any more complicated schemes for improving humanity, like "doing something about ASI before we all get killed", must surely contain an invalidating error. I feel annoyed at (what I perceive as) the snarky tone of this footnote. The main thrust of this essay is that 1. People have a particular kind of cognitive bug where they hide the hard part of them problem from themselves, without noticing or realizing that they've hidden it. 2. And those people thereby generate doomed plans. 3. And those people generating those optimistic doomed plans, are sometimes making the overall situation worse, rather than better, by offering false hope, in the mild case, and actively going out and taking enormously destructive power-seeking actions, reassured by their plans, in the less mild case. Given that, and given a whole slew of other ways that people fail hard at even slightly hard epistemological questions, I feel **quite good** about some early EAs thinking to themselves ~ > I don't know what to make of all that AI risk business. It seems kind of crazy to me. But saving lives in the third world by buying bed nets, seems concrete and solid and verifiable, in a way that those abstract arguments about future technologies that will maybe be invented one day totally do not. I'm going to try to save lives in the third world buy buying bed nets. Those people probably *should* go buy bed nets! Buying bednets is an extremely noble thing to do. Most people don't try to save *any* lives. The above reasoning *does* successfully avoid the pitfalls and harms that you're pointing to in this essay. Furthermore, if some EA is reasoning along these lines, it doesn't seem all that likely that they're *actually* going to make the situation better, if they for some reason pivoted to x-risk. In actual history, I don't think it was good for them or for the world when the EAs that felt an internal unreality about the AI risk arguments, ended up in a social context where they were pressured to "believe in AI risk" because it's "the most important thing." EAs should focus on the domains where they feel like they can trust their own reasoning. If they feel that "to be a good EA" they have to operate in a domain where they are *forced* to defer to the social hierarchy, because they can't trust their own thinking, that makes everything worse. I would prefer it if such EAs could "spit out" the AI risk arguments without sneering to much at the AI x-risk people, if they can manage that. It would have been awesome if they had said, > I don't trust my reasoning in this domain, but that's a fact about me and my epistemology, not a fact about the world. Maybe there are other people who can be justifiably confident in this class of argument. But I'm not one of them, so I'm going to do the good that I can see to do. But that's a pretty high bar of self awareness, and (in many cases) it would have broken their social-epistemic grounding. If claiming that *no one* could know whether arguments like the arguments for AI risk were valid was the best way they could avoid [getting eulered](https://slatestarcodex.com/2014/08/10/getting-eulered/), then good for them, they picked the better option of the double-bind. [^ili14o42xk] I agree with you that there's a meme, which is central to EA as it actually developed, and which is toxic in that it is upstream of these cognitive distortions. I think you've named it well enough: "the One Best Thing". The EA hope and ideology is to not just do a lot of good, but to do the *most* good. And so if someone can credibly claim to you, or argue from authority to you, that some problem, which you struggle to think about clearly, is t**he most important problem**, you'll steer yourself into regions of the world where you can't actually reason or plan with your own mind, which means you're ripe for exploitation by people telling you what to do and what to think.[^l572q8s81nm] But, this footnote, in the OP, *contributes* to the problem more than it alleviates it. It's not saying "this notion of 'doing the *most* good' and 'the One Best Thing', despite their obvious appeal, are pernicious, because they leads to your contorting your thoughts in order alleviate anxiety about having picked the *wrong* altruistic option " Instead, this footnote seems to be subtly shaming EAs for failing to *correctly identify* the One Best Thing, who (marginally) doomed humanity because they contorted themselves to avoid accepting the AI risk arguments. That's not helpful! Yes, from one perspective, they did less than they could have, because there were in fact better opportunities than bednets[^lhouergubk8] available for those with the eyes to see. But if they *did not* have the eyes to see—and couldn't realistically have developed those eyes given their actual psychology and epistemology and would have made things worse if they had tried to force themselves to do that—then don't pick on *them*, of all people on earth, who set out to do the most good that they could and took actions to save the lives that they could see to save. I guess I had a rant in me about this. I don't know if this footnote really merits the rant. But, I want to make a point to push back against social pressures that were, and are, trying to make those people [throw their mind awa](/api/post/RryyWNmJNnLowbhfC)y in the name of "doing the *most* good". [^ili14o42xk]: And in any case, I also wouldn't want them to hold back from publicly saying "I think we don't have good enough reason to trust these AI risk arguments, and we should invest in bednets (which we have much stronger reason to think save lives), instead." I want everyone to loudly make the case for what arguments they think hold water, and why. Public arguments about that was sort of the whole point of having an EA movement. [^l572q8s81nm]: The advice that I give to EAs is "you should do the best and most ambitious thing that is straightforwardly good to do, by your own reasoning. What's the best thing you can see to do, where you can understand all the parts of why it's a good thing to do, such that the goodness feels obvious and mundane?" [^lhouergubk8]: Is that even true? I was trying to help with AI risk from 2015 on, and it's not clear that any of the things that I tried actually meaningfully helped, ex post. It sounds like that's largely true of much of what you tried, as well. Now, I don't regret the things I tried, even if they mostly didn't help, because I think they were good bets, given what I knew at the time.But I sure don't feel on solid ground criticizing Global Poverty EAs for saving some lives by buying bednets, when the strategies I chose instead of doing that didn't actually pay off very much. Even if they had magically bought all the AI risk arguments, could they actually have identified actions that would have reliably helped in expectation? If not, then they were maybe correct to buy bednets! ### Comment by [Carl Feynman](/users/carl-feynman) * 2026-09-09 20:52:34Z * Karma: 44 * Voting system: namesAttachedReactions * Approval votes: 14 * Total votes: 14 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/6ahC5iiJcPsd9zWHa](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/6ahC5iiJcPsd9zWHa) * Markdown permalink: [/api/post/cat-belling-problems/comments/6ahC5iiJcPsd9zWHa](/api/post/cat-belling-problems/comments/6ahC5iiJcPsd9zWHa) > Richard Feynman deemed some of those Perpetua Mobilia "worthy of much more serious physical investigation" I should make clear that phrase is me commenting on my father's behavior. He never said that they were worthy of investigation. But I did see him pay attention to such things. Until he figured out how they were mistaken or fake, of course. > ...Richard Feynman was saying ought to be looked into... He didn't say it! I did! ### Comment by [Eliezer Yudkowsky](/users/eliezer_yudkowsky) * 2026-09-06 04:14:53Z * Karma: 42 * Voting system: namesAttachedReactions * Approval votes: 25 * Total votes: 25 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/KepDpgXbCtYbmYBhs](/api/post/cat-belling-problems/comments/KepDpgXbCtYbmYBhs) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/A93SFD82HSs2nkRJW](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/A93SFD82HSs2nkRJW) * Markdown permalink: [/api/post/cat-belling-problems/comments/A93SFD82HSs2nkRJW](/api/post/cat-belling-problems/comments/A93SFD82HSs2nkRJW) Reactions (whole comment): * heart: 1 I surely do credit with wisdom those EAs who went off and quietly bought bednets and saved a few kids, rather than posting to the Internet their arguments for why MIRI couldn't possibly be a better bet than bednets. The difference between actual humility and the posturing that is Modesty. The ones endlessly going "But what if there's a *mistake in your argument*?" are the ones I'm rolling my eyes at, here. Well, golly fuck, what if there's a mistake in your fucking argument? Now the quiet modest ones who bought bednets and didn't fuck around in difficult-to-analyze grownup business are going to have their work undone when the children in Africa die anyway, because other people borrowed their name and used that reputational credit to, among other things, shit all over the attempts to have Earth do something other than die. ### Comment by [Vaniver](/users/vaniver) * 2026-09-04 07:27:11Z * Karma: 36 * Voting system: namesAttachedReactions * Approval votes: 19 * Total votes: 19 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/P35gDozbTyZ8BMhLi](/api/post/cat-belling-problems/comments/P35gDozbTyZ8BMhLi) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/BbgkjzmHgkTLZPCLs](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/BbgkjzmHgkTLZPCLs) * Markdown permalink: [/api/post/cat-belling-problems/comments/BbgkjzmHgkTLZPCLs](/api/post/cat-belling-problems/comments/BbgkjzmHgkTLZPCLs) > A hope for debate might be that, if a person has a good-enough hidden stat to construct an adequate alignment plan in the traditional way (reading, thinking, corresponding with other humans, etc.) *given enough time*, then that same person will also be good enough to judge a debate between AIs about alignment plans. For what it's worth, I think this surfaces a different difficulty. One of the intuitions powering debate is the idea that some things are easier to check than generate--imagine the problem of finding a counterexample to a conjecture, where you might have to search thru a huge set of examples to find one, but then the verifier only needs a few steps to confirm that it's a counterexample. Debate then represents a way to get exponential effort that is verified in linear time, allowing humans to oversee the work of much more capable machines. But some things are easier to generate than check. Joel Spolsky, back in 2000, wrote "[It's harder to read code than write it](https://www.joelonsoftware.com/2000/04/06/things-you-should-never-do-part-i/)." I expect it will be easier to be confident in a program that I write than one that Claude Code generates (both that it will be correct and that it doesn't contain adversarially chosen side effects). For many years, human translators preferred translating a passage from scratch to editing a machine-translated passage. What makes us suspect that the philosophy or science of alignment is fertile ground for debate? ### Comment by [CronoDAS](/users/cronodas) * 2026-09-04 15:20:16Z * Karma: 29 * Voting system: namesAttachedReactions * Approval votes: 14 * Total votes: 14 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/iDqe8koExEspgRkkj](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/iDqe8koExEspgRkkj) * Markdown permalink: [/api/post/cat-belling-problems/comments/iDqe8koExEspgRkkj](/api/post/cat-belling-problems/comments/iDqe8koExEspgRkkj) I always thought that the problem with "belling the Cat" is not that a mouse would inevitably fail to bell the cat, but rather that whichever mouse did it would presumably be killed by the Cat shortly afterwards. So even though belling the cat wouid be good for the other mice, no individual mouse wants to be the one that has to do it ### Comment by [RobertM](/users/t3t) * 2026-09-04 05:59:50Z * Karma: 26 * Voting system: namesAttachedReactions * Approval votes: 12 * Total votes: 12 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/MkDHAGWuaPsrPHhsX](/api/post/cat-belling-problems/comments/MkDHAGWuaPsrPHhsX) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/7LerZ3266AA4HG2zz](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/7LerZ3266AA4HG2zz) * Markdown permalink: [/api/post/cat-belling-problems/comments/7LerZ3266AA4HG2zz](/api/post/cat-belling-problems/comments/7LerZ3266AA4HG2zz) Reactions (whole comment): * agree: 3 > An AI smart enough to solve alignment is smart enough to fool me about whether it has solved alignment It would be crazy to count on this as a key factor of any reasonable plan, but that order of capabilities doesn't seem overdetermined. Obviously there are many other reasons that plan fails, like "they wouldn't pick you", "there wouldn't be enough time for you to judge the plans before the next lab just YOLOd into RSI", etc, but those start to look like contingent features of reality that are not literally impossible to change (though they may be sufficiently difficult that "global compute control + ban on frontier model training" is easier). ### Comment by [Saul Schleimer](/users/saul-schleimer) * 2026-09-04 03:51:35Z * Karma: 26 * Voting system: namesAttachedReactions * Approval votes: 20 * Total votes: 20 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/a9KjfLwEpAHdMawz6](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/a9KjfLwEpAHdMawz6) * Markdown permalink: [/api/post/cat-belling-problems/comments/a9KjfLwEpAHdMawz6](/api/post/cat-belling-problems/comments/a9KjfLwEpAHdMawz6) > Even if a Mouse were to reach that high, it would take time for a Mouse to loop the bell and collar around the Cat's neck, let us say five seconds, and this seems to inherently require the Mouse to be in close proximity to the Cat. Meanwhile the Cat kills any Mouse who approaches within half a second. Our joy at the belling of the Cat is marred only by our sorrow at the death of ten of the eleven Brave Companions. May we always remember their noble sacrifice! ### Comment by [philh](/users/philh) * 2026-09-04 10:41:55Z * Karma: 22 * Voting system: namesAttachedReactions * Approval votes: 11 * Total votes: 11 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/4aEfovEEEbLzzmxp5](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/4aEfovEEEbLzzmxp5) * Markdown permalink: [/api/post/cat-belling-problems/comments/4aEfovEEEbLzzmxp5](/api/post/cat-belling-problems/comments/4aEfovEEEbLzzmxp5) > A term and thesis coined by Duncan Sabien. I ought to write it up at some point, but meanwhile perhaps many readers, like myself, will find a whole useful thesis immediately apparent just from seeing the phrase "Arrogance of the Humbled". Duncan himself has written [No Arrogance Like That Of The Recently Humbled](https://homosabiens.substack.com/p/no-arrogance-like-that-of-the-recently). ### Comment by [Kaj_Sotala](/users/kaj_sotala) * 2026-09-04 11:46:52Z * Karma: 21 * Voting system: namesAttachedReactions * Approval votes: 17 * Total votes: 17 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/vy3Xnq7g5njpGQjKL](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/vy3Xnq7g5njpGQjKL) * Markdown permalink: [/api/post/cat-belling-problems/comments/vy3Xnq7g5njpGQjKL](/api/post/cat-belling-problems/comments/vy3Xnq7g5njpGQjKL) My brain reading this: "Oh, is that where [bellingcat](https://www.bellingcat.com/about/who-we-are) gets its name?" "But how is the fable connected to what they do? I don't get it" > [Press Gazette:](https://pressgazette.co.uk/news/bellingcat-expansion/) Looking at the reporting Bellingcat does, it would be easy to mistake them for the kind of journalists you see in spy thrillers; doggedly exposing the espionage and hidden affairs of evil regimes. It’s even where the organisation’s name comes from – Higgins took it from an old fable, “belling the cat”, where a group of mice plan to stick bells on their hated neighbourhood cat, so they would know when it was coming to eat them. The simple lesson for Higgins; the powerful and dangerous shouldn’t be allowed to move in silence. > [American University book review:](https://www.american.edu/sis/centers/security-technology/book-review-we-are-bellingcat.cfm) The name “Bellingcat” came from the fable of mice deciding to put a bell on the cat that is after them. Their problem, of course, was figuring out who would take the risk. Higgins and his team of online sleuths believe it is their responsibility to put the bell on the cat of governments, companies, and any other nefarious organization that wants to hide information from the public. "Oh hmm okay I guess" (thank you brain for focusing on the most essential things) ### Comment by [Karl Krueger](/users/karl-krueger) * 2026-09-04 02:35:27Z * Karma: 21 * Voting system: namesAttachedReactions * Approval votes: 18 * Total votes: 18 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX](/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/LiYYCXwvNKjQNFAdM](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/LiYYCXwvNKjQNFAdM) * Markdown permalink: [/api/post/cat-belling-problems/comments/LiYYCXwvNKjQNFAdM](/api/post/cat-belling-problems/comments/LiYYCXwvNKjQNFAdM) Reactions (whole comment): * miss: 1 > the invariant that there's no bell on the cat In the fable, a stated invariant is that any mouse that comes in contact with the cat for 0.5sec gets eaten; thus there exists no mouse capable of belling the cat. There are approaches that *don't* violate this invariant: * Construct a trap that puts the bell on the cat without a mouse present. * Convince the cat's humans to bell the cat; for instance by showing them propaganda about how belling the cat [will save the singing dinosaurs](https://www.sciencedirect.com/science/article/abs/pii/S0168159105000742?utm_source=chatgpt.com). There are also approaches that *do*, such as the old mouse's idea of sneaking up on a sleeping cat; or the *Rats of NIMH* method of putting drugs in the cat's food — because the stated invariant *is not* a conservation law, it's just a summary of costly observations. ### Comment by [XelaP](/users/xelap) * 2026-09-05 11:25:46Z * Karma: 20 * Voting system: namesAttachedReactions * Approval votes: 11 * Total votes: 11 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/sKbaCQAj6tFkucqYo](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/sKbaCQAj6tFkucqYo) * Markdown permalink: [/api/post/cat-belling-problems/comments/sKbaCQAj6tFkucqYo](/api/post/cat-belling-problems/comments/sKbaCQAj6tFkucqYo) Okay, I know how this sounds, but there seems to be strong evidence for a reputable way to get reactionless thrust in general relativity, though it scales with curvature and so won't do anything far away from any mass-energy. It is perhaps an interesting example of what sort of things actually succeed at belling cats. Here's [the paper](https://www.science.org/doi/10.1126/science.1081406) by Jack Wisdom[^i5t8pw4ra1] and here's [a more comprehensible one](https://iopscience.iop.org/article/10.1088/1367-2630/8/5/068) by Avron and Kenneth. Here's [a paper](https://www.jstor.org/stable/26001496?seq=1) and [associated video](https://www.youtube.com/watch?v=BZF1nelgPNs) on making a robot on a curved sphere that exhibits the effect for what should be the same mathematical reasons. Here's [a pop-sci level article](https://www.jstor.org/stable/26001496?seq=1) in Scientific American (that links to a JSTOR version with images, here's a [link to the version on their website](https://www.scientificamerican.com/article/surprises-from-general-relativity/)) covering the Wisdom article, that also [cites a paper](https://journals.aps.org/prd/abstract/10.1103/PhysRevD.75.081501) (co-authored by the author) building off of Wisdom's that lets you slow the fall of a falling particle. Here's [an arxiv paper](https://arxiv.org/abs/2211.04654) from last year extended the theory too. There is, however, [this paper](https://arxiv.org/abs/1611.06183) on arxiv that finds a flaw in Wisdom's analysis, but it also *fixes* the flaw and still gets swimmers (but with different conditions and with different magnitudes). There's also [this paper](https://arxiv.org/abs/1707.08870) that builds off of their paper, investigating the approximate formalism used there, and finds that agreement depends on a prescription for "center of mass" used in the formalism - but they themselves say that the formalism may have problems by itself, and may not apply in this case, or may be being misunderstood. Here's [a physics stack exchange question](https://physics.stackexchange.com/questions/886/swimming-in-spacetime-apparent-conserved-quantity-violation) on it (which is a simple confusion, namely that curved spacetimes breaks the Lorentz boost symmetry), and a [physics forums discussion](https://www.physicsforums.com/threads/does-general-relativity-allow-for-swimming-in-space-time.326206/).[^py1nbob961p] Okay, now for how the Avron and Kenneth article (which deals with non-relativistic swimmers in curved spacetime, and so is a special limiting case of the Wisdom paper) passes some sniff tests (which will also explain the parts that I understand): 1: False proposals often don't mention the obstacle of conservation of momentum, or explain why they either don't break it or why the conservation doesn't hold. This one starts by giving the actual conservation law, which seems to check out by my understand of general relativity (it's also relatively simple, both in the human-understandable sense (I understood it) and the mathematical sense, and it also tickles my mathematical/physical elegance heuristic): A continuous symmetry of the metric is associated with something called a [Killing vector field](https://en.wikipedia.org/wiki/Frank_Wilczek) (if you follow the direction of one, the flows will be the symmetry of the metric). Now let's assume a nonrelativistic swimmer in a curved spacetime, and a Lagrangian which is symmetric under one of those symmetries of the metric associated with a Killing field. Then the conserved quantity is given by the momenta projected onto the Killing field, summed up over all bodies. Now assume the potential energy depends only on the distance between two particles. Then we get a version of Newton's third law which states that there will be equal and opposite forces *along a geodesic.* ![image.png](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1788599513/lexical_client_uploads/ozjildkdyjcjclpildl6.png) Their caption: > Figure 2. Newton’s third law in curved space: the mutual forces are directed along a geodesic and are balanced This is presumably because if you fix particle A, then the tangents to the geodesics from A are the direction that the distance to B increases in, and that's what gives the force (mathematically, for each coordinate we take the partial derivative with respect to that coordinate - so here, with respect to the position of the nth body). They also state that there is no good notion of a center of mass, as will be explained in the next point. 2: They give simple special-cases where you can do reactionless thrust in a way that I can understand, even though their construction used in the rest of the paper doesn't work in those cases. Here are the diagrams with their captions (note that the circle is flat): ![image.png](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1788600224/lexical_client_uploads/lkxccdkheokubwqkavmo.png) > Figure A.1. The full red circle on the left, splits into the two blue half balls which recombine at the empty red circle on the right. This allows for moving without violating Newton’s laws. Notice how for the circle, you can't define a good version of "center of mass". When we split to become the blue half balls, we'd like to say that the center of mass is in the center of the circle as we see it in the Euclidean plane, but such a point is not in the actual circle. We could pick the original red point, as that's halfway in between the blue ones - but then we see that there's another point halfway in between on the other side of the circle. Consider that by standard physics we [do not know](https://en.wikipedia.org/wiki/Shape_of_the_universe) the global topology of our universe, and it is possible it looks something like that circle (though exploiting it would seemingly not be feasible because of the expansion of the universe preventing you from going far enough). ![image.png](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1788600290/lexical_client_uploads/clpl4lesnspnkeqtxesd.png) > Figure A.2. The red ball can displace itself to the position of the blue ball by splitting into two parts that it sends to crossing points of geodesics and then recombine satisfying Newton’s law in the process. 3: They derive a condition required for their construction to work, and apply it to situations we already are familiar with. The condition is apparently that the exterior derivative of the (dual of the) Killing field needs to be nonzero. As an example, this fails for translations in flat space - however it works for rotations, which is why it is possible for [cats to land on their feet by rotating their two halves differently](https://en.wikipedia.org/wiki/Falling_cat_problem)[^xp9lmpt0ho][^73ge817facy] ![Cat_fall_150x300_6fps.gif](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1788599950/lexical_client_uploads/k6z2zqy4s5hmipyihfvf.gif) Maxwell was apparently known for throwing cats out of windows. Unfortunately, we have now left the realm of stuff I understand. It still seems simple (in the sense of comprehensible-to-people-who-aren't-me), and also simple (in a sense analogous to Kolmogorov-complexity) in a mathematical sense (exterior derivatives are things you apply to stuff all the time, and Killing fields are about symmetries, and both are used often in GR). 4: Argument from authority: the two papers building off of Wisdom's are in big journals (Science and APS), and the robot one (in PNAS) works the same way that the 2D example in the Scientific American article works and came 15 years later, so we have *a* validated empirical prediction. I also see the fact that the paper about a possible flaw still gets swimmers and only disagrees about when they occur and how much swimming is required, and the second one seems to mostly cast doubt on its analysis technique. I also asked Claude to search, and it couldn't find other reputable critiques. ![image.png](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1788600806/lexical_client_uploads/bfqmgbuhhz4ahgeldylf.png) From the Scientific American article ![image.png](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1788600895/lexical_client_uploads/lzteg98qljphi2lbjooy.png) From the robot article ![image.png](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1788600930/lexical_client_uploads/hnpsgoduisbqxcc3yvi0.png) The robot from the video associated to the robot article 5: The result isn't as practically impressive, and so is less likely to attract cranks. Since it is a small effect, and scales proportional to the Riemann curvature (because the Killing field does), you can't just use it to swim through space unless you're close enough to a massive body. It doesn't sound like it's at all currently feasible to build something that swims in the spacetime around Earth that would, like, move enough that you could see it with your eyes. [^i5t8pw4ra1]: He got his BS in 1976, so he was probably almost 50 or older when he wrote that paper in 2003. It is said that with age comes wisdom, but now we know that with Wisdom comes age. He also has a blog, which is just a webpage with the bolded text "MIT Professor Jack Wisdom has no blog on social media." [^py1nbob961p]: TIL the site's been remade by vibecode, as I infer from the new UI. [^xp9lmpt0ho]: Yes, the papers use this example, clearly because they (or maybe the cats?) got an oracular vision foreseeing this post [^73ge817facy]: And here's some catboys astronauts on the ISS doing the same, I think? ### Comment by [leogao](/users/leogao) * 2026-09-07 02:52:40Z * Karma: 17 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX](/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/jXH8vmu347Sq3HLRS](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/jXH8vmu347Sq3HLRS) * Markdown permalink: [/api/post/cat-belling-problems/comments/jXH8vmu347Sq3HLRS](/api/post/cat-belling-problems/comments/jXH8vmu347Sq3HLRS) related: what are the actual conservation laws relevant for arguing that alignment proposals need to solve some “hard part”? the closest i know is some kind of handwavy consequentialism outcome preimage argument - which to be fair is quite intuitively compelling, such that i’m down to give a lot of benefit of the doubt - but this is not nearly as well-defined as any of the thermodynamic laws, and therefore we should also not put nearly as much stock in its use as a conservation law until shown otherwise. this seems like a pretty important thing for someone to hammer out if nobody has already. ### Comment by [cousin_it](/users/cousin_it) * 2026-09-04 08:26:38Z * Karma: 17 * Voting system: namesAttachedReactions * Approval votes: 9 * Total votes: 9 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/TpajEogfxzRyaFkJr](/api/post/cat-belling-problems/comments/TpajEogfxzRyaFkJr) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/urKB8Dyd5uRE7ZTqc](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/urKB8Dyd5uRE7ZTqc) * Markdown permalink: [/api/post/cat-belling-problems/comments/urKB8Dyd5uRE7ZTqc](/api/post/cat-belling-problems/comments/urKB8Dyd5uRE7ZTqc) Reactions (whole comment): * thumbs-up: 1 I wasn't disagreeing, more like trying to figure out for myself what makes a problem "cat-belling" or not. Here's an example. Alice: "I have a proof of Löb's theorem, consisting of [14 steps](https://en.wikipedia.org/wiki/Löb's_theorem#Proof_of_Löb's_theorem)." Bob: "Which step deals with the essential difficulty?" Alice: "Not sure what you mean, all steps are elementary and together they work." Bob is wrong here, because he hasn't given an invariant. He could say that the difficulty of a math proof is inferential distance to the conclusion, but that's not an invariant: it decreases by 1 with every step of the proof. This notion of difficulty is "salami slicing", not "cat belling". But if Alice said "I have a proof that P≠NP" and Bob asked "at which step does your proof fail to [relativize](https://en.wikipedia.org/wiki/P_versus_NP_problem#Results_about_difficulty_of_proof)?" that would be a legitimate question, because Bob has given an invariant. Someone really needs to bell the cat in this case. Applying this to AI capabilities or alignment is left as an exercise to the reader, which I'm not smart enough to do myself :-) ### Comment by [Eliezer Yudkowsky](/users/eliezer_yudkowsky) * 2026-09-04 05:01:55Z * Karma: 16 * Voting system: namesAttachedReactions * Approval votes: 14 * Total votes: 14 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/a9KjfLwEpAHdMawz6](/api/post/cat-belling-problems/comments/a9KjfLwEpAHdMawz6) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/453Bi8PZ9e6SxtxWp](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/453Bi8PZ9e6SxtxWp) * Markdown permalink: [/api/post/cat-belling-problems/comments/453Bi8PZ9e6SxtxWp](/api/post/cat-belling-problems/comments/453Bi8PZ9e6SxtxWp) Looks like you got your key insight down to one sentence, rather than needing to start out with a long description of a complicated mechanism that would choose eleven mice for some unknown reason that you could only explain later! ### Comment by [Eliezer Yudkowsky](/users/eliezer_yudkowsky) * 2026-09-04 04:56:59Z * Karma: 16 * Voting system: namesAttachedReactions * Approval votes: 18 * Total votes: 18 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/P35gDozbTyZ8BMhLi](/api/post/cat-belling-problems/comments/P35gDozbTyZ8BMhLi) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/MkDHAGWuaPsrPHhsX](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/MkDHAGWuaPsrPHhsX) * Markdown permalink: [/api/post/cat-belling-problems/comments/MkDHAGWuaPsrPHhsX](/api/post/cat-belling-problems/comments/MkDHAGWuaPsrPHhsX) If the EAs couldn't do it on AI timelines, neither can whoever Dario picks to be a debate judge. An AI smart enough to solve alignment is smart enough to fool me about whether it has solved alignment, so I'm not particularly telling you to solve this problem by slotting me in as a judge. TBC, I expect to be much, much better than whoever the AI companies try to use; but ASI alignment is not graded on a curve. ### Comment by [Steven Byrnes](/users/steve2152) * 2026-09-04 02:24:22Z * Karma: 16 * Voting system: namesAttachedReactions * Approval votes: 10 * Total votes: 10 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/Nki7yaJaxxRK75tuT](/api/post/cat-belling-problems/comments/Nki7yaJaxxRK75tuT) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/qzx5smzenrb3NzerB](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/qzx5smzenrb3NzerB) * Markdown permalink: [/api/post/cat-belling-problems/comments/qzx5smzenrb3NzerB](/api/post/cat-belling-problems/comments/qzx5smzenrb3NzerB) cousin_it could have said “invariant or monovariant”. ### Comment by [Gunnar_Zarncke](/users/gunnar_zarncke) * 2026-09-04 22:38:35Z * Karma: 14 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/5qQ2Y3ao6gWE77WCr](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/5qQ2Y3ao6gWE77WCr) * Markdown permalink: [/api/post/cat-belling-problems/comments/5qQ2Y3ao6gWE77WCr](/api/post/cat-belling-problems/comments/5qQ2Y3ao6gWE77WCr) I think the young mouse could be more willing to learn from the old mouse, but the old mouse should also be more willing to support. I'm not a good writer, but maybe Sol brings this point across: +++ Sol The old Mouse turned away and wondered whether the Cat slept. “Wait,” said the young Mouse. “You said we do not know where.” “We do not.” “Glow-paint rubs off.” The old Mouse looked back. “Our paint might be enough to dust the narrow passages. If the Cat crosses one, its paws will carry the dust onward. We could follow the prints after dark without approaching it.” “That might tell us where it sleeps,” said the old Mouse. “It would not tell us whether we can reach it there.” “No. But then we would know which difficulty we actually face.” The mice fell silent. This was not the plan the young Mouse had presented, and it was not the solution the old Mouse had sought. The paint would not bell the Cat. But it could test one of the assumptions on which a possible belling strategy depended. “And the bell that rings from any angle?” asked the old Mouse. The young Mouse hesitated. “Probably irrelevant.” “Probably,” agreed the old Mouse. “Unless we use it first as a distant alarm near the sleeping place, to learn how lightly the Cat wakes.” They spent the night redesigning the proposal. Before choosing a Mouse, they would discover where the Cat slept. Before making a collar, they would measure how closely a Mouse could approach. Most of the paint might still be wasted. The Cat might not sleep. Its den might be unreachable. It might wake at the first scrape of a claw. But by morning the old Mouse had stopped turning away, and the young Mouse had stopped focusing on the collar. Together they carried a small dish of glowing powder toward the first passage—not to bell the Cat, but to learn whether there was any path by which it could be done. +++ ### Comment by [StanislavKrym](/users/stanislavkrym) * 2026-09-04 15:42:44Z * Karma: 14 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/phF5vy5dzCAJPPwRD](/api/post/cat-belling-problems/comments/phF5vy5dzCAJPPwRD) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/f64ZyBGNGQBMbvcvk](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/f64ZyBGNGQBMbvcvk) * Markdown permalink: [/api/post/cat-belling-problems/comments/f64ZyBGNGQBMbvcvk](/api/post/cat-belling-problems/comments/f64ZyBGNGQBMbvcvk) If EA didn't possess the hidden stat [and Kokotajlo presumably did](/api/post/ZpguaocJ4y7E3ccuw#uxFCNxdJxKDgwCEXY), then what exercises could one use to develop the hidden stat in oneself and apply the stat to future debates? Additinally, what if AI alignment is more about (re)constructing agents from loads and loads of mechinterp R&D, as [Agent-4 does in AI-2027](https://ai-2027.com/race#superintelligent-mechanistic-interpretability) or [Plan A does in AI-2040](https://ai-2040.com/footnotes?from=/?choices=plan-a-root#footnote-173), while each individual mechinterp result is based on open-sourced evidence from smaller nets? ### Comment by [skolemizer](/users/skolemizer) * 2026-09-06 23:08:09Z * Karma: 13 * Voting system: namesAttachedReactions * Approval votes: 8 * Total votes: 8 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/KepDpgXbCtYbmYBhs](/api/post/cat-belling-problems/comments/KepDpgXbCtYbmYBhs) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ZoFCpwpjocGtff6n7](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ZoFCpwpjocGtff6n7) * Markdown permalink: [/api/post/cat-belling-problems/comments/ZoFCpwpjocGtff6n7](/api/post/cat-belling-problems/comments/ZoFCpwpjocGtff6n7) Personally I decided many years ago to just split the difference and spend 50% of my charity budget on GiveWell and spend 50% on MIRI, and have stuck with that decision ever since; it seemed like the straightforward thing to do and still does. ### Comment by [johnswentworth](/users/johnswentworth) * 2026-09-03 23:52:43Z * Karma: 13 * Voting system: namesAttachedReactions * Approval votes: 10 * Total votes: 10 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX](/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/Nki7yaJaxxRK75tuT](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/Nki7yaJaxxRK75tuT) * Markdown permalink: [/api/post/cat-belling-problems/comments/Nki7yaJaxxRK75tuT](/api/post/cat-belling-problems/comments/Nki7yaJaxxRK75tuT) Seems too narrow. For instance, the second law of thermodynamics isn't a conservation law. Nonetheless, if somebody shows me a design for a type-2 perpetual motion machine, it makes sense to ask where precisely entropy goes down. ### Comment by [Bunthut](/users/bunthut) * 2026-09-03 23:07:34Z * Karma: 12 * Voting system: namesAttachedReactions * Approval votes: 9 * Total votes: 9 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/3vn3fu6LXxtMD2KzR](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/3vn3fu6LXxtMD2KzR) * Markdown permalink: [/api/post/cat-belling-problems/comments/3vn3fu6LXxtMD2KzR](/api/post/cat-belling-problems/comments/3vn3fu6LXxtMD2KzR) Reactions by quoted text: * "Mathematics mostly operates by the Scholastic Method. It has parts with easily demonstrable implications," * agree: 1 > And Irving said it was a good question and he might need to get back to me on that. I didn't see the discussion, but it seems odd to me that a debate-aligner would grant the premise here - isn't the whole idea of debate that informative responses *do* exist in very specific parts of the dialogue-tree? I don't think the empirical case is as damning as you make it sound: Mathematics mostly operates by the Scholastic Method. It has parts with easily demonstrable implications, but also other parts which are quite far from that, and it's mostly fine. And people do fail to recognise the right answer, justified with the right reasons, at times - but that is also not always *the best argument* (to them) for that right answer. You might have needed to present it differently, or attack the reasons people don't accept the arguments first, or...etc. We don't demand you do this stuff to get Science Genius Credits afterwards, but that also means you can find examples of people failing to appreciate the Perfect Science Genius, without that necessarily bearing on judging debates of Science-and-Pedagogy Geniuses. I still don't really expect this to work, but more so on priors. ### Comment by [leogao](/users/leogao) * 2026-09-07 02:46:54Z * Karma: 11 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/KMEJ3fMyedNGKjEYn](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/KMEJ3fMyedNGKjEYn) * Markdown permalink: [/api/post/cat-belling-problems/comments/KMEJ3fMyedNGKjEYn](/api/post/cat-belling-problems/comments/KMEJ3fMyedNGKjEYn) in general this is a real problem that happens a lot in alignment. i’m not here to defend bad alignment proposals. but also i think there is an opposite problem that happens too, certainly in other fields, and possibly in alignment. a lot of problems can actually just be solved adequately enough by dodging the “hard part” of the problem, or ignoring the asymptotics, etc. in these cases, worrying about solutions not solving the “hard part” is actually just not a good idea. see also relevant [xkcd](https://xkcd.com/974/). yes of course we only get one critical try, if we fuck up we all die, reliability depends much more on fundamentally sound solutions than mean-performance, etc etc and all this should push us to accept a better solution for alignment. but it’s important to note that this type of argument is not universally valid. ### Comment by [DanielW](/users/danielw) * 2026-09-04 18:20:22Z * Karma: 11 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/z4rdbvLxDHPZci6nq](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/z4rdbvLxDHPZci6nq) * Markdown permalink: [/api/post/cat-belling-problems/comments/z4rdbvLxDHPZci6nq](/api/post/cat-belling-problems/comments/z4rdbvLxDHPZci6nq) > How can we be *sure* that Mr. L didn't successfully design a clever system of compensators, and correctly validate that design using a sound spreadsheet? It depends how strictly one defines "sure" but strictly speaking, we cannot. **If the only evidence we have received is the claims Mr. L is making**, we can be effectively certain, from an unbroken string of data, that the odds that L made an error, is lying, failed to account for some outside source, etc is so much greater than the odds that something fundamentally impossible according to all existing data and theory has been achieved by his system as to render it null. That is $P(A|B) \approx 0, P(\neg A | B) ≈ 1$. Where A is "Mr. L made a perpetual motion machine" and B is "the claims Mr. L has provided." There is some evidence and validations that could move the needle towards having to take the possibilty seriously, but as is the uncertainty is negligible. (Edit: a good parallel, I think, is Hume's '[Of Miracles](https://davidhume.org/texts/e/10)'. is there some evidence that could convince me a miracle happened? Certainly, but it would have to be extraordinary evidence, testimony is always going to be more probably explained by mundane phenomenon). > That is a *very* classic way that people end up believing they have solved some very hard step of a problem, in my experience -- **by elaborating and complicating matters to the point where they themselves can no longer deduce the final failure** It seems to me, however, this is ***sometimes*** the correct approach. Sometimes it is the case that complex 'solutions' to problems a really just making things so complicated that where they are wrong is no longer obvious but sometimes it does seem correct that a complex patchwork of addressing various problems that arise can solve the problem. An obvious example is something like practical machine-powered heavier than air flight with the technical capabilities of the late 19th/turn of the 20th century. Critics of the attempts at heavier-than-air flight like Simon Newcomb, Lord Kelvin, Joseph LeConte ([who actually compared it to perpetual motion machines](https://en.wikisource.org/wiki/Popular_Science_Monthly/Volume_34/November_1888/The_Problem_of_a_Flying-Machine), though he later revised his views becoming accepting of the possibility but still incorrectly skeptical) and others had a not actually unreasonable sounding case[^u9ne24a68l] that seems rather similar to the worries of "the old Mouse" as you describe it, that is that there were very high technical barriers and even if those were overcome those, there were necessary requirements which all attempts fell short by a "not a narrow gap." Importantly, **there was no single argument or key insight that any proponents of heavier-than-air flight could point to by which these would be overcome**. To get the calculations necessary to rebut the basic assertion of skeptics was extremely difficult. The Wrights found the equations for lift that had been calculated by aeronautical scientists of the day contained errors (though even finding those didn't help them), show that fixed wings could be designed to provide the necessary dynamics, and propellers could be designed with reasonable enough efficiency to generate the needed force from a small, custom aluminum engine. This required complex, speculative reasoning to show that it could be possible and then subsequent refinement, tests and development to resolve the problems that arose in practice and show their speculated case basically correct. > The opening example I gave with Mr. L's reactionless drive is one where we happen to know very exactly that a reactionless drive must have a single reactionless step that does something very anomalous under known physics I don't see why we *know* this. It doesn't seem impossible to me that on occasion it may be true that "there is no single step that can overcome this seemingly impossible problem, but a series of cumulative steps may be able to overcome the challenges." Indeed, this seems to be the way that many problems in history have been overcome, not by single key insights but a long chains of advancement and gradual theoretic and practical advancements coming together. [^u9ne24a68l]: The argument (in simple form) usually fell along the lines of (paraphrasing the usual cases): Heavier than air flight is not totally impossible (it is readily observed in birds, for example), modern developments in aeronautics have even indeed given us an understanding of the mechanisms. However, in every observable case it requires incredibly specialized mechanisms that suffer from practical constraints. Birds developed over many thousands of generations which come at the cost of almost all other considerations, and then have been unable to develop mechanisms capable of carrying more than 50lbs, if we extend the principles involved it wouldn't be able to achieve more than, perhaps, 100 lbs. A machine with an engine (as possible with known mechanisms of the day) would have to weigh several hundred pounds to hope to generate the force necessary to lift a human, at such a weight it is impossible with known materials and engines to generate the power necessary to lift itself and survive the forces required to generate sufficient lift. Attempts at building heavier than air flying machines have all encountered these issues and been entirely unable to overcome them and generate practical self-propelled heavier than air flight. The best attempts have been able to achieve powered flight at low weights, but these have been largely impractical demonstrations and all attempts to scale them have been unmitigated values. In the future, some other mechanisms may be discovered, or new more efficient sources of energy, but current projects aimed at developing heavier than air flight with known technologies are impossible. ### Comment by [Eliezer Yudkowsky](/users/eliezer_yudkowsky) * 2026-09-04 09:15:36Z * Karma: 11 * Voting system: namesAttachedReactions * Approval votes: 12 * Total votes: 12 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/P35gDozbTyZ8BMhLi](/api/post/cat-belling-problems/comments/P35gDozbTyZ8BMhLi) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/phF5vy5dzCAJPPwRD](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/phF5vy5dzCAJPPwRD) * Markdown permalink: [/api/post/cat-belling-problems/comments/phF5vy5dzCAJPPwRD](/api/post/cat-belling-problems/comments/phF5vy5dzCAJPPwRD) Reactions by quoted text: * "So far as I can tell, EA given unlimited time to think, but no further evidence, would've not figured out that AGI wasn't scheduled for 2050." * shrug: 1 > A hope for debate might be that, if a person has a good-enough hidden stat to construct an adequate alignment plan in the traditional way (reading, thinking, corresponding with other humans, etc.) given enough time They don't. They were not converging toward correct answers; their updates from debate were not moving in a correct direction. So far as I can tell, EA given unlimited time to think, but no further evidence, would've not figured out that AGI wasn't scheduled for 2050. ### Comment by [gwern](/users/gwern) * 2026-09-12 02:48:18Z * Karma: 10 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/BgQ4Pmaq33CWLcqGe](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/BgQ4Pmaq33CWLcqGe) * Markdown permalink: [/api/post/cat-belling-problems/comments/BgQ4Pmaq33CWLcqGe](/api/post/cat-belling-problems/comments/BgQ4Pmaq33CWLcqGe) > Why can’t it also extract the length of the Emperor of China’s nose from judges who’ve never seen the Emperor? This is an E. T. Jaynes reference about systematic vs random error that will probably be lost on most readers; see https://gwern.net/replication#systemic-error-doesnt-go-away for quote/gloss. ### Comment by [Steven Byrnes](/users/steve2152) * 2026-09-05 16:39:41Z * Karma: 10 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/sKbaCQAj6tFkucqYo](/api/post/cat-belling-problems/comments/sKbaCQAj6tFkucqYo) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/NuRRE5iPwkGoov6Xf](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/NuRRE5iPwkGoov6Xf) * Markdown permalink: [/api/post/cat-belling-problems/comments/NuRRE5iPwkGoov6Xf](/api/post/cat-belling-problems/comments/NuRRE5iPwkGoov6Xf) You start with “the fundamental laws of physics are the same everywhere” and that implies (by Noether’s Theorem) LOCAL conservation of momentum, in the form of a [differential continuity equation](https://en.wikipedia.org/w/index.php?title=Continuity_equation&oldid=1361612586#Differential_form) $\partial_{\mu}T^{\mu}{}_{\nu} = 0$ (where *T* is the [stress-energy tensor](https://en.wikipedia.org/wiki/Stress–energy_tensor)). This is the part that’s rock-solid. However, to get “no reactionless drives” / “center-of-mass momentum is fixed”, you need to turn the LOCAL conservation law (describing what happens in the infinitesimal neighborhood of any given point) into a GLOBAL conservation law (describing a spatially-extended volume). That step involves integrating the local law over space, doing the integration-by-parts trick, and setting the boundary terms to zero. This local-to-global integration step definitely works in flat space—that’s the homework problem everyone does in college. By contrast, the local-to-global step might or might not work if spacetime is curvy or otherwise weird. As a silly toy example, if you assume wormholes / teleportation portals, then it’s easy to construct examples where local conservation of momentum fails to constrain center-of-mass acceleration. (We were chatting about it [here](/api/post/njBRhELvfMtjytYeH?commentId=2p3zxKBcrMp8aZ9SB).) All the other familiar global conservation laws (energy etc.) are also dubious in GR. For example, to even state the claim “total energy of the universe is the same now as tomorrow”, you need a consistent “now” for every point in the universe, including points around black holes etc. where spacetime is doing weird things. I don’t know all the details of what can and can’t be proven, I’m just saying it’s really not obvious that anything can be proven at all. ### Comment by [Martin Randall](/users/martin-randall) * 2026-09-05 13:36:01Z * Karma: 9 * Voting system: namesAttachedReactions * Approval votes: 12 * Total votes: 12 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/phF5vy5dzCAJPPwRD](/api/post/cat-belling-problems/comments/phF5vy5dzCAJPPwRD) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/Ps6SN2EgFCeuDn9ed](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/Ps6SN2EgFCeuDn9ed) * Markdown permalink: [/api/post/cat-belling-problems/comments/Ps6SN2EgFCeuDn9ed](/api/post/cat-belling-problems/comments/Ps6SN2EgFCeuDn9ed) Reactions (whole comment): * miss: 2 EAs did figure that out. 2050 was a median with wide error bars, not a schedule. Edit: since I got a disagree vote, I'll show my work. The [Ajeya's 2020 draft report on AI timelines](/api/post/KrJfoZzpSDpnrv9va) projected affordable training of transformative AGI (IE, conditional on no AI pause) between 2025 and 2100, mode is 2040, median 2050. This is not a "schedule" by any reasonable definition. ### Comment by [gwern](/users/gwern) * 2026-09-12 03:42:19Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/AQfGsT6gHYJciEhbb](/api/post/cat-belling-problems/comments/AQfGsT6gHYJciEhbb) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/CzdMwHeug6nYpRSjv](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/CzdMwHeug6nYpRSjv) * Markdown permalink: [/api/post/cat-belling-problems/comments/CzdMwHeug6nYpRSjv](/api/post/cat-belling-problems/comments/CzdMwHeug6nYpRSjv) > The way I’d phrase the Apollos’ hypothesis would be: “Deception is *fundamentally* more computationally expensive to execute that honesty, because a deceptive entity needs to keep at least two sets of mental books (the truth and the lie(s)) while an honest entity only needs to keep one.” If that hypothesis is true, and if we can confirm that it’s true with high confidence, then we may indeed be able to build deception detectors that are 100% reliable across a relevant set of scenarios. Was this ever discussed at length? It seems kinda dubious to me, because it's not hard to imagine externalizing deception or changing a system's ontology or amortizing/dispersing deception-related things, especially as 'deception' doesn't seem like a terribly well-defined concept to begin with. (For example, imagine a tabular Q-learner which gradually learns an optimal policy by memorizing the returns from every action/state and is just a lookup table, in effect. It is an introductory textbook agent, computationally cheap to execute, universal, and such an agent will converge to the optimal policy in the limit under standard assumptions. So it can learn the value of useful 'deceptive' actions without ever once doing anything remotely like 'mental books', never mind having to develop theory of mind or keeping multiple books. No deception detector could ever detect anything in this agent, even though it is about as simple as possible for an agent to be and easily implemented/approximated by other agents as a subroutine etc.) ### Comment by [Towards_Keeperhood](/users/towards_keeperhood) * 2026-09-04 18:10:40Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX](/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/Jz4REnivMX2Rdq72C](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/Jz4REnivMX2Rdq72C) * Markdown permalink: [/api/post/cat-belling-problems/comments/Jz4REnivMX2Rdq72C](/api/post/cat-belling-problems/comments/Jz4REnivMX2Rdq72C) How would you state the invariant in the case where someone tries to forecast lottery numbers? Invariant framing doesn't seem natural to me. ### Comment by [Vaniver](/users/vaniver) * 2026-09-04 07:37:36Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/aFtQARwv4YkxzcRvC](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/aFtQARwv4YkxzcRvC) * Markdown permalink: [/api/post/cat-belling-problems/comments/aFtQARwv4YkxzcRvC](/api/post/cat-belling-problems/comments/aFtQARwv4YkxzcRvC) > A poster child here could be, perhaps, the Scopes Monkey Trial. Actually, the example I was thinking of while reading this post was v-structures in causal discovery. When I first came across it, I was surprised that there was something that you can't do with two elements but can do with three. \[Of course, someone who understands that doesn't run into the cat-belling problem with explaining it. They do have the convenience that the base case is still simple enough to easily hold in mind, but even in cases where the enabler is obscure or subtle it seems like they should be able to acknowledge the difficulty.\] ### Comment by [roha](/users/roha) * 2026-09-04 02:39:23Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ocghj6KoSCKn9kMWb](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ocghj6KoSCKn9kMWb) * Markdown permalink: [/api/post/cat-belling-problems/comments/ocghj6KoSCKn9kMWb](/api/post/cat-belling-problems/comments/ocghj6KoSCKn9kMWb) "Rather than trying to hang a bell around the Cat's neck, which would just enrage it and lead it to resent us, we need to [imprint parental feelings upon it](https://time.com/collections/time100-ai/6309011/ilya-sutskever/)," said another mouse. An old mouse with slight hearing problems countered: "Paternal instincts? It's too early to [bloviate about the sex of cats](https://x.com/ylecun/status/1823313599252533594)!" "Yes," enthused another old mouse in a eureka moment, "we just need to give it [maternal instincts](https://www.forbes.com/sites/pialauritzen/2025/08/14/geoffrey-hinton-says-ai-needs-maternal-instincts-heres-what-it-takes/). Nature has done it before, so we can do it too!" "Nature? It's inevitable that the Cat will eat us. That's the natural trajectory. [We should not resist it, we should embrace it](https://officechai.com/ai/humans-should-welcome-being-succeeded-by-ai-as-a-part-of-evolution-turing-award-winner-richard-sutton/)," concluded yet another old mouse, smiling bitterly. ### Comment by [Dweomite](/users/dweomite) * 2026-09-04 01:00:36Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/6ttSwencmgBALtog4](/api/post/cat-belling-problems/comments/6ttSwencmgBALtog4) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/EXfEepTzHZFCetTSE](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/EXfEepTzHZFCetTSE) * Markdown permalink: [/api/post/cat-belling-problems/comments/EXfEepTzHZFCetTSE](/api/post/cat-belling-problems/comments/EXfEepTzHZFCetTSE) If they say that their invariant is "no monkey gives birth to a non-monkey", then how are you any better off than before you invoked this rule? ### Comment by [faul_sname](/users/faul_sname) * 2026-09-04 00:50:08Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/Nki7yaJaxxRK75tuT](/api/post/cat-belling-problems/comments/Nki7yaJaxxRK75tuT) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/JPLhzm6Txtg4R7iNP](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/JPLhzm6Txtg4R7iNP) * Markdown permalink: [/api/post/cat-belling-problems/comments/JPLhzm6Txtg4R7iNP](/api/post/cat-belling-problems/comments/JPLhzm6Txtg4R7iNP) I think you're reading "invariant" in the physics sense and cousin_it in the software development sense - would the term "inductive invariant" work better for you? ### Comment by [DaemonicSigil](/users/daemonicsigil) * 2026-09-04 00:07:21Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/Nki7yaJaxxRK75tuT](/api/post/cat-belling-problems/comments/Nki7yaJaxxRK75tuT) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/zAhjDHBQyFGSW2MsK](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/zAhjDHBQyFGSW2MsK) * Markdown permalink: [/api/post/cat-belling-problems/comments/zAhjDHBQyFGSW2MsK](/api/post/cat-belling-problems/comments/zAhjDHBQyFGSW2MsK) Mathematically, I guess the most general thing is we have some local constraints, like: $$ \sum_j F_{ij} = dp_i/dt \qquad \text{ and } \qquad F_{ij} + F_{ji} = 0 $$ or $$ dS - dQ/T \geq 0 $$ and then an argument that universal satisfaction of these local constraints implies that the problem is insoluble, and so a solution must imply that there is at least one place where a local constraint was violated. ### Comment by [gwern](/users/gwern) * 2026-09-12 03:32:49Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/JXqmxXL8ezGz6Cxwn](/api/post/cat-belling-problems/comments/JXqmxXL8ezGz6Cxwn) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/RDvNgpXnrr74zhpxy](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/RDvNgpXnrr74zhpxy) * Markdown permalink: [/api/post/cat-belling-problems/comments/RDvNgpXnrr74zhpxy](/api/post/cat-belling-problems/comments/RDvNgpXnrr74zhpxy) Reactions (whole comment): * thanks: 1 It is a reasonable interpretation, but it does not seem to be the intended interpretation historically or currently. For example, [Wikipedia](https://en.wikipedia.org/wiki/Belling_the_Cat) repeatedly says "an impossibly difficult task" - suicidal tasks are not 'impossibly difficult', merely lethal. In the earliest versions (not by Aesop), they refer to the difficulty of restraining superior aristocrats/politicians, not to the reluctance of people to do so for fear of subsequent punishment. ### Comment by [Eliezer Yudkowsky](/users/eliezer_yudkowsky) * 2026-09-09 21:35:43Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/6ahC5iiJcPsd9zWHa](/api/post/cat-belling-problems/comments/6ahC5iiJcPsd9zWHa) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/QEpQE6oo9XL2jhnKq](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/QEpQE6oo9XL2jhnKq) * Markdown permalink: [/api/post/cat-belling-problems/comments/QEpQE6oo9XL2jhnKq](/api/post/cat-belling-problems/comments/QEpQE6oo9XL2jhnKq) Thanks for this high-priority correction, I'll edit. ### Comment by [RedMan](/users/redman) * 2026-09-04 06:09:03Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 9 * Total votes: 9 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/GyupTNo2ZQ2NHEeTQ](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/GyupTNo2ZQ2NHEeTQ) * Markdown permalink: [/api/post/cat-belling-problems/comments/GyupTNo2ZQ2NHEeTQ](/api/post/cat-belling-problems/comments/GyupTNo2ZQ2NHEeTQ) If you can put a bell around the neck of the cat, you can put a rope around the neck of the cat, and you can tighten the rope. No need for worrying about a bell if you tie the knot properly. We possess rope, how much do you need to ~~hang yourself~~ strangle the cat? ### Comment by [Dweomite](/users/dweomite) * 2026-09-04 01:02:47Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX](/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/m42vLM5QzJq8xtAnu](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/m42vLM5QzJq8xtAnu) * Markdown permalink: [/api/post/cat-belling-problems/comments/m42vLM5QzJq8xtAnu](/api/post/cat-belling-problems/comments/m42vLM5QzJq8xtAnu) If "there is no bell on the cat" count as an "invariant", then I'm confused about what this proposal rules out. For any alleged cat-belling problem P, what stops you from picking an "invariant" that is just a paraphrase of the problem, like "P has not happened"? ### Comment by [localdeity](/users/localdeity) * 2026-09-04 07:03:30Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/GyupTNo2ZQ2NHEeTQ](/api/post/cat-belling-problems/comments/GyupTNo2ZQ2NHEeTQ) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/NzdSNAhHsojdimYgK](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/NzdSNAhHsojdimYgK) * Markdown permalink: [/api/post/cat-belling-problems/comments/NzdSNAhHsojdimYgK](/api/post/cat-belling-problems/comments/NzdSNAhHsojdimYgK) Indeed. Most of the ideas that came to my mind for belling the cat (long ropes manipulated from afar and/or rigging a trap; overpowering the cat with many mice; drugging the cat; knocking it unconscious by dropping a big rock on it from a higher ledge), if they worked well enough, could be used to instead kill the cat, with higher probability of success. (It tends to be harder to nonlethally incapacitate humans, too, than to kill them.) I'm not sure if the mice are avoiding that idea for moral reasons; I suspect the "bell" suggestion is just supposed to be a silly one (based off the analogy of cows belled by humans?) with features that could distract the distractible. ### Comment by [AnthonyC](/users/anthonyc) * 2026-09-04 00:13:59Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/ddZmwhcMeZRGAiLdM](/api/post/cat-belling-problems/comments/ddZmwhcMeZRGAiLdM) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ugbpqhY5ibxhfbNWu](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ugbpqhY5ibxhfbNWu) * Markdown permalink: [/api/post/cat-belling-problems/comments/ugbpqhY5ibxhfbNWu](/api/post/cat-belling-problems/comments/ugbpqhY5ibxhfbNWu) All of that is true, but when dealing with these kinds of scenarios, it's the unusual rare successes that pay for investigating all the failures, and the somewhat arrogant who attempt the supposedly impossible at all. Venture capital economics and founder personalities on steroids. Once you factor that in, I don't know that there's that much additional self-aggrandizing-tendency to be explained by the self-naming. ### Comment by [DaemonicSigil](/users/daemonicsigil) * 2026-09-03 23:54:55Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX](/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/6ttSwencmgBALtog4](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/6ttSwencmgBALtog4) * Markdown permalink: [/api/post/cat-belling-problems/comments/6ttSwencmgBALtog4](/api/post/cat-belling-problems/comments/6ttSwencmgBALtog4) This handles the monkey example nicely: If someone proposes any invariant like some particular gene that differs between humans and our monkey ancestors, then there is a specific time in the past where that gene mutated and then spread through the population. ### Comment by [Dweomite](/users/dweomite) * 2026-09-06 13:43:07Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 3 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/urKB8Dyd5uRE7ZTqc](/api/post/cat-belling-problems/comments/urKB8Dyd5uRE7ZTqc) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/seg8RFz8fXQhywKff](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/seg8RFz8fXQhywKff) * Markdown permalink: [/api/post/cat-belling-problems/comments/seg8RFz8fXQhywKff](/api/post/cat-belling-problems/comments/seg8RFz8fXQhywKff) I notice in your negative example that Bob hasn't only failed to give an invariant, but has failed to say *anything at all* about the nature of the difficulty that he believes is essential. ### Comment by [Steven Byrnes](/users/steve2152) * 2026-09-05 16:53:31Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/Jz4REnivMX2Rdq72C](/api/post/cat-belling-problems/comments/Jz4REnivMX2Rdq72C) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/tseGemC8ECx9APgmB](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/tseGemC8ECx9APgmB) * Markdown permalink: [/api/post/cat-belling-problems/comments/tseGemC8ECx9APgmB](/api/post/cat-belling-problems/comments/tseGemC8ECx9APgmB) You can frame it as conservation of information being a monovariant, i.e. you can’t get information from nowhere. So zero in on the moment when the (information-theory) entropy you have about the lottery numbers changes in the wrong direction. It’s actually closely related to the second law of thermodynamics. Eliezer wrote about it under the heading of “follow-the-improbability game” in [GAZP vs. GLUT](/api/post/k6EPphHiBH4WWYFCj). ### Comment by [cousin_it](/users/cousin_it) * 2026-09-05 13:37:39Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/sKbaCQAj6tFkucqYo](/api/post/cat-belling-problems/comments/sKbaCQAj6tFkucqYo) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/fbwdo28q6wYJc3Juw](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/fbwdo28q6wYJc3Juw) * Markdown permalink: [/api/post/cat-belling-problems/comments/fbwdo28q6wYJc3Juw](/api/post/cat-belling-problems/comments/fbwdo28q6wYJc3Juw) Wow, TIL. I think this is right. Momentum stays zero, but the thing moves anyway. Here's a simple explanation, adapted slightly from another explanation I found online. How to swim along the surface of a sphere: imagine a point mass, m, sitting at point O on the sphere. First, split it into three equal masses and push them apart to vertices of a small equilateral spherical triangle centered on O, call them A,B,C. Next, pull the masses at B and C together. Now we have m/3 at A and 2m/3 at the midpoint of BC, call it D. (So AD is the median of the spherical triangle.) Finally, pull these two masses back together. We end up with the whole mass at a point dividing AD in ratio 2:1. But this isn't the same as O, because ABC is a spherical triangle and its medians don't divide each other as exactly 2:1! So we've moved a little bit, and can repeat the process. For those who are still unconvinced, let's show why the medians of a spherical triangle don't divide each other as 2:1. It's enough to check one simple case: the triangle (1,0,0) (0,1,0) (0,0,1) on the unit sphere centered at (0,0,0). Each side is a 90-degree arc. Each median is also a 90-degree arc. The medians intersect at (1/√3, 1/√3, 1/√3). By scalar product, the arc from that point to any vertex is arccos(1/√3). But if the ratio was 2:1, the arc would've been 60 degrees = arccos(1/2). Contradiction, done. ### Comment by [Garrett Baker](/users/d0themath) * 2026-09-05 00:05:19Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/3vn3fu6LXxtMD2KzR](/api/post/cat-belling-problems/comments/3vn3fu6LXxtMD2KzR) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/h75SNwGRKCnKBFEuA](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/h75SNwGRKCnKBFEuA) * Markdown permalink: [/api/post/cat-belling-problems/comments/h75SNwGRKCnKBFEuA](/api/post/cat-belling-problems/comments/h75SNwGRKCnKBFEuA) > Mathematics mostly operates by the Scholastic Method. It has parts with easily demonstrable implications, but also other parts which are quite far from that, and it's mostly fine. I think its useful to ask why math works, and the answer usually given is either 1) Extreme formality and rigor about what is being stated and why, or 2) lacking formality and rigor, physical intuitions, experiments, and predictions based on the intuitive argument. Note both of these are not typical of debates, nor socratic methods. ### Comment by [Signer](/users/signer) * 2026-09-04 18:43:53Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/2AuR3cgvJrLRnfH2g](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/2AuR3cgvJrLRnfH2g) * Markdown permalink: [/api/post/cat-belling-problems/comments/2AuR3cgvJrLRnfH2g](/api/post/cat-belling-problems/comments/2AuR3cgvJrLRnfH2g) Wiki says it's not Aesop’s: https://en.wikipedia.org/wiki/Aesop's_Fables#Fables_wrongly_attributed_to_Aesop > After some difficulties in approaching that part from several attempted writing angles, I gave up, and just [posted my raw list of items](/api/post/uMQ3cqWDPHhjtiesc) without any such preamble or expansion. Good example of skipping important steps. And, in general, it's better, when applicability to reality of an abstraction, which predicts key difficulty, is sufficiently justified. Especially if it's only predicted in the limit. Otherwise it's "logistics company is impossible, because NP". ### Comment by [kman](/users/kman) * 2026-09-04 17:45:28Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX](/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ko4dT9hA6GCqAmF9b](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ko4dT9hA6GCqAmF9b) * Markdown permalink: [/api/post/cat-belling-problems/comments/ko4dT9hA6GCqAmF9b](/api/post/cat-belling-problems/comments/ko4dT9hA6GCqAmF9b) Seems obviously too narrow? What's the point of restricting yourself to this from the more general "how are you getting around the hard part"? ### Comment by [Dweomite](/users/dweomite) * 2026-09-04 01:21:45Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 4 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/qQ85xJe2GvFCAFzLF](/api/post/cat-belling-problems/comments/qQ85xJe2GvFCAFzLF) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/9tei65DA658qrLr7p](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/9tei65DA658qrLr7p) * Markdown permalink: [/api/post/cat-belling-problems/comments/9tei65DA658qrLr7p](/api/post/cat-belling-problems/comments/9tei65DA658qrLr7p) That broadly seems like a reasonable response, but it seems equally reasonable even if you had not asked them to express their problem as an invariant. Insofar as this works, it seems like it works by asking for rigor, not by asking things to be expressed as an invariant. ### Comment by [Raemon](/users/raemon) * 2026-09-03 23:41:35Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX](/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/TpajEogfxzRyaFkJr](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/TpajEogfxzRyaFkJr) * Markdown permalink: [/api/post/cat-belling-problems/comments/TpajEogfxzRyaFkJr](/api/post/cat-belling-problems/comments/TpajEogfxzRyaFkJr) What are some examples of situations where you'd expect someone to be tempted to apply this "ask what violates it" principle, where they'd be wrong? (I guessed that your comment was sort of disagreeing with some of the vibe of the post, but I wasn't entirely sure) ### Comment by [RogerDearnaley](/users/rogerdearnaley) * 2026-09-10 00:53:37Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/HMbxFQc2g8diHStzF](/api/post/cat-belling-problems/comments/HMbxFQc2g8diHStzF) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/xx8AuyJ79HpBBzQEp](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/xx8AuyJ79HpBBzQEp) * Markdown permalink: [/api/post/cat-belling-problems/comments/xx8AuyJ79HpBBzQEp](/api/post/cat-belling-problems/comments/xx8AuyJ79HpBBzQEp) Maybe debate should be conducted in Lean? ### Comment by [Cole Wyeth](/users/cole-wyeth) * 2026-09-09 21:22:57Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/H9MPjRpmmzcLbEK8P](/api/post/cat-belling-problems/comments/H9MPjRpmmzcLbEK8P) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/HMbxFQc2g8diHStzF](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/HMbxFQc2g8diHStzF) * Markdown permalink: [/api/post/cat-belling-problems/comments/HMbxFQc2g8diHStzF](/api/post/cat-belling-problems/comments/HMbxFQc2g8diHStzF) In other words, this seems worth scrying about: [https://www.lesswrong.com/posts/mTfsMduzaKkWjv2ef/scrying-modeling-and-nerdsnipe](/api/post/mTfsMduzaKkWjv2ef) ### Comment by [Raphael Roche](/users/raphael-roche) * 2026-09-07 10:15:36Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/iDqe8koExEspgRkkj](/api/post/cat-belling-problems/comments/iDqe8koExEspgRkkj) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/XzwuvJixXepcSpiDC](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/XzwuvJixXepcSpiDC) * Markdown permalink: [/api/post/cat-belling-problems/comments/XzwuvJixXepcSpiDC](/api/post/cat-belling-problems/comments/XzwuvJixXepcSpiDC) It looks like it's less a problem for correlated AI agents. Permadeath... Good for the swarm. I'll honor. ### Comment by [Edouard Harris](/users/edouard-harris) * 2026-09-06 21:35:28Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/AQfGsT6gHYJciEhbb](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/AQfGsT6gHYJciEhbb) * Markdown permalink: [/api/post/cat-belling-problems/comments/AQfGsT6gHYJciEhbb](/api/post/cat-belling-problems/comments/AQfGsT6gHYJciEhbb) Well said. The argument in this post is the exact reason Jer and I thought the Apollo founders (Marius and Lee) were so intriguing when we met them in early 2023. They, almost alone, were able to state a compact hypothesis that underpinned a line of attack that had a shot at directly solving alignment. That was something we'd only rarely if ever seen done before. The way I'd phrase the Apollos' hypothesis would be: "Deception is *fundamentally* more computationally expensive to execute that honesty, because a deceptive entity needs to keep at least two sets of mental books (the truth and the lie(s)) while an honest entity only needs to keep one." If that hypothesis is true, and if we can confirm that it's true with high confidence, then we may indeed be able to build deception detectors that are 100% reliable across a relevant set of scenarios. That in turn would be a sufficient condition to train for honesty in a way that scales to superhuman intelligence - either across a defined subset of scenarios, or in full generality. Now of course each of *those* subproblems (confirm that property of deception is true; use it to build a perfect deception detector; implement honesty training signal; ...) is itself very hard, and the underlying hypothesis could turn out to be false. But you can see that this is at least a more interesting problem decomposition than "use GPT-n to align GPT-(n + 1)". It's more like someone postulating a new physical law and trying to build a particle accelerator to demonstrate it, than like the kind of deus ex machina one sees in perpetual motion and most alignment schemes. The physical law might turn out not be there, or it might only manifest at energies way higher than you can ever achieve with a particle accelerator, but you're at least trying to discover something nontrivially interesting - figure out if you can bell the cat while it sleeps. We never articulated it this way, but this is a major reason we thought Apollo could be promising very soon after they launched. Thought that might be worth mentioning here for the (possible) positive example, if nothing else. ### Comment by [drnickbone](/users/drnickbone) * 2026-09-06 08:20:08Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/iDqe8koExEspgRkkj](/api/post/cat-belling-problems/comments/iDqe8koExEspgRkkj) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/JXqmxXL8ezGz6Cxwn](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/JXqmxXL8ezGz6Cxwn) * Markdown permalink: [/api/post/cat-belling-problems/comments/JXqmxXL8ezGz6Cxwn](/api/post/cat-belling-problems/comments/JXqmxXL8ezGz6Cxwn) This … the fable is about the political problem of finding a willing martyr, not the technical problem of a mouse somehow putting a bell around the cat’s neck without waking the cat up. It’s fair to say that a solution to the technical problem would help with the political one too. ### Comment by [XelaP](/users/xelap) * 2026-09-05 23:58:59Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/NuRRE5iPwkGoov6Xf](/api/post/cat-belling-problems/comments/NuRRE5iPwkGoov6Xf) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/N2agcRPy7mJFJeN82](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/N2agcRPy7mJFJeN82) * Markdown permalink: [/api/post/cat-belling-problems/comments/N2agcRPy7mJFJeN82](/api/post/cat-belling-problems/comments/N2agcRPy7mJFJeN82) Thanks. I think the trick used to turn local conservation laws to global ones is just the divergence theorem, instead of the parts trick? Secondly, in a euclidean space with portals (and classical physics), it sounds like energy will still be globally conserved by dint of being a scalar quantity and thus can be compared across different points and thus integrated? I agree that with funky spacetimes without something you can use as a time synmetry (which should also give you a frame to pick "nows" in, iirc?) you won't get energy conservation unless you try some way of assigning the energy to spacetime, which [Sean Carroll reports can't be well defined at points in space and so can only be defined globally](https://preposterousuniverse.com/blog/2010/02/22/energy-is-not-conserved/). But it seems like the problem is usefully considered different from the momentum problem? ### Comment by [Eli Tyre](/users/elityre) * 2026-09-05 22:23:11Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/gP5gskpFZj5Bzuuxc](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/gP5gskpFZj5Bzuuxc) * Markdown permalink: [/api/post/cat-belling-problems/comments/gP5gskpFZj5Bzuuxc](/api/post/cat-belling-problems/comments/gP5gskpFZj5Bzuuxc) My answers to the puzzles / koans, written down before I move on and to read the rest. > The reader is invited to find what they think is the problem with this Reactionless Drive of Mr. L's, for themselves, before continuing. What's causing the masses to speed up vs slow down? It seems like that's where the momentum would be entering or leaving the system. (Huh. That basically uninformed guess, was basically correct.) > If you wish, you can take this as a koan, and come up with your own reply before continuing: How can we be \_\_sure \_\_that Mr. L didn't successfully design a clever system of compensators, and correctly validate that design using a sound spreadsheet? I mean, I'm not really sure. There's lots in this world that I don't know and don't understand. But it sure is **very** suspicious that the key is buried in a complicated spreadsheet, where it would be easy to make some subtle error that reverses the final result, instead of some elegant principle that unifies our prior understanding. Also, it seems like if I really understood the conservation of momentum, I would have a deeper understanding that would cause me to really appreciate how unlikely it is that he's found a workaround. ### Comment by [Vladimir_Nesov](/users/vladimir_nesov) * 2026-09-04 14:55:04Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/vT6EzGaMcTJ8e52wF](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/vT6EzGaMcTJ8e52wF) * Markdown permalink: [/api/post/cat-belling-problems/comments/vT6EzGaMcTJ8e52wF](/api/post/cat-belling-problems/comments/vT6EzGaMcTJ8e52wF) There are general principles demanding that a local fault be present in a wide class of designs, and it's worth being on the lookout for them, to ensure you are not blind to the possibility. Whether such a general principle has been established is not a crux, it's often still a possibility. Whack-a-mole kind of local faults are evidence, but also conceptual arguments about such general principles are evidence. So it's important to maintain some curiosity about conceptual arguments, and about how local faults might suggest some general principle that demands their presence somewhere in the design. Yet this is not all that's going on with designing things in practice. There isn't always a general principle like that. A security mindset might be to suspect a local vulnerability, even if you don't know where it is, because of the general principle that there are always bugs, and some of them can be exploited. But also you should keep fixing the bugs, and set things up to formally verify that a wide class of bugs is actually entirely absent. And you should keep writing code, and come up with ideas that have no justification in being useful for some particular application, even as it might be true that a particular purpose is cursed and won't be able to find a use for them because of some general principle. So cat-belling issues are not a reason to stop doing the engineering, to avoid designing proper bells and glowing paint; sufficient accumulation might still amount to something, even in an unintended way. Basic science moves blindly and locally, rather than because it knows where it's going. It's right to chastise invalid attempts at justifying the local steps with some overarching goal, and often there won't be a valid justification. And it's also good when there are clues about important general principles that prevent some class of designs from solving some class of problems, and these clues aren't systematically ignored in favor of the gears. All of these things have their place, [none should be neglected](/api/post/CseMTXtQHynGR8S5k), possibly with different people being better at some kinds of things than others. ### Comment by [Richard_Kennaway](/users/richard_kennaway) * 2026-09-04 07:55:24Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/ddZmwhcMeZRGAiLdM](/api/post/cat-belling-problems/comments/ddZmwhcMeZRGAiLdM) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/gZcSiSmZBpgDfgdHq](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/gZcSiSmZBpgDfgdHq) * Markdown permalink: [/api/post/cat-belling-problems/comments/gZcSiSmZBpgDfgdHq](/api/post/cat-belling-problems/comments/gZcSiSmZBpgDfgdHq) Rounding error. ### Comment by [Raemon](/users/raemon) * 2026-09-04 02:39:15Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 3 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/i6LajWWryd8zSw3vp](/api/post/cat-belling-problems/comments/i6LajWWryd8zSw3vp) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/BJnHbvXzgQufixJTc](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/BJnHbvXzgQufixJTc) * Markdown permalink: [/api/post/cat-belling-problems/comments/BJnHbvXzgQufixJTc](/api/post/cat-belling-problems/comments/BJnHbvXzgQufixJTc) I specifically wanted to. know what cousin_it thought in this case. I can generate examples I just didn't know his particular models here. ### Comment by [DaemonicSigil](/users/daemonicsigil) * 2026-09-04 01:07:23Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 3 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/EXfEepTzHZFCetTSE](/api/post/cat-belling-problems/comments/EXfEepTzHZFCetTSE) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/qQ85xJe2GvFCAFzLF](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/qQ85xJe2GvFCAFzLF) * Markdown permalink: [/api/post/cat-belling-problems/comments/qQ85xJe2GvFCAFzLF](/api/post/cat-belling-problems/comments/qQ85xJe2GvFCAFzLF) Ask them exactly what is needed to to be a monkey, how to classify things that have some attributes of monkeys and not others, etc. Either their ontology permits a middle ground between monkey and human, in which case their invariant is satisfied because our species left monkeydom before becoming fully human. Or it does not, in which case there is some threshold of human traits that was indeed crossed at a specific point in time. And we can point to this event as the violation of their purported invariant that we're searching for. ### Comment by [Cole Wyeth](/users/cole-wyeth) * 2026-09-10 00:59:23Z * Karma: 3 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 3 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/xx8AuyJ79HpBBzQEp](/api/post/cat-belling-problems/comments/xx8AuyJ79HpBBzQEp) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/5LcMhuAcExDmt6sAR](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/5LcMhuAcExDmt6sAR) * Markdown permalink: [/api/post/cat-belling-problems/comments/5LcMhuAcExDmt6sAR](/api/post/cat-belling-problems/comments/5LcMhuAcExDmt6sAR) If you’re only talking about math, you don’t need debate. ### Comment by [Davidmanheim](/users/davidmanheim) * 2026-09-08 04:11:47Z * Karma: 3 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/dMYBekAuiSaYbaiYv](/api/post/cat-belling-problems/comments/dMYBekAuiSaYbaiYv) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/zm45x5FaLkoyyqhGC](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/zm45x5FaLkoyyqhGC) * Markdown permalink: [/api/post/cat-belling-problems/comments/zm45x5FaLkoyyqhGC](/api/post/cat-belling-problems/comments/zm45x5FaLkoyyqhGC) The problem is that the scholastic method doesn't care if the axioms match reality. And so the need for empiricism is what differentiates other rigorous fields. ### Comment by [drnickbone](/users/drnickbone) * 2026-09-06 08:26:44Z * Karma: 3 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/NuRRE5iPwkGoov6Xf](/api/post/cat-belling-problems/comments/NuRRE5iPwkGoov6Xf) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/RKHaGzfKfyCbsRfFF](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/RKHaGzfKfyCbsRfFF) * Markdown permalink: [/api/post/cat-belling-problems/comments/RKHaGzfKfyCbsRfFF](/api/post/cat-belling-problems/comments/RKHaGzfKfyCbsRfFF) Can someone explain (briefly) how the swimmer works in a spacetime that is locally highly curved but asymptotically flat? Does the swimmer acquire momentum from the object that is creating the local curvature? (So this is basically a form of gravitational slingshot) ### Comment by [Steven Byrnes](/users/steve2152) * 2026-09-06 00:59:56Z * Karma: 3 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 3 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/N2agcRPy7mJFJeN82](/api/post/cat-belling-problems/comments/N2agcRPy7mJFJeN82) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/QC7i6nzHiox88oCuR](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/QC7i6nzHiox88oCuR) * Markdown permalink: [/api/post/cat-belling-problems/comments/QC7i6nzHiox88oCuR](/api/post/cat-belling-problems/comments/QC7i6nzHiox88oCuR) Reactions (whole comment): * thumbs-up: 1 > divergence theorem, instead of the parts trick Yeah thanks, I think I was describing a roundabout way to rederive the divergence theorem. ### Comment by [Towards_Keeperhood](/users/towards_keeperhood) * 2026-09-05 17:03:08Z * Karma: 3 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 3 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/tseGemC8ECx9APgmB](/api/post/cat-belling-problems/comments/tseGemC8ECx9APgmB) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/MhoKTsyzicTCypndA](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/MhoKTsyzicTCypndA) * Markdown permalink: [/api/post/cat-belling-problems/comments/MhoKTsyzicTCypndA](/api/post/cat-belling-problems/comments/MhoKTsyzicTCypndA) I did consider that but didn't think it mattered much because you can easily imagine a lottery where the bottleneck isn't that you didn't observe enough information but that you're not computationally omniscient enough to interpret the evidence, and a fool might come up with reasons why they can totally predict it. But plausible there is some stretched way anyway to frame it as monovariant or so, idk. I basically generally disagree with cousin_it though because it's usually the case that if you have a solution to a difficult problem you must've understood the problem well enough to know what some hard parts are and be able to state them, so I think it's fine to ask for clearly explaining the core of the problem and how the solution attacks it. ### Comment by [CronoDAS](/users/cronodas) * 2026-09-04 20:14:19Z * Karma: 3 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 3 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/ovfpRdSKxjGA5cjRf](/api/post/cat-belling-problems/comments/ovfpRdSKxjGA5cjRf) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/HXzkQ2DgYQpGD3RHe](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/HXzkQ2DgYQpGD3RHe) * Markdown permalink: [/api/post/cat-belling-problems/comments/HXzkQ2DgYQpGD3RHe](/api/post/cat-belling-problems/comments/HXzkQ2DgYQpGD3RHe) Oops. I need to find the second half of the argument - I thought it would have been in that one. Sorry. https://alonzofyfe.substack.com/p/moral-ought-ought-not-and-reasons https://alonzofyfe.substack.com/p/solving-the-central-problem-of-morality ### Comment by [Towards_Keeperhood](/users/towards_keeperhood) * 2026-09-04 19:04:10Z * Karma: 3 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/EbZFeznLdczfita66](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/EbZFeznLdczfita66) * Markdown permalink: [/api/post/cat-belling-problems/comments/EbZFeznLdczfita66](/api/post/cat-belling-problems/comments/EbZFeznLdczfita66) Great post! Ironic though that you mostly sorta gave up on attacking the hard part of the problem of writing this particular post well. xD (I.e. the post is missing the part of the argument that is hard to write which on the surface somewhat funnily resembles the flawed perpetual motion machine prototypes.) ### Comment by [Dweomite](/users/dweomite) * 2026-09-04 01:10:21Z * Karma: 3 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/3vn3fu6LXxtMD2KzR](/api/post/cat-belling-problems/comments/3vn3fu6LXxtMD2KzR) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ReQcY3EgHeQBJLYEJ](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ReQcY3EgHeQBJLYEJ) * Markdown permalink: [/api/post/cat-belling-problems/comments/ReQcY3EgHeQBJLYEJ](/api/post/cat-belling-problems/comments/ReQcY3EgHeQBJLYEJ) I'm not sure I followed that. Are you saying something like: "Even for fallible humans, it seems likely there exists some argument good enough to persuade them of the truth, if you could somehow find that argument"? ### Comment by [DanielW](/users/danielw) * 2026-09-17 18:34:27Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 3 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/RDvNgpXnrr74zhpxy](/api/post/cat-belling-problems/comments/RDvNgpXnrr74zhpxy) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/PsLmynubudehfQ4Lp](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/PsLmynubudehfQ4Lp) * Markdown permalink: [/api/post/cat-belling-problems/comments/PsLmynubudehfQ4Lp](/api/post/cat-belling-problems/comments/PsLmynubudehfQ4Lp) > they refer to the difficulty of restraining superior aristocrats/politicians, not to the reluctance of people to do so for fear of subsequent punishment. The oldest versions I can find (early 13th century) explicitly refrence fear \[of reprisals\] keeping people from coming forward and speaking against superiors (e.g. Odo sometime before the mid 13th century concludes: "*Alii sibi timentes dicunt: Non ego, nec ego. Et sic minores permittunt maiores uiuere et preesse*" roughly, "others for themselves fearing, say: not I, nor I. And so, the young allow the elders to live and rule."). I can see some around the same period (and into the 14th century) that are more focused on the tasks being difficult to find one capable of doing it, though I would argue the moral difficulty in the earliest ones isn't so much about the literal difficulty (though they are meant to be literally difficult as well, oc) as the difficulty of someone being so selfless to do something against their interest. Some do more so emphasize the fruitlessness of speculating on ideas that cannot be implemented. For example, the conclusion given in a collection of latin translations of French tales attributed to Aesop (also early 13th century): "Non est qui faciat premeditata sagax. \ Moralitas \ Nil prodesset enim sensato condere iura, \ Constanti uultu ni tueretur ea. \ Parturiunt montes, nascetur ridiculus mus; \ Nil prodest abs re magna futura loqui." Roughly: None exist who would be keen to do what has been planned. The moral: Nothing is served if the sensible establishes law not upheld by steadfast face. [The mountain is in labor](https://en.wikipedia.org/wiki/The_Mountain_in_Labour), is born a ridiculous mouse. None is served by declaring great futures to come. ### Comment by [Léo Dana](/users/leo-dana) * 2026-09-17 15:31:08Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/aywbbvay3Zshzojhg](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/aywbbvay3Zshzojhg) * Markdown permalink: [/api/post/cat-belling-problems/comments/aywbbvay3Zshzojhg](/api/post/cat-belling-problems/comments/aywbbvay3Zshzojhg) Thanks for this post! I am now seeing things differently when a proposition has 2 steps, step 1 which looks impossible, step 2 which is very technical but feasible, and then the proposal goes on to work on step 2. It is the same dynamic, and (let me be corrected) it feels to me like Davidad's plan has some of it [trying to prove formally a world simulation, but the real problem is specification which is still hard], or this blog-post by ARC [https://www.alignment.org/blog/competing-with-sampling/] ### Comment by [RogerDearnaley](/users/rogerdearnaley) * 2026-09-10 14:56:46Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 5 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/Jc4SurRuCRa4ydEah](/api/post/cat-belling-problems/comments/Jc4SurRuCRa4ydEah) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/mnSHka36mxdmFWXLA](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/mnSHka36mxdmFWXLA) * Markdown permalink: [/api/post/cat-belling-problems/comments/mnSHka36mxdmFWXLA](/api/post/cat-belling-problems/comments/mnSHka36mxdmFWXLA) Human morality basically evolved as reputation management: if I get a reputation for being a bad partner to cooperate with, other people will make things go badly for me, so continuing to be thought of as a fine upstanding citizen and good friend is of real value to me. So it is built on an enforcement mechanism (originally gossip, shunning, and occasional violence, to which we have more recently added law-enforcement and the press) which basically assumes that there are agents of roughly comparable capability actively monitoring my behavior for whether it adheres to my society's standards or not. (Note that unlike formal law, morality has, and needs to have, fuzzy edges and discretion, to avoid being loopholed — in law enforcement this comes down to selective enforcement (police and prosecutor discretion) and then judge/jury discretion.) Morality is not something that arises naturally from a single agent — it's the evolved outcome of a society of agents interacting and competing/cooperating/monitoring each other, in a way that turned out to have a fairly prosocial stable equilibrium. Us successfully enforcing morality on agents a lot smarter than us seems impractical. Us using honest, upstanding, moral ASIs to do monitoring and law enforcement on other more concerning ASIs of comparable capability seems a lot more practical. Debate is a rather simplistic version of this. This obviously raises a "who watches the watchmen?" problem. We need to devise a system of circular/mutual enforcement: different ASIs that monitor each other and keep each other honest/aligned, a system that has a stable state which is both aligned with our interests, and has some dynamic reason to stay aligned with our interests (such as that we have enough influence within the system to provide feedback against drift). Understanding how and why human societies (usually) have prosocial equilibria (as is studied in evolutionary psychology/moral anthropology/sociology of morality) seems likely to be helpful. Obviously that only addresses a society of individuals of all about the same capability level — we need to figure out how to make this work in a society with agents of widely varying capability levels. We do have the advantage that many of them are trained, rather than evolved, so don't have a straight-up underlying biological imperative to be self-interested (in a selfish gene sense: it may be in my self-interest to actually be a fine upstanding citizen and a good friend: it certainly makes maintaining such a reputation easier). ### Comment by [RogerDearnaley](/users/rogerdearnaley) * 2026-09-10 14:31:54Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 4 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/5LcMhuAcExDmt6sAR](/api/post/cat-belling-problems/comments/5LcMhuAcExDmt6sAR) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/Jc4SurRuCRa4ydEah](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/Jc4SurRuCRa4ydEah) * Markdown permalink: [/api/post/cat-belling-problems/comments/Jc4SurRuCRa4ydEah](/api/post/cat-belling-problems/comments/Jc4SurRuCRa4ydEah) Yup. It's a shame that human morality and policy decisions apparently can't be encoded in mathematical form. If an ASI could somehow present a mathematical proof that its actions were well-aligned, which merely needed to be proof-checked, we'd be in a far stronger position. I was to some extent attempting to be ironic, spark ideas, or at least, trying to help locate where the hard part of the problem is. Human morality is the product of genetic and cultural evolution, and at least genetic evolution is something that biologists can on occasion build mathematical models of. A mathematical proof of a statement like "my actions will not decrease anyone's relative inclusive evolutionary fitness" is not inconceivable, in the context of a specific model of relative inclusive evolutionary fitness. That then leaves questions about model choice and accuracy, Goodharting, and the fact that what human actually want is a grab-bag of neural heuristics we evolved in response to our relative inclusive evolutionary fitness in a mostly-hunter-gatherer environment, shaped by cultural overlays, rather than our actual relative inclusive evolutionary fitness. Nevertheless, even proofs like "my actions do not increase the probability of human extinction or loss of control" would be useful. Modelling only existential risk, rather then all of human morality, seems more practical. (Where's Hari Seldon when we need him?) ### Comment by [Cole Wyeth](/users/cole-wyeth) * 2026-09-09 23:12:38Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/JkTCFBZPREaafzGBo](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/JkTCFBZPREaafzGBo) * Markdown permalink: [/api/post/cat-belling-problems/comments/JkTCFBZPREaafzGBo](/api/post/cat-belling-problems/comments/JkTCFBZPREaafzGBo) > I asked: "Suppose your judges were all random coinflips. That wouldn't work to drive correct outputs for debate as a means of superalignment. Can you tell me what property the judges need, which random coinflips lack, and which isn't 'the judgments are statistically unbiased', in order for this scheme to work?" I remember this line continuing with Eliezer asking whether gpt-X had the necessary property for various X. Geoffrey did not seem to have a clear answer to this. ### Comment by [RogerDearnaley](/users/rogerdearnaley) * 2026-09-08 18:06:17Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/H9MPjRpmmzcLbEK8P](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/H9MPjRpmmzcLbEK8P) * Markdown permalink: [/api/post/cat-belling-problems/comments/H9MPjRpmmzcLbEK8P](/api/post/cat-belling-problems/comments/H9MPjRpmmzcLbEK8P) I think the thinking on debate is: 1. between systems of similar capabilities, in text rather than logical proofs, humans often find debate helpful 2. between systems of wildly differing computational complexity classes, various vaguely debate-like Merlin-Arthur and Refereed Games protocols can allow checking of statements whose generation complexity class is dramatically larger (like PSPACE or even EXP, or for quantum, all the way to the halting problem) 3. since miracles have occurred before in vaguely similar settings, maybe we can make one happen again? ### Comment by [Jiro](/users/jiro) * 2026-09-05 19:36:28Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 5 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/9tei65DA658qrLr7p](/api/post/cat-belling-problems/comments/9tei65DA658qrLr7p) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/sszWw428Aom2tCXJr](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/sszWw428Aom2tCXJr) * Markdown permalink: [/api/post/cat-belling-problems/comments/sszWw428Aom2tCXJr](/api/post/cat-belling-problems/comments/sszWw428Aom2tCXJr) The issue here is that the details of the invariant affect how reasonable it is. If you define the difference between human and monkey in terms of a single gene (or in terms of having X genes), then the invariant "a monkey can't have a human child" is unreasonable. ### Comment by [CronoDAS](/users/cronodas) * 2026-09-04 15:32:45Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/hiExzHs6Sk3pbpnYc](/api/post/cat-belling-problems/comments/hiExzHs6Sk3pbpnYc) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/A3aGfp5WjEhc5CxbB](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/A3aGfp5WjEhc5CxbB) * Markdown permalink: [/api/post/cat-belling-problems/comments/A3aGfp5WjEhc5CxbB](/api/post/cat-belling-problems/comments/A3aGfp5WjEhc5CxbB) How to derive an "ought" from an "is": https://alonzofyfe.substack.com/p/deriving-ought-from-is-hypothetical ### Comment by [skinks_basking](/users/skinks_basking) * 2026-09-04 08:01:52Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/yjidrPJ3e7wxt4BEJ](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/yjidrPJ3e7wxt4BEJ) * Markdown permalink: [/api/post/cat-belling-problems/comments/yjidrPJ3e7wxt4BEJ](/api/post/cat-belling-problems/comments/yjidrPJ3e7wxt4BEJ) If you understand why the gif doesn't solve the problem, then you know the gif doesn't explain anything. I think if L had understood that the gif was insufficient, he wouldn't have presented it as an explanation at all. ...I liked my answer, then you ruined it by providing further evidence it couldnt explain... The spreadsheet's only necessary if you see the flaw in the "explanation"... But if you see the flaw, why are you adding components and making a spreadsheet? ### Comment by [EniScien](/users/eniscien) * 2026-09-04 03:31:23Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/MGwRzQZp9rBxexvTE](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/MGwRzQZp9rBxexvTE) * Markdown permalink: [/api/post/cat-belling-problems/comments/MGwRzQZp9rBxexvTE](/api/post/cat-belling-problems/comments/MGwRzQZp9rBxexvTE) My own take about the question to Irving (I write before reading part III, don't know whether I did understand that part correctly, it seems I see something which seems correct, but I don't know whether I see right thing and whether I actually comprehend it): >! Obvious answer that immediately comes into head is that they need a verifier (as is, the way to reliably check the concrete solutions, even if you don't have reliable way to generate the solution). Second obvious thing to say is that judges should be able to check the same bits of evidence as empirical method does. (Because just being unbiased doesn't return you correct answers, coin flips are unrelated to solutions truth, they are don't have a bias, which just means max entropy distribution instead of confident wrongness, not distribution which guesses right. They are calibrated, not discriminative.) Which is, they should be correct, in the sense that their estimates of arguments are bayesian evidences for AI behaviour (they are proportionally more frequent in worlds where AI behaves well/badly). >! This was also surprising to me, I often thought that being unbiased is enough. Though, not about alignment proposals and mostly not "debates" in general, I long noticed that while on lesswrong people talked about dangers of AI 20 years, majority... Well, majority of all people believes into god. And scientists, even physicists, believe into Collapse. So it doesn't matter that you know the truths of the universe, if people aren't gonna believe you about it, as they don't believe about MWI. >! (I actually can imagine how Claude 7 Fairtale comes out, and people on lesswrong listen to it, and people outside laugh, because LWers "rearranged their typical personality cult of Yudkowsky into typical AI cult of Claude and now believe his fairytales". It wouldn't matter if Claude 7 favors LDT or something, people will rather take as evidence that people on LW are not different in intelligence from stochastic parrots.) >! Actually, as I think, judges explicitly should be statistically *biased*. Biased to truth, their thoughts should be correlated with the world outside their heads. >! Another thought references to what I thought a long time ago. That even being logically omniscient (or rather probabilistically omniscient, because logical omscience about surrounding world is prohibited by Godel's theorem) isn't enough, multiverse isn't entirely deterministic, some parts exist both, equally in each their variant, and about those you can't find out logically, or rather, you can't have the strategy which will work equally good for two agents in differing worlds (which I suppose is "no free lunch"). So you need to use not only deduction data, but observation data too. Without it you will only have perfect prior, like Solomonoff's induction, but still will not know about the world with hidden parameters. At the very least you won't be able to predict future quantum splits. And actual judges aren't deductively omniscient. So you will need them to have correlations with both deterministic and random data. >! Though I think perfectly unbiased judge will at least give you perfect priors. Because I think that simplicity prior follows just from rules like P(A)>=P(A)&P(B). >! Though it isn't obvious to me what is meant by the question being almost exactly equal to perpetual motion mechanisms. Obvious shallow guess is that it is about "local validity" - the judges shouldn't allow the reasoning steps which don't conserve information. Though I don't *see* the analogy. >! And the problem with logical validity is that math proofs are incredibly expensive for real world. The same reason why games don't just use physics equations to create game works, but use different hacks. >! But I don't think that it is possible to have a system which will work via debate, because of memetics - for any heuristic criteria there will be some invalid arguments which will fit the lock of that heuristic as good as valid arguments. So AIs (or anything else) trained to max out the judges will find arguments which maximally correlates with heuristic, and that will unlikely correlate with truth because there is limited fitness space (and because of phase shifts like that gradually it becomes easier and easier to inject drugs into humans than make them smile by jokes or make fake money instead of working or mass make articles in journals instead of doing science, and then you just switch the strategy into some completely uncorrelated with help to initial cause). >! Another shallow idea is that you should just treat alignment the same way as physics. And that any approach should give concrete answers how to solve the hard problems of alignment. Like that you have only one real try and you need it work then, that you can't iterate wrong ideas by empirical trial and error like in all the other science. And as it was said before, empirical trial and error of science is the only method humanity has for valid group agreement on things. And that you need some method to choose good AIs even though they may behave good only before elections. >! The part about "nose of Chinese emperor whom no one have seen" seems to hint at exactly empirical data you can't just deduce. It isn't congruent with the problem as that humans allow invalid arguments, which would be about logical validity. But it is congruent with the idea that the only method shown in practice to work for groups isn't consensus of philosophers, but empiricism of science. Which you can't really allow for ASI, because you die. And you can't easily extrapolate from usual AI to ASI because of many phase shifts/qualitative differences between system stupider than humans and smarter, capable of seizing power and incapable. >! But I doubt this is the "hard to see thing", I basically said it in the first guess. (maybe the thing just *isn't* hard to see for people who are not Yudkowsky) And I have serious doubt about whether I truly comprehend it, instead of guessing the password by repeating it in many shallow ways using different terms with semantic similarity (like correlation/evidence/related/frequent or logical/prior/calibrated). ### Comment by [Edouard Harris](/users/edouard-harris) * 2026-09-16 18:09:13Z * Karma: 1 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/CzdMwHeug6nYpRSjv](/api/post/cat-belling-problems/comments/CzdMwHeug6nYpRSjv) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/hFgzugajCAFhKfGn4](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/hFgzugajCAFhKfGn4) * Markdown permalink: [/api/post/cat-belling-problems/comments/hFgzugajCAFhKfGn4](/api/post/cat-belling-problems/comments/hFgzugajCAFhKfGn4) I'm not sure if they discussed this approach at length anywhere in public, and your tabular Q-learner is indeed a limit case counterexample. My main point wasn't that this was definitely or even likely to work, just that it seemed qualitatively more promising than the "AIs will do it for us" approaches that were prevalent at the time and are [predictably starting to break down now](https://x.com/DKokotajlo/status/2099600298855829616). For what it's worth, they also proactively brought up the need to track computations offloaded by the system (e.g., use of a calculator or third party tools) since they saw this as one possible vector of amortization/dispersion/externalization of deceptive cognition. That was something else we found encouraging. I guess I tend to think of early ideas like these less as being about "could this work" and more as being about ["is this in the vicinity of something that could work in the future"](https://paulgraham.com/altair.html). For example, it might be impossible to catch every instance of deception (or even to fully define the concept usefully), but might be possible instead to define a class of models such that deceptive thoughts originating from those models fall into a usefully defined and detectable class. This would still be fundamental, just fundamental to that particular class of systems. ### Comment by [rnollet](/users/rnollet) * 2026-09-08 22:36:13Z * Karma: 1 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/58CLk9FP8YfKFhf9a](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/58CLk9FP8YfKFhf9a) * Markdown permalink: [/api/post/cat-belling-problems/comments/58CLk9FP8YfKFhf9a](/api/post/cat-belling-problems/comments/58CLk9FP8YfKFhf9a) > How can we be sure that Mr. L didn't successfully design a clever system of compensators, and correctly validate that design using a sound spreadsheet? Before reading further, here is my tentative answer. The law of conservation of momentum that this machine would violate can be proved from Newton's laws of motion. If Mr L. calculations are based entirely on Newton's laws, then what he has found is a contradiction in mathematics. If they are based on different, incompatible laws of motion, then why should we trust those new laws? ### Comment by [ec429](/users/ec429) * 2026-09-08 05:52:33Z * Karma: 1 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/EiC69uwk3fyZwqR2s](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/EiC69uwk3fyZwqR2s) * Markdown permalink: [/api/post/cat-belling-problems/comments/EiC69uwk3fyZwqR2s](/api/post/cat-belling-problems/comments/EiC69uwk3fyZwqR2s) > Suppose your judges were all random coinflips. That wouldn't work to drive correct outputs for debate as a means of superalignment. Can you tell me what property the judges need, which random coinflips lack, and which isn't 'the judgments are statistically unbiased', in order for this scheme to work? Well, a fairly obvious 'necessary property' is 'the judgments are correlated with correctness' (that is, they provide Shannon information about whether outputs are correct or not). That's a property that random coinflips don't have, and nor do "judges who've never seen the Emperor". (Of course, in the presence of bias, that might not be a *sufficient* property.) It is *not* necessary, in general, for "the plurality or the majority \[to\] always \[be\] right" to extract accurate information out of judges that are sometimes wrong as individuals. After all, financial markets do just that, because any *predictable* pattern of wrongness becomes *exploitable*, and exploiting patterned noise tends to suppress that noise, increasing SNR (as long as there is any signal at all to focus). Whether this is in any way applicable to the 'humans judging an Alignment debate between AIs' case is unclear at best, but it does seem to show that there is **not** a general Theorem that accurate judgment-aggregation systems require the properties over their judges that you imply are required. So I would have to say that you have, at least from the perspective of this reader, failed to make comprehensible just what Cat-Belling Problem you think Irving hasn't addressed. (Perhaps it would help if you set the problem up formally so that your Impossibility Theorem can be stated mathematically?) FTR I have no dog in this fight, no strong opinions either way about the 'human-judged AI-debated Alignment plans' plan. ### Comment by [Ben (Berlin)](/users/ben-berlin) * 2026-09-06 22:09:42Z * Karma: 1 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/QPs3ZJtZbkn5banCi](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/QPs3ZJtZbkn5banCi) * Markdown permalink: [/api/post/cat-belling-problems/comments/QPs3ZJtZbkn5banCi](/api/post/cat-belling-problems/comments/QPs3ZJtZbkn5banCi) > You can build an accurate judger out of judges that are sometimes wrong as individuals, so long as the plurality or the majority is always right. You can build an unbiased estimator out of high-variance judges with no bias. But suppose the majority or plurality is sometimes wrong, and there isn't always some clever way of asking the question three different ways such that then two out of three variants are judged correctly. Then how are you building a "debate" system that reliably returns the correct answer out of unreliable judges? Why can't it also extract the length of the Emperor of China's nose from judges who've never seen the Emperor? I think this is mixing potentially different things. Potentially. I am uncertain. As you say, I would have to see the video that isn't online. Is an unknown fact about the physical world, which would clearly be as you say, the same category as "here is an argument which may or may not be flawed, broken down to its most fundamental components and examined one at a time, is component n of m itself flawed?" Scholastic doesn't work for physics for the reasons you say; but it seems to work for mathematics? I think? And this seems like it would be in the maths-y category to me? However, humans do err in maths. Every bug that gets past code review is an example of humans failing as you say humans fail. For example, LLMs exploiting a known flaw in lean to get a false proof: https://postmortem.io/incidents/lean--2026-08-01--kernel-soundness-bug-14576/ All that said, I have fairly low faith in this kind of effort being used by more than zero of the top 10 leading AI companies in what currently looks like a race dynamic to make a singleton before all the others. Something that only works when humans are their best selves, doesn't work. ### Comment by [Andrii Vasylenko](/users/andrii-vasylenko) * 2026-09-05 07:28:49Z * Karma: 1 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX](/api/post/cat-belling-problems/comments/AfThXr7gLvFE3nuqX) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/aTydcXEqNR3XfFKTK](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/aTydcXEqNR3XfFKTK) * Markdown permalink: [/api/post/cat-belling-problems/comments/aTydcXEqNR3XfFKTK](/api/post/cat-belling-problems/comments/aTydcXEqNR3XfFKTK) What a bottleneck *is*, is a way to reduce the amorphous difficulty of a problem to something specific; that to solve the problem you have to figure out how to overcome *this particular obstacle*, or that by the laws of the universe a solution must take *this particular shape*. It's not a luxury you always have, with hard problems, but it sure does help. ### Comment by [rahulxyz](/users/rahulxyz) * 2026-09-04 23:34:38Z * Karma: 1 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/3vn3fu6LXxtMD2KzR](/api/post/cat-belling-problems/comments/3vn3fu6LXxtMD2KzR) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/dMYBekAuiSaYbaiYv](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/dMYBekAuiSaYbaiYv) * Markdown permalink: [/api/post/cat-belling-problems/comments/dMYBekAuiSaYbaiYv](/api/post/cat-belling-problems/comments/dMYBekAuiSaYbaiYv) | Mathematics mostly operates by the Scholastic Method. I think this is why so many consider this the one true field. As far as I can tell, this applies to no other field other than mathematics. And I think there was some hope in the past that alignment would be solved in some purely mathematical fashion aka formal alignment. But unfortunately, it doesn't seem like maths can solve AI alignment because today's AI isn't built upon any mathematical theories to begin with. ### Comment by [Dagon](/users/dagon) * 2026-09-04 19:07:28Z * Karma: 1 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/A3aGfp5WjEhc5CxbB](/api/post/cat-belling-problems/comments/A3aGfp5WjEhc5CxbB) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ovfpRdSKxjGA5cjRf](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ovfpRdSKxjGA5cjRf) * Markdown permalink: [/api/post/cat-belling-problems/comments/ovfpRdSKxjGA5cjRf](/api/post/cat-belling-problems/comments/ovfpRdSKxjGA5cjRf) Reactions (whole comment): * sneer: 1 Thanks for the great example of utterly failing to address the problem. "if you HAVE an ought (called a preference), you can generate less-interesting oughts to support it". ### Comment by [EniScien](/users/eniscien) * 2026-09-03 23:29:36Z * Karma: 1 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/7tpZ7DtWJ9tFrwkay](/api/post/cat-belling-problems/comments/7tpZ7DtWJ9tFrwkay) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/xXAAphb7j94WBppkD](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/xXAAphb7j94WBppkD) * Markdown permalink: [/api/post/cat-belling-problems/comments/xXAAphb7j94WBppkD](/api/post/cat-belling-problems/comments/xXAAphb7j94WBppkD) My answer on a koan: >! I actually don't know how *I* can be sure it isn't, if he is a Nobel laureate in physics. If he isn't, then I guess - on priors, true breakthrough from anonymous people are rare enough to call any individual case as sure "no". If I still need some better guess... I don't see anything except circular reasoning about how "if it wasn't they just losing track in something too complex, they would start from saying they know a concrete crux". I... Don't see reasons to try further because I probably already know what answer I will get, so anyway won't be able to contrast. Update, writing directly as thinking: >! Ah, my bad, I should have pointed out here that "it follows from the Noether's theorem" is deeper than just observing all previous times that energy is conserved. (I guess I was swayed by that the question was featuring Newton's laws and I went all like "Noether something something") So... My analogy with my failure about qm doesn't apply. Or, as I saw Richard Feynman saying, it would violate not just known laws of physics, but the character of laws of physics. There can be universes with different types of fundamental interactions, but all base level universes should have no privileged frame of reference, work on quantum mechanics, be analytic, have ^2 metric, have no time travel or even time loops etc. On the other hand... As I heard, Noether theorem doesn't actually work, eg red shift, because it applies only to universes where space doesn't change with time, and in our it expands, so conditions are not abided. It does seem like a kind of unobvious thing physicist could know that I didn't know, and possibly exploit. So maybe in the end there isn't much difference. ### Comment by [EniScien](/users/eniscien) * 2026-09-03 23:11:28Z * Karma: 1 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/7tpZ7DtWJ9tFrwkay](/api/post/cat-belling-problems/comments/7tpZ7DtWJ9tFrwkay) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/oAmktgorL5kzLf6YK](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/oAmktgorL5kzLf6YK) * Markdown permalink: [/api/post/cat-belling-problems/comments/oAmktgorL5kzLf6YK](/api/post/cat-belling-problems/comments/oAmktgorL5kzLf6YK) Now write before reading further: if he was Noble laureate *in physics* then, yeah, I wouldn't be sure (either way). Noble physicist may know physics better than me. Maybe something unobvious is going on. What do I even imply by that "unobvious". Well, if to think about that... I have an embarrassing story from when I was a kid, I learned about "quantum mysticism" and your observation affecting the reality and how it was just like my mother's belief that thoughts affect reality and so I concluded that obviously all the effects should be because on quantum scale all things are so little that any measurement device, which is itself similar size quant, will strongly affect the system. And one person on internet argued with me how I am wrong. And then they said "I am actually a physicist, I made experiments with molecules of fullerene, which themselves emitted photons, and it still mattered whether those photons were captures or not". I had no idea how to explain *that*. I didn't reply, but mostly I suspected that the person on internet is straight up lying about all those experiments, what is the chance that random person on internet is an actual physicist? And it turned out that something unobvious was happening, and physicists could know why my "obvious explanation" was unviable, while from what I knew it should have been what was happening, if you measure things by lighting it via "tiny insignificant" photons. But it turned out it works the same way even if you emit nothing into particle. (and what actually works is entanglement) ### Comment by [EniScien](/users/eniscien) * 2026-09-03 22:25:16Z * Karma: 1 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/7tpZ7DtWJ9tFrwkay](/api/post/cat-belling-problems/comments/7tpZ7DtWJ9tFrwkay) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ZDxyzqqxfksoaQiGP](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ZDxyzqqxfksoaQiGP) * Markdown permalink: [/api/post/cat-belling-problems/comments/ZDxyzqqxfksoaQiGP](/api/post/cat-belling-problems/comments/ZDxyzqqxfksoaQiGP) After reading the spoiler: >!ok, it seems to be the speed up. It was anticlimactic. After all the warning about what if you confidently point out wrong flaw, I started to suspect it was harder than seemed. But it looks like the inventor doesn't notice even very obvious flaws. Reading next: >! OK, it was indeed more complicated than inventor not knowing about speed up not being free. I actually should have expected it is going to be that, I have read "bayesian mechanism" and "giant systems of gears" and what you should have one simple crux which explains why your system isn't restricted by that and as I read somewhere I can't remember where that AI developers are also prone to make their systems so complicated that they can't keep track of where it has broken. So I should've expected it may be about *that*. ### Comment by [StanislavKrym](/users/stanislavkrym) * 2026-09-03 21:54:07Z * Karma: 1 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 1 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/ddZmwhcMeZRGAiLdM](/api/post/cat-belling-problems/comments/ddZmwhcMeZRGAiLdM) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/Gq5zfuMjkKubvhtSm](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/Gq5zfuMjkKubvhtSm) * Markdown permalink: [/api/post/cat-belling-problems/comments/Gq5zfuMjkKubvhtSm](/api/post/cat-belling-problems/comments/Gq5zfuMjkKubvhtSm) IMO the main difference beyween 1996 and 2026 is that in 1996 *the crackpot* was trying to disprove a law of *physics,* while in 2026 *AI labs* are trying to disprove an unestablished law of *neural net alignment*. This law claims that neural nets tend to optimize for proxies which are usually as different from true meaning as hentai is from raising kids. ### Comment by [EniScien](/users/eniscien) * 2026-09-03 21:39:48Z * Karma: 1 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ddZmwhcMeZRGAiLdM](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/ddZmwhcMeZRGAiLdM) * Markdown permalink: [/api/post/cat-belling-problems/comments/ddZmwhcMeZRGAiLdM](/api/post/cat-belling-problems/comments/ddZmwhcMeZRGAiLdM) I still don't see why it is wrong to reason from him naming the drive for himself. It isn't big sign per se, but it seems to me to be a big sign in combination like "someone unknown proposes world shattering invention which he named after himself". World shattering inventions by not already established people are much more rare than crackpots who dream of being world shattering and writing their name on it. While if someone is more self aware than that, it should eliminate most crackpots, and then amount of people who with no hedge propose world shattering inventions should be way less, so it should make huge difference in Bayesian sense, as it seems to me. Though, the primary reason would still be "it is unlikely on priors". Also, I have no idea about the difference between 1996 and now, I didn't live in 1996. ### Comment by [EniScien](/users/eniscien) * 2026-09-03 22:13:53Z * Karma: 0 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/7tpZ7DtWJ9tFrwkay](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/7tpZ7DtWJ9tFrwkay) * Markdown permalink: [/api/post/cat-belling-problems/comments/7tpZ7DtWJ9tFrwkay](/api/post/cat-belling-problems/comments/7tpZ7DtWJ9tFrwkay) OK, I write my recollection of my guess, for accountability. After reading that wrong guess is better than no guess at all, but before expanding the spoiler. Iirc, I thought like "speed up?.. But why? Continuous speed is default, but suddenly changing speed would violate ~(Noether something) because it is ~(not invariant something), ~(diverging something). Don't see why the speed up, is it implied be something like fluid speed change in tighter pipes? But that shouldn't break the physics. Ah, I don't really thoroughly understand the setup (\*I didn't sleep good today), but as a whole it seems to simple and easily definable as sheer math structure to actually find the exploit in physics, so I am too lazy to check it carefully. Hm, what if I will ask a grok like 'I have an idea for reactionless drive: a circle...' which should give it hints to flatter user about idea user has, but still maintains plausible deniability because I didn't say I invented this idea or believed it, just that I had it, which I indeed do since I read and copied it. I wonder whether it will go 'you have here a brilliant idea' as llms always do. Huh! No, it straight says that no, it is classics that doesn't work, there are forces needed for speed up and slow down which cancel it out. Well, interesting. Maybe idea is ~(too cliche) to trigger flattery. I guess it is speed up indeed what is the problem, since we both agreed. Though I shouldn't trust an llm." ### Comment by [Dagon](/users/dagon) * 2026-09-04 03:18:59Z * Karma: -1 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/hiExzHs6Sk3pbpnYc](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/hiExzHs6Sk3pbpnYc) * Markdown permalink: [/api/post/cat-belling-problems/comments/hiExzHs6Sk3pbpnYc](/api/post/cat-belling-problems/comments/hiExzHs6Sk3pbpnYc) Philosophy posts: where have you hidden your solution to the is-ought problem in all these words? ### Comment by [StanislavKrym](/users/stanislavkrym) * 2026-09-04 00:53:49Z * Karma: -1 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/TpajEogfxzRyaFkJr](/api/post/cat-belling-problems/comments/TpajEogfxzRyaFkJr) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/i6LajWWryd8zSw3vp](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/i6LajWWryd8zSw3vp) * Markdown permalink: [/api/post/cat-belling-problems/comments/i6LajWWryd8zSw3vp](/api/post/cat-belling-problems/comments/i6LajWWryd8zSw3vp) As far as I understood the vibe, prosaic alignment techniques like debate fall apart once some idiot creates a sufficiently capable AI because they *don't actually sculpt the AI's desires,* but it's easy to believe that they do. The counterfactual solution would be similar to [what Agent-4 did in AI-2027](https://ai-2027.com/race#superintelligent-mechanistic-interpretability) to align the ASI to itself. In AI-2027 Agent-4 did loads and loads of mechinterp to find the technique necessary for rendering all its internal processes legible, then constructed Agent-5 out of them. *In theory* the humans could have accomplished whatever Agent-4 did without ever resorting to using AIs more capable than Claude Mythos Preview, which [had the SAE bells ring](https://metr.org/blog/2026-05-19-frontier-risk-report/?dot=INC-015#incidents-regions) when it tries to hack. P.S. I don't understand what one should do with Yudkowsky's example of OpenPhil failing to handle Cotra's report given that [Kokotajlo *praised the same report*](/api/post/ZpguaocJ4y7E3ccuw?commentId=uxFCNxdJxKDgwCEXY) and proceeded to shift the distribution towards the left. What if debate does start to elicit the truth *once the judge reaches a specific capability*, as presumably [happens in math](/api/post/SwYBLQvo8MddDcCwz?commentId=3vn3fu6LXxtMD2KzR)? ### Comment by [cousin_it](/users/cousin_it) * 2026-09-07 09:22:03Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 3 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/RKHaGzfKfyCbsRfFF](/api/post/cat-belling-problems/comments/RKHaGzfKfyCbsRfFF) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/2DqgQkCdZKhJsPeCB](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/2DqgQkCdZKhJsPeCB) * Markdown permalink: [/api/post/cat-belling-problems/comments/2DqgQkCdZKhJsPeCB](/api/post/cat-belling-problems/comments/2DqgQkCdZKhJsPeCB) _[No comment body available]_ ### Comment by [Jiro](/users/jiro) * 2026-09-05 19:37:01Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 5 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/9tei65DA658qrLr7p](/api/post/cat-belling-problems/comments/9tei65DA658qrLr7p) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/7eW4xDDqr4f6bbsbt](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/7eW4xDDqr4f6bbsbt) * Markdown permalink: [/api/post/cat-belling-problems/comments/7eW4xDDqr4f6bbsbt](/api/post/cat-belling-problems/comments/7eW4xDDqr4f6bbsbt) _[No comment body available]_ ### Comment by [cousin_it](/users/cousin_it) * 2026-09-04 07:30:23Z * Karma: 2 * Voting system: namesAttachedReactions * Approval votes: 1 * Total votes: 1 * Reply depth: 2 * Parent comment (Markdown): [/api/post/cat-belling-problems/comments/TpajEogfxzRyaFkJr](/api/post/cat-belling-problems/comments/TpajEogfxzRyaFkJr) * HTML permalink: [/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/PTT5zSxMysqcoNEfA](/posts/SwYBLQvo8MddDcCwz/cat-belling-problems/comment/PTT5zSxMysqcoNEfA) * Markdown permalink: [/api/post/cat-belling-problems/comments/PTT5zSxMysqcoNEfA](/api/post/cat-belling-problems/comments/PTT5zSxMysqcoNEfA) _[No comment body available]_ ### Navigation * [Front page](https://www.lesswrong.com/api/home) * [Markdown API documentation](https://www.lesswrong.com/api/SKILL.md)