I've been writing a comprehensive Lakatos-style dialogue on Yudkowsky's Bayesianism. There have been rather a lot of criticisms of Yudkowsky's Bayesianism over the years, both directly related to him and to the general subjective Bayes as optimal reasoner programme; I get the sense that, for the people on this website, the kind of overarching, technical takedown on all the specifics is not quite useful, but instead something focused on the parts of his Bayesianism that have actually been passed down.
I would like to ask you to tell me what those are. I would summarize Yud's take as, roughly,
0. Think laws, not tools. Bayes is not one instrument among others, to be picked up when convenient and put down when awkward; it is the law that governs whatever you pick up. Everything below follows from taking that seriously.
1. The ideal reasoner is Bayesian.
1.1 What is meant by 'ideal' is that no other procedure can 'do better'; a la Carnot's engine. The comparison is his: "Bayes' rule is to reasoning as the Carnot cycle is to engines: Nobody can be a perfect Bayesian, but Bayesian reasoning is still the theoretical ideal."
1.2 What is meant by 'do better' is that "you can't get a higher expected score by making any different update", with the expectation being taken over the prior, and the score being log-loss.
1.21 Equivalently: "You can't extract any more evidence from an observation than is given by its likelihood ratio."
1.22 This is not a "normative fact", per se; the idea "begins life as a descriptive assertion, not as a normative assertion". But: "If you want to assign higher probability to the correct hypothesis, it's a short step from that preference to regarding Bayesian updates as a normative ideal". (The reading this invites, and which is what actually gets passed down: if you reason by something other than a Bayes update, you are assigning lower probability to the correct hypothesis. Note that this does not follow from 1.2 - an expected-score claim under your own prior is not an accuracy claim about the truth - but it is what the passage is for.)
1.23 He is aware of the distinction between proper and strictly proper scoring rules, and treats the logarithmic score as the working case.
1.3 We may also mean that breaking with the ideal makes you vulnerable to a sure loss by Dutch booking: "anything that is not Bayesian must fail one of the coherency tests."
1.4 A third meaning is by the complete class theorem's converse; an admissible decision-rule is Bayes or generalized-Bayes, so ideal reasoners are Bayesian.
1.5 A fourth is uniqueness: "Given what you knew, and what you saw, the maximally accurate state of belief for you to be in is completely pinned down." Hard to find in practice, known in principle. There is only ever one answer.
1.6 And so, definitionally: "Bayesian reasoner" is "the technically precise codeword that we use to mean rational mind."
1.7 This ideal is fully general in the sense that it, at the very least, must bound why Einstein discovered relativity (so, at the very least, general in the sense of function-spaces, logical uncertainty, old-evidence). It explains all reasoning.
1.71 Any theoretical gap in the current Bayesian formalism is to be explained by further Bayesianism, not by a retreat from Bayesianism (law-likeness).
2. It is more precisely a subjectivist Bayesianism qua Jaynes.
2.1 Uncertainty is 'in the map, not in the territory'.
2.11 Priors ought to be coherent probability distributions, not weighing functions. A prior that is data-dependent or which does not normalize to a probability distribution is non-Cox and is thus Dutch-bookable.
2.12 The prior is a summary of 'all known prior information', not properties of the posterior.
2.2 A single case has no probability at all except somebody's credence in it. Coverage of posterior intervals is irrelevant; frequency concerns with repeated sampling are nonsense.
2.21 The frequentist's objective/subjective distinction is confused rather than merely austere: the 'objective' long-run statement is itself obtained as a limit over the probabilities of finite sequences, every intermediate one of which he calls meaningless.
2.22 The positive replacement: a single-case probability is cashed out by a logloss correspondence against what is subsequently observed.
3. Distance from the Bayesian is distance from the ideal.
3.1 'Whatever approximation you use, both its failures and its successes are explainable in Bayesian terms'.
3.11 Frequentist/statistical tools work to the extent that they track the ideal Bayesian calculation, and fail to the extent that they depart.
3.111 If a Frequentist tool looks like it outperforms the Bayesian law, extra, subjective information has been smuggled in. (Licensed by 1.5: since the maximally accurate state is unique, an apparent improvement on it must be an illusion or a theft.)
3.12 This does not imply Bayesian 'tools' are more useful directly; but, the Bayesian way 'suggests the path' to the ideal - knowledge of the law "helps you get as close to the ideal efficiency as you can."
3.2 As such, it has sense to 'look for Bayes-structure'.
3.21 Bayes-structure is nontrivial, generically; the Bayesian version of the thing that works is 'why' it works.
3.22 Thus, if something has trivial Bayes-structure, in that the prior corresponds to no prior or the approximation taken is one that makes no Bayesian sense to take, it will malfunction in proportion to the departure.
4. We may justify 'rational' behaviour for human behaviour and scientific inference on the basis that it is Bayesian.
4.1 Stopping-rules must be made irrelevant - two researchers with the same data must give the same conclusions, regardless of the process to get the data.
4.2 Randomness has no power - "there is no beauty in entropy, nor strength from noise". An algorithm is improved by randomization only where some step was already doing worse than chance.
4.21 Wherever the environment cares only about your actions and not your algorithm, anything improvable by randomization is further improvable by derandomization.
4.22 One exception is granted: superintelligent or cryptographic adversaries, where entropy acts as an antidote to intelligence. Otherwise, without adversariality, no.
4.3 All the knowledge contained in the data is contained in the likelihood of the data; procedures that depend on anything else are to be disrecommended.
4.31 Frequentist error rates are irrelevant; confidence intervals are meaningless, since repeated sampling is incoherent. The reference class is fixed by the experimenter's intentions and is therefore not a fact about the world.
4.32 Inference must not change on the basis of likelihood-preserving experiment design; all conclusions depend on the probabilistic element alone, not on extraneous elements.
4.4 An improvement would be to report full likelihood functions, rather than p-values, tests, estimates or intervals.
4.5 The sinister misdeeds p-values are meant to prevent "are just flatly mathematically impossible in the first place under this system."
4.51 The ground given is conservation of expected evidence: you would have to know in advance which direction you would update, which is impossible. "If you start out thinking it's 70% probable that some coin is fair, nothing you can possibly plan to do by gathering more data... can result in you expecting for that analysis to make you believe on average that the coin is not 70% probably fair."
4.6 Frequentist/statistical theory is thus only relevant because of computational-approximate concerns; pragmatic, not epistemic or quite scientific ones. "For so long as we do not have infinite computing power, there may yet be a place in science for non-Bayesian statistics."
5. Apparent counterexamples are faults in the counterexample, rather than faults in the methodology. (Pre-emptive monster-barring/adjustment.)
5.0 The general policy: shown a purported paradox, "look for the division by zero; or the infinity that is assumed rather than being constructed as the limit of a finite operation" - something illegal. "Trust Bayes. Bayes has earned it."
5.1 Infinite sets lead to trouble, so only priors on finite sets approximating the infinite set are coherent.
5.11 The pathologies of infinite-set Bayesianism are illusions, as the finite-set Bayesian is fine.
5.2 Simultaneously, the Bayesian must place no zero- or one- probabilities anywhere.
5.21 In particular, insofar as Einstein was bounded by the ideal Bayesian, the ideal Bayesian must place a nonzero probability on the equations of GR.
5.22 Technically speaking, zero and one are not probabilities, per se. The fact that the probability formalism has them is a kind of gap; it is sensible to imagine a theory that is rid of them that is better for it.
5.23 Consequently [this is an entailment I am drawing, not a claim he states]: the Bayesian contains the correct model of the world, and concerns of 'model misspecification' are pragmatic, of no concern to the ideal epistemologist.
5.3 In particular, Solomonoff solves all such troubles.
5.31 Indeed, the Solomonoff prior gives a unified proof of Ockham's principle in some general way.
5.32 Despite UTM-relativity, we know roughly what 'good' UTMs look like, and anyhow the constant washes out with enough data, which is fine for our ideal. Abusing the freedom requires constructing "a downright embarrassing Universal Turing Machine" - though he concedes that fully objective priors are not to be had by deduction, not "without principles that are unknown to me and beyond the scope of Solomonoff induction."
5.33 It is "something that bootstraps to good epistemology rather than being all of good epistemology by itself."
Which of these do you actually hold? Any places where your reading of him differs from mine? There are ~20 years of stuff in here, after all.
Gelman, Bayesian Workflow:
A sometimes annoying habit of Bayesians is to take non-Bayesian methods and give them Bayesian interpretations. For example, “maximum likelihood is just Bayesian inference with a flat prior,” “fixed effects are just random effects with the group-level variance set to infinity,” “lasso is just regression with an exponential prior,” or “confidence intervals are just posterior intervals assuming flat priors.”
This practice can be useful in giving insight into the scenarios in which a statistical procedure will apply, but it can also be misleading. An example of the latter is the interpretation of lasso (Tibshirani 1996) mentioned just above: lasso is just regression with an exponential prior. Lasso is an approach for obtaining more stable regression estimates that pulls coefficient estimates toward or all the way to zero. The lasso estimate can be viewed as an approximate posterior mode assuming independent exponential priors on the coefficients, and we see some value in making this connection—but this does not make lasso a Bayesian method. The problem is that lasso is intended for use in problems of moderate and high dimensions. In such problems, the mode is not necessarily or even usually a good summary of the posterior distribution, because of the phenomenon of concentration of measure. As such, in high dimensional problems, the formulation, “a penalty is just a log prior density,” stops being useful or accurate. This does not mean that lasso is a bad idea, or that a Bayesian version is necessarily better; some of lasso’s desirable computational and applied properties directly arise from it being a mode rather than a full posterior.
More generally, any statistical method can be evaluated on its own, in non-Bayesian terms, by studying its statistical properties under various assumptions. Yes, least squares regression can be derived as the maximum likelihood estimate assuming independent normally-distributed errors with equal variance. But you can use least squares without making those assumptions, and you can evaluate the robustness of least squares under various alternative assumptions. That said, the theoretical derivation of least squares can provide some insight and practical guidance. For example, if you have reason to believe that the model errors have some particular covariance structure, you could consider incorporating that into an improved estimate. Conversely, if least squares estimation does not seem to be performing well, you could go back and check its implicit assumptions.
Still, careful Bayesian interpretations offer potential benefits. In applied settings where estimates are problematically noisy, leading to replication failures (Button et al. 2013), we can get some perspective by recognizing that simple means, differences, and least squares regression correspond to Bayesian estimates assuming flat priors. When the posterior distribution includes implausible parameter values, it can make sense to add further information. This can take the shape of an informative prior. Another way to say this: When the estimate is too large to believe, this implies that inference can be improved by incorporating prior information.
Yudkowsky, Searching for Bayes-Structure:
In fact, any part of a cognitive process that contributes usefully to truth-finding must have at least a little Bayesian structure - must harmonize with Bayes, at some point or another - must partially conform with the Bayesian flow, however noisily - despite however many disguising bells and whistles - even if this Bayesian structure is only apparent in the context of surrounding processes. Or it couldn't even help.
(...)
But perhaps it is not quite as exciting to see something that doesn't look Bayesian on the surface, revealed as Bayes wearing a clever disguise, if: (a) you don't unravel the mystery yourself, but read about someone else doing it (Newton had more fun than most students taking calculus), and (b) you don't realize that searching for the hidden Bayes-structure is this huge, difficult, omnipresent quest, like searching for the Holy Grail.
It's a different quest for each facet of cognition, but the Grail always turns out to be the same. It has to be the right Grail, though - and the entire Grail, without any parts missing - and so each time you have to go on the quest looking for a full answer whatever form it may take, rather than trying to artificially construct vaguely hand-waving Grailish arguments. Then you always find the same Holy Grail at the end.
(...)
This left me in a bit of a pickle when it came to trying to explain in advance where I was going. I know from experience that if I say, "Bayes is the secret of the universe," some people may say "Yes! Bayes is the secret of the universe!"; and others will snort and say, "How narrow-minded you are; look at all these other ad-hoc but amazingly useful methods, like regularized linear regression, that I have in my toolbox."
(...)
To see through the surface adhockery of a cognitive process, to the Bayesian structure underneath - to perceive the probability flows, and know how, not just know that, this cognition too is Bayesian - as it always is - as it always must be - to be able to sense the Force underlying all cognition - this, is the Bayes-Sight.
So, to Gelman, the "Bayes-structure" is only sometimes helpful, but sometimes misleading as to why the thing works and in what ways it does, but to Yudkowsky, it is the Holy Grail and the only reason why these methods work is because of the Bayesian explanation underneath (Neatly, they both mention the same example as demonstrating the opposite things).
In my experience as a statistician, I think Gelman has the right of it (it is certainly the case with LASSO; I think Castillo is essentially correct that the LASSO posterior is a "useless object" that does away with why people use it), but I wonder about some more orthodox Bayesian perspectives on this contrast?