I've been writing a comprehensive Lakatos-style dialogue on Yudkowsky's Bayesianism. There have been rather a lot of criticisms of Yudkowsky's Bayesianism over the years, both directly related to him and to the general subjective Bayes as optimal reasoner programme; I get the sense that, for the people on this website, the kind of overarching, technical takedown on all the specifics is not quite useful, but instead something focused on the parts of his Bayesianism that have actually been passed down.
I would like to ask you to tell me what those are. I would summarize Yud's take as, roughly,
0. Think laws, not tools. Bayes is not one instrument among others, to be picked up when convenient and put down when awkward; it is the law that governs whatever you pick up. Everything below follows from taking that seriously.
1. The ideal reasoner is Bayesian.
1.1 What is meant by 'ideal' is that no other procedure can 'do better'; a la Carnot's engine. The comparison is his: "Bayes' rule is to reasoning as the Carnot cycle is to engines: Nobody can be a perfect Bayesian, but Bayesian reasoning is still the theoretical ideal."
1.2 What is meant by 'do better' is that "you can't get a higher expected score by making any different update", with the expectation being taken over the prior, and the score being log-loss.
1.21 Equivalently: "You can't extract any more evidence from an observation than is given by its likelihood ratio."
1.22 This is not a "normative fact", per se; the idea "begins life as a descriptive assertion, not as a normative assertion". But: "If you want to assign higher probability to the correct hypothesis, it's a short step from that preference to regarding Bayesian updates as a normative ideal". (The reading this invites, and which is what actually gets passed down: if you reason by something other than a Bayes update, you are assigning lower probability to the correct hypothesis. Note that this does not follow from 1.2 - an expected-score claim under your own prior is not an accuracy claim about the truth - but it is what the passage is for.)
1.23 He is aware of the distinction between proper and strictly proper scoring rules, and treats the logarithmic score as the working case.
1.3 We may also mean that breaking with the ideal makes you vulnerable to a sure loss by Dutch booking: "anything that is not Bayesian must fail one of the coherency tests."
1.4 A third meaning is by the complete class theorem's converse; an admissible decision-rule is Bayes or generalized-Bayes, so ideal reasoners are Bayesian.
1.5 A fourth is uniqueness: "Given what you knew, and what you saw, the maximally accurate state of belief for you to be in is completely pinned down." Hard to find in practice, known in principle. There is only ever one answer.
1.6 And so, definitionally: "Bayesian reasoner" is "the technically precise codeword that we use to mean rational mind."
1.7 This ideal is fully general in the sense that it, at the very least, must bound why Einstein discovered relativity (so, at the very least, general in the sense of function-spaces, logical uncertainty, old-evidence). It explains all reasoning.
1.71 Any theoretical gap in the current Bayesian formalism is to be explained by further Bayesianism, not by a retreat from Bayesianism (law-likeness).
2. It is more precisely a subjectivist Bayesianism qua Jaynes.
2.1 Uncertainty is 'in the map, not in the territory'.
2.11 Priors ought to be coherent probability distributions, not weighing functions. A prior that is data-dependent or which does not normalize to a probability distribution is non-Cox and is thus Dutch-bookable.
2.12 The prior is a summary of 'all known prior information', not properties of the posterior.
2.2 A single case has no probability at all except somebody's credence in it. Coverage of posterior intervals is irrelevant; frequency concerns with repeated sampling are nonsense.
2.21 The frequentist's objective/subjective distinction is confused rather than merely austere: the 'objective' long-run statement is itself obtained as a limit over the probabilities of finite sequences, every intermediate one of which he calls meaningless.
2.22 The positive replacement: a single-case probability is cashed out by a logloss correspondence against what is subsequently observed.
3. Distance from the Bayesian is distance from the ideal.
3.1 'Whatever approximation you use, both its failures and its successes are explainable in Bayesian terms'.
3.11 Frequentist/statistical tools work to the extent that they track the ideal Bayesian calculation, and fail to the extent that they depart.
3.111 If a Frequentist tool looks like it outperforms the Bayesian law, extra, subjective information has been smuggled in. (Licensed by 1.5: since the maximally accurate state is unique, an apparent improvement on it must be an illusion or a theft.)
3.12 This does not imply Bayesian 'tools' are more useful directly; but, the Bayesian way 'suggests the path' to the ideal - knowledge of the law "helps you get as close to the ideal efficiency as you can."
3.2 As such, it has sense to 'look for Bayes-structure'.
3.21 Bayes-structure is nontrivial, generically; the Bayesian version of the thing that works is 'why' it works.
3.22 Thus, if something has trivial Bayes-structure, in that the prior corresponds to no prior or the approximation taken is one that makes no Bayesian sense to take, it will malfunction in proportion to the departure.
4. We may justify 'rational' behaviour for human behaviour and scientific inference on the basis that it is Bayesian.
4.1 Stopping-rules must be made irrelevant - two researchers with the same data must give the same conclusions, regardless of the process to get the data.
4.2 Randomness has no power - "there is no beauty in entropy, nor strength from noise". An algorithm is improved by randomization only where some step was already doing worse than chance.
4.21 Wherever the environment cares only about your actions and not your algorithm, anything improvable by randomization is further improvable by derandomization.
4.22 One exception is granted: superintelligent or cryptographic adversaries, where entropy acts as an antidote to intelligence. Otherwise, without adversariality, no.
4.3 All the knowledge contained in the data is contained in the likelihood of the data; procedures that depend on anything else are to be disrecommended.
4.31 Frequentist error rates are irrelevant; confidence intervals are meaningless, since repeated sampling is incoherent. The reference class is fixed by the experimenter's intentions and is therefore not a fact about the world.
4.32 Inference must not change on the basis of likelihood-preserving experiment design; all conclusions depend on the probabilistic element alone, not on extraneous elements.
4.4 An improvement would be to report full likelihood functions, rather than p-values, tests, estimates or intervals.
4.5 The sinister misdeeds p-values are meant to prevent "are just flatly mathematically impossible in the first place under this system."
4.51 The ground given is conservation of expected evidence: you would have to know in advance which direction you would update, which is impossible. "If you start out thinking it's 70% probable that some coin is fair, nothing you can possibly plan to do by gathering more data... can result in you expecting for that analysis to make you believe on average that the coin is not 70% probably fair."
4.6 Frequentist/statistical theory is thus only relevant because of computational-approximate concerns; pragmatic, not epistemic or quite scientific ones. "For so long as we do not have infinite computing power, there may yet be a place in science for non-Bayesian statistics."
5. Apparent counterexamples are faults in the counterexample, rather than faults in the methodology. (Pre-emptive monster-barring/adjustment.)
5.0 The general policy: shown a purported paradox, "look for the division by zero; or the infinity that is assumed rather than being constructed as the limit of a finite operation" - something illegal. "Trust Bayes. Bayes has earned it."
5.1 Infinite sets lead to trouble, so only priors on finite sets approximating the infinite set are coherent.
5.11 The pathologies of infinite-set Bayesianism are illusions, as the finite-set Bayesian is fine.
5.2 Simultaneously, the Bayesian must place no zero- or one- probabilities anywhere.
5.21 In particular, insofar as Einstein was bounded by the ideal Bayesian, the ideal Bayesian must place a nonzero probability on the equations of GR.
5.22 Technically speaking, zero and one are not probabilities, per se. The fact that the probability formalism has them is a kind of gap; it is sensible to imagine a theory that is rid of them that is better for it.
5.23 Consequently [this is an entailment I am drawing, not a claim he states]: the Bayesian contains the correct model of the world, and concerns of 'model misspecification' are pragmatic, of no concern to the ideal epistemologist.
5.3 In particular, Solomonoff solves all such troubles.
5.31 Indeed, the Solomonoff prior gives a unified proof of Ockham's principle in some general way.
5.32 Despite UTM-relativity, we know roughly what 'good' UTMs look like, and anyhow the constant washes out with enough data, which is fine for our ideal. Abusing the freedom requires constructing "a downright embarrassing Universal Turing Machine" - though he concedes that fully objective priors are not to be had by deduction, not "without principles that are unknown to me and beyond the scope of Solomonoff induction."
5.33 It is "something that bootstraps to good epistemology rather than being all of good epistemology by itself."
Which of these do you actually hold? Any places where your reading of him differs from mine? There are ~20 years of stuff in here, after all.