Hmm. I'm sorry to have to say this, but I think this post is not very good.
In summary, I would say that this article is discussing a set of very common, well-trodden issues (Bayesian model checking/expansion/selection, generally speaking problems of insufficient 'model space') in a very nontechnical way (which is of course a fine approach on its own, but it's strange that the thing that this is talking about is not referenced or referred to as the thing it's called anywhere, instead everything here is presented as if it was a novel first-principles contribution), in such a way that proposes the most wishy-washy solution to the problem (just ignoring it) without reference to any alternatives and without really discussing the disadvantages of this approach.
To go line by line: first, though the title and the framing of this article surround priors, as you kind of point out, the prior isn't really what's changing here, but the likelihood (and the part of the prior that interacts with it). You can consider yourself to have, in theory, a larger prior that you just weren't aware you should be considering. So, the problem that this is dealing with is the classic "I don't actually believe my likelihood is correct" model, for which there are a billion proposed solutions, which you don't really touch on for some reason (to name a few, "M-open" analyses, everything in the Gelman BDA chapter on model expansion, nonparametric methods which to be fair you hint at with 'infinite-dimensional', etc).
To briefly explain some of these: one classic approach is to consider your prior attempt some model , with some prior probability, and your next one an , and your final posterior will be an averaging out of each, with some prior model credences. In this specific case, you could probably nest both of the likelihoods into a single model, somehow. The most common applied approach, as many have pointed out, is to just split the data and do whatever exploratory analyses you want on the head of the table or some such approach.
The problem with this classic solving-by-not-solving approach is that, well, you are not being an actual Bayesian anymore; if your (technical) prior depends on observing the data, your procedure no longer actually follows the laws of probability, no longer benefits from Cox's theorem, has no guarantee relating to Dutch books and does not result in decision-optimal rules (roughly the same reasons orthodox Bayesians reject empirical-Bayes methods, which are a non-Bayesian half-solution to this issue, kind of). So, well, I don't see the point.
There are good ways and bad ways to be approximately Bayesian, and this particular one seems no good to me, at least without further argument (especially when avoiding it is so easy and common. Just data-split!). Double-counting methods always look good when you are in a thought experiment and are doubling down on what you know to be the right answer, the problem is that they are doubling down on the wrong answers for no reason, too; it seems perfectly reasonable to me to admit that 50/50 posterior after seeing just a little bit of data, then some model refinement happens, and your posterior upon seeing the full data looks more 'reasonable'.
So, I don't know. These are good (classic) questions, but I am forced to disagree with the notion that this is a good answer.
Thanks for the reply, sorry I just saw this. It was indeed my goal to talk about existing ideas in a nontechnical way, which is why I didn't frame things in terms of model expansion, etc.. Beyond that however, I am confused by your reply, as it seems to make little contact with my intended argument. You state that I recommend "just ignoring" the issue, and suggest that I endorse double-counting as OK. Can you explain what parts of the post led you to believe that was my recommendation? Because that is very much not my intended message!
(I stress that I'm not trying to be snarky. The goal of the post is to be a non-technical explanation, and I don't want to change that. But if the post reads as you suggest, I interpret that as a failure of the post, and I'd like to fix that.)
Thanks for replying. Given that it's been a month, sadly, I don't fully remember all the details of why I wrote what I wrote in my initial comment, but I'll try to roughly rewrite my objections in a more specific way so that you get where it makes contact with your post ("if I had more time, I would've written a shorter letter"). Forgive me if it was somehow hard to understand, English is my second language.
My first issue: the post is titled "Good if make prior after data instead of before". Yet, the post's driving example is a situation where the (marginal) prior probability of what you're interested in doesn't actually change, but instead is coupled to a larger model with a larger probability space where the likelihood is different at these different points. So, what you're talking about isn't really post-hoc changes to the prior, but something like model expansion, as you write in the comment.
In the context of methodologies for Bayesian model expansion, there is a lot of controversy and much ink has been spilled, because being ad-hoc and implicitly accepting a data-driven prior/selected model leads to incoherence; the decision-procedure you now derive from this is not actually Bayesian in the sense that it satisfies all the nice properties people expect of Bayesian decision rules and Bayesian reasoning, it just vaguely follows Bayes's rule for conditioning. When you write
So the only practical way to get good results is to first look at the data to figure out what categories are important, and then to ask yourself how likely you would have said those categories were, if you hadn’t yet seen any of the evidence.
you are sidestepping all of these issues (what I called "solving by not solving") and accepting incoherence as OK. And, well, this can be a fine approach - being approximately incoherent can be approximately no problem. But, I think that the post not only fails to address the negatives of this particular approach, positioning it as kind of the only thing you can reasonably do (which is in itself a sufficiently large problem), but fails to consider any other ones (A classic objection to this type of methodology in a canonical introductory textbook, providing one of the alternatives I mentioned, is here, for example, in which the idea is to have a model flexible and general enough that it can learn in essentially any situation; I mentioned other methods in the comment). Do you not see the incoherence of a data-driven prior as bad somehow?
To be clear, the other approach you consider of "never change your model/prior after seeing the data, even if your model makes no sense, your posterior is stuck as it is" is also bad for all the obvious model misspecification reasons. But, at the very least it is coherent (and, of course, by data-splitting you get to enjoy this coherency without being rigid at the cost of a little data, so there's another approach, much less technical to explain than the nonparametric approach mentioned prior). This is my main problem with the article, really: it proposes just this one idea among several without discussing its positives or negatives in relation to any of the other ones.
My point with this article endorsing "double-counting" is that one way in which this approach (roughly summarized as "construct the model after seeing the data, pretending like you haven't seen the data") is that, in comparison to either a nonparametric approach or some M-open idea like model mixing or stacking, it will privilege the particular model which you happened to construct on the basis of the data more so than a fully coherent theoretical approach.
An easy way to see this is to imagine if you were to try this approach while being knowledgeable about all possible models you could have picked (i. e. in model averaging, they wave at a similar critique to this idea in this other intro); in this representation, instead of observing the data and updating yourself towards one particular model representation which fits best with the data, your method is to set one model's probability to 1 and all others to zero, which is a rather extreme version of double-counting[1].
So, in my perspective, a good version of this article would not talk about anything being "the only practical way to get good results", and would situate this idea alongside all the other ones in this vein which have been discussed for decades, or at least sort of gestures at the more common approaches you consider sensible and that you think you can explain nontechnically (hopefully referenced by their names), and at the bare minimum it should explain the pros and cons of what it advocates with more balance. Admittedly, this is a much harder article to write, because the issue has become nuanced, and I would not know how to write it non-technically, at least immediately. However, the issue seems to be nuanced, at least to me, and this level of simplification misleads more than it helps.
In the original comment I decided to talk more about how easy it is to make double-counting methods seem arbitrarily good by way of constructing examples where you know the truth in advance, since of course it looks better if you get to the truth twice as fast, but the double-counting when the data happens to be misleading gets you doubly wrong too, but this objection seems kind of petty and irrelevant compared to the other ones, in hindsight.
Thanks for the response. But again, your response makes almost no contact with the content of the post. You give general comments on the problems of double-counting by using information from the data in arbitrary ways to change your prior, seem to assume I am endorsing that, and then criticize the post for not solving those problems. But... I am not suggesting using information from the prior in arbitrary ways. I suggest doing so in a specific way, under specific rules that are designed to avoid the problems of double-counting. Those rules might fail somehow, but your response demonstrates no understanding that I've suggested any rules at all.
Fair enough. When I stated "a prior that depends on the observed data is not coherent", I relied on it being true, when I should've been more specific. To my understanding, the idea you are proposing goes, roughly, like:
You have some idealized prior
Call the pre-observation model Model 1, the data-dependent refinement Model 2. Specifically, say the refinement occurs only after one has observed the realization
The Dutch Book is this. I ask you your
Say
I imagine one responds with the notion that the idea is only acceptably non-Bayesian, because we are all mortal finite beings who can only ever do with approximating Bayesian rationality, for which this is a simple, sensible and suitable procedure. My earlier posts were arguing that, no, while it is known that this can be done, it is disrecommended (Technically, it would be fine if you would refine in the same way regardless of what you observe, but this seems to run opposite of the post). Obviously you are not charged with 'solving those problems' in full generality. But, when you say
So the only practical way to get good results is to first look at the data to figure out what categories are important, and then to ask yourself how likely you would have said those categories were, if you hadn’t yet seen any of the evidence.
Then I think you are wrong; you can do any of those other things I mentioned, which apply to any such "I observed the data and think my model was wrong"-type problems while being more sensible either in Bayesian terms (e. g., not incurring in incoherence) or in pragmatic terms (not double-counting and easily tricking yourself into overconfidence; I don't know what makes this any procedure different from asking for an intuitive posterior, one would have to prove that it is not possible to bisect tendentiously until some arbitrary 'constant-ish' probability. In fact, it seems like it starts at a very flat probability, so I am confused about when one stops).
Thank you, this is much more specific!
You have some idealized prior in your head, which you cannot access directly but can feel out and compare against a working prior. You observe the data. The data tells you how finely to partition the working prior. You then fill in the cells with what you judge you would have said before observing anything. Do this until is constant-ish inside each cell.
Yes, that's roughly correct. I wouldn't endorse the last sentence, but I don't think that's a crux for your criticism.
I believe your Dutch Book argument is equivalent to the statement that the above procedure is not decision-theoretic optimal. I of course agree with that. The only decision-theoretic optimal procedure is to state your full prior (without any approximation) and make decisions based on that.
But as far as I can tell, you acknowledge the following propositions:
(Incidentally, these alternate procedures are not so different. For example, if you follow my procedure but only make the discretization more fine that is literally an example of model expansion.)
So what criticisms remain? As far as I can tell, you are arguing that the post errs in two ways:
For the first criticism, I reiterate that this post is targeting a general audience, and I have made every effort to avoid unnecessary technical terminology. There is a place in the world for posts targeting a general audience, so I can't accept this criticism.
For the second point, you have taken that quote out of context. (And added misleading bold text.) In the post, it is clearly for how to make a specific categorization for a specific example:
Technically, the fix to the first model is simple: Make P[data | aliens] lower. But the reason it’s lower is that I have additional prior information that I forgot to include in my original prior. If I just assert that P[data | aliens] is much lower than P[data | no aliens] then the whole formal Bayesian thing isn’t actually doing very much—I might as well just state that I think P[aliens | data] is low. If I want to formally justify why P[data | aliens] should be lower, that requires a messy recursive procedure where I sort of add that missing prior information and then integrate it out when computing the data likelihood.
I don’t think that technical fix is very good. While it’s technically correct (har-har) it’s very unintuitive. The better solution is what I did in the second model: To create a finer categorization of the space of things that might be true, such that the probability of the data is constant-ish for each term.
The thing is: Such a categorization depends on the data. Without seeing the actual data in our world, I would never have predicted that we would have so many pilots that report seeing tic-tacs. So I would never have predicted that I should have categories that are based on how much people might hallucinate evidence or how much aliens like to mess with us. So the only practical way to get good results is to first look at the data to figure out what categories are important, and then to ask yourself how likely you would have said those categories were, if you hadn’t yet seen any of the evidence.
Not only am I not claiming that this is the only way to deal with the general problem, I'm not even claiming that it's the only way to deal with this specific example. I am stating that the the only way to choose a better discretization for this example is to look at the data.
More broadly, the posts never claims that data-dependent discretizations are the only solution. All I've claimed is that there are often perfectly good reason to change your prior after looking at the data. And as far as I can tell, you accept that this is true. So what am I missing? How do I square this with the very broad and very sharp criticisms you made in your first comment?
A prompt reply! Wonderful. Thank you. Apologies for the bold, it would indeed have been more proper in any academic context to append "(emphasis mine)". A crass error.
I wouldn't endorse the last sentence, but I don't think that's a crux for your criticism.
An agreeable amendment to the post, since that line was copied from it (including the '-ish'):
The better solution is what I did in the second model: To create a finer categorization of the space of things that might be true, such that the probability of the data is constant-ish for each term. (emphasis mine)
Either you apply the procedure with this stopping rule for when to stop discretizing, or it boils down to "discretize until the subjective prior and likelihood lead to agreement with what you imagine the posterior intuitively ought to look like", which seems empty; one may as well state the subjective posterior without rationalizing it, all the prior is doing is adding discretization error (you indeed do state that just eliciting the posterior directly is not cogent, but if the procedure only runs on intuition, then I don't know that this is any different).
I believe your Dutch Book argument is equivalent to the statement that the above procedure is not decision-theoretic optimal.
No. Suboptimal is "you could have done better", of course you could. Dominated is "you lose something for free every state." It may be true that the theoretical full prior is preferable to the coarse prior; the data-dependent in-between is incoherent, and an easy rectification is available which does not lead to such problems. This is not a claim about optimality.
The only decision-theoretic optimal procedure is to state your full prior (without any approximation) and make decisions based on that.
No. Proceeding with the coarse prior is coherent and undominated, as are averaging, stacking, nonparametrics. Also coherent: data splitting. Etc. It is sometimes true that it is too expensive or inconvenient to be fully, coherently Bayesian; typically when this is argued some binding argument is provided that things will still 'work out okay' somehow (say, logical induction, or most proper approximations generally), showing that one is being appropriately approximately Bayesian. I do not believe there is such a bound here, it seems to me as though this produces arbitrary answers in general (through careful picking of level-sets, say). It is sometimes hard to notice that by loosening one of the Bayesian screws one has been led to something that produces arbitrary, uncheckable answers, but this seems like one of those cases.
if you follow my procedure but only make the discretization more fine that is literally an example of model expansion
No. Model 1 gives a different likelihood ratio for the data when compared to Model 2 (9:1 vs. 1:1). Model expansion qua Gelman means the model agrees with the simpler model as a special case; that's what makes it different from just changing the model. If you marginalize the finer model here, you recover a different model. Indeed, it is possible to do model expansion coherently, while this is not coherent, thus it is not that.
The comparison is fruitful in what the post does not recommend: when Gelman recommends expansion as a sensible way to make inferences from data, there are a trillion little sense checks he suggests so that one doesn't fool themselves into modelling towards a foregone conclusion (prior and posterior predictive checks, prior sensitivity analysis, there's a whole diagram), and indeed he is constantly careful in making sure not to give yourself too many 'researcher degrees of freedom', to have a coherent workflow where each step is checkable and has sense. Of course, you don't need to explain the whole thing, but sufficient warning that this is dangerous territory seems at least appropriate if not just correct.
This is closer to 'model criticism', which he makes no qualms about separating from the properly, fully Bayesian part of an analysis. Though it still isn't that, because there's a methodical aspect to it which (something like) "the posterior looks implausible" does not capture.
There is a place in the world for posts targeting a general audience, so I can't accept this criticism.
Fair enough. Still, "this is one particular way to address this well-known problem; others are A, B or C" seems non-technical and cogent. When presenting one method among many, it seems best to present weaknesses as well as strengths. If you observe the other comments, it seems as though some other commenters seem to believe that this is perfectly Bayesian and in little to no tension with the formalism. I believe this to be a reasonable reading of what was written, and that something else should have been written instead.
Not only am I not claiming that this is the only way to deal with the general problem, I'm not even claiming that it's the only way to deal with this specific example. I am stating that the the only way to choose a better discretization for this example is to look at the data.
I don't know. Is that true? If you imagine that the prior is fixed, then that the likelihood is fixed, then the posterior it produces is the posterior it produces.
If the probabilities are for betting, these are plainly worse than the coarse ones. If they are meant to be more accurate, in the Bayesian sense, they are not - see the accuracy theorems. If they are for convincing, I don't know that they convince anyone - it seems to me like they might be able to help you convince yourself that your answer is Bayesian in some sense that it is not, but I don't know that you can in any sense argue persuasively against someone who simply dismisses the post-data model out of hand as post-hoc and ad-hoc. Maybe you can state the 'rules' more plainly? You seemed to agree with my description and that one allows any conclusion (really, you disagreed with the only part that could have disallowed it).
If we are allowed to complain about the man's prior, a simple predictive check would have sufficed (his opinion does not change after seeing the data, so even before seeing anything the model may be coherently rejected). To be frank I am not convinced that it is impossible to consider "what if some witnesses are unreliable" without necessarily seeing the data. Then, whichever refinement one produces is perfectly coherent and possibly much better, insofar as it incorporates whichever other relevant prior information may work out.
All I've claimed is that there are often perfectly good reason to change your prior after looking at the data. And as far as I can tell, you accept that this is true. So what am I missing? How do I square this with the very broad and very sharp criticisms you made in your first comment?
There are "often good reasons for changing the prior after seeing the data", yes. The problem is that your suggestion is that it should be fine to do this in an ad hoc way until the posterior either reaches that arbitrary stopping point (with the near-constant likelihood) or until the Bayes rule is made to agree with intuition (which seems to be the stopping rule in the example?). Just because I concede "it is sometimes good to surgically open a man up" does not mean I agree to clefting the poor Bayesian in twain.
Other procedures preserve coherence (as above) perfectly well (or require a psychic diachronic bookie, unlike this one), a third type have either efficiency or robustness guarantees that make the model modification at least somewhat convincing. With your method, I expect an aliens-believer to roll his eyes and produce his own bisection that accords with his principles.
My initial comment was fairly vague, yes. As I said, it is a general fact that such a data-dependent prior produces incoherency, so I thought that the existence statement would be enough to convince one that maybe this has a pretty hefty (and avoidable) cost. I am glad to see that producing the counterexample instead of saying it exists seems to have helped communicate the sense in which this is bad. And, fair enough, I was incautious in saying that you abide by any generic double-counting.
On the quote being out of context: a more general claim appears before any aliens are in mention (for real problems, I've come to believe that refusing to change your prior after you see the data often leads to tragedy). But at the end, somehow, the thesis stops referring to "real problems" and only to "aliens". The sentence is the post's last line, following "So". The idea is that you meant to say that this refers not only to specifically the toy "aliens" example, but also specifically in picking a discretization? If that is truly the intention, I humbly ask that you consider appending your post with "at least for picking a discretization in this aliens example". I don't think I am alone in supposing you meant to say more than this.
Which of the following propositions do you accept?
I never suggested iterating my procedure, and have now explicitly stated that I do not endorse iterating the procedure.
If you have a model that is piecewise constant over that interval, and you split that interval in two, that is equivalent to adding a new parameter controlling the relative weight for the two sub-intervals.
Stating your full mental prior without any approximation is often difficult.
The actual procedure I suggest is reasonable in some cases.
The actual procedure I suggest is reasonable in the toy case given in the post.
There are multiple ways in which one might compare different statistical procedures.
The decision-theoretic optimal procedure, if the prior is known, is to compute the posterior and choose the decision with maximum expected utility. This is optimal in the sense that averaged over many latent variable / data samples from the prior / likelihood, this procedure produces decisions with higher expected utility than any other procedure.
The criteron you suggest, being non-dominated, means that there is no other procedure that is at least as good for some prior/likelihood, but strictly better for at least one prior/likelihood.
That criterion is not universal. I have not endorsed it. There are other possible criteria.
An alternative, also valid way to compare statistical procedures is to imagine that you have a fixed mental model and dataset and you want to approximate that fixed posterior as accurately as possible with some bounded amount of effort.
Carefully choosing a data-dependent discretization may well do better under that criteron than just choosing a single coarse discretization.
1. Agreed, minor comment. Iterating the partitioning or choosing a finer partition are identical. It is hard to interpret "such that the probability of the data is constant-ish for each term" as anything but a rule for how to partition. If you disavow it, fine. Frankly, that was more of a steelman, since at least with that procedure there's some definition on what a 'reasonable' partition looks like - without it, it's all arbitrary (a kind of reversed reverse bayes).
2. True. Also irrelevant. You have added a new parameter such that the models are no longer nested (observe the likelihood ratios). Thus, it is not model expansion, it is just a different model. In fact, if you were to do this coherently, the parameter would have a prior, instead of having been fixed a priori by introspection. 'a+x' is a model expansion of the model 'x'; '1+x' is just a different model.
3. True.
4. True; if the data-dependent partition is obtained identically regardless of what the data happens to be, the procedure is coherent and the resulting numbers are probabilities (and in fact one imagines that they are probably better, insofar as the prior information from the resulting introspection was accurate). Otherwise, unclear; would have to study the properties of the other numbers.
5. Not particularly. A 1:1 likelihood ratio is converted to 1:9 by pure post-hoc introspection. If Model 2 refines over Model 1, then Model 2 returns the exact same thing on the Model 1 level of analysis; marginalizing should return the Model 1 numbers; they are not. As stated prior, the numbers obtained are not Bayesian probabilities per se, and have unknown epistemic properties. If I am to judge "does it look more reasonable at the end", yes. If I am to judge "is it a reasonable thing to do", no. The thought-exercise has produced numbers and the numbers do not mean what they say they mean. It is unclear to me what they do, or are for.
6. Yes.
7. Indeed.
8. Different theorem. That's admissibility/the complete class theorem. I am stating non-domination in the Dutch Book sense, which is a different thing. Your prior and likelihood are fixed and I get a sure gain and you a sure loss at no risk without extra information. No other priors or likelihoods appear.
9. I don't know what to do about this one. The point of my post was, essentially, "If you are using Bayesian probabilities, you are using them for these reasons; these reasons do not apply" (dutch-bookability, etc.) "If you are just using the probability numbers as helpful pistimetric guidelines, then they are pragmatically very dubious in pretty much every formal sense I can imagine". If I understand your meaning, you seem to be going in direction 2? At which point one wonders what properties these numbers/this approximation has which makes it worth abandoning whatever credit one gives to the properties of direction 1. The coarse prior's coarse posterior retains the virtues of being Bayesian. What virtues do these numbers have?
10. Of course.
11. It 'may well do better', it is possible that it can. The question is how one evaluates if it might, or if there is any guarantee that it could, or if there is any sense that it may. There are many procedures that amount to data-driven priors (empirical Bayes, some design-based inference, some 'default' priors in software); what is done when one proposes such a procedure is, to please the Bayesians, verify that the thing in some decent sense approximates something actually Bayesian, and to please the Frequentists, show that it has good empirical properties and can verifiably improve predictions. It is unclear to me that this is the case even in the toy example; the inference obtained is not convincing. When presented with a probability obtained from this method I see no reason to trust it more than an elicited posterior.
While you claim to accept 6 and 10, your answer to 9 seems to suggest that for something labeled Bayesian, you do in fact believe your criterion is universal, that no one should be allowed to use any Bayesian procedure unless it satisfies your particular notion of optimality. Respectfully, I think it's indisputably true that:
The goal of approximating a single / fixed mental posterior with bounded effort is completely reasonable and coherent.
There are many cases where that goal is best served by choosing your prior in a data-dependent way. If your data happens to show you that the exact amount of prior weight you give to the the [0.90 to 0.91] intervals is crucial, then you would give that interval some extra thought. This will improve performance on the above goal, relative to picking a fixed partition.
In short, I continue to believe my post is correct.
There is no "particular notion of optimality", nor is it "my criterion", these are just the properties of Bayesian conditioning. What is provided here is instead a rationalization, which has no truth-tracking properties, Bayesian or otherwise. Approximating a single fixed posterior is coherent (which is to say, it is legitimate to elicit it). Justifying an approximation to a posterior by pretending to perform an update on old evidence after changing the likelihood is not; the resulting numbers are not meaningful (notice Mateusz's comment elsewhere stating that deconditioning should return the deconditioned probabilities, which is what fails here for the likelihood). "Without seeing the poll results, I never would've predicted that I should control for age and class. The answer is much less insane (p<0.05) if I report the likelihood I would have assigned" - π-hacking. When one is being Bayesian, one should strive to be Bayesian.
To put it another way, the post goes: The (Bayes) machine is provided poor inputs (misspecified prior and likelihood) and one receives a poor output. The proposal is to kick the machine (change the marginal likelihood post-hoc); one provides an example where the output is less indecent. All of the screws come loose (coherence, the defining property of doing Bayes, no longer holds). Possibly the internal gears are still fine (it has completely undefined truth-tracking properties). It seems as though it can just as easily take good inputs and return bad outputs, or in fact output just about anything (it can produce any, arbitrary posterior odds for the arbitrary likelihood ratios). No mention is made to nonviolent methods, and the risks are not properly warned against.
We may have reached the limits of this debate, but I'll state my position one last time: I believe that the abstraction of "approximate the posterior you would find if you had an infinite amount of time to think about it, while using a less-than-infinite amount of time" is completely reasonable, and in fact is a good model of a situation people often face in practice. And I believe that under that abstraction, making your prior after looking at the data (rather than before) is often a good idea, provided you do it carefully. I think you have offered no real counterargument, but instead gesturing at abstract definitions. But those abstract definitions are just ways of smuggling in a different abstraction, under which the problem I am trying to solve doesn't exist. I understand that a solution built for one abstraction may have poor properties under a different one, but that doesn't impress me.
I appreciate you taking so much time to discuss this, and I'd like to take something productive from this discussion, but I find it very hard, because you seem determined to show that the post is "fractally wrong", because you seem unwilling to make any durable concessions, and because you seem to interpret every statement I make in the least-charitable possible way. If you're truly trying to help me make the post better, that will be impossible unless you adopt a less adversarial posture.
Look, let me explain my perspective, here, then. The discussion took place on 3 entirely different 'epochs' (one 7 months ago, then another 6 months ago), and between one of these you directly told me to take a different angle; what you perceive as me being "determined to show that the post is fractally wrong" is, to me, just a result of rereading the full thing and finding different problems.
I will concede that my first approach was misguided; I was more appalled by the combination of title and closing thesis as not properly discussing the cons, and thought the useful thing to post for others to read was, roughly, the costs related to going down this road (incoherence, that is, credence-numbers that are not obtained by Bayes updating and as such do not have the advantageous properties one purportedly uses and defends Bayes numbers through, and of the kind that is hard to repair or argue is 'approximately' coherent in some implicitly-Bayesian sense) and the many unmentioned alternatives. I still think a better article would do both of those; I do not think this would make it too technical, especially the former.
Then you told me the thing has 'specific rules'; you specifically used the word 'rules'. I went and looked! That's been the attempt, here, with what you call a 'different abstraction'. My conclusion so far is that I have found no viable, good formalization of this idea; it either collapses back in the coarse model (insofar as one wants to keep coherence) or it has no properties that I can state as good or bad (if one allows the marginal likelihood to shift away). I think that the non-technicality of the post has led to an idea that is unsalvageable, at least as a procedure with 'rules', in the current form. If you wanted me to discuss it as a practical, pragmatic device for thinking about things in a pseudo-Bayesian way, then, I think I get to pout at the fact that you pointed me in the exact opposite direction. (Not only that, but, when I stated my understanding of the thing, you took it upon yourself to agree that it was, at least roughly, correct!)
So, okay. I will put away the snark and the rhetoric and just state my problems with the thing directly. Afterwards I will elaborate on what I believe can be patched, what I believe lacks nuance and warning as presented (of the kind that you have displayed in these comments after), and what I cannot help but conclude is just misguided, or at least requires substantial refinement. For whatever reason I think it is best to express these in pairs. Apologies for length.
Coherence is hard. Doing a full analysis of any problem, even simple, informal checks, while staying fully within the Bayesian formalism is sometimes basically impossible.
However, it is not impossible. Sometimes, with a little difficulty, it can be done. This often requires care; it sometimes requires accepting things that seem unintuitive or even nonsensical. This is because most of the results that point to Bayesianism as the thing to do, the reasons why one performs Bayesian updating and not Dempster-Shafer updating or Jeffrey updating or ad hoc numerical reasoning, usually come in the form of 'if and only if's; the Dutch book converse, the CCT converse, the accuracy theorems, Cox's theorem, forecasts-as-martingales, all of these uniquely identify probabilities-as-credences or Bayes updating as, specifically and uniquely, the thing to do; if you end up doing something else, you lose the properties instantly.
This difficulty means that, often, approximating coherence is essentially required.
If and only ifs are very fragile things; you step away from the formalism in one subtle way and all of these things have suddenly broken into pieces somewhere. What you must ensure, then, is that most of the pieces are still there. Sometimes this can be done formally; I mentioned logical induction; there it is shown that one is only boundedly Dutch-Booked in a relatively benign way. Safebayes is a more pragmatic one; it is very clear to see how that is very approximately Bayesian. Empirical Bayes as approximate hierarchical Bayes.
Sometimes the approximation looks right, and becomes unsalvageable very quickly, in some mathematical sense. This is the case for many objective priors; everyone can nod at the univariate Jeffreys' rule, but not even Jeffreys defended the multivariate one. You may fondly remember the right-Haar prior, but the left-Haar inspires little fondness. Very classically, nobody puts up a fight to defend flat priors anymore. And so on. I tried putting your procedure in some formalizations and this seemed to be the case; but you (now) reject them, so I will not reiterate.
Sometimes, there is no formal procedure to analyse the proximity of.
Other times, the approximate-incoherence is incurred from some (sometimes necessary) pragmatic data-procedure. You change the prior after the sampler runs for too long, you pick the best model out of a bag after looking at a predictive score, you throw out the model because the predictions look kind of bad; nobody (in the 21st century) would dispute that these things happen.
However, what is missing in your post and not missing whenever Bayesian statisticians discuss these things is that when you do this you are treading very dangerous ground. So, generally, the best of them (which is to say, Gelman's cohort, Stephen Walker, Rubin, etc), if one cannot work out some alternative justification or method to do these things that is still within the formalism (as with Gelman in the model expansion cases, and in Rubin's case something as trivial as randomization), at the very least you run a bunch of checks to see that you are really not shooting yourself in the foot somewhere.
This is because, when you do abandon coherence in a way that does not give you some kind of 'bounded' error, even in some conceptual sense (as in your case, where the marginal likelihood shifts arbitrarily; I will reiterate that the problem is not the prior becoming more detailed, but the marginal likelihood shifting alongside it in a way that does not reduce the model back to it), you really are playing with fire. The numbers obtained by such a method do not have the philosophical clarity of the numbers obtained with Bayesian updating, even if you could imagine that you might have obtained the numbers through some Bayesian update that was not performed.
I understand that you may raise an eyebrow at this, but, the analogy with p-hacking is relevant: you do a little tweak here or there, in ways that might make a lot of common sense, hell, sometimes, it really would be better if you had ran the analysis you wish you had ran after you saw the data; and, really, if the null your tweak led to be rejected really was wrong, then one may feel some solace that no damage was done. But, still, the '0.05' number just no longer means a 0.05 error rate; the number has ceased having a meaning. You may run some garden-of-forking-paths analysis that might elucidate what that number might mean, but it is no longer the p-value it purported itself to be.
Besides, frankly, sometimes you have to put common sense and other statistical properties over following Bayesian edicts directly.
So, lacking a formalization through which to analyse approximation quality and observing that the procedure can output ultimately arbitrary results (thus even a conceptual approach to approximation seems hard), since they are no longer clean Bayes-rule probabilities obtained from a prior and posterior update, what are they for?
It is easy to make this kind of procedure look good if one picks out examples where it happens to reach the correct answer (when the null is false and you performed a terrible analysis, p-hacking can look absurdly, genuinely convincing!). I don't know about you, but, for me, and I think for most statisticians, we like to look at what happens if the person doing it is maybe a little tendentious someplace; maybe not too disciplined here or there. A statistical (or a Bayesian) justification can often serve as a little "certificate of formal correctness" for people who will then no longer squint when they do some things that they really should squint at.
So, how does this procedure fare in this respect? I honestly think that the infamous "Bayesian arguments for God" are, frankly, not too far away from this; you could probably put them in exactly this language, or derive them in this way. "I first took a prior and likelihood, did an update, it did not seem sensible to me; I then found the little tweak to the model that made it make sense with my intuition, so now I have obtained a Bayesian justification for my belief"; I think you know how bad this can be. I am not saying this to imply that the idea fares as poorly as that in general, I imagine in most situations it acts moderately more benignly, but that it is a natural thing to reach for, and easy to imagine that you are doing nothing wrong through, regardless of how right you are (or aren't).
Of course, you say that things should only be done in the right way, in the post. You do not, I think, specify how to do this; honestly the bisection you performed in the aliens example would look pretty tendentious to an aliens-believer, if you imagine that a few hundred years from now we discover that aliens really are real I think you know that the toy example looks pretty bad. A demonstration of a good method survives the particular example being wrong; one can at least be wrong gracefully, or at least where the flaw lies can be shown to subside in the specifications of the method.
In any event, this is a good model for something that is encountered in practice.
I think that you are right, and indeed you are right that people do some version of this, often. But, one virtue I can name of the Bayesians of the 21st century at least is that they do not do this, typically; if their model was wrong, they are fine eating the bullet, without performing any rigmarole that puts their objectivity into question, they are fine with putting "money on the table". Here, you salvage a broken inference by a method that simply would not convince a skeptic. But, the problem that you salvage lies in the likelihood, not in your prior; just allow yourself to observe some more data and you'll be fine, or split the sample first, or cautiously observe through a prior predictive check that your model allows no information to be gained regardless of what happens, or in fact just obtain a better prior (I believe witness unreliability to be the Hume thesis on miracles in the 1700s, so it can certainly be prior to the data in this example to consider it).
Ultimately, when I looked for rules, all I found was the rough sense that the posterior should match one's intuition. Of course intuition is often right and good, possibly often better than our Bayesian calculations. But, ultimately, I think that performing this kind of calculation to affirm one's intuitions is at best just an elucidation of at least part of why you think you think the things you think; at worst, and to be frank, very typically in this framing at least, a way to smuggle vague intuition as formally justifiable insight. The best cases of Bayesian inference are when you perform it and convince yourself out of an intuition that was actually wrong; spiritually, at least, this seems to basically prohibit this, there is always an out, always a way to slice up the likelihood to get numbers that you are more concordant with.
I think I will finish on that note. On how the idea can be patched: As Mateusz mentioned earlier, having the partition be data-independent solves the coherency issue entirely (qua Savage). Or, at least, a formal justification that binds one not to change the likelihood too much. Or, in some sense an argument that one will not easily mislead themselves, that one will not be able to convince themselves of false conclusions as easily as true ones, or a sensitivity analysis of some kind that shows, okay, maybe your new flimsy post-hoc model is not too flimsy. Or, as a last resort, if you want a number for an intuition, you can elicit a full posterior from your intuitions via all the other methods in a way that doesn't make it look like reasoning was performed when it really wasn't. Mostly, if the idea has formalizable value, it ought to be formalized, at least partially; if it can't be, then it should at least be argued that one does not easily mislead themselves through it.
On how the post can be improved, I already mentioned a fair few. If you think that your last paragraph is really only specifically about the toy aliens example and nothing else, I think it should be stated, since it is very unusual for articles to end on a note about the examples rather than what they illustrate. I also think alternatives and risks should be mentioned, if one cannot improve the idea as above to avoid them; a post on a heuristic is useless without showing where and how the heuristic fails, otherwise one falls off cliffs. Reframe the post as being about changing your model post-data, since that's what actually changes the inference, not the prior (the prior changing is the most benign part of this, ultimately; it just opens the door to an entirely different model, which does not match on the marginal).
Bunglesome to respond to 'you are tendentiously trying to show this is fractally wrong' with the longest reply yet, but, you put me between Scylla and Charybdis by asking me for clearly stating a full, coherent position in as neutral a light as possible. Would've written a shorter letter, etc.
Addendum: Possibly some of this comes down to a difference in intuition: you imagine your procedure as mimicking a clean, contiguous cut of a 1d curve; there it can at least kind of be imagined what a sensible cut may look like. Insofar as one applies this procedure to questions like your example, I am imagining that the cut is really happening on some infinite-dimensional, abstract space - God only knows what reasonable cuts look like, I have no idea if my particular one is random, tendentious, an improvement or a bias. Certainly I trust no-one else's.
Bayes Theorem is so simple that it's felt like if you get it, say if you're gone through the Bayes rule guide, then you get it. Posts like Why I’m not a Bayesian and this one show that there's more to think about. That's pretty cool when you are a Bayesian and it's a cornerstone of your epistemology. Good job here.
A question for the author, is the position here something like the actual data/question at hand influences how you should calculate your priors, but only previously held info should be fed into the calculation?
I sort of flicked past the equations, seeing that they were vaguely intuitive and kinda redundant to the text, until the first Huh heading. At that point, I sighed deeply and headed back to the top to actually read the equations. Will report back.
Reporting back: yep worth reading the equations slowly
This is great. I don't remember the last time a post made me flip so hard from "the thesis is obviously false" to "the thesis is obviously true"! And not just because I didn't understand it, but also because I learned a thing. (Tho part was a pedagogically useful misunderstanding.)
This post seems to fundamentally misunderstand how Bayesian reasoning works.
First of all, the opening "paradox" isn't one. If you have an inconclusive prior belief, and then you apply inconclusive data, you should have an inconclusive result. That's not weird. Why is it surprising?
Secondly, the argument that follows is adding more conditions to the thesis that muddy the waters. P(A, B) is always less than P(A) if P(B) > 0. The probability that a coin lands on heads and that it's tuesday is always less than the probability that a coin lands on heads. That fact does not tell you anything about the probability of the coin itself, and trying to include the fact that coins land face up on tuesdays less than coins land face up into your beliefs on the nature of coins, you will go mad because those things are not related and the probability relationships are tautological.
Thirdly, the post ignores marginal likelihood, the denominator of Bayes theorem. If there are alternative explanations for evidence[1], then it is weak evidence, and should not result in a large change in belief. It's not just about how likely the evidence is given a thesis, it's about how likely some evidence is given that thesis as opposed compared to all others. P(cough|COVID-19) is very high, almost 1. But that's not really worth much because P(cough|any other respiratory disease) is also very high, almost 1, and 1/1=1, so the posterior = the prior.
Finally and most importantly, the thesis is that priors should include past evidence, which is like... yeah. That's how bayesianism works. You calculate the posterior given a new piece of data, and that posterior becomes your prior for the next piece of data. This is Bayesian Updating.
When you want to decide whether to believe something, you don't have to start from the most naive argument from first principles every single time. Beliefs are refined over time by evidence. When you see new evidence, your reasoning doesn't start from what you knew when you were a baby, it starts from what you knew the moment before exposure to new evidence.
The 5th figure is incorrect and should be like what I show here. Then you will not get the nonsensical P[data|aliens] = 50%.
There are two kinds of errors the piece makes:
1. Probabilities do not add to 100%, which is the one I just pointed out.
2. Probabilities can be quite far off. The Baysian method assumes you can get close, and refines the probability. If you cannot get close, i.e. if initial data samples are far off the assumed probability, then the Baysian method does not apply and you'll have to use the Gaussian method, which requires a lot more samples.
Applying Gaussian reasoning to UAP problem:
1. There are 1.5 million pilots in the world (source: Google)
2. 800 official reports to Pentagon's AARO investigation => 800/1,500,000 = 0.053% probability of something that appears unnatural based on current understanding. P[aliens]<<0.053% in light of missing artifacts.
3. 120,000 sightings including non-pilots (assumed to have less expertise) => 8% probability of something that appears to non-experts as unnatural, so a much lower confidence number, but it explains the widespread "belief" in UFO/UAP.
Other Possibilities P[unknown other]
The author divides up possibility space into 4 categories. An interesting one to expand further is "There are aliens. But they stay hidden until humans get interested in space travel. And after that, they let humans take confusing grainy videos." In fact there were unnatural phenomena apparently seen earlier, which were identified tentatively as...
1. evidence for survival of death (seeing ghosts => ceremonial burian with artifacts needed in afterlife)
2. evidence for "High Gods" or demons
3. dreams, visions, hallucinations (not a popular majority explanation)
These are the priors of H. Sapiens collectively, and as data is being collected from H. Sapiens (mostly), they should be included in analytical priors regardless of the opinion of the investigator - because they are the opinions of the data source and inseparable from the data.
A very interesting side point from Hunter-Gatherers and the Origins of Religion - PubMed is that while many superstitious beliefs evolve naturally, High Gods do not fit the pattern of naturally evolved beliefs. They were produced or imposed in some other way. Either by high gods themselves, or by humans on other humans. Either one fits into the P[unknown other] category. You can of course dismiss this or assign P approaching 0. But your argument will appear invalid to 80% of Americans (and much of the rest of the world). Perhaps you do not care about them. But if you are going to sell to them or your work is supported by grants from them, you have a very limited future if you do not take beliefs you do not share into account. The simple solution is to just discard the Baysian method as inapplicable when the data source itself has a wide variety of undecidable beliefs.
Undecidable meaning you cannot get them to agree because each belief if used as a prior causes data to be interpreted in a way to reinforce the belief.
The figure you are referring to does not need to add up to 100%, since it is showing P[data | aliens] and P[data | no aliens].
P[data | aliens] and P[not data | aliens] need to add to 100%, but that is not on the graph.
As an extreme case where P[A | B] + P[A | C] != 1, consider A = coin did not land on its edge, B = the coin is ordinary, C = the coin is weighted to land heads twice as often as tails.
Then P[A | B] = 0.9999 and P[A | C] = 0.9999 would be reasonable values.
Pragmatically, in data analysis tasks, what you do is a separate preliminary data collection that you only use to decide the priors (the whole data analysis structure, really) and then collect data again on which you run the actual analysis. This applies to non-Bayesian data analysis as well. This duplicated data collection helps you stay objective and not sneak into the prior any information which would not be Bayes-kosher to glean from the data. Of course it's less efficient because you are not using all the data in the final analysis.
I liked this!
Isn’t this post an elaborate way of saying that today’s posteriors are tomorrow’s priors?
As in- all posteriors eventually get baked into the prior.
How I would pack in one sentence the intuition for why 'changing priors' is ok: "Sometimes the number you wrote down as your prior is not your real prior, and looking at the data helps you sort that out."
Yes, that's how Bayesianism is supposed to work. It's called Bayesian Updating.
You don't wake up every day with a child's naivete about whether the sun will rise or not, you have a prior belief that is refined by knowledge combined with the weight of previous evidence.
Then, upon observing that the sun did in fact rise on this new morning, your belief that the sun rises every day gets that much stronger going into the next day.
Interestingly, this is broadly consistent with Leonard Savage's view of priors, at least according to Ken Binmore https://link.springer.com/article/10.1007/s41412-017-0056-1
Everybody would presumably agree that it would be better for Alice to consult her gut feelings when she has more evidence rather than less. For each possible future course of events, she should therefore ask herself what subjective probabilities her gut would come up with after experiencing these events. In the likely event that these posterior probabilities turn out to be inconsistent with each other, she should then take account of the confidence she has in her current snap judgements to massage her posterior probabilities until they become consistent. After the massaging is over, Alice would have eliminated the possibility of being surprised by a black-swan event. She would already have taken account of the impact that all future information might have on the unformalized internal model that she uses in determining her beliefs. She would not only have adjusted the subjective probabilities she attaches to events in the small world with which she begins, but potentially expanded her state space to a larger small world in light of the new possibilities that black-swan events commonly suggest (Gillies 2001; Williamson 2003).
The end-product of Alice’s massaging process is therefore a bunch of consistent posteriors defined on a state space that she will never need to revise. With the consistency axioms of Savage’s theory of subjective probability, her massaged posteriors can all be deduced by Bayes’ rule from a single prior. In this story, Bayes’ rule is therefore reduced to a mere book-keeping tool that saves Alice from having to remember all her massaged posterior probabilities. The prior that Savage attributes to Alice therefore squeezes all the juice that can be squeezed from the disorderly set of impressions with which she comes to the problem.
In short, the procedure is something like:
Experimentally, can I say frequentist actually have a better representation of the "event" to observe? It doesn't require the "observer" to make any prior assertion about the distribution. This is especially true when we are gathering more and more data with ease. Those data might not be of high quality. But the volume can simply "wash out" those quality issues as long as the collection method is not biased. I always have this nagging feeling that bayesian is not a very practical tool when it comes to experimental science.
They say you’re supposed to choose your prior in advance. That’s why it’s called a “prior”. First, you’re supposed to say say how plausible different things are, and then you update your beliefs based on what you see in the world.
For example, currently you are—I assume—trying to decide if you should stop reading this post and do something else with your life. If you’ve read this blog before, then lurking somewhere in your mind is some prior for how often my posts are good. For the sake of argument, let’s say you think 25% of my posts are funny and insightful and 75% are boring and worthless.
OK. But now here you are reading these words. If they seem bad/good, then that raises the odds that this particular post is worthless/non-worthless. For the sake of argument again, say you find these words mildly promising, meaning that a good post is 1.5× more likely than a worthless post to contain words with this level of quality.
If you combine those two assumptions, that implies that the probability that this particular post is good is 33.3%. That’s true because the red rectangle below has half the area of the blue one, and thus the probability that this post is good should be half the probability that it’s bad (33.3% vs. 66.6%)
(Why half the area? Because the red rectangle is ⅓ as wide and ³⁄₂ as tall as the blue one and ⅓ × ³⁄₂ = ½. If you only trust equations, click here for equations.)(Why half the area? Because the red rectangle is ⅓ as wide and ³⁄₂ as tall as the blue one and ⅓ × ³⁄₂ = ½. If you only trust equations, click here for equations.)
It’s easiest to calculate the ratio of the odds that the post is good versus bad, namely
It follows that
and thus that
Alternatively, if you insist on using Bayes’ equation:
Theoretically, when you chose your prior that 25% of dynomight posts are good, that was supposed to reflect all the information you encountered in life before reading this post. Changing that number based on information contained in this post wouldn’t make any sense, because that information is supposed to be reflected in the second step when you choose your likelihood
p[good | words]. Changing your prior based on this post would amount to “double-counting”.In theory, that’s right. It’s also right in practice for the above example, and for the similar cute little examples you find in textbooks.
But for real problems, I’ve come to believe that refusing to change your prior after you see the data often leads to tragedy. The reason is that in real problems, things are rarely just “good” or “bad”, “true” or “false”. Instead, truth comes in an infinite number of varieties. And you often can’t predict which of these varieties matter until after you’ve seen the data.
Aliens
Let me show you what I mean. Say you’re wondering if there are aliens on Earth. As far as we know, there’s no reason aliens shouldn’t have emerged out of the random swirling of molecules on some other planet, developed a technological civilization, built spaceships, and shown up here. So it seems reasonable to choose a prior it’s equally plausible that there are aliens or that there are not, i.e. that
Meanwhile, here on our actual world, we have lots of weird alien-esque evidence, like the Gimbal video, the Go Fast video, the FLIR1 video, the Wow! signal, government reports on unidentified aerial phenomena, and lots of pilots that report seeing “tic-tacs” fly around in physically impossible ways. Call all that stuff
data. If aliens weren’t here, then it seems hard to explain all that stuff. So it seems likeP[data | no aliens]should be some low number.On the other hand, if aliens were here, then why don’t we ever get a good image? Why are there endless confusing reports and rumors and grainy videos, but never a single clear close-up high-resolution video, and never any alien debris found by some random person on the ground? That also seems hard to explain if aliens were here. So I think
P[data | aliens]should also be some low number. For the sake of simplicity, let’s call it a wash and assume thatSince neither the prior nor the data see any difference between aliens and no-aliens, the posterior probability is
See the problem?
(Click here for math.)
Observe that
where the last line follows from the fact that
P[aliens] ≈ P[no aliens]andP[data | aliens] ≈ P[data | no aliens]. Thus we have thatWe’re friends. We respect each other. So let’s not argue about if my starting assumptions are good. They’re my assumptions. I like them. And yet the final conclusion seems insane to me. What went wrong?
Assuming I didn’t screw up the math (I didn’t), the obvious explanation is that I’m experiencing cognitive dissonance as a result of a poor decision on my part to adopt a set of mutually contradictory beliefs. Say you claim that Alice is taller than Bob and Bob is taller than Carlos, but you deny that Alice is taller than Carlos. If so, that would mean that you’re confused, not that you’ve discovered some interesting paradox.
Perhaps if I believe that
P[aliens] ≈ P[no aliens]and thatP[data | aliens] ≈ P[data | no aliens], then I must accept thatP[aliens | data] ≈ P[no aliens | data]. Maybe rejecting that conclusion just means I have some personal issues I need to work on.I deny that explanation. I deny it! Or, at least, I deny that’s it’s most helpful way to think about this situation. To see why, let’s build a second model.
More aliens
Here’s a trivial observation that turns out to be important: “There are aliens” isn’t a single thing. There could be furry aliens, slimy aliens, aliens that like synthwave music, etc. When I stated my prior, I could have given different probabilities to each of those cases. But if I had, it wouldn’t have changed anything, because there’s no reason to think that furry vs. slimy aliens would have any difference in their eagerness to travel to ape-planets and fly around in physically impossible tic-tacs.
But suppose I had divided up the state of the world into these four possibilities:
No aliens + normal peopleNo aliens + weird peopleNormal aliensWeird aliensIf I had broken things down that way, I might have chosen this prior:
Now, let’s think about the empirical evidence again. It’s incompatible with
no aliens + normal people, since if there were no aliens, then normal people wouldn’t hallucinate flying tic-tacs. The evidence is also incompatible withnormal alienssince is those kinds of aliens were around they would make their existence obvious. However, the evidence fits pretty well withweird aliensand also withno aliens + weird people.So, a reasonable model would be
If we combine those assumptions, now we only get a 10% posterior probability of aliens.
Now the results seem non-insane.
(math)
(math)
To see why, first note that
since both
normal aliensandno aliens + normal peoplehave near-zero probability of producing the observed data.Meanwhile,
where the second equality follows from the fact that the data is assumed to be equally likely under
no aliens + weird peopleandweird peopleIt follows that
and so
Huh?
I hope you are now confused. If not, let me lay out what’s strange: The priors for the two above models both say that there’s a 50% chance of aliens. The first prior wasn’t wrong, it was just less detailed than the second one.
That’s weird, because the second prior seemed to lead to completely different predictions. If a prior is non-wrong and the math is non-wrong, shouldn’t your answers be non-wrong? What the hell?
The simple explanation is that I’ve been lying to you a little bit. Take any situation where you’re trying to determine the truth of anything. Then there’s some space of things that could be true.
In some cases, this space is finite. If you’ve got a single tritium atom and you wait a year, either the atom decays or it doesn’t. But in most cases, there’s a large or infinite space of possibilities. Instead of you just being “sick” or “not sick”, you could be “high temperature but in good spirits” or “seems fine except won’t stop eating onions”.
(Usually the space of things that could be true isn’t easy to map to a small 1-D interval. I’m drawing like that for the sake of visualization, but really you should think of it as some high-dimensional space, or even an infinite dimensional space.)
In the case of aliens, the space of things that could be true might include, “There are lots of slimy aliens and a small number of furry aliens and the slimy aliens are really shy and the furry aliens are afraid of squirrels.” So, in principle, what you should do is divide up the space of things that might be true into tons of extremely detailed things and give a probability to each.
Often, the space of things that could be true is infinite. So theoretically, if you really want to do things by the book, what you should really do is specify how plausible each of those (infinite) possibilities is.
After you’ve done that, you can look at the data. For each thing that could be true, you need to think about the probability of the data. Since there’s an infinite number of things that could be true, that’s an infinite number of probabilities you need to specify. You could picture it as some curve like this:
(That’s a generic curve, not one for aliens.)
To me, this is the most underrated problem with applying Bayesian reasoning to complex real-world situations: In practice, there are an infinite number of things that can be true. It’s a lot of work to specify prior probabilities for an infinite number of things. And it’s also a lot of work to specify the likelihood of your data given an infinite number of things.
So what do we do in practice? We simplify, usually by limiting creating grouping the space of things that could be true into some small number of discrete categories. For the above curve, you might break things down into these four equally-plausible possibilities.
Then you might estimate these data probabilities for each of those possibilities.
Then you could put those together to get this posterior:
That’s not bad. But it is just an approximation. Your “real” posterior probabilities correspond to these areas:
That approximation was pretty good. But the reason it was good is that we started out with a good discretization of the space of things that might be true: One where the likelihood of the data didn’t vary too much for the different possibilities inside of
A,B,C, andD. Imagine the likelihood of the data—if you were able to think about all the infinite possibilities one by one—looked like this:This is dangerous. The problem is that you can’t actually think about all those infinite possibilities. When you think about four four discrete possibilities, you might estimate some likelihood that looks like this:
If you did that, that would lead to you underestimating the probability of
A,B, andC, and overestimating the probability ofD.This is where my first model of aliens went wrong. My prior
P[aliens]was not wrong. (Not to me.) The mistake was in assigning the same value toP[data | aliens]andP[data | no aliens]. Sure, I think the probability of all our alien-esque data is equally likely given aliens and given no-aliens. But that’s only true for certain kinds of aliens, and certain kinds of no-aliens. And my prior for those kinds of aliens is much lower than for those kinds of non-aliens.Technically, the fix to the first model is simple: Make
P[data | aliens]lower. But the reason it’s lower is that I have additional prior information that I forgot to include in my original prior. If I just assert thatP[data | aliens]is much lower thanP[data | no aliens]then the whole formal Bayesian thing isn’t actually doing very much—I might as well just state that I thinkP[aliens | data]is low. If I want to formally justify whyP[data | aliens]should be lower, that requires a messy recursive procedure where I sort of add that missing prior information and then integrate it out when computing the data likelihood.(math)
Mathematically,
But now I have to give a detailed prior anyway. So what was the point of starting with a simple one?
I don’t think that technical fix is very good. While it’s technically correct (har-har) it’s very unintuitive. The better solution is what I did in the second model: To create a finer categorization of the space of things that might be true, such that the probability of the data is constant-ish for each term.
The thing is: Such a categorization depends on the data. Without seeing the actual data in our world, I would never have predicted that we would have so many pilots that report seeing tic-tacs. So I would never have predicted that I should have categories that are based on how much people might hallucinate evidence or how much aliens like to mess with us. So the only practical way to get good results is to first look at the data to figure out what categories are important, and then to ask yourself how likely you would have said those categories were, if you hadn’t yet seen any of the evidence.