Thanks for putting this together! Permit me to promote PDGs a bit, which I think are under-sold here :)
First: motivating a "degree of inconsistency" scoring-function semantics, valued in
Why not simply declare that everything of import is a set/category/TM? In tension with generality is the desire to expose minimal knobs and buttons to the modeler. A common complaint about probability is that it's too difficult for an agent to maintain a joint probability distribution about all variables of interest---and the burden of selecting a lower semi-continuous function on distributions (over which variables?) is WAY higher, since there are far more choices to make, and less standard guidance on how to make them. PDG semantics are arguably the natural way to answer these questions and produce a belief
A challenge for @davidad: can you motivate a scenario that really requires stepping outside of the sub-class of beliefs generated by PDGs, and is clearly better-modeled by a different lower semi-continuous function?
Finally, a couple of quick technical points/corrections:
Thanks for your thoughtful engagement!
On syntax vs semantics, I fully agree that your work is the state of the art of how to produce a belief
On convexity, I was going off of Lemma A.1 of "Probabilistic Dependency Graphs", but now I see that this is guaranteed of the full semantics only under the condition
My current view is that negative inconsistency/incompatibility is an anti-pattern, because it means we no longer have the property that the "expected value" of
Thank you for your generous recognition and changes to the post.
To connect some dots, recall that
I sympathize with your suspicion of negative belief, and I advise anyone who isn't 100% sure of their footing to assume
Finally: I like your answer to my challenge :)
I think there might be a PDG angle, but I haven't fully worked it out yet. Here are my thoughts so far:
I'd like to add another point here regarding burden. The fact that davidad didn't make this point is surprising to me, so I may be misunderstanding something here - please let me know if so!
I want to push back on your point that "the burden of selecting a lower semi-continuous function on distributions (over which variables?) is WAY higher, since there are far more choices to make, and less standard guidance on how to make them" - to do so I will draw an analogy. Consider a problem which asks the user for a function satisfying some properties on some space
This of course does not discount the utility of PDGs. As you note, their clearly interpretable presentation makes them a very natural[2] way to construct belief functionals, and since they are a more structured type there are probably theorems that will be easier to prove about PDGs than about belief functionals (especially since there should be things that are provable about PDGs but are not true or even expressible for general belief functionals). Furthermore, since PDGs carry more structure than belief functionals, it is more clear how to describe/understand the dynamics of the former (such as LIR) as opposed to the latter (is there a natural way to do something like LIR in this context?). But it does seem to me that belief functionals are in some sense more natural / more fundamental than PDGs (on account of being a broader/less-structured/less-artificial type), so I hope that the two can inform each other: PDGs are fertile ground for interesting belief dynamics, and the theoretical constraints of working at the belief functional level might help prune those dynamics down to the most essential ones!
More freedom does not necessarily mean less modeling burden (and certainly not less burden of choice). Making 300 decisions in sequence is a greater burden than making the first 10 and having the others automatically handled by context. Low-level programming languages are burdensome despite (or perhaps because of) the freedom that they expose to the modeler. And universal constructions are appealing in part because they allow you to specify important information arguably without making any (unjustified) choices at all (i.e., very little modeling burden)! As davidad points out, this burden is a matter of syntax, i.e., the interface for constructing valid objects. Do the constraints of that interface provide helpful, simplifying guidance? Or does working within the constraints impose (computational) difficulties on the modeler? My experience is that the structure of PDGs does both: it makes specifying "natural" beliefs easier and "unnatural" ones more difficult.
My understanding is that @davidad wants to sidestep syntax and focus on delivering an IR (intermediate representation) for beliefs: a simple common compilation target that retains some useful structure (in this case, additive combination). This effectively deflects my point about the modeler's burden, since his aim is instead to reduce the burden of a compiler that targets the IR. This allows him to endorse PDGs as one particular (nice) language for articulating inconsistency functionals, casting PDG semantics as the way to compile them. This makes a lot of sense to me, and the mathematical simplicity is definitely an appealing advantage over PDGs for this purpose.
However, I'm not 100% convinced that this is a good idea. The class of PDGs is itself already a useful IR, since it unifies so much, and it intentionally occupies a different point in the design space. The structure can indeed create some burden for certain modelers (e.g., for one who wants to arrive at a certain preconceived belief functional, but now, annoyingly, has to figure out how to write it down in terms of local conditional probabilities and confidences), but this exercise also has nice side-effects (e.g., gives you justification for your loss function), often yielding significant interpretability benefits. As you point out, the map from PDGs to belief functionals is non-injective; in my view, that is because the mapping loses something that is worth tracking in a belief state. There's technically a difference between specifying P(X) and P(Y|X) vs P(Y) and P(X|Y) (all with the same confidence), even though we all agree they are semantically equivalent. Making this distinction (and others in the same class) can be relevant for revising your beliefs (e.g., if you determine that your mechanism for forming conditional beliefs had a flaw). My concern is that direct specification of an inconsistency functional could short-circuit the modeling process that PDGs are designed to elicit.
FWIW, it's not really clear to me how to map the IR/higher-language analogy into this situation, but when I try to it feels more like belief functionals are the higher-level language and PDGs are a particular lower-level implementation thereof (and from this perspective it doubly makes sense why the map isn't injective, since there are a wide variety of ways to compile a given higher-level program into a given lower-level representation); I don't really see how it could be the case that PDGs are less burdensome than belief functionals given that if you want to specify your belief functional with a PDG, you're always free to do so, and given that you're also free to do so with simpler techniques (e.g. all of the items under "...to express beliefs" in the original post, which are IMO much simpler to express with belief functionals that aren't PDGs).
Regardless, I do find your argument at the end ("in my view [...] had a flaw)") quite compelling; I think we're on the same page about belief-revision being an extremely important aspect of belief-modeling, and if there's in fact not a clear way to do this sort of credit assignment for belief functionals then it does make them significantly less attractive than PDGs for understanding belief dynamics.
A second, different kind of answer to your challenge is a stochastic PDE. The PDG formalism is finitary (finite node set
As a concrete example, for the SPDE
Fair enough, although, if we're modeling a continuum of variables
Thanks again for the two nice concrete examples!
Re: d20
If you use Bayesianism + CDT, and the coin has already landed: Yes, you never take the d20 bet.
Why CDT is well-motivated for Bayesianism by default: Because CDT motivates taking bets according to one's subjective probabilities. Without CDT (or something close), non-Bayesian agents can choose to decline bets that would naively be good according to their subjective probabilities. This makes it harder to apply Dutch book arguments to them (to argue they should be Bayesian). See my recent post, "Simple Dutch books versus Sleeping Beauty halfers".
As my post suggests, it is worth considering EDT+FNC as a quasi-Bayesian anthropic decision theory, as the main alternative to CDT+SIA. Here's how EDT+FNC reasons on the d20:
"I have no information other than the problem setup; there is no FNC update to make so far. Evidentially on my choosing Heads, Omega has made it Tails, meaning I lose. Similarly, I lose evidentially on choosing Tails. Evidentially on the d20 bet, I'm not guaranteed to lose. Therefore I should pick the d20 bet."
So they pick the d20 bet despite assigning at least 50% subjective probability to at least one of Heads and Tails.
Since FNC is not even Bayesian (dynamic inconsistency as pointed out by Stuart Armstrong), Bayesianism does not combine well with EDT. See Stuart Armstrong's Anthropics: Full Non-indexical Conditioning (FNC) is inconsistent.
I think one route towards finding alternatives to Bayesianism is to consider decision theories with Dutch book resistance and/or ex ante optimality, and think of what probabilities/beliefs go well with them.
I really appreciate this comment, because I must admit I was not even previously aware of FNC, and I think FNC+EDT solves my problem of completing the corresponding decision theory for my notion of beliefs.
I already was favorable to Halpern’s MWER as a decision rule, but MWER leaves the
This is the most immediately appealing notion of imprecise beliefs I've seen. "Level of inconsistency" feels like a much more natural primitive to me than credal sets or minimax decision rules. I'd be very keen to see some toy applications where it handles the issues well. Bonus points if they're not especially exotic.
So far as this is true that all formalizations of belief are full subcategories of davidad!beliefs... why has this only been posted today???
Also, what about stuff like ... various logics? AGM theory? Do you not count them as "beliefs"? They are less probability-like but Dempster-Shafer functions are very probability-like and I don't see them here. I recall that Halpern (and a coauthor?)'s Generalized Expected Utility generalized almost every decision rule (?) but (some decision rule using?) Dempster-Shafer was the main interesting exception. Is something similar going on here?
Also, it seems to me that you're equating beliefs with functions judging consistency of probability distributions with them. So if you believe that X and Y are independent then this is expressed by a belief function that judges accordingly. But I would expect that in some cases (not in the probabilistic independence case) this equates differently expressed propositional attitudes (different senses/intensions) because of having the same references/extensions, for the purpose of determining a belief function. Do you consider this an issue at all? My guess is that keeping the intentional difference in mind is relevant for handling ontological crisis-shaped stuff.
In any case, hooray for expanding the domain fo discourse to reveal the structure already there but hidden!
To your second point, very much yes. I am also working on a much bigger framework for world-modeling, still along the lines of Safeguarded AI TA1.1, which takes this notion of beliefs as a central ingredient in its semantics. The syntax is most of the work. By syntax/semantics I mean the same thing as sense/referent and intension/extension.
The purpose of the semantics is to ground judgments about observational equivalence about belief states, but when we are exploring the hypothesis space about infinite-dimensional
Richardson’s PDGs are in my opinion the best syntax published to date.
Dempster-Shafer functions express beliefs via their credal sets. AGM theory and epistemic logics carry only true/false beliefs about propositions, which are quite degenerate but can still be embedded as full subcategories of credal sets (
Not all credal sets satisfy the inclusion/exclusion rule that characterizes DS belief functions. Also, they arguably combine differently (Dempster's rule vs conditioning pointwise).
Definitely the right way of turning a DS function into a belief
Suppose there is an unfair coin, which you know to be > 5% unfair (but not exactly how much or in which direction), and you must choose between three options: bet on Heads, bet on Tails, or bet instead on a fair d20 coming up at least 12. Exercise for the reader: Every probability distribution you could possibly believe makes choosing the d20 irrational.
Why do you believe this? There are plenty of probability distributions that work here and would imply choosing d20.
The unfairness constraint on the coin is extremely weak and only excludes a small range of probability distributions. The fairness constraint on the d20 is much stronger but essentially irrelevant. It is also not entirely clear what "fair" means in this context, especially since you mention Omega immediately afterward.
For example, a coin that lands Heads if and only if I bet Tails before the flip is obviously unfair by the required margin, and a distribution that assigns significant probability to such a hypothesis would make it clearly not irrational to roll the fair d20 instead.
You seem to be saying in the last sentence that it is literally impossible to believe such a distribution. Why do you believe that?
There exists no probability distribution, neither about the outcome (
For all actual probability distributions, either the bet on Heads strictly dominates the d20, or the bet on Tails strictly dominates the d20. This is intended as a reductio for representing your belief as a probability distribution — of course I agree with you that in fact it is rational to bet on the d20 in this situation.
Makes sense that probabilities on only outcomes or biases aren't rich enough to imagine Omega messing with you, but is there some slightly richer probabilistic model that works fine?
E.g. if Omega is predicting you before you even choose and rigging the coin, maybe your hypotheses need to be UDT-style universes that take the "you" program as an input. Learning that Omega is messing with you could be done by updating to place a higher probability on some universes rather than others.
On the one hand, hypotheses about the entire universe are much more extravagant than hypotheses about a single binary variable. But I don't think it's crazy that in order to imagine Omega messing with you, you can't think about the coin in perfect isolation.
That’s a reasonable idea, but if you work through it, you will nonetheless find that if your belief state is represented by a single probability distribution, then when you compute the expected value of Heads and Tails, one (or both) of them will exceed 0.45.
Of course, you could say that your probability distribution is about Omega’s policy, and that it puts all (or most of) its mass on “Heads iff I bet Tails”, and then you can say that your payoff, instead of being an expected value, is some richer pairing of your policy with Omega’s policy. This is roughly the orthodox LessWrong way of handling Newcomblike problems, which is also a departure from orthodox Bayesian decision theory. But then there is still no consistent way to formalize your epistemic state about the coin’s outcome as a probability distribution about the coin’s outcome, even though there clearly is a rational epistemic state to have about the coin’s outcome.
The domain of a probability distribution can be any set whatsoever, and the distribution itself is a measure over a sigma-algebra of subsets of that domain. It's common for it to be a nice set like the real numbers or some subset thereof, or maybe tuples of real numbers, but nothing requires that.
In particular, it is perfectly valid for one element of that domain to be a world-model such as "the coin is biased to always land opposite of your bet (and if you didn't make a bet it will always land tails)", as well as many others. The domain is even allowed to include elements such as "none of the above". The distribution can then express your credence that the real world behaves according to one of those world-models (or none of the above).
The usual axioms of probability then constrain your rational credences in various outcomes, such as credence in winning certain bets.
Here is a very simple probability distribution over world models that makes it rational to roll the d20:
P("the coin is biased to always land opposite of your bet (and if you didn't make a bet it will always land tails)") = 1. There is no need for any other elements in the domain, and this is allowed by the axioms of a probability space and definition of distribution.
It is clear that this satisfies the condition of coin bias > 0.05, since the coin's outcome is deterministic. Consequently P(winning | bet Heads) = P(winning | bet Tails) = 0, but P(winning | roll d20) = 0.45 (by definition of fair d20 and application of probability rules).
I see what you mean, but then the question becomes: what is the formal foundation for what counts as a “world model”, beyond "it's a string in the English language that I can use to make predictions"? That is exactly the question my thinking about formal epistemology is working toward answering.
(The orthodox Bayesian answer is that a “world model” is nothing other than a single joint probability distribution about all conceivable variables at once, and that any outer “hyperprior” such as the one you named can be “marginalized away”. So no matter how many variables you include, this forces you to, in particular, act in accordance with a joint probability distribution about states and outcomes. Therefore, what I have shown is that the orthodox Bayesian view of “world models” forces you into an irrational choice in some Newcomblike problems.)
But then there is still no consistent way to formalize your epistemic state about the coin’s outcome as a probability distribution about the coin’s outcome, even though there clearly is a rational epistemic state to have about the coin’s outcome.
If there's a natural way to condense your information about the coin into a single probability, you can do it just as well starting from a distribution over UDT-style universes. Like if you think it should be 50/50 because that's your best guess if you forget the information about your precise action, you can give yourself a low-information distribution over policies and then marginalize over it to see what happens to the coin.
I don't know what you mean by "actual probability distributions". The more I read here, the less I understand what you mean by it.
Why are you only allowing distribution of that extremely specific form? Nothing in the usual definitions of a probability distribution require that.
I think you may have a very much more restrictive view on what sorts of things are allowed to be "probability distributions" than almost everyone else I've ever conversed with on the subject.
My view on what counts as an "actual probability distribution" is the Kolmogorov axioms, which have been standard across all of mathematics worldwide since the 1950s.
I think what you say might still be true, in that most rationalists who are not mathematicians use the term "probability distribution" to refer to a whole arsenal of conditional probability distributions about various topics, without really considering whether these assemble into a single joint probability distribution about any well-defined outcome space and/or state space.
In a way, I see what I am doing here as proposing a formal foundation for epistemological content that is already widely handled informally (perhaps with an unwarranted sense that the informal handling already has formal foundations in mathematical probability theory).
Could you make it work using a conditional probability? i.e. P(heads | do("bet tails")) = 0, P(tails | do("bet heads")) = 1 . I saw in the post that you said that there was no distribution that could express this, but I am unsure why conditionals cannot be used.
Conditional distributions are perfectly valid beliefs in my sense (see Definition 8), but conditional distributions are not actually probability distributions about anything — rather, a conditional distribution of
This paucity of probability alone for epistemology is also what inspired Judea Pearl to invent Structural Causal Models, the source of the do-notation you used. Pearl’s Causal Hierarchy also makes the point that ordinary probability distributions are inadequate to express causal beliefs.
The distribution JBlack proposes is not over the outcome or bias. It is over possible coin behaviours. If you want an explicit set, you could use this one:
({Coin chosen adversarially} union
{Coin has bias
You seem to be saying that if you define your belief to be a probability distribution over an explicit state space, that then you can't believe anything that involves information beyond that state space. That seems tautological to the point of uselessness.
Complex beliefs are common. It is nice when we can simplify them to a distribution over state spaces, but if you want to use examples that involve Omega-style self-reference, you are unlikely to be able to simplify them easily. (And if you consider the example of solomonoff induction, then the probability distribution is over all halting Turing machines.)
“Coin chosen adversarially” is not just a more complex state space or a hierarchical hyperprior, it is an exit from probability entirely. (It is “demonic nondeterminism”, which is a form of uncertainty that is not expressible via probability.)
Functional Decision Theory and Logical Decision Theory, although obviously the right sort of direction, have never been given proper formal definitions. I believe that part of the reason for this is that to do so requires a fundamentally nonprobabilistic notion of belief state. For example, Logical Inductors, which were a step toward this, have a notion of belief state which is nonprobabilistic (I think it is a kind of partial prevision, which assigns to some gambles a price, without demanding these prices be complete or obey the laws of probability proper).
"Coin chosen adversarially" is an element of a state space. The agent assigns a probability to it being in different possible worlds.
JBlack and I do not seem to have made arguments that depend on the agent's decision theory. (For all we've said, it might be ignoring the world model and acting randomly.) We have only discussed probability distributions that admit representing the Omega situation you brought up.
You would likely want to have your agent able to consider multiple different possible worlds, and evaluate the probability of being in any one of them. The typical highly general formalization of this is Solomonoff induction.
Even if you use the full machinery of Solomonoff induction, or any other process for Bayesian inference over large hypothesis classes, there is still a marginal probability distribution for the outcome of the coin, and therefore the d20 bet remains always strictly dominated by the bet on Heads, or on Tails, or both, regardless of how large the joint distribution was.
You have a few options for where to sample the predictions.
I think the best one is p(outcome | prospective agent action). i.e. consider what would happen if you take an action.
You could compute p(outcome), implicitly assuming that the agent's actions will not impact the outcome it is about to observe. This will often be false.
You could also compute p(outcome) assuming that the agent follows it's existing policy. This kind of self-prediction seems likely to lead to some strange situations, and also assumes that the policy has already decided on a course of action before we sample the prediction about the outcome (in which case why are we calculating probabilities of outcomes, if they can't inform the prediction?)
I think that the purpose of modelling the world is mostly to answer questions of the form "what happens if I do X?", and that form 1 works best. If we use form 1, then there will be a distribution for p(Heads | Agent chooses to bet on heads) and p(Tails | Agent chooses to bet on tails). A truing machine is perfectly capable of outputting 0 for both of those (or any other arbitrary number). If the inference process has incorporated the fairness of the d20, it will return p(d20 > 12 | Agent bets on d20) = 8/20.
If you use 2, you end up with an agent that can't handle Newcomblike situations. More severely, you end up with an agent that does not know "If I drive into that building, something bad will happen", as it is ignoring its action in the prediction. 3 probably depends on how you handle the self reference.
The offer doesn't even need to be coming from Omega for the d20 to be a permissible choice
Did you mean to introduce Omega before this?
Richard Ngo challenged me to set a time box and write down as many of the most important features of my formal epistemology as I can in one sitting. Here goes.
Motivation: Where probability distributions fail...
...to express beliefs
...to make safety tradeoffs
Beliefs, according to davidad
Definition 1. Given a state space , we define probability space as the space of all conceivable probability distributions on .
Definition 2. A belief about ( ) is a functional
which is lower semicontinuous (meaning that ).
Following Richardson, we interpret as the level of inconsistency between the conceivable probability distribution and one's belief .
Slogan: When one has multiple beliefs at the same time, one's overall belief is simply the sum of its parts.
All other known notions of belief are full subcategories
Bayesian beliefs
Definition 3. If one has a prior (big if), then one's prior belief is
Bayesian updating
Definition 4. If one has a belief over a hypothesis space, and the total state space of the situation is the product of hypothesis space and data space , then one's prior belief about is the inverse-image functor
Definition 5. An observation is a closed subset .
Example 6. If one observes that , this is the subset .
Definition 7. The belief that corresponds to an observation is
Definition 8. If one has a conditional distribution , then one's conditional belief is
Example 9. If one has a likelihood function , then the previous definition applies with and :
Definition 10. A Bayesian reasoner with prior and likelihood function has belief state
Theorem 11. The two components of a Bayesian reasoner's initial belief state, only one of which is a probability distribution, simply sum when lifted into the belief space, forming the joint belief:
However, this would be meaningless unless we could recover the Bayesian update by also summing the observations in the belief space.
Theorem 12. Whenever the Bayesian posterior is well-defined, then
with a constant denoting the level of incompatibility of the observation with the Bayesian reasoner's initial belief state, namely the prior predictive surprisal, .
Notice that the orthodox Bayesian update's “normalization” to a full-mass posterior distribution silently subtracts away the constant — the amount of incompatibility between the observed data and the statistical model as a whole — which is Deborah Mayo's criticism of Bayesian epistemology in a nutshell.
Bayesian beliefs have fully trivial categorical structure (no Bayesian belief implies any other Bayesian belief except itself), so Bayesian beliefs about are a full subcategory of , but the content of this is just that is injective (its post-inverse is ).
Infra-Bayesian beliefs (Kosoy and Appel)
Definition 13 (Kosoy and Appel). A homogenous ultracontribution is a nonempty topologically-closed convex down-closed subset of subprobability space, .
Theorem 14. Homogenous ultracontributions (ordered by ) form a full subcategory of . Specifically, given , the corresponding belief is
and given a belief , the corresponding subset of subprobability space is
and . Furthermore, is a member of iff is convex in probability space, which every satisfies.
Definition 15 (Kosoy and Appel). A homogenous ultradistribution is a homogenous ultracontribution whose intersection with the full-mass face is nonempty ( ).
Of course, is also a full subcategory of , since is a full subcategory of .
MWER (Halpern and Leung)
Theorem 16. Homogenous ultradistributions are exactly isomorphic to the belief states of Halpern and Leung's MWER framework (Minimax Weighted Expected Regret), via
Theorem 17. The updating process of simply adding , which is equivalent to that of Theorem 12, is also equivalent to Halpern and Leung's prescribed belief-updating procedure on the domain of their belief states.
Note: via the transform, this updating process is also consonant with the famous Multiplicative Weight Update family of algorithms (although I am not yet confident about whether e.g. AdaBoost is literally a special case of it).
Probabilistic dependency graphs (Richardson and Halpern)
Richardson and Halpern's PDGs, a common generalization of Bayes nets and factor graphs, take their semantics in functionals , and these functionals are always lower semicontinuous (though Richardson does not explicitly prove this), so therefore every PDG denotes a belief in my sense. Furthermore, the combination of PDGs with overlapping variables denotes exactly the sum of their beliefs, so the semantics of PDGs is a monoidal functor from PDGs into beliefs.
For a wide parameter regime of PDGs ( on every edge, or equivalently ) — assumed by some but not all PDG theorems [correction due to Richardson himself in the comments] — the beliefs they denote are convex, and thus immediately satisfy the exp-convexity criterion to be transformed into homogenous ultracontributions. However, even these PDGs do not typically denote homogenous ultradistributions, unless their semantics are “normalized” (by subtracting ).
PDGs can also express independence beliefs, which are non-convex while remaining lower semicontinuous.
Credal sets (Cozman)
Definition 18 (Cozman). A credal set is a nonempty closed convex set of probability distributions, .
Theorem 19. Credal sets about (ordered by ) form a full subcategory of . Specifically, given , the corresponding belief is
and given a belief , the corresponding credal set is
which is topologically closed by lower semicontinuity of .
Previsions (Goubault-Larrecq)
Definition 20 (Goubault-Larrecq). A gamble about is a Borel function . A prevision about is a functional such that and . An upper prevision is a prevision which is sub-additive: . A continuous upper prevision is an upper prevision which is Scott-continuous (for every directed family with least upper bound , ).
Definition 21. A coherent continuous upper prevision is a continuous upper prevision satisfying .
Theorem 22. Coherent continuous upper previsions about form a full subcategory of . Specifically, given the coherent continuous upper prevision , the corresponding belief is
and given a belief , the corresponding upper prevision is
The monad (Mio, Sarkis, and Vignudelli)
Definition 23 (Mio, Sarkis, and Vignudelli). The monad is defined on sets as the set of non-empty finitely-generated down-closed convex sets of subprobability distributions .
Mio, Sarkis, and Vignudelli prove that this monad is presented by the equational theory of semilattices equipped with finite probabilistic choice and , making in a strong sense the smallest semantic universe that can simultaneously interpret finite nondeterminism, finite probability, and partiality (and partiality is, in turn, needed to interpret either inconsistent beliefs or nonterminating probabilistic programs).
Theorem 24. ordered by inclusion is a full subcategory of , assuming is a Polish space. Specifically, this condition implies that finitely-generated convex sets are topologically closed, which makes a full subcategory of .