Working with John Wentworth is confusing and overwhelming at times.[1] The guy has a lot of information and models in his head and he doesn’t always make good guesses about where I’m starting from. That’s normal, but he makes things worse by using a strange, slapdash vocabulary he cobbled together from a half-a-dozen different disciplines. He also sometimes uses words in strange ways.
He and David Lorell have this sign on their office door:
If you knock on the door and ask for wise teachings, they bop you on the head with a roll of wrapping paper. It doesn’t hurt, but I haven’t seen anyone become obviously wiser as a result.
John also has a tendency to compress his statements so much that they don’t say very much, or they are wide open for misinterpretation.
For example:
“Even if you know everything about a system, there will still be uncertainty left.”[2]
Huh?
Explaining this is going to go better with some concrete examples.
Example: The Fish Pond
Imagine that you have two adjacent ponds. The first pond has a large but finite number of fish; the other is empty. The fish are very well-behaved. They do not breed, eat, poop, or die; they remain static in number and weight.
You’re interested in how much these fish weigh. You decide to model the distribution of fish-weights using a normal distribution, with the mean and the variance as latent variables.[3] Eyeballing the fish from the edge of the pond, you decide on your best initial guess for the mean and the variance of the weights of the fish. This is your prior.
Then, you pull the fish out of the first pond, one at a time. For each fish, you weigh it, you do a Bayesian update to your model of the fish-weights, and you throw that fish in the second pond. You’re done with it. As you work through the fish, your model fits reality better and better; the uncertainty about the mean and variance shrinks as you go.
You measure the last fish and update your model accordingly. You now know every weight in the system exactly. And yet, your model still has a mean and a variance in it, and you still have a posterior spread.
When John got to this point in the story, I said: “That’s dumb, you know everything. The only reason you still have uncertainty is because you are bull-headedly insisting on using a probabilistic model beyond its usefulness. The uncertainty is in the model, not the world. You should now throw the model away and just use the pond truth.”[4]
John half-agreed. The uncertainty is in the model, not the world. But should we really throw the model away? Probably not. We’ll come back to that after another example.
Example: Temperature of a gas
Here’s another example, parallel to the first: John’s favorite, a container of an ideal gas.
Here’s how he explained it in a long comment thread on a post of Richard Ngo’s. Richard had asked, about a different example, “What would resolve the uncertainty that remains after you have conditioned on the entire low-level state of the physical world?”
John replied, “nothing would resolve that uncertainty.” He went on:
The simplest concrete example I know is the Boltzmann distribution for an ideal gas - not the assorted things people say about the Boltzmann distribution, but the actual math, interpreted as Bayesian probability. The model has one latent variable, the temperature T, and says that all the particle velocities are normally distributed with mean zero and variance proportional to T. Then, just following the ordinary Bayesian math: in order to estimate T from all the particle velocities, I start with some prior P[T], calculate P[T|velocities] using Bayes' rule, and then for ~any reasonable prior I end up with a posterior distribution over T which is very tightly peaked around the average particle energy... but has nonzero spread. There's small but nonzero uncertainty in T given all of the particle velocities. And in this simple toy gas model, those particles are the whole world, there's nothing else to learn about which would further reduce my uncertainty in T.
When John and I tried to talk about this one in person I quickly got tangled up. He said, “If you know where every particle is and how fast it’s going, you know the total energy. Total energy is not the same as temperature, though.”
And then I blinked because I only got far enough in physics to learn that average kinetic energy and temperature are the same thing, they’re just a unit conversion, but not far enough to find out that was yet another lie promulgated by Big Pedagogy.
Later I talked to an LLM about it for a while and pierced the veil of the conspiracy. The equation I kinda sorta remembered relates temperature to the average energy according to the model distribution, not the average energy my particular particles have. This is the same gap as the mean parameter and the sample mean of the fish.
So then why bother?
So our model sucks, in some ways. It’s fuzzy where reality is both crisp and known.[5] Why are we bothering with this way of world-modeling? In other words, what’s the virtue in abstraction, when we have options?
It’s the compression, obviously. Why pay for full fidelity when a compact generator will do? Compression is often a good bargain: lots of efficiency without much loss of precision or predictiveness.[6]
And later, when we’re dealing with abstract concepts that are much harder to model than finite fish-ponds, we’re going to have to accept some fuzziness for other reasons anyway, so we might as well get used to having latents with uncertainty in them.
Coping with Uncertainty
Ngo’s post, linked above, argues that Bayesian reasoning seems busted to him, in part because its binary-valued propositions seem overly rigid. He proposes fuzzy truth-values as a partial fix. “I’d rather be vaguely right than precisely wrong.”
John countered that you don't have to leave or complicate the Bayesian frame. You're already working with probabilistic models; those models have latents; the values of those latents often stay uncertain even when you know everything about the physical system you’re modeling. We don’t need to insert the fuzziness Richard wanted into the proposition truth-values, it’s already there in the latents.
Even the meaning of a vague sentence can be treated as a latent. Pulling from another comment on Richard’s post, “humanity will be extinct in 100 years” doesn’t need a fuzzy truth-value if the meaning of the sentence is itself a construct in the listener’s model, carrying uncertainty that no amount of information will resolve.
I’m not sure Richard bought it, and it seems unlikely that my post here will convince him either. My main goal is to take a point of John’s that was buried in a comment thread and pull it up to a top-level post so it’s easier to refer to in the future.
How much uncertainty do we have to put up with?
So what did John actually mean, when he said that even if you know everything, there’s uncertainty?
Well, he wasn’t wrong, he just didn’t use quite enough words. (Story of my life, these days.)
I’d rephrase it as, “If you’re using a probabilistic model with latent variables in it, your model will still have remaining uncertainty even after you’ve seen the low-level state of the whole world.” I admit this is not as catchy. It no longer sounds like a “wise teaching,” it just sounds obviously true.
How much uncertainty are you left with? Well, it depends on how well your model fits the data, how much entropy the data has, how much data you’ve seen… all the usual stuff. But the point is, you never drive the uncertainty down to zero.
I asked in advance about “ragging on John” as a theme for my posts, and he was very enthusiastic. He loves criticism and disagreements. “Contra Wentworth is a very popular genre, go for it!” he said.
So I told Eliezer about this opening, and he got excited and immediately wanted to guess what John might have meant by that provocative statement. We quickly ruled out that it was some kind of Heisenberg uncertainty thing – my own first guess – and then Eliezer came up with the following three guesses.
“Logical uncertainty is when you know everything about the universe, and therefore you know that the number 2033 is written on a whiteboard, but you don't know if that number is prime or composite.”
“ Indexical uncertainty is when you're in a quantum universe or you're on a computer, and the computer program that is you is about to be cloned and one of you is going to see a coin come up heads, and one of you is going to see a coin come up tails, or equivalently, you're flipping a quantum coin that therefore branches the universe, and therefore even though you know everything about the universe's prior state, you don't know what you will see happen next. And that's like, where am I going to end up, or indexical uncertainty.”
“And the third one is metaphysical uncertainty, where you know where all the atoms are in the universe, and you're like why does this universe exist and not some other universes? Or which universes exist by how much? Or do any other universes exist? And possibly this form of uncertainty is reducible to logical uncertainty, but if so, I haven't done so yet.”
None of these were right.
Sometimes one of these guys can explain the other one to me, but not this time.
When John introduced this example, he grumbled to himself that a long-tailed distribution would be better, but carried on with a normal distribution after all for inscrutable John-reasons.
When I explained about the fish pond to Eliezer, he was grumpy about the same thing I was, and the same answer applied to his objection.
Eliezer: ”Sounds like a whack choice of prior, but I guess you could maybe know everything in the universe and then end up with uncertainty 'cause you have a whack choice of prior and non-normative uncertainty about the universe.”
Gretta: “What? No, it doesn't... Which prior you picked doesn't matter very much.”
Eliezer: “Once I know everything about all the fish in the pond my mean and variance should now be points. If I have instead chosen to update a Gaussian prior off of them, that's me being stupid.”
Gretta: “You can think of no reasons why you might want to model the fish pond using a probabilistic model?”
Eliezer: “Only to save on computing power....”
Gretta: “The probabilistic model is a nice compression of the data, is the point.”
And he agreed that none of his first three theories about John’s meaning turned out to be correct.
Working with John Wentworth is confusing and overwhelming at times.[1] The guy has a lot of information and models in his head and he doesn’t always make good guesses about where I’m starting from. That’s normal, but he makes things worse by using a strange, slapdash vocabulary he cobbled together from a half-a-dozen different disciplines. He also sometimes uses words in strange ways.
He and David Lorell have this sign on their office door:
If you knock on the door and ask for wise teachings, they bop you on the head with a roll of wrapping paper. It doesn’t hurt, but I haven’t seen anyone become obviously wiser as a result.
John also has a tendency to compress his statements so much that they don’t say very much, or they are wide open for misinterpretation.
For example:
“Even if you know everything about a system, there will still be uncertainty left.”[2]
Huh?
Explaining this is going to go better with some concrete examples.
Example: The Fish Pond
Imagine that you have two adjacent ponds. The first pond has a large but finite number of fish; the other is empty. The fish are very well-behaved. They do not breed, eat, poop, or die; they remain static in number and weight.
You’re interested in how much these fish weigh. You decide to model the distribution of fish-weights using a normal distribution, with the mean and the variance as latent variables.[3] Eyeballing the fish from the edge of the pond, you decide on your best initial guess for the mean and the variance of the weights of the fish. This is your prior.
Then, you pull the fish out of the first pond, one at a time. For each fish, you weigh it, you do a Bayesian update to your model of the fish-weights, and you throw that fish in the second pond. You’re done with it. As you work through the fish, your model fits reality better and better; the uncertainty about the mean and variance shrinks as you go.
You measure the last fish and update your model accordingly. You now know every weight in the system exactly. And yet, your model still has a mean and a variance in it, and you still have a posterior spread.
When John got to this point in the story, I said: “That’s dumb, you know everything. The only reason you still have uncertainty is because you are bull-headedly insisting on using a probabilistic model beyond its usefulness. The uncertainty is in the model, not the world. You should now throw the model away and just use the pond truth.”[4]
John half-agreed. The uncertainty is in the model, not the world. But should we really throw the model away? Probably not. We’ll come back to that after another example.
Example: Temperature of a gas
Here’s another example, parallel to the first: John’s favorite, a container of an ideal gas.
Here’s how he explained it in a long comment thread on a post of Richard Ngo’s. Richard had asked, about a different example, “What would resolve the uncertainty that remains after you have conditioned on the entire low-level state of the physical world?”
John replied, “nothing would resolve that uncertainty.” He went on:
When John and I tried to talk about this one in person I quickly got tangled up. He said, “If you know where every particle is and how fast it’s going, you know the total energy. Total energy is not the same as temperature, though.”
And then I blinked because I only got far enough in physics to learn that average kinetic energy and temperature are the same thing, they’re just a unit conversion, but not far enough to find out that was yet another lie promulgated by Big Pedagogy.
Later I talked to an LLM about it for a while and pierced the veil of the conspiracy. The equation I kinda sorta remembered relates temperature to the average energy according to the model distribution, not the average energy my particular particles have. This is the same gap as the mean parameter and the sample mean of the fish.
So then why bother?
So our model sucks, in some ways. It’s fuzzy where reality is both crisp and known.[5] Why are we bothering with this way of world-modeling? In other words, what’s the virtue in abstraction, when we have options?
It’s the compression, obviously. Why pay for full fidelity when a compact generator will do? Compression is often a good bargain: lots of efficiency without much loss of precision or predictiveness.[6]
And later, when we’re dealing with abstract concepts that are much harder to model than finite fish-ponds, we’re going to have to accept some fuzziness for other reasons anyway, so we might as well get used to having latents with uncertainty in them.
Coping with Uncertainty
Ngo’s post, linked above, argues that Bayesian reasoning seems busted to him, in part because its binary-valued propositions seem overly rigid. He proposes fuzzy truth-values as a partial fix. “I’d rather be vaguely right than precisely wrong.”
John countered that you don't have to leave or complicate the Bayesian frame. You're already working with probabilistic models; those models have latents; the values of those latents often stay uncertain even when you know everything about the physical system you’re modeling. We don’t need to insert the fuzziness Richard wanted into the proposition truth-values, it’s already there in the latents.
Even the meaning of a vague sentence can be treated as a latent. Pulling from another comment on Richard’s post, “humanity will be extinct in 100 years” doesn’t need a fuzzy truth-value if the meaning of the sentence is itself a construct in the listener’s model, carrying uncertainty that no amount of information will resolve.
I’m not sure Richard bought it, and it seems unlikely that my post here will convince him either. My main goal is to take a point of John’s that was buried in a comment thread and pull it up to a top-level post so it’s easier to refer to in the future.
How much uncertainty do we have to put up with?
So what did John actually mean, when he said that even if you know everything, there’s uncertainty?
Well, he wasn’t wrong, he just didn’t use quite enough words. (Story of my life, these days.)
I’d rephrase it as, “If you’re using a probabilistic model with latent variables in it, your model will still have remaining uncertainty even after you’ve seen the low-level state of the whole world.” I admit this is not as catchy. It no longer sounds like a “wise teaching,” it just sounds obviously true.
How much uncertainty are you left with? Well, it depends on how well your model fits the data, how much entropy the data has, how much data you’ve seen… all the usual stuff. But the point is, you never drive the uncertainty down to zero.
I asked in advance about “ragging on John” as a theme for my posts, and he was very enthusiastic. He loves criticism and disagreements. “Contra Wentworth is a very popular genre, go for it!” he said.
So I told Eliezer about this opening, and he got excited and immediately wanted to guess what John might have meant by that provocative statement. We quickly ruled out that it was some kind of Heisenberg uncertainty thing – my own first guess – and then Eliezer came up with the following three guesses.
None of these were right.
Sometimes one of these guys can explain the other one to me, but not this time.
When John introduced this example, he grumbled to himself that a long-tailed distribution would be better, but carried on with a normal distribution after all for inscrutable John-reasons.
I did not actually think of the pun at the time.
That’s weird. Reality isn’t usually crisp and known. Let’s not get used to it.
When I explained about the fish pond to Eliezer, he was grumpy about the same thing I was, and the same answer applied to his objection.
And he agreed that none of his first three theories about John’s meaning turned out to be correct.