"Coin chosen adversarially" is an element of a state space. The agent assigns a probability to it being in different possible worlds.
JBlack and I do not seem to have made arguments that depend on the agent's decision theory. (For all we've said, it might be ignoring the world model and acting randomly.) We have only discussed probability distributions that admit representing the Omega situation you brought up.
You would likely want to have your agent able to consider multiple different possible worlds, and evaluate the probability of being in any one of them. The typical highly general formalization of this is Solomonoff induction.
The distribution JBlack proposes is not over the outcome or bias. It is over possible coin behaviours. If you want an explicit set, you could use this one:
({Coin chosen adversarially} union
{Coin has bias
You seem to be saying that if you define your belief to be a probability distribution over an explicit state space, that then you can't believe anything that involves information beyond that state space. That seems tautological to the point of uselessness.
Complex beliefs are common. It is nice when we can simpli...
In the 3blue1brown subreddit, LLM posts about the Riemann hypothesis and other crackpot-bait make up about half of the posts there. (It is a low traffic subreddit.)
I think that DRY still matters. Code repetition usually means that a change in one place will need to be reflected in the copies. You (or the LLM) needs to be aware of this implicit constraint, in order to ensure that the changes are made in all the necessary places. With enough intelligence & time/compute, you can handle this kind of extra difficulty, but it makes things harder. When enough of these extra difficulties stack up, eventually it will be enough to harm development speed/quality.
It does however require infinite compute. (I've also gotten the impression that all approximations of Solomonoff induction are also very compute intensive, unless you allow for very loose definitions of approximation.)
Perhaps more relevant to the quoted sentence is that it takes advantage of a very cleverly non-flat prior.
The figure you are referring to does not need to add up to 100%, since it is showing P[data | aliens] and P[data | no aliens].
P[data | aliens] and P[not data | aliens] need to add to 100%, but that is not on the graph.
As an extreme case where P[A | B] + P[A | C] != 1, consider A = coin did not land on its edge, B = the coin is ordinary, C = the coin is weighted to land heads twice as often as tails.
Then P[A | B] = 0.9999 and P[A | C] = 0.9999 would be reasonable values.
Here are the relevant quotes:
- Gather proposals for a hundred RCTs ...
- Randomly pick 5% of the proposed projects, fund them as written, and pay off the investors who correctly predicted what would happen.
- Take the other 95% of the proposed projects, give the investors their money back, and use the SWEET PREDICTIVE KNOWLEDGE [to take useful actions]
Other than the difference in the portion of the markets you run (1/20 vs 1/1000), this is equivalent.
(It does not discuss liquidity costs, just the the randomization as a way to avoid having to take many random actions.)
Just for the record, Dynomight proposed this back in 2022: https://dynomight.net/prediction-market-causation/#commit-to-randomization. (I assume that the idea has been around for longer.)
(Also I would phrase it as being able to use the same money to trade on all 1000 of the markets at once. I think that is equivalent to your free loan.)
I was able to deduce them by
making a scatter-plot of Colleen vs Liboulen's predictions. You can see that this plot has the points on a "flattened prism" in 3 directions, and manually count the shifts and see that each of the underlying components has 10 possible values.
Once you have that structure, you can pick out points on the extremes and use their slopes to calculate some of the relevant slopes. Finally, I brought in Bella's info and used that to work out the remaining stats. (I used chatGPT for some help throwing together some linear regressions
At this point I am throwing everything that I found in a linear regression, because I ran out of time. My pick is:
Candidate 11, with an estimated 0.91 chance of success.
Candidates 19 and 7 would be my next choices, with 0.87 and 0.85 estimated chances of success respectively.
If I had had more time to work on this, I would have like to look at:
A summary of some interesting results. I am leaving how I found some of this out for now, for brevity's sake.
I have manage to extract 6 integer variables that range from 1-10.
3 of them are from the components of (Coleen, Linestra, Liboulen, Bella), the other 3 are from (Fizz, Ister, Ziqual).
Each of them has a very similar histogram, sort of like a truncated normal distribution. A linear regression of them with Holly gives approximately 1 as their coefficient, except for 1 variable (which I am calling X2 for now) which has a coefficient of roughly -1.
All of
A few miscellaneous observations:
Ister, Ziqual and Fizz seem to have some pretty deterministic structure connecting them.
Ister always predicts an integer between 51 and 60 inclusive.
Ziqual's prediction is equal to (Ister - 50) * (Integer from 1 to 10) - (one of 0, 1). Multipliers in the 5 to 7 range are most common.
Fizz's prediction is less than or equal to (Ister's prediction + 10). Fizz's prediction is greater than or equal to 44.
Separately, a scatterplot of Liboulen and Colleen's predictions has a lot of structure: [Scatterplot removed since it se
Story of a mostly homeless guy who scammed Isaac King out of $300. Isaac sued in small claims court on principle, did all the things, and none of it mattered.
This link goes to Sarah's tweet, not to Isaac's story.
This is not what the article says. It says that BC is re-criminalizing hard drugs.
I am in BC, and have not heard anything about decriminalizing marijuana. I get the sense that it being legal is generally popular. Complaints about drug users are common here, but they are usually not talking about weed.
I would expect that player 2 would be able to win almost all of the time for most normal hash functions, as they could just play randomly for the first 39 turns, and then choose one of the 2^8 available moves. It is very unlikely that all of those hashes are zero. (For commonly used hashes, player 2 could just play randomly the whole game and likely win, since the hash of any value is almost never 0.)
In addition to the object level reasons mentioned by plex, misleading people about the nature of a benchmark is a problem because it is dishonest. Having an agreement to keep this secret indicates that the deception was more likely intentional on OpenAI's part.
Based on the quote from Kirkpatrick, It looks like a clear example of preference falsification, but I do not see any reason to believe that it is internalized preference falsification. Did I miss how the submissive apes were internalizing the preference to not mate? The sentence "This is an easy to understand example of an important general fact about humans: we can be threatened into internalized preference falsification, i.e. preference inversion." makes me think that you intended it as an example of primates internalizing a preference falsification. It ...
As an example of how Manifold reacted to a (crude) attempt at manipulation:
Dr P (a Manifold user) would create and bet yes on markets for "Will Trump be president on [some date]?" for various dates where there was no plausible way trump would be president. Other users quickly noticed and set up limit orders to capture this source of free money. Eventually Dr. P's bets were cancelled out quickly enough that they had little to no effect on the probability, and it became hard to find one of those bets profit from. Eventually Dr P gave up and their account bec...
One thing that I have seen on manifold is markets that will resolve at a random time, with a distribution such that at any time, their expected duration (from the current day, conditional on not having already resolved) is 6 months. They do not seem particularly common, and are not quite equivalent to a market with a deadline exactly 6 months in the future. (I can't seem to find the market.)
You have a few options for where to sample the predictions.
-
-
-
... (read more)I think the best one is p(outcome | prospective agent action). i.e. consider what would happen if you take an action.
You could compute p(outcome), implicitly assuming that the agent's actions will not impact the outcome it is about to observe. This will often be false.
You could also compute p(outcome) assuming that the agent follows it's existing policy. This kind of self-prediction seems likely to lead to some strange situations, and also assumes that the policy has already decided on a c