The word "overconfident" seems overloaded. Here are some things I think that people sometimes mean when they say someone is overconfident:
How much does this overloading matter? I'm not sure, but one worry is that it allows people to score cheap rhetorical points by claiming someone else is overconfident when in practice they might mean something like "your probability distribution is wrong in some way". Beware of accusing someone of overconfidence without being more specific about what you mean.
In addition to your 1-6, I have also seen people use "overconfident" to mean something more like "behaving as though the process that generated a given probabilistic prediction was higher-quality (in terms of Brier score or the like) than it really is."
In prediction market terms: putting more money than you should into the market for a given outcome, as distinct from any particular fact about the probabilit(ies) implied by your stake in that market.
For example, suppose there is some forecaster who predicts on a wide range of topics. And their forecasts are generally great across most topics (low Brier score, etc.). But there's one particular topic area -- I dunno, let's say "east Asian politics" -- where they are a much worse predictor, with a Brier score near random guessing. Nonetheless, they go on making forecasts about east Asian politics alongside their forecasts on other topics, without noting the difference in any way.
I could easily imagine this forecaster getting accused of being "overconfident about east Asian politics." And if so, I would interpret the accusation to mean the thing I described in the first 2 paragraphs of this comment, rather than any of 1-6 in the OP.
Note that the objection here does not involve anything about the specific values of the forecaster's distributions for east Asian politics -- whether they are low or high, extreme or middling, flat or peaked, etc. This distinguishes it from all of 1-6 except for 4, and of course it's also unrelated to 4.
The objection here is not that the probabilities suffer from some specific, correctable error like being too high or extreme. Rather, the objection is that forecaster should not be reporting these probabilities at all; or that they should only report them alongside some sort of disclaimer; or that they should report them as part of a bundle where they have "lower weight" than other forecasts, if we're in a context like a prediction market where such a thing is possible.
Moore & Schatz (2017) made a similar point about different meanings of "overconfidence" in their paper The three faces of overconfidence. The abstract:
Overconfidence has been studied in 3 distinct ways. Overestimation is thinking that you are better than you are. Overplacement is the exaggerated belief that you are better than others. Overprecision is the excessive faith that you know the truth. These 3 forms of overconfidence manifest themselves under different conditions, have different causes, and have widely varying consequences. It is a mistake to treat them as if they were the same or to assume that they have the same psychological origins.
Though I do think that some of your 6 different meanings are different manifestations of the same underlying meaning.
Calling someone "overprecise" is saying that they should increase the entropy of their beliefs. In cases where there is a natural ignorance prior, it is claiming that their probability distribution should be closer to the ignorance prior. This could sometimes mean closer to 50-50 as in your point 1, e.g. the probability that the Yankees will win their next game. This could sometimes mean closer to 1/n as with some cases of your points 2 & 6, e.g. a 1/30 probability that the Yankees will win the next World Series (as they are 1 of 30 teams).
In cases where there isn't a natural ignorance prior, saying that someone should increase the entropy of their beliefs is often interpretable as a claim that they should put less probability on the possibilities that they view as most likely. This could sometimes look like your point 2, e.g. if they think DeSantis has a 20% chance of being US President in 2030, or like your point 6. It could sometimes look like widening their confidence interval for estimating some quantity.
When I accuse someone of overconfidence, I usually mean they're being too hedgehogy when they should be being more foxy.
On Plan A vs. Plan S
Plan A scales to top-expert-level AI in ~5 years, then to superintelligence once very high confidence in alignment is achieved; in the scenario this is 5 more years later, for a total of 10 years from the deal beginning until superintelligence. (More on the scaling strategy in https://ai-2040.com/supplements/capability-scaling-strategy.)
The version of Plan S that we’re most sympathetic to involves a halt on AI capabilities for a minimum of ~3-5 years with intent to eventually scale to top-expert-level AI and then superintelligence, but much slower than in Plan A.
(The below is all my own view, other authors may disagree on the details.)
I think Plan S is a big improvement over Plans D, C, or B but worse than Plan A.
The main reason that I prefer Plan A to Plan S is that I think the risk of the international slowdown/pause deal declining is very significant: i.e. either dissolving or having its effectiveness becoming highly impaired. I think that the risk is roughly 35-40% within 5 years, with high error bars (and with deal dissolution and deal impairment contributing roughly equally to that tootal). More on this in https://ai-2040.com/supplements/deal-decline.
Given a fixed amount of time bought, it’s better to scale capabilities as long as this can be done with high confidence in safety, in order to get useful work out of AIs and study AIs that are closer to being able to take over. If the deal dissolves 3 years into Plan A, you’ve used improve AIs to achieve lots of useful alignment research, decision-making/epistemics improvements, etc. If it dissolves 3 years into Plan S, you’re better off than in Plan D because of the increased time you’ve bought, but you’ve gotten much less useful work out of your AIs.
Another key questions is whether the chance of deal decline is higher in Plan A or S: my guess is that it’s higher in Plan S because better AIs allow you to develop technologies that stabilize the deal, though I’m not confident as the AI progress could be destabilizing in other ways; if we went all out Plan-D-style that would likely be more destabilizing to the deal than Plan A even pre-TED-AI, so faster isn’t always better.
The main upside of Plan S is that the initial pause phase is simpler than Plan A and thus harder to mess up; in particular, in Plan A the risk of catastrophe due to a misjudgment that led to scaling too fast is higher probability. If you want to eventually resume scaling, you will need to transition to a regime that can handle this. But the extra time you’ve bought and the slower pace at which you might scale should help with managing this risk.
If the risk of deal decline were very low, this would provide more reason to do Plan S instead of A and I’d think that they were pretty close in value with Plan S potentially being better. Even then, I’d think that Plan S should aim to scale to superintelligence eventually and probably within around 30-100 years, because there are other reasons to scale at some point besides deal decline: background risks such as pandemics and nuclear war, and covert projects.
(crossposted from https://x.com/eli_lifland/status/2075734827832401959 with minor edits)
I think it's much easier to achieve Plan S than Plan A.
Plan S is simple "Don't train new AI until there's very strong consensus it's a good idea."
Plan A requires continuously making nuanced judgment calls about what counts as safe, while generally maintaining momentum on training more powerful AIs (and leaving lots of dry tinder around). It's essentially an unsolved problem to make regulations careful enough to distinguish good vs bad safety cases, and I don't think the Plan A documents had particularly good ideas for now to do so.
I think Plan A basically only makes sense if we get pleasantly surprisingly good governance (which would be a marked departure from the governance we currently seem on track to have).
We've seen examples of Plan S for cloning/eugenics and (sorta) for nuclear power, so it's not like it's obviously intractable.
I realize Plan A / Plan S is a spectrum. But one of the main things I'd want to see to feel safer is interrupting the momentum of the AI labs and transitioning the world to "training a new AI is treated as a dangerous, careful endeavor." I think this requires multiple years (vs the approximately 1 year pause in Plan A).
Some things that'd update me include seeing a) a significant reduction in US gov corruption after the 2028 election, one way or another, b) seeing "AI for epistemics" actually begin to play out.
I don't think AI for epistemics is actually that bottlenecked on better AIs (community notes didn't require LLMs at all). I think it's more just "actually bothering to design social media and other infrastructure for epistemics at all.")
Plan S is better at averting permanent disempowerment (it enables more chances to take this specific problem seriously), even as perhaps it's modestly worse than Plan A at averting extinction. In the futures where Plan A doesn't lead to extinction, it still almost certainly ends in permanent disempowerment (the future of humanity only gets breadcrumbs of the reachable universe).
(My impression is that Yudkowsky/Soares expect extinction where I expect permanent disempowerment. And I expect permanent disempowerment to be likely averted in the futures where they expect extinction to be averted.)
I think Plan A is significantly worse than Plan S.
I am uncertain that even the best versions of the sorts of control/superficial alignment techniques portrayed in the plan will be sufficient to make sure nothing catastrophic happens, when "Top-Expert-Dominating AI" is on the table.
And it does not seem to me that every time, across years and dozens of companies, the best possible versions of said alignment techniques will be implemented. It seems very plausible that something of the flavor of Anthropic and OpenAI training against the CoT will happen, except much more dangerous, since "Top-Expert-Dominating AI".
I do not think that scaling to what Plan A portrays, will speed up the time to "alignment is solved" sufficiently to outweigh the risk; in my experience, the sorts of things current AIs are good at, or are on track to become good at, are not the limiting factor in alignment research, especially the sorts of exceptional alignment research that bring us substantially closer to "alignment is solved".
I agree that deal breakdown is a big problem, one that people should be working on post-pause. I do not think getting closer to the edge of existentially dangerous capabilities helps with that.[1]
If anything, it might make a deal weaker, since it goes past a natural Schelling fence.
[crossposted from EA Forum]
Reflecting a little on my shortform from a few years ago, I think I wasn't ambitious enough in trying to actually move this forward.
I want there to be an org that does "human challenge"-style RCTs across lots of important questions that are extremely hard to get at otherwise, including (top 2 are repeated from previous shortform):
Edited to add: I no longer think "human challenge" is really the best way to refer to this idea (see comment that convinced me); I mean to say something like "large scale RCTs of important things on volunteers who sign up on an app to randomly try or not try an intervention." I'm open to suggestions on succinct ways to refer to this.
I'd be very excited about such an org existing. I think it could even grow to become an effective megaproject, pending further analysis on how much it could increase wisdom relative to power. But, I don't think it's a good personal fit for me to found given my current interests and skills.
However, I think I could plausibly provide some useful advice/help to anyone who is interested in founding a many-domain human-challenge org. If you are interested in founding such an org or know someone who might be and want my advice, let me know. (I will also be linking this shortform to some people who might be able to help set this up.)
--
Some further inspiration I'm drawing on to be excited about this org:
Votes/considerations on why this is a good or bad idea are also appreciated!
I'm confused why these would be described as "challenge" RCTs, and worry that the term will create broader confusion in the movement to support challenge trials for disease. In the usual clinical context, the word "challenge" in "human challenge trial" refers to the step of introducing the "challenge" of a bad thing (e.g., an infectious agent) to the subject, to see if the treatment protects them from it. I don't know what a "challenge" trial testing the effects of veganism looks like?
(I'm generally positive on the idea of trialing more things; my confusion+comment is just restricted to the naming being proposed here.)
Thanks, I agree with this and it's probably not good branding anyway.
I was thinking the "challenge" was just doing the intervention (e.g. being vegan), but agree that the framing is confusing since it refers to something different in the clinical context. I will edit my shortforms to reflect this updated view.
Just made a bet with Jeremy Gillen that may be of interest to some LWers, would be curious for opinions:
[cross-posting from blog]
I made a spreadsheet for forecasting the 10th/50th/90th percentile for how you think GPT-4.5 will do on various benchmarks (given 6 months after the release to allow for actually being applied to the benchmark, and post-training enhancements). Copy it here to register your forecasts.
If you’d prefer, you could also use it to predict for GPT-5, or for the state-of-the-art at a certain time e.g. end of 2024 (my predictions would be pretty similar for GPT-4.5, and end of 2024).
You can see my forecasts made with ~2 hours of total effort on Feb 17 in this sheet; I won’t describe them further here in order to avoid anchoring.
There might be a similar tournament on Metaculus soon, but not sure on the timeline for that (and spreadsheet might be lower friction). If someone wants to take the time to make a form for predicting, tracking and resolving the forecasts, be my guest and I’ll link it here.
(epistemic status: exploratory)
I think more people into LessWrong in high school - college should consider trying Battlecode. It's somewhat similar to The Darwin Game which was pretty popular on here and I think generally the type of people who like LessWrong will both enjoy and be good at Battlecode. (edited to add: A short description of Battlecode is that you write a bot to beat other bots at a turn-based strategy game. Each unit executes its own code so communication/coordination is often one of the most interesting parts.)
I did it with friends for 6 years (junior year of high school - end of undergrad), and I think it at least helped me gain legible expertise in strategizing and coding quickly, but plausibly also helped me pick up skills in these areas as well as teamwork.
If any students are interested (I believe PhD students can qualify as well but may not be worth their time), there's still 2/3 weeks left in this year's game which is plenty of time. If you're curious to learn more about my experiences with Battlecode, see the README and postmortem here.
Feel free to comment or DM me if you have any questions.