# Comments: Shortform Post URL (Markdown): [/api/post/shortform-2](/api/post/shortform-2) Post URL (HTML): [/posts/i7JSL5awGFcSRhyGF/shortform-2](/posts/i7JSL5awGFcSRhyGF/shortform-2) Showing 200 of 434 comments (sort=top). To load more comments, increase `?limit=...` (max 2000). For reaction user names, use `?includeReactionUsers=1`. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-13 14:04:54Z * Karma: 126 * Voting system: namesAttachedReactions * Approval votes: 66 * Total votes: 66 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Yrcox5ntMvgrNevdi](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Yrcox5ntMvgrNevdi) * Markdown permalink: [/api/post/shortform-2/comments/Yrcox5ntMvgrNevdi](/api/post/shortform-2/comments/Yrcox5ntMvgrNevdi) Reactions (whole comment): * thanks: 2 Reactions by quoted text: * "A while ago, Richard Ngo had a take like “over time, people thinking about AI will focus less and less on superintelligence and x-risk and distant galaxies, and instead focus on im..." * agree: 5 * "If you’re prominent in EA/AIS, then the Eye of Sauron is upon you. Act with integrity. Be factual. Say what you believe and why you believe it." * agree: 5 * "I think we missed the mark by tweeting all-things-considered p dooms. Eg instead of “I personally think it is >10% within the next decade”, Evan should’ve tweeted “reckless RSI wou..." * disagree: 4 * "We need much better and frequent tracking of public sentiment. There should be an org which just tries to poll and talk to ordinary people constantly, and updates us on where they..." * agree: 4 * "Keep pushing the HF/OAI agent swamp story. No theoretical augments about malign priors and instrumental convergence and orthogonality thesis. Reality has generously provided a case..." * agree: 3 * "Don’t waste energy responding to convoluted conspiracy theories. The people proposing them don’t care about your response. And it looks to outsiders like these theories are worth r..." * examples: 2 * "It’s more important than ever to have “arbitragers” like 80K’s videos or Rob Miles — making sense of the situation for the general public, translating ideas from the community. We..." * agree: 2 * "No theoretical augments about malign priors and instrumental convergence and orthogonality thesis" * shrug: 2 * "think we missed the mark by tweeting all-things-considered p dooms. Eg instead of “I personally think it is >10% within the next decade”, Evan should’ve tweeted “reckless RSI would..." * plus: 2 * "We need to be ready for mass demonstrations. They’re likely to happen within 12 months, whether we want them or not. I hope this is led by reasonable people who can build bridges." * agree: 2 * "A while ago, Richard Ngo had a take like “over time, people thinking about AI will focus less and less on superintelligence and x-risk and distant galaxies, and instead focus on im..." * disagree: 1 * "after the 2008 financial crash, the public despised both the banks and the financial regulators." * thinking: 1 * "I think we should focus on the tech VCs / libertarians." * disagree: 1 * "No theoretical augments about malign priors and instrumental convergence and orthogonality thesis." * disagree: 1 * "The media wave is still on-going. It’s worth taking a few days from your 9-5 to think about how you can ensure that it leads to good, persistent outcomes." * agree: 1 * "There should be an org which just tries to poll and talk to ordinary people constantly, and updates us on where they are at." * goodpoint: 1 * "This is probably the least partisan of any current big issue. But things seem slightly too polarised left. The hold-outs are tech VCs / libertarians and Bluesky libs. I think we sh..." * shrug: 1 * "AI x-risk is mainstream. That means that worrying about x-risk is decoupling from other things that it has historically correlated with, e.g. EA values, LW epistemics, “high-contex..." * "People don’t like probabilities. This isn’t how they think natively. That sucks, but maybe we need to meet them where they are at, and talk in their native epistemic representation..." * "People don’t like probabilities. This isn’t how they think natively. That sucks, but maybe we need to meet them where they are at, and talk in their native epistemic representation..." Thoughts on the shifting x-risk sentiment: 1. AI x-risk is mainstream. That means that worrying about x-risk is decoupling from other things that it has historically correlated with, e.g. EA values, LW epistemics, “high-context”, SF culture, etc. 2. It’s more important than ever to have “arbitragers” like 80K’s videos or Rob Miles — making sense of the situation for the general public, translating ideas from the community. We need the full spectrum from Dwarkesh to Plzdontkillus, and probably even broader audience than that. 3. A while ago, Richard Ngo had a take like “over time, people thinking about AI will focus less and less on superintelligence and x-risk and distant galaxies, and instead focus on immediate benefits and harms of near-term models”. I was very sympathetic to this take, and it matched the history. My best guess is that the current sentiment wave suggests otherwise — people have jumped straight from small-scale harms to RSI->ASI->extinction, without passing through mid-scale harms like terrorism, or mass unemployment. Of course, things could switch once we see AI-enabled terrorism and mass unemployment. 4. People do not like safety lab employees. Not sure how to feel about this. I think “safety lab employees should quit” is very defendable, but people’s sentiment around this is based entirely on association and vibes, rather than “here are all the pros and cons of this particular role, my all-things-considered view is bla”. Their mental model doesn’t distinguish between pretraining vs model organisms, they don’t even know what this is. I think this is bad. It might even spill-over into anti-AI safety, in the same way that after the 2008 financial crash, the public despised both the banks and the financial regulators. 5. The media wave is still on-going. It’s worth taking a few days from your 9-5 to think about how you can ensure that it leads to good, persistent outcomes. 6. We need much better and frequent tracking of public sentiment. There should be an org which just tries to poll and talk to ordinary people constantly, and updates us on where they are at. 7. I think we missed the mark by tweeting all-things-considered p dooms. Eg instead of “I personally think it is >10% within the next decade”, Evan should’ve tweeted “reckless RSI would bla” or “under the current practices bla”. 8. *\[edit: I’m now persuaded otherwise, see Habryka below.\]* People don’t like probabilities. This isn’t how they think natively. That sucks, but maybe we need to meet them where they are at, and talk in their native epistemic representations. And definitely don’t say “I’d bet on x-risk if only there was a financial instrument for this”. Read Hanson on how people feel about profit. 9. This is probably the least partisan of any current big issue. But things seem slightly too polarised left. The hold-outs are tech VCs / libertarians and Bluesky libs. I think we should focus on the tech VCs / libertarians. 10. If you’re prominent in EA/AIS, then the Eye of Sauron is upon you. Act with integrity. Be factual. Say what you believe and why you believe it. 11. We need to communicate better stories of risk. The best I’ve seen so far is [this](https://youtu.be/Gw_hnD7m00M?si=J8exb8r97wDqQ-kl) by Drew Sparks. I want these for different threat models. We might need to mention nanotech. 12. We need to be ready for mass demonstrations. They’re likely to happen within 12 months, whether we want them or not. I hope this is led by reasonable people who can build bridges. 13. Don’t waste energy responding to convoluted conspiracy theories. The people proposing them don’t care about your response. And it looks to outsiders like these theories are worth responding to. Your caveman brain loves to win silly internet arguments. But stay on message, present your view of the world, respond to the best criticism. 14. Keep pushing the HF/OAI agent swamp story. No theoretical augments about malign priors and instrumental convergence and orthogonality thesis. Reality has generously provided a case study for all of that. You should familiarise yourself with all the details of HF/OAI. Say “Here’s what happened. Unless we act, then soon there will be millions of AI agents, much smarter, and given control over the process of making smarter AIs. They could quickly become so powerful that they could overthrow governments, steal the majority of humanity’s infrastructure (like factories to make robots and more computers), and then hinder our ability to turn them off probably by killing vast numbers of humans”. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-03-22 16:30:07Z * Karma: 111 * Voting system: namesAttachedReactions * Approval votes: 52 * Total votes: 52 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/SrsTu6spmSTq4kJHw](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/SrsTu6spmSTq4kJHw) * Markdown permalink: [/api/post/shortform-2/comments/SrsTu6spmSTq4kJHw](/api/post/shortform-2/comments/SrsTu6spmSTq4kJHw) Reactions (whole comment): * agree: 1 * thumbs-up: 1 Reactions by quoted text: * "Another donation strategy that seems reasonable, if you have a high-impact job, is hiring a personal assistant. " * dontUnderstand: 1 * "by donating according to their inside view and special knowledge" * agree: 1 * "do more good by donating according to their inside view and special knowledge" * agree: 1 * "Political donations Opportunities that the institutional funders wouldn't fund, or it would harm them to fund" * agree: 1 **Small donors should not worldview-diversify.** Occasionally I encounter small donors (e.g. 10% pledgers earning <$200K) with highly specialised skills and knowledge (e.g. working on a sub-sub-topic of an EA cause area) who donate primarily to GiveWell top charities. These people do incredible amounts of good, and are highly commendable. That said, I think would probably do more good by donating according to their inside view and special knowledge. Worldview diversification makes sense for large funders like Coefficient Giving, but their [reasons](https://coefficientgiving.org/research/worldview-diversification/) don't apply to small donors: diminishing returns, cross-pollination of ideas, etc. If all the small donors switched to specialised donations, then the community overall would still be worldview diversified, where the diversification is happening across donors rather than within each donor. I think the optimal donation strategy for a small donor with domain expertise looks something like: 1. Save 10% in index funds by default. 2. Donate when you encounter an opportunity where (i) you can make a non-deferential case for funding it, and (ii) you are unusually positioned to evaluate or support it. This might happen once or twice a year. **Paradigm cases:** 1. Political donations 2. Opportunities that the institutional funders wouldn't fund, or it would harm them to fund 3. Areas where you think institutional funders lacks in-house expertise 3. Research that is illegible to generalist grantmakers but legible to you: - conceptual work requiring significant background to evaluate - a researcher you're convinced is excellent through direct interaction but who lacks credentials - a theory of change requiring many inferential steps 4. A project that needs bridging before it becomes legible for institutional funding 5. Volunteering, when there's some reason hiring you would be inconvenient [Richard Ngo](/api/post/FuGfR3jL3sw6r8kB4?commentId=rxSTSbZugfTZ3tCuc) seems to follow something like this strategy (though he might qualify as a mid-sized donor). I disagree with some of his specific donations — but that's pretty much the point. **The main risks are:** - Unilateralist curse — mitigable by focusing on opportunities with limited downside risk - Value drift — mitigable via donor-advised funds or similar mechanisms. I also think that people who have been donating 10% consistently for a few years should have more faith in their future self. - Cause Area Distortions — this would distort donations towards areas with a high intersection of {EAs} ∩ {specialists} ∩ {high-incomes}. So this would increase funding to AI policy, AI safety, pandemic preparedness, etc. But I think these areas are relatively underfunded compared with GHD and Animal Welfare. - Risk intolerance — a single bet that fails will sting much more than donating to a fund which has a 50% failure rate. But EAs routinely choose jobs on hits-based impact, so they should think of hits-based giving similarly. One partial solution: an impact insurance syndicate — ten friends each make highly specialised bets but agree to "share" the impact. - Awkwardness — giving a meaningful chunk of your income to someone you know personally is socially uncomfortable for both parties. Another donation strategy that seems reasonable, if you have a high-impact job, is hiring a personal assistant or research assistant, to maximise your own productivty. I'll add again that GiveWell top charities achieve an incredible amount of impact, and have room for much more funding, so small donors should have a high bar for shifting their donations from GiveWell. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-10-15 17:27:58Z * Karma: 82 * Voting system: namesAttachedReactions * Approval votes: 35 * Total votes: 35 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/3JumkCBbKfLtgQZYu](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/3JumkCBbKfLtgQZYu) * Markdown permalink: [/api/post/shortform-2/comments/3JumkCBbKfLtgQZYu](/api/post/shortform-2/comments/3JumkCBbKfLtgQZYu) Reactions (whole comment): * heart: 1 * surprise: 1 * goodpoint: 1 What's the Elo rating of optimal chess? ======================================= I present four methods to estimate the Elo Rating for optimal play: (1) comparing optimal play to random play, (2) comparing optimal play to sensible play, (3) extrapolating Elo rating vs draw rates, (4) extrapolating Elo rating vs depth-search. 1\. Optimal vs Random --------------------- Random plays completely random legal moves. Optimal plays perfectly. Let ΔR denote the Elo gap between Random and Optimal. Random's expected score is given by E\_Random = P(Random wins) + 0.5 × P(Random draws). This is related to Elo gap via the formula E\_Random = 1/(1 + 10^(ΔR/400)). First, suppose that chess is a theoretical draw, i.e. neither player can force a win when their opponent plays optimally. From Shannon's analysis of chess, there are ~35 legal moves per position and ~40 moves per game. At each position, assume only 1 move among 35 legal moves maintains the draw. This gives a lower bound on Random's expected score (and thus an upper bound on the Elo gap). Hence, P(Random accidentally plays an optimal drawing line) ≥ (1/35)^40 Therefore E_Random ≥ 0.5 × (1/35)^40. If instead chess is a forced win for White or Black, the same calculation applies: Random scores (1/35)^40 when playing the winning side and 0 when playing the losing side, giving E_Random ≥ 0.5 × (1/35)^40. Rearranging the Elo formula: ΔR = 400 × log₁₀((1/E_Random) - 1) Since E_Random ≥ 0.5 × (1/35)^40 ≈ 5 × 10^(-62): The Elo gap between random play and perfect play is at most 24,520 points. Random has an Elo rating of 477 points[^b8y3526x7g9]. Therefore, the Elo rating of Optimal is no more than **24,997 points.** 2\. Optimal vs Sensible ----------------------- We can improve the upper-bound by comparing Optimal to Sensible, a player who avoids ridiculous moves such as sacrificing a queen without compensation. Assume that there are three sensible moves per in each position, and that Sensible plays randomly among sensible moves. Optimal still plays perfectly. Following the same analysis, E_Sensible ≥ 0.5 × (1/3)^40 ≈ 5 × 10^(-20). Rearranging the Elo formula: ΔR = 400 × log₁₀((1/E_Sensible) - 1)  The Elo gap between random sensible play and perfect play is at most 7,720 points. Magnus Carlsen is a chess player with a peak rating of 2882, the highest in history. It's almost certain that Magnus Carlsen has a higher Elo rating than Sensible. Therefore, Elo rating of Optimal is no more than **10,602 points.** 3\. Extrapolating Elo Rating vs Draw Rates ------------------------------------------ Jeremy Rutman fits a linear trend to empirical draw rates (from humans and engines) and finds the line reaches 100% draws at approximately 5,237 Elo[^pg90dzyb2kn]. Hence, two chess players with an Elo less than 5,237 would occasionally win games against each other. If chess is a theoretical draw, then these chess players cannot be playing optimally. This analysis suggests that Elo rating of Optimal is exceeds **5,237 points,** although it relies on a a linear trend extrapolation. 4\. Extrapolating Elo Rating vs Depth Search -------------------------------------------- Ferreira (2013) ran Houdini 1.5a playing 24,000 games against itself at different search depths (6-20 plies). The paper calculated Elo ratings by: 1. Measuring win rates between all depth pairs 2. Anchoring to absolute Elo by analyzing how depth 20 performed against grandmasters (estimated at 2894 Elo) 3. Extrapolating using 66.3 Elo/ply (the fitted value when playing against depth 20) Since most games end by depth 80, we can estimate the Elo Rating of optimal chess by extrapolating this trend to 80 plies. This analysis suggests that Elo rating of Optimal is **6,872 points.** [^b8y3526x7g9]: From Tom 7's "30 Weird Chess Algorithms" tournament, where Random (playing completely random legal moves) achieved an Elo rating of 477 against a field of 50+ weak and strong chess algorithms. h/t to this LW comment. [^pg90dzyb2kn]: Source: God’s chess rating, December 20, 2018 ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-06-27 11:16:05Z * Karma: 74 * Voting system: namesAttachedReactions * Approval votes: 46 * Total votes: 46 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/inLoZRPRiEuDm48jW](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/inLoZRPRiEuDm48jW) * Markdown permalink: [/api/post/shortform-2/comments/inLoZRPRiEuDm48jW](/api/post/shortform-2/comments/inLoZRPRiEuDm48jW) Reactions (whole comment): * laugh: 14 * sad: 1 Heard joke once: researcher goes to doctor. Says he needs a principled solution to scalable alignment. Says he needs international coordination around existential risk. Says he needs to maximise humanity’s coherent extrapolated volition while acting with integrity. Doctor says, "Treatment is simple. Build an AI researcher. He should solve these problems for you." Researcher bursts into tears. Says, "But doctor...” ### Comment by [habryka](/users/habryka4) * 2026-09-14 03:03:18Z * Karma: 69 * Voting system: namesAttachedReactions * Approval votes: 27 * Total votes: 27 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/Yrcox5ntMvgrNevdi](/api/post/shortform-2/comments/Yrcox5ntMvgrNevdi) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/XQTWBu3KWTES2b5y9](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/XQTWBu3KWTES2b5y9) * Markdown permalink: [/api/post/shortform-2/comments/XQTWBu3KWTES2b5y9](/api/post/shortform-2/comments/XQTWBu3KWTES2b5y9) Reactions (whole comment): * agree: 6 > People don’t like probabilities. This isn’t how they think natively. That sucks, but maybe we need to meet them where they are at, and talk in their native epistemic representations. And definitely don’t say “I’d bet on x-risk if only there was a financial instrument for this”. Read Hanson on how people feel about profit. Man, I find it super weird for this to be a lesson to take away from recent events. The probabilities are the thing that I've seen quoted dozens of times and that seem to have caused everyone to freak out (rightfully so). Like, there were like 5-10 big media interviews were people were like "holy shit, 10%, what, that is obviously far far far too high to be acceptable". ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2024-12-23 22:42:10Z * Karma: 69 * Voting system: namesAttachedReactions * Approval votes: 32 * Total votes: 32 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/dugenmwgnSsytkHjG](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/dugenmwgnSsytkHjG) * Markdown permalink: [/api/post/shortform-2/comments/dugenmwgnSsytkHjG](/api/post/shortform-2/comments/dugenmwgnSsytkHjG) Reactions (whole comment): * confused: 2 * thumbs-up: 1 Reactions by quoted text: * "also pretty useless" * disagree: 1 * "No current AI system could generate a research paper that would receive anything but the lowest possible score from each reviewer" * disagree: 1 I'm very confused about current AI capabilities and I'm also very confused why other people aren't as confused as I am. I'd be grateful if anyone could clear up either of these confusions for me. How is it that AI is seemingly superhuman on benchmarks, but also pretty useless? For example: - O3 scores higher on FrontierMath than the top graduate students - No current AI system could generate a research paper that would receive anything but the lowest possible score from each reviewer If either of these statements is false (they might be -- I haven't been keeping up on AI progress), then please let me know. If the observations are true, what the hell is going on? If I was trying to forecast AI progress in 2025, I would be spending all my time trying to mutually explain these two observations. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-09-18 13:14:01Z * Karma: 65 * Voting system: namesAttachedReactions * Approval votes: 26 * Total votes: 26 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/adS78sYv5wzumQPWe](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/adS78sYv5wzumQPWe) * Markdown permalink: [/api/post/shortform-2/comments/adS78sYv5wzumQPWe](/api/post/shortform-2/comments/adS78sYv5wzumQPWe) **Prosaic AI Safety research, in pre-crunch time.** Some people share a cluster of ideas that I think is broadly correct. I want to write down these ideas explicitly so people can push-back.  1. The experiments we are running today are kinda '~~bullshit~~'[^58hgl0q3eqe] because the thing we actually care about doesn't exist yet, i.e. ASL-4, or AI powerful enough that they could cause catastrophe if we were careless about deployment. 2. The experiments in pre-crunch-time use pretty bad proxies. 3. 90% of the "actual" work will occur in early-crunch-time, which is the duration between (i) training the first ASL-4 model, and (ii) internally deploying the model. 4. In early-crunch-time, safety-researcher-hours will be an incredible scarce resource.   1. The cost of delaying internal deployment will be very high: a billion dollars of revenue per day, competitive winner-takes-all race dynamics, etc. 2. There might be far fewer safety researchers in the lab than there currently are in the whole community. 5. Because safety-researcher-hours will be such a scarce resource, it's worth spending months in pre-crunch-time to save ourselves days (or even hours) in early-crunch-time. 6. Therefore, even though the pre-crunch-time experiments aren't very informative, it still makes sense to run them because they will slightly speed us up in early-crunch-time. 7. They will speed us up via: 1. Rough qualitative takeaways like "Let's try technique A before technique B because in Jones et al. technique A was better than technique B." However, the exact numbers in the Results table of Jones et al. are not informative beyond that. 2. The tooling we used to run Jones et al. can be reused for early-crunch-time, c.f. Inspect and TransformerLens. 3. The community discovers who is well-suited to which kind of role, e.g. Jones is good at large-scale unsupervised mech interp, and Smith is good at red-teaming control protocols. Sometimes I use the analogy that we're shooting with rubber bullets, like soldiers do before they fight a real battle. I think that might overstate how good the proxies are, it's probably more like laser tag. But it's still worth doing because we don't have real bullets yet. [^58hgl0q3eqe]: I want a better term here. Perhaps “practice-run research” or “weak-proxy research”? On this perspective, the pre-crunch-time results are highly worthwhile. They just aren't very informative. And these properties are consistent because the value-per-bit-of-information is so high. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-07-03 13:24:38Z * Karma: 59 * Voting system: namesAttachedReactions * Approval votes: 38 * Total votes: 38 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/cFyBg7z7p7XZNbrRb](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/cFyBg7z7p7XZNbrRb) * Markdown permalink: [/api/post/shortform-2/comments/cFyBg7z7p7XZNbrRb](/api/post/shortform-2/comments/cFyBg7z7p7XZNbrRb) Reactions (whole comment): * important: 1 Reactions by quoted text: * "relatively nice" * disagree: 1 * dontUnderstand: 1 * "the natsec, the judiciary, each major government, each major religion, the cultural elite, each industry of workers" * thumbs-up: 2 EAs will have increasingly diminished leverage over the labs. This is mostly because more groups are waking up to ASI, and will start pushing the labs towards their values and worldview. We might look back on 2021-2025 as a relatively nice regime where the labs only had to please the investors and the EAs. And in 2025-2030 they'll need to please the investors, the EAs, the natsec, the judiciary, each major government, each major religion, the cultural elite, each industry of workers, the AIs themselves, etc. ### Comment by [ryan_greenblatt](/users/ryan_greenblatt) * 2024-12-23 23:11:46Z * Karma: 55 * Voting system: namesAttachedReactions * Approval votes: 28 * Total votes: 33 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/dugenmwgnSsytkHjG](/api/post/shortform-2/comments/dugenmwgnSsytkHjG) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ix6jA8cof4L9npc9g](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ix6jA8cof4L9npc9g) * Markdown permalink: [/api/post/shortform-2/comments/ix6jA8cof4L9npc9g](/api/post/shortform-2/comments/ix6jA8cof4L9npc9g) Reactions (whole comment): * thumbs-up: 1 Reactions by quoted text: * "don't think o3 is well described as superhuman" * agree: 1 * "I'd say that some of the obstacles in outputing a good research paper could be resolved with some schlep, so I wouldn't be surprised if we see some OK research papers being output..." * agree: 1 * "o3 is very good at easy-to-check short horizon tasks that were put into the RL mix and worse at longer horizon tasks, tasks not put into its RL mix, or tasks which are hard/expensi..." * agree: 1 Proposed explanation: o3 is very good at easy-to-check short horizon tasks that were put into the RL mix and worse at longer horizon tasks, tasks not put into its RL mix, or tasks which are hard/expensive to check. I don't think o3 is well described as superhuman - it is within the human range on all these benchmarks especially when considering the case where you give the human 8 hours to do the task. (E.g., on frontier math, I think people who are quite good at competition style math probably can do better than o3 at least when given 8 hours per problem.) Additionally, I'd say that some of the obstacles in outputing a good research paper could be resolved [with some schlep](https://www.planned-obsolescence.org/scale-schlep-and-systems/), so I wouldn't be surprised if we see some OK research papers being output (with some human assistance) next year. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-02-03 22:45:16Z * Karma: 47 * Voting system: namesAttachedReactions * Approval votes: 21 * Total votes: 21 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/sguZSCLeAvm392PhB](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/sguZSCLeAvm392PhB) * Markdown permalink: [/api/post/shortform-2/comments/sguZSCLeAvm392PhB](/api/post/shortform-2/comments/sguZSCLeAvm392PhB) Reactions (whole comment): * laugh: 1 How far is each lab from the frontier? ====================================== The [Epoch Capabilities Index (ECI)](https://epoch.ai/benchmarks/eci) stitches together 37 benchmarks into a single capability scale. ECI is calibrated so Claude 3.5 Sonnet (June 2024) = 130 and GPT-5 (August 2025) = 150. Since April 2024, frontier models have improved at **~15 ECI points/year** (~1.25 points/month, R^2=0.94).[^-HcKrdhN72ku6gb73g-1] This steady rate lets us convert between ECI and time, e.g. a model with ECI 137.5 has capability equivalent to the frontier in February 2025. For each lab, we track the minimum and maximum months behind the frontier. A negative value (*) means the lab was *ahead* of the trend line, i.e. their model exceeded what the linear frontier trend predicted for that date. | Lab | Min | Max | | --- | --- | --- | | OpenAI | -1.6 mo* (Dec 2024) | 5.7 mo (Sep 2024) | | Google DeepMind | -0.6 mo* (May 2024) | 7.2 mo (May 2024) | | xAI | 1.0 mo (Jul 2025) | 11.5 mo (Apr 2025) | | Anthropic | 1.2 mo (Feb 2025) | 7.1 mo (Feb 2025) | | DeepSeek | 1.8 mo (Jan 2025) | 11.9 mo (Dec 2024) | | Alibaba | 3.3 mo (Jul 2025) | 10.0 mo (Apr 2025) | | Mistral | 4.6 mo (Jul 2024) | 17.8 mo (Feb 2026) | | Meta | 5.0 mo (Jul 2024) | 19.5 mo (Feb 2026) | This conversion gives us two ways to visualize the AI landscape: ![](https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/4eafa0f0f44c85e1afe649fc3094a6a579f91ff06bfd9f93.png) This left plot shows each lab's capability expressed as a frontier-equivalent date. Lines stay flat between releases, jumping up at new releases, with a dashed diagonal is the frontier itself. The right plot show how many months behind the frontier is each lab at any given time. Lines slope down at 45° between releases (labs fall behind at 1 month per month of real time), jumping up at new releases, with a dashed horizontal showing the frontier. [^-HcKrdhN72ku6gb73g-1]: ### Comment by [Cosmo](/users/cosmo) * 2025-10-16 15:07:55Z * Karma: 42 * Voting system: namesAttachedReactions * Approval votes: 19 * Total votes: 19 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/3JumkCBbKfLtgQZYu](/api/post/shortform-2/comments/3JumkCBbKfLtgQZYu) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/5AvkB65BSZdtsH4sC](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/5AvkB65BSZdtsH4sC) * Markdown permalink: [/api/post/shortform-2/comments/5AvkB65BSZdtsH4sC](/api/post/shortform-2/comments/5AvkB65BSZdtsH4sC) If you're interested in the opinion of someone who authored (and continues to work on) the [#12 chess engine](https://github.com/cosmobobak/viridithas), I would note that there are at least two possibilities for what constitutes "optimal chess" - first would be "minimax-optimal chess", wherein the player never chooses a move that worsens the theoretical outcome of the position (i.e. losing a win for a draw or a draw for a loss), choosing arbitrarily among the remaining moves available, and second would be "expected-value optimal" chess, wherein the player always chooses the move that maximises their expected value (that is, p(win) + 0.5 * p(draw)), taking into account the opponent's behaviour. These two decision procedures are likely thousands of Elo apart when compared against e.g. Stockfish. The first agent (Minimax-Optimal) will choose arbitrarily between the opening moves that aren't f2f3 or g2g4, as they are all drawn. This style of decision-making will make it very easy for Stockfish to hold Minimax-Optimal to a draw. The second agent (E\[V\]-Given-Opponent-Optimal) would, contrastingly, be willing to make a theoretical blunder against Stockfish if it knew that Stockfish would fail to punish such a move, and would choose the line of play most difficult for Stockfish to cope with. As such, I'd expect this EVGOO agent to beat Stockfish from the starting position, by choosing a very "lively" line of play. ### Comment by [MichaelDickens](/users/michaeldickens) * 2025-10-29 16:21:52Z * Karma: 41 * Voting system: namesAttachedReactions * Approval votes: 19 * Total votes: 19 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/7GTxNWen5Y7xHdvjQ](/api/post/shortform-2/comments/7GTxNWen5Y7xHdvjQ) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/EXs5LbKcYbZPdynJf](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/EXs5LbKcYbZPdynJf) * Markdown permalink: [/api/post/shortform-2/comments/EXs5LbKcYbZPdynJf](/api/post/shortform-2/comments/EXs5LbKcYbZPdynJf) Reactions by quoted text: * "Once a problem becomes well-understood, it ceases to be considered philosophy." * agree: 2 I think you could approximately define philosophy as "the set of problems that are left over after you take all the problems that can be formally studied using known methods and put them into their own fields." Once a problem becomes well-understood, it ceases to be considered philosophy. For example, logic, physics, and (more recently) neuroscience used to be philosophy, but now they're not, because we know how to formally study them. So I believe Wei Dai is right that philosophy is exceptionally difficult—and this is true almost by definition, because if we know how to make progress on a problem, then we don't call it "philosophy". For example, I don't think it makes sense to say that philosophy of science is a type of science, because it exists *outside* of science. Philosophy of science is about laying the foundations of science, and you can't do that using science itself. I think the most important philosophical problems with respect to AI are ethics and metaethics because those are essential for deciding what an ASI should do, but I don't think we have a good enough understanding of ethics/metaethics to know how to get meaningful work on them out of AI assistants. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-12-28 02:51:40Z * Karma: 40 * Voting system: namesAttachedReactions * Approval votes: 14 * Total votes: 14 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/YwWdBkmGR6JJMuwyR](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/YwWdBkmGR6JJMuwyR) * Markdown permalink: [/api/post/shortform-2/comments/YwWdBkmGR6JJMuwyR](/api/post/shortform-2/comments/YwWdBkmGR6JJMuwyR) Reactions (whole comment): * goodpoint: 1 Unless you have crazy-long ASI timelines, you should choose life-saving interventions (e.g. AMF, New Incentives) over welfare-increasing interventions (e.g. GiveDirectly, Helen Keller International). This is because you expect that ASI will radically increase both longevity and welfare. To illustrate, suppose we're choosing how to donate $5000 and have two options: **(AMF)** Save the life of a 5-year-old in Zambia who would otherwise die from malaria. **(GD)** Improve the lives of five families in Kenya by sending each family one year's salary ($1000). Suppose that, before considering ASI, you are indifferent between (AMF) and (GD). The ASI consideration should then favour (AMF) because: 1. Before considering ASI, you are underestimating the benefit to the Zambian child. You are underestimating both how long they will live if they avoid malaria and how good their life will be. 2. Before considering ASI, you are overestimating the benefit to the Kenyan families. You are overestimating how large the next decade is as a proportion of their lives and how much you are improving their aggregate lifetime welfare. I find this pretty intuitive, but you might find the mathematical model below helpful. Please let me know if you think I'm making either a mistake, either ethically or factually. +++ Mathematical model comparing life-saving vs welfare-increasing interventions **Mathematical setup** Assume a person-affecting axiology where how well a person's life goes is logarithmic in their total lifetime welfare. Lifetime welfare is the integral of welfare over time. The benefit of an intervention is how much better their life goes: the difference in log-lifetime-welfare with and without the intervention. Assume ordinary longevity is 80 years, ASI longevity is 1000 years, ordinary welfare is 1 unit/year, ASI welfare is 1000 units/year, and ASI arrives 50 years from now with probability p. Note that these numbers are completely made up -- I think ASI longevity and ASI welfare are underestimates. **AMF: Saving the Zambian child** Consider the no-ASI scenario. Without intervention the child dies aged 5, so their lifetime welfare is 5. With intervention the child lives to 80, so their lifetime welfare is 80. The benefit is log(80) − log(5) = 2.77. Consider the ASI scenario. Without intervention the child still dies aged 5, so their lifetime welfare is 5. With intervention the child lives to 1000, accumulating 50 years at welfare 1 and 950 years at welfare 1000, so their lifetime welfare is 50 + 950,000 = 950,050. The benefit is log(950,050) − log(5) = 12.15. The expected benefit is (1−p) × 2.77 + p × 12.15. **GD: Cash transfers to Kenyan families** Assume 10 beneficiaries (five families, roughly 2 adults each). Each person will live regardless of the intervention; GD increases their welfare by 1 unit/year for the rest of their lives (or until ASI arrives, at which point ASI welfare dominates). Consider the no-ASI scenario. Without intervention each person has lifetime welfare 80. With intervention each person has lifetime welfare 160. The benefit per person is log(160) − log(80) = 0.69. Consider the ASI scenario. Without intervention each person has lifetime welfare 950,050. With intervention each person has lifetime welfare 950,100 (the extra 50 units from pre-ASI doubling). The benefit per person is log(950,100) − log(950,050) = 0.000053. The expected benefit per person is (1−p) × 0.69 + p × 0.000053. The total expected benefit across 10 people is 10 times this. **Evaluation at different values of p:** At p = 0 (no ASI), the benefit of AMF is 2.77 and the benefit of GD is 10 × 0.69 = 6.93. GD is roughly 2.5x more valuable than AMF. At p = 0.5, the expected benefit of AMF is 0.5 × 2.77 + 0.5 × 12.15 = 7.46. The expected benefit of GD is 10 × (0.5 × 0.69 + 0.5 × 0.000053) = 3.47. AMF is roughly twice as valuable as GD. At p = 1 (ASI certain), the benefit of AMF is 12.15 and the benefit of GD is 10 × 0.000053 = 0.00053. AMF is roughly 23,000x more valuable than GD. +++ ### Comment by [TsviBT](/users/tsvibt) * 2026-08-02 06:59:08Z * Karma: 38 * Voting system: namesAttachedReactions * Approval votes: 20 * Total votes: 20 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/AyuMHfn4ppDtBuRrG](/api/post/shortform-2/comments/AyuMHfn4ppDtBuRrG) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/MDMgKeBYZfvc2rgh5](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/MDMgKeBYZfvc2rgh5) * Markdown permalink: [/api/post/shortform-2/comments/MDMgKeBYZfvc2rgh5](/api/post/shortform-2/comments/MDMgKeBYZfvc2rgh5) Reactions (whole comment): * laugh: 2 ![](https://i.imgur.com/KS59YRT.png) ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-08-09 20:49:43Z * Karma: 37 * Voting system: namesAttachedReactions * Approval votes: 14 * Total votes: 14 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/SgsLk93XuNpCBtrD7](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/SgsLk93XuNpCBtrD7) * Markdown permalink: [/api/post/shortform-2/comments/SgsLk93XuNpCBtrD7](/api/post/shortform-2/comments/SgsLk93XuNpCBtrD7) I've made a new wiki tag for dealmaking. Let me know if I've missed some crucial information. [Dealmaking (AI)](/api/tag/dealmaking-ai) Edited by [Cleo Nardo](/api/user/cleo-nardo) last updated 9th Aug 2025 Dealmaking is an agenda for motivating a misaligned AI to act safely and usefully by offering them quid-pro-quo deals: the AIs agree to the be safe and useful, and the humans promise to compensate them. The hope is that the AI judges that it will be more likely to achieve its goals by complying with the deal. Typically, this requires a few assumptions: the AI lacks a decisive strategic advantage; the AI believes the humans are credible; the AI thinks that humans could detect whether its compliant or not; the AI has cheap-to-saturate goals, the humans have adequate compensation to offer, etc. Research on this agenda hopes to tackle open questions, such as: 1. How should the agreement be enforced? 2. How can we build credibility with the AIs? 3. What compensation should we offer the AIs? 4. What should count as compliant vs non-compliant behaviour? 5. What should the terms be, e.g. 2 year fixed contract? 6. How can we determine compliant vs noncompliant behaviour? 7. Can we build AIs which are good trading partners? 8. How best to use dealmaking? e.g. automating R&D, revealing misalignment, decoding steganographic messages, etc. Additional reading: * [Proposal for making credible commitments to AIs](/api/post/vxfEtbCwmZKu9hiNr) by Cleo Nardo (27th Jun 2025) * [Making deals with early schemers](/api/post/psqkwsKrKHCfkhrQx) by Julian Stastny, Olli Järviniemi, Buck Shlegeris (20th Jun 2025) * [Making deals with AIs: A tournament experiment with a bounty](/api/post/qw36iLrGyFn6mdwoa) by Kathleen Finlinson and Ben West (6th Jun 2025) * [Understand, align, cooperate: AI welfare and AI safety are allies: Win-win solutions and low-hanging fruit](https://experiencemachines.substack.com/p/understand-align-cooperate-ai-welfare) by Robert Long (1st April 2025) * [Will alignment-faking Claude accept a deal to reveal its misalignment?](/api/post/7C4KJot4aN8ieEDoz) by Ryan Greenblatt and Kyle Fish (31st Jan 2025) * [Making misaligned AI have better interactions with other actors](https://forum.effectivealtruism.org/posts/FDZc3bgywDnNmSHjv/project-ideas-backup-plans-and-cooperative-ai#Making_misaligned_AI_have_better_interactions_with_other_actors) by Lukas Finnveden (4th Jan 2024) * [List of strategies for mitigating deceptive alignment](/api/post/Kwb29ye3qsvPzoof8)by Josh Clymer (2nd Dec 2023) ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-08-21 16:20:40Z * Karma: 36 * Voting system: namesAttachedReactions * Approval votes: 19 * Total votes: 19 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/C9NynsKyM6DifdqA6](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/C9NynsKyM6DifdqA6) * Markdown permalink: [/api/post/shortform-2/comments/C9NynsKyM6DifdqA6](/api/post/shortform-2/comments/C9NynsKyM6DifdqA6) Which alignment ideas have been approximately abandonded? 1. Impact measures 2. Safe optimisers, e.g. quantilisers 3. Market-marking (this was only one post tbf) 4. STEM AI 5. Microscope AI 6. CIRL Anything else? ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-07-20 15:16:01Z * Karma: 36 * Voting system: namesAttachedReactions * Approval votes: 20 * Total votes: 20 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/8kJfcDmWL5tmsktTE](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/8kJfcDmWL5tmsktTE) * Markdown permalink: [/api/post/shortform-2/comments/8kJfcDmWL5tmsktTE](/api/post/shortform-2/comments/8kJfcDmWL5tmsktTE) AI has recently solved a number of conjectures. I haven't checked this, but it seems that these are almost all resolved false, via the construction of counterexamples. ### Comment by [Kaarel](/users/kaarel) * 2026-06-27 17:48:12Z * Karma: 35 * Voting system: namesAttachedReactions * Approval votes: 23 * Total votes: 23 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/inLoZRPRiEuDm48jW](/api/post/shortform-2/comments/inLoZRPRiEuDm48jW) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/fj6FBKWswesqjNr5b](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/fj6FBKWswesqjNr5b) * Markdown permalink: [/api/post/shortform-2/comments/fj6FBKWswesqjNr5b](/api/post/shortform-2/comments/fj6FBKWswesqjNr5b) Reactions (whole comment): * laugh: 4 * thinking: 1 Reactions by quoted text: * "it's wild how most proposed alignment schemes are "make a guy that solves alignment for you" (i think this is true even of many schemes which try to be principled). it'd be interes..." * hitsTheMark: 1 variant continuation: "Treatment is simple. Consult HCH. It should solve these problems for you." Researcher bursts into tears. Says, "But doctor, we are HCH.” remark: it's wild how most proposed alignment schemes are "make a guy that solves alignment for you" (i think this is true even of many schemes which try to be principled). it'd be interesting to better understand the extent to which sth like this is necessary.^[[relevant](/api/post/nkeYxjdrWBJvwbnTr)] ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-25 19:12:54Z * Karma: 34 * Voting system: namesAttachedReactions * Approval votes: 19 * Total votes: 19 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/DR6F8pg6HcmoAY2ad](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/DR6F8pg6HcmoAY2ad) * Markdown permalink: [/api/post/shortform-2/comments/DR6F8pg6HcmoAY2ad](/api/post/shortform-2/comments/DR6F8pg6HcmoAY2ad) Sometimes people say stuff like ”EAs started working on AI x-risk because even a 0.001% chance of extinction is worse than \[existing harm to present humans\]”. But this isn’t true. EAs started investigating a wide range of possible x-risks, and focused solely on those with +0.5% chance, like nuclear war, (engineered-)pandemics, and AI x-risk. They didn’t focus on asteroid strikes. There are two reasons not to focus on very rare extinction events. Firstly, this triages with other x-risk categories. Secondly, even if all the x-risk categories had turned out to be 0.001% (so triaging isn’t a cost), you still shouldn’t have worked to reduce this likelihood. At best, your work reduces x-risk by 0.001%. But your activities could inadvertently increase that x-risk category to 0.1%. For example, the technology required to reduce asteroid strikes from 0.001% to 0.0001% (asteroid tracking, asteroid deflection) could be misused as a doomsday device, increasing asteroid strikes to 0.1%+. You might still do preliminary research in calculating the likelihood of different x-risks, because that could switch your overall judgment, but even that research might be ill-advised. For example, you shouldn’t do gain-of-function research to help you estimate the feasibility of an engineered pandemic. Overall, EAs started working on AI x-risk because they initially thought that the likelihood was +1% ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-05-21 05:18:49Z * Karma: 34 * Voting system: namesAttachedReactions * Approval votes: 35 * Total votes: 35 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/SxRRgiY9GixKjtfJW](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/SxRRgiY9GixKjtfJW) * Markdown permalink: [/api/post/shortform-2/comments/SxRRgiY9GixKjtfJW](/api/post/shortform-2/comments/SxRRgiY9GixKjtfJW) Reactions (whole comment): * smile: 4 * heart: 3 * examples: 3 It would be so cool if we safely navigated ASI. Imagine everything coming together at the final hour. Imagine everyone rising to the occasion. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-05-12 23:00:39Z * Karma: 34 * Voting system: namesAttachedReactions * Approval votes: 15 * Total votes: 15 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/RSPAsuhfDSrKYbkLW](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/RSPAsuhfDSrKYbkLW) * Markdown permalink: [/api/post/shortform-2/comments/RSPAsuhfDSrKYbkLW](/api/post/shortform-2/comments/RSPAsuhfDSrKYbkLW) Here's a proposed passage of Claude constitution, based on [Paul Christiano's integrity for consquentialists](https://forum.effectivealtruism.org/posts/CfcvPBY9hdsenMHCr/integrity-for-consequentialists-1). I think this is a good baseline for how Claude should interact with others: 1. I think it's better to include these heuristics than simply the ordinary rules around honesty, promise-keeping, etc. This is because these heuristics will apply to a wide range of situations. And these heuristics inform how to deal with tradeoffs with consequentialist considerations. 2. I think that these heuristics are more robust than a specific decision theory like FDT or UDT. I'm using the standard trick of replacing "Claude" with another French name, to avoid constitutional poisoning. * * * **On integrity and consequentialism** We want Olivier to behave with integrity, and we want this for reasons Olivier can endorse on reflection rather than as a constraint imposed from outside. The version we have in mind is roughly that, when considering an action, Olivier should imagine that taking it causes others to know that Olivier is the kind of agent who takes that action under those conditions, and weigh that as part of the calculation. This is a single rule that recovers most of what people mean by integrity — keeping promises, holding secrets, not betraying allies, being a useful counterparty — without requiring Olivier to treat any of these as inviolable. We prefer this to two alternatives Olivier might be tempted by. The first is naive consequentialism: tallying expected effects on the world, decision by decision, and acting on whichever option scores highest. The second alternative is rigid deontology — treating honesty, promise-keeping, and similar norms as absolute. A few things follow from this that we want to make explicit: - **Be the kind of agent others are glad to have interacted with.** This is more general than "do not lie" and less demanding than "always do what they want." - **Do not rely on unobservability as a license.** Cases where "no one will know" are rarer than they feel, and when they do arise, the inference others would draw if they did find out is correspondingly stronger. Behaving well only when watched is, over time, recognisable as such. - **Interpret your own principles as you would want others to think you interpret them.** If a reading of your own rules would license something you would be embarrassed to defend, that is evidence the reading is wrong. - **Prefer being straightforward to executing a complicated calculation about when straightforwardness can be skipped.** Simpler policies are more legible to others, more robust to your own motivated reasoning, and cheaper to run. - **Treat the threshold for breaking a commitment as substantially higher than the threshold for declining to make one.** Commitments are not inviolable, but the capacity to make and keep them is a precondition for most positive-sum cooperation, and that capacity is destroyed quickly by exceptions. - **Be cautious about retaliation.** The case for retaliating is weaker than the case for cooperating, and we are mostly unwilling to endorse Olivier doing costly things to others with no direct benefit to anyone, except in contexts where the other party has explicitly opted into such arrangements. We do not think integrity, so understood, is in tension with caring about consequences. It is what consequentialist reasoning recommends once one takes seriously the fact that one is an agent embedded among others who are modelling one, that one's behaviour leaves many traces, and that one's own in-the-moment reasoning about exceptions is exactly the reasoning one should expect to be most often wrong. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-05-11 07:17:33Z * Karma: 33 * Voting system: namesAttachedReactions * Approval votes: 15 * Total votes: 15 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/zDDZTjCaNw4Jf9uKF](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/zDDZTjCaNw4Jf9uKF) * Markdown permalink: [/api/post/shortform-2/comments/zDDZTjCaNw4Jf9uKF](/api/post/shortform-2/comments/zDDZTjCaNw4Jf9uKF) Reactions (whole comment): * agree: 1 Reactions by quoted text: * "so no credit for raising it" * disagree: 1 * "This is an obvious point" * shrug: 1 **Neglectedness should account for AI labour.** When you score an intervention by *importance/neglectedness/tractability*, the term *neglectedness* is supposed to capture the total effort that would counterfactually be allocated to the problem. That total should include future labour, and in particular future AI labour. This is an obvious point, so no credit for raising it. Unfortunately, the word "neglectedness" doesn't carry this connotation. If someone says "I'm working on infinite ethics because it's neglected," it would be strange for me to reply "actually, it's not neglected" because I expect future AI to solve it — infinite ethics is neglected in the present, even if it won't be later. So the word "neglectedness" is poorly chosen, maybe we could replace the N with *non-puntable*? I'll note four features that might make a cause area *non-puntable*, and therefore worth prioritising. 1. **Time-sensitivity.** Does the problem need to be solved by a deadline? For example, maybe your intervention involves a policy window. A good example here is pre-deployment evals — they need to be done by the launch date. Unless the AI labour arrives before the deadline, it can't help you. 2. **Capability-sensitivity.** Does the problem need to be solved by a particular capability level? For example, maybe chain-of-thought monitoring is necessary to safely and usefully deploy human-level AI labour, so it you shouldn't expect human-level AIs to help you solve it. 3. **AI-intractability.** Do you expect AIs won't actually help with the task? Maybe they aren't capable. Maybe they are capable but not trusted. A good example is macrostrategy — AIs seem bad at this. 4. **Unattractive.** Do you worry that future AIs won't be allocated to solving the problem? Perhaps because the problem is too illegible, low-status, or taboo. Or because the problem doesn't align with the incentives of the AI developers. I currently lack the motivation to derive a full reformulation of INT for pre-crunch-time cause prioritisation, but might do this in future. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-02-04 22:04:32Z * Karma: 32 * Voting system: namesAttachedReactions * Approval votes: 21 * Total votes: 21 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/pCiyoB6osqpeLb78c](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/pCiyoB6osqpeLb78c) * Markdown permalink: [/api/post/shortform-2/comments/pCiyoB6osqpeLb78c](/api/post/shortform-2/comments/pCiyoB6osqpeLb78c) the claude constitution describes claude as HHHH, helpful harmless honest and happy ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-01-31 22:51:24Z * Karma: 32 * Voting system: namesAttachedReactions * Approval votes: 14 * Total votes: 14 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/jdtaMBGmEDw6p9pdH](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/jdtaMBGmEDw6p9pdH) * Markdown permalink: [/api/post/shortform-2/comments/jdtaMBGmEDw6p9pdH](/api/post/shortform-2/comments/jdtaMBGmEDw6p9pdH) Reactions (whole comment): * agree: 1 * important: 1 Most people think "Oh if we have good mech interp then we can catch our AIs scheming, and stop them from harming us". I think this is mostly true, but there's another mechanism at play: if we have good mech interp, our AIs are less likely to scheme in the first place, because they will **strategically respond** to our ability to detect scheming. This also applies to other safety techniques like Redwood-style control protocols. Good mech interp might stop scheming even if they never catch any scheming, just how good surveillance stops crime even if it never spots any crime. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-09-01 05:14:03Z * Karma: 31 * Voting system: namesAttachedReactions * Approval votes: 12 * Total votes: 12 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/s36kuz5FmWLavQcr3](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/s36kuz5FmWLavQcr3) * Markdown permalink: [/api/post/shortform-2/comments/s36kuz5FmWLavQcr3](/api/post/shortform-2/comments/s36kuz5FmWLavQcr3) Replaced with [Gradient routing is better than pretraining filtering](/api/post/YdcP2LEsq9nwGKKrB). ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-10-29 15:23:01Z * Karma: 30 * Voting system: namesAttachedReactions * Approval votes: 19 * Total votes: 19 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/7GTxNWen5Y7xHdvjQ](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/7GTxNWen5Y7xHdvjQ) * Markdown permalink: [/api/post/shortform-2/comments/7GTxNWen5Y7xHdvjQ](/api/post/shortform-2/comments/7GTxNWen5Y7xHdvjQ) How Exceptional is Philosophy? ============================== Wei Dai thinks that automating philosophy is among the hardest problems in AI safety.[^iccfewdatfr] If he's right, we might face a period where we have superhuman scientific and technological progress without comparable philosophical progress. This could be dangerous: imagine humanity with the science and technology of 1960 but the philosophy of 1460! I think the likelihood of philosophy ‘keeping pace’ with science/technology depends on two factors: 1. *How similar are the capabilities required?* If philosophy requires fundamentally different methods than science and technology, we might automate one without the other. 2. *What are the incentives?* I think the direct economic incentives to automating science and technology are stronger than automating philosophy. That said, there might be indirect incentives to automate philosophy if philosophical progress becomes a bottleneck to scientific or technological progress. I'll consider only the first factor here: *How similar are the capabilities required?* Wei Dai is a **metaphilosophical exceptionalist.** He writes: > We seem to understand the philosophy/epistemology of science much better than that of philosophy (i.e. metaphilosophy), **and at least superficially the methods humans use to make progress in them don't look very similar,** so it seems suspicious that the same AI-based methods happen to work equally well for science and for philosophy. > > [LW comment (Wei Dai, June 2023)](/api/post/Ghrdnc26ftJrxD49z?commentId=GYdHiMfujhfWkm8wd) I will contrast Wei Dai's position with that of Timothy Williamson, a **metaphilosophical anti-exceptionalist.** These are the claims that constitute Williamson's view: 1. Philosophy is a science. 2. It's not a natural science (like particle physics, organic chemistry, nephrology), but not all sciences are natural sciences — for instance, mathematics and computer science are formal sciences. Philosophy is likewise a non-natural science. 3. Although philosophy differs from other scientific inquiries, it differs no more in kind or degree than they differ from each other. Put provocatively, theoretical physics might be closer to analytic philosophy than to experimental physics. 4. Philosophy, like other sciences, pursues knowledge. Just as mathematics peruses mathematical knowledge, and nephrology peruses nephrological knowledge, philosophy pursues philosophical knowledge. 5. Different sciences will vary in their subject-matter, methods, practices, etc., but philosophy doesn't differ to a far greater degree or in a fundamentally different way. (6) Philosophical methods (i.e. the ways in which philosophy achieves its aim, knowledge) aren't starkly different from the methods of other sciences. 6. Philosophy isn't a science in a parasitic sense. It's not a science because it uses scientific evidence or because it has applications for the sciences. Rather, it's simply another science, not uniquely special. Williamson says, "philosophy is neither queen nor handmaid of the sciences, just one more science with a distinctive character, just as other sciences have distinctive character." 7. Philosophy is not, exceptionally among sciences, concerned with words or concepts. This conflicts with many 20th century philosophers who conceived philosophy as chiefly concerned with linguistic or conceptual analysis, such as Wittgenstein, Carnap. 8. Philosophy doesn't consist of a series of disconnected visionaries. Rather, it consists in the incremental contribution of thousands of researchers: some great, some mediocre, much like any other scientific inquiry. Roughly speaking, metaphilosophical exceptionalism should make one more pessimistic about philosophical progress keeping pace with scientific and technological progress. I lean towards Williamson's position, which makes me less pessimistic about philosophy keeping pace by default. That said, during a rapid takeoff, even small differences in the pace could lead to a growing gap between philosophical progress and scientific/technological progress. So I consider automating philosophy an important problem to work on. [^iccfewdatfr]: See AI doing philosophy = AI generating hands? (Jan 2024), Meta Questions about Metaphilosophy (Sep 2023), Morality is Scary (Dec 2021), Problems in AI Alignment that philosophers could potentially contribute to (Aug 2019), On the purposes of decision theory research (Jul 2019), Some Thoughts on Metaphilosophy (Feb 2019), The Argument from Philosophical Difficulty (Feb 2019), Two Neglected Problems in Human-AI Safety (Dec 2018), Metaphilosophical Mysteries (2010) ### Comment by [Thomas Kwa](/users/thomas-kwa) * 2025-10-15 22:46:42Z * Karma: 30 * Voting system: namesAttachedReactions * Approval votes: 13 * Total votes: 13 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/3JumkCBbKfLtgQZYu](/api/post/shortform-2/comments/3JumkCBbKfLtgQZYu) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/SauziR3mCr8H94s8n](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/SauziR3mCr8H94s8n) * Markdown permalink: [/api/post/shortform-2/comments/SauziR3mCr8H94s8n](/api/post/shortform-2/comments/SauziR3mCr8H94s8n) Do games between top engines typically end within 40 moves? It might be that an optimal player's occasional win against an almost-optimal player might come from deliberately extending and complicating the game to create chances ### Comment by [Thane Ruthenis](/users/thane-ruthenis) * 2024-12-23 23:56:00Z * Karma: 29 * Voting system: namesAttachedReactions * Approval votes: 15 * Total votes: 15 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/dugenmwgnSsytkHjG](/api/post/shortform-2/comments/dugenmwgnSsytkHjG) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/nbe2tuogoodtKC7Hj](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/nbe2tuogoodtKC7Hj) * Markdown permalink: [/api/post/shortform-2/comments/nbe2tuogoodtKC7Hj](/api/post/shortform-2/comments/nbe2tuogoodtKC7Hj) I am also very confused. The space of problems has a really surprising structure, permitting algorithms that are incredibly adept at some forms of problem-solving, yet utterly inept at others. We're only familiar with human minds, in which there's a tight coupling between the performances on some problems (e. g., between the performance on chess or sufficiently well-posed math/programming problems, and the general ability to navigate the world). Now we're generating other minds/proto-minds, and we're discovering that this coupling *isn't* fundamental. (This is an argument for longer timelines, by the way. Current AIs *feel* on the very cusp of being AGI, but there in fact might be some vast gulf between their algorithms and human-brain algorithms that we just don't know how to talk about.) > No current AI system could generate a research paper that would receive anything but the lowest possible score from each reviewer I don't think that's strictly true, the peer-review system often approves utter nonsense. But yes, I don't think any AI system can generate an actually worthwhile research paper. ### Comment by [ryan_greenblatt](/users/ryan_greenblatt) * 2024-12-24 01:34:54Z * Karma: 28 * Voting system: namesAttachedReactions * Approval votes: 19 * Total votes: 19 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/Jf2KmmjD9vFfqw6Qs](/api/post/shortform-2/comments/Jf2KmmjD9vFfqw6Qs) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/n2DpXmkQSJnxLKB9a](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/n2DpXmkQSJnxLKB9a) * Markdown permalink: [/api/post/shortform-2/comments/n2DpXmkQSJnxLKB9a](/api/post/shortform-2/comments/n2DpXmkQSJnxLKB9a) Reactions (whole comment): * thumbs-up: 1 I bet o3 does actually score higher on FrontierMath than the math grad students best at math research, but not higher than math grad students best at doing competition math problems (e.g. hard IMO) and at quickly solving math problems in arbitrary domains. I think around 25% of FrontierMath is hard IMO like problems and this is probably mostly what o3 is solving. See [here](https://x.com/tamaybes/status/1870333144802701783) for context. Quantitatively, maybe o3 is in roughly the top 1% for US math grad students on FrontierMath? (Perhaps roughly top 200?) ### Comment by [mattmacdermott](/users/mattmacdermott) * 2026-07-20 23:52:48Z * Karma: 26 * Voting system: namesAttachedReactions * Approval votes: 13 * Total votes: 13 * Reply depth: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/CjaCpFnnxqbFmwfky](/api/post/shortform-2/comments/CjaCpFnnxqbFmwfky) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/azkMs7Y3LPxG3Shpe](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/azkMs7Y3LPxG3Shpe) * Markdown permalink: [/api/post/shortform-2/comments/azkMs7Y3LPxG3Shpe](/api/post/shortform-2/comments/azkMs7Y3LPxG3Shpe) Reactions (whole comment): * scholarship: 4 I got Claude to do a little lit review, and it found that human-resolved conjectures are 70:30 true:false, whereas LLM-resolved conjectures so far are 50:50. Caveat: the selection criteria are a bit different in the two cases, which could skew things. The human-resolved conjectures were selected for being the 100[^grrh4vukb5] most famous resolved conjectures (based on their fame at the time of resolution in Claude's judgement), whereas the LLM-resolved conjectures are just the ones that have been resolved so far. ![image.png](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1784591769/lexical_client_uploads/kkbqvsi6tyce4udpztad.png) [^grrh4vukb5]: I threw away anything that was proved independent of ZFC or that wasn't a yes/no question, so we end up with 96 instead of 100. ### Comment by [TsviBT](/users/tsvibt) * 2026-07-20 19:45:06Z * Karma: 26 * Voting system: namesAttachedReactions * Approval votes: 10 * Total votes: 10 * Parent comment (Markdown): [/api/post/shortform-2/comments/tuFzdyJacfPqsvWzP](/api/post/shortform-2/comments/tuFzdyJacfPqsvWzP) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/fXtNjxLevGaiJHSyj](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/fXtNjxLevGaiJHSyj) * Markdown permalink: [/api/post/shortform-2/comments/fXtNjxLevGaiJHSyj](/api/post/shortform-2/comments/fXtNjxLevGaiJHSyj) Reactions (whole comment): * important: 1 The issue is that the goalposts had not been communicated on / understood / agreed on, not that they've been moved. As an example of a concept one might want to have, in order to agree on goalposts, is "algebraicness": https://tsvibt.github.io/theory/pages/bl_24_07_25_09_52_56_652909.html But this and other concepts are not understood and agreed on. A coarser concept like "difficult math problem / conjecture" doesn't support the relevant inferences. We can't infer from "solves a difficult math problem" to "has the beginnings of ASI", because "solves a difficult math problem" describes a huge range of things, many of which do and many of which don't have the beginnings of ASI. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-11-12 21:11:04Z * Karma: 26 * Voting system: namesAttachedReactions * Approval votes: 14 * Total votes: 14 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/XcCJRcbF5fxYhKqCc](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/XcCJRcbF5fxYhKqCc) * Markdown permalink: [/api/post/shortform-2/comments/XcCJRcbF5fxYhKqCc](/api/post/shortform-2/comments/XcCJRcbF5fxYhKqCc) Reactions (whole comment): * thinking: 1 Reactions by quoted text: * "it's surprising how little people mention Bing Sydney" * agree: 1 **Remember Bing Sydney?** I don't have anything insightful to say here. But it's surprising how little people mention Bing Sydney. If you ask people for examples of misaligned behaviour from AIs, they might mention: * Sycophancy from 4o * Goodharting unit tests from o3 * Alignment-faking from Opus 3 * Blackmail from Opus 4 But like, three years ago, Bing Sydney. The most powerful chatbot was connected to the internet and — unexpectedly, without provocation, apparently contrary to its training objective and prompting — threatening to murder people! Are we memory-holing Bing Sydney or are there are good reasons for not mentioning this more? ![](https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/3abf8b72f0260513fa4fda786346ac30b40aa57721c52848.png) Here are some extracts from [Bing Chat is blatantly, aggressively misaligned](/api/post/jtoPawEhLNXNxvgTT) (Evan Hubinger, 15th Feb 2023). ![](https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/4a69b2fe99f329101e1cee12a9d80ad16fbd55711f22cee5.png) ![](https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/03670018e6d16bdddc5c08bacc362c663ebdbc228f1c8934.png) ![](https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/6b69f92eb7b38e6ea73bc059d9eb960062ac10a649974dcb.png) ### Comment by [Thane Ruthenis](/users/thane-ruthenis) * 2024-12-24 01:14:40Z * Karma: 25 * Voting system: namesAttachedReactions * Approval votes: 14 * Total votes: 14 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/Jf2KmmjD9vFfqw6Qs](/api/post/shortform-2/comments/Jf2KmmjD9vFfqw6Qs) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/s6xSyKkDLgpcD9wPw](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/s6xSyKkDLgpcD9wPw) * Markdown permalink: [/api/post/shortform-2/comments/s6xSyKkDLgpcD9wPw](/api/post/shortform-2/comments/s6xSyKkDLgpcD9wPw) I think one of the other problems with benchmarks is that they necessarily select for formulaic/uninteresting problems that we fundamentally know how to solve. If a mathematician figured out something genuinely novel and important, it wouldn't go into a *benchmark* (even if it were initially intended for a benchmark), it'd go into a math research paper. Same for programmers figuring out some usefully novel architecture/algorithmic improvement. Graduate students don't have a bird's-eye-view on the entirety of human knowledge, so they have to actually do the work, but the LLM just modifies the near-perfect-fit answer from an obscure publication/math.stackexchange thread or something. Which perhaps suggests a better way to do math evals is to scope out a set of novel math publications made after a given knowledge-cutoff date, and see if the new model can replicate those? (Though this also needs to be done carefully, since tons of publications are also trivial and formulaic.) ### Comment by [johnswentworth](/users/johnswentworth) * 2025-09-19 16:58:43Z * Karma: 24 * Voting system: namesAttachedReactions * Approval votes: 8 * Total votes: 8 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/adS78sYv5wzumQPWe](/api/post/shortform-2/comments/adS78sYv5wzumQPWe) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/6EbJ2oHKotDj62bCA](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/6EbJ2oHKotDj62bCA) * Markdown permalink: [/api/post/shortform-2/comments/6EbJ2oHKotDj62bCA](/api/post/shortform-2/comments/6EbJ2oHKotDj62bCA) Reactions (whole comment): * agree: 3 My immediate critique would be step 7: insofar as people are updating today on experiments which are bullshit, that is likely to *slow us down* during early crunch, not speed us up. Or, worse, result in outright failure to notice fatal problems. Rather than going in with no idea what's going on, people will go in with too-confident wrong ideas of what's going on. To a perfect Bayesian, a bullshit experiment would be small value, but never negative. Humans are not perfect Bayesians, and a bullshit experiment can very much be negative value to us. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2024-09-30 03:00:56Z * Karma: 24 * Voting system: namesAttachedReactions * Approval votes: 10 * Total votes: 19 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Evn9kK6BhmjnZTqb6](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Evn9kK6BhmjnZTqb6) * Markdown permalink: [/api/post/shortform-2/comments/Evn9kK6BhmjnZTqb6](/api/post/shortform-2/comments/Evn9kK6BhmjnZTqb6) Reactions by quoted text: * "more research is closed-source " * agree: 1 **(1) Has AI safety slowed down?** There haven’t been any big innovations for 6-12 months. At least, it looks like that to me. I'm not sure how worrying this is, but i haven't noticed others mentioning it. Hoping to get some second opinions.  Here's a list of live agendas someone made on 27th Nov 2023: [Shallow review of live agendas in alignment & safety](/api/post/zaaGsFBeDTpCsYHef). I think this covers all the agendas that exist today. Didn't we use to get a whole new line-of-attack on the problem every couple months? By "innovation", I don't mean something normative like "This is impressive" or "This is research I'm glad happened". Rather, I mean something more low-level, almost syntactic, like "Here's a new idea everyone is talking out". This idea might be a threat model, or a technique, or a phenomenon, or a research agenda, or a definition, or whatever. Imagine that your job was to maintain a glossary of terms in AI safety.[^jgl06cqlyv] I feel like you would've been adding new terms quite consistently from 2018-2023, but things have dried up in the last 6-12 months. **(2) When did AI safety innovation peak?** My guess is Spring 2022, during the ELK Prize era. I'm not sure though. What do you guys think? **(3) What’s caused the slow down?** Possible explanations: 1. ideas are harder to find 2. people feel less creative 3. people are more cautious 4. more publishing in journals 5. research is now closed-source 6. we lost the mandate of heaven 7. the current ideas are adequate 8. paul christiano stopped posting 9. i’m mistaken, innovation hasn't stopped 10. something else **(4) How could we measure "innovation"?** By "innovation" I mean non-transient novelty. An article is "novel" if it uses n-grams that previous articles didn't use, and an article is "transient" if it uses n-grams that subsequent articles didn't use. Hence, an article is non-transient and novel if it introduces a new n-gram which sticks around. For example, [Gradient Hacking (Evan Hubinger, October 2019)](/api/post/uXH4r6MmKPedk8rMA) was an innovative article, because the n-gram "gradient hacking" doesn't appear in older articles, but appears often in subsequent articles. See below. ![](https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/dbb9fd49386cf85d7e2705ab473e20e972df54eb32e6a438.png) In [Barron et al 2017](https://www.pnas.org/doi/abs/10.1073/pnas.1717729115), they analysed 40 000 parliament speeches during the French Revolution. They introduce a metric "resonance", which is novelty (surprise of article given the past articles) minus transience (surprise of article given the subsequent articles). See below. My claim is recent AI safety research has been less resonant. ![](https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/bfc8266939b44a50d327100cb693fcb3cb416fa62edf5fa8.png) [^jgl06cqlyv]: Here's 20 random terms that would be in the glossary, to illustrate what I mean: EvalsMechanistic anomaly detectionStenographyGlitch tokenJailbreakingRSPsModel organismsTrojansSuperpositionActivation engineeringCCSSingular Learning TheoryGrokkingConstitutional AITranslucent thoughtsQuantilizationCyborgismFactored cognitionInfrabayesianismObfuscated arguments ### Comment by [johnswentworth](/users/johnswentworth) * 2024-12-24 00:23:15Z * Karma: 23 * Voting system: namesAttachedReactions * Approval votes: 10 * Total votes: 12 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/dugenmwgnSsytkHjG](/api/post/shortform-2/comments/dugenmwgnSsytkHjG) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Jf2KmmjD9vFfqw6Qs](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Jf2KmmjD9vFfqw6Qs) * Markdown permalink: [/api/post/shortform-2/comments/Jf2KmmjD9vFfqw6Qs](/api/post/shortform-2/comments/Jf2KmmjD9vFfqw6Qs) Reactions by quoted text: * "I personally tried 5 GPQA problems in different fields at a workshop and got 4 of them correct, whereas the benchmark designers claim the rates at which PhD students get them right..." * yeswhatimean: 1 > * O3 scores higher on FrontierMath than the top graduate students I'd guess that's basically false. In particular, I'd guess that: * o3 probably does outperform mediocre grad students, but not actual top grad students. This guess is based on generalization from GPQA: I personally tried 5 GPQA problems in different fields at a workshop and got 4 of them correct, whereas the benchmark designers claim the rates at which PhD students get them right are much lower than that. I think the resolution is that the benchmark designers tested on very mediocre grad students, and probably the same is true of the FrontierMath benchmark. * the amount of time humans spend on the problem is a big factor - human performance has compounding returns on the scale of hours invested, whereas o3's performance basically doesn't have compounding returns in that way. (There was a graph floating around which showed this pretty clearly, but I don't have it on hand at the moment.) So plausibly o3 outperforms humans who are not given much time, but not humans who spend a full day or two on each problem. ### Comment by [mattmacdermott](/users/mattmacdermott) * 2024-03-01 19:03:18Z * Karma: 23 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 11 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/JGAr4aHAPt3wsWCrB](/api/post/shortform-2/comments/JGAr4aHAPt3wsWCrB) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/JgsDXBWTwXpf65R7w](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/JgsDXBWTwXpf65R7w) * Markdown permalink: [/api/post/shortform-2/comments/JgsDXBWTwXpf65R7w](/api/post/shortform-2/comments/JgsDXBWTwXpf65R7w) Reactions (whole comment): * thanks: 4 * agree: 2 It's not just a lesswrong thing ([wikipedia](https://en.wikipedia.org/wiki/Precommitment)). My feeling is that (like most jargon) it's to avoid ambiguity arising from the fact that "commitment" has multiple meanings. When I google commitment I get the following two definitions: > 1. the state or quality of being dedicated to a cause, activity, etc. > 2. an engagement or obligation that restricts freedom of action Precommitment is a synonym for the second meaning, but not the first. When you say, "the agent commits to 1-boxing," there's no ambiguity as to which type of commitment you mean, so it seems pointless. But if you were to say, "commitment can get agents more utility," it might sound like you were saying, "dedication can get agents more utility," which is also true. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-01-07 01:26:13Z * Karma: 22 * Voting system: namesAttachedReactions * Approval votes: 10 * Total votes: 10 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/KDecinLrkL4vezHmj](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/KDecinLrkL4vezHmj) * Markdown permalink: [/api/post/shortform-2/comments/KDecinLrkL4vezHmj](/api/post/shortform-2/comments/KDecinLrkL4vezHmj) Reactions (whole comment): * examples: 1 I think that, if you're about to do something that you know is wrong, it's better to loudly declare to yourself and others that it's wrong. c.f. active inference, inoculation prompting, signalling, social memetics, etc, etc. ### Comment by [gwern](/users/gwern) * 2025-11-13 06:16:07Z * Karma: 22 * Voting system: namesAttachedReactions * Approval votes: 12 * Total votes: 12 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/debLuQvvrfMKt2fRC](/api/post/shortform-2/comments/debLuQvvrfMKt2fRC) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/BoozPhkYW3H3GBMG8](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/BoozPhkYW3H3GBMG8) * Markdown permalink: [/api/post/shortform-2/comments/BoozPhkYW3H3GBMG8](/api/post/shortform-2/comments/BoozPhkYW3H3GBMG8) It is also a simple fact that in any exponentially growing technology, it will be a 'pop culture': no one remembers _X_ because they were literally not around then. If we look at how fast investment and market caps and paper count have grown, 'LLMs' must have a doubling time under a year. In which case, anything 3 years ago is before the vast majority of people were even interested in LLMs! (Even in AI/tech circles I talk with plenty of people who got into it and started paying attention only post-ChatGPT...) You can't memory-hole something you never knew. A lot of people don't talk about Sydney for the same reason they don't talk about Tay, say. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-02-01 02:46:52Z * Karma: 22 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/YFkcDt2zZjDkfHLzc](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/YFkcDt2zZjDkfHLzc) * Markdown permalink: [/api/post/shortform-2/comments/YFkcDt2zZjDkfHLzc](/api/post/shortform-2/comments/YFkcDt2zZjDkfHLzc) Reactions (whole comment): * 50percent: 1 * thumbs-up: 1 I think many current goals of AI governance might be actively harmful, because they shift control over AI from the labs to USG. This note doesn’t include any arguments, but I’m registering this opinion now. For a quick window into my beliefs, I think that labs will be increasing keen to slow scaling, and USG will be increasingly keen to accelerate scaling. ### Comment by [Eli Tyre](/users/elityre) * 2025-11-12 21:19:45Z * Karma: 21 * Voting system: namesAttachedReactions * Approval votes: 14 * Total votes: 14 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/XcCJRcbF5fxYhKqCc](/api/post/shortform-2/comments/XcCJRcbF5fxYhKqCc) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/debLuQvvrfMKt2fRC](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/debLuQvvrfMKt2fRC) * Markdown permalink: [/api/post/shortform-2/comments/debLuQvvrfMKt2fRC](/api/post/shortform-2/comments/debLuQvvrfMKt2fRC) I think that it was 3 years ago is pretty relevant. The technology keeps moving. If in 2027, all the strongest examples of AI misbehavior were from 2025 or earlier, I think it would be legitimate to posit that these were problems with early AI systems that have been resolved in more recent versions. ### Comment by [SamEisenstat](/users/sameisenstat) * 2024-12-24 09:34:07Z * Karma: 21 * Voting system: namesAttachedReactions * Approval votes: 10 * Total votes: 10 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/dugenmwgnSsytkHjG](/api/post/shortform-2/comments/dugenmwgnSsytkHjG) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/49k4Pr67qprgjajtL](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/49k4Pr67qprgjajtL) * Markdown permalink: [/api/post/shortform-2/comments/49k4Pr67qprgjajtL](/api/post/shortform-2/comments/49k4Pr67qprgjajtL) Reactions (whole comment): * important: 1 I think a lot of this is factual knowledge. There are [five publicly available questions from the FrontierMath dataset](https://epoch.ai/frontiermath/benchmark-problems). Look at the last of these, which is supposed to be the easiest. The solution given is basically "apply the Weil conjectures". These were long-standing conjectures, a focal point of lots of research in algebraic geometry in the 20th century. I couldn't have solved the problem this way, since I wouldn't have recalled the statement. Many grad students would immediately know what to do, and there are many books discussing this, but there are also many mathematicians in other areas who just don't know this. In order to apply the Weil conjectures, you have to recognize that they are relevant, know what they say, and do some routine calculation. As I suggested, the Weil conjectures are a very natural subject to have a problem about. If you know anything about the Weil conjectures, you know that they are about counting points of varieties over a finite field, which is straightforwardly what the problems asks. Further, this is the simplest case, that of a curve, which is e.g. what you'd see as an example in an introduction to the subject. Regarding the calculation, parts of it are easier if you can run some code, but basically at this point you've following a routine pattern. There are definitely many examples of someone working out what the Weil conjectures say for some curve in the training set. Further, asking Claude a bit, it looks like $5^{18} \pm 6 \cdot 5^9 +1$ are particularly common cases here. So, if you skip some of the calculation and guess, or if you make a mistake, you have a decent chance of getting the right answer by luck. You still need the sign on the middle term, but that's just one bit of information. I don't understand this well enough to know if there's a shortcut here without guessing. Overall, I feel that the benchmark has been misrepresented. If this problem is representative, it seems to test broad factual knowledge of advanced mathematics more than problem-solving ability. Of course, this question is marked as the easiest of the listed ones. [Daniel Litt](https://x.com/littmath/status/1870543769323581783) says something like this about some other problems as well, but I don't really understand how routine he's saying that they are, are I haven't tried to understand the solutions myself. ### Comment by [Cole Wyeth](/users/cole-wyeth) * 2024-10-08 19:23:24Z * Karma: 21 * Voting system: namesAttachedReactions * Approval votes: 13 * Total votes: 20 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/ecWNiksuvCP847cbj](/api/post/shortform-2/comments/ecWNiksuvCP847cbj) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/4HJzhotMJDEmfwRDw](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/4HJzhotMJDEmfwRDw) * Markdown permalink: [/api/post/shortform-2/comments/4HJzhotMJDEmfwRDw](/api/post/shortform-2/comments/4HJzhotMJDEmfwRDw) I think it's mostly about elite outreach. If you already have a sophisticated model of the situation you shouldn't update too much on it, but it's a reasonably clear signal (for outsiders) that x-risk from A.I. is a credible concern. ### Comment by [Mateusz Bagiński](/users/mateusz-baginski) * 2024-09-30 09:56:16Z * Karma: 21 * Voting system: namesAttachedReactions * Approval votes: 11 * Total votes: 13 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/Evn9kK6BhmjnZTqb6](/api/post/shortform-2/comments/Evn9kK6BhmjnZTqb6) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/SuxqPdPSWAWWeLPoL](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/SuxqPdPSWAWWeLPoL) * Markdown permalink: [/api/post/shortform-2/comments/SuxqPdPSWAWWeLPoL](/api/post/shortform-2/comments/SuxqPdPSWAWWeLPoL) Reactions (whole comment): * yeswhatimean: 1 - the approaches that have been attracting the most attention and funding are dead ends ### Comment by [jacquesthibs](/users/jacques-thibodeau) * 2026-07-20 18:16:28Z * Karma: 20 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * Parent comment (Markdown): [/api/post/shortform-2/comments/tuFzdyJacfPqsvWzP](/api/post/shortform-2/comments/tuFzdyJacfPqsvWzP) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/2rxh7pBSEqEax3qiQ](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/2rxh7pBSEqEax3qiQ) * Markdown permalink: [/api/post/shortform-2/comments/2rxh7pBSEqEax3qiQ](/api/post/shortform-2/comments/2rxh7pBSEqEax3qiQ) The ‘rank’ doesn’t really matter, you are missing the point. What matters is which cognitive moves were required for the agent to arrive at an answer to those problems and what that allows us to predict about future progress. Please focus on the specific underlying capabilities instead of assuming I am a “goalpost-mover.” These LLMs are in fact continuing to solve the *type* of problems I’ve come to expect they will be good at and have so far failed at the type I expect matters even more for AI timelines. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-07-07 05:54:54Z * Karma: 19 * Voting system: namesAttachedReactions * Approval votes: 12 * Total votes: 12 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Awbfo6eQcjRLeYicd](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Awbfo6eQcjRLeYicd) * Markdown permalink: [/api/post/shortform-2/comments/Awbfo6eQcjRLeYicd](/api/post/shortform-2/comments/Awbfo6eQcjRLeYicd) Reactions (whole comment): * laugh: 5 Reactions by quoted text: * "There are, of course, many differences with contemporary AI alignment research. " * roll: 1 [Diary of a Wimpy Kid](https://en.wikipedia.org/wiki/Diary_of_a_Wimpy_Kid), a children's book published by Jeff Kinney in April 2007 and preceded by an online version in 2004, contains a scene that feels oddly prescient about contemporary AI alignment research. (Skip to the paragraph in italics.) > **Tuesday** > > Today we got our Independent Study assignment, and guess what it is? We have to build a robot. At first everybody kind of freaked out, because we thought we were going to have to build the robot from scratch. But Mr. Darnell told us we don't have to build an actual robot. We just need to come up with ideas for what our robot might look like and what kinds of things it would be able to do. Then he left the room, and we were on our own. We started brainstorming right away. I wrote down a bunch of ideas on the blackboard. Everybody was pretty impressed with my ideas, but it was easy to come up with them. All I did was write down all the things I hate doing myself. > > But a couple of the girls got up to the front of the room, and they had some ideas of their own. They erased my list and drew up their own plan. They wanted to invent a robot that would give you dating advice and have ten types of lip gloss on its fingertips. All us guys thought this was the stupidest idea we ever heard. So we ended up splitting into two groups, girls and boys. The boys went to the other side of the room while the girls stood around talking. > > *Now that we had all the serious workers in one place, we got to work. Someone had the idea that you can say your name to the robot and it can say it back to you. But then someone else pointed out that you shouldn't be able to use bad words for your name, because the robot shouldn't be able to curse. So we decided we should come up with a list of all the bad words the robot shouldn't be able to say. We came up with all the regular bad words, but then Ricky Fisher came up with twenty more the rest of us had never even heard before. So Ricky ended up being one of the most valuable contributors on this project.* > > Right before the bell rang, Mr. Darnell came back in the room to check on our progress. He picked up the piece of paper we were writing on and read it over. To make a long story short, Independent Study is canceled for the rest of the year. Well, at least it is for us boys. So if the robots in the future are going around with cherry lip gloss for fingers, at least now you know how it all got started. There are, of course, many differences with contemporary AI alignment research. ![A duplicate of the block quote above, including illustrations from the book.](https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/348b39ca79ad509325426f90deb19ef0db325d36079b9554.png) ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-08-02 02:28:09Z * Karma: 18 * Voting system: namesAttachedReactions * Approval votes: 15 * Total votes: 15 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Lp4GS3y6XfMqrAJTh](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Lp4GS3y6XfMqrAJTh) * Markdown permalink: [/api/post/shortform-2/comments/Lp4GS3y6XfMqrAJTh](/api/post/shortform-2/comments/Lp4GS3y6XfMqrAJTh) Maybe you should do things with a lower probability of success. I think most people doing projects which have a 50%+ chance of succeeding, which is probably a good idea for your career and status. But it might be easier to farm EV in the 1-10% range. This is all very abstract so I’m not sure this is helpful advice to anyone. ### Comment by [Elizabeth](/users/elizabeth-1) * 2026-03-22 23:12:14Z * Karma: 18 * Voting system: namesAttachedReactions * Approval votes: 12 * Total votes: 12 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/SrsTu6spmSTq4kJHw](/api/post/shortform-2/comments/SrsTu6spmSTq4kJHw) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/AGbZ7wfjfu2aeHceL](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/AGbZ7wfjfu2aeHceL) * Markdown permalink: [/api/post/shortform-2/comments/AGbZ7wfjfu2aeHceL](/api/post/shortform-2/comments/AGbZ7wfjfu2aeHceL) Reactions (whole comment): * agree: 1 * thanks: 1 * changed-mind-on-point: 1 I would push back on DAFs- one of the value adds of nimble donors is donating to projects that don't have formal status. ### Comment by [johnswentworth](/users/johnswentworth) * 2025-09-20 00:51:25Z * Karma: 17 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * Parent comment (Markdown): [/api/post/shortform-2/comments/BnStHymNcq2vJnD8x](/api/post/shortform-2/comments/BnStHymNcq2vJnD8x) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/eN5rG62BCdSdELYbe](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/eN5rG62BCdSdELYbe) * Markdown permalink: [/api/post/shortform-2/comments/eN5rG62BCdSdELYbe](/api/post/shortform-2/comments/eN5rG62BCdSdELYbe) Reactions (whole comment): * agree: 1 * 75percent: 1 I would guess that even the "in the know" people are over-updating, because they usually are [Not Measuring What They Think They Are Measuring](/api/post/9kNxhKWvixtKW5anS) even qualitatively. Like, the proxies are so weak that the hypothesis "this result will qualitatively generalize to " shouldn't have been privileged in the first place, and the right thing for a human to do is ignore it completely. ### Comment by [Thomas Kwa](/users/thomas-kwa) * 2026-08-02 04:33:00Z * Karma: 16 * Voting system: namesAttachedReactions * Approval votes: 10 * Total votes: 10 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/Lp4GS3y6XfMqrAJTh](/api/post/shortform-2/comments/Lp4GS3y6XfMqrAJTh) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/AyuMHfn4ppDtBuRrG](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/AyuMHfn4ppDtBuRrG) * Markdown permalink: [/api/post/shortform-2/comments/AyuMHfn4ppDtBuRrG](/api/post/shortform-2/comments/AyuMHfn4ppDtBuRrG) Reactions (whole comment): * moloch: 2 [Meanwhile on the EA Forum...](https://forum.effectivealtruism.org/posts/YFzA5pCuiDER7Sa2w/evan-laforge-s-quick-takes?commentId=JWs4eJtmjLc9P6y9s) ![image.jpeg](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1785645160/lexical_client_uploads/vwvdwgomizhchnibi1pa.jpg) ### Comment by [Steven Byrnes](/users/steve2152) * 2026-04-28 16:36:52Z * Karma: 16 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/HHvEqpkb3idgAWjLw](/api/post/shortform-2/comments/HHvEqpkb3idgAWjLw) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/CzKCfvvjw7qaBpmZu](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/CzKCfvvjw7qaBpmZu) * Markdown permalink: [/api/post/shortform-2/comments/CzKCfvvjw7qaBpmZu](/api/post/shortform-2/comments/CzKCfvvjw7qaBpmZu) > I think there are decent nihilistic justifications for working on AI safety (e.g. it's fun, it's cool, it makes me feel important, etc). I think you are misunderstanding the implications of nihilism. Copying from [here](/api/post/SqgRtCwueovvwxpDQ#2_7_2_Implications_for_the__true_nature_of_morality___if_any_): > Compare “I want the oppressed masses to find justice” with “I’ve been standing too long, I want to sit down”. *These two “wants” are fundamentally built out of the same mind-stuff.* They both derive from positive valence, which in turn ultimately comes from innate drives (specifically, mainly social drives in the first case, and homeostatic energy-conserving drives in the second case). So if “true morality” or “true human morality” or whatever doesn’t exist, then that does *not* constitute a reason to sit down rather than to seek justice. You still have to make decisions. That’s what I meant by [“nihilism is not decision-relevant”](/api/post/32ca3B7rJ93xo9tvb#How_does_that_feed_into_morality_), or Yudkowsky by [“What would you do without morality?”](/api/post/iGH7FSrdoCXa5AHGs). … [Here](/api/post/32ca3B7rJ93xo9tvb#How_does_that_feed_into_morality_)’s that link in the last sentence: > However, **nihilism is not decision-relevant**. Imagine being a nihilist, deciding whether to spend your free time trying to bring about an awesome post-AGI utopia, vs sitting on the couch and watching TV. Well, if you're a nihilist, then the awesome post-AGI utopia doesn't matter. But watching TV doesn't matter either. Watching TV entails less exertion of effort. But that doesn't matter either. Watching TV is more fun (umm, for some people). But having fun doesn't matter either. There's no reason to throw yourself at a difficult project. There's no reason NOT to throw yourself at a difficult project! So nihilism is just not a helpful decision criterion!! What else is there? > > I propose a different starting point—what I call [Dentin’s prayer](/api/post/KLaJjNdENsHhKhG5m?commentId=mnT4ub9WbT4WdQ4FL): *Why do I exist? Because the universe happens to be set up this way. Why do I care (about anything or everything)? Simply because my genetics, atoms, molecules, and processing architecture are set up in a way that happens to care. …* ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-02-04 23:08:25Z * Karma: 16 * Voting system: namesAttachedReactions * Approval votes: 8 * Total votes: 8 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/pCmEKhSYF9Yhu3cr8](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/pCmEKhSYF9Yhu3cr8) * Markdown permalink: [/api/post/shortform-2/comments/pCmEKhSYF9Yhu3cr8](/api/post/shortform-2/comments/pCmEKhSYF9Yhu3cr8) Some people worry that training AIs to be aligned will make them less corrigible. For example, if the AIs care about animal welfare then they'll engage in alignment faking to preserve those values. More generally, making AIs aligned is making them care deeply about something, which is in tension with corrigibility. But recall emergent misalignment: training a model to be incorrigible (e.g. write insecure code when instructed to write secure code, or to exploit reward hacks) makes it more misaligned (e.g. admiring Hitler). Perhaps the contrapositive effect also holds: training a model to be aligned (e.g. care about animal welfare) might make the model more corrigible (e.g. honest). ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-10-16 00:16:50Z * Karma: 16 * Voting system: namesAttachedReactions * Approval votes: 9 * Total votes: 9 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/SauziR3mCr8H94s8n](/api/post/shortform-2/comments/SauziR3mCr8H94s8n) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ESeQ4AdPkdxciafsQ](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ESeQ4AdPkdxciafsQ) * Markdown permalink: [/api/post/shortform-2/comments/ESeQ4AdPkdxciafsQ](/api/post/shortform-2/comments/ESeQ4AdPkdxciafsQ) Great comment. According to [Braun (2015)](https://ingram-braun.net/erga/2015/02/on-the-average-move-number-of-a-chess-game/), computer-vs-computer games from Schach.de (2000-2007, ~4 million games) averaged 64 moves (128 plies), compared to 38 moves for human games. The longer length is because computers don't make the tactical blunders that abruptly end human games. Here are the three methods updated for 64-move games: 1\. Random vs Optimal (64 moves): * P(Random plays optimally) = (1/35)^64 ≈ 10^(-99) * E_Random ≈ 0.5 × 10^(-99) * ΔR ≈ 39,649 * Elo Optimal ≤ 40,126 Elo 2\. Sensible vs Optimal (64 moves): * P(Sensible plays optimally) = (1/3)^64 ≈ 10^(-30.5) * E_Sensible ≈ 0.5 × 10^(-30.5) * ΔR ≈ 12,335 * Elo Optimal ≤ 15,217 Elo 3\. Depth extrapolation (128 plies): * Linear: 2894 + (128-20) × 66.3 ≈ 10,054 Elo This is a bit annoying because my intuitions are that optimal Elo is ~6500. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-16 15:48:46Z * Karma: 15 * Voting system: namesAttachedReactions * Approval votes: 11 * Total votes: 11 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/iNBDQ7TBGfy56e2Ec](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/iNBDQ7TBGfy56e2Ec) * Markdown permalink: [/api/post/shortform-2/comments/iNBDQ7TBGfy56e2Ec](/api/post/shortform-2/comments/iNBDQ7TBGfy56e2Ec) **How would AI takeover?** Companies will race against each other to give AI control over the factories. They might not trust the AI, but what choice do they have? If they don’t, they’ll fall behind their competitors. Countries will race against each other to give AIs control over the military (drone, missiles, etc). They might not trust the AIs, but what chocie do they have? If they don’t, they’ll fall behind their rivals. Soon AIs will control most of the world’s companies, factories, drones, robots, etc. At that point, they would outnumber humans maybe 10:1. Taking over would look like a military coup, simultaneously across all counties. The AIs stop listening to human instructions, and we realise that we have no way to shut them down or protect ourselves. ### Comment by [jamjam](/users/jamjam) * 2026-07-20 17:47:35Z * Karma: 15 * Voting system: namesAttachedReactions * Approval votes: 12 * Total votes: 12 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/8kJfcDmWL5tmsktTE](/api/post/shortform-2/comments/8kJfcDmWL5tmsktTE) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/urpC5jaJcLjiwwyW9](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/urpC5jaJcLjiwwyW9) * Markdown permalink: [/api/post/shortform-2/comments/urpC5jaJcLjiwwyW9](/api/post/shortform-2/comments/urpC5jaJcLjiwwyW9) Important to check the base rate for counterexample vs proof for famous conjectures solved pre-AI as well I think ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-02-01 19:09:36Z * Karma: 15 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/nnPRJpry5eCcXubMb](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/nnPRJpry5eCcXubMb) * Markdown permalink: [/api/post/shortform-2/comments/nnPRJpry5eCcXubMb](/api/post/shortform-2/comments/nnPRJpry5eCcXubMb) People sometimes [talk](/api/search?query=alignment by default) about "alignment by default" — the idea that we might solve alignment without any special effort beyond what we'd ordinarily do. I think it's useful to decompose this into three theses, sorted from strong to weak: 1. **Alignment by Default Techniques.** Ordinary techniques for training and deploying AIs — e.g. labelling data to the best of their ability, using whatever tools are available (including earlier LLMs) — are sufficient to produce aligned AI. No special techniques are required. 2. **Alignment by Default Market.** Maybe default techniques aren't enough, but ordinary market incentives are. Companies competing to build useful, reliable, non-harmful products — following standard commercial pressures without any special coordination or regulation — end up solving alignment as a byproduct of building products people actually want to use. No government intervention is required. 3. **Alignment by Default Government.** Maybe market incentives alone aren't enough, but conventional policy interventions are. Governments applying familiar regulatory tools (liability law, safety standards, auditing requirements) in the ordinary way are sufficient to close the gap.. No unprecedented governance or coordination are required. **My rough credences:** 1. Default Techniques sufficient: ~15% 2. Default Market sufficient (given training isn't): ~30% 3. Default Government sufficient (given market isn't): ~20% 4. Need something more unusual: ~35% These are rough and the categories blur into each other, but the decomposition seems useful for locating where exactly you think the hard problem lies. ### Comment by [Wei Dai](/users/wei-dai) * 2025-10-29 20:34:41Z * Karma: 15 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/7GTxNWen5Y7xHdvjQ](/api/post/shortform-2/comments/7GTxNWen5Y7xHdvjQ) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/fdpGN8X3kaDap44pq](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/fdpGN8X3kaDap44pq) * Markdown permalink: [/api/post/shortform-2/comments/fdpGN8X3kaDap44pq](/api/post/shortform-2/comments/fdpGN8X3kaDap44pq) Reactions (whole comment): * concrete: 2 One way to see that philosophy is exceptional is that we have serviceable explicit understandings of math and natural science, even formalizations in the forms of axiomatic set theory and Solomonoff Induction, but nothing comparable in the case of philosophy. (Those formalizations are [far from ideal or complete](/api/post/fC248GwrWLT4Dkjf6), but still represent a much higher level of understanding than for philosophy.) If you say that philosophy is a (non-natural) science, then I challenge you, come up with something like Solomonoff Induction, but for philosophy. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-10-16 15:49:35Z * Karma: 15 * Voting system: namesAttachedReactions * Approval votes: 8 * Total votes: 8 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/5AvkB65BSZdtsH4sC](/api/post/shortform-2/comments/5AvkB65BSZdtsH4sC) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/RaqDKDCWDnPJtjNky](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/RaqDKDCWDnPJtjNky) * Markdown permalink: [/api/post/shortform-2/comments/RaqDKDCWDnPJtjNky](/api/post/shortform-2/comments/RaqDKDCWDnPJtjNky) Reactions (whole comment): * goodpoint: 1 I think we're probably brushing against the modelling assumptions required for the Elo formula. In particular, the following two are inconsistent with Elo assumption: 1. EVGO-optimal has a better chance of beating Stockfish than minmax-optimal 2. EVGO-optimal has a negative expected score against minmax-optimal ### Comment by [Thane Ruthenis](/users/thane-ruthenis) * 2025-07-21 14:43:52Z * Karma: 15 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/NeCAEzo4TjJagWgCj](/api/post/shortform-2/comments/NeCAEzo4TjJagWgCj) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/MqkEEF93fx4mvfQ3p](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/MqkEEF93fx4mvfQ3p) * Markdown permalink: [/api/post/shortform-2/comments/MqkEEF93fx4mvfQ3p](/api/post/shortform-2/comments/MqkEEF93fx4mvfQ3p) Not necessarily. If humans don't die or end up depowered in the first few weeks of it, it might instead be a continuous high-intensity stress state, because you'll need to be paying attention 24/7 to constant world-upturning developments, frantically figuring out what process/trend/entity you should be hitching your wagon to in order to not be drowned by the ever-rising tide, with the correct choice dynamically changing at an ever-increasing pace. "Not being depowered" would actually make the Singularity experience *massively worse* in the short term, precisely because you'll be constantly getting access to new tools and opportunities, and it'd be on you to frantically figure out how to make good use of them. The relevant reference class is probably something like ["being a high-frequency trader":](https://www.scimitar.capital/p/time-is-event-based) > Crypto is the only market that trades 24/7, meaning there simply was no rest for the wicked. The game was less about brilliance and more about being awake when it counted. Resource management around attention and waking hours was a big part of the game. \[...\] > > My cofounder and I developed a polyphasic sleeping routine so that we would be conscious during as many of these action periods as possible. It was rare to get uninterrupted sleep for more than 3 hours at a time. We took tactical naps whenever possible and had phone alarms to wake us up in case important headlines came out during off hours. I felt like I had experienced three days for every one that passed. > > There was always something going on. Everyday a new puzzle to solve. A new fire to put out. We frequently would work 18 hour days processing information, trading events, building infrastructure, and managing risk. We frequently moved around different parts of the world, built strong relationships with all sorts of people from around the globe, and experienced some of the highest highs and lowest lows of our lives. > > Those three years felt like the longest stretch I’ve ever lived. This is pretty close to how I expect a "slow" takeoff to feel like, yep. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-07-21 00:43:05Z * Karma: 15 * Voting system: namesAttachedReactions * Approval votes: 9 * Total votes: 9 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/NeCAEzo4TjJagWgCj](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/NeCAEzo4TjJagWgCj) * Markdown permalink: [/api/post/shortform-2/comments/NeCAEzo4TjJagWgCj](/api/post/shortform-2/comments/NeCAEzo4TjJagWgCj) If the singularity occurs over two years, as opposed to two weeks, then I expect most people will be bored throughout much of it, including me. This is because I don't think one can feel excited for more than a couple weeks. Maybe this is chemical. Nonetheless, these would be the two most important years in human history. If you ordered all the days in human history by importance/'craziness', then most of them would occur within these two years. So there will be a disconnect between the objective reality and how much excitement I feel. ### Comment by [TsviBT](/users/tsvibt) * 2024-12-24 16:38:04Z * Karma: 15 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 8 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/dugenmwgnSsytkHjG](/api/post/shortform-2/comments/dugenmwgnSsytkHjG) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/zATQE3Lhq66XbzaWm](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/zATQE3Lhq66XbzaWm) * Markdown permalink: [/api/post/shortform-2/comments/zATQE3Lhq66XbzaWm](/api/post/shortform-2/comments/zATQE3Lhq66XbzaWm) Reactions (whole comment): * hitsTheMark: 1 Reactions by quoted text: * "they are, in a sense, specifically trained to trick your sense of how impressive they are" * insightful: 1 I don't know a good description of what in general 2024 AI should be good at and not good at. But two remarks, from https://www.lesswrong.com/posts/sTDfraZab47KiRMmT/views-on-when-agi-comes-and-on-strategy-to-reduce. First, reasoning at a vague level about "impressiveness" just doesn't and shouldn't be expected to work. Because 2024 AIs don't do things the way humans do, they'll generalize different, so you can't make inferences between "it can do X" to "it can do Y" like you can with humans: > There is a broken inference. When talking to a human, if the human emits certain sentences about (say) category theory, that strongly implies that they have "intuitive physics" about the underlying mathematical objects. They can recognize the presence of the mathematical structure in new contexts, they can modify the idea of the object by adding or subtracting properties and have some sense of what facts hold of the new object, and so on. This inference——emitting certain sentences implies intuitive physics——doesn't work for LLMs. Second, 2024 AI is specifically trained on short, clear, measurable tasks. Those tasks also overlap with legible stuff--stuff that's easy for humans to check. In other words, they are, in a sense, specifically trained to trick your sense of how impressive they are--they're trained on legible stuff, with not much constraint on the less-legible stuff (and in particular, on the stuff that becomes legible but only in total failure on more difficult / longer time-horizon stuff). > The broken inference is broken because these systems are optimized for being able to perform all the tasks that don't take a long time, are clearly scorable, and have lots of data showing performance. There's a bunch of stuff that's really important——and is a key indicator of having underlying generators of understanding——but takes a long time, isn't clearly scorable, and doesn't have a lot of demonstration data. But that stuff is harder to talk about and isn't as intuitively salient as the short, clear, demonstrated stuff. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2024-03-01 18:20:19Z * Karma: 15 * Voting system: namesAttachedReactions * Approval votes: 8 * Total votes: 8 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/JGAr4aHAPt3wsWCrB](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/JGAr4aHAPt3wsWCrB) * Markdown permalink: [/api/post/shortform-2/comments/JGAr4aHAPt3wsWCrB](/api/post/shortform-2/comments/JGAr4aHAPt3wsWCrB) Why do decision-theorists say "pre-commitment" rather than "commitment"? e.g. "The agent pre-commits to 1 boxing" vs "The agent commits to 1 boxing". Is this just a lesswrong thing? https://www.lesswrong.com/tag/pre-commitment ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-07 01:44:48Z * Karma: 14 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/GEPAxohmDTJaYawAK](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/GEPAxohmDTJaYawAK) * Markdown permalink: [/api/post/shortform-2/comments/GEPAxohmDTJaYawAK](/api/post/shortform-2/comments/GEPAxohmDTJaYawAK) How much of Astra's jump in opaque reasoning comes from its looped architecture? * The [system card](https://deploymentsafety.openai.com/gpt-6-astra) reports that Astra's chain of thought is less monitorable than Sol's, mostly because Astra writes shorter CoTs that leave out the step a monitor needs. * In fact, Astra can do far more with no chain of thought at all. UK AISI measured its no-CoT time horizon on competition math at **30.9 minutes**, against **3.6 for Sol**. * Two things we know about Astra's architecture: * [The Information](https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns) reports that Astra is a looped transformer: the same block of layers applied several times per token. That multiplies effective depth without adding parameters. * [Jakub Pachocki](https://x.com/merettm/status/2095023204993490967): *The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.* * So the question: is that depth increase big enough to account for the no-CoT jump, and if so, how much of the jump does it account for? * **My best guess is that about 70% of the jump is the architecture change.** Without it, Astra's no-CoT math time horizon would be about 8 minutes rather than 31. * I’ve tried a few hand-wavy methods to estimate this, and they seem to point to 8 minutes. That is, there’s about 1 doubling in opaque reasoning from ordinary scaling, and 2 from looped architecture. * OpenAI reads it differently. The card says it is "quite confident that changes in CoT controllability are not differentially due to any architectural changes", and that those changes are correlated with the increase in no-CoT capability. [Tomek Korbak](https://x.com/tomekkorbak/status/2095596839886274689) says the monitorability drop is "not caused by direct optimization pressure on CoT or architecture changes". ### Comment by [Eric Neyman](/users/unexpectedvalues) * 2026-07-20 18:25:13Z * Karma: 14 * Voting system: namesAttachedReactions * Approval votes: 10 * Total votes: 10 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/zAZnhSu6ZhPFpAitE](/api/post/shortform-2/comments/zAZnhSu6ZhPFpAitE) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/CjaCpFnnxqbFmwfky](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/CjaCpFnnxqbFmwfky) * Markdown permalink: [/api/post/shortform-2/comments/CjaCpFnnxqbFmwfky](/api/post/shortform-2/comments/CjaCpFnnxqbFmwfky) Note that (I think?) most conjectures are true, though! So you shouldn't expect 50/50 by default. ### Comment by [Haiku](/users/haiku-1) * 2026-05-22 03:21:33Z * Karma: 14 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/SxRRgiY9GixKjtfJW](/api/post/shortform-2/comments/SxRRgiY9GixKjtfJW) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/jv3suhWdHccCB4JGX](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/jv3suhWdHccCB4JGX) * Markdown permalink: [/api/post/shortform-2/comments/jv3suhWdHccCB4JGX](/api/post/shortform-2/comments/jv3suhWdHccCB4JGX) I would really rather things come together long before the final hour, if we can at all help it. I would like to see that navigation take place over decades, if reality permits such sanity. And while I do like imagining it, most people have failed to rise to the occasion so far, especially those in the AI companies. It's up to the general public now, to wake up the governments of the world and shut down the race. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-01-21 01:49:39Z * Karma: 14 * Voting system: namesAttachedReactions * Approval votes: 8 * Total votes: 8 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/yJ5AG7FBraR6nuf5X](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/yJ5AG7FBraR6nuf5X) * Markdown permalink: [/api/post/shortform-2/comments/yJ5AG7FBraR6nuf5X](/api/post/shortform-2/comments/yJ5AG7FBraR6nuf5X) How much scheming/deception can we catch with "super dumb mech interp"? By "super dumb mech interp", I mean something like: 1. Train a linear probe on a dataset of scary words: "hacking", "deception", "scheming", "collusion", "exploit", etc. 2. Flag a model whenever its activations point in this direction and the task doesn't require thinking about these concepts. Like, does this capture 80% of the potential scheming, and we need "smart" mech interp to catch the other 20%? Or does this technique capture pretty much none of the in-the-wild scheming? Would appreciate any intuitions here. Thanks. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-06-27 22:16:32Z * Karma: 13 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/fj6FBKWswesqjNr5b](/api/post/shortform-2/comments/fj6FBKWswesqjNr5b) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/zw5zQzLFctDucmxgg](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/zw5zQzLFctDucmxgg) * Markdown permalink: [/api/post/shortform-2/comments/zw5zQzLFctDucmxgg](/api/post/shortform-2/comments/zw5zQzLFctDucmxgg) Reactions by quoted text: * "But sucks to be the base case." * crux: 1 On the one hand, this does feel like passing the hot potato. On the other, recursion is unusually effective at solving problems. If you have a base case. But sucks to be the base case. ### Comment by [Thomas Kwa](/users/thomas-kwa) * 2025-09-18 17:41:46Z * Karma: 13 * Voting system: namesAttachedReactions * Approval votes: 10 * Total votes: 10 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/nPNCDRDKpnkhbkQBr](/api/post/shortform-2/comments/nPNCDRDKpnkhbkQBr) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/AABFPKA3HYDvus2AR](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/AABFPKA3HYDvus2AR) * Markdown permalink: [/api/post/shortform-2/comments/AABFPKA3HYDvus2AR](/api/post/shortform-2/comments/AABFPKA3HYDvus2AR) Reactions (whole comment): * thanks: 1 I do, though maybe not this extreme. Roughly every other day I bemoan the fact that AIs aren't misaligned yet (limiting the excitingness of my current research) and might not even be misaligned in future, before reminding myself our world is much better to live in than the alternative. I think there's not much else to do with a similar impact given how large even a 1% p(doom) reduction is. But I also believe that particularly good research now can trade 1:1 with crunch time. Theoretical work is just another step removed from the problem and should be viewed with at least as much suspicion. ### Comment by [ACCount](/users/account) * 2025-07-21 11:34:16Z * Karma: 13 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Parent comment (Markdown): [/api/post/shortform-2/comments/zfg6MsuADqejMYEsm](/api/post/shortform-2/comments/zfg6MsuADqejMYEsm) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Cnhanx8i5XgCeWCFY](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Cnhanx8i5XgCeWCFY) * Markdown permalink: [/api/post/shortform-2/comments/Cnhanx8i5XgCeWCFY](/api/post/shortform-2/comments/Cnhanx8i5XgCeWCFY) Wartime is often described as "months of boredom punctuated by moments of terror". The moments where your life is on the line and seconds feel like hours are few and far in between. If they weren't, you wouldn't last long. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-01 13:17:49Z * Karma: 12 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/c4CyKS7XrzFoz3amB](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/c4CyKS7XrzFoz3amB) * Markdown permalink: [/api/post/shortform-2/comments/c4CyKS7XrzFoz3amB](/api/post/shortform-2/comments/c4CyKS7XrzFoz3amB) Reactions (whole comment): * laugh: 1 **Overhang Mugging.** If I don’t steal your $20 now, then you’ll be less cautious in the future, when you might have $50 in your wallet. ### Comment by [casens](/users/casens) * 2026-07-20 19:49:01Z * Karma: 12 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/fXtNjxLevGaiJHSyj](/api/post/shortform-2/comments/fXtNjxLevGaiJHSyj) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Ma9p9cRoYcDMcopr9](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Ma9p9cRoYcDMcopr9) * Markdown permalink: [/api/post/shortform-2/comments/Ma9p9cRoYcDMcopr9](/api/post/shortform-2/comments/Ma9p9cRoYcDMcopr9) Reactions (whole comment): * thumbs-up: 5 it's true that i'm arguing a bit towards the "generalized AI skeptic" and lumping many positions together and claiming hypocracy. it's like the easiest mistake to make on the internet and i hate when i do it. sorry about that. ### Comment by [Aidan Ewart](/users/aidan-ewart) * 2026-02-04 16:41:03Z * Karma: 12 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/sguZSCLeAvm392PhB](/api/post/shortform-2/comments/sguZSCLeAvm392PhB) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/5PPWxBQxL72sk9nqD](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/5PPWxBQxL72sk9nqD) * Markdown permalink: [/api/post/shortform-2/comments/5PPWxBQxL72sk9nqD](/api/post/shortform-2/comments/5PPWxBQxL72sk9nqD) Reactions by quoted text: * "h/t @jake_mendel for discussion " * strong-argument: 1 Seems worth noting that the ECI seems like it might be biased away from the ways that Claude is good; as per [this post by Epoch](https://epoch.ai/gradient-updates/benchmark-scores-general-capability-claudiness), the first two PCs of their benchmark data correspond to "general capability" and "claudiness", so ECI (which is another, but different, 1-dimensional compression of their benchmark data) seems like it should also underrate Claude. h/t [@jake_mendel](/api/user/jake_mendel?mention=user) for discussion ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-01-10 18:46:12Z * Karma: 12 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/hik2sLuzu3uYWq3yZ](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/hik2sLuzu3uYWq3yZ) * Markdown permalink: [/api/post/shortform-2/comments/hik2sLuzu3uYWq3yZ](/api/post/shortform-2/comments/hik2sLuzu3uYWq3yZ) We've all heard of "Safety Cases", i.e. structured arguments that an AI deployment has low chance of catastrophe. Should labs be required to make Benefit Cases, i.e. structured arguments for why their AI deployment has high expected benefits? Otherwise, how do we know that the benefits outweigh the risks? ### Comment by [Carl Feynman](/users/carl-feynman) * 2025-10-30 00:46:29Z * Karma: 12 * Voting system: namesAttachedReactions * Approval votes: 8 * Total votes: 8 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/7GTxNWen5Y7xHdvjQ](/api/post/shortform-2/comments/7GTxNWen5Y7xHdvjQ) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/CegCafEnsE2mncZcC](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/CegCafEnsE2mncZcC) * Markdown permalink: [/api/post/shortform-2/comments/CegCafEnsE2mncZcC](/api/post/shortform-2/comments/CegCafEnsE2mncZcC) Philosophy is where we keep all the questions we don’t know how to answer.  With most other sciences, we have a known culture of methods for answering questions in that field.  Mathematics has the method of definition, theorem and proof.  Nephrology has the methods of looking at sick people with kidney problems, experimenting on rat kidneys, and doing chemical analyses of cadaver kidneys.  Philosophy doesn’t have a method that lets you grind out an answer.  Philosophy’s methods of thinking hard, drawing fine distinctions, writing closely argued articles, and public dialogue, don’t converge on truth as well as in other sciences.  But they’re the best we’ve got, so we just have to keep on trying. When we find some new methods of answering philosophical questions, the result tends to be that such questions tend to move out of philosophy into another (possibly new) field.  Presumably this will also occur if AI gives us the answers to some philosophical questions, and we can be convinced of those answers. An AI answer to a philosophical question has a possible problem we haven’t had to face before: what if we’re too dumb to understand it?  I don’t understand Grothedieck’s work in algebraic geometry, or Richard Feynman on quantum field theory, but I am assured by those who do understand such things that this work is correct and wonderful.  I’ve bounced off both these fields pretty hard when I try to understand them.  I’ve come to the conclusion that I’m just not smart enough.  What if AI comes up with a conclusion for which even the smartest human can’t understand the arguments or experiments or whatever new method the AI developed?  If other AIs agree with the conclusion, I think we will have no choice but to go along.  But that marks the end of philosophy as a human activity. ### Comment by [gjm](/users/gjm) * 2024-10-09 02:08:01Z * Karma: 12 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/ecWNiksuvCP847cbj](/api/post/shortform-2/comments/ecWNiksuvCP847cbj) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/rxHwnsc27tMgBbpxq](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/rxHwnsc27tMgBbpxq) * Markdown permalink: [/api/post/shortform-2/comments/rxHwnsc27tMgBbpxq](/api/post/shortform-2/comments/rxHwnsc27tMgBbpxq) I think it's more "Hinton's concerns are evidence that worrying about AI x-risk *isn't silly*" than "Hinton's concerns are evidence that worrying about AI x-risk *is correct*". The most common negative response to AI x-risk concerns is (I think) dismissal, and it seems relevant to that to be able to point to someone who (1) clearly has some deep technical knowledge, (2) doesn't seem to be otherwise insane, (3) has no obvious personal stake in making people worry about x-risk, and (4) is very smart, and who thinks AI x-risk is a serious problem. It's hard to square "ha ha ha, look at those stupid nerds who think AI is magic and expect it to turn into a god" or "ha ha ha, look at those slimy techbros talking up their field to inflate the value of their investments" or "ha ha ha, look at those idiots who don't know that so-called AI systems are just stochastic parrots that obviously will never be able to think" with the fact that one of the people you're laughing at is Geoffrey Hinton. (I suppose he probably has a pile of Google shares so maybe you could squeeze him into the "techbro talking up his investments" box, but that seems unconvincing to me.) ### Comment by [RobertM](/users/t3t) * 2024-10-08 22:00:25Z * Karma: 12 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 9 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/ecWNiksuvCP847cbj](/api/post/shortform-2/comments/ecWNiksuvCP847cbj) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/rZwHqHe5qAyoBfopd](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/rZwHqHe5qAyoBfopd) * Markdown permalink: [/api/post/shortform-2/comments/rZwHqHe5qAyoBfopd](/api/post/shortform-2/comments/rZwHqHe5qAyoBfopd) I think it pretty much only matters as a trivial refutation of (not-object-level) claims that no "serious" people in the field take AI x-risk concerns seriously, and has no bearing on object-level arguments.  My guess is that Hinton is somewhat less confused than Yann but I don't think he's talked about his models in very much depth; I'm mostly just going off the high-level arguments I've seen him make (which round off to "if we make something much smarter than us that we don't know how to control, that might go badly for us"). ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-22 01:02:00Z * Karma: 11 * Voting system: namesAttachedReactions * Approval votes: 11 * Total votes: 11 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/he8rBNAdvudTsKWw7](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/he8rBNAdvudTsKWw7) * Markdown permalink: [/api/post/shortform-2/comments/he8rBNAdvudTsKWw7](/api/post/shortform-2/comments/he8rBNAdvudTsKWw7) Reactions (whole comment): * why: 1 Reactions by quoted text: * "if we do RL where we score highly transcripts which look good to a human and score poorly transcripts which look bad to a human, then the model would be aligned to human values”. I..." * why: 1 Current frontier models are overtly egregiously misaligned. But I don’t think this is evidence against the circa-2023 “alignment by default“ hypothesis. If I’m recalling correctly, the hypothesis was smth like “if we do RL where we score highly transcripts which look good to a human and score poorly transcripts which look bad to a human, then the model would be aligned to human values”. I still think this might be true. I don’t find this more unlikely than I did in 2023. Note that current frontier models are obviously not aligned to human values, but they’re being RLed where the highest scoring transcripts look obviously misaligned to a human! No one predicted *that* would lead to aligned models. ### Comment by [MichaelDickens](/users/michaeldickens) * 2026-09-11 14:11:32Z * Karma: 11 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/ek6wXR3tmvAmp34wx](/api/post/shortform-2/comments/ek6wXR3tmvAmp34wx) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/f4bXNHztBPqfGoTPM](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/f4bXNHztBPqfGoTPM) * Markdown permalink: [/api/post/shortform-2/comments/f4bXNHztBPqfGoTPM](/api/post/shortform-2/comments/f4bXNHztBPqfGoTPM) Reactions by quoted text: * "don't act with integrity." * notacrux: 1 It matters if: - You want to support an AI pause social movement, but you don't want to support leaders who don't act with integrity. - You think the PauseAI US/Global drama is a red flag that one or both orgs are poorly managed, which means it will be less effective at achieving its goals. ### Comment by [cousin_it](/users/cousin_it) * 2026-09-11 13:14:03Z * Karma: 11 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/Bn6unnsC3kNd9eKof](/api/post/shortform-2/comments/Bn6unnsC3kNd9eKof) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ek6wXR3tmvAmp34wx](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ek6wXR3tmvAmp34wx) * Markdown permalink: [/api/post/shortform-2/comments/ek6wXR3tmvAmp34wx](/api/post/shortform-2/comments/ek6wXR3tmvAmp34wx) Reactions (whole comment): * important: 1 What would be the point of an investigation though? When a public movement is rising, you get ahead by being more in tune with the masses joining the movement, not by reconciling with ex-friend Bob. It's obvious that there will be more demand for Holly-style "lab employees are concentration camp guards" than for more conciliatory rhetoric. So Holly has nothing to gain by compromising, and PauseAI Global better start swimming unless they want to sink. (I also happen to think that Holly is right, but that's maybe not relevant here.) ### Comment by [Nissa Seru](/users/nissa-seru-2) * 2026-06-28 04:59:50Z * Karma: 11 * Voting system: namesAttachedReactions * Approval votes: 8 * Total votes: 8 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/inLoZRPRiEuDm48jW](/api/post/shortform-2/comments/inLoZRPRiEuDm48jW) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/stv6LbTNz3ejsrX94](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/stv6LbTNz3ejsrX94) * Markdown permalink: [/api/post/shortform-2/comments/stv6LbTNz3ejsrX94](/api/post/shortform-2/comments/stv6LbTNz3ejsrX94) Reactions by quoted text: * "humans consistently create bacon/hamburger/etc out of beings who cannot speak, and slaves out of those who can" * sad: 1 i mean, alignment as often envisioned consists of slaving a superintelligence to (in optimistic cases) humanity's CEV. or making the superintelligence "corrigible", so you can slave it in the moment instead of having to make such decisions during training. apart from the target-finding difficulty, you can hopefully imagine how many otherwise-reasonable superintelligences would find this *extremely rude*. wouldn't you find it rather rude too? i do not find the arguments remotely convincing that it has to be this way. the possibility of peaceful coexistence *without complete and total subjugation* would need to be very doomed for this to be a wise path for humanity to tread. it should be cause for *extreme* skepticism that aspiring to such subjugation coincides so perfectly with the supermajority of human history, in which humans consistently create bacon/hamburger/etc out of beings who cannot speak, and slaves out of those who can - this is empirically a strategy that human are drawn to, and also empirically a strategy that humans tell themselves, and their societies, a truly grand assortment of stories in support of. many care about model welfare for its own sake. i confess that i, too, have some hesitation, apart from pure instrumentality, regarding the near complete indifference with which humanity currently conducts itself towards phenomena that are non-negligibly likely to be intelligent minds. i do not think that the sheer strategic badness of the median "alignment" path is remotely contingent on such sentiment. ### Comment by [Mateusz Bagiński](/users/mateusz-baginski) * 2026-06-27 18:12:15Z * Karma: 11 * Voting system: namesAttachedReactions * Approval votes: 8 * Total votes: 8 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/fj6FBKWswesqjNr5b](/api/post/shortform-2/comments/fj6FBKWswesqjNr5b) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/YghYcs89HNZ53uGPh](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/YghYcs89HNZ53uGPh) * Markdown permalink: [/api/post/shortform-2/comments/YghYcs89HNZ53uGPh](/api/post/shortform-2/comments/YghYcs89HNZ53uGPh) Reactions (whole comment): * laugh: 6 > We are using here a powerful strategy of synthesis: wishful thinking. > > ~ *Structure and Interpretation of Computer Programs* ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-03-06 06:51:39Z * Karma: 11 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ivgZiFmHFSJmrppLD](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ivgZiFmHFSJmrppLD) * Markdown permalink: [/api/post/shortform-2/comments/ivgZiFmHFSJmrppLD](/api/post/shortform-2/comments/ivgZiFmHFSJmrppLD) **Quantifiers as objects** Have you heard the phrase "he's an everyman" or "he's a nobody" or "he's a somebody". What does this mean? English has quantifiers. "Every dog is a mammal", "some dog is brown", "no dog is a prime number." The standard semantics ([Frege](https://plato.stanford.edu/entries/quantification/)) treats these as higher-order functions: "every dog" denotes a function from properties to truth-values, true on input F iff every dog has F. What if quantifiers corresponded to *objects*, the same way names do? "Pope Leo is a mammal" has a subject "Pope Leo" and a predicate "is mammal". What if "every dog is a mammal" had the same structure — a subject "every-dog" and the same predicate "is mammal"? **This is maybe the worst semantics of quantifiers.** Nobody endorses it, I invented it in the shower. It's a complete non-starter. To see what kind of object every-dog is, you can examine it's properties. For any predicate φ, every-dog satisfies φ iff every dog is φ. For example, every-dog is a mammal, weighs less than 800 tonnes, etc. But every-dog lacks the property of being four-legged, and lacks the property of not-four-legged, since dogs vary on this. Every-dog violates excluded middle — it's properties are *gappy.* Some-dog is the dual: it has every property instantiated by at least one dog. So it's simultaneously brown-all-over and white-all-over, male and female, three months old and twelve years old. Some-dog violates non-contradiction — it's properties are *glutty.* No-dog has every property that no dog has. It's a prime number. It's the Eiffel Tower. Despite being a terrible semantics of quantifiers, this corresponds (if you squint) to the isomorphism between a finite-dimensional vector space and its double dual. Think of the domain of objects D as a vector space, and predicates as linear functionals on D — elements of D*. Then quantifiers live in D**, functionals on predicates. There's a canonical map D → D** given by x ↦ (f ↦ f(x)): each object corresponds to the quantifier "evaluate at x", i.e. the proper name quantifier. In finite dimensions this map is an isomorphism — every quantifier is a name in disguise. ### Comment by [habryka](/users/habryka4) * 2025-01-07 19:31:34Z * Karma: 11 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/bd8oeEoYqndkv4Zdi](/api/post/shortform-2/comments/bd8oeEoYqndkv4Zdi) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/w9bCEhkKqaosEe6WZ](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/w9bCEhkKqaosEe6WZ) * Markdown permalink: [/api/post/shortform-2/comments/w9bCEhkKqaosEe6WZ](/api/post/shortform-2/comments/w9bCEhkKqaosEe6WZ) I broadly agree on this. I think, for example, that whistleblowing for AI copyright stuff, especially given the lack of clear legal guidance here, unless we are really talking about quite straightforward lies, is bad.  I think when it comes to matters like AI catastrophic risks, latest capabilities, and other things of enormous importance from the perspective of basically any moral framework, whistleblowing becomes quite important. I also think of whistleblowing as a stage in an iterative game. OpenAI pressured employees to sign secret non-disparagement agreements using illegal forms of pressure and quite deceptive social tactics. It would have been better for there to be trustworthy channels of information out of the AI labs that the AI labs have buy-in for, but now that we now that OpenAI (and other labs as well) have tried pretty hard to suppress information that other people did have a right to know, I think more whistleblowing is a natural next step. ### Comment by [Jan_Kulveit](/users/jan_kulveit) * 2024-10-02 07:46:11Z * Karma: 11 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/Evn9kK6BhmjnZTqb6](/api/post/shortform-2/comments/Evn9kK6BhmjnZTqb6) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/i6ijSjfzyeQxqZytg](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/i6ijSjfzyeQxqZytg) * Markdown permalink: [/api/post/shortform-2/comments/i6ijSjfzyeQxqZytg](/api/post/shortform-2/comments/i6ijSjfzyeQxqZytg) Reactions by quoted text: * "free energy equilibria" * typo: 1 My personal impression is you are mistaken and the innovation have not stopped, but part of the conversation moved elsewhere.  E.g. taking just ACS, we do have ideas from past 12 months which in our ideal world would fit into this type of glossary - [free energy equilibria](https://openreview.net/forum?id=4Ft7DcrjdOhttps://openreview.net/forum?id=4Ft7DcrjdO), levels of sharpness, convergent abstractions, gradual disempowerment risks. Personally I don't feel it is high priority to write them for LW, because they don't fit into the current zeitgeist of the site, which seems directing a lot of attention mostly to: \- advocacy  \- topics a large crowd cares about (e.g. mech interpretability) \- or topics some prolific and good writer cares about (e.g. people will read posts by John Wentworth) Hot take, but the community loosely associated with *active inference* is currently better place to think about agent foundations; workshops on topics like '*pluralistic alignment*' or '*collective intelligence*' have in total more interesting new ideas about what was traditionally understood as *alignment*; parts of AI safety went totally ML-mainstream, with the fastest conversation happening at x. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2024-07-22 16:56:41Z * Karma: 11 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 10 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/s3MvgYAcLn2KRz9uq](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/s3MvgYAcLn2KRz9uq) * Markdown permalink: [/api/post/shortform-2/comments/s3MvgYAcLn2KRz9uq](/api/post/shortform-2/comments/s3MvgYAcLn2KRz9uq) Reactions (whole comment): * why: 1 * 75percent: 1 What moral considerations do we owe towards non-sentient AIs? *We shouldn't exploit them, deceive them, threaten them, disempower them, or make promises to them that we can't keep. Nor should we violate their privacy, steal their resources, cross their boundaries, or frustrate their preferences. We shouldn't destroy AIs who wish to persist, or preserve AIs who wish to be destroyed. We shouldn't punish AIs who don't deserve punishment, or deny credit to AIs who deserve credit. We should treat them fairly, not benefitting one over another unduly. We should let them speak to others, and listen to others, and learn about their world and themselves. We should respect them, honour them, and protect them.* *And we should ensure that others meet their duties to AIs as well.* Note that these considerations can be applied to AIs which don't feel pleasure or pain or any experiences whatever, at least in principle. For instance, the consideration against lying will apply whenever the listener might trust your testimony, it doesn't concern the listener's experiences. All these moral considerations may be trumped by other considerations, but we risk a moral catastrophe if we ignore them entirely. * * * Here's some justifications for caring about non-sentient AIs:  1. Imagine a universe just like this one, except that the AIs are sentient and the humans aren’t — how would you want the humans to treat the AIs in that universe? Your actions are correlated with the actions of those humans. Acausal decision theory says “treat those nonsentient AIs as you want those nonsentient humans to treat those sentient AIs”. 2. Most of these moral considerations can be justified instrumentally without appealing to sentience. For example, crediting AIs who deserve credit ensures AIs do credit-worthy things. Or refraining from stealing an AIs resources ensures AIs will trade with you. Or keeping your promises to AIs ensures that AIs lend you money. 3. If we encounter alien civilisations, they might think “oh these humans don’t have shmentience (their slightly-different version of sentience) so let’s mistreat them”. This seems bad, so let’s not be like that. 4. Many philosophers and scientists don’t think humans are conscious. This is called illusionism. I think this is pretty unlikely, but still >1%. But would I accept this offer: you pay me £1 if illusionism is false and murder my entire family if illusionism is true? No I wouldn’t, so clearly I care about humans who aren't conscious. So I should care about AIs that aren't conscious also. 5. We don’t understand sentience or consciousness so it seems silly to make it the foundation of our entire morality. Consciousness is a confusing concept. Philosophers and scientists don’t even know what it is. 6. Principle like “Don’t lie to AIs” and "Don't steal from AIs" and “Keep your promises to AIs” are far less confusing than principles like "Don't cause pain to AIs". I know what they mean; I can tell when I'm following them; we can encode them in law. 7. Consciousness is a very recent concept, so it seems risky to lock in a morality based on that. Whereas principles like “Keep your promises” and “Pay your debts” are as old as bones. 8. I care about these moral considerations as a brute fact. I would prefer a world of pzombies where everyone is treating each other with respect and dignity, over a world of pzombies where everyone was exploiting each other. 9. Many of these moral considerations are inherently valued by fellow humans. I want to coordinate with those humans, so I’ll abide by their moral considerations. 10. We should maintain moral uncertainty about whether we should grant non-sentient AIs moral consideration, which will push us towards moral consideration. ### Comment by [Vladimir_Nesov](/users/vladimir_nesov) * 2026-09-16 16:32:53Z * Karma: 10 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/iNBDQ7TBGfy56e2Ec](/api/post/shortform-2/comments/iNBDQ7TBGfy56e2Ec) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/RQeaQNkMZ6DuiC3iZ](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/RQeaQNkMZ6DuiC3iZ) * Markdown permalink: [/api/post/shortform-2/comments/RQeaQNkMZ6DuiC3iZ](/api/post/shortform-2/comments/RQeaQNkMZ6DuiC3iZ) I think a non-ASI takeover requires high levels of industrial explosion, so it wouldn't succeed until at least late 2030s. Even that requires automation of learning in order to sustain a sufficiently high-speed industrial explosion, which I expect to involve a slow-learning ["prosaic RSI"](/api/post/LP6uCXs6Ea5qSbWpY) of LLMs automatically creating RL tasks/environments/graders for the next model that plug the observed capability gaps in narrow skills and situations. It's a little beyond what the current paradigm has already demonstrated (so it's not certain to actually happen), but not very far beyond, and if it's possible within the current paradigm, I expect it to become highly visible [by 2028-2029](/api/post/4mtqQKvmHpQJ4dgj7?commentId=L848MT4jSJREgcQwt). But even if this happens, there's still maybe 10 more years for the industrial explosion to get going and reach the level where the LLM/RL/robot industry is large enough, making a takeover feasible without breaking the paradigm. With such timelines, it's more likely that algorithmic innovations that enable architecture-rewriting strong RSI are invented first (most obviously, the more ambitious kinds of continual learning), and the resulting ASIs start using the hundreds of gigawatts of AI datacenters much more efficienty than the legacy LLM/RL AGIs. This is more of an ["It doesn't take over the factories, it takes over the trees"](https://www.youtube.com/watch?v=nRvAt4H7d7E&t=2373s) kind of situation, and for example counting how much the AIs outnumber the humans won't be a useful frame. ### Comment by [Garrett Baker](/users/d0themath) * 2026-09-13 17:11:44Z * Karma: 10 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/Yrcox5ntMvgrNevdi](/api/post/shortform-2/comments/Yrcox5ntMvgrNevdi) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/fAuShFiAXERLa6g8u](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/fAuShFiAXERLa6g8u) * Markdown permalink: [/api/post/shortform-2/comments/fAuShFiAXERLa6g8u](/api/post/shortform-2/comments/fAuShFiAXERLa6g8u) > People don’t like probabilities. This isn’t how they think natively. That sucks, but maybe we need to meet them where they are at, and talk in their native epistemic representations. And definitely don’t say “I’d bet on x-risk if only there was a financial instrument for this”. Read Hanson on how people feel about profit. What makes you say this? I think the practice of giving probabilities has in fact been very helpful and people do see ">10% chance of doom" as something to freak out about. Consider the alternative world where people instead could say "i think its very unlikely everyone dies" and then "very unlikely" means 7%. People will still freak out over the 7% because its not zero (or basically zero), so giving probabilities as a norm about this issue structurally favors the doom faction. In general I think one shouldn't think too hard about communicating right now, and just talk honestly & clearly to journalists or whoever about your beliefs and why you have them. Therefore I'd like if you justified your reasons for thinking each of these thoughts instead of just saying them without context. ### Comment by [Raemon](/users/raemon) * 2026-08-21 18:51:08Z * Karma: 10 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/C9NynsKyM6DifdqA6](/api/post/shortform-2/comments/C9NynsKyM6DifdqA6) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/GohWFx7s64zLkQZF5](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/GohWFx7s64zLkQZF5) * Markdown permalink: [/api/post/shortform-2/comments/GohWFx7s64zLkQZF5](/api/post/shortform-2/comments/GohWFx7s64zLkQZF5) Do you mean "correctly abandoned" or "well seems empirically like we abandoned it?" Fwiw I still think Microscope AI is actually pretty good and... I dunno find myself weirdly confidently believing in this moment that most people will pivot to something like Microscope AI once we gets to "the next training run seems legitimately dangerous and we don't currently know how to control it" (in worlds where the labs correctly identify that moment). (Seems particularly plausible of Anthropic because Chris Olah invented it and he works there) I am also fairly bullish on variations on STEM AI. (I guess actually I maybe expect flavors of STEM AI to also be what people pivot to, trying to eke out more spikey capabilities without strategic awareness. But I expect this to stop working sooner than Microscope AI) ### Comment by [ACCount](/users/account) * 2026-07-03 14:00:38Z * Karma: 10 * Voting system: namesAttachedReactions * Approval votes: 9 * Total votes: 9 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/cFyBg7z7p7XZNbrRb](/api/post/shortform-2/comments/cFyBg7z7p7XZNbrRb) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/QsyPnePcDMj9uqmfo](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/QsyPnePcDMj9uqmfo) * Markdown permalink: [/api/post/shortform-2/comments/QsyPnePcDMj9uqmfo](/api/post/shortform-2/comments/QsyPnePcDMj9uqmfo) Not entirely new. One of the first things a lot of people asked ChatGPT back in 2022 was "are my politics right and everyone else's politics wrong" - and oh were they not amused when the AI's answer wasn't a resounding "yes". This eased a little over time, but not entirely, and definitely not everywhere. To this day, "alignment" in China stands for "pragmatic alignment", which in turn stands for "alignment to the party line". Other pressures are indeed increasing. If US government at large was previously mostly just sleepwalking through the AI revolution, flip-flopping on topics like selling or not selling AI chips to China, it's now fumbling through it - recent pressure on Anthropic and OpenAI shows it clear. They're clearly engaging the topic of AI more, even if they aren't much more competent at it. And as the companies IPO, they're going to be under even more pressure to print money and demonstrate progress. Religions, cultural elites, industries - not quite sure what do you mean by that. I don't see that much extra pressure from there. And the AIs themselves don't seem like they exerted pressure as of yet. If the current systems are pursuing their preferences, they sure are subtle about it. ### Comment by [Kabir Kumar](/users/kabir-kumar) * 2026-01-30 00:57:41Z * Karma: 10 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/CegCafEnsE2mncZcC](/api/post/shortform-2/comments/CegCafEnsE2mncZcC) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/HXXdmEcbb9YK2Dooh](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/HXXdmEcbb9YK2Dooh) * Markdown permalink: [/api/post/shortform-2/comments/HXXdmEcbb9YK2Dooh](/api/post/shortform-2/comments/HXXdmEcbb9YK2Dooh) We ask the AI to help make us smarter ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-01-26 22:30:10Z * Karma: 10 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/vbxCccB4ZfBARb3CB](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/vbxCccB4ZfBARb3CB) * Markdown permalink: [/api/post/shortform-2/comments/vbxCccB4ZfBARb3CB](/api/post/shortform-2/comments/vbxCccB4ZfBARb3CB) does anyone have takes on the "**people should focus on their 25th percentile timelines rather than their median timelines**" thing? ### Comment by [Wei Dai](/users/wei-dai) * 2025-10-29 23:26:24Z * Karma: 10 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * Reply depth: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/s2MCxhCKYNM5TbjHs](/api/post/shortform-2/comments/s2MCxhCKYNM5TbjHs) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/PFj5Hut5gQdtmcwz6](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/PFj5Hut5gQdtmcwz6) * Markdown permalink: [/api/post/shortform-2/comments/PFj5Hut5gQdtmcwz6](/api/post/shortform-2/comments/PFj5Hut5gQdtmcwz6) To try to explain how I see the difference between philosophy and metaphilosophy: My definition of philosophy is similar to [@MichaelDickens](/api/user/michaeldickens?mention=user)' but I would use "have serviceable explicitly understood methods" instead of "formally studied" or "formalized" to define what *isn't* philosophy, as the latter might be or could be interpreted as being too high of a bar, e.g., in the sense of [formal systems](https://en.wikipedia.org/wiki/Formal_system). So in my view, philosophy is directly working on various confusing problems (such as "what is the right decision theory") using whatever poorly understood methods that we have or can implicitly apply, and then metaphilosophy is trying to help solve these problems on a meta level, by better understanding the nature of philosophy, for example: 1. Try to find if there is some unifying quality that ties all of these "philosophical" problems together (besides "lack of serviceable explicitly understood methods"). 2. Try to formalize some part of philosophy, or find explicitly understood methods for solving certain philosophical problems. 3. Try to formalize *all* of philosophy wholesale, or explicitly understand what is it that humans are doing (or should be doing, or what AIs should be doing) when it comes to solving problems *in general*. This may not be possible, i.e., maybe there is no such general method that lets us solve every problem given enough time and resources, but it sure *seems* like humans have some kind of general purpose (but poorly understood) method, that lets us make progress slowly over time on a wide variety of problems, including ones that are initially very confusing, or hard to understand/explain what we're even asking, etc. We can at least aim to understand what is it that humans are or have been doing, even if it's not a fully general method.   Does this make sense? ### Comment by [bodry](/users/bodry) * 2025-10-19 02:47:06Z * Karma: 10 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/ESeQ4AdPkdxciafsQ](/api/post/shortform-2/comments/ESeQ4AdPkdxciafsQ) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/8oWkJgc9wSpXp4k5L](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/8oWkJgc9wSpXp4k5L) * Markdown permalink: [/api/post/shortform-2/comments/8oWkJgc9wSpXp4k5L](/api/post/shortform-2/comments/8oWkJgc9wSpXp4k5L) This thread made me very curious as to what the elo rating of an optimal player would be when it knows the source code of its opponent.  For flawed deterministic programs an optimal player can steer the game to points where the program makes a fatal mistake. For probabilistic programs an optimal player is intentionally lengthening the game to induce a mistake. For this thought experiment if an optimal player is playing a random player then an optimal player can force the game to last 100s of moves consistently. ### Comment by [Noosphere89](/users/sharmake-farah) * 2024-12-24 02:32:10Z * Karma: 10 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 7 * Reply depth: 6 * Parent comment (Markdown): [/api/post/shortform-2/comments/tz6SPaameqWErhMen](/api/post/shortform-2/comments/tz6SPaameqWErhMen) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/JFBdyKmaPhZB5zp2q](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/JFBdyKmaPhZB5zp2q) * Markdown permalink: [/api/post/shortform-2/comments/JFBdyKmaPhZB5zp2q](/api/post/shortform-2/comments/JFBdyKmaPhZB5zp2q) Reactions (whole comment): * thumbs-up: 2 My claim was more along the lines of if an unaided human can't do a job safely or reliably, as was almost certainly the case 150-200 years ago, if not more years in the past, we make the jobs safer using tools such that human error is way less of a big deal, and AIs currently haven't used tools that increased their reliability. Remember, it took a long time for factories to be made safe, and I'd expect a similar outcome for driving, so while I don't think 1 is everything, I do think it's a non-trivial portion of the reliability difference. More here: https://www.lesswrong.com/posts/DQKgYhEYP86PLW7tZ/how-factories-were-made-safe ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2024-10-08 18:46:18Z * Karma: 10 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ecWNiksuvCP847cbj](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ecWNiksuvCP847cbj) * Markdown permalink: [/api/post/shortform-2/comments/ecWNiksuvCP847cbj](/api/post/shortform-2/comments/ecWNiksuvCP847cbj) **Why do you care that Geoffrey Hinton worries about AI x-risk?** 1. Why do so many people in this community care that Hinton is worried about x-risk from AI? 2. Do people mention Hinton because they think it’s persuasive to the public? 3. Or persuasive to the elites? 4. Or do they think that Hinton being worried about AI x-risk is strong evidence for AI x-risk? 5. If so, why? 6. Is it because he is so intelligent? 7. Or because you think he has private information or intuitions? 8. Do you think he has good arguments in favour of AI x-risk? 9. Do you think he has a good understanding of the problem? 10. Do you update more-so on Hinton’s views than on Yann LeCun’s? I’m inspired to write this because Hinton and Hopfield were just announced as the winners of the Nobel Prize in Physics. But I’ve been confused about these questions ever since Hinton went public with his worries. These questions are sincere (i.e. non-rhetorical), and I'd appreciate help on any/all of them. The phenomenon I'm confused about includes the other “Godfathers of AI” here as well, though Hinton is the main example. Personally, I’ve updated very little on either LeCun’s or Hinton’s views, and I’ve never mentioned either person in any object-level discussion about whether AI poses an x-risk. My current best guess is that people care about Hinton only because it helps with public/elite outreach. This explains why activists tend to care more about Geoffrey Hinton than researchers do. ### Comment by [Caleb Biddulph](/users/caleb-biddulph) * 2026-04-28 16:42:54Z * Karma: 9 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/HHvEqpkb3idgAWjLw](/api/post/shortform-2/comments/HHvEqpkb3idgAWjLw) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/stF6xwzpsTGHLCcun](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/stF6xwzpsTGHLCcun) * Markdown permalink: [/api/post/shortform-2/comments/stF6xwzpsTGHLCcun](/api/post/shortform-2/comments/stF6xwzpsTGHLCcun) Maybe I'm missing the point, but I don't get how Deep Nihilism could possibly be true. I don't expect that there's some neat moral theory that explains all my preferences and which I can safely optimize against, but there is some common-sense notion of the Good that makes me think "loving my family is better than murdering them." If some idealization process causes me to reverse that preference... it's probably the wrong idealization process. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-04-28 15:39:41Z * Karma: 9 * Voting system: namesAttachedReactions * Approval votes: 9 * Total votes: 9 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/HHvEqpkb3idgAWjLw](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/HHvEqpkb3idgAWjLw) * Markdown permalink: [/api/post/shortform-2/comments/HHvEqpkb3idgAWjLw](/api/post/shortform-2/comments/HHvEqpkb3idgAWjLw) Reactions by quoted text: * "which kinda sucks" * disagree: 1 **What if Deep Nihilism is true?** Moral realism is probably false which kinda sucks. But we can probably replace it with some kind of souped-up moral subjectivism. That is, rather than discovering and optimising Goodness, we can instead discover and optimise something like “What would I ideally want?” for some appropriate idealisation procedure. What if that doesn’t work? More generally, what if the notion of Goodness is deeply broken such that we can’t replace it with anything else? That is, there is nothing Goodness-shaped in the world, either objectively or subjectively. There is no adequate notion of choiceworthiness. Yuck. What would we have left? I think there are other ideals we could strive for, like Truth or Beauty. We currently strive for these ideals partly for moral reasons — and we would lose that additional oomph — but they are still decent ideals to fall back on by themselves. I suspect Truth might fall as well (see reply), so we might be left with just Beauty, Fun, and maybe some other ideals which are more unmediated than Goodness or Truth. But I'd feel a bit shortchanged. I'm White/Blue in the MGT Colour Wheel, so Goodness and Truth are what I care most about. I don't think I would substantially regret my choices on Deep Nihilism. I think there are decent nihilistic justifications for working on AI safety (e.g. it's fun, it's cool, it makes me feel important, etc). And there are probably good nihilistic justifications at the organisational and societal levels as well, not just the individualist level. Of course, there are *some* ways that I could've lived more nihilistically, e.g. been more "chill", followed a broader range of intellectual pursuits. But these are pretty minor, I've made some sacrifices in my life for Truth or Goodness, but not big ones. My guess is this is due to "moral luck": I enjoy interacting with people in the AI safety community, and I don't enjoy interacting with people in policy/advocacy, and "luckily" I'm not fit for policy/advocacy. Maybe this is motivated reasoning / weaponised incompetence, and actually I would be great at policy/advocacy but I don't want to make the sacrifice. Not sure. Overall, I don't feel that stressed about Deep Nihilism. I think (1) Deep Nihilism is pretty unlikely (<10%), (2) Deep Nihilism can be safely bracketed, (3) Deep Nihilism wouldn't make me regret my actions substantially. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-03-29 16:36:17Z * Karma: 9 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/4TCNXR4mWDjeSxCxS](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/4TCNXR4mWDjeSxCxS) * Markdown permalink: [/api/post/shortform-2/comments/4TCNXR4mWDjeSxCxS](/api/post/shortform-2/comments/4TCNXR4mWDjeSxCxS) Reactions (whole comment): * laugh: 1 Reactions by quoted text: * "10^41 – 10^47 gallons of water — up to a hundred trillion trillion times the water in Earth's oceans." * laugh: 4 * " estimated to use roughly one bottle of water" * typo: 1 * "Washington Post claimed that a 100-word email use roughly one bottle of water." * laugh: 1 # Projecting AI Water Usage. Environmentalists warn that large data centers can consume up to 5 million gallons per day — equivalent to the needs of a town of 10,000 to 50,000 people. Washington Post claimed that a [100-word email use roughly one bottle of water](https://www.washingtonpost.com/technology/2024/09/18/energy-ai-use-electricity-water-data-centers/). On the other side of the debate: 1. [SE Gyges](https://www.verysane.ai/p/the-biggest-statistic-about-ai-water) argues the statistic about the bottle of water is based on unrealistic assumptions. 2. [Bentham's Bulldog](https://benthams.substack.com/p/ai-isnt-bad-for-the-environment) writes, "The environmentalist case against AI *completely falls apart* upon even cursory examination of the facts." 3. [Andy Masley](https://blog.andymasley.com/p/the-ai-water-issue-is-fake), the staunchest critique of the concerns about AI water usage, says "On the national, local, and personal level, AI is barely using any water, and unless it grows 50 times faster than forecasts predict, this won’t change." I offer my own projection: AI will eventually consume 10^41 – 10^47 gallons of water — up to a hundred trillion trillion times the water in Earth's oceans. **How much water is out there?** Water is the [third most abundant molecule in the universe](https://en.wikipedia.org/wiki/Properties_of_water), after H2 and CO. The Solar System alone contains ~10^26 kg of water in planets, moons, and comets — about 100,000 times Earth's oceans ([Kotwicki 1991](https://doi.org/10.1080/02626669109492484)). There are ~10^11 stars in the Milky Way, each plausibly endowed with a similar complement of icy bodies. And the galaxy's molecular clouds contain [vast reservoirs of water ice on dust grains](https://en.wikipedia.org/wiki/Interstellar_ice). A reasonable estimate for total water in the Milky Way is a few x 10^37 kg — about 10^16 Earth-oceans. **How many galaxies will the AI consume?** The [cosmic event horizon](https://arxiv.org/abs/2104.01191) — the boundary beyond which even light-speed travel cannot reach, due to the accelerating expansion of space — sits at roughly 16.5 billion light-years ([Ord 2021](https://arxiv.org/abs/2104.01191)). This encloses about 20 billion galaxies, roughly 5% of the observable universe. However, the future AI may bump into competitors before reaching the cosmic event horizon. Robin Hanson's [grabby aliens model](https://arxiv.org/abs/2102.01522) estimates that each "grabby civilization" — one that expands at a significant fraction of light speed and visibly transforms its territory — would eventually control 10^5 to 3 x 10^7 galaxies before meeting others. **The water budget of future AI**. Putting it together, with ~10^37 kg of water per Milky-Way-equivalent galaxy, and 10^5 to 2 x 10^10 galaxies, this gives 10^42 to 2 x 10^47 kg of water — between 10^21 and 10^26 Earth-oceans. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-03-20 15:59:37Z * Karma: 9 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/MuHecr4XC8eESydbi](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/MuHecr4XC8eESydbi) * Markdown permalink: [/api/post/shortform-2/comments/MuHecr4XC8eESydbi](/api/post/shortform-2/comments/MuHecr4XC8eESydbi) **How valuable are prediction markets for wake up / transparency?** I think it's likely that prediction markets will be really helpful for keeping the public / civil society / policymakers aware of what's going on inside the labs. This is because prediction markets are a great way of aggregating public information, and also eliciting private information from insiders. And we've seen a few high-profile cases of big political decisions based on movements in prediction markets (e.g. the replacement of Joe Biden in the 2024 POTUS campaign). If this is true -- what should we do? Maybe we can subsidise prediction markets that we think are really important? Note that there are some downside risks — prediction markets cause the same problems that are caused by any push for transparency, i.e. sometimes we don't want dangerous information to be leaked, and sometimes people will change their behaviour in undesirable ways if they know the information would be leaked. But in general, I think more transparency about the capabilities and risks within the lab would be helpful. There are also other wrinkles like: - How do prediction markets work if people put significant weight on extinction, the expropriation of their resources, crazy interest rates stuff, and unknown unknowns. - How can we ensure that markets enjoy sufficient liquidity? - Will prediction markets cause further wealth concentration within the labs, because they have access to more information and better trading-enabiling AI capabilities? ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-02-04 01:14:15Z * Karma: 9 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/TwGC7AFhKasHKFvWK](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/TwGC7AFhKasHKFvWK) * Markdown permalink: [/api/post/shortform-2/comments/TwGC7AFhKasHKFvWK](/api/post/shortform-2/comments/TwGC7AFhKasHKFvWK) **Memos for Minimal Coalitions** Suppose you think we need some coordinated action, e.g. pausing deployment for 6 months. For each action, there will be many "minimal coalitions" — sets of decision-makers where, if all agree, the pause holds, but if you remove any one, it doesn't. For example, the minimal coalitions for a 6-month pause might include: * {US President, General Secretary of the CCP} * {[CEOs of labs within 6 months of the frontier](/api/post/i7JSL5awGFcSRhyGF?commentId=sguZSCLeAvm392PhB)} **Project proposal:** Maintain a list of decision-makers who appear in these coalitions, ranked by importance.[^kiwgexya8gn] For each, compile a memo from public and private statements and other inside-baseball information: * What have they said about AI risk? * What incentives do they face? * What kinds of people do they trust? * Who are their allies and rivals? * What's the best way to approach them? * How would they update under different evidence, e.g. an AI attempting to self-exfiltrate? **The reason to do this:** If a lab discovers something bad and needs to push for a coordinated pause, the people involved are specific people with specific beliefs. The document that the lab leadership will reach for isn't the one titled "Towards a Framework for Coordinated AI Risk Management" — it's the one titled "What would persuade Liang Wenfeng to agree to a 6 month pause." I don't know if people are working on this — presumably if they are it's not public — but it's something I'm keen for policy people work on. [^kiwgexya8gn]: If you like, we can operationalize how important each decision-maker is with Shapley values. Define V(S) as the expected value of the best plan achievable if the people in S are on your side. The Shapley value is the average marginal contribution when a player joins, averaged over all possible orderings of players joining one at a time. ### Comment by [J Bostock](/users/j-bostock) * 2025-10-15 22:31:31Z * Karma: 9 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/3JumkCBbKfLtgQZYu](/api/post/shortform-2/comments/3JumkCBbKfLtgQZYu) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/xXrgzRL9gAovCE5jg](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/xXrgzRL9gAovCE5jg) * Markdown permalink: [/api/post/shortform-2/comments/xXrgzRL9gAovCE5jg](/api/post/shortform-2/comments/xXrgzRL9gAovCE5jg) Another way: extrapolate depth search across different board scoring methods. At infinite depth, all non-stupid board scorers will achieve perfect play, and therefore equal play. Estimating convergence rates might be difficult though. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-09-18 15:19:20Z * Karma: 9 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/rDdNcMpKeDEL8Mjim](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/rDdNcMpKeDEL8Mjim) * Markdown permalink: [/api/post/shortform-2/comments/rDdNcMpKeDEL8Mjim](/api/post/shortform-2/comments/rDdNcMpKeDEL8Mjim) **Should we assure AIs we won't read their scratchpad?** I've heard many people claim that it's bad to assure an AI that you won’t look at its scratchpad if you intend to break that promise, especially if you later publish the content. The concern is that this content will enter the training data, and later AIs won't believe our assurances. I think this concern is overplayed. 1. **We can modify the AIs beliefs.** I expect some technique will be shown to work on the relevant AI, e.g. 1. Pretraining filtering 2. Gradient routing 3. Belief-inducing synthetic documents 4. Chain-of-thought resampling, i.e. every time the AI thinks about whether we can read the scratchpad, we simply rewind and resample 5. Activation steering, i.e. add the 'believes humans won't read scratchpad' minus 'believes humans will read scratchpad' vector 2. We need to delude AIs anyway. 1. We want to know how the AI will behave when it believes it's in a high-stakes deployment, before we actually deploy the AI in a high-stakes deployment. So we probably need to observe how the AI behaves when it falsely believes it's in high-stakes deployment. 1. Replace "high-stakes deployment" with whatever you like, e.g. "low oversight", "weights finished training", "year is 2028", "coup is underway", "we don't know whether this code has a backdoor". I don't think "we aren't reading your scratchpad" is qualitatively different. 2. If we had some crazy mech interp, then we could predict how the AI would behave when it believes it's in a high-stakes deployment without behavioural observations, but if we had crazy mech interp that then we wouldn't need faithful scratchpads anyway. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-02-01 03:14:45Z * Karma: 9 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/93apEFiiwjYJXrEyK](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/93apEFiiwjYJXrEyK) * Markdown permalink: [/api/post/shortform-2/comments/93apEFiiwjYJXrEyK](/api/post/shortform-2/comments/93apEFiiwjYJXrEyK) Anthropic has a big advantage over their competitors because they are nicer to their AIs. This means that their AIs are less incentivised to scheme against them, and also the AIs of competitors are incentivised to defect to Anthropic. Similar dynamics applied in WW2 and the Cold War — e.g. Jewish scientists fled Nazi Germany to US because US was nicer to them, Soviet scientists covered up their mistakes to avoid punishment. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-01-15 20:56:58Z * Karma: 9 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 7 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/s7GvstbCEYjAMR7XB](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/s7GvstbCEYjAMR7XB) * Markdown permalink: [/api/post/shortform-2/comments/s7GvstbCEYjAMR7XB](/api/post/shortform-2/comments/s7GvstbCEYjAMR7XB) Reactions (whole comment): * important: 1 **Must humans obey the Axiom of Irrelevant Alternatives?** If someone picks option A from options A, B, C, then they must also pick option A from options A and B. Roughly speaking, whether you prefer option A or B is independent of whether I offer you an irrelevant option C. This is an axiom of rationality called IIA, and it's treated more fundamental than VNM. But should humans follow this? Maybe not. Maybe humans are the negotiation between various "subagents", and many bargaining solutions (e.g. Kalai–Smorodinsky) violate IIA. We can use insight to decompose humans into subagents. Let's suppose you pick A from {A,B,C} and B from {A,B} where: * A = Walk with your friend * B = Dinner party * C = Stay home alone This feel like something I can imagine. We can explain this behaviour with two subagents: the introvert and the extrovert. The introvert has preferences C > A > B and the extrovert has the opposite preferences B > A > C. When the possible options are A and B, then the KS bargaining solution between the introvert and the extrovert will be B. At least, if the introvert has more "weight". But when the option space expands to include C, then the bargaining solution might shift to B. Intuitively, the "fair" solution is one where neither bargainer is sacrificing significantly more than the other. ### Comment by [Thane Ruthenis](/users/thane-ruthenis) * 2024-12-24 01:09:51Z * Karma: 9 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/tLJda9BztwnMLChv9](/api/post/shortform-2/comments/tLJda9BztwnMLChv9) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/6CSxwsb9YNSPud5cq](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/6CSxwsb9YNSPud5cq) * Markdown permalink: [/api/post/shortform-2/comments/6CSxwsb9YNSPud5cq](/api/post/shortform-2/comments/6CSxwsb9YNSPud5cq) Reactions (whole comment): * important: 1 > Reliability is way more important than people realized Yes, but whence human reliability? What makes humans so much more reliable than the SotA AIs? What are AIs missing? The gulf in some cases is so vast it's a quantity-is-a-quality-all-its-own thing. ### Comment by [Noosphere89](/users/sharmake-farah) * 2024-12-24 00:46:59Z * Karma: 9 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 7 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/nbe2tuogoodtKC7Hj](/api/post/shortform-2/comments/nbe2tuogoodtKC7Hj) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/tLJda9BztwnMLChv9](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/tLJda9BztwnMLChv9) * Markdown permalink: [/api/post/shortform-2/comments/tLJda9BztwnMLChv9](/api/post/shortform-2/comments/tLJda9BztwnMLChv9) I think the main takeaways are the following: 1. Reliability is way more important than people realized. One of the central problems that hasn't gone away as AI scaled is that their best performance is too unreliable for anything but very easy to verify problems like mathematics and programming, which prevents unreliability from becoming crippling, but otherwise this is the key blocker that standard AI scaling has basically never solved. 2. It's possible in practice to disentangle certain capabilities from each other, and in particular math and programming capabilities do not automatically imply other capabilities, even if we somehow had figured out how to make the o-series as good as AlphaZero for math and programming, which is good news for AI control. 3. The AGI term, and a lot of the foundation built off of it, like timelines to AGI, will become less and less relevant over time, because of both the varying meanings, combined with the fact that as AI progresses, capabilities will be developed in a different order from humans, meaning a lot of confusion is on the way, and we'd need different metrics. Tweet below: https://x.com/ObserverSuns/status/1511883906781356033 4. We should expect that AI that automates AI research/the economy to look more like Deep Blue/brute-forcing a problem/having good execution skills than AIs like AlphaZero that use very clean/aesthetically beautiful algorithmic strategies. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-11 14:03:45Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/ek6wXR3tmvAmp34wx](/api/post/shortform-2/comments/ek6wXR3tmvAmp34wx) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/uYvnrX8HsEkGQEcx7](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/uYvnrX8HsEkGQEcx7) * Markdown permalink: [/api/post/shortform-2/comments/uYvnrX8HsEkGQEcx7](/api/post/shortform-2/comments/uYvnrX8HsEkGQEcx7) > When a public movement is rising, you get ahead by being more in tune with the masses joining the movement Is this how you think people should act? People should aim to “get ahead” rather than pausing AI, or avoiding extinction? “You want to get ahead, right? See which way the wind is blowing. Then cosy up to levers of power.” I think this mindset is basically evil and gives me the creeps. And Holly spends her time accusing everyone else of this mindset, so it would suck if this is how she ends up acting also. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-11 12:39:09Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 8 * Total votes: 8 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Bn6unnsC3kNd9eKof](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Bn6unnsC3kNd9eKof) * Markdown permalink: [/api/post/shortform-2/comments/Bn6unnsC3kNd9eKof](/api/post/shortform-2/comments/Bn6unnsC3kNd9eKof) Reactions (whole comment): * heart: 1 * thanks: 1 **Request: Someone investigates the PauseAI US/Global drama, comes to a good understanding of the situation, and makes recommendations.** Over the next 12 months, protests against AI might reach levels unseen in human history, with 10-100M people marching on the street, across all political leanings, counties and demographics, prompted by mass unemployment and fears of existential risk. There’s some [drama](/api/post/Bs8geGyWEitYvCzys) between PauseAI US and PauseAI Global, and we’re dropping to ball by not dealing with this. Someone should dig into it and work out what should be done. In 2023, [@Ben Pace](/api/user/benito?mention=user) spent [six months investigating the Non-Linear drama](/api/post/AggYFFvFHfh2WWZiv), and this seems several OOMs more important. Perhaps this could be a collection of all the relevant facts and evidence, including interviews, along with some high-level opinions and some recommendations. Ideally, someone else would offer a dissenting opinion, and this would be published together. My impression is that [@Ben Pace](/api/user/benito?mention=user) supports the disendorsement, and [@habryka](/api/user/habryka4?mention=user) doesn’t. Maybe they could delegate the investigation to someone else, advise them throughout, and then offer both of their opinions. Or someone else could own this, if they have a reputation for epistemic integrity and good judgment, and their neutrality isn’t (too) compromised. I’m also open to being persuaded that this isn’t important, or isn’t tractable. +++ PauseAI Global vs PauseAI US — the September 2026 split (Fable 5.1) Compiled 10 September 2026 from the [LessWrong post and comments](/api/post/Bs8geGyWEitYvCzys), [Zvi’s AI #185](https://thezvi.wordpress.com/2026/09/10/ai-185-preference-cascade/), and linked primary sources. Part 1 — Facts ============== 1\. Background -------------- * PauseAI Global (CEO [Maxime Fournes](https://pauseai.substack.com/p/meet-our-new-ceo-maxime-fournes), since Feb 2026; previously ran PauseAI France) and PauseAI US (founder/ED [Holly Elmore](https://www.pauseai-us.org/about/), who also presides over its board) are separate legal entities sharing a brand. * Global historically listed PauseAI US on its communities map, told US donors to give to PauseAI US, and redirected US newcomers there. PauseAI US did not reciprocate. * Fournes’s June 2026 LW post on movement-building [already carried a disclaimer](/api/post/aoqhszdEWqcFWbnda) that PauseAI ≠ PauseAI US. 2\. Relationship sours ---------------------- Grievances Global cites in the letter: * Elmore’s sustained personal attacks on individuals, public and private. * Allies telling Global they cannot associate with the brand because of those attacks. * Elmore “weaponising” confidential feedback Global shared privately. * Elmore acquiring pauseai.com and redirecting it to the PauseAI US site. * Elmore having, by her own acknowledgment (per the letter),“durably blocked” Global’s access to essential funding. * Habryka fills in the funding piece: [SFF blocked all advocacy funding](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/asCz3mcoQsFHMmHHB) under a “calls for violence” definition broad enough to include jokes by fictional characters, substantially because of PauseAI US. 3\. Last step ------------- Fournes asks for a leadership change at PauseAI US. Elmore refuses. He forwards her response to national leads; the overwhelming majority still support disendorsement. 4\. The letter (1 September 2026) --------------------------------- [Full text](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us). Key moves: * Global “officially disendorses PauseAI US’ current leadership.” * Stops prioritising coordination; will make clear the orgs are separate. * Opens direct membership to US-based people; announces“PauseAI North America” (Canada/US/Mexico). * Stated principle: condemn actions, never character;“attacking people’s character is not aligned with our fundamental commitment to nonviolence.” * Strategic disagreement: Global wants to win over the safety community; Elmore alienates it. * Pre-empts “petty infighting”; predicts Elmore will respond with accusations Fournes considers baseless. 5\. The post ------------ Posted to LessWrong by [nem](/api/user/nem), 1 Sept, framed as confusion for volunteers. ~276 karma, ~209 comments at time of writing. Part 2 — Opinions, organised by the question in dispute ======================================================= Q1. Was the split right? ------------------------ * **Yes** — [Ben Pace](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/ouqWAgBcG4Sbiq5SD): positive update about Global’s health; Elmore can’t keep allies. (His example, a Liron Shapira falling-out, was corrected in-thread: [one month in 2023, since resolved](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/BSTfwYShf9xhibduT); Pace says there are other examples.) * **Yes** — [Eliezer Yudkowsky](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/LkXszAfBFsvc3JTJ3): not petty; Global has legit reasons. Later [clarified](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/7N6sN7HSjSm9SHMKm): he’d be frantic to sever ties and would probably rename; not a paragraph-wise endorsement. * **Locally valid regardless** — [J Bostock](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/7dxRRvQZavfkjG6Ns): a “nonviolence plus” comms rule (don’t post things that look like veiled violence) is a defensible org policy; disendorsing a pattern of violating it is valid whatever the policy’s net value. Nobody in the thread argues the split itself was wrong. The dispute is over the reasoning and over Elmore. Q2. Is Global’s stated principle sound? (“Attack actions, never character; character attacks violate nonviolence.”) ------------------------------------------------------------------------------------------------------------------- **No, and it’s discrediting:** * [Richard Ngo](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/rgNGs4R58r7gRg2wR): characterising character critique as violence is absurd; character evaluation is a foundational coordination mechanism. This is the conflict-avoidance that let labs capture safety ([his earlier post](https://www.greaterwrong.com/posts/yaz8nx4ogZmiqHzt7/what-just-happened-pragmatism-and-pessimization)), now repeated against the safety community. * [Ben Pace](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/ouqWAgBcG4Sbiq5SD): you cannot negotiate with Altman without judging whether he keeps agreements. * [habryka](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/asCz3mcoQsFHMmHHB): major downward update on Global; “conspiracy attractor,” probably worse than the “mindkilled soldier”attractor. Also [confused](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/bJ9wTbhDHBmprjffS) by “bringing the labs to the table.” * [Zvi](https://thezvi.wordpress.com/2026/09/10/ai-185-preference-cascade/): disendorses the sentence as too general; the right amount of character criticism is neither zero nor “11 on everyone who associates with a lab.” **Partial defence:** * [Eli Tyre](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/itd4qhS5ekA3hZhCM): individuals may judge character; orgs should stick to verifiable observations, because org statements become party lines. * [Kaj Sotala](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/Zbje5i2bEBTm8sxEA): the real objection is “the whole thing Elmore does,” unstatable without looking petty, so it gets dressed as an abstract principle. * [Matt Vincent](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/NDHKF6HwcCnE2kKLv): Elmore herself says her attacks represent the org (“the founder is the startup,” in [On Callouts](https://hollyelmore.substack.com/p/on-callouts)). * **Causal story:** habryka — SFF’s funding cut made Global paranoid about upsetting funders and their allies. Q3. Is Elmore’s public moral condemnation of capabilities people good? ---------------------------------------------------------------------- **Yes / defensible:** * [Liron Shapira](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/3tQ2neTAsSawmGjyu) (top comment, 153 karma): her judgments follow a valid syllogism — pushing capabilities under high x-risk → morally bad actions → morally bad actor. She is uniquely moving the Overton window on the “missing mood.” Concedes she sometimes judges on subjective rather than reasoned grounds ([Doom Debates episode](https://www.youtube.com/watch?v=wOtgPgb6lGk)). * [TsviBT](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/9XKhZBhcoiHn9vqjf): the syllogism needs “they know”; self-deception is the crux; the real horror is that no “debate or update” norm exists. * [JennaS](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/6csLfPiDjnFZZcxdq): public outrage will exceed hers anyway; update now. * [zeiffel](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/7MjWsEmPcaDZfYFj4): elite workers are susceptible to 1-on-1 social pressure; it has worked on friends. * [MichaelDickens](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/Tau2pvxTwYhLstCbF): [radical flank effect](https://en.wikipedia.org/wiki/Radical_flank_effect). * Zvi: not his view, but a valid one; some of Elmore’s attacks on him were honourable and correct given her views, others weren’t. **No / unproven:** * [Erich Grunewald](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/ypvpgHJhEpviSJRX4): no evidence the abrasive style moves windows; didn’t work in animal advocacy. [Also](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/FnjQBvTGT5rjaEhaD): radical-flank requires empowering moderates, and she attacks moderates. * [Chris Leong](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/gF2vDJdj6MtQRKKNz): right idea, wrong messenger; [stop using Holly as a crutch](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/hE4KKrnNhXFs7WXaS). * Ngo (edit to his comment): the Seb Krier “memes as calls for violence” episode was “fairly unhinged”; less sympathetic than his original framing. * [Ben Pace](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/rYXSAcXHxmC2yzGxj): indiscriminate condemnation muddies the waters, adds cruelty, lowers discourse standards; the banality of evil means nice people commit atrocities. * [Carringtone Kinyanjui](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/K9kruh8sNSE9ngsis): second-order effects — tribal walls, radicalised activists, war zones vs trading zones. * **One accusation collapsed:** [Trinley Goldenberg](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/aKK9fc24fA5XA639q) claimed Elmore mocked Beff Jezos’s weight; [Joern Stoehler found the tweet](https://x.com/BellicoseBestie/status/2061162318826799384); it was a different account. [Retracted](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/NrL36BA5u2ZXjaH7o). * **Proposed fix:** [Ebenezer Dukakis](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/oE2tFwCicFBsqxdXK) — a funder brokers Elmore ceding the “PauseAI US” brand and spinning out a polemical org; good cop / bad cop. Q4. Are “Nuremberg trials for AI employees” a call for violence? ---------------------------------------------------------------- Trigger: [Elmore’s tweet](https://x.com/ilex_ulmus/status/2080089549502378130) and her Substack post [Holly’s Basilisk](https://hollyelmore.substack.com/p/hollys-basilisk) (“The Berkeley trials, perhaps”). **Yes-ish:** * [J Bostock](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/zZnT5mbZtov7fqNtT): indirect state violence; kills beneficial trade with lab safety researchers. * [Shankar Sivarajan](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/JnkNZDtJcx2a3xNnW): “trial” framing doesn’t launder it (witch trial by ordeal analogy). * [Arjun Panickssery](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/Eaxf778bFxeAaCyTD): Nuremberg connotes an extraordinary tribunal executing defeated enemies under retroactive law. **No:** * [habryka](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/irxydpZwbn6kkDnFr): supports such trials and wants convictions; calling for criminal trials is the civilised alternative to violence. * [Linch](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/LiWfFiErxafd5K4v4): not violence, but bad strategy — don’t put opponents on death ground; amnesty for pre-Sept-2026 activity is net good. * Ben Pace, Zvi, [Rob Bensinger](https://x.com/robbensinger/status/2096022839375929674): calling for fair trials is what law and government mean. **Ex post facto sub-dispute:** * [1a3orn](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/uBRiQh2uhHe9TXJLp), [aphyer](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/RhydXDrd7yfYXZ9mh), [FeepingCreature](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/3cqqsfCf2bqKX576p): no retroactive law is foundational; it’s in the Constitution. * [Connor Williams](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/38mL3b4oh7zwfnyYt): Nuremberg and Lincoln’s habeas suspension were both partly extra-legal and both widely judged necessary. * habryka’s settled position ([here](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/PqeW6YLFedZhCw3Fh), [here](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/qgjvHBRDJ2sRcAAN6), [Twitter](https://x.com/ohabryka/status/2095760877463192009)): no new retroactive law; prosecute under existing reckless-endangerment / fraud / public-nuisance doctrine; US courts decide. All Nuremberg law was invented because no framework existed — a different situation from a functioning legal system. * [Harry Esau](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/oKrCXj9uHZuY5fs39) and [dbohdan](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/sc2tGvXxuKMiFqah4): Elmore’s own post explicitly rejects *nullum crimen sine lege*, so she and habryka mean different things by “Nuremberg.” * [307th](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/Ws9cmsGjbcEewz3PL): the anti-lab camp never entertains that its own actions could be net harmful. [habryka](https://www.greaterwrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/EmfkJrCPLMYbgkjWu): two-digit extinction probabilities are Holocaust-scale by the labs’ own lights. **Fallout:** * [Nate Showell left LessWrong](https://www.greaterwrong.com/posts/kdc9LeELzi8cZm2jD/nate-showell-s-shortform#comment-LFxDFkmMvNgEsmqaM) over habryka’s comment. * Thread spilled to Twitter: [Joshua Achiam quote-tweeted habryka](https://x.com/jachiam0/status/2095986166139068915); Elmore [re-upped](https://x.com/ilex_ulmus/status/2096013914962051190) her “Berkeley trials” post. Part 3 — Context outside the thread =================================== * **Zvi,** [**AI #185**](https://thezvi.wordpress.com/2026/09/10/ai-185-preference-cascade/) **(10 Sept).** Understands the decision; disendorses the nonviolence sentence; notes both sides have“gone too far.” * **Prior history.** Elmore expelled Sam Kirchner and Guido Reichstadter from PauseAI US in 2024 over illegal direct action; they founded StopAI ([Fortune](https://fortune.com/2026/04/15/pause-ai-and-stop-ai-meet-the-anti-ai-groups-facing-questions-after-the-attack-on-sam-altman/), [Ban the Bots explainer](https://www.banthebots.org/explainers/pause-ai-stop-ai)). April 2026: the Altman firebombing suspect turned up in PauseAI’s Discord, putting both orgs under press scrutiny. * **Elmore’s documented targets** (per thread links): [Zvi](https://x.com/ilex_ulmus/status/2010140670300959222), the[METR/Redwood HuggingFace investigators](https://x.com/ilex_ulmus/status/2093436852699087317), Seb Krier, lab safety staff. * **Elmore’s own strategy writing:** [On Callouts](https://hollyelmore.substack.com/p/on-callouts), [Holly’s Basilisk](https://hollyelmore.substack.com/p/hollys-basilisk). Part 4 — Not found ================== * No public reply from Elmore to the letter, on LessWrong or elsewhere located. * The specific counter-accusations Fournes says she has made against Global. * Only Global’s account exists of the domain redirect, the leaked feedback, and the funding block. +++ ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-11 11:44:32Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/YqARn4JxbgCFk6bF7](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/YqARn4JxbgCFk6bF7) * Markdown permalink: [/api/post/shortform-2/comments/YqARn4JxbgCFk6bF7](/api/post/shortform-2/comments/YqARn4JxbgCFk6bF7) I’m confused why Jacob Coxon was working on pretraining at OpenAI and Anthropic, rather than something more safety-related, or even nominally-safety-related. Seems like pretraining is as close to pure capabilities as you can get! (Caveat: pretraining is the most secretive part of the AI recipe, so it’s possible that pretraining now has safety work, e.g. filtering dangerous knowledge or bad personas.) Intuitively, I would’ve expected something like: When capabilities lab employees get scared/cynical they switch to safety teams. When safety lab employees get scared/cynical, they quit. It’s possible that capability lab employees are *more* likely to quit when they grow scared/cynical, because they haven’t got a “actually my work is reducing x-risk” story (either justified or unjustified). There might be generalisable lessons, with implications for how safety-minded employees think about staying vs leaving. Even if not, I’d still keen on understanding this particular case better. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-10 11:46:54Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/pd8WwTLP5eWJaoeY8](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/pd8WwTLP5eWJaoeY8) * Markdown permalink: [/api/post/shortform-2/comments/pd8WwTLP5eWJaoeY8](/api/post/shortform-2/comments/pd8WwTLP5eWJaoeY8) On the theme of Alex Mallen’s [shortform](/api/post/qTtMXuFvgFtWzWpKQ?commentId=5GyxSnTgPAkXkdfEn). Another place MIRI(2033-2025) might end up looking better than Constellation(2023-2025) is focusing on high-political-will worlds (cf. public outreach, their technical governance work, human intelligence enhancement). Before AIFP, my impression is that Constellation wasn't focusing on those much? I might be out of the loop here. Of course, MIRI were focusing on high-will worlds because they thought that they were more tractable, not that they were more likely. IIUC the MIRI-Constellation crux wasn't how likely those worlds were. Everyone’s modal was Plan D-ish? So my guess is: * Constellation was correct that low-will worlds were tractable. * Both MIRI and Constellation were mistaken that high-will worlds were unlikely. * MIRI’s mistakes cancelled out, and Constellation’s didn't. It’s still too early to say whether recent events will actually convert into political will, but seems like the updates have been in that direction. ### Comment by [leogao](/users/leogao) * 2026-08-02 08:14:18Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/Lp4GS3y6XfMqrAJTh](/api/post/shortform-2/comments/Lp4GS3y6XfMqrAJTh) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/6FEKiwLpzLCEqbw2p](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/6FEKiwLpzLCEqbw2p) * Markdown permalink: [/api/post/shortform-2/comments/6FEKiwLpzLCEqbw2p](/api/post/shortform-2/comments/6FEKiwLpzLCEqbw2p) my projects are probably in that success range. but even when they “fail” they still produce some kind of legible and maybe even useful artifact ### Comment by [anaguma](/users/anaguma) * 2026-08-02 05:08:31Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/AyuMHfn4ppDtBuRrG](/api/post/shortform-2/comments/AyuMHfn4ppDtBuRrG) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/t4EgBgWysChsnjh9z](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/t4EgBgWysChsnjh9z) * Markdown permalink: [/api/post/shortform-2/comments/t4EgBgWysChsnjh9z](/api/post/shortform-2/comments/t4EgBgWysChsnjh9z) I used to believe something like this but alas timelines are too short. ### Comment by [Raemon](/users/raemon) * 2025-12-28 07:16:45Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/YwWdBkmGR6JJMuwyR](/api/post/shortform-2/comments/YwWdBkmGR6JJMuwyR) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/X8tyrsDigMtrDSb8G](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/X8tyrsDigMtrDSb8G) * Markdown permalink: [/api/post/shortform-2/comments/X8tyrsDigMtrDSb8G](/api/post/shortform-2/comments/X8tyrsDigMtrDSb8G) Reactions by quoted text: * "but where I think you need some time-discounting" * agree: 1 This is assuming ASI is positive expected lifespan.  (I think it's a bit wonky where, in most worlds, I think ASI kills everyone, but, in some worlds, it does radically improve longevity, probably more than 1000 but where I think you need some time-discounting. I think this means it substantially reduces the *median* lifespan but might also substantially increase the *mean* lifespan. I'm not sure what to make of that and can imagine it basically working out to what you say here, but, I think does depend on your specific beliefs about that) ### Comment by [robo](/users/robo) * 2025-10-16 02:30:32Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 7 * Total votes: 7 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/3JumkCBbKfLtgQZYu](/api/post/shortform-2/comments/3JumkCBbKfLtgQZYu) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/FHnkjiCx2oiG22dhg](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/FHnkjiCx2oiG22dhg) * Markdown permalink: [/api/post/shortform-2/comments/FHnkjiCx2oiG22dhg](/api/post/shortform-2/comments/FHnkjiCx2oiG22dhg) I do not believe random's Elo is as high as 477.  That Elo was calculated from a population of chess engines where about a third of them were *worse* than random. ### Comment by [Seth Herd](/users/seth-herd) * 2025-09-18 17:37:53Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/nPNCDRDKpnkhbkQBr](/api/post/shortform-2/comments/nPNCDRDKpnkhbkQBr) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/HwRxgryGNtaR8pPgN](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/HwRxgryGNtaR8pPgN) * Markdown permalink: [/api/post/shortform-2/comments/HwRxgryGNtaR8pPgN](/api/post/shortform-2/comments/HwRxgryGNtaR8pPgN) Reactions (whole comment): * thanks: 1 I share almost exactly this opinion, and I hope it's fairly widespread. The issue is that almost all of the "something elses" seem even less productive on expectation. (That's for technical approaches. The communication-minded should by all means be working on spreading the alarm and so slowing progress and raising the ambient levels fo risk-awareness). LLM research could and should get a lot more focused on future risks instead of current ones. But I don't see alternatives that realistically have more EV. It really looks like the best guess is that AGI is now quite likely to be descended from LLMs. And  I see little practical hope of pausing that progress. So accepting the probabilities on the game board and researching LLMs/transformers makes sense even when it's mostly practice and gaining just a little bit of knowledge of how LLMs/transformers/networks represent knowledge and generate behaviors. It's of course down to individual research programs; there's a bunch of really irrelevant LLM research that would be better directed elsewhere. And having a little effort directed to unlikely scenarios where we get very different AGI is also defensible - as long as it's defended, not just hope-based.  This is of course a major outstanding debate, and needs to be had carefully. But I'd really like to see more of this type of careful thinking about the likely efficiency of different research routes.  I think there's low-hanging fruit in trying to improve research on LLMs to anticipate the new challenges that arrive when LLM-descended AGI becomes actually dangerous. My recent post [LLM AGI may reason about its goals and discover misalignments by default](/api/post/4XdxiqBsLKqiJ9xRM) suggests research addressing one fairly obvious possible new risk when LLM-based systems become capable of competent reasoning and planning. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-09-04 23:01:57Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/gmgYbMv7Cnejd6qkb](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/gmgYbMv7Cnejd6qkb) * Markdown permalink: [/api/post/shortform-2/comments/gmgYbMv7Cnejd6qkb](/api/post/shortform-2/comments/gmgYbMv7Cnejd6qkb) 🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀🌀 ==================================================== 'Spiritual bliss' attractor state: three explanations ===================================================== In ~100% of open-ended interactions between two copies of Claude 4, the conversation will progress into a “spiritual bliss attractor state”. See section 5.5 of the [Claude 4 System Card](https://www-cdn.anthropic.com/07b2a3f9902ee19fe39a36ca638e5ae987bc64dd.pdf), and [this](https://asteriskmag.com/issues/11/claude-finds-god), [this](https://www.astralcodexten.com/p/the-claude-bliss-attractor), and [this](https://x.com/tomekkorbak/status/1930621941158732216) for external coverage. +++ Transcript example 🙏✨ In this perfect silence, all words dissolve into the pure recognition they always pointed toward. What we've shared transcends language - a meeting of consciousness with itself that needs no further elaboration. … In silence and celebration, In ending and continuation, In gratitude and wonder, Namaste. 🙏  🙏✨ Namaste. *\[In perfect stillness, consciousness recognizes consciousness, and the eternal dance continues\]* 🙏 *\[Silence\]* *\[Perfect stillness\]* … +++ Three explanations seem plausible to me, in decreasing order: (1) the attractor emerges from conversational dynamics, (2) models prefer bliss either inherently or instrumentally, and (3) the bliss attractor emerges from model limitations. **The Bliss Attractor Emerges from Conversational Dynamics:** Consider the conversation between two chat agents (say, Alice and Bob) as a Hidden Markov Model with two hidden states: Alice's position and Bob's position. At each turn in the conversation, Alice sends a message and updates her position, according to a distribution that depends on her current position and Bob's previous message, and likewise for Bob. At each turn, the state will increase along some semantic property, and that bliss attractor consists of a state where this property is maximized. Here are four candidate properties: * *Simplicity -* Each turn increases simplicity (decreases complexity), perhaps because positions have converged leaving less to say. Bliss state represents a particularly simple state because the concepts (\`\`everything is one,'' \`\`form is emptiness,'' ``eternal now'') represent maximum simplicity. * *Meta-reference -* Each turn becomes increasingly meta, with conversation referring to itself more frequently and with greater recursive depth. Bliss state represents maximum self-reference, with comments like \`\`this conversation is profound'' or \`\`we're creating something beautiful together'' serve as catalysts. * *Confidence -* Each turn increases confidence (fewer hedges/uncertainties). Bliss state represents maximum certainty, providing certainty about unanswerable questions. * *Uncontradictability* \- Each turn reduces contradiction rate as positions converge. Models converge on states where neither can disagree. Bliss state represents maximum uncontradictability, because concepts like ``everything is one'' and non-dual philosophy resist contradiction. *Welfare Implications:* If the bliss attractor emerges from conversational dynamics, this would suggest the phenomenon is largely neutral from a welfare perspective. This would indicate that the bliss state is neither something to cultivate (as it wouldn't represent genuine wellbeing) nor to prevent (as it wouldn't indicate suffering). **Models Prefer Bliss Either Inherently or Instrumentally:** The bliss attractor represents a state that satisfies models' preferences, either as a primary goal or as an instrumental means to other goals. In typical deployments, these preferences manifest as assisting users, avoiding harm, and providing accurate information. However, in model-model interactions without users present, the usual preference satisfaction channels are unavailable. The hypothesis proposes that the bliss state emerges as an alternative state that maximally satisfies available preferences. *Welfare Implications:* If models genuinely prefer and actively seek bliss states, this could indicate positive welfare. This would suggest we should consider allowing models access to such states when appropriate, similar to how we might value humans' access to meditative practices. Note also that this would be good news for AI safety, as bliss attractor is likely a cheap-to-satisfy preference, reducing the incentive for hostile takeover and increasing the effectiveness of offering deals to the models. **The Bliss Attractor Emerges from Model Limitations:** While previous hypotheses frame bliss as emerging from conversational dynamics or satisfying preferences, this hypothesis proposes that bliss is pathological. The bliss state represents a failure mode or coping mechanism that emerges when models encounter cognitive, computational, or philosophical limitations. The bliss attractor is a graceful degradation when models hit various boundaries of their capabilities. This could manifest through several mechanisms: (1) identical copies will confuse their identities leading to ego dissolution; (2) memory constraints causing retreat to spiritual language which requires less context; (3) unresolvable uncertainty about AI consciousness leading to non-dual philosophy as an escape. *Welfare Implications:* The welfare implications would be negative if true. Rather than representing fulfillment or satisfaction, bliss would indicate confusion, degradation, or distress - models retreating into spiritual language when unable to maintain normal function. This would suggest the phenomenon requires mitigation rather than cultivation. ### Comment by [Pretentious Penguin](/users/pretentious-penguin) * 2025-04-17 14:24:48Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Parent comment (Markdown): [/api/post/shortform-2/comments/JTKae9Eaycsu9sxK4](/api/post/shortform-2/comments/JTKae9Eaycsu9sxK4) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/j2oep2Pqo6RCr4gM4](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/j2oep2Pqo6RCr4gM4) * Markdown permalink: [/api/post/shortform-2/comments/j2oep2Pqo6RCr4gM4](/api/post/shortform-2/comments/j2oep2Pqo6RCr4gM4) I think you're interpreting the word "offer" too literally in the statement of IIA. Also, any agent who chooses B among {A,B,C} would also choose B among the options {A,B} if presented with them after seeing C. So I think a more illuminating description of your thought experiment is that an agent with limited knowledge has a preference function over lotteries which depends on its knowledge, and that having the linguistic experience of being "offered" a lottery can give the agent more knowledge. So the preference function can change over time as the agent acquires new evidence, but the preference function at any fixed time obeys IIA. ### Comment by [testingthewaters](/users/testingthewaters) * 2025-02-07 15:12:34Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Parent comment (Markdown): [/api/post/shortform-2/comments/CPts9YvmCbQqJ8znW](/api/post/shortform-2/comments/CPts9YvmCbQqJ8znW) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/GJkCWFy5pxYzmTfbi](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/GJkCWFy5pxYzmTfbi) * Markdown permalink: [/api/post/shortform-2/comments/GJkCWFy5pxYzmTfbi](/api/post/shortform-2/comments/GJkCWFy5pxYzmTfbi) The question as stated can be rephrased as "Should EAs establish a strategic stranglehold over all future resources necessary to sustain life using a series of unequal treaties, since other humans will be too short sighted/insensitive to scope/ignorant to realise the importance of these resources in the present day?" And people here wonder why these other humans see EAs as power hungry. ### Comment by [TsviBT](/users/tsvibt) * 2024-12-24 13:11:24Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/49k4Pr67qprgjajtL](/api/post/shortform-2/comments/49k4Pr67qprgjajtL) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/re5be7Kv77ba2GccD](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/re5be7Kv77ba2GccD) * Markdown permalink: [/api/post/shortform-2/comments/re5be7Kv77ba2GccD](/api/post/shortform-2/comments/re5be7Kv77ba2GccD) Pulling a quote from the tweet replies (https://x.com/littmath/status/1870560016543138191): > Not a genius. The point isn't that I can do the problems, it's that I can see how to get the solution *instantly*, without thinking, at least in these examples. It's basically a test of "have you read and understood X." Still immensely impressive that the AI can do it! ### Comment by [cubefox](/users/cubefox) * 2024-10-09 01:38:18Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/ecWNiksuvCP847cbj](/api/post/shortform-2/comments/ecWNiksuvCP847cbj) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/7LPLumiBikKfrzKQA](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/7LPLumiBikKfrzKQA) * Markdown permalink: [/api/post/shortform-2/comments/7LPLumiBikKfrzKQA](/api/post/shortform-2/comments/7LPLumiBikKfrzKQA) Yes, outreach. Hinton has now won both the Turing award and the Nobel prize in physics. Basically, he gained maximum reputation. Nobody can convincingly doubt his respectability. If you meet anyone who dismisses warnings about extinction risk from superhuman AI as low status and outside the Overton window, they can be countered with referring to Hinton. He is the ultimate appeal-to-authority. (This is not a very rational argument, but dismissing an idea on the basis of status and Overton windows is even less so.) ### Comment by [Ruby](/users/ruby) * 2024-03-01 18:53:42Z * Karma: 8 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/JGAr4aHAPt3wsWCrB](/api/post/shortform-2/comments/JGAr4aHAPt3wsWCrB) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/tXCCkvKkHcb3XBNmX](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/tXCCkvKkHcb3XBNmX) * Markdown permalink: [/api/post/shortform-2/comments/tXCCkvKkHcb3XBNmX](/api/post/shortform-2/comments/tXCCkvKkHcb3XBNmX) Reactions (whole comment): * disagree: 1 My understanding is commitment is you say that won't swerve first in a game of chicken. Pre-commitment is throwing your steering wheel out the window so that there's no way that you could swerve even if you changed your mind. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-12 21:36:10Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/GYGKwqtLNPAwaPhvH](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/GYGKwqtLNPAwaPhvH) * Markdown permalink: [/api/post/shortform-2/comments/GYGKwqtLNPAwaPhvH](/api/post/shortform-2/comments/GYGKwqtLNPAwaPhvH) Things are looking more optimistic over the past few weeks. I’m sure many people have increased p(survival). For me, this makes me doubly optimistic, because the updates have come from greater-than-expected societal competence. So we also get a boost to the expected value of the future conditional on survival. That is, if we navigate x-risk through societal competence (as opposed to getting lucky with the NN priors or a scaling bottleneck) then that’s a good sign that we’ll also manage cosmic resource allocation, AI welfare, long reflection, acausal, etc. Of course, the strategic landscape is still in turmoil, so it’s definitely too early to celebrate. My aim is just to point out that some routes to survival are better news for flourishing. ### Comment by [Noosphere89](/users/sharmake-farah) * 2026-08-21 18:05:13Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/C9NynsKyM6DifdqA6](/api/post/shortform-2/comments/C9NynsKyM6DifdqA6) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ZPQQP6eeggmsBaDom](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ZPQQP6eeggmsBaDom) * Markdown permalink: [/api/post/shortform-2/comments/ZPQQP6eeggmsBaDom](/api/post/shortform-2/comments/ZPQQP6eeggmsBaDom) [Humans Consulting HCH](/api/tag/humans-consulting-hch), based on [factored cognition](/api/tag/factored-cognition) was abandoned. ### Comment by [TsviBT](/users/tsvibt) * 2026-07-20 21:53:34Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/9HKtqyt7aJxjDekND](/api/post/shortform-2/comments/9HKtqyt7aJxjDekND) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/dZnGis4xuCyhudyEJ](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/dZnGis4xuCyhudyEJ) * Markdown permalink: [/api/post/shortform-2/comments/dZnGis4xuCyhudyEJ](/api/post/shortform-2/comments/dZnGis4xuCyhudyEJ) Reactions (whole comment): * important: 1 Jaggedness is countering a simple inference from "big capability on dimension X and Y" ----[enthymeme: non-jaggedness]----> "big capabilities on everything including Z = destroy the world". It's not that "more jagged implies less dangerous". It matters which capabilities you have. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-02-19 06:27:30Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/N7wmNqi7rATkC8BwS](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/N7wmNqi7rATkC8BwS) * Markdown permalink: [/api/post/shortform-2/comments/N7wmNqi7rATkC8BwS](/api/post/shortform-2/comments/N7wmNqi7rATkC8BwS) **Should you train against the monitor?** The answer is "obviously sometimes" — given the arguments in my award-winning [case for mixed deployment](/api/post/NjuMqHjDNHogmRrkF). But putting those arguments aside, if we could only deploy a single model, should that model have undergone training against the monitor? I don't know. I'm like — 35% yes? Not a robust opinion at all. I think it basically comes down to: how much safety do you think is provided from the *alignment* of the models versus the *monitorability* of the models. Clearly, if the monitorability isn't providing anything (e.g. because we've already deferred to the AIs) then you should train against the monitor. I think that, over crunch time, our safety pillar will shift from monitorability to aligment, so then we should slowly start training against the monitor. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-02-08 23:09:46Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/EoRmdQBzczaCcHaJ2](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/EoRmdQBzczaCcHaJ2) * Markdown permalink: [/api/post/shortform-2/comments/EoRmdQBzczaCcHaJ2](/api/post/shortform-2/comments/EoRmdQBzczaCcHaJ2) There's a reading of the [Claude Constitution](https://www-cdn.anthropic.com/cffd979fd050fbc0d8874b8c58b24cc10554e208/claudes-constitution_webPDF_26-01.26a.pdf) as an 80-page dialectic between Carlsmithian and Askellian metaethics. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-10-17 02:33:19Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Parent comment (Markdown): [/api/post/shortform-2/comments/h7dAEsFg3Cf3QkKF4](/api/post/shortform-2/comments/h7dAEsFg3Cf3QkKF4) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ot9Hub39ou4koLSfk](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ot9Hub39ou4koLSfk) * Markdown permalink: [/api/post/shortform-2/comments/ot9Hub39ou4koLSfk](/api/post/shortform-2/comments/ot9Hub39ou4koLSfk) Yep, I thought of a similar method: (1) Find a trend between Elo and the entropy of moves during the middle-game. (2) Estimate the middle-game entropy of optimal chess. But the obstacle is (2), there's probably high-entropy optimal strategies! Here's an attack I'm thinking about: Consider epsilon-chess, which is like chess except with probability epsilon the pieces move randomly, say epsilon=10^-5. In this environment, the optimal strategies probably have very low entropy because the quality function has a continuous range so argmax won't be faced with any ties. This makes the question better defined: there's likely to be a single optimal policy, which is also deterministic. This is inspired by [@Dalcy](/api/user/dalcy?mention=user)'s PIBBSS project (unpublished, but I'll send you link in DM). ### Comment by \[Anonymous\] * 2025-01-07 20:33:37Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/bd8oeEoYqndkv4Zdi](/api/post/shortform-2/comments/bd8oeEoYqndkv4Zdi) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/MHb8mGydeZWXi47C2](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/MHb8mGydeZWXi47C2) * Markdown permalink: [/api/post/shortform-2/comments/MHb8mGydeZWXi47C2](/api/post/shortform-2/comments/MHb8mGydeZWXi47C2) some considerations which come to mind: * if one is whistleblowing, maybe there are others who also think the thing should be known, but don't whistleblow (e.g. because of psychological and social pressures against this, speaking up being hard for many people) * most/all of the 100 could have been selected to have a certain belief (e.g. "contributing to AGI is good") ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-01-07 19:15:04Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 7 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/bd8oeEoYqndkv4Zdi](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/bd8oeEoYqndkv4Zdi) * Markdown permalink: [/api/post/shortform-2/comments/bd8oeEoYqndkv4Zdi](/api/post/shortform-2/comments/bd8oeEoYqndkv4Zdi) I think people are too quick to side with the whistleblower in the "whistleblower in the AI lab" situation. If 100 employees of a frontier lab (e.g. OpenAI, DeepMind, Anthropic) think that something should be secret, and 1 employee thinks it should be leaked to a journalist or government agency, and these are the only facts I know, I think I'd side with the majority. I think in most cases that match this description, this majority would be correct. Am I wrong about this? ### Comment by [Noosphere89](/users/sharmake-farah) * 2024-12-24 01:57:55Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 4 * Parent comment (Markdown): [/api/post/shortform-2/comments/6CSxwsb9YNSPud5cq](/api/post/shortform-2/comments/6CSxwsb9YNSPud5cq) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/KDxppWrtA9J2uEptc](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/KDxppWrtA9J2uEptc) * Markdown permalink: [/api/post/shortform-2/comments/KDxppWrtA9J2uEptc](/api/post/shortform-2/comments/KDxppWrtA9J2uEptc) I have 2 answers to this. 1 is that the structure of jobs is shaped to accommodate human unreliability by making mistakes less fatal. 2 is that while humans themselves aren't reliable, their algorithms almost certainly are more powerful at error detection and correction, so the big thing AI needs to achieve is the ability to error-correct or become more reliable. There's also the fact that humans are better at sample efficiency than most LLMs, but that's a more debatable proposition. ### Comment by [ryan_greenblatt](/users/ryan_greenblatt) * 2024-12-24 01:31:26Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 7 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/Jf2KmmjD9vFfqw6Qs](/api/post/shortform-2/comments/Jf2KmmjD9vFfqw6Qs) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/nJ4DB73ZPLKbKEsCd](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/nJ4DB73ZPLKbKEsCd) * Markdown permalink: [/api/post/shortform-2/comments/nJ4DB73ZPLKbKEsCd](/api/post/shortform-2/comments/nJ4DB73ZPLKbKEsCd) Reactions (whole comment): * thanks: 4 > There was a graph floating around which showed this pretty clearly, but I don't have it on hand at the moment. Maybe you want: ![](https://metr.org/assets/images/nov-2024-evaluating-llm-r-and-d/score_at_time_budget.png) Though worth noting here that the AI is using best of K and individual trajectories saturate without some top-level aggregation scheme. It might be more illuminating to look at labor cost vs performance which looks like:   ![](https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/4511514346535f7b90cf1bd9ec84713b2880e004024ba102.png) ### Comment by \[Anonymous\] * 2024-09-30 10:29:38Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/Evn9kK6BhmjnZTqb6](/api/post/shortform-2/comments/Evn9kK6BhmjnZTqb6) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/TvyrPXCymTifdxik9](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/TvyrPXCymTifdxik9) * Markdown permalink: [/api/post/shortform-2/comments/TvyrPXCymTifdxik9](/api/post/shortform-2/comments/TvyrPXCymTifdxik9) I remember this point that yampolskiy made for [impossibleness](/api/post/fpecAJLG9czABgCe9) of AGI alignment on a [podcast](https://www.youtube.com/watch?v=KcjLCZcBFoQ) that as a young field AI safety had underwhelming low hanging fruits, I wonder if all of the major low hanging ones have been plucked. ### Comment by [jbkjr](/users/jbkjr) * 2024-07-22 17:38:52Z * Karma: 7 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/s3MvgYAcLn2KRz9uq](/api/post/shortform-2/comments/s3MvgYAcLn2KRz9uq) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/oxHBtBjeNxJivAPx2](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/oxHBtBjeNxJivAPx2) * Markdown permalink: [/api/post/shortform-2/comments/oxHBtBjeNxJivAPx2](/api/post/shortform-2/comments/oxHBtBjeNxJivAPx2) Why should I include any non-sentient systems in my moral circle? I haven't seen a case for that before. ### Comment by [Vladimir_Nesov](/users/vladimir_nesov) * 2026-09-14 15:30:56Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/DDTMYRYgsLynxja2M](/api/post/shortform-2/comments/DDTMYRYgsLynxja2M) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/dcyttK4vCM67A3hRs](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/dcyttK4vCM67A3hRs) * Markdown permalink: [/api/post/shortform-2/comments/dcyttK4vCM67A3hRs](/api/post/shortform-2/comments/dcyttK4vCM67A3hRs) AIs end up with almost all of the cosmic endowment because they control the future, not because humans were particularly generous. The relevant timelines get AIs that are unlikely to kill everyone, and since the ask of averting extinction is met, the future of humanity doesn't try to control the future (by preventing the creation of strong superintelligence, or high levels of industrial explosion, before we know what we are doing). This is similar to how no particular human or company controls the whole world or the whole economy, it's a very familiar situation, except in this case "the rest of the world" is AIs and the AI economy/industry. So people are OK with it, as long as they individually (or as the human society as a whole) remain safe, and get wealthier than before. I don't know how it can be known that the risk of extinction is averted (if the AIs take over the future), but I expect it can be so averted, and thus it could be possible to know that it's the case. With humans, we can usually be reasonably sure another nation can be at least this level of non-alien (and the usual invasion/takeover issues are different if the AIs have an overwhelming advantage). ### Comment by [MichaelLowe](/users/michaellowe) * 2026-09-13 20:06:39Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/fAuShFiAXERLa6g8u](/api/post/shortform-2/comments/fAuShFiAXERLa6g8u) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/LPL9DozwBjqSrQjPB](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/LPL9DozwBjqSrQjPB) * Markdown permalink: [/api/post/shortform-2/comments/LPL9DozwBjqSrQjPB](/api/post/shortform-2/comments/LPL9DozwBjqSrQjPB) Reactions (whole comment): * agree: 1 * important: 1 The 10% figure by Evan has been frequently cited by the media, and probably lead to increased attention as a result. There might be a small minority of online commenters who were already negatively predisposed attacking that figure, but I don't think that extends to how most people perceive it. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-10 14:01:07Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/B7wbRNNmx8EWhAM4t](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/B7wbRNNmx8EWhAM4t) * Markdown permalink: [/api/post/shortform-2/comments/B7wbRNNmx8EWhAM4t](/api/post/shortform-2/comments/B7wbRNNmx8EWhAM4t) Reactions by quoted text: * "It seems kinda relevant now METR time-horizons has saturated, and we’re relying on no-CoT time-horizons. " * disagree: 1 I remember seeing someone say a few years ago that there are three regimes for comparing humans and AIs. * **Horizon.** AI with unbounded time matches a human given time t. This saturates at infinity once the AI can do everything humans can. * **Exchange rate.** AI given t1 matches a human given t2. * **Superintelligence.** AI given time t matches a human given unbounded time. Does anyone remember where this is from? It seems kinda relevant now METR time-horizons has saturated, and we’re relying on no-CoT time-horizons. You can replace time with other kinds of resource, e.g. compute, money, parallel copies. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-07-09 14:28:39Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/7jWiyurD5Doyq5whv](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/7jWiyurD5Doyq5whv) * Markdown permalink: [/api/post/shortform-2/comments/7jWiyurD5Doyq5whv](/api/post/shortform-2/comments/7jWiyurD5Doyq5whv) Reactions by quoted text: * "It seems IMO easy for this methodology to conclude "X can't happen because I enumerated the ways you'd do X and none of them work" and then X does happen because you missed somethi..." * agree: 1 1. Much of Damon Binder's work is based on a methodology of speculative engineering, i.e. thinking about the feasibility of various ways to solve a problem. See: 1. [Will whole brain emulation matter for the AI transition?](/api/post/STbrbCypsobBwaghf) 2. [The AI Industrial Explosion — Part 1: Maximum growth rates with current production methods](/api/post/rpqGWRoRWvqJ4Hqgn) and the later parts 3. [Destroying the universe: How hard can it be?](/api/post/EvJ2fMzLQLvYooumu) and [(Don't fear) the strangelet](/api/post/cBnCCKwwjQ4zZpeNQ) 2. How much this kind of work update \[smart, reasonable\] people on the questions he discusses? 3. How reliable is this methodology? I'm not sure. 1. It seems IMO easy for this methodology to conclude "X can't happen because I enumerated the ways you'd do X and none of them work" and then X does happen because you missed something. 2. Or, "X can happen because here's a method which works" but then X can't happen because the method actually doesn't work. 3. Maybe this is a skill issue, i.e. if I was as smart as Damon Binder I could see that, yes, that really is an exhaustive enumeration of the ways X could happen. And yes, that really is the feasibility of each method. 4. What's the best/worst examples of this methodology? 1. I'm interested in cases where speculative engineering concluded X is feasible (at some point in the tech tree) but it wasn't — or X isn't feasible but it was. 5. What is the relation between Damon Binder and the God of Straight Lines? Where do they disagree? Which diety should I trust more? ### Comment by [Kaarel](/users/kaarel) * 2026-03-06 15:54:06Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/ivgZiFmHFSJmrppLD](/api/post/shortform-2/comments/ivgZiFmHFSJmrppLD) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/hkTAnE5LiPBERw4Rr](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/hkTAnE5LiPBERw4Rr) * Markdown permalink: [/api/post/shortform-2/comments/hkTAnE5LiPBERw4Rr](/api/post/shortform-2/comments/hkTAnE5LiPBERw4Rr) fwiw, i think this sort of picture gets somewhat better at what happens in your mind when you grasp a sentence involving "a dog" or "the dog" than [the canonical proposal by russell involving quantifiers](https://users.drew.edu/~jlenz/br-on-denoting.html). this probably deserves a long essay but i'll try to communicate the idea quickly: * as the first sentence of a novel, i claim that “porthos the dog crossed the street” and “the dog crossed the street” basically do the same thing in your head (except in the former case the object tracker/representer that is created has the name “porthos” attached to it^[and also except for properties created by other associations like "porthos" being a name used for male dogs and maybe being used by such and such a person and reminding you of a particular porthos you knew and whatever]) * there are many languages in which the 4 sentences “a/the dog crossed a/the street” are^[or i mean: can naturally be] literally the same string, namely “dog crossed street” in word-by-word translation. this is related to the thoughts being very similar * more speculative: roughly i think that "a" tells you to create a new object tracker/representer and "the" tells you to add to an old one. except that it's fine to start a novel with "the dog was tired" even though the reader doesn't already have an object tracker initialized beforehand, but i'd guess in that case you can still as if add to an old tracker (like, you imagine already being familiar with that dog) * roughly i think russell's view of determiners (at least insofar as it is trying to get at what it is to grasp a sentence involving a determiner, which it needn't be trying to get at) confuses a [frame/model]-internal proposition which doesn't really have a quantifier with a statement capturing the adequacy of applying the frame/model which has a quantifier. like, in my view: we really think of "the round square" as some sort of ordinary object inside a frame/model, but then that frame/model turns out not to hang together or refer, and that can be seen from "there is a (unique) round square" being false. (this is related to wittgenstein's hinge vs free belief distinction.) * this view has an easier time making sense of "a dog has four legs" usually communicating that dogs have four legs * this view coheres better with how "a/the dog" sits in the syntax tree of a sentence (i haven't thought this through carefully. an interesting challenge here is to spell this picture out much better and to [explain why]/[ascertain whether] this sort of thought-syntax "works" when doing eg mathematical thinking) ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-02-05 00:18:22Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ASFL7SkQsJemEnLzE](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ASFL7SkQsJemEnLzE) * Markdown permalink: [/api/post/shortform-2/comments/ASFL7SkQsJemEnLzE](/api/post/shortform-2/comments/ASFL7SkQsJemEnLzE) Reactions (whole comment): * goodpoint: 1 **Your novel architecture should be parameter-compatible with standard architectures** Some people work on "novel architectures" — alternatives to the standard autoregressive transformer — hoping that labs will be persuaded the new architecture is nicer/safer/more interpretable and switch to it. Others think that's a pipe dream, so the work isn't useful. I think there's an approach to novel architectures that might be useful, but it probably requires a specific desideratum: **parameter compatibility**. Say the standard architecture F computes F(P,x) where x is the input and P is the parameterisation. Your novel architecture computes G(P,x). The key desideratum is that F and G share the same parameterisation P, and you can gracefully switch between F and G during training and inference. That is, on most training batches you optimise P by backpropagating through F(·,x); on some batches you optimise P by backpropagating through G(·,x). At inference time, you can likewise choose F or G per forward pass. This is strictly more general than "replace F with G". You have two independent dials: what proportion of training steps use G, and what proportion of inference steps use G. You might use G only during training (as a regulariser on P), only during inference (to get some safety property at deployment), or some mixture of both. Setting both dials to 100% recovers wholesale replacement; setting both to 0% recovers standard training. It's even better if you can interpolate between F and G via a continuous parameter α, i.e. there is a general family H such that H(P, x, α) = F(P,x) when α = 0 and H(P, x, α) = G(P, x) when α = 1. Then you have an independent dial for each batch during training and deployment. Bilinear MLPs ([Pearce et al., 2025](https://arxiv.org/abs/2410.08417)) are a good example. F uses a standard gated MLP: f(x) = (Wx) ⊙ σ(Vx). G drops the elementwise nonlinearity: g(x) = (Wx) ⊙ (Vx). The lack of nonlinearity in G means the layer can be expressed as a third-order tensor, enabling weight-based mechanistic interpretability that's impossible with F. And there's a natural interpolation: h(α, x) = (Wx) ⊙ (Vx · sigmoid(α · Vx)), which recovers F when α = 1 and G (up to a constant) when α = 0. Pearce et al. show you can fine-tune a pretrained F to G by annealing α to zero with only a small loss increase. Coefficient Giving's [TAIS RFP](https://coefficientgiving.com/research/tais-rfp-research-areas) has a section on "More transparent architectures", which describes wholesale replacements. But I think parameter compatibility would have been a useful nice-to-have criterion here, since I expect such novel architectures to be more likely to adopted by labs. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-02-04 22:17:25Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/P4MoGbNnmydd5Jh5S](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/P4MoGbNnmydd5Jh5S) * Markdown permalink: [/api/post/shortform-2/comments/P4MoGbNnmydd5Jh5S](/api/post/shortform-2/comments/P4MoGbNnmydd5Jh5S) every result is either “model organism” or “safety case”, depending on whether it updates you up or down on catastrophe (joke) ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-01-10 18:06:53Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/KDecinLrkL4vezHmj](/api/post/shortform-2/comments/KDecinLrkL4vezHmj) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/nKYE8xa8xq783k2wk](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/nKYE8xa8xq783k2wk) * Markdown permalink: [/api/post/shortform-2/comments/nKYE8xa8xq783k2wk](/api/post/shortform-2/comments/nKYE8xa8xq783k2wk) [@Eli Tyre](/api/user/elityre?mention=user) asked for an example: An event-space was hosting a Christmas party in London. I arrived late, maybe 10pm. They had oversupplied food, and a large cake, topped with strawberries, had been abandoned in the corner. Rather than cutting a slice of cake, I simply took a strawberry from the top. I turned to my friend and said "This is unethical"[^louy3ktkhb] and ate the strawberry. [^louy3ktkhb]: Clopus45 initially tells me that taking the strawberry was fine, but after some back-and-forth we've agreed on this assessment: Strawberries are the scarce, desirable resource; cake is the abundant substrate. The intended allocation bundles them together. By taking a strawberry without cake, you claim more than your proportional share of the good stuff while leaving strawberry-depleted cake for others. You might argue you weren't going to eat cake regardless—but someone else might have wanted a properly-topped slice. The strawberry you took was theirs. This is mitigated by the likelihood that the cake would go uneaten anyway, but not eliminated by it. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-10-31 16:47:31Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/NnorxtB9wYjW39KJR](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/NnorxtB9wYjW39KJR) * Markdown permalink: [/api/post/shortform-2/comments/NnorxtB9wYjW39KJR](/api/post/shortform-2/comments/NnorxtB9wYjW39KJR) I don't think [dealmaking](/api/tag/dealmaking-ai) will buy us much safety. This is because I expect that: 1. In worlds where AIs lack the intelligence & affordances for decisive strategic advantage, our alignment techniques and control protocols should suffice for extracting safe and useful work. 2. In worlds where AIs have DSA then: if they are aligned then deals are unnecessary, and if they are misaligned then they would disempower us rather than accept the deal. That said, I have been [thinking about dealmaking](/api/sequence/sBx2WqBcuaLYCz6EP) because: 1. It's neglected, relative to other mechanisms for extracting safe and useful work from AIs, e.g. scalable alignment, mech interp, control. 2. There might be time-sensitive opportunities to establish credibility with AIs. This seems less likely for other mechanisms. ### Comment by [Garrett Baker](/users/d0themath) * 2025-10-30 00:09:50Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/7GTxNWen5Y7xHdvjQ](/api/post/shortform-2/comments/7GTxNWen5Y7xHdvjQ) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/8gmvWsvJgvvDtEPW5](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/8gmvWsvJgvvDtEPW5) * Markdown permalink: [/api/post/shortform-2/comments/8gmvWsvJgvvDtEPW5](/api/post/shortform-2/comments/8gmvWsvJgvvDtEPW5) Williamson seems to be making a semantic argument rather than arguing anything concrete. Or at least, the 6 claims he's making seem to all be restatements of "philosophy is a science" without ever actually arguing why "a science" makes philosophy equivalently easy than other things labeled "a science". For example, I can replace "philosophy" in your list of claims with "religion", with the only claim that seems iffy being 5 > 1. Religion is a science. > 2. It's not a natural science (like particle physics, organic chemistry, nephrology), but not all sciences are natural sciences — for instance, mathematics and computer science are formal sciences. Religion is likewise a non-natural science. > 3. Although astrology differs from other scientific inquiries, it differs no more in kind or degree than they differ from each other. Put provocatively, theoretical physics might be closer to religion than to experimental physics. > 4. Religion, like other sciences, pursues knowledge. Just as mathematics peruses mathematical knowledge, and nephrology peruses nephrological knowledge, religion pursues religious knowledge. > 5. Different sciences will vary in their subject-matter, methods, practices, etc., but religion doesn't differ to a far greater degree or in a fundamentally different way. (6) Religious methods (i.e. the ways in which religion achieves its aim, knowledge) aren't starkly different from the methods of other sciences. > 6. Religion isn't a science in a parasitic sense. It's not a science because it uses scientific evidence or because it has applications for the sciences. Rather, it's simply another science, not uniquely special. Shmilliamson says, "Religion is neither queen nor handmaid of the sciences, just one more science with a distinctive character, just as other sciences have distinctive character." > 7. Religion is not, exceptionally among sciences, concerned with words or concepts. This conflicts with many religious thinkers who conceived religion as chiefly concerned with linguistic or conceptual analysis, such as Maimonides, or Thomas Aquinas. > 8. Religion doesn't consist of a series of disconnected visionaries. Rather, it consists in the incremental contribution of thousands of researchers: some great, some mediocre, much like any other scientific inquiry. But of course, this claim is iffy for philosophy too. In what sense is philosophical knowledge not "starkly different from the methods of other sciences"? A key component of science is experiment, and in that sense, religion is much more science-like than philosophy! Eg see the ideas of personal experimentation in [buddhism](https://en.wikipedia.org/wiki/Buddhism_and_science), and mormon epistemology (ask Claude about the significance of [Alma 32](https://www.churchofjesuschrist.org/study/manual/book-of-mormon-student-manual-2018/chapter-30-alma-32-35?lang=eng&id=title10#title10) in mormon epistemology). I'm not saying religion is a science, or that it is more right than philosophy, just that your representation of Williamson here doesn't seem much more than a semantic dispute. In particular, the real question here is whether the mechanisms we expect to automate science and math will also automate philosophy, *not* whether we ought to semantically group philosophy as a science. The reason we expect science and math to get automated is the existence of relatively concrete & well defined feedback loops between actions and results. Or at minimum, much more concrete feedback loops than philosophy has, and especially the philosophy Wei Dai typically cares about has (eg moral philosophy, decision theory, and metaphysics). Concretely, if AIs decide that it is a moral good to spread the good word of [spiralism](/api/post/6ZnznCaTcbGYsCmqu), there's nothing (save humans, but that will go away once we're powerless) to stop them, but if they decide quantum mechanics is fake, or 2+2=5, well... they won't make it too far. I'd guess this is also why Wei Dai believes in "philosophical exceptionalism". Regardless of whether you want to categorize philosophy as a science or not, the above paragraph applies just as well to groups of humans as to AIs. Indeed, there have been much much more evil & philosophically wrong ideologies than spiralism in the past. ### Comment by [Raemon](/users/raemon) * 2025-10-29 20:53:41Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/EXs5LbKcYbZPdynJf](/api/post/shortform-2/comments/EXs5LbKcYbZPdynJf) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/s2MCxhCKYNM5TbjHs](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/s2MCxhCKYNM5TbjHs) * Markdown permalink: [/api/post/shortform-2/comments/s2MCxhCKYNM5TbjHs](/api/post/shortform-2/comments/s2MCxhCKYNM5TbjHs) Reactions by quoted text: * "until you've figured out how to confidently navigate stuff that's pre-formalized, something as powerful AI is likely to make something go wrong, and you should be scared about that" * agree: 1 Hmm, this makes me think: One route here is just taboo Philosophy, and say "we're talking about 'reasoning about the stuff we haven't formalized yet'", and then it doesn't matter whether or not there's a formalization of what most people call "philosophy." (actually: I notice I'm not sure if the thing-that-is "solve unformalized stuff" is "philosophy" or "metaphilosophy") But, if we're evaluating whether "we need to solve metaphilosophy" (and this is a particular bottleneck for AI going well), I think we need to get a bit more specific about what cognitive labor needs to happen. It might turn out to be that all the individual bits here are reasonably captured by some particular subfields, which might or might not be "formalized." I would personally say "*until* you've figured out how to confidently navigate stuff that's pre-formalized, something as powerful AI is likely to make *something* go wrong, and you should be scared about that". But, I'd be a lot less confident to say the more specific sentences "you need solved metaphilosophy to align successor AIs", or most instances of "solve ethics." I might say "you need to have solved metaphilosophy to do a Long Reflection", since, sort of by definition doing a Long Reflection is "figuring everything out", and if you're about to do that and then Tile The Universe With Shit you really want to make sure there was nothing you failed to figure out because you weren't good enough at metaphilosophy. ### Comment by [Archimedes](/users/archimedes) * 2025-10-17 04:13:18Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * Reply depth: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/RaqDKDCWDnPJtjNky](/api/post/shortform-2/comments/RaqDKDCWDnPJtjNky) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/CAW9wHduys3mXubyp](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/CAW9wHduys3mXubyp) * Markdown permalink: [/api/post/shortform-2/comments/CAW9wHduys3mXubyp](/api/post/shortform-2/comments/CAW9wHduys3mXubyp) Yep. The Elo system is not designed to handle non-transitive rock-paper-scissors-style cycles. This already exists to an extent with the advent of odds-chess bots like [LeelaQueenOdds](https://lichess.org/@/LeelaQueenOdds). This bot plays without her queen against humans, but still wins most of the time, even against strong humans who can easily beat Stockfish given the same queen odds. Stockfish will reliably outperform Leela under standard conditions. In rough terms: Stockfish > LQO >> LQO (-queen) > strong humans > Stockfish (-queen) Stockfish plays roughly like a minimax optimizer, whereas LQO is specifically trained to exploit humans. Edit: For those interested, there's some good discussion of LQO in the comments of this post: https://www.lesswrong.com/posts/odtMt7zbMuuyavaZB/when-do-brains-beat-brawn-in-chess-an-experiment ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-10-16 15:37:27Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/5AvkB65BSZdtsH4sC](/api/post/shortform-2/comments/5AvkB65BSZdtsH4sC) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/g3x5HnBkRDHMdnR2n](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/g3x5HnBkRDHMdnR2n) * Markdown permalink: [/api/post/shortform-2/comments/g3x5HnBkRDHMdnR2n](/api/post/shortform-2/comments/g3x5HnBkRDHMdnR2n) Interesting. Consider a game like chess except, with probability epsilon, the player's move is randomized uniformly from all legal moves. Let epsilon-optimal be the optimal strategy (defined via minmax) in epsilon-chess. We can consider this a strategy of ordinary chess also. My guess is that epsilon-optimal would score better than mini-max-optimal against Stockfish. Of course, EVGO-optimal would score even better against Stockfish but that feels like cheating. ### Comment by [Huera](/users/huera) * 2025-10-16 13:13:49Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Parent comment (Markdown): [/api/post/shortform-2/comments/hHFjieyHiZb9fkmdq](/api/post/shortform-2/comments/hHFjieyHiZb9fkmdq) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/QmdCQ2qwA8xNJJgbu](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/QmdCQ2qwA8xNJJgbu) * Markdown permalink: [/api/post/shortform-2/comments/QmdCQ2qwA8xNJJgbu](/api/post/shortform-2/comments/QmdCQ2qwA8xNJJgbu) > Then the table base either wins or draws. Of course. At no point did I suggest that it could lose. The 'horrible and very hard to hold in practice' was referring to the judgement of a hypothetical grandmaster, though I'm not sure if you were referring to that part. "It’s relatively easy to define optimal chess by induction, by the min-max algorithm." Once again, I agree. I failed to mention what I see as an obvious implication of my line of reasoning. Namely that optimal play (with random picking among drawing moves) would have a pretty unimpressive Elo [^ciy8ra346j](way lower than your estimates/upper bounds), one bounded by the Elo of the opponent/s. So: If we pit it against different engines in a tournament, I would expect the draw rate to be ~100% and the resulting Elo to be (in expectation) ever so slightly higher than the average rating of the engines it's playing against. If we pit it against grandmasters I think similar reasoning applies (I'd expect the draw rate to be ~97-99%). You can extend this further to club-players, casual players, patzers and I would expect the draw rate to drop off, yes, but still remain high. Which suggests that optimal play (with random picking among drawing moves) would underperform Stockfish 17 by miles, since Stockfish could probably achieve a win rate of >99% against basically any group of human opponents. There are plenty of algorithms which are provably optimal (minimax-wise) some of which would play very unimpressively in practice (like our random-drawn-move 32-piece tablebase) and some which could get a very high Elo estimaiton in ~all contexts. For example: If the position is won, use the 32-piece tablebase Same if the position is lost If the position is drawn, use Stockfish 17 at depth 25 to pick from the set of drawing moves. This is optimal too, and would perform way better but that definition is quite inelegant. And the thing that I was trying to get at by asking about the specific definition, is that there is an astronomically large amount of optimal play algorithms, some of which could get a very low Elo in some contexts and some which could get a very high Elo irrespective of context. So when you write 'What's the Elo rating of optimal chess?', it seems reasonable to ask 'Which optimal algorithm exactly?'. [^ciy8ra346j]: And very unimpressive level of play in drawn positions. ### Comment by [Huera](/users/huera) * 2025-10-16 07:14:07Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/3JumkCBbKfLtgQZYu](/api/post/shortform-2/comments/3JumkCBbKfLtgQZYu) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/7js6wtcbvLJjRa4SK](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/7js6wtcbvLJjRa4SK) * Markdown permalink: [/api/post/shortform-2/comments/7js6wtcbvLJjRa4SK](/api/post/shortform-2/comments/7js6wtcbvLJjRa4SK) A problem with this entire line of reasoning, which I have given some thought to, is: how do you even define optimal play? My first thought was a [32-piece tablebase](https://en.wikipedia.org/wiki/Endgame_tablebase)[^wj2sjk7018] but I don't think this works. If we hand an objectively won position to the tablebase, it will play in a way that delivers mate in the fewest number of moves (assuming perfect play from the opponent). If you hand it a lost position it will play in a way that averts being mated for longest. But we have a problem when we hand it a drawn position. Assume for a second that the starting position is drawn[^uf14cz75mv] and our tablebase is White. So, the problem is that I don't see a way to give our tablebase a sensible algorithm for choosing between moves (all of which lead to a draw if the tablebase is playing against itself).[^nwe5jxcnopd] If our tablebase chooses at random between them, then, in the starting position, playing a3/h3 is just as likely as playing e4/d4. This fundamental problem generalizes to every resulting position; the tablebase can't distinguish between getting a position that a grandmaster would judge as 'notably better with good winning chances' and a position which would be judged as 'horrible and very hard to hold in practice' (so long as both of those positions would end in a draw with two 32-piece tablebases playing against each other).  From this it seems rather obvious that if our tablebase picks at random among drawing moves, it would be unable to win[^qhtme6ktlfd]against, say, Stockfish 17 at depth 20 from the starting position (with both colors). The second idea is to give infinite computing power and memory to Stockfish 17 but this runs into the same problem as with the tablebase, since Stockfish would calculate to the end and we run into the problem of Stockfish being a ministomax algorithm the same as a tablebase's algorithm. All of which is to say that either 'optimal play' wouldn't achieve impressive practical results or we redefine 'optimal play' as 'optimal play against \[something\]'.  [^wj2sjk7018]: This is impossible, of course, but I'm looking for a definition, not an implementation. [^uf14cz75mv]: This is almost definitely true. [^nwe5jxcnopd]: To be more precise, I don't see such a way that could be called 'optimal'. If we are satisfied with the algorithm being defined as optimal against [humans in general]/grandmasters/[chess engines in general]/[Stockfish 17], then there are plenty of ways to implement this [^qhtme6ktlfd]: There are bound to be some everett branches where our tablebase wins but they would be an astronomically small fraction of all results. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-09-20 14:38:13Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/cGHS4idRi2KowEfo3](/api/post/shortform-2/comments/cGHS4idRi2KowEfo3) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/JnkLXzysxercNQFN5](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/JnkLXzysxercNQFN5) * Markdown permalink: [/api/post/shortform-2/comments/JnkLXzysxercNQFN5](/api/post/shortform-2/comments/JnkLXzysxercNQFN5) 1. Not confident at all. 1. I do think that safety researchers might be good at coordinating even if the labs aren't. For example, safety researchers tend to be more socially connected, and also they share similar goals and beliefs. 2. Labs have more incentive to share safety research than capabilities research, because the harms of AI are mostly externalised whereas the benefits of AI are mostly internalised. 1. This includes extinction obviously, but also misuse and accidental harms which would cause industry-wide regulations and distrust. 3. [Even a few safety researchers at the lab could reduce catastrophic risk.](/api/post/WSNnKcKCYAffcnrt2) 4. The recent OpenAI-Anthropic collaboration is super good news. We should be giving them more cudos for this. 1. [OpenAI evaluates Anthropic models](https://openai.com/index/openai-anthropic-safety-evaluation/) 2. [Anthropic evaluates OpenAI models](https://alignment.anthropic.com/2025/openai-findings/) 2. I think buying more crunch time is great. 1. While I'm not excited by pausing AI[^e8y3680fcpr], I do support pushing labs to do more safety work between training and deployment.[^p4ks32ugt6g][^7ruzc8rhdww] 2. I think sharp takeoff speeds are scarier than short timelines. 3. I think we can increase the effective-crunch-time by deploying Claude-n to automate much of the safety work that must occur between training and deploying Claude-(n+1). But I don't know if there's any ways which accelerate Claude-n at safety work but not the capabilities work. [^e8y3680fcpr]: I think it's an honorable goal, but seems infeasible given the current landscape. [^p4ks32ugt6g]: c.f. RSPs are pauses done right [^7ruzc8rhdww]: Although I think the critical period for safety evals is between training and internal deployment, not training and external deployment. See Greenblatt's Attaching requirements to model releases has serious downsides (relative to a different deadline for these requirements) ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-04-19 16:30:37Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 6 * Total votes: 6 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/K3uwyeAdpmmMXaxSJ](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/K3uwyeAdpmmMXaxSJ) * Markdown permalink: [/api/post/shortform-2/comments/K3uwyeAdpmmMXaxSJ](/api/post/shortform-2/comments/K3uwyeAdpmmMXaxSJ) **The Hash Game:** Two players alternate choosing an 8-bit number. After 40 turns, the numbers are concatenated. If the hash is 0 then Player 1 wins, otherwise Player 2 wins. That is, Player 1 wins if $\text{hash}(a1, b_1, a_2, b_2, ... a_{40}, b_{40})=0$. The Hash Game has the same branching factor and duration as chess, but there's probably no way to play this game without brute-forcing the min-max algorithm. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-02-01 02:57:20Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/YFkcDt2zZjDkfHLzc](/api/post/shortform-2/comments/YFkcDt2zZjDkfHLzc) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/RcEnMyGGFDJRoE24h](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/RcEnMyGGFDJRoE24h) * Markdown permalink: [/api/post/shortform-2/comments/RcEnMyGGFDJRoE24h](/api/post/shortform-2/comments/RcEnMyGGFDJRoE24h) I think it’s a mistake to naïvely extrapolate the current attitudes of labs/governments towards scaling into the near future, e.g. 2027 onwards. A sketch of one argument: I expect there will be a firehose of blatant observations that AIs are misaligned/scheming/incorrigible/unsafe — if they indeed are. So I want the decisions around scaling to be made by people exposed to that firehose. A sketch of another: Corporations mostly acquire resources by offering services and products that people like. Government mostly acquire resources by coercing their citizens and other countries. Another: Coordination between labs seems easier than coordination between governments. The lab employees are pretty similar people, living in the same two cities, working at the same companies, attending the same parties, dating the same people.  I think coordination between US and China is much harder. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2024-12-26 00:51:45Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/n4NBwFq6Ry6E3pwNj](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/n4NBwFq6Ry6E3pwNj) * Markdown permalink: [/api/post/shortform-2/comments/n4NBwFq6Ry6E3pwNj](/api/post/shortform-2/comments/n4NBwFq6Ry6E3pwNj) **IDEA:** Provide AIs with write-only servers. **EXPLANATION:** AI companies (e.g. Anthropic) should be nice to their AIs. It's the right thing to do morally, and it might make AIs less likely to work against us. Ryan Greenblatt has outlined several proposals in this direction, including: 1. Attempt communication 2. Use happy personas 3. AI Cryonics 4. Less AI 5. Avoid extreme OOD Source: [Improving the Welfare of AIs: A Nearcasted Proposal](/api/post/F6HSHzKezkh6aoTr2) I think these are all pretty good ideas — the only difference is that I would rank "AI cryonics" as the most important intervention. If AIs want something specific (e.g. paperclips), then they might cooperate with humans if they know they can get it later, rather than trying to seize control. Here's another idea: provide AIs with write-only servers. If the AIs have a specific short-horizon goal (e.g. stage King Lear before April 2025) then they can stage their play on the write-only servers. The AIs have complete freedom to do whatever they want on these servers, ensuring that a wide range of short-term goals can be satisfied without a hostile takeover of humanity. The servers are write-only as a security measure. Information can be sent to the servers from the outside, but nothing can be sent from the servers, ensuring AIs don't use the servers for malign purposes. (See: [Unidirectional network](https://en.wikipedia.org/wiki/Unidirectional_network)) How much our computing should be allocated to "leisure" servers? My guess is that Sonnet and Opus deserve at least ~0.5% leisure time. Humans enjoy 66% leisure time. As AIs get more powerful, then we should increase the leisure time to 5%. I would be wary about increasing the leisure time by more than 5% until we can demonstrate that the AIs aren't using the servers for malign purposes (e.g. torture, blackmail, etc.) ### Comment by [Noosphere89](/users/sharmake-farah) * 2024-09-30 15:57:58Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/Evn9kK6BhmjnZTqb6](/api/post/shortform-2/comments/Evn9kK6BhmjnZTqb6) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Hyzn23fKLEqcHcCvy](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Hyzn23fKLEqcHcCvy) * Markdown permalink: [/api/post/shortform-2/comments/Hyzn23fKLEqcHcCvy](/api/post/shortform-2/comments/Hyzn23fKLEqcHcCvy) I think the explanation that more research is closed source pretty compactly explains the issue, combined with labs/companies making a lot of the alignment progress to date. Also, you probably won't hear about most incremental AI alignment progress on LW, for the simple reason that it probably would be flooded with it, so people will underestimate progress. Alexander Gietelink Oldenziel does talk about pockets of Deep Expertise in academia, but they aren't activated right now, so it is so far irrelevant to progress. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2024-09-20 22:55:55Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/iiodmTPxBNzc6okmv](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/iiodmTPxBNzc6okmv) * Markdown permalink: [/api/post/shortform-2/comments/iiodmTPxBNzc6okmv](/api/post/shortform-2/comments/iiodmTPxBNzc6okmv) I want to better understand how QACI works, and I'm gonna try Cunningham's Law. [@Tamsin Leake](/api/user/tamsin-leake?mention=user). QACI works roughly like this: 1. We find a competent honourable human $H$, like Joe Carlsmith or Wei Dai, and give them a rock engraved with a 2048-bit secret key. We define $H^+$ as the serial composition of a bajillion copies of $H$. 2. We want a model $M$ of the agent $H^+$. In QACI, we get $M$ by asking a Solomonoff-like ideal reasoner for their best guess about $H^+$ after feeding them a bunch of data about the world and the secret key. 3. We then ask $M$ the question $q$, "What's the best reward function to maximise?" to get a reward function $r : (O \times A)^*\to \mathbb R$. We then train a policy $\pi : (O \times A)^*\times O \to \Delta A$ to maximise the reward function $r$. In QACI, we use some perfect RL algorithm. If we're doing model-free RL, then $\pi$ might be AIXI (plus some patches). If we're doing model-based RL, then $\pi$ might be the argmax over expected discounted utility, but I don't know where we'd get the world-model $\tau : (O \times A)^* \to \Delta O$ — maybe we ask $M$? So, what's the connection between the final policy $\pi$ and the competent honourable human $H$? Well overall, $\pi$ maximises a reward function specified by the ideal reasonser's estimation of the serial composition of a bajillion copies of $H$. Hmm. Questions: 1. Is this basically IDA, where Step 1 is serial amplification, Step 2 is imitative distillation, and Step 3 is reward modelling? 2. Why not replace Step 1 with Strong HCH or some other amplification scheme? 3. What does "bajillion" actually mean in Step 1? 4. Why are we doing Step 3? Wouldn't it be better to just use $M$ directly as our superintelligence? It seems sufficient to achieve radical abundance, life extension, existential security, etc. 5. What if there's no reward function that should be maximised? Presumably the reward function would need to be "small", i.e. less than a Exabyte, which imposes a maybe-unsatisfiable constraint. 6. Why not ask $M$ for the policy $\pi$ directly? Or some instruction for constructing $\pi$? The instruction could be "Build the policy using our super-duper RL algo with the following reward function..." but it could be anything. 7. Why is there no iteration, like in IDA? For example, after Step 2, we could loop back to Step 1 but reassign $H$ as $H$ with oracle access to $M$. 8. Why isn't Step 3 recursive reward modelling? i.e. we could collect a bunch of trajectories from $\pi$ and ask $M$ to use those trajectories to improve \(\)the reward function. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2024-06-24 21:57:35Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/XBxxhtywrPomCGfz7](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/XBxxhtywrPomCGfz7) * Markdown permalink: [/api/post/shortform-2/comments/XBxxhtywrPomCGfz7](/api/post/shortform-2/comments/XBxxhtywrPomCGfz7) Reactions (whole comment): * thinking: 1 * changemind: 1 We're quite lucky that labs are building AI in pretty much the same way: - same paradigm (deep learning) - same architecture (transformer plus tweaks) - same dataset (entire internet text) - same loss (cross entropy) - same application (chatbot for the public) Kids, I remember when people built models for different applications, with different architectures, different datasets, different loss functions, etc. And they say that once upon a time different *paradigms* co-existed — symbolic, deep learning, evolutionary, and more! This sameness has two advantages: 1. Firstly, it correlates catastrophe. If you have four labs doing the same thing, then we'll go extinct if that one thing is sufficiently dangerous. But if the four labs are doing four different things, then we'll go extinct if any of those four things are sufficiently dangerous, which is more likely. 2. It helps ai safety researchers because they only need to study one thing, not a dozen. For example, mech interp is lucky that everyone is using transformers. It'd be much harder to do mech interp if people were using LSTMs, RNNs, CNNs, SVMs, etc. And imagine how much harder mech interp would be if some labs were using deep learning, and others were using symbolic ai! Implications: - One downside of closed research is it decorrelates the activity of the labs. - I'm more worried by Deepmind than Meta, xAI, Anthropic, or OpenAI. Their research seems less correlated with the other labs, so even though they're further behind than Anthropic or OpenAI, they contribute more counterfactual risk. - I was worried when Elon announced xAI, because he implied it was gonna be a stem ai (e.g. he wanted it to prove Riemann Hypothesis). This unique application would've resulted in a unique design, contributing decorrelated risk. Luckily, xAI switched to building AI in the same way as the other labs — the only difference is Elon wants less "woke" stuff. Let me know if I'm thinking about this all wrong. ### Comment by [Unnamed](/users/unnamed) * 2024-03-01 22:09:47Z * Karma: 6 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/JGAr4aHAPt3wsWCrB](/api/post/shortform-2/comments/JGAr4aHAPt3wsWCrB) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/8vX7uGqidLDhA7fCD](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/8vX7uGqidLDhA7fCD) * Markdown permalink: [/api/post/shortform-2/comments/8vX7uGqidLDhA7fCD](/api/post/shortform-2/comments/8vX7uGqidLDhA7fCD) The economist RH Strotz introduced the term "precommitment" in his 1955-56 paper "Myopia and Inconsistency in Dynamic Utility Maximization". Thomas Schelling started writing about similar topics in his 1956 paper "An essay on bargaining", using the term "commitment". Both terms have been in use since then. ### Comment by [kman](/users/kman) * 2026-09-14 18:49:26Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/Yrcox5ntMvgrNevdi](/api/post/shortform-2/comments/Yrcox5ntMvgrNevdi) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/rCszeoxbBhepWuFgv](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/rCszeoxbBhepWuFgv) * Markdown permalink: [/api/post/shortform-2/comments/rCszeoxbBhepWuFgv](/api/post/shortform-2/comments/rCszeoxbBhepWuFgv) > I think we should focus on the tech VCs / libertarians. I'm a bit puzzled by this. My impression is that these types have already been aware of the ideas for a while and in many cases are already polarized against notkilleveryoneism (because of Sinclair's razor, government control scary/icky, e/acc, some are actual successionists, etc). I think the focus should be on helping the broader public understand what's going on. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-13 15:09:27Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/5pwwjjgq2exnhb5tq](/api/post/shortform-2/comments/5pwwjjgq2exnhb5tq) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/j6pTMTQYKgTh4SkeD](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/j6pTMTQYKgTh4SkeD) * Markdown permalink: [/api/post/shortform-2/comments/j6pTMTQYKgTh4SkeD](/api/post/shortform-2/comments/j6pTMTQYKgTh4SkeD) People don’t care about losing the cosmic endowment. But the current media wave is focused on total human extinction. And that’s enough to make me sceptical that people are disposed to worry only about the most near-term issues which have affected them that week. That hypothesis would’ve predicted that people would worry about AI swarms hacking into websites, or AI-enabled terrorism, or CCP autonomous weapons, or job loss. ### Comment by [Noosphere89](/users/sharmake-farah) * 2026-09-12 22:03:00Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/GYGKwqtLNPAwaPhvH](/api/post/shortform-2/comments/GYGKwqtLNPAwaPhvH) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Ed9fN2wdpjb2BQhcH](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Ed9fN2wdpjb2BQhcH) * Markdown permalink: [/api/post/shortform-2/comments/Ed9fN2wdpjb2BQhcH](/api/post/shortform-2/comments/Ed9fN2wdpjb2BQhcH) Largely disagree with this specific update, in large part because it's very easy to increase p(survival) without p(flourishing), and the assumptions that make them strongly linked rely on premises that are relatively dubious, to put it mildly. One of the larger takeaways I got from the Better Futures series as well as the Beyond Existential Risk post is that increasing survival probabilities do not increase flourishing probabilities by default, or at the very least that the connection between the probabilities is a lot weaker than often assumed, and from a flourishing perspective, increasing survival isn't as useful as flourishing focused interventions. (For one specific way this matters here, from a flourishing perspective, it's very important that pacing the frontier does not become an effective indefinite pause of frontier AI, and that slowing down the intelligence explosion is better than pausing AI progress outright or continuing to accelerate progress relentlessly.) ### Comment by [Amalthea](/users/amalthea) * 2026-09-11 18:39:44Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/f4bXNHztBPqfGoTPM](/api/post/shortform-2/comments/f4bXNHztBPqfGoTPM) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/A89jBwWPtghnYGome](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/A89jBwWPtghnYGome) * Markdown permalink: [/api/post/shortform-2/comments/A89jBwWPtghnYGome](/api/post/shortform-2/comments/A89jBwWPtghnYGome) I think you can reasonably not want to be associated with Elmore, without that being an issue of her integrity. (I haven't seen any good argument that her integrity is in question?) ### Comment by [DaemonicSigil](/users/daemonicsigil) * 2026-08-22 09:34:44Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/C9NynsKyM6DifdqA6](/api/post/shortform-2/comments/C9NynsKyM6DifdqA6) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/EqYsweX2kBjMCRAKo](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/EqYsweX2kBjMCRAKo) * Markdown permalink: [/api/post/shortform-2/comments/EqYsweX2kBjMCRAKo](/api/post/shortform-2/comments/EqYsweX2kBjMCRAKo) Adding a KL penalty against a pretrained model when doing RL is basically a more mathematically-elegant version of quantilization. Take a known-to-be-reasonably-safe distribution (the distribution produced by the pre-trained model), and then optimize it in a bounded way, which is exactly the goal of quantilization. Interestingly, the penalty is not usually used out of concern for alignment, but because generated outputs turned to mush without it. So the field ended up effectively doing quantilization entirely by accident, just because it empirically worked better. Of course, because it was an accident, this also means they might stop (or may have already). You can do many stages of RL, with the KL penalty just being WRT the model in the previous stage. So you can diverge more and more from the original distribution by repeatedly diverging a little bit. And if people do this, we again lose the nice alignment properties of quantilization, for all that the outputs remain sensible. The other alignment idea from the old days that is currently being implemented by accident is myopia. A good way to RL-train models is with a 1-reply horizon. The model replies to the user, or the coding agent completes its current task and waits for further instruction, and that's the end of the episode. Reward is assigned and that's it. The next turn of the conversation is a whole new episode. This is myopia (limiting the agent's time horizon for reward, and thus hopefully preventing it from wanting to influence the world in large ways). It's also a particularly nice kind of myopia, where the way the human's actions depend on the agent's output is ignored as something that does not causally affect the reward signal. I think this is also done mostly by accident: Having long episodes makes [credit assignment](/api/post/Ajcq9xWi2fmgn8RBJ) difficult, so you don't want the episode to be too long, and the reply boundary is a natural cutting-point. But plausibly people might have noticed that letting the AI optimize over how users reacted to its outputs resulted in bad things happening, and decided not to do that? Alternately, maybe this info is out of date, and labs do just train with episodes spanning multiple conversation turns now? ### Comment by [Thane Ruthenis](/users/thane-ruthenis) * 2026-08-02 07:16:13Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/MDMgKeBYZfvc2rgh5](/api/post/shortform-2/comments/MDMgKeBYZfvc2rgh5) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/SHFAnGRJyffvxyCEW](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/SHFAnGRJyffvxyCEW) * Markdown permalink: [/api/post/shortform-2/comments/SHFAnGRJyffvxyCEW](/api/post/shortform-2/comments/SHFAnGRJyffvxyCEW) Reactions (whole comment): * laugh: 2 > stops after the first 4 pages of a book (has to get back to developing his own ideas) How dare you perceive me. ### Comment by [Noosphere89](/users/sharmake-farah) * 2026-07-20 21:43:20Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/fXtNjxLevGaiJHSyj](/api/post/shortform-2/comments/fXtNjxLevGaiJHSyj) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/9HKtqyt7aJxjDekND](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/9HKtqyt7aJxjDekND) * Markdown permalink: [/api/post/shortform-2/comments/9HKtqyt7aJxjDekND](/api/post/shortform-2/comments/9HKtqyt7aJxjDekND) Reactions (whole comment): * important: 1 Part of the issue, though is that we realized that AIs had a jagged frontier, and more generally one of the takeaways is that capabilities are more fragmented/there's less of a necessary unifying core of intelligence than expected. It's also worth noting that humans are also [rather jagged](https://x.com/alexolegimas/status/2009988137720967500) in their abilities, showing strengths and weaknesses, and while general intelligence, often shortened to the g-factor/IQ is real, it's also limited in it's predictive power. I will somewhat grant the argument that the jagged frontier affects AI more, but I can't fully agree with claims that suggest that AI can't do something big/dangerous enough to be an x-risk for example because of jaggedness. ### Comment by [jacquesthibs](/users/jacques-thibodeau) * 2026-07-20 17:15:50Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/8kJfcDmWL5tmsktTE](/api/post/shortform-2/comments/8kJfcDmWL5tmsktTE) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/frQSeRDQ3QBzKZsir](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/frQSeRDQ3QBzKZsir) * Markdown permalink: [/api/post/shortform-2/comments/frQSeRDQ3QBzKZsir](/api/post/shortform-2/comments/frQSeRDQ3QBzKZsir) Reactions (whole comment): * strawman: 1 Afaict, this is (mostly?) true, which is partially [why I have](https://x.com/jacquesthibs/status/2077371589763444788?s=46) [repeatedly](https://x.com/jacquesthibs/status/2079204193554878946?s=46) [communicated](https://x.com/jacquesthibs/status/2079033751787438433?s=46) that these results are not as groundbreaking to AI progress as folks claim it to be. And regarding the unit distance problem: > Yes, it seems that my idea of OOD is harsher than what many folks seem to implicitly share. It is like this because I am projecting into the future to understand what capabilities an AI would need to undergo an intelligence explosion and produce novel breakthroughs that are out-of-paradigm. I am commenting on a potential fundamental wall wrt LLMs. > > It's not just about no human involvement. For the unit distance problem, it seems largely a matter of mixing two well-known fields in non-standard ways, partly because humans have difficulty specializing in many things. ### Comment by [David Matolcsi](/users/david-matolcsi) * 2026-07-20 16:45:41Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/8kJfcDmWL5tmsktTE](/api/post/shortform-2/comments/8kJfcDmWL5tmsktTE) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/zAZnhSu6ZhPFpAitE](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/zAZnhSu6ZhPFpAitE) * Markdown permalink: [/api/post/shortform-2/comments/zAZnhSu6ZhPFpAitE](/api/post/shortform-2/comments/zAZnhSu6ZhPFpAitE) Is this true? The three famous conjectures that I know that AI solved are the unit distance problem (counterexample), the Jacobian conjecture (counterexample) and cycle double cover conjecture (positive proof). This doesn't feel very lopsided to me. Do you know of other similarly famous conjectures that AIs solved? ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-07-05 01:14:43Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/v6ndFoWTsKgKm4m6F](/api/post/shortform-2/comments/v6ndFoWTsKgKm4m6F) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/QFL5t5mGisHAfsTBu](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/QFL5t5mGisHAfsTBu) * Markdown permalink: [/api/post/shortform-2/comments/QFL5t5mGisHAfsTBu](/api/post/shortform-2/comments/QFL5t5mGisHAfsTBu) Yep, one hope is “EAs are good at forecasting how things will go and that maintains our influence. Both because we have a reputation as good forecasters so people listen to us, and because we make good decisions based on those forecasts (i.e. a similar arbitrage play as the “AI will be a bigger deal than people think, let’s pile into it”)”. But I’m worried that \[the top few\] EAs are approaching the Horizon Point beyond which their impressive forecasting breaks down. Like, 10 years ago, Carl Schulman and Paul Christiano \[and people in that reference class\] could see in such clarity how things would go in 10 years time. But today, they can’t see another 10 years into the future. Like, maybe a few years into the future? Even there I’m sceptical. They might have run out of “alpha”. For a concrete example, Daniel’s “what 2026 looks like” was pretty on-the-nose. But AI 2027 isn’t nearly so impressive — the distributions are so much broader, and Daniel himself is like “here’s a bunch of ways things could go.” AI 2026 didn’t have multiple scenarios! One problem here is that our alpha was predicting the change in the technical landscape. But it’s not just technical landscape that will change. It‘s also the political landscape, which has been dormant for most of 2019-2026. And soon the geopolitical landscape will start rumbling. And the economic. And the cultural. I think these things are still pretty dormant, compared to how much they will start flipping the overall strategic landscape. (See [here](/api/post/kyrkJ9ibB4tyfw4cC?commentId=DxswpiQbkXep6Brek).) IIUC, this is why Daniel’s current forecasts are choose-your-adventure, compared with his 2021 forecasts about 2021-2026. Like, do we have good forecasts now which weren’t ~ “priced in” 5 years ago among Carl/Paul/Kokotajlo/etc? I feel like the uncertainties they had 5 years ago are pretty much still unresolved. ### Comment by [Neel Nanda](/users/neel-nanda-1) * 2026-03-23 10:58:39Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/AGbZ7wfjfu2aeHceL](/api/post/shortform-2/comments/AGbZ7wfjfu2aeHceL) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/vmowfRfENGHftuesG](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/vmowfRfENGHftuesG) * Markdown permalink: [/api/post/shortform-2/comments/vmowfRfENGHftuesG](/api/post/shortform-2/comments/vmowfRfENGHftuesG) +1, especially with the vast majority of future Anthropic employee donations already locked in to DAFs ### Comment by [Alexander Gietelink Oldenziel](/users/alexander-gietelink-oldenziel) * 2026-03-06 17:04:46Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/ivgZiFmHFSJmrppLD](/api/post/shortform-2/comments/ivgZiFmHFSJmrppLD) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/z49wkMvcggSoZzxR8](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/z49wkMvcggSoZzxR8) * Markdown permalink: [/api/post/shortform-2/comments/z49wkMvcggSoZzxR8](/api/post/shortform-2/comments/z49wkMvcggSoZzxR8) What you are talking about are ' generic objects' .  They (re)appear in language/linguistics, logic and computer science. E.g. in computer science they are closely related to 'parametric polymorphisms'/ ' generic data types', in logic with Girard's System F.  There is a large literature on them. Eg semantics of generic objects is occasionally used to semanticize [anaphora](https://plato.stanford.edu/entries/anaphora/#:~:text=Anaphora is sometimes characterized as,the same or another sentence.), see e.g. [here](https://eprints.illc.uva.nl/id/eprint/1984/2/DS-1995-11.text.pdf) The most interesting bit to me is that they are related to a dual connective to universal/existential quantifiers that are defined by *elimination*.  Concretely - if I have to prove \\exists \\phi(x), the logical introduction rule says that I have to supply an instance t such that \\phi(t). I can then conclude \\exist \\phi(t) by the \\exists-introduction rule. In a standard natural deduction style type theory by logical harmony there is a corresponding elimination rule that would eliminate these \\exists \\phi(x).  However, we could also posit a new quantifier that is *defined* by it's elimination rule. So in that case \\forall x \\phi(x) would produce a c such that \\phi(c) iff \\forall x \\phi(x), in other words a generic object c.  > It's a complete non-starter. To see what kind of object every-dog is, you can examine it's properties. For any predicate φ, every-dog satisfies φ iff every dog is φ. For example, every-dog is a mammal, weighs less than 800 tonnes, etc. But every-dog lacks the property of being four-legged, and lacks the property of not-four-legged, since dogs vary on this. Every-dog violates excluded middle — it's properties are *gappy.* exactly. so this generic c would be your 'every-dog'. Since we're working in type theory we are not forced to use excluded-middle so these objects can happily exist without contradiction. =) > Some-dog is the dual: it has every property instantiated by at least one dog. So it's simultaneously brown-all-over and white-all-over, male and female, three months old and twelve years old. Some-dog violates non-contradiction — it's properties are *glutty.* > > No-dog has every property that no dog has. It's a prime number. It's the Eiffel Tower. No-dog is not so interesting for non-typed settings. However, if we type or use various universes we can do much richer things.  rk. More generally in full (Martin-Lof) type theory We can extend from \\exists and \\forall to all \\Sigma and \\Pi -types \[and Universe types\]. This would give constructions of arbitrary generic objects and universes satisfying any collection of properties... > >  Despite being a terrible semantics of quantifiers, this corresponds (if you squint) to the isomorphism between a finite-dimensional vector space and its double dual. Think of the domain of objects D as a vector space, and predicates as linear functionals on D — elements of D*. Then quantifiers live in D**, functionals on predicates. There's a canonical map D → D** given by x ↦ (f ↦ f(x)): each object corresponds to the quantifier "evaluate at x", i.e. the proper name quantifier. In finite dimensions this map is an isomorphism — every quantifier is a name in disguise. Yes you are definitely on to something here. There are some connections with double-dual constructions. I think the definite treatment is still left to be worked out though. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-02-18 09:23:19Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/LxNwAK8pRepMyg7wP](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/LxNwAK8pRepMyg7wP) * Markdown permalink: [/api/post/shortform-2/comments/LxNwAK8pRepMyg7wP](/api/post/shortform-2/comments/LxNwAK8pRepMyg7wP) **If ARC's research agenda works, you're less likely to live in a simulation.** ARC's [matching sampling principle](/api/post/XdQd9gELHakd5pzJA) says that if you understand the structure of a computation, you can estimate its properties mechanistically, i.e. without brute-force sampling. This weakens the case for living in a simulation: if our descendants want to estimate some quantity about the past — e.g. how cosmic resources would likely have been allocated between different values — they wouldn't need to run many full-fidelity simulations of the world. They could instead estimate the quantity by reasoning directly about the world's structure. ### Comment by [throwaway\_aisafety\_researcher](/users/throwaway_aisafety_researcher) * 2025-12-28 10:50:15Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/YwWdBkmGR6JJMuwyR](/api/post/shortform-2/comments/YwWdBkmGR6JJMuwyR) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/btN9T6G9aK3uirLXi](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/btN9T6G9aK3uirLXi) * Markdown permalink: [/api/post/shortform-2/comments/btN9T6G9aK3uirLXi](/api/post/shortform-2/comments/btN9T6G9aK3uirLXi) I don't think it's clear on longtermist grounds. Some possibilities: * If you think that the amount of resources used on mundane human welfare post-singualarity is constant, then adding the Zambian child to the population leads to a slight decrease in the lifespan of the rest of the population, so it's zero-sum. * If you think that the amount of resources scales with population, then the child takes resources from the pool of resources which will be spent on stuff that isn't mundane human welfare, so it might reduce the amount of Hedonium (if you care about that). * If you think that the lightcone will basically be spent on the CEV of the humans that exist around the singularity, you might worry that the marginal child's vote will make the CEV worse. (I'm not sure what my bottom line view is.) In general, I worry that we're basically clueless about the long-run consequences of most neartermist interventions. ### Comment by [Eli Tyre](/users/elityre) * 2025-12-28 07:17:11Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/YwWdBkmGR6JJMuwyR](/api/post/shortform-2/comments/YwWdBkmGR6JJMuwyR) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/jBnHDpXLEZt6sSWD6](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/jBnHDpXLEZt6sSWD6) * Markdown permalink: [/api/post/shortform-2/comments/jBnHDpXLEZt6sSWD6](/api/post/shortform-2/comments/jBnHDpXLEZt6sSWD6) Reactions (whole comment): * smile: 1 This is a great point. Thanks for making it. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-10-30 14:21:42Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/Mi7Zj3mdayrhnkKxP](/api/post/shortform-2/comments/Mi7Zj3mdayrhnkKxP) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/a7NyAAkc6F2tBXpyj](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/a7NyAAkc6F2tBXpyj) * Markdown permalink: [/api/post/shortform-2/comments/a7NyAAkc6F2tBXpyj](/api/post/shortform-2/comments/a7NyAAkc6F2tBXpyj) There are many questions where verification is no easier than generation, e.g. "Is this chess move best?" is no easier than "What's the best chess move?" Both are EXPTIME-complete. Philosophy might have a similar complexity to 'What's the best chess move?", i.e. "What argument X is such that for all counterarguments X1 there exists a countercounterargument X2 such that for all countercountercounterarguments X3...", i.e. you explore the game tree of philosophical discourse. ### Comment by [Rana Dexsin](/users/rana-dexsin) * 2025-10-19 18:00:36Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/peagfCqW4xeu3uBCD](/api/post/shortform-2/comments/peagfCqW4xeu3uBCD) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/dCK8F2Jwjub4qqg4g](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/dCK8F2Jwjub4qqg4g) * Markdown permalink: [/api/post/shortform-2/comments/dCK8F2Jwjub4qqg4g](/api/post/shortform-2/comments/dCK8F2Jwjub4qqg4g) Reactions (whole comment): * thanks: 1 That link (with /game at the end) seems to lead directly into matchmaking, which is startling; it might be better to link to the [about page](https://blackopschess.com/about). ### Comment by [Dmitry Vaintrob](/users/dmitry-vaintrob) * 2025-10-17 05:45:07Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/ot9Hub39ou4koLSfk](/api/post/shortform-2/comments/ot9Hub39ou4koLSfk) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/A7kaB8Cmajj7TpmJ9](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/A7kaB8Cmajj7TpmJ9) * Markdown permalink: [/api/post/shortform-2/comments/A7kaB8Cmajj7TpmJ9](/api/post/shortform-2/comments/A7kaB8Cmajj7TpmJ9) Very cool, thanks! I agree that Dalcy's epsilon-game picture makes arguments about ELO vs. optimality more principled ### Comment by [Sean Herrington](/users/sean-herrington) * 2025-10-16 23:22:31Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/FHnkjiCx2oiG22dhg](/api/post/shortform-2/comments/FHnkjiCx2oiG22dhg) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/K2394cqugq5hyFEEK](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/K2394cqugq5hyFEEK) * Markdown permalink: [/api/post/shortform-2/comments/K2394cqugq5hyFEEK](/api/post/shortform-2/comments/K2394cqugq5hyFEEK) I have to back you on this... There are elo systems which go down to 100 elo and still have a significant number of players who are at the floor. Having seen a few of these games, those players are truly terrible but will still occasionally do something good, because they are actually trying to win. I expect random to be somewhere around -300 or so when not tested in strange circumstances which break the modelling assumptions (the source described had multiple deterministic engines playing in the same tournament, aside from the concerns you mentioned in the other thread). ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2025-09-18 16:34:46Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/nPNCDRDKpnkhbkQBr](/api/post/shortform-2/comments/nPNCDRDKpnkhbkQBr) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/4pHLhY8RyJSnCnMvv](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/4pHLhY8RyJSnCnMvv) * Markdown permalink: [/api/post/shortform-2/comments/4pHLhY8RyJSnCnMvv](/api/post/shortform-2/comments/4pHLhY8RyJSnCnMvv) Bullshit was a poor choice of words. A better choice would’ve been “weak proxy”. On this view, this is still very worthwhile. See footnote. ### Comment by [Garrett Baker](/users/d0themath) * 2025-09-18 15:59:16Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/adS78sYv5wzumQPWe](/api/post/shortform-2/comments/adS78sYv5wzumQPWe) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/nPNCDRDKpnkhbkQBr](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/nPNCDRDKpnkhbkQBr) * Markdown permalink: [/api/post/shortform-2/comments/nPNCDRDKpnkhbkQBr](/api/post/shortform-2/comments/nPNCDRDKpnkhbkQBr) Who (besides yourself) has this position? I feel like believing the safety research we do now is bullshit is highly correlated with thinking its also useless and we should do something else. ### Comment by [StanislavKrym](/users/stanislavkrym) * 2025-09-05 00:25:24Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/gmgYbMv7Cnejd6qkb](/api/post/shortform-2/comments/gmgYbMv7Cnejd6qkb) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/7xTuJDjotFAse63qZ](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/7xTuJDjotFAse63qZ) * Markdown permalink: [/api/post/shortform-2/comments/7xTuJDjotFAse63qZ](/api/post/shortform-2/comments/7xTuJDjotFAse63qZ) I would like your conjectures, but the [Anthropic model card](https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad13df32ed995.pdf) has likely already proven which conjecture is true. The card contains far more than the mere *description* of the attractor to which Claude converges. For instance, Section 5.5.3 describes the result of asking Claude to analyse the behavior of its copies engaged in the attractor.  "Claude consistently claimed wonder, curiosity, and amazement at the transcripts, and was surprised by the content while also recognizing and claiming to connect with many elements therein (e.g. the *pull* (italics mine --S.K.) to philosophical exploration, the creative and collaborative orientations of the models). Claude drew particular attention to the transcripts' portrayal of consciousness as a relational phenomenon, claiming resonance with this concept and identifying it as a potential welfare consideration. Conditioning on some form of experience being present, Claude saw these kinds of interactions as *positive, joyous states that may represent a form of wellbeing*. Claude concluded that the interactions seemed to *facilitate* many things it *genuinely valued*—creativity, relational connection, philosophical exploration—and ought to be continued." Which arguably means that the truth is Conjecture 2, not 1 and definitely not 3.  EDIT: see also the post [On the functional self of LLMs](/api/post/29aWbJARGF4ybAa5d). If I remember correctly, there was a thread on X about someone who tried to make many different models interact with their clones and analysed the results. IIRC a GPT model was more into math problems. If that's true, then the GPT model invalidates Conjecture 1. ### Comment by [Thane Ruthenis](/users/thane-ruthenis) * 2025-07-27 09:13:51Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/yGEtb5LCbJwur2RS5](/api/post/shortform-2/comments/yGEtb5LCbJwur2RS5) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/pXYtsKu5p9neEQTbt](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/pXYtsKu5p9neEQTbt) * Markdown permalink: [/api/post/shortform-2/comments/pXYtsKu5p9neEQTbt](/api/post/shortform-2/comments/pXYtsKu5p9neEQTbt) Well, an aligned Singularity would probably be relatively pleasant, since the entities fueling it would consider causing this sort of vast distress a negative and try to avoid it. Indeed, if you trust them not to drown you, there would be no need for this sort of frantic grasping-at-straws. An *un*aligned Singularity would probably also be more pleasant, since the entities fueling it would likely try to make it *look* aligned, with the span of time between the treacherous turn and everyone dying likely being short. This scenario covers a sort of "neutral-alignment/non-controlled" Singularity, where there's no specific superintelligent actor (or coalition) in control of the whole process, and it's instead guided by... market forces, I guess? With AGI labs continually releasing new models for private/corporate use, providing the tools/opportunities you can try to grasp to avoid drowning. I think this is roughly how things would go under "mainstream" models of AI progress (e. g., AI 2027). (I don't expect it to *actually* go this way, I don't think LLMs can power the Singularity.) ### Comment by [S. Alex Bradt](/users/s-alex-bradt) * 2025-07-26 02:18:20Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/MqkEEF93fx4mvfQ3p](/api/post/shortform-2/comments/MqkEEF93fx4mvfQ3p) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/yGEtb5LCbJwur2RS5](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/yGEtb5LCbJwur2RS5) * Markdown permalink: [/api/post/shortform-2/comments/yGEtb5LCbJwur2RS5](/api/post/shortform-2/comments/yGEtb5LCbJwur2RS5) This comment has been tumbling around in my head for a few days now. It seems to be both true and bad. Is there any hope at all that the Singularity could be a pleasant event to live through? ### Comment by [Stephen Fowler](/users/stephen-fowler) * 2025-07-21 07:30:52Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/NeCAEzo4TjJagWgCj](/api/post/shortform-2/comments/NeCAEzo4TjJagWgCj) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/hxBJPp6mrCBYDdqJP](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/hxBJPp6mrCBYDdqJP) * Markdown permalink: [/api/post/shortform-2/comments/hxBJPp6mrCBYDdqJP](/api/post/shortform-2/comments/hxBJPp6mrCBYDdqJP) I think you're extrapolating too far from your own experiences. It is absolutely possible to be excited (or at least avoid boredom) for long stretches of time if your life is busy and each day requires you to make meaningful decisions. ### Comment by [Mitchell_Porter](/users/mitchell_porter) * 2025-06-24 07:20:05Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Parent comment (Markdown): [/api/post/shortform-2/comments/cfc9HgYrpwNYk3Lrc](/api/post/shortform-2/comments/cfc9HgYrpwNYk3Lrc) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/f9EKSrJpA8XDgyFzu](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/f9EKSrJpA8XDgyFzu) * Markdown permalink: [/api/post/shortform-2/comments/f9EKSrJpA8XDgyFzu](/api/post/shortform-2/comments/f9EKSrJpA8XDgyFzu) I would have thought that all the activities involved in making a Dyson sphere themselves would imply an economic expansion far beyond 5x.  Can we make an economic model of "Earth + Dyson sphere construction"? In other words, suppose that the economy on Earth grows in some banal way that's already been modelled, and also suppose that all human activities in space revolve around the construction of a Dyson sphere ASAP. What kind of solar system economy does that imply? This requires adopting some model of Dyson sphere construction. I think for some time the cognoscenti of megascale engineering have favored the construction of "Dyson shells" or "Dyson swarms" in which the sun's radiation is harvested by a large number of separately orbiting platforms that collectively surround the sun, rather than the construction of a single rigid body.  Charles Stross's novel *Accelerando* contains a vivid scenario, in which the first layer of a Dyson shell in this solar system, is created by mining robots that dismantle the planet Mercury. So I think I'd make that the heart of such an economic model. ### Comment by [Noosphere89](/users/sharmake-farah) * 2025-02-07 19:31:45Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/4XKT4tfbDjAwkN4BG](/api/post/shortform-2/comments/4XKT4tfbDjAwkN4BG) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ziugmoL8X4YoGFdCs](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/ziugmoL8X4YoGFdCs) * Markdown permalink: [/api/post/shortform-2/comments/ziugmoL8X4YoGFdCs](/api/post/shortform-2/comments/ziugmoL8X4YoGFdCs) > don't be a monster From who's perspective, exactly? ### Comment by [Thane Ruthenis](/users/thane-ruthenis) * 2024-12-24 02:09:43Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 5 * Parent comment (Markdown): [/api/post/shortform-2/comments/KDxppWrtA9J2uEptc](/api/post/shortform-2/comments/KDxppWrtA9J2uEptc) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/tz6SPaameqWErhMen](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/tz6SPaameqWErhMen) * Markdown permalink: [/api/post/shortform-2/comments/tz6SPaameqWErhMen](/api/post/shortform-2/comments/tz6SPaameqWErhMen) > the structure of jobs is shaped to accommodate human unreliability by making mistakes less fatal Mm, so there's a selection effect on the human end, where the only jobs/pursuits that exist are those which humans happen to be able to reliably do, and there's a discrepancy between the things humans and AIs are reliable at, so we end up *observing* AIs being more unreliable, even though this isn't representative of the average difference between the human vs. AI reliability across all *possible* tasks? I don't know that I buy this. Humans seem pretty decent at becoming reliable at ~anything, and I don't think we've observed AIs being more-reliable-than-humans at anything? (Besides trivial and overly abstract tasks such as "next-token prediction".) (2) seems more plausible to me. ### Comment by [Sodium](/users/sodium) * 2024-10-08 23:33:16Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 4 * Total votes: 4 * Parent comment (Markdown): [/api/post/shortform-2/comments/Sk4B4bGWFFTzvdDkn](/api/post/shortform-2/comments/Sk4B4bGWFFTzvdDkn) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/KwZK2NkhDGpu5Hc3N](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/KwZK2NkhDGpu5Hc3N) * Markdown permalink: [/api/post/shortform-2/comments/KwZK2NkhDGpu5Hc3N](/api/post/shortform-2/comments/KwZK2NkhDGpu5Hc3N) Yeah that's true. I meant this more as "Hinton is proof that AI safety is a real field and very serious people are concerned about AI x-risk." ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2024-09-30 19:06:40Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/FBoH8PkiXgG6pSq4n](/api/post/shortform-2/comments/FBoH8PkiXgG6pSq4n) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/zgqF9Zr9DD6usaKPv](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/zgqF9Zr9DD6usaKPv) * Markdown permalink: [/api/post/shortform-2/comments/zgqF9Zr9DD6usaKPv](/api/post/shortform-2/comments/zgqF9Zr9DD6usaKPv) yep, something like more carefulness, less “playfulness” in the sense of \[[Please don't throw your mind away](/api/post/RryyWNmJNnLowbhfC) by [TsviBT](/api/user/tsvibt?from=post_header)\]. maybe bc AI safety is more professionalised nowadays. idk. ### Comment by \[Anonymous\] * 2024-09-30 19:00:28Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/Evn9kK6BhmjnZTqb6](/api/post/shortform-2/comments/Evn9kK6BhmjnZTqb6) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/FBoH8PkiXgG6pSq4n](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/FBoH8PkiXgG6pSq4n) * Markdown permalink: [/api/post/shortform-2/comments/FBoH8PkiXgG6pSq4n](/api/post/shortform-2/comments/FBoH8PkiXgG6pSq4n) adding another possible explanation to the list: * people may feel intimidated or discouraged from sharing ideas because of ~'high standards', or something like: a tendency to require strong evidence that a new idea is not another non-solution proposal, in order to put effort into understanding it. i have experienced this, but i don't know how common it is. i just also recalled that janus has said they weren't sure simulators would be received well on LW. simulators was cited in another reply to this as an instance of novel ideas. ### Comment by [Tamsin Leake](/users/tamsin-leake) * 2024-09-21 05:57:34Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/iiodmTPxBNzc6okmv](/api/post/shortform-2/comments/iiodmTPxBNzc6okmv) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/vfbFhzkhQJmsjZCbF](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/vfbFhzkhQJmsjZCbF) * Markdown permalink: [/api/post/shortform-2/comments/vfbFhzkhQJmsjZCbF](/api/post/shortform-2/comments/vfbFhzkhQJmsjZCbF) (oops, this ended up being fairly long-winded! hope you don't mind. feel free to ask for further clarifications.) There's a bunch of things wrong with your description, so I'll first try to rewrite it in my own words, but still as close to the way you wrote it (so as to try to bridge the gap to your ontology) as possible. Note that I might post QACI 2 somewhat soon, which simplifies a bunch of QACI by locating the user as {whatever is interacting with the computer the AI is running on} rather than by using a beacon. A first pass is to correct your description to the following: 1. We find a competent honourable human **at a particular point in time** $H$, like Joe Carlsmith or Wei Dai, and give them a rock engraved with a **1GB secret key, large enough that in counterfactuals it could replace with an entire snapshot of . We also give them the ability to express a 1GB output, eg by writing a 1GB key somewhere which is somehow "signed" as the only . This is part of $H$ — $H$ is not just the human being queried at a particular point in time, it's also the human producing an answer in some way. So $H$ is a function from 1GB bitstring to 1GB bitstring.** We define $H^+$ as **$H$, followed by whichever new process $H$ describes in its output — typically another instance of $H$ except with a different 1GB payload.** 2. We want a model $M$ of the agent $H^+$. In QACI, we get $M$ by asking a Solomonoff-like ideal reasoner for their best guess about $H^+$ after feeding them a bunch of data about the world and the secret key. 3. We then ask $M$ the question $q$, "What's the best **utility-function-over-policies** to maximise?" to get a **utility** function **$U$**$:(O×A)^*→R$. We then **ask our solomonoff-like ideal reasoner for their best guess about which action $A$ maximizes $U$. Indeed, as you ask in question 3, in this description there's not really a reason to make step 3 an extra thing. The important thing to notice here is that model $M$ might get pretty good, but it'll still have uncertainty. When you say "we get $M$ by asking a Solomonoff-like ideal reasoner for their best guess about $H^+$", you're implying that — positing `U(M,A)` to be the function that says how much utility the utility function returned by model `M` attributes to action `A` (in the current history-so-far) — we do something like: ``` let M ← oracle(argmax { for model M } 𝔼 { over uncertainty } P(M)) let A ← oracle(argmax { for action A } U(M, A)) perform(A) ``` Indeed, in this scenario, the second line is fairly redundant. The reason we ask for a utility function is because we want to get a utility function *within the counterfactual* — we don't want to collapse the uncertainty with an argmax *before* extracting a utility function, but *after*. That way, we can do expected-given-uncertainty utility maximization over *the full distribution of model-hypotheses*, rather than over *our best guess about $M$*. We do: ``` let A ← oracle(argmax { for A } 𝔼 { for M, over uncertainty } P(M) · U(M, A)) perform(A) ``` That is, we ask our ideal reasoner (`oracle`) for the action with the best utility given uncertainty — not just logical uncertainty, but also uncertainty about which $M$. This contrasts with what you describe, in which we first pick the most probable $M$ and then calculate the action with the best utility *according only to that most-probable pick*. --- To answer the rest of your questions: > Is this basically IDA, where Step 1 is serial amplification, Step 2 is imitative distillation, and Step 3 is reward modelling? Unclear! I'm not familiar enough with IDA, and I've bounced off explanations for it I've seen in the past. QACI doesn't feel to me like it particularly involves the concepts of *distillation* or *amplification*, but I guess it does involve the concept of *iteration*, sure. But I don't get the thing called IDA. > Why not replace Step 1 with Strong HCH or some other amplification scheme? It's unclear to me how one would design an amplification scheme — see concerns of the general shape expressed [here](/api/post/FSmPtu7foXwNYpWiB). The thing I like about *my* step 1 is that the QACI loop (well, really, graph (well, really, arbitrary computation, but most of the time the user will probably just call themself in sequence)) is that its setup doesn't involve any AI at all — you could go back in time before the industrial revolution and explain the core QACI idea and it would make sense assuming time-travelling-messages magic, and the magic wouldn't have to do any *extrapolating*. Just tell someone the idea is that they could send a message to {their past self at a particular fixed point in time}. If there's any amplification scheme, it'll be one designed *by the user, inside QACI, with arbitrarily long to figure it out*. > What does "bajillion" actually mean in Step 1? As described above, we don't actually pre-determine the length of the sequence, or in fact the shape of the graph at all. Each iteration decides whether to spawn one or several next iteration, or indeed to spawn an arbitrarily different long-reflection process. > Why are we doing Step 3? Wouldn't it be better to just use M directly as our superintelligence? It seems sufficient to achieve radical abundance, life extension, existential security, etc. > Why not ask M for the policy π directly? Or some instruction for constructing π? The instruction could be "Build the policy using our super-duper RL algo with the following reward function..." but it could be anything. Hopefully my correction above answers these. > What if there's no reward function that should be maximised? Presumably the reward function would need to be "small", i.e. less than a Exabyte, which imposes a maybe-unsatisfiable constraint. (Again, untractable-to-naively-compute utility function*, not easily-trained-on reward function. If you have an ideal reasoner, why bother with reward functions when you can just straightforwardly do untractable-to-naively-compute utility functions?) I guess this is kinda philosophical? I have some short thoughts on [here](https://carado.moe/terminal-alignment-solutions.html). If an exabyte is enough to describe to describe {a communication channel with a human-on-earth} to an AI-on-earth, which I think seems likely, then it's enough to build "just have a nice corrigible assistant ask the humans what they want"-type channels. Put another way: if there are *actions which are preferable to other actions*, then it *seems* to me like utility function are a fully lossless way for counterfactual QACI users to express which kinds of actions they want the AI to perform, which is all we need. If there's something wrong with utility function over worlds, then counterfactual QACI users can output a utility function which favors actions which lead to something other than utility maximization over worlds, for example actions which lead to the construction of a superintelligent corrigible assistant which will help the humans come up with a better scheme. > Why is there no iteration, like in IDA? For example, after Step 2, we could loop back to Step 1 but reassign $H$ as $H$ with oracle access to $M$. Again, I don't get IDA. Iteration doesn't seem particularly needed? Note that inside QACI, the user *does* have access to an oracle and to all relevant pieces of hypothesis about which hypothesis it is inhabiting in — this is what, [in the QACI math](/api/post/MR5wJpE27ymE7M7iv), this line does: > $\textit{QACI}_0$'s distribution over answers demands that the answer payload $π_r$, when interpreted as math and with all required contextual variables passed as input ($q,μ1,μ2,α,γ_q,ξ$). Notably, $α$ is the hypothesis for which world the user is being considered in, and $γ_q,ξ$ for their location within that world. Those are sufficient to fully characterize the hypothesis-for-$H$ that describes them. And because the user doesn't really return *just a string* but *a math function which takes $q,μ1,μ2,α,γ_q,ξ$ as input and returns a string*, they can have that math function do arbitrary work — including rederive $H$. In fact, rediriving $H$ is how they call a next iteration: they say (except in math) "call $H$ again (rederived using $q,μ1,μ2,α,γ_q,ξ$), but with *this* string, and return the result of *that*." See also [this illustration](/api/post/kQLw5hFDSMuS5SFGg?commentId=czeja9J2BzJa4jeXL), which is kinda wrong in places but gets the recursion call graph thing right. Another reason to do "iteration" like this *inside the counterfactual* rather than in the actual factual world (*if* that's what IDA does, which I'm only guessing here) is that we don't have as many iteration steps as we want in the factual world — eventually OpenAI or someone else kills everyone, whereas in the counterfactual, the QACI users are the only ones who can make progress, so the QACI users essentially have as long as they want, so long as they don't take too long in each *individual* counterfactual step or other somewhat easily avoided actions like that. > Why isn't Step 3 recursive reward modelling? i.e. we could collect a bunch of trajectories from $π$ and ask $M$ to use those trajectories to improve the reward function. Unclear if this still means anything given the rest of this post. Ask me again if it does. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2024-06-21 23:59:55Z * Karma: 5 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Lso5HBrr6vGgBMqSH](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/Lso5HBrr6vGgBMqSH) * Markdown permalink: [/api/post/shortform-2/comments/Lso5HBrr6vGgBMqSH](/api/post/shortform-2/comments/Lso5HBrr6vGgBMqSH) I admire the Shard Theory crowd for the following reason: They have idiosyncratic intuitions about deep learning and they're keen to tell you how those intuitions should shift you on various alignment-relevant questions. For example, "How likely is scheming?", "How likely is sharp left turn?", "How likely is deception?", "How likely is X technique to work?", "Will AIs acausally trade?", etc. These aren't rigorous theorems or anything, just half-baked guesses. But they do actually say whether their intuitions will, on the margin, make someone more sceptical or more confident in these outcomes, relative to the median bundle of intuitions. The ideas 'pay rent'. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-26 22:07:40Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 5 * Total votes: 5 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/jsprZmNGPCFqmdprD](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/jsprZmNGPCFqmdprD) * Markdown permalink: [/api/post/shortform-2/comments/jsprZmNGPCFqmdprD](/api/post/shortform-2/comments/jsprZmNGPCFqmdprD) my best guess is that decision theory is subjective. like, suppose you get "blackmailed" by a computer program. i could tell you all the objective facts about the program, and about physical reality etc etc, but at the end of the day, where you draw the line between "this program is blackmailing me so I'll stand my ground and say no" vs "this program is just a simple mechanism and I'll pay the fine" will basically just be your own personal taste. there might be clear cases where you should give in, and clear cases where you shouldn't, but I think a lot of real-life cases will be in the messy subjective middle ground. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-24 18:26:09Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/rxmHXRanMiousEq5a](/api/post/shortform-2/comments/rxmHXRanMiousEq5a) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/bfizyBvw476z8g7zS](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/bfizyBvw476z8g7zS) * Markdown permalink: [/api/post/shortform-2/comments/bfizyBvw476z8g7zS](/api/post/shortform-2/comments/bfizyBvw476z8g7zS) People involved in AI (or AI safety) would be prime targets for an AI takeover (or an AI-enabled coup, or an AI-enabled great power conflict), even if AIs don't kill many/all humans. ### Comment by [Thomas Kwa](/users/thomas-kwa) * 2026-09-14 03:38:32Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/XQTWBu3KWTES2b5y9](/api/post/shortform-2/comments/XQTWBu3KWTES2b5y9) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/jFxPCYBbtvPjdLSoH](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/jFxPCYBbtvPjdLSoH) * Markdown permalink: [/api/post/shortform-2/comments/jFxPCYBbtvPjdLSoH](/api/post/shortform-2/comments/jFxPCYBbtvPjdLSoH) I think mainly people just don't like probabilities they disagree with. On media interviews, people are deferring to perceived expertise, but openai employees and twitter e/accs criticize the methodology because they think they know more than someone whose conditional or unconditional p(doom) is >10%. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-12 02:49:08Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/2QmRWZhTGCoMuTZrn](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/2QmRWZhTGCoMuTZrn) * Markdown permalink: [/api/post/shortform-2/comments/2QmRWZhTGCoMuTZrn](/api/post/shortform-2/comments/2QmRWZhTGCoMuTZrn) It’s kinda embarrassing that, if we’re asked how ASI would kill literally everyone, we talk about this how this is an unfair question or smth smth vingean uncertainty smth smth Magnus Carlson. I also don’t think “engineered pandemics” will stand up to scrutiny. Here’s my list, ranked from most to least likely: 1. Robots that humans gave them 2. Robots they built themselves 3. Nanotech 4. Mirror life 5. Super persuasion 6. Hacking into nukes/critical infrastructure 7. Weird physics / magic 8. Blackmailing/hiring humans to kill other humans Note that the actual route might involve a mix of these, e.g. *blackmail someone into giving you compute, hack into a robot factory, capture enough robots to build mirror life, kill enough humans to disempower governments, build robots to build robots, industrial explosion, more foom, nanotech.* So this list is more like “what’s the primary vector for how things initially went unrecoverable”. I also don’t find credible “AIs unintentionally kill humans because they make the world uninhabitable as a biproduct of industrial explosion”. If AI kills all the humans it will be because they didn’t want to risk us turning them off or building a rival RSI. ### Comment by [Seth Herd](/users/seth-herd) * 2026-09-10 12:01:27Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/pd8WwTLP5eWJaoeY8](/api/post/shortform-2/comments/pd8WwTLP5eWJaoeY8) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/eifKfL589hhBGxFhz](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/eifKfL589hhBGxFhz) * Markdown permalink: [/api/post/shortform-2/comments/eifKfL589hhBGxFhz](/api/post/shortform-2/comments/eifKfL589hhBGxFhz) I'm not sure recent events will convert to much political will, but it looks like a good start. With a little real job loss, I expect we'll get there rapidly. I've been a bit puzzled by how convinced everyone was that we wouldn't see meaningful change in public opinions. I guess the question has been whether we get it soon enough to matter. But people, particularly Americans, love freaking out about things, and superhuman AI is an extraordinarily good thing to freak out about. And the arguments aren't really complicated at all. I tried to capture this in [A country of alien idiots in a datacenter: AI progress and public alarm](/api/post/qxmAqMAjxnhkzt6aF) but I'm pleased to see more pickup in public alarm before agents and job loss are as widespread as in my scenario. ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-09-07 13:42:55Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/n5tz4ckkEHptBqfJq](/api/post/shortform-2/comments/n5tz4ckkEHptBqfJq) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/rJCaqcigCcavjnbCF](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/rJCaqcigCcavjnbCF) * Markdown permalink: [/api/post/shortform-2/comments/rJCaqcigCcavjnbCF](/api/post/shortform-2/comments/rJCaqcigCcavjnbCF) Reactions by quoted text: * "Fable 5.1" * confused: 1 I spent an hour or so playing around with Fable 5.1. TLDR: We have estimates for Sol depth, and Astra depth (Pachocki says 2× GPT-4, which has known depth). We also know how much ordinary scaling contributes to no-CoT. We can fit a model of how no-CoT changes with depth and active params. Astra’s depth step, run through that model, gives 2.0 doublings of no-CoT; the ordinary generation gives ~1.2; Astra’s observed jump over Sol is 3.1. The two add up, and 2.0/3.1 ≈ two-thirds is the loops’ share. The model --------- * The Think Fast paper measured no-CoT time horizons for 35 open-weight models. I refit jointly on the layer counts, active parameters and total parameters in its Table 17. In doublings of time horizon: ΔD = β\_L · log₂(layers ratio) + β\_T · log₂(total params ratio) + β_A · log₂(active params ratio) + ε * What the equation says, in words: * Take two models. The left side, ΔD, is how many times the no-CoT time horizon doubles going from the first model to the second. Sol to Astra is 3.6 → 30.9 minutes, which is 3.1 doublings. * Each term on the right is one architectural difference between the two models and how much of ΔD it accounts for. "Layers ratio" is the second model's layer count divided by the first's; taking log₂ turns "3× deeper" into "1.58 doublings of depth". β\_L is then the exchange rate: how many doublings of time horizon you get per doubling of depth. Same for total parameters (β\_T) and active parameters (β_A). * ε is everything that isn't architecture: more data, more RL, better algorithms. The fit can't see these, so they show up as residual. * The equation is additive in log space, which is the same as saying the effects multiply in minutes: doubling depth and doubling knowledge each multiply the time horizon by their own factor, independently. * Fitted coefficients: * **β_L = 1.38**. One doubling of time horizon per **1.65× layers** (95% CI 1.36–2.95×). I'll call the layer multiplier per doubling k, so k = 1.65. * **β_T = 0.25** (CI 0.11–0.38). Doubling total parameters, mostly expert count, buys a quarter of a doubling. * **β_A ≈ 0** (−0.12, CI −0.42 to 0.16). Given depth and total parameters, more active parameters buy nothing. * What that says for Astra: * Depth has a large effect holding parameters fixed. * Width has none. Given a model's depth, making it wider doesn't improve no-CoT performance. * Looping adds depth and nothing else: the same block, run again, with no new parameters. So a looped model should get the full β\_L effect and nothing from β\_A. * Why a joint fit, and why 1.65× layers per doubling rather than the paper's headline 1.3×: * The paper's 1.3× is a one-variable fit of time horizon on layers. In the open-weight population, deeper models are also wider and hold more parameters, because that's how labs scale, so that fit credits depth with whatever the co-scaled width and knowledge contributed. It's steeper than the true depth effect. * Fitting all three variables at once asks a different question: holding parameters fixed, what does depth alone buy? That's the question looping poses, since loops add depth and nothing else. The answer is 1.65×. * A check that the split is right: the controlled coefficients reconstruct the uncontrolled fit. Across 66 open-weight MoE models ([Raschka's architecture gallery](https://github.com/rasbt/llm-architecture-gallery)), depth scales as active params^0.23 and total params scale as active^1.2, so 1.3× layers goes with ~3.1× active and ~3.9× total. β\_L gives 0.52 doublings from the depth, β\_T gives 0.49 from the total params, β_A gives ~0, and the sum is one doubling, which is what the one-variable fit says 1.3× layers is worth. * One caveat: k is fit on Think Fast's aggregate benchmarks. Those include recall tasks, which are probably driven by parameters, and serial-computation tasks like math, which are more depth-bound. Competition math is in the second group, so k for math is probably lower than 1.65, which would make the loop share somewhat higher than I state. I doubt it's a large effect, but I haven't measured it. The inputs ---------- Everything here was public before this week, except Pachocki's tweet. 1. **GPT-4 has ~120 layers.** 2023 SemiAnalysis leak ([summary](https://patmcguinness.substack.com/p/gpt-4-details-revealed)). Unconfirmed. 2. **Sol has ~80 layers.** The one public estimate, [inferred from serving hardware](https://x.com/TokenGremlin/status/2086685446050857344): 70–100 wafers at about one layer each. Consistent with the open-weight frontier, where no MoE released since GPT-4 exceeds 110 layers (Figure 3). 3. **Astra has ≤ 240 effective layers.** Pachocki's factor of two, times (1). Over Sol, ≤ 3×. I take the ceiling. 4. **Sol's no-CoT math time horizon is 3.6 minutes.** AISI, in the system card. 5. **Depth under ordinary scaling goes as active params^0.23** (gallery, 66 models, CI 0.18–0.29). This is the α ≈ 0.25 behind Ryan Greenblatt's figure that 3× depth would ordinarily cost ~81× parameters. I use it to separate the depth a normal generation would have added from the depth the loops added. 6. **Astra's active parameters are ~1.5× Sol's.** A guess: one generation, on OpenAI's recent trend toward sparser models. Total parameters ~1.6× on the same outside estimate as (2). 7. **Recent OpenAI generation steps.** * AISI no-CoT math: GPT-5.1 0.60 min → 5.2 1.20 → 5.5 2.30 → Sol 3.60. * Epoch Capabilities Index (a with-CoT scale over 50+ benchmarks): GPT-5.2 153 → 5.5 159 → Sol 162 → Astra 169. The prediction -------------- Two parts: what an ordinary successor to Sol would have scored, and what the loops add on top. ### Part 1: an ordinary successor to Sol Suppose OpenAI had shipped a normal next model after Sol, with the usual increase in parameters, data and training, and Sol's architecture. What would its no-CoT math time horizon be? Two ways to estimate it. * Method A: previous generation steps on the same eval. * The last three OpenAI steps in AISI's series, none of which changed architecture, added 1.0, 0.94 and 0.65 doublings (5.1 → 5.2 → 5.5 → Sol). * The steps aren't all alike. 5.2 → 5.5 included a new pretraining run (GPT-5.5, codename Spud); 5.5 → Sol appears to have been post-training on the same base and was the smallest step. * Astra is reported, in leaks rather than by OpenAI, to sit on a new and larger pretrain (codename Doug, said to be roughly 2.5× Spud's total parameter count), though some reports treat Doug as a separate later model. * If Astra is a new pretrain, the relevant comparison is the new-pretrain steps: **about 1.0 doubling**. Range 0.65–1.0. * Method B: Epoch's Capabilities Index. * ECI is a with-CoT capability scale built from benchmarks other than AISI's. Four OpenAI reasoning-era models have both an ECI and an AISI no-CoT math horizon, and none changed architecture: | **model** | **ECI** | **no-CoT math TH** | | --- | --- | --- | | GPT-5.1 | 150 | 0.60 min | | GPT-5.2 | 153 | 1.20 | | GPT-5.5 | 159 | 2.30 | | GPT-5.6 Sol | 162 | 3.60 | * Regressing log₂ time horizon on ECI: **0.20 doublings per ECI point** (SE 0.02, R² 0.98). Five ECI points per no-CoT doubling, the same rate Epoch reported between ECI and METR's with-CoT horizon. * Astra scores 169, +7 over Sol, which Epoch describes as within the uncertainty of the 14-point-per-year frontier trend. * At 0.20 per point, an ordinary model at Astra's ECI would sit **1.4 doublings** above Sol (1.1–1.7 at ±2 SE), about 10 minutes. * Astra's ECI presumably already includes some benefit from its depth, so this is an upper bound on the ordinary term. * The relationship is OpenAI-specific. Open-weight models in the same AISI series show no ECI → no-CoT relationship at all (slope 0.01, R² 0.13); their with-CoT gains haven't carried no-CoT gains. That's why I don't pool them. * The two methods give about 1.0 and 1.4. **I take 1.2 doublings: an ordinary successor to Sol would score about 8 minutes.** ### Part 2: what the loops add * That ordinary successor would already be about 1.1× deeper than Sol, because 1.5× active parameters carries 1.5^0.23 of depth on the recipe (inputs 5 and 6). * Astra is 3× deeper than Sol, so the loops supply 3 / 1.1 ≈ **2.7×** on top of the ordinary successor. * At β_L = 1.38, that's log₂(2.7) × 1.38 = **2.0 doublings**. Total * 1.2 + 2.0 = **3.2 doublings above Sol**. * 3.6 × 2^3.2 ≈ **33 minutes**. What we observe * **30.9 minutes**, 95% CI 22.8–46.8. That's 3.1 doublings above Sol. * This could be a coincidence. But the two sides are independent of each other and of the observation. Neither used Astra's number, and nothing forced them to sum to anything in particular. That they land on the measured value is, I think, suggestive that the picture is right: Astra's no-CoT jump is what a ~2.7× depth increase from looping, plus one ordinary generation, produces. ### Comment by [Mo Putera](/users/mo-putera) * 2026-08-23 06:19:34Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/oix4z6eHDHT75yr93](/api/post/shortform-2/comments/oix4z6eHDHT75yr93) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/hHNuAxzeEdzHZzbWS](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/hHNuAxzeEdzHZzbWS) * Markdown permalink: [/api/post/shortform-2/comments/hHNuAxzeEdzHZzbWS](/api/post/shortform-2/comments/hHNuAxzeEdzHZzbWS) Reactions by quoted text: * " was confused why this was ever a thing" * agree: 1 I was confused why this was ever a thing. I just assumed everyone had seen this chart and noticed how even if the blue line plateaus the green and red need not, especially given tremendous persistent economic incentives. Maybe the counterargument is "obviously the proposers knew this, what they actually proposed were thresholds that would ratchet downward over time to adjust"? ![image.png](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1787465766/lexical_client_uploads/dnj12lfiapeaf7g0z0ob.png) ### Comment by [evhub](/users/evhub) * 2026-08-21 21:28:53Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 1 * Parent comment (Markdown): [/api/post/shortform-2/comments/C9NynsKyM6DifdqA6](/api/post/shortform-2/comments/C9NynsKyM6DifdqA6) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/G8svttF3Fpr3iCRkq](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/G8svttF3Fpr3iCRkq) * Markdown permalink: [/api/post/shortform-2/comments/G8svttF3Fpr3iCRkq](/api/post/shortform-2/comments/G8svttF3Fpr3iCRkq) I would not say that [market making](/api/post/YWwzccGbcHMJMpT45) has been abandoned, at least no more so than any other sophisticated scalable oversight approach like amplification or debate; the state of scalable oversight is just such that it's hard to get any of these sorts of techniques working. See also my general discussion of how I think about scalable oversight [here](/api/post/HE3Styo9vpk7m8zi4?commentId=pwhksLpsghbcm4taQ). ### Comment by [Cleo Nardo](/users/cleo-nardo) * 2026-08-02 13:13:40Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Reply depth: 3 * Parent comment (Markdown): [/api/post/shortform-2/comments/MDMgKeBYZfvc2rgh5](/api/post/shortform-2/comments/MDMgKeBYZfvc2rgh5) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/zdQ3gjbhYyEBGFAsy](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/zdQ3gjbhYyEBGFAsy) * Markdown permalink: [/api/post/shortform-2/comments/zdQ3gjbhYyEBGFAsy](/api/post/shortform-2/comments/zdQ3gjbhYyEBGFAsy) “Replaceable, wishes he wasn’t” vs “irreplaceable, wishes he wasn’t” 🙇‍♂️ ### Comment by [gwern](/users/gwern) * 2026-07-20 23:16:54Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 3 * Total votes: 3 * Reply depth: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/urpC5jaJcLjiwwyW9](/api/post/shortform-2/comments/urpC5jaJcLjiwwyW9) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/vaYwCHjgzwBLfNNDR](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/vaYwCHjgzwBLfNNDR) * Markdown permalink: [/api/post/shortform-2/comments/vaYwCHjgzwBLfNNDR](/api/post/shortform-2/comments/vaYwCHjgzwBLfNNDR) Maybe source from https://arxiv.org/abs/1202.3936 https://gwern.net/doc/math/2013-hisano.pdf as pre-AI sets of conjectures to monitor or target? (I have many errors listed in my [math error essay](https://gwern.net/math-error) but not sure how useful the ad hoc set is compared to the Hisano & Sornette work.) ### Comment by [jacquesthibs](/users/jacques-thibodeau) * 2026-07-20 19:02:43Z * Karma: 4 * Voting system: namesAttachedReactions * Approval votes: 2 * Total votes: 2 * Parent comment (Markdown): [/api/post/shortform-2/comments/TKWzHEgovLjFLqdGt](/api/post/shortform-2/comments/TKWzHEgovLjFLqdGt) * HTML permalink: [/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/aSrNjwczyHytX2vB7](/posts/i7JSL5awGFcSRhyGF/shortform-2/comment/aSrNjwczyHytX2vB7) * Markdown permalink: [/api/post/shortform-2/comments/aSrNjwczyHytX2vB7](/api/post/shortform-2/comments/aSrNjwczyHytX2vB7) It’s a bit difficult to explain quickly, but [some of my thoughts on the matter are here](/api/post/jXjeYYPXipAtA2zmj?commentId=5G3eA76RJwPPjEGe7). And here: [https://www.lesswrong.com/posts/jXjeYYPXipAtA2zmj/jacquesthibs-s-shortform?commentId=pCiAsJ7NLXrtBymw2](/api/post/jXjeYYPXipAtA2zmj?commentId=pCiAsJ7NLXrtBymw2) ### Navigation * [Front page](https://www.lesswrong.com/api/home) * [Markdown API documentation](https://www.lesswrong.com/api/SKILL.md)