# User: Wei Dai
Profile URL (HTML): [/users/wei-dai](/users/wei-dai)
Profile URL (Markdown): [/api/user/wei-dai](/api/user/wei-dai)
* Karma: 48365
* Alignment Forum karma: 3581
* Posts: 153
* Comments: 5530
* Member since: 2009-03-06 19:59:52Z
Bio
---
If anyone wants to have a voice chat with me about a topic that I'm interested in (see my recent post/comment history to get a sense), please contact me via PM.
My main "claims to fame":
* Created the first general purpose open source cryptography programming library ([Crypto++](https://www.cryptopp.com/), 1995), motivated by AI risk and what's now called "defensive acceleration".
* Published one of the first descriptions of a cryptocurrency based on a distributed public ledger ([b-money](http://www.weidai.com/bmoney.txt), 1998), predating Bitcoin.
* Proposed [UDT](https://wiki.lesswrong.com/wiki/Updateless_decision_theory), combining the ideas of updatelessness, policy selection, and evaluating consequences using logical conditionals.
* First to argue for pausing AI development based on the technical difficulty of ensuring AI x-safety ([SL4 2004](http://sl4.org/archive/0410/10098.html), [LW 2011](https://www.lesswrong.com/posts/73SotZnDbsYpxfnuQ/some-thoughts-on-singularity-strategies)).
* [Identified](https://www.lesswrong.com/posts/w6d7XBCegc96kz4n3/the-argument-from-philosophical-difficulty) current and future philosophical difficulties as core AI x-safety bottlenecks, potentially insurmountable by human researchers, and [advocated](https://www.lesswrong.com/posts/vrnhfGuYTww3fKhAM/three-approaches-to-friendliness) for research into [metaphilosophy](https://www.lesswrong.com/w/meta-philosophy) and AI philosophical competence as possible solutions.
[My Home Page](http://www.weidai.com)
Top Posts
---------
### [Legible vs. Illegible AI Safety Problems](/api/post/legible-vs-illegible-ai-safety-problems)
By [Wei Dai](/users/wei-dai)
2025-11-04 21:39:07Z
* Karma: 413
* Curated
* Tags: [AI Risk](/w/ai-risk), [AI](/w/ai) (Frontpage)
Read more: [/api/post/legible-vs-illegible-ai-safety-problems](/api/post/legible-vs-illegible-ai-safety-problems)
### [A tale from Communist China](/api/post/a-tale-from-communist-china)
By [Wei Dai](/users/wei-dai)
2020-10-18 17:37:42Z
* Karma: 303
* Tags: [History](/w/history), [Censorship](/w/censorship), [China](/w/china), [Politics](/w/politics) (Personal Blog)
Read more: [/api/post/a-tale-from-communist-china](/api/post/a-tale-from-communist-china)
### [Morality is Scary](/api/post/morality-is-scary)
By [Wei Dai](/users/wei-dai)
2021-12-02 06:35:06Z
* Karma: 291
* Tags: [Ethics & Morality](/w/ethics-and-morality), [Human-AI Safety](/w/human-ai-safety), [AI](/w/ai) (Frontpage)
Read more: [/api/post/morality-is-scary](/api/post/morality-is-scary)
Recent Posts
------------
### [The Long (Self-)Correction](/api/post/the-long-self-correction-2)
By [Wei Dai](/users/wei-dai)
2026-07-24 21:01:47Z
* Karma: 267
* Curated
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/the-long-self-correction-2](/api/post/the-long-self-correction-2)
### [Increasing AI Strategic Competence as a Safety Approach](/api/post/increasing-ai-strategic-competence-as-a-safety-approach)
By [Wei Dai](/users/wei-dai)
2026-02-03 01:08:13Z
* Karma: 57
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/increasing-ai-strategic-competence-as-a-safety-approach](/api/post/increasing-ai-strategic-competence-as-a-safety-approach)
### [My 2003 Post on the Evolutionary Argument for AI Misalignment](/api/post/my-2003-post-on-the-evolutionary-argument-for-ai)
By [Wei Dai](/users/wei-dai)
2026-01-06 20:45:20Z
* Karma: 37
* Tags: [Evolution](/w/evolution), [AI](/w/ai) (Frontpage)
Read more: [/api/post/my-2003-post-on-the-evolutionary-argument-for-ai](/api/post/my-2003-post-on-the-evolutionary-argument-for-ai)
### [A Conflict Between AI Alignment and Philosophical Competence](/api/post/a-conflict-between-ai-alignment-and-philosophical-competence)
By [Wei Dai](/users/wei-dai)
2025-12-27 21:32:07Z
* Karma: 78
* Tags: [AI-Assisted Alignment](/w/ai-assisted-alignment), [Corrigibility](/w/corrigibility-1), [Meta-Philosophy](/w/meta-philosophy), [Metaethics](/w/metaethics-1), [Philosophy](/w/philosophy-1), [AI](/w/ai) (Frontpage)
Read more: [/api/post/a-conflict-between-ai-alignment-and-philosophical-competence](/api/post/a-conflict-between-ai-alignment-and-philosophical-competence)
### [Relitigating the Race to Build Friendly AI](/api/post/relitigating-the-race-to-build-friendly-ai)
By [Wei Dai](/users/wei-dai)
2025-12-03 11:34:13Z
* Karma: 109
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/relitigating-the-race-to-build-friendly-ai](/api/post/relitigating-the-race-to-build-friendly-ai)
### [Please, Don't Roll Your Own Metaethics](/api/post/please-don-t-roll-your-own-metaethics)
By [Wei Dai](/users/wei-dai)
2025-11-12 22:17:04Z
* Karma: 172
* Tags: [Meta-Philosophy](/w/meta-philosophy), [Metaethics](/w/metaethics-1), [Philosophy](/w/philosophy-1), [World Optimization](/w/world-optimization) (Frontpage)
Read more: [/api/post/please-don-t-roll-your-own-metaethics](/api/post/please-don-t-roll-your-own-metaethics)
### [Problems I've Tried to Legibilize](/api/post/problems-i-ve-tried-to-legibilize)
By [Wei Dai](/users/wei-dai)
2025-11-09 10:27:22Z
* Karma: 151
* Tags: [AI Risk](/w/ai-risk), [AI](/w/ai) (Frontpage)
Read more: [/api/post/problems-i-ve-tried-to-legibilize](/api/post/problems-i-ve-tried-to-legibilize)
### [Legible vs. Illegible AI Safety Problems](/api/post/legible-vs-illegible-ai-safety-problems)
By [Wei Dai](/users/wei-dai)
2025-11-04 21:39:07Z
* Karma: 413
* Curated
* Tags: [AI Risk](/w/ai-risk), [AI](/w/ai) (Frontpage)
Read more: [/api/post/legible-vs-illegible-ai-safety-problems](/api/post/legible-vs-illegible-ai-safety-problems)
### [Trying to understand my own cognitive edge](/api/post/trying-to-understand-my-own-cognitive-edge)
By [Wei Dai](/users/wei-dai)
2025-11-03 08:49:29Z
* Karma: 74
* Tags: [Intellectual Progress (Individual-Level)](/w/intellectual-progress-individual-level), [Rationality](/w/rationality) (Frontpage)
Read more: [/api/post/trying-to-understand-my-own-cognitive-edge](/api/post/trying-to-understand-my-own-cognitive-edge)
### [Wei Dai's Shortform](/api/post/wei-dai-s-shortform)
By [Wei Dai](/users/wei-dai)
2024-03-01 20:43:15Z
* Karma: 10
* Tags: None (Personal Blog)
Read more: [/api/post/wei-dai-s-shortform](/api/post/wei-dai-s-shortform)
Recent Comments
---------------
### Comment by [Wei Dai](/users/wei-dai) on [PauseAI Has 'officially disendorsed' PauseAI-US](/api/post/pauseai-has-officially-disendorsed-pauseai-us)
* 2026-09-11 22:56:16Z
* Karma: 2
* Total votes: 1
* Comment URL (Markdown): [/api/post/pauseai-has-officially-disendorsed-pauseai-us/comments/XCQpFqqkH2tgcZvL7](/api/post/pauseai-has-officially-disendorsed-pauseai-us/comments/XCQpFqqkH2tgcZvL7)
* Comment URL (HTML): [/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/XCQpFqqkH2tgcZvL7](/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/XCQpFqqkH2tgcZvL7)
Your notion of "morally bad actor" does not seem to require/imply that the person "know what they're doing is bad", because you could convert someone who doesn't know this into being a friend, by convincing them that what they're doing is bad, either through arguments or by applying social pressure. This seems to be what Liron and Holly are in part trying to do?
(I can also imagine Liron/Holly using the term in a different way, which also seems reasonable, more in line with my example of a torturer AI, which is that this agent is doing something really bad according to our morality, and so we should mobilize/coordinate to do something about it.)
### Comment by [Wei Dai](/users/wei-dai) on [PauseAI Has 'officially disendorsed' PauseAI-US](/api/post/pauseai-has-officially-disendorsed-pauseai-us)
* 2026-09-11 17:57:09Z
* Karma: 1
* Total votes: 2
* Comment URL (Markdown): [/api/post/pauseai-has-officially-disendorsed-pauseai-us/comments/CwtmYn94SDonHeeZD](/api/post/pauseai-has-officially-disendorsed-pauseai-us/comments/CwtmYn94SDonHeeZD)
* Comment URL (HTML): [/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/CwtmYn94SDonHeeZD](/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us/comment/CwtmYn94SDonHeeZD)
>I think this doesn’t actually follow. It does matter whether they know what they’re doing is bad. If they know and continue, they’re a morally bad actor.
(Assuming you mean this as a biconditional "if and only if".) By this standard, if a future AI decides to torture a bunch of humans in the course of taking over the world (and this isn't bad according to their values/decision theory/philosophy), they're not a morally bad actor. Do you endorse this?
### Comment by [Wei Dai](/users/wei-dai) on [Personal statement on joining the OpenAI nonprofit board](/api/post/personal-statement-on-joining-the-openai-nonprofit-board)
* 2026-09-11 01:47:09Z
* Karma: 23
* Total votes: 10
* Comment URL (Markdown): [/api/post/personal-statement-on-joining-the-openai-nonprofit-board/comments/AJRZPQWDEcCYhxL3u](/api/post/personal-statement-on-joining-the-openai-nonprofit-board/comments/AJRZPQWDEcCYhxL3u)
* Comment URL (HTML): [/posts/82z6FvbYRdjYjqigK/personal-statement-on-joining-the-openai-nonprofit-board/comment/AJRZPQWDEcCYhxL3u](/posts/82z6FvbYRdjYjqigK/personal-statement-on-joining-the-openai-nonprofit-board/comment/AJRZPQWDEcCYhxL3u)
> In general, it doesn't seem that bad policy to me that some thoughtful people should try to specialize in joining powerful institutions and talk sense to them at the cost of their public voice, while others should remain fully independent and try to become honest and unbiased public thought-leaders.
Isn't it crazy that our world makes these choices mutually exclusive, and on a meta level, everyone just takes it in stride?
BTW who are some remaining public thought-leaders with views similar to Paul's? I often wonder "what does Paul think about this development or idea?" and I'm not sure whose posts/comments to look up for the closest substitute. I guess Geoffrey Irving comes to mind but he doesn't post/comment nearly as much as Paul did.
Here's a list made by Perplexity[^hv5s00d96yp], but none really fit a combo of prolific poster/commenter, independent voice, and focus on scalable pragmatic alignment, that Paul represented before joining USG:
* **Rohin Shah** for the most regular technical commentary and best coverage of the whole practical-alignment portfolio.
* **Jan Leike** for the closest direct continuation of scalable oversight and automated alignment.
* **Geoffrey Irving** for debate and formal scalable-oversight protocols.
* **Evan Hubinger** for deception/inner-alignment pressure-testing.
* **Buck Shlegeris** for control, monitoring, and the “what if the model is actively trying to bypass oversight?” perspective.
* **Beth Barnes / METR** for empirical evaluations and evidence standards.
[^hv5s00d96yp]: prompt: who are some people with views similiar to Paul Christiano, especially on the need for scalable alignment? who among them posts/comments the most?
### Comment by [Wei Dai](/users/wei-dai) on [Wei Dai's Shortform](/api/post/wei-dai-s-shortform)
* 2026-08-27 22:57:14Z
* Karma: 21
* Total votes: 7
* Comment URL (Markdown): [/api/post/wei-dai-s-shortform/comments/q2o6a5qJ93kex7vwr](/api/post/wei-dai-s-shortform/comments/q2o6a5qJ93kex7vwr)
* Comment URL (HTML): [/posts/HbkNAyAoa4gCnuzwa/wei-dai-s-shortform/comment/q2o6a5qJ93kex7vwr](/posts/HbkNAyAoa4gCnuzwa/wei-dai-s-shortform/comment/q2o6a5qJ93kex7vwr)
Some points (cleaned up by LLM) I made during an out-of-band chat with [@Raemon](/api/user/raemon?mention=user) (all of which, IIUC, either didn't update him much or are not cruxy for him):[^4h5ve0zvxgw]
**Criticism is rare** For example, UDT didn't get much criticism — see [UDT shows that decision theory is more puzzling than ever](/api/post/wXbSAKu2AcohaK2Gt), where I had to find a lot of the issues with it myself. (Daniel wrote up the commitment race problem, but I had found and talked about it previously.) So, like some other LW commenters, I don't think hobby-horsing is bad: someone making even a not-that-central critique is at least better than nothing, if it corrects some error in the post — like I think my comment to Richard did. (Reviews don't fix this: there are very few of them compared to comments, and writing an overall assessment is a different skill from bug hunting — I've never managed to write a single annual review myself.)
**Critics can't reliably tell which of their critiques are "central."** Part of my implicit model: critics typically won't share the same understanding of an author's idea as the author himself — either they haven't put in as much effort to understand it, or the idea is actually nonsense but the author is too biased to see it. So their sense of what critiques are central is likely to be pretty different from the author's. Asking critics to avoid non-central critiques, explicitly or implicitly via the threat of deletion, makes people less likely to make critiques at all, because they don't want to risk doing something considered wrong / low-status according to local culture (or just waste their effort, or not have fun). The only solution I can see is to encourage people to make all the critiques they can think of, and make the author resilient enough to just ignore the non-central ones, or say something like: "Really appreciate your effort in making a critique. Unfortunately, according to my understanding of my idea (which I do not expect you or others to fully share), I don't think this is central, so I'm not going to address it, at least for now. Feel free to elaborate more if you do strongly suspect it is central, or I may come back to it myself in the future."
**I've had this concern since 2010.** See [Tips and Tricks for Answering Hard Questions](/api/post/SEq8bvSXrzF4jcdS8): "Don't stop at the first good answer." That was the LW1 era. LW2 seems to be going in the opposite direction (not exactly, but at least in an orthogonal direction that takes us further from the destination I envision).
**To clarify my slippery slope concern.** I've been told the target is "authors feel free to moderate their posts the way they would their Facebook wall, but people are encouraged to write up responses elsewhere and those responses show up in Mentioned In." To clarify: that *is* the end of the slippery slope I'm worried about, not something I'm worried about slipping past. I do also worry it could slip a bit further, in the sense of people finding new reasons to ban people — like the reason cited for banning Said, and the hobby horse idea — and being more willing to apply existing such ideas, so banning becomes more frequent (or people become more discouraged to criticize even if the frequency of banning stays the same), even as the hard rules stay constant and site mods don't try to push the culture more.
**A concrete compromise proposal.** A modified version of the compromise from [this comment](/api/post/HbkNAyAoa4gCnuzwa?commentId=a984jNtz8dAsA4hji): create a new type of content similar to shortform but in a different place (so they don't clutter up shortforms) — call them "off-comments" for now. Have "Mentioned In" include these as links, with the "Mentioned In" section fully expanded by default. There's also a collapsed-by-default section at the bottom that inlines all of these off-comments (and their reply threads); when you click a link in "Mentioned In" it expands that collapsed section and navigates to the inlined off-comment. Authors can't ban people from making off-comments (at least without site mod approval), and when they delete a comment the commenter has a button/option to convert that comment into an off-comment.
[^4h5ve0zvxgw]: Writing them down for public record, not inviting or expecting more engagement from the mods.
### Comment by [Wei Dai](/users/wei-dai) on [What just happened? Pragmatism and Pessimization](/api/post/what-just-happened-pragmatism-and-pessimization)
* 2026-08-25 08:59:19Z
* Karma: 15
* Total votes: 8
* Comment URL (Markdown): [/api/post/what-just-happened-pragmatism-and-pessimization/comments/ybHyzGpBDrn98JXQm](/api/post/what-just-happened-pragmatism-and-pessimization/comments/ybHyzGpBDrn98JXQm)
* Comment URL (HTML): [/posts/yaz8nx4ogZmiqHzt7/what-just-happened-pragmatism-and-pessimization/comment/ybHyzGpBDrn98JXQm](/posts/yaz8nx4ogZmiqHzt7/what-just-happened-pragmatism-and-pessimization/comment/ybHyzGpBDrn98JXQm)
> i don't really think of situational awareness as "ai safety people".
Carl Shulman is its Research Director (EDIT: and co-portfolio manager), and used to be a MIRI employee (2010-2013).
+++ prompt: carl shulman's publications on AI safety
%%% llm-output model="Perplexity / GPT 5.6 Terra"
Carl Shulman’s AI-safety work is concentrated in early foundational writing on superintelligent-agent alignment, instrumental convergence, intelligence-explosion dynamics, and governance rather than contemporary empirical alignment research. His publication record also includes adjacent work on digital minds, forecasting, and long-run governance.\[[semanticscholar](https://www.semanticscholar.org/author/Carl-Shulman/3389522)\]\[[alignmentforum](/api/tag/carl-shulman)\]
Core AI-safety publications
---------------------------
| Publication | Year | Coauthors | Main contribution |
| --- | --- | --- | --- |
| *Machine Ethics and Superintelligence* | 2009 | Henrik Jonsson, Nick Tarleton | Argues that ordinary machine-ethics approaches make assumptions—incremental deployment, human-comparable capabilities, human institutional embedding—that may fail for agents at or beyond human level. It calls for advance analysis of the harder alignment problem. \\[[intelligence](https://intelligence.org/files/MachineEthicsSuperintelligence.pdf)\\] |
| *Arms Control and Intelligence Explosions* | 2009 | Stuart Armstrong | An early AI-governance paper: analyzes incentives for unsafe competitive development, winner-take-all dynamics, and the possibility that advanced AI capabilities could also make verification and international agreements more feasible. \\[[intelligence](https://intelligence.org/files/ArmsControl.pdf)\\] |
| *Which Consequentialism? Machine Ethics and Moral Divergence* | 2009 | Henrik Jonsson, Nick Tarleton | Discusses divergence among consequentialist objectives and its implications for constructing moral machines—relevant to the problem of specifying AI goals. It appears in his scholarly bibliography, though it is less directly focused on catastrophic-risk mechanisms than the three items above. \\[[semanticscholar](https://www.semanticscholar.org/author/Carl-Shulman/3389522)\\] |
| *Omohundro’s “Basic AI Drives” and Catastrophic Risks* | 2010 | — | Critically examines the argument that a broad class of goal-directed systems could develop instrumental tendencies such as self-preservation and resource acquisition, and relates those tendencies to catastrophic AI-risk scenarios. \\[[intelligence](https://intelligence.org/files/BasicAIDrives.pdf)\\] |
| *Implications of a Software-Limited Singularity* | 2010/2013 version | Anders Sandberg | Explores implications of a rapid AI transition constrained primarily by software development rather than hardware or physical replication. It is important background for takeoff-speed and strategic-risk debates. \\[[semanticscholar](https://www.semanticscholar.org/author/Carl-Shulman/3389522)\\] |
| *How Hard Is Artificial Intelligence? Evolutionary Arguments and Selection Effects* | 2012 | Nick Bostrom | Addresses evolutionary and selection-effect arguments about the difficulty of developing AI, bearing on forecasts and the plausibility/timing of transformative AI. \\[[semanticscholar](https://www.semanticscholar.org/author/Carl-Shulman/3389522)\\] |
| *Racing to the Precipice: A Model of Artificial Intelligence Development* | 2016 | Stuart Armstrong, Nick Bostrom | Models competitive AI development in which actors trade safety for speed, formalizing “race to the bottom” concerns that remain central to AI governance. \\[[dl.acm](https://dl.acm.org/doi/10.1007/s00146-015-0590-y)\\]\\[[semanticscholar](https://www.semanticscholar.org/author/Carl-Shulman/3389522)\\] |
Especially worth reading
------------------------
If your interest is specifically AI x-risk and alignment, I’d prioritize:
1. **“Omohundro’s ‘Basic AI Drives’ and Catastrophic Risks”** — best for instrumental convergence and the case that misaligned optimization can be dangerous even without malicious intent.\[[intelligence](https://intelligence.org/files/BasicAIDrives.pdf)\]
2. **“Machine Ethics and Superintelligence”** — best for the conceptual distinction between conventional AI ethics and alignment for highly capable autonomous agents.\[[intelligence](https://intelligence.org/files/MachineEthicsSuperintelligence.pdf)\]
3. **“Racing to the Precipice”** — best for strategic dynamics: why safety measures may be underprovided under competition.\[[dl.acm](https://dl.acm.org/doi/10.1007/s00146-015-0590-y)\]
4. **“Arms Control and Intelligence Explosions”** — best for governance implications and the argument that AI may create both unusual destabilization and unusual verification/enforcement opportunities.\[[intelligence](https://intelligence.org/files/ArmsControl.pdf)\]
Related, but broader than safety
--------------------------------
Two later works with Nick Bostrom—*Propositions Concerning Digital Minds and Society* and *Sharing the World with Digital Minds*—are principally about moral status, rights, and social organization involving digital minds. They matter for post-AGI governance and AI welfare, but are not technical alignment papers.\[[80000hours](https://80000hours.org/podcast/episodes/carl-shulman-economy-agi/)\]\[[experiencemachines.substack](https://experiencemachines.substack.com/p/carl-shulman-on-the-moral-status-11b)\]
For a more complete bibliography that includes papers, posts, and conversations, Shulman’s LessWrong publication timeline says formal work can be isolated by filtering its **Format** field to “Paper.”\[[lesswrong](/lw/7ob/timeline_of_carl_shulman_publications)\]
%%% /llm-output
+++
### Comment by [Wei Dai](/users/wei-dai) on [Wei Dai's Shortform](/api/post/wei-dai-s-shortform)
* 2026-08-25 08:40:40Z
* Karma: 8
* Total votes: 5
* Comment URL (Markdown): [/api/post/wei-dai-s-shortform/comments/5vExpeveTnKzur7ti](/api/post/wei-dai-s-shortform/comments/5vExpeveTnKzur7ti)
* Comment URL (HTML): [/posts/HbkNAyAoa4gCnuzwa/wei-dai-s-shortform/comment/5vExpeveTnKzur7ti](/posts/HbkNAyAoa4gCnuzwa/wei-dai-s-shortform/comment/5vExpeveTnKzur7ti)
> I think most attempts to make intellectual progress are likely to fail
I wonder if one issue with people's models of intellectual progress is that ~nobody on LW saw the ~10 years I spent exploring various dead ends / unclear results in anthropic reasoning and decision theory[^y2k11yc5ybk], they just saw UDT 1.0 pop almost out of nowhere. (Perhaps some saw the few months of decision theory discussion on LW before my post, where I got some of the final pieces of the puzzle from Eliezer and Nesov's hints.)
As for Bitcoin, I don't know how long it took Satoshi, but there were hundreds of published papers on digital cash designs of all kinds[^xi5osxot1rl], before Bitcoin came on the scene and took off, but ~everyone just saw Bitcoin and don't know the rest.
Sorry if this is being unfair/uncharitable to you[^4c75ptfc7sc], [@Richard_Ngo](/api/user/ricraz?mention=user), but I suspect this might be what's accounting for our apparent disagreements around intellectual progress.
[^y2k11yc5ybk]: I posted half-baked ideas to my own everything-list mailing list, but didn't get much attention from others, except Hal Finney who liked and wrote up UDASSA on his website, which somehow spread from there even though I wasn't all that happy with it.
[^xi5osxot1rl]: My boss during my internship at Microsoft Research asked me to work on his own design, but I let it quietly fall off my plate because it seemed unpromising, and I had my own ideas at that point.
[^4c75ptfc7sc]: Maybe you've observed similar phenomenon elsewhere and still reached different conclusions than me, or I'm completely misunderstanding where our disagreement lies.
### Comment by [Wei Dai](/users/wei-dai) on [Wei Dai's Shortform](/api/post/wei-dai-s-shortform)
* 2026-08-24 19:46:39Z
* Karma: 4
* Total votes: 2
* Comment URL (Markdown): [/api/post/wei-dai-s-shortform/comments/nimhbCK4JHrqGEqiF](/api/post/wei-dai-s-shortform/comments/nimhbCK4JHrqGEqiF)
* Comment URL (HTML): [/posts/HbkNAyAoa4gCnuzwa/wei-dai-s-shortform/comment/nimhbCK4JHrqGEqiF](/posts/HbkNAyAoa4gCnuzwa/wei-dai-s-shortform/comment/nimhbCK4JHrqGEqiF)
from [https://intelligence.org/2021/11/11/discussion-with-eliezer-yudkowsky-on-agi-interventions/](https://intelligence.org/2021/11/11/discussion-with-eliezer-yudkowsky-on-agi-interventions/) (this is the closet thing Perplexity found)
**Anonymous**
Does pushing for a lot of public fear about this kind of research, that makes all projects hard, seem hopeless?
**Eliezer Yudkowsky**
What does it buy us? 3 months of delay at the cost of a tremendous amount of goodwill? 2 years of delay? What’s that delay for, if we all die at the end? Even if we then got a technical miracle, would it end up impossible to run a project that could make use of an alignment miracle, because everybody was afraid of that project? Wouldn’t that fear tend to be channeled into “ah, yes, it must be a government project, they’re the good guys” and then the government is much more hopeless and much harder to improve upon than Deepmind?
### Comment by [Wei Dai](/users/wei-dai) on [Wei Dai's Shortform](/api/post/wei-dai-s-shortform)
* 2026-08-24 09:37:40Z
* Karma: 5
* Total votes: 3
* Comment URL (Markdown): [/api/post/wei-dai-s-shortform/comments/i3x8BhzFCppnEQL6J](/api/post/wei-dai-s-shortform/comments/i3x8BhzFCppnEQL6J)
* Comment URL (HTML): [/posts/HbkNAyAoa4gCnuzwa/wei-dai-s-shortform/comment/i3x8BhzFCppnEQL6J](/posts/HbkNAyAoa4gCnuzwa/wei-dai-s-shortform/comment/i3x8BhzFCppnEQL6J)
> There's also something that I (and, I infer, Habryka) want to protect—as implicitly expressed in my post it was something like "intellectual progress is in fact a valuable thing which we can aim for". Insofar as we're in a conflict frame with each other, you could view [Wei's original comment](/api/post/9RL9MuGZjzm4q3gKG?commentId=muNxt8kf2AwjKxT4w) as an attack on that thing.
This confuses me. Can you explain what you sense as a possible motivation on my part to attack "intellectual progress is in fact a valuable thing which we can aim for"?
My own interpretation of what the conflict is, is that everyone here wants to support intellectual progress, but have different ideas how to go about it due to being biased due to status seeking, sunken costs, etc., some of which is strong/ingrained enough to make it intractable to fix the underlying mistakes, so we can only fight it out as a conflict. (To be clear I'm not certain about this.)
> On your end, it would ideally involve acknowledging the thing that Habyrka and I are trying to protect, and helping us figure out how to protect it with as few tradeoffs as possible.
Does this still make sense given what I wrote above?
### Navigation
* [Front page](https://www.lesswrong.com/api/home)
* [Markdown API documentation](https://www.lesswrong.com/api/SKILL.md)