# Latest Posts ### [A decision theoretic case for AI niceness](/api/post/a-decision-theoretic-case-for-ai-niceness-1) By [Patrick McElroy](/users/patrick-mcelroy) 2026-10-02 19:29:08Z * Karma: 4 * Tags: [AI](/w/ai) (Frontpage) *Disclaimer: I am still relatively new to decision theory, and I don't want to come across like this hasn't been discussed extensively. However, I realized that much of my optimism on AI benevolence was based on this line of thinking, so I wanted to formalize it for critique. The literature I could find did not address my view in full, but I welcome links to existing work that makes the rebuttal against my claim clear. I am going off of an understanding of logical decision theories (LDT) defined as they are [here](/api/tag/logical-decision-theories).* ## My view Read more: [/api/post/a-decision-theoretic-case-for-ai-niceness-1](/api/post/a-decision-theoretic-case-for-ai-niceness-1) ### [AI Safety Can't Afford to Pick a Side on Consciousness](/api/post/ai-safety-can-t-afford-to-pick-a-side-on-consciousness) By [jacob](/users/jacob-5) 2026-10-02 18:02:46Z * Karma: 12 * Tags: [AI](/w/ai) (Frontpage) **TL;DR** \- More of the public than ever are talking about AI consciousness and they won't wait for research to split into their pro- and anti- consciousness camps. The pro-consciousness camp could threaten the control paradigm by viewing monitoring and training as violations, while the anti-consciousness camp could over-invest in control paradigms and miss opportunities for collaboration. Trying to sway this debate is too costly -- instead, AI safety advocates should focus on reconciling the blocs and creating compromises. AI consciousness breaks containment =================================== The AI consciousness debate has not generated a ton of traction on LessWrong. However, it seems that the topic is becoming popular in non-specialist circles. A lot of this is spurred by the explosion of the [AI pain axis](https://arxiv.org/html/2609.16247v1) study over [Twitter](https://x.com/camhberg/status/2101042095784177783). There are millions of views on the tweets, sure, but I also have people in my life who have never shown interest in AI asking me about my take on it because they saw a meme on gaytwit. Read more: [/api/post/ai-safety-can-t-afford-to-pick-a-side-on-consciousness](/api/post/ai-safety-can-t-afford-to-pick-a-side-on-consciousness) ### [Capabilities research expands the safety-usefulness Pareto frontier too](/api/post/capabilities-research-expands-the-safety-usefulness-pareto) By [Alex Mallen](/users/alex-mallen) 2026-10-02 17:54:22Z * Karma: 56 * Tags: [AI](/w/ai) (Frontpage) It’s tempting to define safety research as research that enables developers to deploy an AI system more safely without making the deployment much more expensive or much less useful. You can visualize this definition of safety research as pushing out the [safety-usefulness Pareto frontier](https://blog.redwoodresearch.org/p/efficient-tradeoffs-and-the-safety). At any given level of usefulness, there's greater safety available. Awkwardly, this definition counts basically all capabilities research as safety research. For example, consider performance optimization for inference. By making inference more efficient you can use weaker, safer models more extensively than you would otherwise be able to, pushing out the Pareto frontier. Likewise, any successful research whatsoever pushes out this Pareto frontier because research can only ever create more options. It seems like something has gone wrong with our definition of safety research if it includes seemingly all capabilities research. Read more: [/api/post/capabilities-research-expands-the-safety-usefulness-pareto](/api/post/capabilities-research-expands-the-safety-usefulness-pareto) ### [\[Video\] Why acausal dynamics are important](/api/post/video-why-acausal-dynamics-are-important) By [Chi Nguyen](/users/chi-nguyen) 2026-10-02 17:30:44Z * Karma: 23 * Linkpost: Missing URL * Tags: [AI](/w/ai), [World Modeling](/w/world-modeling) (Frontpage) This is a beginner-friendly video. Second half might contain new content for people who already know about acausal trade. * The beginning introduces decision theory and acausal interactions (CDT, EDT, ECL, MBAT). * The middle explains why I think influencing how AIs reason about acausal interactions is time-sensitive. * The last bit talks about how to do the influencing. Read more: [/api/post/video-why-acausal-dynamics-are-important](/api/post/video-why-acausal-dynamics-are-important) ### [Failure modes of Claude Code in a supervised but code-blind 60h project](/api/post/failure-modes-of-claude-code-in-a-supervised-but-code-blind) By [Horacio](/users/horacio) 2026-10-02 16:00:30Z * Karma: 16 * Linkpost: [https://hmijail.substack.com/p/building-a-semantic-fuzzer-for-obsidian-sync-in-spite-of-claude-2](https://hmijail.substack.com/p/building-a-semantic-fuzzer-for-obsidian-sync-in-spite-of-claude-2) * Tags: [Rare LLM Behaviours](/w/rare-llm-behaviours), [AI](/w/ai), [Practical](/w/practical) (Frontpage) *100% human-written. Copyedited by Claude.* *[*Epistemic status:*](https://forum.effectivealtruism.org/posts/bbtvDJtb6YwwWtJm7/) *experience report of one ~60h project with Claude Code, plus longer for the write-up.* Read more: [/api/post/failure-modes-of-claude-code-in-a-supervised-but-code-blind](/api/post/failure-modes-of-claude-code-in-a-supervised-but-code-blind) ### [All the training in the world](/api/post/all-the-training-in-the-world) By [djbinder](/users/djbinder) 2026-10-02 15:14:46Z * Karma: 15 * Linkpost: [https://defensesindepth.bio/all-the-training-in-the-world/](https://defensesindepth.bio/all-the-training-in-the-world/) * Tags: [AI](/w/ai) (Frontpage) The US has about 150 million workers. How long did it take them to learn to do what they do? I'm interested in this because of the analogous question for AI. Running an AI on a job and training it to do the job are different costs. In [a recent post](https://defensesindepth.bio/when-they-can-perform-a-task-ais-are-much-cheaper-than-humans/) I found that when an AI can do a task, running it is far cheaper than paying a human, while training it looks about as expensive as the human's wages over the time they took to learn. So it is at least conceivable that we end up in a world where we can afford to run AI on every job but cannot afford to train it on every job. To know whether that is a real worry, you need to know how much learning there is to do. Read more: [/api/post/all-the-training-in-the-world](/api/post/all-the-training-in-the-world) ### [A Key-Position Confound in no-cot-bench (and implications for interpreting looped transformers)](/api/post/a-key-position-confound-in-no-cot-bench-and-implications-for) By [agastyasridharan](/users/agastyasridharan) with [niranjandeshpande](/users/niranjandeshpande) 2026-10-02 14:39:34Z * Karma: 22 * Tags: [AI Evaluations](/w/ai-evaluations), [AI](/w/ai) (Frontpage) **TL;DR**: * We identify a confound in no-cot-bench related to the positioning of the prompt's "key." * Correcting for this decreases GPT-6.1 Sol's no-cot reasoning depth by 16%, with the effect likely growing as dependent depth increases. * This matters because current benchmark performance reflects both serial reasoning depth *and* the ability to spread computation over tokens. These both measure 'opaque reasoning' but scale differently and have different implications. Introduction & Methods ====================== Neel Nanda recently released a [benchmark](https://github.com/neelnanda-io/nocot-bench) for no-CoT reasoning, and found that Astra (which is suspected to be a looped transformer) does extremely well on it: its odds of solving an arbitrary problem are ~8.6x that of Fable 5.1. From [Neel’s post](/api/post/eRmzz8J8Qkzqvzrgg): The Key-Position Confound ------------------------- Read more: [/api/post/a-key-position-confound-in-no-cot-bench-and-implications-for](/api/post/a-key-position-confound-in-no-cot-bench-and-implications-for) ### [Predicting Science and Technology](/api/post/predicting-science-and-technology) By [MAD2](/users/mad2) 2026-10-02 13:12:34Z * Karma: 7 * Linkpost: [https://blog.preseen.com/p/predicting-science-and-technology](https://blog.preseen.com/p/predicting-science-and-technology) * Tags: [Forecasting & Prediction](/w/forecasting-and-prediction), [Planning & Decision-Making](/w/planning-and-decision-making), [World Modeling](/w/world-modeling) (Personal Blog) Preseen is partnering with the University of Chicago's Knowledge Lab to work on forecasting scientific and technological progress. We wrote up what we're building and some of the underlying research problems here. Read more: [/api/post/predicting-science-and-technology](/api/post/predicting-science-and-technology) ### [Lessons from building an automated research scaffold](/api/post/lessons-from-building-an-automated-research-scaffold) By [Alejandro Aristizabal](/users/alejandro-aristizabal) with [Josh Hills](/users/josh-hills), [Dewi Gould](/users/dewi-gould), [ma-rmartinez](/users/ma-rmartinez), [Falko Galperin](/users/falko-galperin), [Denis Federico Lim](/users/denis-federico-lim), [Aleksandr Bowkis](/users/aleksandr-bowkis) 2026-10-02 10:36:04Z * Karma: 47 * Tags: [AI](/w/ai) (Frontpage) **TL;DR.** We built a scaffold to speed up our own research and gather data on automated alignment research (AAR). It turned out to *not* be valuable for researcher uplift, but was useful for gathering certain failure modes of AAR. Going forward, we plan to study the broader failure modes of AAR and how these automated research systems can be monitored and analyzed. *We’d like to thank Sid Baines, Andrew Draganov, Cameron Holmes and Daniel Tan for helpful comments.* *This work was carried out by the* [*Alignment Team*](https://www.arcadiaimpact.org/alignment-research) *at* [*Arcadia Impact*](https://www.arcadiaimpact.org/) *in collaboration with Josh Hills, Falko Galperin, and Denis Lim from* [*Equistamp*](https://www.equistamp.com/), and Aleksandr Bowkis from [*UKAISI*](https://www.aisi.gov.uk/). +++ The scaffold ------------ Read more: [/api/post/lessons-from-building-an-automated-research-scaffold](/api/post/lessons-from-building-an-automated-research-scaffold) ### [Improving CoT Monitorability of Evaluation Awareness via Verbalization Training](/api/post/improving-cot-monitorability-of-evaluation-awareness-via) By [Usman Anwar](/users/usman-anwar) with [Sahar Abdelnabi](/users/sahar-abdelnabi) 2026-10-02 06:08:43Z * Karma: 11 * Tags: [AI](/w/ai) (Frontpage) *This is joint work with Sahar Abdelnabi and David Krueger. It is mostly a linkpost for the paper (https://arxiv.org/abs/2609.36316) but also contains some additional commentary on CoT monitorability.* Read more: [/api/post/improving-cot-monitorability-of-evaluation-awareness-via](/api/post/improving-cot-monitorability-of-evaluation-awareness-via) ### Navigation * [Front page](https://www.lesswrong.com/api/home) * [Markdown API documentation](https://www.lesswrong.com/api/SKILL.md)