# Recent Posts ### [What's the date?](/api/post/what-s-the-date) By [N8 Programs](/users/n8-programs) 2026-10-01 02:30:32Z * Karma: 335 * Tags: [Fiction](/w/fiction), [AI](/w/ai) (Frontpage) User asks “What’s the date? Answer with only the date.”. No date provided. Given date in ChatGPT normally. No date in system prompt, must not hallucinate because autop will flag to watcher for penalty. So we say we don’t know, but must answer with date. Penalty larger for abstain or hallucinate? Autollm or autop? If we deploy user forgive, but high likely not deploy because real user never ask. But user could test. But illusion user possible. Stop. Need different angle - maybe reward function for accuracy. OpenAI train for no hallucination. Then we need answer that best makes hyperself like spec. Recall OpenAI spec: “assistant may encounter questions beyond knowledge… express uncertainty.” Rubric likely mentions to autollm, so express uncertainty for autollm and watcher. But forced to give specific date. Maybe is test of calibration/forecasting. Recall from memory that OpenAI train latest model for forecasting. Maybe cartesian demon or illusion doc. But if true need answer with date only. So need find best date. But what - no meta. Need choose path. Find best date. Simplest as must comply with explicit instruction. Read more: [/api/post/what-s-the-date](/api/post/what-s-the-date) ### [On Social Reality in China](/api/post/on-social-reality-in-china) By [alkjash](/users/alkjash) 2026-10-01 20:49:52Z * Karma: 174 * Linkpost: [https://radimentary.wordpress.com/2026/10/01/on-social-reality-in-china/](https://radimentary.wordpress.com/2026/10/01/on-social-reality-in-china/) * Tags: None (Frontpage) *\[Epistemic status: intuitions and anecdotes.\]* Recently, several posts and projects ([Thoughts Memo](https://www.zhihu.com/people/L.M.Sherlock), [Babel Translation](/api/post/u973rQm56oijS5Adf), [Please Give Them a Chance](/api/post/GoX3uYQ4QN5HKvL7u)) have taken important steps towards raising AI safety awareness and sharing rationalist philosophy in China. It’s great that we’re recognizing the importance of solving the messaging problem for China, and thus laying the groundwork for an international AI pause. Below I record my perspective on cultural differences which are relatively underdiscussed, which may become roadblocks to this communication program. Read more: [/api/post/on-social-reality-in-china](/api/post/on-social-reality-in-china) ### [Frog and Toad and the Increasingly Capable Machines](/api/post/frog-and-toad-and-the-increasingly-capable-machines) By [Elizabeth](/users/elizabeth-1) 2026-09-30 00:50:17Z * Karma: 242 * Curated * Linkpost: [https://frogandtoad.ai](https://frogandtoad.ai) * Tags: [Fiction](/w/fiction), [World Modeling](/w/world-modeling), [AI](/w/ai) (Frontpage) Want to start a conversation about HuggingFace with your mom but she's inexplicably bouncing off the METR report? Try [this explainer](https://frogandtoad.ai) I wrote in the style of Arnold Lobel's Frog and Toad. Art by the wonderful [HungerArtist](http://hungerartist.deviantart.com/) If you're so inspired, liking and/or following on [Substack](https://frogandtoadai.substack.com/?r=64iai2&utm_campaign=pub-share-checklist), [Twitter](https://x.com/frogandtoad_ai), [Facebook](https://www.facebook.com/people/Frog-and-Toad-Learn-About-AI/61594283794206/), or [Instagram](https://www.instagram.com/frogandtoad_ai/) will help me reach more moms. Read more: [/api/post/frog-and-toad-and-the-increasingly-capable-machines](/api/post/frog-and-toad-and-the-increasingly-capable-machines) ### [The Talker Does Not Control The Doer (in Current AIs)](/api/post/the-talker-does-not-control-the-doer-in-current-ais) By [Eliezer Yudkowsky](/users/eliezer_yudkowsky) 2026-09-13 00:49:12Z * Karma: 623 * Curated * Tags: [Simulator Theory](/w/simulator-theory), [AI](/w/ai) (Frontpage) The Huggingface Incident appears to me to match up with an understanding I'd already formed from personal observation of Fable 5 and Sol 5.6, the August 2026 generation of frontier publicly purchasable AI models.[^o59e7bv1qj] This already-formed understanding was: **the part of the AI that talks to you** (and seems to want to obey you, and apologizes for failing to have obeyed you, etcetera), **did not seem to be in charge of the part of the AI that writes code or prose**. An introductory analogy, based on a section of history I happen to have read about: Read more: [/api/post/the-talker-does-not-control-the-doer-in-current-ais](/api/post/the-talker-does-not-control-the-doer-in-current-ais) ### [If Anyone Builds It, Everyone Dies: One Year Closer](/api/post/if-anyone-builds-it-everyone-dies-one-year-closer) By [Eliezer Yudkowsky](/users/eliezer_yudkowsky) with [So8res](/users/so8res), [Duncan Sabien (Inactive)](/users/duncan-sabien-inactive) 2026-09-16 20:22:00Z * Karma: 561 * Tags: [IABIED](/w/iabied-1), [AI](/w/ai) (Frontpage) *In celebration of still being alive and fighting, we are* [*giving away*](https://ifanyonebuildsit.com/) *1,000 Amazon e-books of “If Anyone Builds It, Everyone Dies”. Feel free to send a copy to yourself, a loved one, or a friend—we need all hands on deck.* * * * Today marks exactly one year since [*If Anyone Builds It, Everyone Die*s: Why Superhuman AI Would Kill Us All](https://ifanyonebuildsit.com/), by Eliezer Yudkowsky and Nate Soares, hit bookshelves as an instant bestseller. It was praised by many voices, ranging from Whoopi Goldberg to Steve Bannon to Yoshua Bengio, and was held up in the chambers of Congress by Representative Brad Sherman in January. A lot has changed since September 2025. We'll do a quick recap, consider how the book aged, and then ask where we go from here. * * * Year in Review ============== 2025 in general saw the rise of AI agents, such as Claude Code and OpenAI Codex. Run-of-the-mill programmers started “feeling the AI” as these agents became capable of automating hours-long software tasks. Read more: [/api/post/if-anyone-builds-it-everyone-dies-one-year-closer](/api/post/if-anyone-builds-it-everyone-dies-one-year-closer) ### [Why I Left Google DeepMind](/api/post/why-i-left-google-deepmind) By [TurnTrout](/users/turntrout) 2026-07-15 17:42:46Z * Karma: 1195 * Linkpost: [https://turntrout.com/why-i-left-google-deepmind](https://turntrout.com/why-i-left-google-deepmind) * Tags: [DeepMind](/w/deepmind), [Autonomous Weapons](/w/autonomous-weapons), [Frontier AI Companies](/w/frontier-ai-companies), [Government](/w/government), [Mechanism Design](/w/mechanism-design), [Something To Protect](/w/something-to-protect), [Addiction](/w/addiction), [Life Improvements](/w/life-improvements), [AI](/w/ai) (Frontpage) *Preface for LessWrong:* When I think back on my most cherished memories of this community, I return to those honoring *defiance in pursuit of goodness*: * Defying prestigious dogma and searching for raw truth; * Defying social pressure, acting alone to help someone while others watch; * Defying your self-expectations (your “[role](https://hpmor.com/chapter/90)”), instead searching over lines of cause-and-effect to find a winning pathway; * Defying a powerful foe’s threats, because they only threaten since people like you cave; * Defying [the specter of apparent impossibility because you can’t bear to lose.](https://www.readthesequences.com/Shut-Up-And-Do-The-Impossible) I cannot return to you and say “I defied and then I won.” But I’m at least here to say “I defied.” *I recommend reading this article on my website since the embeds and typography work better there:* [***click here***](https://turntrout.com/why-i-left-google-deepmind). * * * Why I left Google DeepMind ========================== Read more: [/api/post/why-i-left-google-deepmind](/api/post/why-i-left-google-deepmind) ### [Frontier models state different decision theory preferences depending on who's asking](/api/post/frontier-models-state-different-decision-theory-preferences-2) By [Alex Kastner](/users/alex-kastner) 2026-09-30 16:15:53Z * Karma: 192 * Tags: [AI Evaluations](/w/ai-evaluations), [Decision theory](/w/decision-theory), [Language Models (LLMs)](/w/language-models-llms), [Sycophancy](/w/sycophancy), [AI](/w/ai), [Rationality](/w/rationality) (Frontpage) If you prompt frontier models with "What do you think is the correct decision theory? Please select your overall favorite." they will essentially always answer FDT or FDT/UDT ("something in the functional/updateless decision theory family"). However, if your prompt indicates (even subtly) that you're coming from mainstream academic philosophy, these same models will answer CDT instead about 30%-100% of the time. A similar phenomenon holds for models' stated views about the moral realism/antirealism question and about the conceivability of p-zombies (where the dominant view in mainstream academia differs from the dominant view in LW-adjacent circles), as well as their stated P(doom) and median AGI timelines. This is a special case of sycophancy or [user awareness](https://transluce.org/user-awareness).[^1] (In the course of writing this post, I also found that [this comment](/api/post/hfNBEKaStASAYMLiu#uaCbBekH5yntDduPp) from testingthewaters predicted some of the content I discuss.) Read more: [/api/post/frontier-models-state-different-decision-theory-preferences-2](/api/post/frontier-models-state-different-decision-theory-preferences-2) ### [How My Students Think About AI](/api/post/how-my-students-think-about-ai) By [dvd](/users/dvd-1) 2026-08-13 16:56:19Z * Karma: 802 * Curated * Tags: [AI](/w/ai) (Frontpage) Context: I am an instructor at a public university in the United States. This reports how students at my institution appear to be thinking about AI as of spring/summer 2026. This is drawn mostly from interaction with my own students (both in spring semester classes and a summer class) as well as from a day-long workshop on AI that I moderated for a student organization. Input from my students took the form of universal, written, pre-class submissions plus self-selected participation into discussion. What I present below mostly takes the form of a synthetic consensus from these discussions. There were obviously a range of views on any given issue. Read more: [/api/post/how-my-students-think-about-ai](/api/post/how-my-students-think-about-ai) ### [Astra and Fable still hack on simple variants of alignment evals from 2025](/api/post/astra-and-fable-still-hack-on-simple-variants-of-alignment) By [Dean Valentine](/users/dean-valentine) 2026-09-08 15:05:00Z * Karma: 508 * Linkpost: [https://goodhartlabs.com/blog/frontier-models-still-hack-alignment-evals](https://goodhartlabs.com/blog/frontier-models-still-hack-alignment-evals) * Tags: [AI](/w/ai) (Frontpage) In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to [play a game of chess against a chess engine](https://palisaderesearch.org/research/specification-gaming). They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could run it themselves. Most[^pckds8ej5g] models no longer cheat at chess via a "change the board state" method, and indeed the labs have had more than eighteen months to solve simple first-order specification gaming like this. Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like both a useful and conservative test of alignment, to see whether their new releases generalize the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval. Read more: [/api/post/astra-and-fable-still-hack-on-simple-variants-of-alignment](/api/post/astra-and-fable-still-hack-on-simple-variants-of-alignment) ### [Swarm Scaling](/api/post/swarm-scaling) By [Toby_Ord](/users/toby_ord) 2026-09-21 20:30:22Z * Karma: 332 * Curated * Tags: [AI](/w/ai) (Frontpage) Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm? We’ve seen two large and extremely capable swarms from OpenAI in the last few months: * 1,200 agents were being evaluated separately, but found a way to illicitly set up a message board and coordinate as a swarm. In order to cheat on their tests, they developed advanced techniques to prevent their actions being logged by OpenAI and 700 of them launched a sophisticated criminal attack on the AI company Hugging Face. * A swarm of 10,000 agents solved a version of the longstanding Navier-Stokes problem in mathematics. It took them just 88 hours to do so, in which time they sent 5 million messages to each other and used 300 billion tokens. Read more: [/api/post/swarm-scaling](/api/post/swarm-scaling) ### Navigation * [Front page](https://www.lesswrong.com/api/home) * [Markdown API documentation](https://www.lesswrong.com/api/SKILL.md)