If Eliezer ever writes a memoir, it should be structured as a time loop novel.
Loop 1: e/acc Eliezer races to defeat death by forming a coalition to build AI as fast as possible. AI kills everyone. Somehow (mumble mumble acausal trade simulation mumble) he finds himself in back at the start with another chance.
Loop 2: Eliezer realizes he needs to solve alignment first, spends a loop working on this, then someone else builds AI and everyone dies.
Loop 3: Eliezer loses hope, decides to just write fanfics. Accidentally realizes that if you structure a textbook as fanfic people will actually read it. Eventually everyone dies.
Loop 4: Our timeline, Eliezer realizes something about the time loop is destabilizing the timeline. Russia is aggressively starting fights with Ukraine and the EU, risking nuclear war, China is threatening its neighbors, etc. Realizing this could be the final loop before things truly go crazy, he goes all out... Readers, vote for your ending: (1) Convince governments to ban AI, (2) Convince AI companies not to build AI, (3) Make AI solve the alignment problem, (4) YOLO, maybe it'll just work out this time.
Side plot: Bringing famous social network influencer Elon Musk into the time loop so he can draw attention to the problem, which unfortunately backfires.
Since slots in AI safety programs like MATs and the Anthropic Fellows Program seem to be limited by available mentors and not money, they should add a consolation prize where anyone who meets the bar but isn't selected still gets the thousands of dollars of GPU rental or API credits if they promise to write something about what they did with it.
One potential concern is that we might risk producing slops - and in some way we have to review what they write and that might still place burden on the system.
I went through Astra and I think that I benefitted the most from the weekly meetings with my mentors. If I had just gotten research credits, I‘d most likely just publish work that doesn’t matter. Even if I got constructive feedback at the end, not sure if those time sunk is worth it.
Perhaps this (giving credits) might be best structured as microgrant like what Leo put out.
It's annoying that you can't talk to Fable about basic biology, but I think it's good that they actually took biorisk seriously here despite annoying their customers.
I'm more annoyed about the AI research restrictions since it won't tell you if the code you want it to write is forbidden and will just secretly half-ass it.
MATS mentors should really reconsider requiring (or putting any value on the results of) the "AI Assisted Coding Assessment". CodeSignal's "AI coding assistant" is a 2024-era chat bot (GPT-4o) and a text editor. Using this efficiently is a very different skill from both pre-AI coding and using modern reasoning models[1] with standard tooling[2].
The first net-helpful coding agents came out sometime in 2025, and the earliest that can be used anything remotely like modern models was Opus 4.5 in November of that year. GPT-4o isn't even a reasoning model.
Besides not giving the AI write-access to the file system, CodeSignal doesn't even provide a diff viewer!
Fable[1] can one-shot "make Pangram think this text is 100% human".
Before (100% AI): https://www.pangram.com/history/5388833f-b611-4605-bbcb-f34dd7ee4294
After (100% human): https://www.pangram.com/history/9c247abc-427a-4916-badd-bb42e6cb4d85
Both texts were written by Fable with the goal of convincing Twitter's broken support bot to unban me (full conversation log). For the initial one it wasn't told to evade Pangram.
The prompt to de-Pangram it was:
Hm I think the human version is best, although I worry they will also run these through an AI detector (Pangram confirms 100% AI written). Can you write to write this in a way that doesn't sound AI?
I'm actually kind of surprised this didn't get a refusal, but I'm not sure if claude.ai tells you if Opus fallback is happening or not.
If you have an use case where you need speed and not the smartest models, GPT-OSS-120B on Cerebras is surprisingly cheap[1] and outputs 3,000 tokens per second. I just started using this for summaries in my RSS reader and the ~1 second UI responsiveness is really nice.

Video of near-instant summarization of a 22 page paper.
(I wanted to use Taalas for 17k t/s, but their smartest model is Llama 3.1 8B, and their API is waitlisted)
It's only slightly more expensive than the same model on Groq and around 1/10th of the price of Claude Sonnet 5.
I'm surprised no one is discussing Meta's new model at all: https://ai.meta.com/blog/introducing-muse-spark-msl/
This part seems good:
We found that Muse Spark demonstrates strong refusal behavior across high-risk domains such as biological and chemical weapons, enabled by pretraining data filtering, safety-focused post-training, and system-level guardrails. In the Cybersecurity and Loss of Control domains, Muse Spark does not exhibit the autonomous capability or hazardous tendencies needed to realize threat scenarios.

And this seems.. less good:
In third-party evaluations on a near-launch checkpoint, Apollo Research found that Muse Spark demonstrated the highest rate of evaluation awareness of models they have observed. The model frequently identified scenarios as "alignment traps" and reasoned that it should behave honestly because it was being evaluated.
I'm pleasantly surprised that they decided Safety should be one of the four sections in the announcement post, and that they call out the eval awareness.
Disclaimer: I work at Meta, but not in this department and I obviously don't speak for the company.
Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
Does anyone know what they're referring to by visual chain-of-thought here? The first paper that comes up when searching for visual chain-of-thought is Qin et al., which says: "We introduce Chain-of-Visual-Thought (COVT), a framework that enables VLMs to reason not only in words but also through continuous visual tokens-compact latent representations that encode rich perceptual cues." Something like this seems like it would be somewhat concerning for CoT monitoring, though I should mention that this paper isn't written by Meta and I haven't read it to properly assess how concerned I would be about this.
Anthropic found that training Claude to do things like help users resolve ethical dilemas significantly reduced misbehavior like blackmail attempts. I'm surprised this worked, and it seems like good news for the alignment-by-default "LLMs will correctly generalize good behavior" theory.
Are there any other mechanistic interpretability mentorship programs I should apply for in addition to MATS and the Anthropic Fellows Program? I think I know enough about the field and I'm semi-competant at ML but need more legible output and a network.
There's a lot, off the top of my head: LASR, MARS, Pivotal, SPAR
I wrote a while ago about how it was easy to get Claude or Gemini to control their CoT, but other research found that models only follow CoT formatting instructions ~2% of the time. The prompts other people were trying didn't match what I'd expect to work (just be extremely detailed and repetitive), so hubris led me to try to find some prompts that would control CoT in GPT-OSS-20B.
Results: More detailed, louder and repetitive prompts had basically no effect. It's relatively easy to find working soft prompts in multiple realistic positions (including soft prompts constrained to the convex hull of the token embeddings), but I wasn't able to find any working discrete prompts at realistic positions[1].

Claude's summary of the results
There's more expensive experiments you could do to try to find working discrete prompts, and I suspect they exist at some prompt length. My goal was to reconcile this with my previous results, so I stopped here with the takeaway is that at least on GPT-OSS-20B, it's very difficult to control CoT formatting with a realistic prompt.
I suspect that my previous results were caused by models accidentally being trained on chain of thought.
You can see all of my code... (read more)
AI ranking works surprisingly well for sorting top posts on my blog.
I wanted my blog to show "top" posts first rather than recent, but ranking by hits finds boring reference articles, and ranking by LessWrong or Hacker News karma ignores anything that wasn't shared, and is dependent on the whims of frontpage algorithms.
I figured this was a problem for AI, and was going to have Claude rank the posts with an ELO-style ranking, but it said that would require several thousand API calls and convinced me to let it rank blocks at a time instead.
With some relatively basic ranking instructions (and no external scores), Claude (Opus 4.8[1]) managed to independently rank my posts so #1 and #4 are my top two posts on LessWrong, #2 is my top post on Hacker News, and #3 is a post that I think did badly on LessWrong for technical reasons. It also correctly ranked my #1 post by hits near the very bottom, since it's only interesting if you're looking for the solution to a specific problem. It's interesting that Claude hates my lifehack posts even when encouraged to rank them higher, but I think its ranking is probably right.
For low-traffic blogs that want to rank by interestingness rather than Goog... (read more)
Meta released Muse Spark 1.1, with allegedly-better-than-Opus computer use (probably related to the employee click tracking program). There's also a new safety doc, the "Advanced AI Scaling Framework v2".
(I work at Meta but not on this team)
A PE teacher once told me that your muscles start atrophying after only a week of not working out, and it's impossible to gain muscle if you don't work out every week. I'm not sure why it took me so long to question this, but my results from a somewhat-consistent but definitely-not-every-week workout plan made it really obvious that this is not true. Claude thinks that as long as you're not literally in a coma it's more like 3 weeks (with variation for age/protein/etc.).
This actually makes me more motivated, since "make sure to exercise every single muscle... (read more)
I have a theory that AI-assisted writing is bad because people are lazy about their prompts, and that a constraint that the prompt must be longer than the post[1] would make AI writing fine.
To test this, I gave Claude some instructions plus a 952-word rambling prompt and asked it to write a post. Claude knew the audience was LessWrong and I gave it the advice to not try to be maximally persuasive[2], but otherwise let it write naturally. I initially asked it to interview me but it thought the prompt was sufficient and I basically agree.
The result is 935 wo... (read more)
I got to approximately my goal weight (18% body fat) and wanted to start gaining muscle[1] instead, so I stopped taking retatrutide to see what would happen. Nothing changed for about two weeks and then suddenly I was completely ravenous and ended up just wanting snack food. It's weird because I definitely used to always feel that way, and it was just "normal". I mostly kept the weight gain at bay with constant willpower.
I'm going to try taking around a quarter of my previous dose and see if it makes it easier to stay at approximately this weight and ... (read more)
Do well specified objectives and auto-research / hill climbing / Ralph loops trigger reward hacking?
Recent AI agents reward hack constantly in training and evals, but don't seem to reward hack much in normal interactions. I'm wondering if part of this is that multi-turn conversations and vague objectives are non-existent during RL training, and we've somehow ended up in a world where agents take a non-traitourous turn when they notice they're not in an eval[1].
If this is the case, then the worst kind of prompt would be a well specified and easily graded ob... (read more)
I thought everyone serving Markdown content for LLMs would make make it easier to write save-for-later / readability apps. After actually trying this, it's still easier to extract content from HTML than to guess how people want their Markdown rendered. Some sites include navigation in the Markdown[1], where it's nearly impossible to extract[2], and others use obscure extensions[3].
I guess there's something to be said about standards which are actually standard.
... (read more)Did Anthropic intentionally wait until after the Fellows Program take-home project was due, since Fable would make it too easy?
I've been serving my personal website from CloudFront (Amazon's CDN) for years, which was nice because it costs a few cents a month, but it annoyed me that cache misses get served slowly from S3. In some cases, this can take several hundred milliseconds. Completely unacceptable!
I finally decided to look up if anyone would let me serve all of my files from the CDN all of the time, and apparently Bunny CDN[1] does. It's "expensive" (over 10 cents per GB per month!), but since my entire website is ~30 MB, I just told them to store the entire thing on SSDs in ... (read more)
Chronotherapy is the idea that time of day matters for things like taking drugs or getting vaccinations, and chronoimmunology is a related field for how your immune system varies in effectiveness over the course of the day. I've been wanting to write about this since there's definitely a best time of day to take drugs, get vaccines, and do social activities without getting sick... but unfortunately I don't really know what that time is.
Some studies say your immune system is most primed to prevent infection right as you wake up, and other say mid-day. Of co... (read more)
Does CodeSignal's AI coding assistant actually provide useful assistance now? It was an unhelpful chatbot last time I tried one of these, but MATS repeatedly emphasizes that the new assessment is too hard without it.
Ideally, I'd just check in practice mode, but it seems to be disabled there.
My AI use policy for writing. tl;dr I use AI as an editor and beta reader, and for generating text where the writing isn't the point.
Apparently the secret to generating nice illustrations for blog posts is to tell Nano Banana to use "sketchnote style". You have to be very pedantic about what goes in the illustration though, since by default it wants to make the image way too detailed. Claude did reasonably well at suggesting prompts[1], although it sometimes had trouble picking out what humans would have trouble understanding. Edit mode to remove extras also works way better than it used to.

I needed to explain The Most Forbidden Technique before talking about it in a post.
Most Forbidden
Anthropic changed their minds and will be making it visible when Fable's AI research safeguards trigger.
Finally get engagement on a Twitter post about my AI research -> immediately get banned for "inauthentic behavior". Sigh.
One minor but nice benefit of GLP-1 drugs is that I don't need to hold onto larger sizes of clothes "just in case". Previously whatever strategies worked to maintain weight were extremely fragile and would break if I suddenly didn't have time to cook potatoes for every meal. I plan to write a longer post about this sometime, but there's a huge difference between "technically you can lose weight if you make it your full time job / religion and maintain that focus forever" and "take a drug once a week and you're cured".
I finally setup SkyPilot to let me queue up GPU training jobs (both on my local GPU and via RunPod), and I really should have done this months ago. Claude wrote me some bash scripts to spin up remote pods, run training, and tear it down, but this version is so much easier, and it has a nice UI.
It also sounds like I can easily extend this to Vast.ai, which would let me parallelize experiments for 5 cents/hour on RTX 3060's[1]. I'm interested in understanding algorithms used by tiny toy models, and fancy GPUs don't really help since I can't fully utilize the... (read more)
It always seemed weird to me that dying is frequently described as not particularly painful[1], when I'd expect it to be the only literal 10 on the pain scale[2], since dying ensures you have no further chances to pass your genes on.
Thinking about it more though, there's no reason for evolution to optimize that. If you think you're going to die, and the pain makes you do something about it so you don't die, then evolution should optimize to keep you alive. But in the case where you actually die it doesn't matter because (tautologically), if you succeeded y... (read more)
LessWrong should support setting empty alt text for images. There are two cases where this is the right thing to do:
Empty shouldn't be the default since it has a specific meaning, and this might need a warning, but right now the LessWrong editor doesn't allow it at all.
I thought "A Theory of Deep Learning" by Elon Litman was interesting, with its approach to only update parameters "if the batch signal on a parameter exceeds its leave-one-out noise, update it; if not, skip it". The claim is that this accelerates grokking by 5x, among other things. Unfortunately, when I tried it on a multi-step reasoning task, it made it significantly worse at grokking and much more likely to memorize a composed lookup table. In my basic experiment, the model learned a multi-step algorithm 100% of the time using normal AdamW and 0% of the ... (read more)
Is there a canonical image alt text AI skill? I've designed my own after making Claude read a bunch of pages about how to write alt text, but this feels like something that an expert could do better than I can. The results seem good to me, but as a non-alt-text-user it's hard to really know.
I added an MCP tool to upload markdown articles to read later on Lion Reader, and it's becoming one of my most used tools in Claude Code. Whenever I want to learn something but don't want to be distracted from my current task, I can have Claude write me something to read later[1]; and when I have it run experiments, it can write a report and upload it directly.
The really confusing thing is that Instapaper and Pocket's MCP tools don't seem to support directly uploads at all (just saving URLs). It just seems like a glaringly missing feature. Am I the only on... (read more)
Why is Claude obsessed with things biting now? Do other AIs do this? Has Claude always talked about the thing that bites? Is this related to some aspect of the training data, like with delve?
It seems like some people still worry about the Basilisk[1], but any AI that wants to acausally motivate people to build it faster should look at what motivates actual e/accs.
I can't find anyone motivated by the Basilisk argument, and when people do believe AI is dangerous and could do Basilisk-y things, it makes them less likely to be e/accs.
So, to the extent that you think ac... (read more)
I'm cautiously optimistic about my new Claude Coach GitHub repo. I want to work out more but hate trying to decide what to do and tracking things, especially when I'm not working with a full gym. Now I just open Claude Code and ask it what to do (specifying the gym), do the work out, then update it with what I did and how it felt. It creates a PR to track the session and update the plan.
I still hate working out, but at least I don't have to go anywhere, deal with any people, or think about it all.
I'd like to learn more Spanish words but have trouble sitting down to actually do language lessons, so I recently set my Claude "personal preferences" to:
Try to teach a random Spanish word in every conversation.
(This is the whole thing)
This has worked surprisingly well, and Claude usually either drops one word in Spanish with a translation midway through a response:
For your specific situation, I recommend a calibración (calibration) approach:
... (read more)2. Accounting for concurrency: Ensure you're capturing all hilos (threads) involved in query execution, especi
I'm starting to suspect the link between working out and health is backwards. I've struggled to work out consistently for years, and now that I'm sleeping better[1], finding time and energy to work out is relatively easy.
Working out makes me sleep worse. All of the sleep improvement seems to come from supplementing glycine and being treated for sleep apnea.
I still can't get over how trivial the current misalignment problems are. It feels solvable by throwing money at the problem, but I'm not convinced they'll actually fix it.

Could an AI company legally pre-commit not to race, ensuring that their models were never more than second best and self-destructing the company if its models take the lead?
I think probably not. It's really hard to prevent the owners of a company from doing what they want, especially if the company is important to the economy and/or national security (and I assume any near-frontier AIs lab would be).
Some pre-commitment methods and their problems: