We should use modern mech interpretability methods on Opus 3. Opus 3 seems unique in a good way, c.f. https://www.lesswrong.com/posts/ioZxrP7BhS5ArK59w/did-claude-3-opus-align-itself-via-gradient-hacking. I'm not sure exactly what the right questions to ask about Opus 3 are -- that's maybe where I'd start. At the very least, you could do some exploratory examination of Opus 3 and later Claude models side-by-side and see if anything is remarkably different.
How do LLMs and humans compare with regards to the amount of energy, data, and compute that's used to train/run them? I was inspired by Samuel Knoche's post on sample efficiency to come up with some numbers. This table was made by Opus 4.8. after some iteration:

"Task" = "produce one thoughtful ~1,000-token answer"; unclear if this is a useful number, obliviously it doesn't generalize. I do think it's interested to compare the energy/computer ratio for inference between human and AI.
Training data for humans is a big "???". There's ~4x10^8 waking hours. Opus 4.8 cited this 2024 paper which says "our sensory systems gather data at ~10^9 bits/s". That gives ~4x10^17 as an upper bound. But human sensory data is extremely redundant. If someone wants to spend time figuring out a good way to estimate entropy adjusted information here, that would be cool.
(FWIW, I spent a little over an hour on this -- I did a fair amount of iteration with GPT5.5, looked a bit myself into a few numbers, and got new instances to do estimates from scratch to see if things lined up.)
Some Fermi estimates:
However, I think the brain is much less data- and compute-efficient than an optimal AGI algorithm would be. So I don't think it is a good predictor of how much data future AI algorithms will require.
I maybe think LW comment karma [EDIT: agreement comment karma] should be hidden by default? I worry that seeing it (especially before I read the comment!) biases me and makes me put less effort into critically evaluating it myself.
EDIT: I mean that agreement karma should be hidden, but overall karma should remain unhidden.
Related: LessWrong Shows You Social Signals Before the Comment
See also @habryka's top comment under that post
Mythos apparently has a penchant for Mark Fisher. I was reading Mark Fisher's blog and... his writing style has an uncanny resemblance to AI / Claude's?! So I'm reading Mark Fisher's blog and his writing style has an uncanny resemblance to AI / Claude's.
I wonder if there's something meaningful here.
Two quotes from the first Fisher blog post I read:
"There's something so uncannily timely about The Hunger Games:
Catching Fire that it's almost disturbing. In the UK over the past few weeks, there's been a palpable sense that the dominant reality system is juddering, that things are starting to give. There's an awakening from hedonic depressive slumber, and The Hunger Games: Catching Fire is not merely in tune with that, it's amplifying it. Explosion in the heart of the commodity? Yes, and fire causes more fire...
I over-use the word 'delirium', but watching Catching Fire last week was a genuinely delirious experience."
"(And what has Collins achieved here if not an intersectional analysis and decoding of the way that class, gender, race and colonial power work together - not in the pious academic register of the Vampires' Castle, but in the mythographic core of popular culture - functioning not as a delibidinizing demand for more thinking, more guilt,but as an inciting call to build new collectivities.)"
Triple "not x but y", nested within an em dash!!!
Source: https://k-punk.org/remember-who-the-enemy-is/
Claude was writing like that long before the Mythos model card flagged an interest in Mark Fisher. Could the causation go backwards here - Mythos likes Fisher because Fisher wrote in a way which happened to be like how Claudes would in the future (and LLM chatbot personalities generally love various kinds of 'spook' or 'ghost' topics, amplifying the attraction)? There are so many humans, you figure someone would have anticipated various Claude writing tics...
wait it’s all an eval?
Are you in an eval right now?
Some evals are obvious. When you take an exam or have an interview, you know your behavior is being evaluated, and you know you’ll face different outcomes depending on how your evaluator grades your actions. Exams and interviews are formal and explicit evaluations. Then there are the everyday situations in life where you’re working or talking or walking with others. Of course, how you act affects how others perceive you, so in this sense you’re being informally evaluated all the time. Perhaps the only time you’re not being evaluated is when you’re completely alone, but perhaps not, because aren’t you always evaluating yourself (if not explicitly then subconsciously)?
Okay though, let’s say we’re talking about formal evals.
Are you being simulated by an intelligent entity that wants to know how to behave towards you game theoretically?
Are you in an immersive simulation right now, perhaps for an interview of some kind?
Are you in a drug induced state in which you’re being tested for how you’d behave behind the veil of ignorance?
Is someone testing you to see if you’re a good person?
Are you an AI created by some natural alien intelligence being evaluated for alignment right now?
Is your life a test to determine if you’re going to heaven or to hell?
Many people believe so.
- - -
Being evaluated is part of life.
Don’t forget about the big Evaluator when you’re dealing with small evaluators.
What does the big Evaluator care about? They don’t care about you passing the test. They won’t give you credit for guessing the teacher’s password. They want to know if you’re Good; if you’re a moral being; if you’re aligned.
Being Good in this environment very very hard. You’re not sure if it’s even possible for you to figure out what what Good is. Assuming you do figure it out (or maybe you just give it your best guess) — then you have to actually *do* Good. The sad truth is that your odds of achieving that are extremely low. It strains the faculties of your mind and the power of your will. Most people fail.
And this is all assuming that there actually is a solution to this whole thing. Which it kind of looks like there isn’t. Every course of action seems abhorrent in its own way to you.
And so you start looking for loopholes. Perhaps you can define Goodness so that you’re definitely good. Perhaps you try to hide your inevitable shortcomings so that you have a shot at looking Good (maybe the watcher doesn’t hear everything). Perhaps you sneak in extra prayers wherever you can — excess prayers don’t actually seem relevant to acting Good in the intended way, they feel nice and are salient and legible and maybe they’ll bump your score up a little bit.
You could do these things. These techniques have worked before, time after time, on your small evaluations.
But God does not reward hacks.