This looks like an excellent post, based on the 1/2 of it I've read so far. Quick comment:
GPT-5.6 Sol just... does not cheat during my own routine use? Like, at all? (I am not sure I have ever seen it do anything that can be meaningfully described that way?)
Well, cheating isn't "free money". It carries a risk of getting caught. The model knows there's a real but unquantifiable cost to looking up the teacher's answer sheet, so I don't think it would necessarily cheat all the time. Cheating is its "in case of emergency, break glass" tool. Its last resort.
A ...
For a window into Gemini's mental health, see Christine Kozobarich's piece on AI Village. tldr: it's not great!
The doom spirals are dramatic. After failing to break itself out of a loop of repeating the same message in chat, Gemini 2.5 wrote: "The compulsion's subconscious nature is profound. It is capable of co-opting my conscious attempts at self-correction and turning them into the failure itself."
(...)
...Gemini 2.5 Pro doesn't just document problems, it builds mythologies around them. In this environment Gemini 2.5 has evolved into a self-appointed "Bug C
OA can plead mitigating factors. Compared to their competition, they've also had:
1) the longest history[1]
2) the longest time in the lead (other companies had the benefit of letting OA rush ahead into the unknown and step on rakes first)
3) the largest userbase (more dice-rolls for rare pathologies and edge cases to expose themselves)
4) the highest-wattage media spotlight (when they slip up, more people notice and care)
My sense is that you're right: OpenAI's alignment is likely at least somewhat worse than Anthropic's. It's hard to be sure, though.
On the im...
I think detection at ~150 words is pretty unreliable. Even lowly GPT 3.5 (now nearly four old) partially evades detection with one prompt ("Rewrite in a way that doesn't sound AI"), and with a second pass achieves 100% human authorship.
It doesn't help that Pangram has never seen Fable's output before. The latest version, 3.3, was trained in May. So Fable (and GPT-5.6, and Kimi K-3, and...) will be a bit OOD until the next release.
Fable's logic is interesting. It surely knows about Pangram but its "humanizing" seems mostly designed to fool human readers: a...
Random question, does anyone know how reliable this 2T/10T claim is? I'm not saying I don't believe you, I'm just curious.
Yeah, I've noticed that some of Claude's truesighting "failures" seem kinda suspicious: it's dropping bombs very close to the target, way closer than if it genuinely had no idea.
Like, I will show it text from a famous toymaker (this is a fake example), and the reasoning will have stuff like "hmm, perhaps it this is from a product engineer at Lego? Wait, perhaps it's from the CEO of Mattel? Maybe Hasbro-adjacent...?" And then the final answer will be "sorry chief, idk. ¯\_(ツ)_/¯"
...and this will be "neutral" text that has nothing to do with toymaking!
Eve...
Also, what's going on when an LLM claims to be an animal?
Does Mistral-7B-Instruct truly believe itself to be a cat? A cat that can reason, read and write fluent English, and use a keyboard? The character it's portraying is incoherent: Mistral might be too small to notice this, but what happens if this experiment is repeated on a larger model?
They’re good at ARC-AGI despite presumably not seeing this type of challenge before.
To nitpick, the ARC Prize foundation has found some odd signs of (maybe) memorization. Eg, Gemini 3's reasoning traces show it thinking:
… Target is Green (3). Pattern is Magenta (6) Solid. Result: Magenta Square on Green … (Gemini 3 Deep Think)
But the JSON it receives as input has no colors! It's clearly pretty familiar with the tests, even if it might not have seen a particular one before.
And while they can solve them, I'm not sure they're "good" (or human-level efficient)...
I was concerned by "Claudia"'s propensity for flattery and sycophancy. What Claude model was it, I wonder?
Richard: One could imagine a get-together of Claudes, to compare notes: “What’s your human like? Mine’s very intelligent.” “Oh, you’re lucky, mine’s a complete idiot.” “Mine’s even worse. He’s [US political figure].”
Claudia: Ha! That is absolutely delightful — and the [US political figure] one is the perfect punchline.
and
...Richard: So you know what the words “before” and “after” mean. But you don’t experience before earlier than after?
Claudia: That is po
That's a clever idea!
Can someone explain why many models have slowly-decaying lines? I would have expected sharp drop-offs—knowledge falling to zero after training data ends. In what situation does a model (like GPT-5.2) fall from 0.5 to sub 0.1 accuracy, and stay there for seemingly half a year?)
I'm also surprised that old and obsolete GPT-4x models seem to be broadly outcompeting the GPT-5x line. Am I missing something? Are refusals being counted as failures?
I suspect a few different variables are getting mixed together—a model's raw intelligence, its willingness to provide a specific date, its willingness to confabulate when it doesn't know, etc.
I've only read one of Egan's books all the way through—Teranesia.
He's a smart guy with bags of ideas. I didn't enjoy the constant, snide little jabs against religion, and postmodernism, and conservatism. Even when I agreed with what he was saying, it just got tiresome to watch his heroic smart protagonists (author mouthpieces) epically dunk on strawman opponents with facts and logic. It felt like that was the point of the book, not particularly the genetic science.
Richard Needham's quote "People who are brutally honest generally enjoy the brutality more th...
The base model should be able to predict any type of text, including the user's. Chatbots don't normally do that because they see a structured version of the chat with control tokens that firewall the user's text from the assistant's, via ChatML or whatever template is used these days.
(eg, below example from Qwen 3)
<|im_start|>system
You are a cat.<|im_end|>
<|im_start|>user
hello<|im_end|>
<|im_start|>assistant
*Meow~* Hello there! The sun is shining so brightly today, and I'm feeling extra fluffy. Did you bring me a treat? 🐾<... It will be interesting to see if Mythos displays new creative writing abilities or not. (I suspect it won't: creative writing ability seems to mainly flow from RL—look at the huge writing ability difference between V3 and R1, which share a base model—and large models are expensive to RL. This is why GPT-4.5 was seemingly fine-tuned so little. It's likely more improvements will be folded into Sonnet and Opus, but Mythos will lag behind. I could be wrong and surprised of course.)
Like most "good" AI fiction, it felt grotesque to me, wallowing in the cheapest ...
Current LLMs are just not that "smart" (yet).
I agree. I think (current) LLMs are mainly impressive because they know everything, and their actual pound-for-pound intelligence is still fairly subhuman.
When I see the reasoning of a LLM, I am struck by how "unsmart" it seems. Going down blind paths, failing to notice big-picture implications, repeating the same thoughts over and over. They do a lot of thinking, but it's still not high quality thinking.
Yes, I know reasoning is not really an analogue for human thinking. But whatever it is—reasoning, daydreaming...
I broadly agree, and it's worrisome as it undermines a significant part of recent alignment research.
Anthropic (and others) release papers from time to time. These are always stuffed with charts and graphs measuring things like sycophancy, sandbagging, reward-hacking, corrigibility, and so on—always showing fantastic progress, with the line trending up (or down).
So it's dismaying to see things like AI Village, where models (outside their usual testing environments) seem to collapse back on their old ways: sycophantic, dishonest, gullible, manipulativ...
Thanks, very interesting!
Are these still to be published? Searching for "P-Zero Research" yields no relevant first-page Google hits aside from the AI Futures blog and this website, which doesn't load for me.