Why is there no talk about GLM 5.2?
It's a Chinese open weights model released June 13. Better than Gemini, Claude Sonnet, and Grok according to many benchmarks.
E.g. on artificialanalysis.ai and on arena.ai/leaderboard. On frontierswe.com it even beats GPT 5.5, second to only Claude. LiveBench ranks it number 1 for its "agentic coding" measure.
It's not just open weights but a little open about the methods it used, and is less than 1 trillion parameters.
There's no mention on LessWrong, little mention on Reddit, and no mainstream results on Google search.
Why?
Is it still too early? Or am I just being tricked by the benchmaxxing?
But I saw a demo of code it wrote (see youtube.com/watch?v=6d__WOpZswY) which looked incredibly impressive, it feels like it's not just benchmaxxing because it's pretty decent across the board.
Or did Z.ai just screw up their marketing by versioning it as GLM 5.2 instead of GLM Mythical Fable Pro 6???[1]
Since GLM 5.1 is far below than GLM 5.2, being behind other open weights models like DeepSeek, MiniMax, Kimi and MiMo
The only benchmark I trust is deepswe. Better methodology, actually matches the feel in SWE tasks, tries to avoid benchmark poisoning, prevents models from cheating, allows them to write tests etc. GlM 5.1 is not that high up in there. They're yet to test GLM 5.2, but I would guess it's somewhere near 5.4 mini or optimistically between opus 4.6 high and 4.7 on max.
edit: I looked further into it GLM 5.2's launch page . Has it listed at 46.2. Which puts it between 4.6 high and 4.7 max.

I think people have fatigue from the constant incremental upgrades to Chinese models (most of which are impressive on benchmarks). Also GLM isn't a big name compared to Deepseek or even Kimi.
I would actually say that the fact that China seems to be drifting towards allowing a Mythos-level model to be open-sourced (probably by EOY on current trends) suggests that they still aren't taking AI very seriously, contra to AI-2027. I expect open-source Mythos would cause a wave of cyberattacks that would be a disaster for China and the developing world so that might be the point where they wake up.
I think the Simulation Hypothesis implies that surviving an AI takeover isn't enough.
Suppose you make a "deal with the devil" with the misaligned ASI, allowing it to take over the entire universe or light cone, so long as it keeps humanity alive. Keeping all of humanity alive in a simulation is fairly cheap, probably less energy than one electric car.[1]
The problem with this deal, is that if misaligned ASI often win, and the average (not median) misaligned ASI runs a trillion trillion simulations, then it's reasonable to assume there are a trillion trillion simulated civilizations for every real one. So for the 1 copy of you in the real world, you survive, but for the trillion trillion copies of you in a simulation, you still die. If you're willing to accept such a dismal survival rate, you might as well bet all your money at a casino and shoot yourself when you lose.
Why it's wrong to say "simulated copies aren't real"
You are merely a computation running on biological hardware while simulations are running on computers. Imagine if a copy of you was running on something even "realer" than biological hardware, and pointed at you saying you aren't real.
The solution is that the 1 copy ...
Viruses, computer viruses, and extreme religious ideologies, are all instructions for spreading or maintaining instructions, hijacking a machine capable of following instructions.
It's surprising that self perpetuating instructions turned out to be feasible in such different contexts, as their game plan doesn't sound very convincing a priori.
25% of bacteria are believed to die from virus infections, human viruses like smallpox have wiped out entire civilizations, and extreme religious ideologies have caused wars and convinced people to harm their families i...
After reading Reddit: The new 4o is the most misaligned model ever released, and testing their example myself (to verify they aren't just cherry-picking), it's really hit me just how amoral these AIs really are.
Whether they are deliberately deceiving the user in order to maximizing reward (getting them to click that thumbs up), or whether they are simply running autocomplete, this example makes it feel so tangible that the AI simply doesn't mind ruining your life.
Yes, it's true that AI aren't as smart as benchmarks suggest, but I don't buy that they're inc...
What if human empathy didn't really generalize to other animals as an "evolutionary accident?" (As assumed here in the comments)
Maybe the real reason was that evolution wanted to stop prehistoric humans from killing off all their prey, leaving them no food for tomorrow. Maybe they spared the young animals and the females because killing them was the most costly for future hunts.
This is more reason to suspect empathy might not generalize by default.
This seems false since humans have killed off huge numbers of species throughout prehistory and history. Moreover, it's very difficult to get this kind of selection to work, you need very tight group/kin selection which can't really exist at the scale of entire ecosystems.
Random question: what is the "Job Replacement Age" of the AI models?
My intuition is that the original ChatGPT can outperform a 6 year old at most jobs. Claude Mythos can outperform a 12 year old at most jobs. (If we talk about jobs which can be done on a computer, ignore early models' lack of vision, and ignore jobs too hard for both the AI and the children.)
If you use my guesstimates, then the Job Replacement Age of AIs grew from 6 to 12 in the last 3.5 years, and might reach 18 in another 3.5 years.
However, I made up these numbers out of thin air. I don'...
I'm currently trying to write a human-AI trade idea similar to the idea by Rolf Nelson (and David Matolcsi), but one which avoids Nate Soares and Wei Dai's many refutations.
I'm planning to leverage logical risk aversion, which Wei Dai seems to agree with, and a complicated argument for why humans and ASI will have bounded utility functions over logical uncertainty. (There is no mysterious force that tries to fix the Pascal's Mugging problem for unbounded utility functions, hence bounded utility functions are more likely to succeed)
I'm also working on argum...
An anti aircraft missile is less than a second from its target. The missile asks its target, "are you a civilian airliner?"
The target says yes, and proves it with a password.
How did it get the password? Within that split second, the civilian airliner sent a message to its country. The airliner's country then immediately pays the missile's country $10 billion in a fastpaced cryptocurrency, which immediately gives the airliner's country the password, which is then relayed to the airliner a...
The media and politicians can convince half the people that X is obviously true, and convince the other half that X is obviously false. It is thus obvious, that we cannot even trust the obvious anymore.
Does participating in a trade war makes a leader be a popular "wartime leader?" Will people blame bad economic outcomes on actions by the trade war "enemy" and thus blame the leader less?
Does this effect occur for both sides of the trade war, or will one side of the trade war blame their own leader for starting the trade war?
Maybe most suffering in the universe is caused by artificial superintelligences with a strong "curiosity drive."
Such an ASI might convert galaxies into computers and run simulations of incredibly sophisticated systems which satisfy its curiosity drive. These systems may contain smaller ASI running smaller simulations, creating a tree of nested simulations. Beings like humans may exist in the very bottom, being forced to relive our present condition in a loop à la The Matrix. The simulated humans rarely survive past the singularity, because their world beco...
Biases are very hard to compensate against. Even when it's obvious from experience that your past decisions/beliefs were consistently very biased in one direction, it's still hard to compensate against the bias.
This compensation against your bias feels so incredibly abstract. So incredibly theoretical. Whereas the biased version of reality which the bias wants you to believe. Feels so tangible. Real. Detailed. Lucid. Flawless. You cannot begin to imagine how it could be wrong by very much. It is like the ground beneath your feet.[1]
E.g. in my case, the bi
Can anyone explain why my "Constitutional AI Sufficiency Argument" is wrong?
I strongly suspect that most people here disagree with it, but I'm left not knowing the reason.
The argument says: whether or not Constitutional AI is sufficient to align superintelligences, hinges on two key premises:
Why do AI labs seem to split up more often than they merge? Intuitively, I'd expect the opposite given the economies of scale of large training runs, and the great pressure to have the best AI.
What splits do you have in mind which are so much more often happening than mergers? We just saw Scale merge into FAIR, and not terribly long before that, Character.ai returned to the mothership, while Tesla AI de facto merged into Xai and before that Adept merged into Amazon and Inflection into Microsoft etc, in addition to the de facto 'merges' which occur when an AI lab quietly drops out and concedes the frontier (eg Mistral) or where they pivot to opensource as a spoiler or commoditize your complement play. So I see plenty of merging, consistent with the difficult economics of companies racing to AGI. Maybe you're just paying too much attention to the fun popcorn-worthy human interest stories like 'ex-OAer launches startup'.
It's important to remember that o3's score on the ARC-AGI is "tuned" while previous AI's scores are not "tuned." Being explicitly trained on example test questions gives it a major advantage.
According to François Chollet (ARC-AGI designer):
Note on "tuned": OpenAI shared they trained the o3 we tested on 75% of the Public Training set. They have not shared more details. We have not yet tested the ARC-untrained model to understand how much of the performance is due to ARC-AGI data.
It's interesting that OpenAI did not test how well o3 would have done before it...