I read the ARC-AGI-3 paper entirely, and I'm unimpressed.
The "100% human-solvable, <1% AI solved" is basically p-hacking. They cook their metrics to guarantee high human scores and punish any sub-human score. They also prevent measurement of super-human performance, so in practice it's close to a binary metric of "matches best human or not".
There are also a number of incoherences in the stated methodology, but they're non-central.
Their metric is:
Environment must be solved by at least 2/10 humans. Among the successes, pick the median (¤) of actions taken, that's the baseline (per level of the environment), call it b.
Humans are defined as 100% for being the baseline (no analysis of how many humans solve the environment, or whether the average score is 100% or any deeper analysis of human performance).
An environment has n levels. Levels are attempted sequentially, in increasing order of difficulty; solving one unlocks the next one. The environment is solved if all levels are completed.
If a model doesn't solve a level, it scores 0 on that level (and subsequent ones). If it does solve it in m steps, it receives (b/m)² score. (*)
Then take the weighted average of its scores over levels, where level k is weighted k.
(*) If the model is better than human (m < b), its level score is clamped at 1.15, but tbh it doesn't really matter. Also, environment score is clamped at 1 for some reason.
(¤) They say "upper-median best", which doesn't make sense, and their example is the median of people who solve the environment, so I'm going with that interpretation.
There are two problems with this metric:
- Human variance. The baseline might be ultra-optimized, close to optimal, depending on the environment; it might also not. In their empirical evaluation of optimal score (probably from human performance not-first-run?), it's clear that the baseline is very noisy.
- The way it's calculated punishes sub-human performance quadratically for no reason, and upweighs the hardest levels, which means that it's most informative only when approaching human performance.
In addition, they decide to refuse any harness, even though it's obviously the next step to get better on that kind of problem (they also show that ARC-AGI-3 is saturated with a human-made harness, so maybe they just didn't want their benchmark to be obsolete before it was even out).
It's like if they refused CoT on ARC-AGI-1.
I guess solving ARC-AGI-3 will be pretty trivial as soon as a model is RLHF'd to self-harness by default as a first step to any task.
Approximating human scores from the graphs, with Claude's help, I get human average performance in the 40~80% range.
According to Bloomberg, the US and Iran are expected to sign a memorandum of understanding whose core contents are:
- return to pre-war statu quo on territory, sovereignty, strait of Hormuz, nuclear programs, etc.
- $300B reparations from US to Iran
De-paywalled full text here: https://www.bloomberg.com/news/articles/2026-06-16/read-the-14-point-draft-memorandum-between-the-us-and-iran?accessToken=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzb3VyY2UiOiJTdWJzY3JpYmVyR2lmdGVkQXJ0aWNsZSIsImlhdCI6MTc4MTY1NjY2MywiZXhwIjoxNzgyMjYxNDYzLCJhcnRpY2xlSWQiOiJUR1FUVkRUOTZPU0cwMCIsImJjb25uZWN0SWQiOiI4M0Q4RjJERjFDQzA0MDFFQTlBNjg1RjY3N0FGQURERiJ9.Bq9TVRVNXZu1ep06Y3KiLnhlcb9SQ_ZKska1ZbY8yVM&leadSource=uverify%20wall
(Courtesy of pie_flavor, from the ACXD server.)
Read Ra.
Instead of systematization and explicit legibility, Ra chooses an impression of abstract generality which, upon inspection, turns out to be zillions of ad hoc special cases.
I see Ra in my own vibe-coded software: I ask AI to systematically seek generalizations, build from sound principles and strong foundations. But I rarely check which foundations, if any, it builds from - occasionally I catch it in egregious fakery and ad-hoccing and correct it, but often the directive of "find sound principles to build from" is a vibe more than a constraint.
There is nothing Ra-like, for instance, about noticing that software is a fully general force multiplier and trying to invest in or make better software. Ra comes in when you start admiring force multipliers for no specific goal, just because they’re shiny.
I'm guilty of this, which XKCD calls premature optimization, and it's a big drain on my resources and obstacle to finishing software projects.
Anthropic just posted a new post explaining what the economic future will look like. (Technical report)
They explicitly refuse to examine takeoff dynamics and implicitly present things as "business as usual" even in their extreme scenario.
They have a probabilistic model, but don't price in existential risk at all.
This is very disappointing. I read it as normie-washing.