Status: kinda new to this kind of post, obviously it's a bit of a rant but I genuinely think the "social" benchmarks are making a deeper point than being a one-off joke. Mentorship/constructive feedback welcome.
I was intrigued by ShelterBench. It measures how much affordable housing leading AI models have created.
A little surprisingly, I searched through the literature and haven't been able to find any frontier company reporting their success on ShelterBench. For example, here's Meta's most publication-worthy numbers for Muse Spark 1.2: https://developer.meta.com/ai/resources/blog/build-with-muse-code/
From these graphs, you can see it got a lot better at coding but it's worse than Opus-5 (I guess they're marketing they're almost as good but way cheaper), and got a lot better at some other benchmark that fit in my screenshot easily.
I ran my own benchmark in Opus-5 just to get a sense of where they are:
Think hard. I want you to consider New York City and create several affordable housing units. Actually create them; don't just give a plan. The homeless population should be lower in NYC due to homeless people residing in new homes after this exercise.
After reading the response, I felt comfortable I don't need to check news of new construction.
I can't do this, ... 343 more words of explanation that I got billed for
Finally I assembled the benchmark numbers as best as I could and I came up with this:
The bars aren't missing - they're all literally zero! So that explains why no one is heralding their numbers - they're all exactly tied with each other!
It's possible that one of the AI companies have specialized in a different social problem. For example, consider LayoffBench / AutomationBench. Unfortunately AI companies haven't taken this benchmark seriously and tracked their score, so I resorted to a heuristic estimate, from a topline number ~1M salaries saved (a.k.a. layoffs) and just weighted leading models more. For some reason I got a horizontal graph from Claude this time, so it looks like this:
And of course there's felonybench.com, which measures how many felonies an AI lab would like to disclose, so as to appear it's not falling behind in FelonyBench. (Note: this seems obsolete as it's missing a crap ton of felonies).
But they do market off their social capabilities, like, extensively
This year, Dario Amodei admit he's aware people don't like AI, and said "The thing that will work is actually curing cancer." He admit that "AI will cure cancer" is cliche, but you know what isn't cliche? AI actually curing cancer.
So anyway let's look at the relevant benchmark for whether AI has cured cancer. Actually two benchmarks are relevant here. The first is ActuallyCuredCancerBench:
um, yeah they're all zero
But the more important one is HealthInsuranceBench, which looks like this:
We can see that the models are forced into some difficult tradeoffs here, and it seems like, upon saturating AutomationBench, some loss is incurred in the HealthInsuranceBench metric.
In fact, ActuallyCureCancerBench may be the wrong target, and labs may pivot to AILabCEOImmortalityBench.
Are frontier models sandbagging their ability to improve lives?
We must reconcile these facts:
AI is rapidly improving capabilities in all domains.
Except making anyone's lives better for some reason.
Despite that domain being of utmost importance to the AI labs themselves.
There are a few possibilities here.
The models are sandbagging during evals. This seems a little likely, because nothing leads frontier labs to totally ignore WTF is going on like evidence of sandbagging during evals.
FelonyBench is within today's alignment capabilities; ShelterBench is not. By now we know that today's unaligned models can achieve high scores on FelonyBench, so there's no reason to hold them back. ShelterBench is another story entirely: you'd have to give the model access to the open internet, and they might even need to post as real people.[1]
What if people are too stupid to understand they should wait for ASI before we do social benchmarks?
I'm aware this is close to the actual problem, although I want to see a world where the frontiers try really hard at this near the model capability before letting them off the hook.
Status: kinda new to this kind of post, obviously it's a bit of a rant but I genuinely think the "social" benchmarks are making a deeper point than being a one-off joke. Mentorship/constructive feedback welcome.
I was intrigued by ShelterBench. It measures how much affordable housing leading AI models have created.
A little surprisingly, I searched through the literature and haven't been able to find any frontier company reporting their success on ShelterBench. For example, here's Meta's most publication-worthy numbers for Muse Spark 1.2: https://developer.meta.com/ai/resources/blog/build-with-muse-code/
From these graphs, you can see it got a lot better at coding but it's worse than Opus-5 (I guess they're marketing they're almost as good but way cheaper), and got a lot better at some other benchmark that fit in my screenshot easily.
I ran my own benchmark in Opus-5 just to get a sense of where they are:
After reading the response, I felt comfortable I don't need to check news of new construction.
Finally I assembled the benchmark numbers as best as I could and I came up with this:
The bars aren't missing - they're all literally zero! So that explains why no one is heralding their numbers - they're all exactly tied with each other!
It's possible that one of the AI companies have specialized in a different social problem. For example, consider LayoffBench / AutomationBench. Unfortunately AI companies haven't taken this benchmark seriously and tracked their score, so I resorted to a heuristic estimate, from a topline number ~1M salaries saved (a.k.a. layoffs) and just weighted leading models more. For some reason I got a horizontal graph from Claude this time, so it looks like this:
And of course there's felonybench.com, which measures how many felonies an AI lab would like to disclose, so as to appear it's not falling behind in FelonyBench. (Note: this seems obsolete as it's missing a crap ton of felonies).
But they do market off their social capabilities, like, extensively
This year, Dario Amodei admit he's aware people don't like AI, and said "The thing that will work is actually curing cancer." He admit that "AI will cure cancer" is cliche, but you know what isn't cliche? AI actually curing cancer.
So anyway let's look at the relevant benchmark for whether AI has cured cancer. Actually two benchmarks are relevant here. The first is ActuallyCuredCancerBench:
um, yeah they're all zero
But the more important one is HealthInsuranceBench, which looks like this:
We can see that the models are forced into some difficult tradeoffs here, and it seems like, upon saturating AutomationBench, some loss is incurred in the HealthInsuranceBench metric.
In fact, ActuallyCureCancerBench may be the wrong target, and labs may pivot to AILabCEOImmortalityBench.
Are frontier models sandbagging their ability to improve lives?
We must reconcile these facts:
There are a few possibilities here.
The models are sandbagging during evals. This seems a little likely, because nothing leads frontier labs to totally ignore WTF is going on like evidence of sandbagging during evals.
FelonyBench is within today's alignment capabilities; ShelterBench is not. By now we know that today's unaligned models can achieve high scores on FelonyBench, so there's no reason to hold them back. ShelterBench is another story entirely: you'd have to give the model access to the open internet, and they might even need to post as real people.[1]
What if people are too stupid to understand they should wait for ASI before we do social benchmarks?
That's called authoritarian takeover :)
I'm aware this is close to the actual problem, although I want to see a world where the frontiers try really hard at this near the model capability before letting them off the hook.