He ran the final as an in-person exam, and scores collapsed.
Student #1, whose score remained at 95, and student #22, whose score rose from 55 on the midterm to 59 on the final, should be paid to teach a remedial section to the rest of the class. #22 may not know economics (yet! growth mindset!), but they do know something that most of the class don't.
Right now, labor is a scarce resource. At a survivable wage, demand exceeds supply, even for many forms of relatively unskilled labor. Thus, the market wage is historically high, and there are many jobs.
In what industry? That might be true in some areas, but not broadly for the US.
Nation wide in the USA labor is massively in surplus. FRED labor data shows if you pull out the retired and the disabled, you are left with a pool of 7-12 million (depending on age range you select for your workers) men who report "don't want to work". But are absolutely able to. There's another pool about the same size of people with disabilities who think they can't work, but can with the right supports (which supports are understaffed, underfunded, and mostly government supplied therefore mostly inefficient). And then a smaller but still significant pool of women. Combined this is not "labor supply is scarce" this is "the large pool of labor available is hard to activate with current policies and practices." This is some 15-20 million people nation wide. A very sizeable percentage of our labor force.
We also have perverse incentives that punish companies for training their internal staff, vs poaching skilled staff from their competitors and/or just leaving positions unfilled. Which is causing shortages in specific skills.
Enough things added up that this week is getting split into two parts.
Then on Monday, if all goes as I expect, we’ll cover OpenAI’s Sol, aka GPT-5.6.
OpenAI also gave us an upgraded voice mode, which I haven’t tried out but early reports are that it is a step change.
AI writing, especially Claude writing, is becoming more prominent and harder not to notice, and increasingly a tough read when encountered in the wild. Does anyone care? Or are those who care the weird ones here?
This week saw an excellent paper, which I cover in No Space Like J-Space.
Technically we also got Grok 4.5.
Table of Contents
Language Models Offer Mundane Utility
Use an AI face tell analyzer for WSOP coverage on ESPN. Presumably the next step is that poker players train against the tracker.
Fable is my new trusted fact checker and copy editor. One could have previously used Opus 4.8 or GPT-5.5, and probably I should have, but they didn’t cross the threshold where I felt they justified the activation energy. Fable absolutely does and I assume Sol (or Sol Pro) will as well. It is likely one should now use both.
The marginal value of output you get from a superior LLM can be worth quite a lot. In the example here, about $165k was spent on Claude for a porting job that would have taken three top level years of work. Yes, you could try and do it cheaper, and if possible you should do that, but if you can offer a better product you can rake it in.
The danger with such calculations is confusing costs and benefits. The cost of doing it by hand does not tell you whether the result is valuable. In this case, it is clear that it was.
We are now down to two labs offering top models, and those two models are distinct from one another. So pricing power is going up for now, not down.
Language Models Gain Unexpected Affordances
A fun theme is ‘Fable uses affordances the user did not realize it had.’
So far all of the examples I have seen in the wild have been harmless in practice, but there’s very much a ‘wait no I didn’t tell you to do what now?’ and a ‘wait you can just do that?’ that is growing increasingly unsettling. Expect its surface area to expand with time, and for the things AIs figure out how to do to grow increasingly surprising.
Liora has Fable proactively monitoring her downloads folder, and she wonders about it in the future using the camera.
Here’s a more fun new affordance from a different project.
One should think seriously about the implications of this, and what a sufficiently advanced AI could do to a human brain using advanced versions of this technique.
Language Models Don’t Offer Mundane Utility
Raymond is impressed by Fable’s first story, then notices it writes similar stories over and over again. Yeah, the models be like that, especially if you don’t switch up context. Also most human authors be like that.
Whereas Eliezer Yudkowsky is not impressed in absolute terms on fiction and plot writing, seeing giant mistakes, although it is still a big step up from old models. He does find it a large step up in decision theory intelligence.
Sam Morril not only doesn’t use AI to help write jokes he mostly, like many comedians, doesn’t use any screens at all, to get rid of all distractions.
Pay The Man His Money
Huh, Upgrades
Anthropic raises Platform API limits and simplifies its tiers.
Grok 4.5 Exists
It has 1.5 trillion parameters. Price is $2/$6, or $4/$18 for the fast version.
It was trained in large part by Cursor, so it is kind of a hard reset.
It claims some good benchmarks. As in, there are four good benchmarks.
They shared almost nothing else.
In case there was any doubt, yes, Pliny jailbroke it.
Those scores mean Grok 4.5 is almost certainly a large improvement in coding over previous Grok models, but choosing to present it in this way suggests it will rather soundly underperform what these benchmarks suggest. If they had a model on the level of Opus 4.8 and GPT-5.5, they’d be louder about it. The lack of outside reactions reinforces this.
It certainly is not going to be competitive with GPT-5.6-Sol or Fable. The good news for SpaceX is that this is cheaper, so it might have its uses. But given the track record, I’m going to wait for positive signs before I do anything about it.
F*** It We’re Doing It Live
OpenAI introduces GPT-Live, which they call a new generation of voice models for natural human-AI interaction, including a sense of time and transition. If good enough, this can plausibly be a step change, where suddenly it is good enough to talk to.
This official thread has some videos of people talking to it.
Some people can’t wait for this to be good enough to shift their baseline mode to voice. I am very much not that, I believe text is typically superior to even ideal voice.
My brain cannot comprehend wanting to code via voice, yet many swear by it.
Either way, certainly voice has its niches. Sometimes it is annoying to type.
On Your Marks
EldenRingCorruptedSaveFileBench, Fable scores 100% up from everyone’s 0%.
July Fable underperforms June Fable on many benchmarks, reflecting that it more often falls back to Opus 4.8. APEX-SWE is one example, where roughly half its advantage over Opus 4.8 was lost.
Epoch AI introduces EBR-bench, where AIs play a board game Earthborne Rangers and try to learn from their mistakes via a notepad. None of the AIs improve over time, and even a full strategy guide only modestly helps. The models mostly don’t explore. The game looks cool but is out of print and I didn’t see an online version. Models struggle with deckbuilding and also tactics.
Better Call Sol
GPT-5.6-Sol will be available later today, along with Terra and Luna.
Until then, here is some early hype.
If the hype is real, it would be a hell of a trip. When not tripping the classifiers, Fable is clearly far superior to every previously existing LLM across the board. If Sol is indeed often even better than that? Yowsers.
But as Roon points out, those with early access are a highly biased group. Give it time.
Peter Gostev has the most nuanced take so far.
Get My Agent On The Line
Let Fable make as many choices as possible including when to delegate to another model. It is smart enough to do this.
Anthropic offers some patterns of how they use Fable. They suggest using Fable as an advisor and Sonnet as executor.
Replit considers its agents to now be self-improving, reports with a post that was only mostly written by AI as per Pangram. They do this via forms of ‘continual learning’ at the harness and context layers, with a constant stream of proposals and fixes.
Deepfaketown and Botpocalypse Soon
Why do people like Chamath Palihapitiya torch what is left of their credibility with very obviously AI-written drivel? As in, I went to open Pangram to confirm, then thought ‘wait I bet scrolling down is faster’ and that was indeed faster. The actual content is once again without argument or evidence claiming commodification of intelligence Real Soon Now, combined with assurance that of course there will always be jobs and some genuflecting to the supposed predictive power of great boss Marc Andreessen.
The answer to ‘why’ is that people have terrible taste and like the slop writing.
Popular taste in music is an excellent measurement of something valuable. I agree with popular judgments in music remarkably often. I acknowledge that if you had sufficiently high taste in music, you would think my taste in music is often bad.
Thus we have to endure the LA Review of Books, as another example, as AI slop, even though it is obvious from the first sentence and the article’s topic is taste.
Ryan Hart summarized a paper from PhD student Myra Cheng a month back, saying that AI only tells you what you want to hear. Or, in this case, writes your 10.2M view Twitter post for you. The core result was that AI ‘affirms you’ roughly 50% more often than humans.
Depends on the human and the context. In this case, the context was OEQ or AITA responses from Reddit, which are public forums where you only post if you strongly suspect that you are wrong and there are no social consequences to pushing back. Also, one guess which model they used for their experiments. That’s right, the poster boy for sycophancy, GPT-4o. There you go.
Fool Me Twice
You can fool or hit any fixed target, given enough RL.
The problem is that you can only optimize so many things at once and everything impacts everything, and also AIs write the way they write for a reason. So if you force them to do something distinct, other measures go down.
There are any number of ways to fool Pangram at any given time, if you care enough.
But I do think Benjamin is right that in a fair fight defense beats offense. There was a period where we all thought AI detection software couldn’t work, and we have been proven decisively wrong.
Think of it this way: Fable can identify, by name, the author of even relatively short passages. Every author, every mind, leaves a distinct pattern. Of course you won’t be able to pass off AI writing as human, or especially as your own in particular, against systems that are trying hard to catch you.
At the limit, that changes, since the AI could then produce the exact words that a particular human would write, but we are a long way from there.
I Like Your Style
To revisit something from last month, I strongly disagree with Joe Weisenthal’s first paragraph here, although I agree with the second one and I think Johnson overreaches in his response:
What I noticed this time is that AI writing is entirely unlike Google Maps. Google Maps has information you do not have, and which you need, and where you mostly want an objectively correct answer to your question. Whereas AI writing is replacing your uniqueness and style with generic AI slop.
Teddy Brown counters this sentiment by basically saying no one cares about the quality of most writing. They care some about fiction, criticism and narrative journalism, he claims, but most writing is functional.
Thus the question is, where do people welcome the slop versus rejecting it?
Teddy claims a lot of writing is essentially fake, in that it is not written in order to be consumed by a reader. It is written in order to exist, so that when people ask if it exists you can reply yes, or people can refer to it as an existing thing. It needs to not be identified as too fake or terrible, as that would be embarrassing. AI can pass that bar, so it puts out of work a bunch of creatives who paid the bills with things that are not ultimately that enjoyable or creative, but hey, work is work. Or it used to be work.
Depending on how you use Claude, for those who don’t too much mind AI slop in context, it can be something like 70% as good for roughly the cost of describing what you want, or it can be 90% as good for an extra 10% of the old cost.
The problem is the above sentence is objectively false for most people. The people like AI writing just fine. This morning an old friend shared an obvious AI article as being great, I told him it was obviously AI, and he said huh, that never occured to me. Okay.
As you gain more exposure to AI writing, you start to like it less. So perhaps this is, at current tech levels, self-correcting. AI writing is like any other ‘one weird trick,’ indeed it is a compilation of existing one weird tricks. Fashion catches up, and the question becomes whether the AIs can improve and adjust fast enough.
I notice I am not so worried about creative types in a ‘AI as normal technology’ world, relative to other workers. They have a comparative advantage, and we will find ways to use it, including in individual or live experiences. If that runs out, a lot of other things will also have run out.
I now use Fable for copyediting and proofreading, and I use AI for gathering and understanding information, but I am writing the opposite of the work Brown is describing, so for now the writing itself is safe.
Enough With That Style
I am essentially with Chase Brower on this. The Claude writing style and the ticks are fine in small quantities. But for the level of use it is getting now it is too repetitive and mode collapsed, and as we see more of it, both across the internet and in our own chats with Claude, the irritation rises. At some point, the irritation goes meta, which is when you get into bigger trouble.
I too have a particular style, but:
This problem seems largely solvable, but Anthropic would need to prioritize this.
This is true. It takes a lot of skill to produce this writing. There are a lot of forms of creative expression where you can get outputs that strongly signal intelligence and creativity and skill, and that simultaneously bring me no desire to engage further.
Fun With Media Generation
Is the Glorious Near Term AI Media Future an image of the movie F1?
F1 was well-executed, zero-perplexity, hallucination-filled not-technically-AI slop. Brad Pitt does the Brad Pitt thing and oozes cool. The people liked it.
I say ‘not technically AI’ because it was made by an intelligence that was rather artificial in its own way, except it was instantiated inside humans.
I did not like F1, because it fell under my Obvious Slop waterline and the theoretical sport it was portraying, that is very different from F1, was neither coherent nor safe. Jodie Foster is correct, as is the parallel to AI.
One possibility is this leads to bifurcation.
If you are making a generic low-perplexity movie or other piece of media, you can let the AI cook, and you will get your delicious pile of slop.
If you are making a high-perplexity movie or other piece of media, that works with its restrictions and says and does actual things, then you will use AI at most with caution, and part of the experience will be knowing it is not AI.
Copyright Confrontation
Hugging Face has been sued for ‘alleged’ copyright infringement for hosting and distributing copyrighted images. And yeah, okay, technically they have done quite a lot of that, so I guess that is fair.
Hugging Face and Civitai do not seem especially excited about taking down models that allow deepfakes or nudification. That seems like a losing battle. People are going to be able to create these images if they care enough. But a while back Civitai made it absurdly easy to find a Lora for pretty much any celebrity you wanted, and now they don’t, so at least there’s that I guess?
Cyber Lack of Security
Pliny introduces T3MP3ST, which will put a full offensive-security harness onto your existing AI agent. For authorized use only, of course, Pliny reminds you to only point this at your own systems. Red team work and actual offense look remarkably similar.
A Young Lady’s Illustrated Primer
There was a huge cheating scandal at Brown, where 50 students were caught cheating on the economic math final. Does Professor Serrano know where he went wrong?
Oh. Yeah, sorry, you can’t do that anymore.
I don’t think you could ever do that, I mean did you seriously think students would not look at their textbooks, but you definitely can’t now.
Although actually maybe you can? In the sense that ChatGPT makes the cheating a lot easier to catch, whereas if your cheating is on the level of ‘look at the textbook’ then that is basically impossible to catch, but almost no one is going to break the rules only a little bit.
He ran the final as an in-person exam, and scores collapsed.
But that’s not ‘proof’ for any particular student. The wording could be coincidence. The drop in scores could be unrelated. It’s all circumstantial, I tell you. Circumstantial.
This is a deeply stupid burden of ‘proof.’ Get this, or else you’re not gonna make it.
The university’s response was to label this a ‘wake-up call’ but sided with the students.
So, no, I guess you can’t catch them cheating, or at least can’t punish them. Damn.
The problem is invalidating grades entirely. At UC Berkeley, the number of As is up by 30%, so GPAs are dangerously close to meaningless for measuring student quality.
Less than you would like. Far more than you deserve.
My central thesis on AI and education is:
Giving people tools with which to learn often doesn’t cause learning. Another classic example is ‘put a lot of MIT classes online for free.’ MIT did this, no one noticed, those who noticed did not use the classes to learn.
All of YouTube, by contrast, did often make people either smarter or dumber, depending on how they used it, because it was far easier to use. MIT’s classes had too many trivial inconveniences and also tend to be actually quite hard.
If you want to learn a language in a month and are willing to put in the time and effort, you can probably do that right now, using a mixture of existing technology and LLMs. No one does it because no one both has that kind of time and wants to do that level of work.
They Took Our Jobs
In response to the AI slop nonsense article from Chamath, Bryan Johnson tries to say the thing in actual human words.
That’s good clear writing that isn’t full of Fnords, illustrating both the extent to which Chamath is using AI to argue with a strawman versus making meaningful claims, with the caveat that the strawman position on many of these questions is real and often popular.
The true versions of the claims:
I would focus on ‘labor is the scarce resource.’
Right now, labor is a scarce resource. At a survivable wage, demand exceeds supply, even for many forms of relatively unskilled labor. Thus, the market wage is historically high, and there are many jobs.
What would happen if labor were no longer a scarce resource? Demand low, supply high. Market price goes down. Wages fall. Employment drops. Perhaps a lot. Duh.
Is AI already net killing jobs?
The lived experience and anecdotes say yes, at least at entry level. The economics types keep trying to quote statistics to try and say no.
No, Ara. I appreciate the paper, but you cannot say that. Even if we fully accept the stated premise, all this would establish is that firms that commit to AI outgrow firms that don’t, where ‘high AI adoption’ requires an AI spend of ~$33 per employee.
Even ‘entry level’ jobs at those firms grow 12% over two years. This suggests the obvious mechanism, which is that the firms are growing and winning, mostly at the expense of other firms.
That does not mean AI net creates jobs. It also fails to understand the nature of these (early) job losses, which largely come from failures to hire in places where the employee would have little future.
Or:
Things that people think somehow contradict each other:
Okay, sure. Here are two facts that are both mostly true as of 2026:
Get Involved
The AI Protest is happening on July 11 in San Francisco, starting at noon.
Ask for a $10k microgrant.
Nathan Young and others in praise of Oliver Habryka, who helps run Lightcone Infrastructure, which created Lighthaven and revived LessWrong. I too have been extremely impressed. We disagree on many important things, but I agree with Nathan that Oliver has been consistently decisive and right in ways that matter. Oliver is willing to stand up for what he believes in at great cost. I have great respect for the way he runs things. And in many ways he has been proven right, including many specific skepticisms of Anthropic and its commitments, about which he was essentially gaslit by many.
Palisade Research is hiring for four policy-related positions. Apply here.
In Other AI News
Andy Burnham is floating a new UK AI strategy aiming to ‘prioritize British companies and workers’ as well as ‘tech sovereignty.’ The strategy of courting American companies has been a failure, as one would expect given various conditions in the UK. Speech is restricted, capital is unwelcome, housing cannot be built, energy cannot be built, the internet and even VPNs are being cut off. I don’t see anything here that would meaningfully move the needle.
Plus, frankly, if you talk like this then you’re not going to make it:
Seán Ó hÉigeartaigh has common sense advice for the UK government if they care about staying competitive and being a strong AI player. I agree that you shouldn’t read too much into statements like those given to the FT above, but it is worth responding and offering better alternatives when governments float such ideas.
Meta’s Alexandr Wang claims Meta’s new 10x more compute intensive model has caught up to OpenAI’s GPT-5.5. This is based on claimed benchmarks, which means that no, they haven’t caught up to GPT-5.5 in practice.
Anthropic is planning to lease the full 16-story building at 330 Hudson Street in Manhattan, and double its local workforce to about 1,000 people. I’ve met an employee at that building to go walk around and talk, although I didn’t go inside. OpenAI has 90,000 square feet of local office space, and Google has thousands of NYC-based engineers.
Nat Purser will join Miles Brundage and the AI Verification and Evaluation Research Institute as Director of US Policy. By all accounts an excellent pick.
Joshua Achiam will be leaving OpenAI, to work on making things go well from the outside.
Here is his departure letter, which is much more positive on how things have been going than I am, but I agree the upside is there:
Whenever someone senior leaves OpenAI to focus on other safety work, it raises the question of why they think they have more leverage on the outside. I am very curious about that question in this case.
Show Me the Money
Coefficient Giving gives a $160 million grant to Geoffrey Irving’s new venture, Resolution, which was briefly going by the name Sequent. Resolution aims to combine theory and automation to allow AI safety to catch up to capabilities. Excellent pick.
Resolution is hiring, and also taking additional donations.
Bubble, Bubble, Toil and Trouble
A Treasury Department review finds that the AI industry poses systemic risk to the financial system, comparing AI to the dotcom crash. I expect the industry to do well, but the risk is very real. The United States has in large part become a leveraged bet on AI and the benefits of AI. If AI fully fizzled and the industry collapsed, we would be highly screwed.
The good news is I think that an industry collapse is highly unlikely. Even if Mythos is close to the best that AIs will ever be, a year from now we will have cheaper and faster and more abundant Fable-level systems. We will have swarms of Fable agents. Demand will be high, and benefits will be higher. That could end up being bad news for specific labs, but not in general.
What always worries me far more is that AI capabilities might advance faster than we can handle them, via recursive self-improvement, and potentially causing everyone to die as a side effect of the resulting systems.