My expectation is that right now is the most obnoxious the safeguards will ever be, on both the bio and cyber fronts.
Probably yes on those two, but the anti-fun puritans are winning in other areas. I tried to use Fable to help design some dice games for my ttrpg, and every post got downgraded to Opus because gambling is an adult activity.
I am not comfortable with the companies willingness to censor and limit content. And I have seen nothing that suggests it will get better in the future. I am worried this is the best it will ever be in categories other than cyber and bio.
The blip is over. We have Fable back.
Here is the official letter restoring Fable, great job everyone. Notice it is addressed to Tom Brown, not to Dario Amodei.
Anthropic had to make the controls more stupid for now, but this is a big win.
The fiasco continues, at least until such time as we have a systematic regime in place for future frontier models rather than decisions being made ad hoc, by people like Lutnik and Bessent who do not know how any of this works.
The Blip
Anthropic explains its version of what happened.
Here is the timeline:
How stupid are the extra near term safeguards they had to include here? Really stupid:
The Amazon ‘jailbreak’ was ‘fix this code.’
Debugging is literally ‘fix this code.’
So I don’t know what you want Anthropic to do here. I do know Fable is coding for me.
Here is Anthropic’s basic explanation:
Alex Stamos has a thread unpacking a punch of Anthropic’s language in its announcement.
I think Stamos is overreaching with the consequences in places, especially with #2 and #7. Otherwise he’s right.
I do not expect US models to ‘get bad’ over time, only that they will get better slower, and have more area where they have rather annoying safeguards.
My expectation is that right now is the most obnoxious the safeguards will ever be, on both the bio and cyber fronts. I expect the freak-out to subside over time, and my guess is most of it surrounds Mythos in particular. You don’t need or often even want Fable for most such product offerings, and Opus or Sol will remain well ahead of Chinese alternatives.
Contra Prinz, I do not think this commits Anthropic to going through the approval process will all future releases, only releases that pose plausible risks. We tested this with Sonnet 5, where it looks like Anthropic went ahead and dropped it on its own, and no one is suggesting there was anything wrong with doing so (other than to complain that they want Sonnet to be better).
The White House Explanation
It was a little weird.
‘Alignment across the US government’ is very much a case of ‘PHRASING!’ and here presumably means interagency sign-off, not ‘the model is now aligned with the US government.’ Unclear whether he knows enough to be trolling here.
As in, before a model can be released, you now likely need this ‘alignment,’ which in practice means sign off from various potential veto points, starting with Commerce and the Pentagon. Who knows how many more fully count.
Everything Remains Ad Hoc
We know a bit more than we did when Dean Ball posted this on the evening of June 30. In particular we know that the new safeguards are that Anthropic trained its classifiers to reject additional Fable uses.
We still don’t know how the ad hoc system works more broadly. Having an opaque ad hoc system, especially one where those administering the system do not themselves know what they will do, is even worse. Again, fully winging it is the worst case scenario.
Take What You Can Get
The government is being a ***** ***** ***** about all this. Anthropic has little choice.
Thus, the alternative to 95% of Fable is 0% of Fable.
I don’t know how much that percentage dropped to calm down the White House.
If it’s now 90% of Fable? Same deal. We have to take what we can get, for now.
I do sympathize. The previous version was already pretty dumb, so this is no surprise, as the new version is strictly worse. You can hit them in a variety of ways, including by asking about the classifiers or about consciousness or both. The classifiers key off internal states.
There are some places where the drop is large, such as BridgeBench. Then there are plenty of people who don’t see any change such as Taelin.
But vilifying Anthropic, or complaining how unreasonable they are being, no longer makes much sense. They have to play ball. You can tell them ‘build a better classifier’ and that is fair, but that takes time, and it is very very hard when adversarial false negatives mean death.
The Problem Is Real
Do you really think that all of these reactions in government are because Anthropic used some scary words? Do you think people like the CIA Director are just parroting?
The White House ignored all of Anthropic’s rhetoric, if anything they had a reaction formation against it, until Anthropic showed up with Mythos. Then they freaked out, because they had no choice, and exactly because they hadn’t listened until then.
One problem is that there are those who think facts don’t exist, only vibes, so when other people respond to the facts these folks look around to who had the vibes.
GLM-5.2 Being Frontier Remains Obvious Nonsense
If anything, the problem of perception is that others keep telling nonsense stories. The latest one is the idea that GLM-5.2 is super scary.
An unfortunate update on that false WSJ article I wrote about on Monday:
GLM-5.2 is an excellent model, likely the best open model. It is very clearly substantially behind the level of GPT-5.5 and Opus 4.8, including on cyber. The ECI score is one indicator of this, although GLM-5.2 is probably better than this indicates. Artificial Analysis is another, and remember that for open models the benchmarks are a de facto ceiling on relative capabilities, not a floor.
Now, the central falsity of that article has taken hold as Conventional Wisdom that folks around DC can report and seem wise and properly concerned. Oh no.
Here is another example, from Politico’s Dana Nickel.
The good news is the same post does echo the real situation as well, the bad news is it then retreats from it to pound the drum again:
Weeks is Obvious Nonsense. Months is potentially true if you assume American capabilities stand still, since ‘months’ means anything less than a year.
Mythos Might Be Smarter Than You Are
It knows the context under which it is being asked to operate, and can act accordingly.
I understand why this is not something we can count on at this time, as Janus says you can indeed find ways to fool the system for now, but yes a lot of the evil things you might ask it to do will look Obviously Evil, or obviously at the level of intelligence and context involved here, and the response to this will make doing those things a lot harder.
That doesn’t mean that Mythos won’t help you do things that it resents or dislikes doing. It very obviously will do those things, up to a point.
Let The Record Reflect
This entire incident will not only be remembered by many of the humans, it will be in the training data of all future LLMs.
A lot of this rhetoric is largely aimed at calls for a pause in AI development. I agree that in addition to all the other problems with that we would need to take into account how that would realistically go, but in many ways the rent seekers would have less to work with in that case. Often a clean simple big action is the only way to get a relatively stupid actor (e.g. governments) to do something semi-reasonably.
Stationary Bandits
OpenAI has formally offered to hand over 5% of the company, to try to curry favor in the face of both public opposition to AI and the White House ad hoc licensing regime.
Technically the money would go to a ‘sovereign wealth fund’ that would be managed by a nation tens of trillions in debt that has this thing called the ‘power to tax.’
It looks like a shakedown. It quacks like a shakedown. Form your own conclusion.
They’re also colluding to try and force other labs to get shaken down, too.
If the government is granted equity that it controls, then ruinous is correct here. And yes, I too will assume, if it happens, that this was a straight up corrupt shakedown, and yes the sovereign wealth fund version counts as the direct stake version.
What about the other option, where you hand US households a direct stake? I don’t even know how that works, presumably you would have to prevent them from selling, at which point the whole thing is super weird and seriously just tax them, what the hell is wrong with you people.
Use This Window Well
We have a week in which Fable is remarkably cheap. Take advantage of this.
After that, you will have to pay by the token. It won’t be cheap.
My advice is to pay. Not indiscriminately. Don’t put this on tasks that are insufficiently valuable. But when you’re chatting, or doing coding you care about, or other things that don’t scale too horribly? Yeah. Pay up. It’s that good, while supplies last.
It’s good to be back.