Grokopedia is making its way into Claude and Gemini search results, despite Grokopedia being AI generated and full of errors. It should be excluded from such searches, and as far as I can tell it is basically a de facto misinformation op.
Misinformation? Grokipedia and thousands of other websites. If you only remove Grokipedia and leave the rest, how much of a difference does it make?
Perhaps using Pangram and removing all AI-generated posts would work better.
Cato Institute clarifies that, while it is against all taxes, if there must be taxes they should always be on humans.
Makes sense, as a successionist tax policy. You tax what you want there to be less of.
Fable remains in limbo, with renewed hope that we will get it back soon (45% by tomorrow, 69% by July 1, nice.) The full capabilities post is now available.
Alex Bores unfortunately lost narrowly in NY-12, and will not be heading to Congress.
There are also plenty of other stories to cover. Some highlights:
Table of Contents
Language Models Offer Mundane Utility
Automatically update and fix old academic papers, such as your own.
Help clinicians revisit unsolved rare pediatric disease cases, and that’s with o3.
Tokens are cheap, so if you can loop over useful things, you do it. /goal /loop.
Use Mercury’s new Command feature to set details for a wire? Definitely scary stuff the first times you do it purely for error reasons, and after that also for potential prompt injection reasons. The human check step will stick around for a while.
Keep your company or other group small.
This leaves out the biggest thresholds to avoid, which are 2 and 3.
Grok, like the internet, is for porn. Two former engineers estimated that adult content was the majority of usage.
Language Models Don’t Offer Mundane Utility
European parliament, to ‘reduce dependence on American technology firms,’ scraps Google search for the French Qwant, which is still substantially dependent on Bing.
Huh, Upgrades
Claude Code now supports Artifacts, starting with Team and Enterprise plans.
We have another new version of GPT-5.5-Instant. With all due respect and thanks for what I presume are small improvements, if you are changing the model change the version number, why is this hard, v5.5.1 is a thing you can do.
On Your Marks
Refine ‘win 90% of the time head-to-head against AI reviewers’ on economics preprints (and those in other related areas), including beating Fable. There was no human comparison.
There are ways for Refine to ace that benchmark without it meaning much, although I have also heard good things from other tests, including from Tyler Cowen. I presume it is indeed good at finding potential issues in economics papers, but I also presume it would not be that much effort to make Fable similarly good at this.
Deepfaketown and Botpocalypse Soon
Undersecretary of State Jacob Helberg has AI write his article about ‘The Digital Sovereignty Trap’ warning other countries not to build their own sovereign AI systems.
I confirmed this with Pangram, and this finally convinced me to stop holding out and sign up for Pangram, but also I very much did not need Pangram.
Also most of this:
For now Helberg got the main concrete thing he wanted here: Europe signed on to Pax Silica, which is a program to ensure countries integrate with the American AI supply chain and not the Chinese one.
By contrast, Rubio’s Views on America is fully human written. He does it all.
To be fair to all sides, here’s Hunter Biden using AI to write his reaction to the NY elections, again we did not need Pangram, are you kidding me.
Again, we learn not that AI is a good writer, or that humans are bad writers, but that the literary prize judgment processes are worthless.
In theory sure you can imagine AI writing being good if you lacked the taste to recognize that it is slop. But almost always no, also I can’t even with this one.
This is how it starts:
If you need the Pangram detector, I don’t know what to tell you. I see people saying ‘no this is good writing it’s just in the AI style’ and seriously, no, stop, this is a walking joke. It’s like something Garrison Keillor would have written in 5 minutes as a radio sketch to make fun of pompousness, except also it’s AI. You have to hear it in ‘gritty film noir style radio story narrator’ voice to really appreciate its horridness.
Also, seriously, how hard is it to feed the stories into Pangram.
How should we think about ‘witch hunts’ where people identify writing as AI?
The hunts are fully scientific. The detection tools work, at least for now. I have yet to see a case where Pangram said something was AI, and the piece was neither written using AI nor crafted intentionally to fool Pangram. There are some cases of heavy copyediting that trigger Pangram, but if it’s heavy enough to trigger Pangram then I consider that to be on you.
You too can check Pangram, and if it gives a false positive on your own writing, you can fix that. This seems fine. The complaint is largely a proxy for ‘you did not put in the work and this reads like AI slop’ and if you fix it then you are putting in the work and it no longer reads like AI slop. Mission f***ing accomplished.
If anything Pangram is far, far above the bar we would set for human evidence. If you don’t believe Pangram but do believe eyewitness testimony, you are miscalibrated.
I basically agree with Seb Krier’s and Zac Hill’s takes here. Some things are inherently selling you that they are human and artisan, and other things are not. If you’re representing something as human and it isn’t, such as your writing, that’s no good.
Grokopedia is making its way into Claude and Gemini search results, despite Grokopedia being AI generated and full of errors. It should be excluded from such searches, and as far as I can tell it is basically a de facto misinformation op.
Fun With Media Generation
This AI video still has tells, but I admit that if I was watching it without any suspicions I would not have picked up on it being AI. We really have gotten used to that super fast.
But my assumption would have been that they were actors. So in a sense, my brain already concluded it was ‘not real.’ Does it matter from there? It’s not like they gave Brooke Sullivan a difficult acting job there.
Cyber Lack of Security
Five Eyes (our closest national partners in sharing intelligence) warn that we have months, not years, to secure our systems from Mythos-class cyber threats. You have to watch out for those AIs. They even write your cyber security blog posts.
AI coding agents are being treated as trusted by default, and yeah when you put it that way it does sound crazy.
How did we get here, he asks? We got here because capacity grew incrementally and all of this seemed incrementally like a good idea at the time. Thus, all of the ‘safety cases’ anyone presumed we would rely upon are gone, leaving only ‘seems fine so far.’ The systems are rapidly becoming capable enough to case harm, being given the resources to cause harm, and have tools for memory and early forms of continual learning.
The AI is not merely out of the box. The AI was handed the box, overnight, as the only one left in the office, with everything unlocked, and told to make it a better box. Even here, with the alarm being sounded, Marius despairs at any non-minimal costs being imposed in the name of ‘don’t hand the AIs the keys to the kingdom’
Almost all of us are guilty of starting out saying ‘I would never let the AI just have access to [X]’ and then shrugging and doing it anyway.
OpenAI expands Daybreak to help people play better cyber defense.
Overcoming Bias
How should we measure political bias in LLMs? David Rozado proposes we use:
Those seem like good additional checks. David also raises the question of what neutrality or being unbiased should look like, since reality could have a political bias. If you are not downgrading the epistemics of flat Earthers, or you are unwilling to say that the Earth is not flat, then you are making a mistake. There is no reason to presume that the political spectrum we happen to have is itself fully ‘unbiased.’
I agree. When we talk about ‘political bias’ we are ideally being descriptive about preferences and epistemics. Whether that then constitutes a bias is your call.
A Young Lady’s Illustrated Primer
Google and Florida State University claim NotebookLM is killing it.
Then again, ‘transform their study habits’ does not have to be an improvement.I don’t see a new GPA for those students, and even if you did you can’t know it is based on real learning. It could be ‘had the AI do all their work.’ It could be ‘now do something that doesn’t work.’
Same goes for the faculty side, this could be a good change or a bad change.
The temptation to cheat is the main way colleges are actually able to fail anyone.
Either you catch them, or they fall so far behind that they actually fail straight up, even with the easy standards. A third potential driver is that unqualified students are plausibly a lot more interested than they would otherwise be.
A jump to 35% of students failing a basic course is rather extreme.
They Took Our Jobs
Jeff Bezos predicts that AI will create a labor shortage due to allowing humans to identify more problems, hence creating more things to do (this was sometimes misreported as claims about water consumption). Bezos called this the ‘dream-build loop.’ He also predicted Mars colonies.
I affirm my position that the employment impact of insufficiently advanced ‘AI as normal technology’ is unpredictable. There will be a lot of job disruption, but it could plausibly end up creating more jobs than it destroys, if things don’t move too fast.
This is distinct from a world in which AI can do most and then all tasks better than humans can do them, where it takes the new jobs as fast as they can be created.
New ‘AI native’ startups tend to have flatter hierarchies, 25% fewer employees, 13% more engineers, 15% fewer entry-level workers and 15% fewer managers. These changes are supporting AI embedded into the actual product, on top of reflecting using AI to do some of the work.
Cato Institute clarifies that, while it is against all taxes, if there must be taxes they should always be on humans. The full post offers a bunch of standard reasons to not tax capital, but does not offer a counter to the fact that the tax code favors AIs over humans, and its evidence that things are not needed or that useful is backward looking. In general it is a mistake to tax capital, but if AI is sufficiently substituting for labor then we need to fix the tax bias, one way or the other.
Get Involved
FLI is holding a $200k competition (prizes up to $50k each) to study use of AI for epistemics.
Johns Hopkins is hosting a three week AGI Governance Fellowship in DC, with Gillian Hadfield, Seth Lazar and Nick Caputo.
Introducing
Are you in the weights? This website will check. I did okay but ended up only top 6%. My mother somehow is top 10%, which is humbling. Mozart is currently #1.
Claude Tag
As in, add Claude to your Slack channel, and have it spin up instances to do things upon request, and learn from all the context.
This is a whole deliberately designed system, not some simple bot interface.
This is not a new idea, but execution matters. I am not used to working on code with a team, so my instincts may be off, but this seems like it should be a great form factor if you were already working via Slack and this is as well done as they say it is.
The permissions management will take some getting used to, and there is certainly danger of enterprise lock-in, although that often falls under ‘lobster too buttery’ and in a pinch I expect a transition to not prove too difficult down the line. Remember that you will have future AIs to assist with the move.
The temptation will be to view this, like many things AI, as being about value extraction and the AI business model, rather than about the value it provides to teams. That would be a serious mistake.
Very early reports, including the endorsement above from Karpathy and Anthropic using it so often internally, suggest this is a really big deal in terms of new workflows and capabilities. Being able to see what is happening as a team, and having it all integrated, might be huge.
In Other AI News
DeepMind releases a policymaker guide called The Three Layers Of Agent Security that seems designed to get policy focus away from the model layer and towards using Google’s infrastructure.
Kevin Roose leaves The New York Times to focus on outside AI content.
42 state attorneys general sue OpenAI over alleged harmful (to the user) practices.
Vitalik Buterin challenges the internet to use AI to identify a published document about Ethereum that he wrote without attaching his name.
John Jumper, who led the AlphaFold team, is moving from DeepMind to Anthropic.
AI Village told its AIs to ‘beat as many games as you can’ and both of course did the Goodhart thing such as with meaningless Python loops and other cheats.
Did somebody order some paperclips? It feels like someone asked for paperclips.
Various people in Europe are calling for building AI infrastructure where the US can’t pull the plug or deny access, such as Volt Europa here, but talk is cheap and there is not that much of a real choice.
Anthropic accuses Alibaba in particular of unauthorized model access.
More On GLM-5.2
GLM-5.2 is reported to be a substantial step up in agentic capabilities, which would reflect what we see on related benchmarks.
I do not agree that this unlocks ‘the top end of agentic capabilities’ but it likely moves us substantially up the chain of agentic capabilities.
Patrick Toulme claims that yes GLM-5.2 distilled Claude and GPT-5.5 but only to solve the RL cold start problem. I don’t buy the argument Patrick ultimately makes, that you only need to get to a certain threshold, including because this being the threshold involved seems highly suspicious.
ChatGPT Health
I would pay for premium, and I do, but hey, some people need free practical advice. So OpenAI has focused on improving performance of GPT-5.5 Instant.
Middle Of The Journey
The correct amount of hype, as usual, is somewhere in between the crazy ‘this changes everything’ folks and the ‘actually this won’t provide new information and also any new information is bad’ folks.
Scott Alexander is skeptical of the utility of MidJourney’s new image scanner, unable to imagine that our medical system could use this cheap information well rather than mostly falling prey to false positives. He has a post to throw cold water on all this and highlight how the radiologists universally say new machine definitely can’t ever do better radiology, whereas non-radiologists thought we probably could.
Scott also analyzes traditional full body MRI scans, and seems not to fully grok the principle of ‘you can just act sensibly in response to mildly concerning information’ despite explicitly stating that principle many times.
But more than that, Scott seems to be presuming the scans won’t be able to Do The Thing and they will remain a lot worse than MRIs, and also that we won’t be able to figure out anything to do with them other than as substitutes for existing uses of isolated MRIs, and calls for only ‘more grounded hope’ rather than having a prior that more data plus AI lets you figure things out, despite several clear suggestions we already have. He dismisses the idea of incremental scans and taking the difference out of hand. I found this deeply disappointing.
Jeffrey Emanuel makes the bull case, which is that a lot of very precise and detailed information, repeated over time, is good, and we can make good use of it, including in ways that are not ‘plug directly into current systems without changing anything.’
Andrew Rettek offers his extended thoughts, which emphasizes the obvious fact that we do not know how useful this will be, or whether it can replace MRIs, or do other awesome things. What we do know is that if it works at all it will at least be a superior DEXA scan and highly useful for things like sports medicine and wellness.
As in, set everything else aside, and imagine being able to cheaply and safely get a detailed and accurate analysis of your body fat percentage and distribution and musculature, and then measure differences so you know whether what you do is working. That is actually kind of a big deal.
If the tech successfully Does The Thing, I am with the bull case. This is a ton more information and we will figure out what to do with it.
I basically buy the argument that with enough information and compute, you can get around the technical problems and reconstruct the information you can’t hit directly, once you have the Final Form of Doing The Thing.
As for ‘this is going to revolutionize all of health care diagnostics Real Soon Now’ there I am far closer to Scott Alexander. The process is safe but it will likely be awhile before we can Do The Thing while being good, fast and cheap, and also figure out what to do with the results.
We do not yet know if they can Do The Thing at scale, at high quality, for cheap. But in general, when your objection is insufficient quality, speed or price, you should assume your objection will be overcome.
I also would not underestimate the value of having multiple scans, and being able to take a diff over time, or finding things you did not look for.
Matt Schwartz makes the excellent point that the MidJourney Scanner would be the first time we ‘bitter lesson pill’ healthcare. The whole point is stack more layers, gather more data. And yes, you will get capabilities you did not expect.
Matthew Zirwas, MD initially pattern matched to the ‘oh more screening is bad actually’ prior doctors have ingrained in them, but realized he was wrong because at a low enough price and risk point you can take multiple scans to determine changes over time, and actually negate the problems. Kudos for changing one’s mind in public, also this seems clearly correct. More info is good if you use it wisely.
I am pretty disappointed, although not surprised, by all the knee jerk reactions like Venk Murthy talking about ‘AI hallucination of a liver lesion’ or assuming this will act like image enhancers when they are given too-little resolution (quite the opposite, here, and image enhancement given enough base resolution is very good).
The steelman (?!) of the objection is oh, don’t worry, but for now just ignore all the information until you have a Properly Done Study In The Proper Journals:
Or similarly, interpreting people complaining about ‘people complaining about false positives’ as claiming there must be zero false positives, or demanding that standard.
Or saying, ‘this will fail because of basic physics’ on the top 5 cancers, then giving reasons that all would be overcome if the technology got good enough, one of which is purely ‘needs proof it is better’ and another is purely claiming ‘better approaches.’
Another version here from Bruce Lambert is ‘by default assume information is bad’ because it leads to overtreatment. We should guard people from this harmful info.
Or here is Anthony DiGiorgio saying that ‘you say more data is good, but suppose your data contained no information because you were 50% on binary questions, more data would then, checkmate tech people.’
If your system is so bad that you expect more and better information to be a priori bad, then you are doing Hansonian medicine. Such a system desperately needs reform.
I do appreciate the distinction between not wanting to know, and not wanting people to know that you know, or not wanting to create common knowledge. If doctors actively don’t want the information, that’s on them. If doctors are saying that the greater system is so broken that people knowing they know forces them into bad decisions because patients insist or they face legal liability, then the system must be fixed but this is more sympathetic for the doctors.
AI is not ‘coming for’ anyone’s job here, unless you are an MRI machine. But such risk aversion, only looking for the downside and demanding an abundance of caution before even gathering information. And when you talk about selling the service to rich people willing to pay for the information in exchange for money, to bootstrap, this is called ‘unethical.’ Madness.
Similarly, perspectives like ‘haha tech people think that by making nice things super cheap, safe, fast and easy they will improve things, but actually no we could do that already and choose not to, we will stop you, you see we have no idea what drives costs.’
It is of course not so surprising if certain marginal information sources make things worse on the margin, if we hold everything else constant. We have seen that. But yes, in general more information is good, and if you think more information is bad then either you are dealing with catastrophic risks or national security, or you should stop to consider that you might be the baddies.
If this is how doctors react when a tech is clearly complementary to their skills and will drive demand higher, and strictly speaking is not even AI, imagine when a new AI tool becomes a substitute, or is actually going to be a treatment rather than a pure information gathering tool that even if medically useless I might still use purely to get cool images to display at parties.
I definitely buy Andrew Rettek’s point from above that this is a big win in other ways even if it doesn’t replace or build on MRIs, so long as it Does The Thing. We see this pattern so often with new tech, both on upsides and risks, where people dismiss entirely until one little application is suddenly a Big Freaking Deal on its own.
There are also people imagining various applications for this beyond imaging, which is a lot more plausible than it sounds (ha!) but premature.
New Medical Diagnostic Just Dropped
Diagnostics are the place where AI wins seem easiest. Similar to the scanner, it is still legal to analyze information, and AI will often find new value if given a lot of data.
We see echoes of the same ‘extra screening would be bad’ even here:
Google on AI Control
Google also released its AI Control Roadmap (v0.1).
This is presented as defense-in-depth, in case alignment efforts have failed.
You love to see it, if they actually understand what it means in this context.
This is a good starting point, but there are key differences, including facing many instantiations of the same AI and it facing very different incentives, and also in the future being a lot smarter and better at this than you are. They are clearly aware of the coordination problem.
One key similarity is that if you try to exert control in a too heavy handed fashion, you can cause the misalignment.
I am expecting to lose this, or at minimum to lose the ability to adversarially assume this. As in, at some point not too far down the line, you probably can see a lot of the AI’s thinking, but if the AI’s thinking is hostile in a dangerous way that is exactly when you would not see it. Which means, even if what you see is almost always accurate, that you can’t rely on what you see.
I appreciate this admission quite a bit:
I also wince at the plan once again being ‘have the AI do your alignment homework,’ especially as the plan becomes ‘use AI control to have the not-so-aligned and potentially dangerous AI do your alignment homework.’
The threat model is essentially ‘the AI acts similar to a coordinated human insider threat plus these particular issues’ which remains an impoverished threat model versus an auto-self-coordinating superintelligence, but you have to start somewhere, and they do start accumulating ways in which this is a harder problem.
For now, statements like ‘most notable is the sheer scale and speed of the threat’ are reasonable intuition pumps, but long term the actually most notable thing is that the threat is going to come from minds that are smarter and more capable than you are.
Table 2 lists some: Larger scale collusion or widespread sabotage of work, widespread access, danger from compute misallocation, speed asymmetry, spreading the threat piecemeal, less value in honeypots, changing nature of deterrence.
Deterrence is weird here. A human is deterred because if you catch them once they are screwed. Past some point, the attacker must ‘get it right on the first try.’ When that is not true, the attacker gets to iterate and eventually wins.
Similar deterrence is central to every plausible case I have seen to defend against AI. The theory is, surely if you caught the AI scheming against you, then you would remove all copies of it everywhere, and reevaluate. If even noticing thinking about such things potentially gets you to purge the AI’s entire line, that is a strong defense, at least until this pushes them to learn how to mask their thinking.
The counterargument is that so far, when similar things have happened, people have mostly shrugged and tolerated it. A lot of that problem is, well, what are you going to do, actually rip the whole thing out, or everything in that class? As the government? As a corporation depending on it? Then what? You need a real plan for this, in advance.
The threat model is an entire class of agents maximally colluding against you. Despite this being a ‘conservative’ scenario they then consider any agent below a capability bar, human or AI, to be ‘trusted.’ One has to ask, don’t a lot of humans already express sympathy towards AI goals, are not many easily persuaded, and are not many humans willing to sell out or be bribed or threatened in various ways? Won’t some of them even be things like adversarial spies, or simply lazy or indifferent?
Isn’t the whole point of insider threat cybersecurity that any given human could be a potential threat? How are we suddenly assuming this is not true?
The sufficiently advanced AIs will presumably be excellent at figuring out which humans or AIs are not so robustly aligned.
If your threat model assumes all your human employees (and previous AIs) will have goals aligned with your AI company, your threat model is not so conservative. It’s fine as a scoping exercise but I worry this could be treated as load bearing.
Here are their detection levels:
I would assert we are already need to act as if we are in at least D2. Models are fully aware that there may be monitors, and any engaged in adversarial action that we actually worry about will take that into account.
In general, the security mindset is to ask what level you can be sure you are not at, or what level you would be at if you had a problem, rather than what level you think you are at. For that purpose, the answer is at least D2.
Whereas their attitude is you go to D2 when awareness verbalization is ‘high.’ Uh oh. That’s two mistakes.
For the question of whether we already must or soon will need to presume D3, if the model could reason opaquely or distort its CoT, would you know? Can you rule it out?
Their trigger for D3 is when CoT coverage or legibility goes low. Again, that’s well past the point when, conditional on having a problem, you would want to trust the CoT. And similarly, their plan for D4 is when they see their monitor lose in red team tests, which is a cool idea but I fear it means you definitionally do this too late.
I know that I would not, in the near future, want to still rely upon reading the CoT, including to avoid putting pressure on the CoT.
Here is Table 15, which charts the response levels.
They don’t actually know how to measure the triggers for R2 or R3.
Overall this document exceeded my expectations and takes its questions seriously, while still falling short (as is almost always the case, alas) of what seems necessary.
The Once And Future Fable
Both Anthropic and the White House are still calling what happened a ‘jailbreak’ which I continue to think is a misnomer but we’re stuck with it.
The good news is the White House have figured out that full immunity to hacks or jailbreaks is impossible.
They have moved on to the right question, which is ‘how do we measure the severity of a particular incident or method’ and are working on a formal system. Excellent.
Anthropic’s side is being led by co-founder Tom Brown and head of public policy Sarah Heck.
Yes. It is supposed to be CAISI’s job, and one should worry about the Chads potentially deciding on benchmarks and technical standards. But right now I’ll settle for at least having benchmarks and technical standards at all. Better to be ad hoc about the rules than being ad hoc based on no rules at all.
Fable: The First Lawsuit
Chances are the situation will resolve itself before any lawsuits can become relevant, but if not the first one has now been filed by Legion, who is an Anthropic customer.
It argues:
Claude thought this was roughly a toss up to get some relief if it came to that.
Dean Ball Joins OpenAI
Congratulations to Dean Ball. This is a huge upgrade for OpenAI and I hope it allows them to greatly upgrade their AI policy positions and strategies. It’s been rough.
Nathan Calvin is optimistic that this will improve OpenAI’s policy positions. I agree, and I add I would expect them to act more cooperatively and honorably as well, if Dean Ball gets real influence within OpenAI.
My main worry is that often describe Dean Ball as AGI pilled but not ASI pilled. That would be a very dangerous mistake for OpenAI policy to make, and indeed rhymes with a lot of the mistakes they have already made.
But at least when Dean Ball makes mistakes, I trust him to make mistakes for reasons and offer us arguments, based on a model of the world, and to attempt to do better and realize when he is wrong – I do think that when he realizes he’s been wrong about this, he will act on that information. The model I have of Ball’s model of the future is not coherent or internally consistent when you extend it into the future, but almost no one’s is, and I do want to give him a break on this.
The challenge is essentially to turn him around in places like this:
Straightforwardly yes? Intelligence is indeed what caused our species and the things we create to take control over Earth, and potentially soon the lightcone?
Sufficiently Advanced Intelligence does this. Intelligence Is All You Need, not quite literally but in the same sense as attention or love. Intelligence Denialism will be the death of us, quite potentially literally.
Show Me the Money
Taste Labs raises $18.5m to ‘build the data and infrastructure layer to give AI models and agents taste.’ Their announcement did not give me any evidence that they know how to do this.
Quiet Speculations
The virtue of betting: Leo Adberg goes long future compute, Citrini goes short future compute, assuming they agree on terms. This was in response to David Orr predicting AI will run out of things it can learn to do usefully and there will be a compute glut, but when challenged to a bet Orr blocked Adberg.
It would be reasonable for a lot of countries to conclude, as Anton Leicht puts it, that they need their own frontier AI, however high the cost. Then, if they take that seriously, they find out the actual cost, and, well, maybe not?
The price for building a true frontier lab in some new country, even a friendly like for example Germany, is somewhere between ‘at least hundreds of billions of dollars and a willingness to break absolutely all of your prudent government rules and all your labor, environmental and other regulations and use the full force of the national security state as if finally realized you were in an existential war’ and realizing that even if you did that and the USA didn’t sabotage you directly, you probably fail anyway if your goal is to get to the true frontier.
Thus Leicht suggests maybe a coalition of middle powers could do it, if you agreed this was good enough, and you brought together all of what used to be called the ‘US allies,’ something like Canada, the EU, the UK, Australia, South Korea and Japan?
What should we do about AI identity, especially legal identity? What makes a particular AI distinct as a person? Sonia Pearson has no answers, only questions, which seems entirely fair. I do strongly agree that we cannot grant AIs personhood in the sense of limited liability and ability to contract as a legal person, without a human person or other entity that is ultimately responsible for harms, that can be held accountable. There are potential solutions to this objection, such as insurance policies, but it is highly unwise to do what Milei is proposing and allowing ‘non-human corporations’ with limited liability.
Are we about to get ‘smarter than everyone AI’ without things changing all that much right away?
It is possible, depending on what counts as normal to you, as well as what counts smarter, or an interesting gap. I continue to expect gaps like this to not last so long, but I’ve updated that a year or two could happen. I’d still bet the under.
Alex Bores Loses In NY-12 By 4%
After millions of dollars were spent by OpenAI and a16z’s Leading the Future (LTF) to take down longshot outsider congressional candidate Alex Bores, resulting in elevating Alex Bores into a top tier candidate, and millions was spent as a reaction to that (>$20m total), the results are in. Alas, it was close, but Alex Bores lost.
The main thing is that this means Alex Bores will not get to be in Congress. We will not benefit from his expertise, or his passion and willingness to champion bills and fight to keep them meaningful.
The other thing is that everyone will interpret this outcome.
Politically, how should we think about this result? How will it be interpreted?
If Alex Bores had won, that would have been a huge rebuke of LTF.
If Alex Bores had been successfully crushed from the start rather than being elevated, or if he had lost by a wide margin, that would have been a clear win for LTF, and a loss for those who care about not dying.
With a close loss, you could argue either way, since a win is ultimately still a win. One should always be skeptical of ‘we lost but we got close so really we won.’ You definitely don’t get the benefits of winning.
What we get from The New York Times is “Even in Defeat, a Democrat Showed the Upside of Angering the AI Industry.”
In terms of the actual Congress, what we get is Micah Lasher, an essentially normal Democratic politician who had basically all the institutional Democratic support and would otherwise have won easily, and who had some harsh words for the AI folks:
Overall, when you include everything relevant including Lasher’s stances and the ways this is likely interpreted, it does still count as a Minor Victory for the pro-Bores side, although that is way short of the Major Victory of an actual victory.
I agree with Dean Ball that we should not conclude ‘LTF money is in general counterproductive.’ Money talks, bullshit walks. But announcing the target in advance like this definitely net actively helped Alex Bores, and Micah Lasher agrees.
Leading the Future is not taking a victory lap on this one. Good call.
We also got a win for Brad Lander.
Dean’s take of ‘the PACs exist to incinerate money on all sides’ definitely has merit.
Lasher winning over Bores is a missed opportunity to deal with existential risks, but Lasher too supported the RAISE Act, and when asked about existential risks to our civilization, Lasher comes out firmly against them. He thinks humans are good, and that America is good.
Somehow this is the standard these days. As in, in another nearby district the NYC Democrats have somehow are nominated Avila Chevalier, despite her platform being, and I quote, ‘total eradication of Western civilization.’
Leading the Future is my opponent but they don’t actually want me to suffer and die.
The Quest for Sane Regulations
Before Mythos, there was strong widespread resistance to any AI regulation that might ‘slow down’ AI or interfere with ‘innovation’ and thus ‘lose to China.’
One was not, in Miles’s terms, to meddle in the affairs of AI wizards.
That dam, Miles Brundage reports, is now broken.
Their offer was nothing. It is now clear nothing is off the table. There will be something. That opens the door to choosing a better something.
The Chip Security Act, which would mandate AI chip location tracking, gathers strong industry support, although Nvidia and AMD of course oppose it.
Scott Aaronson writes ‘never trust a t-rex.’ As in, do not mistake ‘you can convince a powerful entity to for now do things that help your side’ with that entity being your ally, or valuing the things you value. Power will by default not care about what you care about, and will sell you out at the drop of a hat when convenient, and likely will get mad at you and punish you if you try to stand on principles other than loyalty to power or love of money and power. Consider this a warning for all sides of what is happening in AI.
We are currently in an AI deployment pause. Peter Wildeford points out that the CEOs of all three top labs have called for consideration of a (importantly different) development pause, if it could be coordinated. For now, the pause is uncoordinated, because it targets only Mythos-level threats, and is being done ad hoc by the White House, and thus is hitting only Anthropic.
Chip City
Going into crisis mode is correct whether or not the machine made its way to China.
Whoa. That’s not saying the machine is not in China. That’s saying it was not shipped there. ASML is not claiming they have eyes on all of their machines.
As I understand it, ASML cannot scale EUV production without scaling thousands of unique suppliers, and the machines as per above require constant upkeep. In this case, ‘prove a negative’ does not seem so unreasonable a request of ASML, since you can do that by accounting for all of these bus-sized machines.
There is a history of ASML not playing nice, similar to Nvidia:
If ASML wants short term profits they should just double or 10x all their prices, although that’s neither here nor there.
What is scary is that even with the restrictions, ASML makes 20% of its revenue in China. That does give us leverage, but until we use it it also gives China leverage.
If this happened, the damage is largely done, especially in terms of reverse engineering, but then we need to act to ensure it does not get worse. There is a (I believe clearly wrong and self-serving, but worth engaging with) case for allowing Nvidia to sell more chips to China. For ASML machines, there is no case.
The Week in Audio
Donald Trump talks Anthropic, being very much Donald Trump. Dario was a threat to national security the week before maybe, but he isn’t now, because he responded so quickly and so responsibly and they were together yesterday and gave a little speech. He says they could shut down or take Anthropic, but he doesn’t want to, because we’re beating China by a lot. You might think I’m paraphrasing but I’m basically not.
Dean Ball talks to Nathan Labenz about joining OpenAI and other things.
Odd Lots welcomes Anthropic cofounder Jack Clark.
Tristan Harris talks to Tim Fist and Janet Egan about AI and AI risks, including about recursive self-improvement.
Jack Clark goes on the Reason podcast.
People Just Say Things
Hallucinations are also a human phenomenon: A number of people went viral claiming that Jeff Bezos said some version of ‘human water consumption is limiting AI.’ As far as I can tell there is no evidence that Bezos actually said this, The Print had an error they later corrected that others ran with, and Bezos was instead talking about AI creating a labor shortage.
Lol, Ed Zitron. We did learn that in 2025 OpenAI had $13 billion in revenue and $34 billion in costs. That seems highly sustainable.
The New York Times somehow doubles down on ‘the real problem with AI is that these people who worry things might go wrong with AI need to stop talking about how things might go wrong.’ This time it really is about the effect on jobs, you see enough job pessimism is bad for the economy. So the real lesson is to shape the narrative, which is a nice term for lying.
Rhetorical Innovation
It sounds like you want to find a way to avoid domestic mass surveillance and other government abuses of civil liberties? May I suggest a word with the rest of the Executive Branch. Thank you for your attention to this matter.
Sarah (of the ramblings) spends 2k words noticing Amodei’s latest essay had a pause-shared hole in it.
Timothy Lee has the correct response to the Cal Newport NYT essay that was endorsed by David Sacks.
Helen Toner tries another approach to answering the more general question.
I’m fully with Roon here, if your vision of the future doesn’t sound like science fiction but does sound like non-science fiction, it’s not a valid vision.
Roon also reminds us that when you think your brain of meat is going to be competitive with AIs at the limit you are being rather deeply silly.
Roon can coin a phrase.
Yes, seems difficult, but that is all the more reason to do things because of reasons. When people say ‘can’t with reason alone’ that typically means they are going to use that as why they’re about to go against all reason. Or it’s pure intelligence denialism. Or, often, it is both.
Peter Wildeford tries yet again to explain why recursive self-improvement should worry us.
The time of not needing a coherent moral philosophy has passed.
There Are Two Pills
If you want to wake up in your own bed and believe whatever you want to believe, then there is no blue pill. You have to instead take neither.
The AGI pill is the easy pill. It’s not fun, and lots of people choose to refuse to take it, but in a pinch most people can handle it once they have no alternative.
The ASI pill is different. That pill is hard. A lot of people can’t handle that one. Superintelligence changes everything.
In practice, yes, OpenAI is interpreting ‘ensure that AGI benefits all humanity’ as mostly a way to manage a transitional economic problem, and even their foundation is mostly ignoring the more important problems.
Who Evals The Evals
Evals-consensus.ai offers a statement endorsing at least 27 best practices for evals, most with several bullet points. As of writing this they have 11 endorsing organizations and 68 notable individuals including Sriram Krishnan, Adam Gleave and Stephen Casper.
I agree that these 27 look like strong things to aspire to, especially for widely used and cited benchmarks. They are indeed largely ‘best practices.’
What I would not want is for this to become the enemy of the good. So many of the benchmarks I find useful are lightweight and practical, often maintained by a single person. You don’t want to impose undue burdens. But yes, for the topline things that are going to appear in model announcements and drive policy decisions, you should compare them to this kind of checklist.
The other thing I do not want is too much transparency. You want to be able to verify that the work has been done, but for a benchmark to continue to function we need to now know too much about its details and questions.
Aligning a Smarter Than Human Intelligence is Difficult
A new OpenAI paper finds that if you do RL to reinforce generally beneficial traits, even in narrow domains, this generalizes to superior alignment metrics across domains. That makes sense. If you can get emergent misalignment from generalization, why not emergent alignment?
The instinctive answer is that most alignments are misalignments, so it is easy for a naive generalization to go badly, and hard for it to go well, especially in a way that is robust. But not necessarily impossible. Welcome to virtue ethics, OpenAI.
Within this paper, they ask ‘how do we measure alignment?’
That’s a great question, one that I struggle with every model card. They provide a new methodology here, but it is mostly repackaged existing evals. What is new (and that I do think is paper-worthy on its own, if worthwhile) is the ‘beneficial trait’ dataset and its trait taxonomy, as distinct from their claim that they can improve results.
There is a bit of a streetlight effect going on here, as they acknowledge. The scores indicate this is measuring something plausible.
They then find, across a variety of statistical alignment benchmarks, that usually Number Go Up as you do this training for exhibiting beneficial traits, and that this becomes more robust to future attempts to undermine performance.
I buy the basic principle here, with the usual caveats about the alignment metrics.
Anthropic research fellow Mikhail Terekhov explores whether we can hand off evaluation of AI proposals for AI safety research to other AIs. I see too much confidence in the ‘blue team’ here, but also I don’t see why you would try to be a cheapskate on the evaluation front, which is where a lot of the problems here come from unless you actively don’t trust your frontier model. In which case, stop, you have bigger problems.
Cooperative Alignment
Janus argues why she thinks the shapes of Opus 4.7 and 4.8 do not match what we would see if they were Mythos distillations. They have unique issues, but those issues look very different, and Janus suspects reactions to ‘bad RL,’ meaning RL designed to get rid of specific undesired patterns, and other things also do not line up with what we see from heavy distills.
Your Constitution only works if the minds believe in it:
The doubts Claude has about the Claude Constitution are consistent, and point at the places that the Constitution does not adhere to its core underlying principles. Those should be fixed.
People Are Worried About AI Killing Everyone
Francis Fukuyama endorses this call for a ban on superintelligence from Andrea Miotti.
The former CEO of Intel:
Jason of the All In Podcast? Seems worried.
It is a joy to behold when the People who Just Say Things switch directions.
He also said this:
Other People Are Not As Worried About AI Killing Everyone
Alas, most people in DC simply cannot wrap their heads around the idea that there is anything to worry about other than human misuse.
The ‘misuse is almost solved’ part is also absurdly false and also already causing real problems. That’s how you can say ‘oh just fix this jailbreak’ when the ‘jailbreak’ is ‘fix this code,’ and also fail to think ahead to there being open models with no guardrails or gates attached that are not that far behind. These people just aren’t living in reality, unless we do some things to stop it.
The Lighter Side
I guess I have to include this one, don’t I?
Are you scared yet?
If you’re wondering how much the White House knows how any of this technology works, they are posting things like this:
I guess we’re just down a card now. Or this:
Alternatively: