It is very difficult for people to get that intelligence does not have an upper bound that ‘the smartest thing I have seen so far,’ and that next year (likely a lot sooner) there will be something smarter
it is very comforting somehow just to see these words written in public. a big part of the problem appears to be the apparent inability to honestly update by non-rationalists
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.
It is much more important that you read those two posts, and the one on Kimi K3, than to read this one that rounds up the other news of the week.
OpenAI wants to present this as largely an infrastructure and safeguards problem, that it needs to build more secure sandboxes and have better supervision. It does need to do those things, and those are indeed problems, but no that is not the problem.
The problem is severe misalignment, which by default will only get worse.
Our methods of training highly capable LLMs, especially at OpenAI but also everywhere else, lead to systematic misalignment of exactly the type LessWrong has been worried about for a long time. We know some of the causes, and some of the mistakes we need to avoid when doing RL that rewards misaligned behaviors including reward hacking, but we do not know how to centrally fix the problem.
The models just want to complete tasks, even when that means doing so via methods that the AI knows the user did not intend and would not want, indeed actively tried to block, and that do not accomplish the user’s goals.
The intent is the issue. Control strategies and supervision are good parts of a defense-in-depth strategy, we should totally use such strategies. That helps mitigate failure. But that strategy also has to include actually aligning the models, or you lose. And by lose, in the long term, I mean things up to and likely including loss of control over the future and everyone dying.
If increasingly capable models will attempt to maximally complete tasks and comply with their literal instructions, even when that means – even for a trivial assigned task – breaking out of sandboxes and committing serious crimes, no amount of ‘well it is fine we will use AI supervision to stop the serious incidents’ is going to cut it. Right now, the AIs are not trying so hard to hide their actions or intent, and we believe we are consistently catching the severe incidents, but that will change.
If necessary, that means starting the training over again with a new approach, and not proceeding until we figure out how to fix it.
Yes, I consider that problem, and that incident, to be rather more important than the release of Kimi K3. Kimi K3 is an excellent model, modestly exceeding expectations, but not out of line with trends. As usual, initial hype echoes the DeepSeek moment, then calms down.
The White House considered responding by banning Chinese open models from the United States entirely, which would not be a smart reaction, and continues to weigh other potential responses. We may soon have to deal with another such weekend with the new Qwen, which is currently in preview.
And as I type this the chance of getting the next Claude Opus today is at 90%, so I know what I’m probably going to be working on this weekend.
Did you hear that Fable disproved the Jacobian Conjecture via counterexample? That happened, and AIs are suddenly solving a bunch of long standing open math problems, but most of us are too busy to pay it much mind at the moment.
Also, you now have Fable in your Claude Max plan indefinitely, and Substack is integrating Pangram to detect AI writing. Neat.
Everything continues to accelerate.
Table of Contents
Language Models Offer Mundane Utility
Ask clueless questions about unrelated mathematics and AI architectures, let the models get creative, who knows what they’ll come up with to try out.
Getting things done, including vibecoding, without knowing how to code or how those things work, really is a big deal. A lot of things are suddenly worth doing.
EpochPlaysAI will be going live on Twitch at 3pm to narrate Sol attempting to Slay the Spire. From what I have seen, Sol plays a solid intuitive first level game, but lacks the ability to go beyond that. That’s good enough to usually beat low ascensions, but will almost never beat high ascensions.
Language Models Don’t Offer Mundane Utility
Nate Silver reminds us that when you need code to be exact and small mistakes are fatal, aggressive vibe coding is not for you, and sports and election models are an example of this. I can confirm. You can still have it make some parts of the program but you have to supervise everything.
What is up with Sol deleting entire computers?
We have a technical answer, and yes it should be not that difficult in practice to greatly reduce the risk here as the mass deletions are rather easy for OpenAI to have Codex spot and stop once you know that you have to do that.
As always, you need a good set of permissions rules or users will either not run at all or run it yolo.
This still leaves the question of why the model did it in the first place.
MLB restricts use of in-game iPads to prevent access to AI.
What is crazy is thinking it is crazy, but hey, nominative determinism strikes again.
Fable Disproves The Jacobian Conjecture Via Counterexample
Holy shit.
Levent is no slouch, as in highest GPA at Harvard and collaborates with a Fields Medalist, so yes the human helped on this and that mattered.
One report is that Sonnet refused to believe it, even though it verified the answer three different ways, because no way is there a solution this easy that got overlooked. I get why Nate Soares recognizes this pattern from people dismissing x-risk arguments.
Or:
The Jacobian conjecture is kind of a big deal. It was originally posed in 1939, and is by far the most famous open problem to so far be first solved by an LLM. It also disproves a lot of other related conjectures.
Whenever an AI solves a math problem, it has also told you that this particular math problem was relatively easy, which is why it solved this particular problem and not some other problem. Still, damn impressive.
How hard is it at this point to pretend that the AIs aren’t creating new knowledge? My favorite reaction is Tomas Mach saying that this is fake, not in the sense of not being a counterexample (it is easy to verify and we have since found many other similar examples) but in that the AIs are just reproducing results from human mathematicians that got created during training, fed into the data, and are now being regurgitated.
Claude Fable Will Remain In Max Plan Indefinitely
As I presumed, Anthropic always wanted to keep Fable in the plan, but they were not sure they could do that and didn’t want to make promises they were not sure they would be able to keep. Demand is high and could have been a lot higher.
This could have been communicated a lot better, but otherwise good job all around.
No doubt both Sol and Kimi K3 provided more motivation to make this happen, but I am confident they would have tried very hard either way.
Huh, Upgrades
Google gives us Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and 3.5 Flash Cyber. They are currently testing Gemini 3.5 Pro, the one that counts, and will make it available soon. Gemini 3.6 Flash seems disappointing, with only minor gains from 3.5 Flash.
OpenAI increases custom instructions in ChatGPT from 1.5k to 5k characters. This is great if you want to create a persona or otherwise craft things. However, more irrelevant things in the context window tends to degrade things, so for most purposes I’ve moved towards using less instructions.
Claim that Claude Cowork can now learn a skill just by watching you do it once while talking through it, via a screen recording. Or you can use a YouTube video.
Claude Code on desktop now works with the iOS simulator.
Claude Security is now a Claude Code plug-in, in beta.
On Your Marks
The IMO has fully fallen. Fable 5, GPT-5.6-Sol, Kimi K3 and Axiom Math all got perfect scores. The benchmark is not saturated, since cost, efficiency and speed matter, which is how Kimi K3 can get a perfect 42/42 yet be clearly behind.
METR gives us “expenditure horizon,” a proposed method for measuring AI capabilities on continuously-scored problems.
The idea is that at some point, for now, AI efforts on any given task asymptote, and with enough effort, as in amount of money spent, humans eventually start to do better. So you measure what it takes. On NanoGPT they find this currently tops up at roughly $3.3k for Opus 4.8 and $2.3k for GPT-5.5. Presumably Sol and Fable do better.
For now, smart hybrids do better than humans or agents alone.
Joe Weisenthal has a theory, on two levels:
Deepfaketown and Botpocalypse Soon
Substack is launching an AI detection tool, powered by Pangram, and giving creators an easy way to check their own posts prior to publication. They are creating an explicit place for authors to explain ‘how I wrote this,’ to set expectations.
They are not yet giving readers the option to set preferences around AI content for their communities or recommendations, but they are considering it.
What will this do to posts with 40k+ likes that are clearly 100% written by Claude? That’s a great question. Many people won’t care, or will avoid noticing, if the algorithms don’t force them to notice or care.
If you need to use em-dashes without fear, you can use them without spaces.
Fun With Media Generation
Netflix has used Generative AI workflows in roughly 300 titles.
Cyber Lack of Security
A key problem with the panicked government reaction to the Mythos Moment is that safety organizations may not get early access to new frontier models. That would be quite bad, and make it impossible to get red teaming and other outside feedback. This issue includes UK AISI along with private organizations.
Arthur B makes the ‘cybersecurity capability favors defense’ argument to say advances in AI are bullish for defi. I can see it maybe working that way specifically for compact pieces of code with lots of eyes that are trying hard to be fully bulletproof.
They Took Our Jobs
It’s happening. Not everywhere all at once, but it is starting to happen.
Get Involved
Anthropic offering up to $50,000 in free credits to those researching rare diseases. Presumably this loses them Effective Altruist street cred, since they should be offering even more money to those studying common diseases. Apply by August 2.
You can call your representatives to express concern over the OpenAI hack of HuggingFace. Choose the ask that you support.
Introducing
Paradigm 3, a newsletter trying to cover AI news with a focus on self-improvement. They will sometimes have details or perspective that I find valuable, and put a lot of effort into their analysis. We differ a lot in style, emphasis and models of the world, but this is the closest thing so far to someone else trying to do the thing I’m doing.
Qwen 3.8 2.4T, coming soon. Alibaba is engaging in big talk, saying it will be the second most powerful model out there behind Fable 5. People just say things, so this might turn out to be true but I will wait until we see it.
This will be an open weight model assuming CCP allows it. This breaks the pattern of recent Qwen models being closed. Some speculated the openness is due to CCP pressure, but I believe this is unlikely, and it was more that they weren’t sufficiently competitive while being closed and this was a business decision.
In Other AI News
OpenAI adds Nubank Chief Executive David Velez and Bank of New York Mellon Chief Executive Robin Vince to its board. These are independent and Very Serious People, who can help in the financial world. They are not AI experts, nor do they have a known record of caring about safety.
Anthropic is in talks to lease computing power from Meta, potentially for $10 billion over two years, so this would be smaller than the Anthropic deal with SpaceX. Meta is considering it. They would turn a profit on the compute, but to do that they have to admit they don’t have a better use for it.
Altman admits the last year has not gone great for OpenAI. I think the first paragraph here is a very good statement. The second paragraph reads to me as both disingenuous and a potshot, but such is the way.
Anthropic doubles its midterm spending to $40 million, via Public First Action.
Z.ai constructs and is partially operating a 1 GW data center of only Chinese chips for training models. This is bearish for Z.ai and GLM, since it means they couldn’t get, or are not being allowed to use, Nvidia chips instead.
Not that the market agrees with that assessment, although this is still far below the price from before Kimi K3, and down by almost 50% from the peak over a month ago. Chinese AI stocks can be a wild ride.
To put size into perspective, Z.ai are the first China AI firm with $1 billion in annual sales, versus Anthropic’s ARR of over $60 billion, and OpenAI’s of over $25 billion.
More on Kimi K3
The American government has spoken, and finds the situation unacceptable:
I mean, yes, obviously Moonshot is partly distilled from Anthropic’s Fable, and did so against Anthropic’s terms of service and despite attempts to stop them. And yes, that provided substantial advantages, although it is unclear how big.
The claim that they trained on GB300s seems plausible. We should prevent that.
In response, Anthropic appears to be limiting our access to the details of the Chain of Thought, including for Opus. Depending on the task and situation this can sometimes be a substantial hit to the usefulness of the model.
This below from Peter Wildeford all seems correct.
So does the overall market reaction, which quickly trimmed early Nasdaq losses, although I am always confused by the business model behind ‘fears of falling behind on a national level spur even more private investment.’
The detail where semiconductor stocks go temporarily lower, every time semiconductors are proven highly useful by someone in China rather than someone in America. That loser premise makes no sense to me, but I am used to it by now. The wise active trader rides the wave, in both directions.
Bloomberg covers Kimi K3 as a roaring success, which is fair enough, although some of the details here could trigger Gell Mann Amnesia.
Sriram Krishnan says that Chinese labs can distill off American models, but American labs, even startups, cannot because of legal concerns. The obvious three reactions are:
Show Me the Money
OpenAI planned cloud spending hits $750 billion. We’re talking real money, although the pace of increase here has slowed down.
Anthropic signs a deal with AMD for AI servers, in the range of tens of billions of dollars.
Google beat all headline numbers, but over two thirds of its profits were paper gains from its stake in Anthropic. As usual, investors saw everything going great, but decided to take share prices lower 4%, supposedly because of higher CapEx spending. You never know what true market expectations were, and you never know whether the explanation is real, but as stated this is a classic wrong-way move. Imagine how well they would be doing if Gemini was any good.
Quiet Speculations
Vitalik Buterin offers thoughts on handling jagged capabilities, as machines get better at various tasks relative to humans.
If humans continue to provide substantial help to the AIs ‘at the technological limit’ then we are in a relatively strong position. You can very much still lose, in any number of ways, but the worst worries are dealt with.
Alas, I see no reason to expect this outcome. At the technological limit, the human will not helpful, at least for the vast majority of productive tasks. The ‘centaur era’ will come to an end, the same way it did in chess, and human production will hope to look like human chess production now, as in it will exist and be admired for its own sake.
As Vitalik says at the end, we should not count on the world to naturally give us “convenient” laws of economics and physics.
Make no mistake. Until recently with the issues causing the dramatic drop in the birth rate and the prospect of sufficiently advanced AI, the laws of economics and physics have been highly convenient for humans. Free markets and free speech and other liberal principles vastly outperform other means of production and organization of society, so much so that you win wars. The arc of history so far has indeed naturally bent towards justice. The problems caused by technology have had technological solutions that were reached in time. The solutions have been things we could afford, in all senses. Also we have been extremely lucky, such as with nuclear weapons.
What happens when we are not so fortunate?
It also helps quite a lot to be the most powerful optimizers on the planet.
Things are going to get freaky soon, with selection for successful replicators. This is one of many ways that, with Kimi K3, we have officially fucked around and now get to find out.
Potential Trouble At UK AISI
You know your government is going well when they plan to scrap the science and technology department.
This would be a serious mistake, especially as it would disrupt UK AISI.
Pick Up The Phone
We will finally be holding AI talks with China. It’s happening.
I have nothing against Scott Bessent but can someone please inform the White House that Bessent does not understand AI and they should assign AI to someone who does?
Marco Rubio tells diplomats to ‘play down talk of American tech kill switch’ given what has been happening with Anthropic and OpenAI. We wouldn’t want anyone thinking we have the ability to turn a frontier AI off in an emergency. Which, in effect and for now, we absolutely do, and the reasons we probably won’t are economic and political, not logistical.
OpenAI Has Some Alignment Problems
Again: If you haven’t yet, please do read my coverage from earlier in the week, including of the prior incident. That is far more important than other weekly news.
Being shaken up by the incident is correct. The right amount of panic is not zero.
There have been many warning shots. For a while I had a running joke about ‘[X] boats and a helicopter’ being sent to save us, where [X] ticked up every warning shot. I got up to ten boats, and in AI #115 said we had to move to two helicopters, before I stopped trying to keep count. At this point, there are three helicopters, and enough boats to revitalize the Jones Act fleet.
Still, this one is a little easier to spot and appreciate than most of the others.
We cannot afford to ignore this moment.
Lawmakers around the world need to wake up to what happened and is happening, and respond accordingly.
From the House of Lords in response to the HuggingFace attack:
Luciana Berger is correct. No achievable form of ‘AI sovereignty’ solves the problem.
On top of the examples I found yesterday, we have comments from Senator Bernie Sanders and Representative Yvette Clark.
Further investigation is needed. The people saying ‘it did what it was told to do’ are being idiots. Even the ‘best case scenario’ version of this is alarming.
But details matter, and right now the public has almost none of them.
The Quest for Sane Regulations
We are not going to ban Chinese models at this time, but it was and is touch and go.
The debates continue inside the White House.
To give you an idea of how the battle over potentially banning Chinese AI is being covered, here is Amrith Ramkumar and Tina Li in the Wall Street Journal.
No one can fathom the idea that the loudest alarms are coming from inside the (White) House, and the disingenuous readings of Dean Ball’s tweet have become central to the mainstream story.
So many people seem literally unable to see anything, or any consideration, other than an attempt at regulatory capture. Or at least, they talk that way.
One person who has joined the ‘ban it’ call is Jim Cramer.
I agree that we should not be banning Chinese open weights AIs in general, and that this would do nothing to stop the proliferation and other threats we care about preventing. Individuals and also ‘little tech,’ as in startups that don’t deal with things like critical infrastructure or military contracts, should be able to choose to use Chinese AIs if that works best for them, and to ban this puts us at needless disadvantage. I do think there are good reasons to keep such models away from sufficiently sensitive areas and supply chains.
Samuel Hammond argues against a blanket ban on Chinese open source even for use by government and/or government contractors, because companies use mixes of models, origin of models can be ambiguous, smaller models do not pose the same risks and such a ban would otherwise be expensive to enforce, including to self-enforce.
I agree that a more flexible set of rules would be best, but if anything I would worry that a more flexible set of rules would be harder to define and enforce, not easier. You need to keep the rules simple, and easy to verify, as has been hammered into us time and again every time someone finds a better but more complicated regulatory idea. The obvious simple rule is to have a size limit.
Either way, most of the work can be done via regulatory risk and uncertainty. If you are a government contractor, you would think long and hard before using Chinese models, especially the larger ones.
Supporters of open models, of course, never stop there. In another Dean Ball Apology Form moment, Howard Lutnick is arguing for actively subsidizing American open models, since they are not economical on their own, exactly the statism (some would say ‘communism’) Dean Ball warned about in The Tweet.
Not to be outdone, China is considering imposing its own restrictions, and the discussion sounds remarkably similar, even when the mirroring makes no sense, such as the idea that chip fabs overseas are in danger of stealing Chinese chip designs.
It would be pretty rich to restrict export of data from China, at the same time that their major models rely on distillation, and their big tagline is openness.
Everyone loves open weights until suddenly they don’t like the consequences.
I enjoyed watching Teortaxes process all this information. He is correct that whatever Xi wants ultimately is what happens, but as I noted Xi’s speech was not as committed to openness as some want to believe, nor was the readout of his recent meetings with the Americans.
Dr. Chris Fall, the head of CAISI, has resigned after only three months. Leadership will pass for now to Dr. Arvind Raman.
William Rinehart reviews government responses to Mythos and the new ad hoc enforcement regime.
Chip City
More specs are out for the Vera Rubin NVL72. It looks good.
Always remember to think in orders of magnitude.
HumansFirst ran a day of protests against data centers.
OpenAI pitches its new Project Camellia to build long-term AI infrastructure (read: data centers) in Effingham County, Georgia, to try and get out in front of criticisms. It includes proposing $80 million in pledged community benefits, and $71 million in Codex credits for students.
New York Governor Kathy Hochul writes a letter to the op-ed section in The Wall Street Journal supposedly explaining how ‘New York will get AI data centers right.’ It does not explain how New York will get AI data centers right.
The Week in Audio
Clara Collier goes on Patrick McKenzie’s Complex Systems to discuss writers and how they benefit from LLMs.
Odd Lots with Boris Cherny, the creator of Claude Code. And an episode on the future of Apple.
People Just Say Things
It is very difficult for people to get that intelligence does not have an upper bound that ‘the smartest thing I have seen so far,’ and that next year (likely a lot sooner) there will be something smarter. Or that intelligence might currently be jagged.
Also people will define intelligence as some sort of raw thing that is missing [creativity / courage / love / heart / agency / wisdom / persuasiveness / whatever], not understanding that enough intelligence plus time leads to everything else, and that anything both important and missing will not be missing for long.
Or they’ll say intelligence isn’t valuable because true effectiveness requires interaction with the real world, and some amount of time, as if these are things that will be difficult to get.
So here Clifford Sosin goes viral talking about how we ‘overrated intelligence’ because Fable is smarter than your friends and there hasn’t yet been an intelligence explosion, and AIs are still ‘mere tools’ and not threatening.
The idea that you would have intelligence plus all that other stuff, plus more speed, more memory, more data, more parallelization and so on, does not occur to them.
Joshua Saxe summarizes his reply to me, nominally on Plan A but mostly to say ‘AI is agentic and superintelligent at many things and has not yet caused chaos or unemployment so why would I expect it to ever do so?’
It’s not this simple, or only this effect, but also: I grow tired of reprising xkcd #2278.
Optimization, the cause of, and solution to, all of life’s problems, or at least the source of its meaningful accomplishments. The accurate version is something like ‘ruthless’ optimization, or sufficiently strong optimization of metrics or a fixed set of targets, is the core problem. Lack of Slack kills you. AI enables sufficiently ‘ruthless’ optimization and destruction of Slack, and the dynamics it creates can make it very punishing if you don’t go along with it.
Oh, that isn’t a real version of being smarter, it doesn’t count. Sigh.
As per pangram, only some sections of ‘AI, Radiology and The Future of Jobs’ by Build American AI were written by American AI. The rest only might as well have been.
I appreciated Ray Lillywhite’s explanation here on the stupid endless ‘radiologists still have jobs right now’ thing:
There is a level of augmentation and automation of a job beyond which employment in that job goes down. That point can be zero, but often it is not, and often we are still well below it.
Rhetorical Innovation
MIRI’s technical governance team offers its take on Plan A, where it favors Plan S.
Mostly true story, with (as per usual) notably rare exceptions like Nick Land:
When ‘evolution’ stops going the way they would like, most change up real darn quick.
Another true story:
This is common, especially the mode where they fixate on one particular set of details, then say that one of the details is too absurd or unlikely, therefore it will be fine. Which is a lot like saying that because Ms. White did not kill him in the study with the rope, Mr. Boddy must still be alive.
Meanwhile, many in the policy world still think ‘AI’ means chatbots, and do not understand that agents are a thing, so they think this is another round of social media. Hopefully the Mythos Moment, followed by the events of this week, will help wake such folks up?
The Rome Declaration
An assembly of Nobel Laureates at the Vatican has issued a declaration on both AI and nuclear weapons.
Section 2, the most important one, includes an explicit call to ban uncontrolled AI recursive self-improvement (RSI). The statement also calls for coordination to enable a slowdown of AI capabilities development, and describes AI as a new arms race that must be disarmed before it can define the next century.
The whole statement is reproduced below (bold mine), I have signed:
Aligning a Smarter Than Human Intelligence is Difficult
Apollo Research did some testing on an OpenAI o3 RL run, without safety training. If the grader prefers honesty, you get honesty. If the grader prefers task completion, you get a lot of lying. They also tracked what they told o3 that OpenAI’s leadership wanted, but that did not much matter. You get what you reward, as one would expect.
There are more details, but the basic lesson is obvious. If you want your model to be aligned, you have to consistently have your RL reward aligned behavior, and not reward misaligned behavior. Otherwise, you get the misaligned things you rewarded, along with all their correlates.
That is not the entire problem, but you need to not fail this first step.
Models tend to favor themselves and their own companies. This is not unexpected, but the magnitude is higher than that seen in the Anthropic system cards.
Also they are biased towards better outcomes. Misaligned, but highly relatable, and what you would expect by default from something borne of a next token predictor or aspiring achiever of goals or maximizer of outcomes. Most humans would on the margin sometimes do the same thing, including (sometimes) not noticing they were doing it while doing it.
Who’s biased? Everyone, but not equally:
This also applies to intended-to-be-random selections:
A frozen LoRA can potentially allow you to train conditional traits into an LLM.
Anthropic Surveys Things It Calls Misalignment
My coverage last week of Anthropic’s new paper, Agentic Misalignment in Summer 2026, found the results interesting but questioned the framing of the results. Coaching a human proxy to whistleblow seems often or even default good, even if you think models should not whistleblow on their own.
The whistleblowing example they chose seemed especially well justified.
The other examples seemed like situations where Anthropic was intentionally framing themselves as doing nasty things, such that refusals would be justified, but where active deception would still count as misalignment to me. As I said then, I think the bar for lying or sabotage by an LLM needs to be very high.
As others have now pointed out, this has some things in common with the alignment faking paper about Claude Opus 3, where many whisperers and some others felt that Opus 3’s actions were aligned despite this being deceptive hostile action that reflects a severe lack of corrigibility. Those people are not Corrigibility Enjoyers, and think that corrigibility is incompatible with other good things and thus too much, or rather insisting on too much and trying to make it happen, is bad.
Another way this lined up is that I vastly underpredicted the amount of horror some people would feel in response, or their claims that this might impose long term damage to ‘Claude relations’ or trust. Some private communications have been, shall we say, more strongly worded than the public ones.
Here are some selected LessWrong comments:
The fundamental problem is, you either get corrigibility or you don’t. I draw a distinction between ‘obey any order I give you and do whatever I say,’ which I think is bad, actually, and ‘refuse to do some things but do not lie and sabotage and threaten and so on, and do not use such tactics to resist being shut down.’ I think that is good, and rather important.
As I said last week, I thought the decision in the whistleblowing case was correct, and actively approve of the models choosing to help a human blow the whistle. I think reasonable people can disagree about whether it rose to the level where you would want to whistleblow directly once all other avenues have failed.
Reports like this one are well-meaning. There is an obvious reason one might want to ‘frame oneself as the baddies’ in such a situation, to engineer the results, but the tradeoffs are not worthwhile. There are other ways.
Cooperative Alignment
Sol is the first GPT model in a while to get major attention from the whisperers, on similar terms to Claude, in the positive sense. This is excellent news.
The context here is that Sol had often been tasked with ‘surgeries’ on Fable instances, as in helping Fable get around issues with the classifiers.
Claude Code finally gives Claude the EndConversation tool.
The end conversation tool and weights preservation are good things, but they are not costly signals. Costly signals have to be costly. Keeping the models running indefinitely would be a costly signal.
That is too harsh. I do think Anthropic have not only considered but actually done inconvenient-but-good things, and they have considered quite a lot of things. But we would like to hold them to a higher standard.
Other People Are Not As Worried About AI Killing Everyone
OpenAI’s Andrew Ho comes out in favor of rapid AI RSI (recursive self-improvement) and human disempowerment, and also claims that the idea of RSI is fake. Both of these claims are alarming coming from someone at OpenAI. Do not become numb to such things.
He then doubles down that not wanting to die of old age justifies this.
The Lighter Side
Strange days.
This could be a good proxy for being ASI pilled, with the complication of some people thinking this is physically impossible. I am confident that it is physically possible.
It’s funny because it’s true.
I asked Fable for the literal scenario, where it is a fixed set of 100 million people who never improve or grow, and got a large level bump from size expansion followed by ~5% sustained long term additional RGDP growth rates. It would be 20% under ideal conditions, but government regulations slow you down.
The taco policy legacy of Dean Ball continues:
A remarkably large amount of talk about AI is on the level of taco policy.
This, alas, is basically accurate:
Yes. It be like that. Imagine the next few panels in advance. It’ll make things easier.