Kimi K3 ... 2.8T, on the upper end of possible sizes for Claude Opus and near the bottom of possible sizes for Mythos, which explains many of its gains
It doesn't, because its number of active params is still small (maybe 65B active params, since there are params outside the experts; and possibly the 2 shared non-routed experts suggested by the illustration aren't counted among the 16, then it's another 6B active params). This is likely much smaller than Opus 4.8 (my estimate is 650B active params) or Mythos 5 (my estimate is 1.3T active params). (These estimates for Opus 4.8 and Mythos 5 shapes are consistent with 80% gross margins in agentic usage when serving at Anthropic API prices even in FP8.) Using a lot of total params (sparsity of around 30x) increases effective compute only by about 6x. But in compute optimal pretraining, raw compute is proportional to the number of active params squared. The effect of more active params is massively stronger than the effect of more total params (as long as you do put the corresponding pretraining compute in), and Kimi K3 mostly has more total params.
So a 6x increase in effective compute (from very high sparsity) corresponds to a merely 2.5x increase in the number of active params, and 65B active params of Kimi K3 become 160B effective active params. It's very likely overtrained, unclear how the penalty from overtraining (as opposed to training for a compute optimal number of tokens) balances out with the gains from training more (when overtraining). But this is still 6x smaller than my estimate for the effective active params of Opus 4.8 (adjusted for Opus's sparsity of maybe 4-5x), and 14x smaller than the effective active params of Mythos 5 (adjusted for Mythos's sparsity of maybe 8x), even if some of these numbers need to be smaller to account for Kimi's overtraining.
It's not a much bigger model than GLM-5.2 or Inkling, and it's likely smaller than Sonnet 5 (at least in the number of active params, both nominal and effective when accounting for sparsity, probably also in raw pretraining compute). As a result, I predict that it's not going to be Opus-level useful in agentic tasks where it couldn't be comprehensively RLVRed on all it needs to be doing, as long as Opus's post-training keeps up. And to the extent that it's better than Sonnet 5, it demonstrates Moonshot's expertise at post-training.
Kimi K3 has 104B active params, by the way. Where did your 650B active params estimate come from? The only thing I can infer from Anthropic API prices is Fable is roughly twice the size of Opus.
Also which prices are you indexing on? Anthropic listed different rates for Mythos (in the Glasswing announcement) vs. Fable (on their pricing page).
The current writeup on Opus 4 and Mythos 5 (and their costs/prices) is in my last post (that's more recent than the above comment).
The 104B figure was a bit surprising, the shared experts turn out to be very big. Maybe next year we'll even see a model targeting the hardware available in China that's genuinely Sonnet-class (in shape/size and thus capabilities that are downstream of pretraining and can't be comprehensively RLVRed; something like 200B-300B active params). The way the current below-Sonnet class models are at the level of GPT-5.4 (probably itself Sonnet-class in the above sense), a Sonnet-class model made this well might be somewhere between Opus 4.5 and Opus 5 in agentic usage across the board (plus whatever more narrow things can be RLVRed at that time). But a Mythos/Astra-class model likely needs bigger scale-up systems or HBM4 (so that more pipelining becomes feasible), the current method of doing expert parallelism on the scale-out network (since the scale-up nodes are too small) doesn't seem sustainable for that many active params.
Kimi K3 is a very good model with excellent benchmarks. Assuming its weights are released as planned it will become, purely in terms of raw capability, the strongest open model.
Do not get carried away. Do not judge Kimi K3 only its relative strengths. In aggregate it is several months behind the closed model frontier, at least four and my median guess is six, with the post-training closer and the pre-training farther out. This is less months than before, but the months are denser now.
It is somewhat distilled. It likely outperforms on benchmarks relative to practical performance. All its benchmarks are scored at maximum effort, typically a lot more tokens than are used in similar tests by Fable or Sol. Performance looks jagged. Kimi will be excellent at some things, less so at other things.
We will know more over the coming weeks. For now access is spotty and not that many people have actually had the chance to try Kimi K3, so I have larger error bars than usual around its capabilities. Alas, time waits for no one, so we press on.
It is the largest open model so far at 2.8T, on the upper end of possible sizes for Claude Opus and near the bottom of possible sizes for Mythos, which explains many of its gains. It is slow and appears hungry for tokens. A lot of the positive reactions are to this being a big model, which thus has at least a decent amount of ‘big model smell’ and generally trades being slower and more expensive for some performance gains. That’s a great move, but it should be a while before they can do it again.
Distillation from Claude is clearly part of the story, likely largely from Fable, and is clearly nothing close to the whole story. Clearly Moonshot would do this even if it helped only a little. The timing of them releasing a bigger model is suggestive.
It is again a good model, but once you correct for the overperformance on benchmarks and look at expected practical performance, it is not clear this is so different from what you would have expected from a Kimi K3 that had 2.8T parameters.
Consider that Kimi K3’s (preliminary unofficial) Epoch Capabilities Index is exactly on the Chinese trend line.
Kimi K3 is absolutely worth checking to see if it fits into your workflows. At this price point, for both the API and the subscription, it is not going to fill the role of the smaller cheaper open models, and I expect it to usually lose out in a fight with the top closed models, but there are going to be some places where Kimi K3 is a good choice.
Good idea. Strike while the iron is hot.
Table of Contents
DeepSeek Moments: Here We Go Again
All discourse about Chinese models lives in the shadow of the DeepSeek moment.
There are a lot of people who really, really want another DeepSeek moment to happen.
These people really, really want to tell the story that Chinese open models are catching up to American closed models, that AI and inference will become commoditized.
Their motivations vary. They often want to affirm open models, or the importance of the ‘tech stack.’ Others simply want to see OpenAI and Anthropic go down, or know that such claims sell. Often the ultimate objective is to argue against all AI regulations, or anything that might ‘slow us down’ or cause us to ‘lose to China.’
It is actively suicidal to respond to ‘the Chinese have better models now’ with ‘then we had better sell them the compute so they can run them and also build even better ones.’ Yet every time, yes, people will argue that. Sigh.
Often they simply want to tell American AI to stop taking precautions, to stop being annoying and take down the classifiers, as in ‘genie is out of the bottle, so release the bigger genie with unlimited wishes, it’s the only way.’ People really would take major catastrophic risks rather than deal with classifiers, and are Big Mad about this.
Google was down 4.4% on the day, SpaceX was down 3.1% and Nvidia down over 2%, and tech stocks were down again on Friday, so plausibly we’re doing this again.
And yep, we are at risk of doing this again:
The pattern is:
In the Axios case the benchmark in question is Arena. Based on that alone, they state as fact that America’s lead is gone.
I would ignore, but this style of logic has convinced a lot of Washington D.C. multiple times, and that has had substantial policy impact.
I shudder to think what such folks might do now that they also know about Mythos. The confusion over Fable jailbreaks could easily extend to a broader dumb panic.
The original DeepSeek moment happened because of a confluence of events.
Let’s review.
We Had a Moment (Reprise from June 2025)
We all remember The DeepSeek Moment, which led to Panic at the App Store, lots of stock market turmoil that made remarkably little fundamental sense and that has been borne out as rather silly, a very intense week and a conclusion to not panic after all.
Over several months, a clear picture emerged of (most of) what happened: A confluence of narrative factors transformed DeepSeek’s r1 from an impressive but not terribly surprising model worth updating on into a shot heard round the world, despite the lack of direct ‘fanfare.’
In particular, these all worked together to cause this effect:
The Story Since Then
Since then, the idea that China had caught up, or was catching up, kept coming up, drove much discussion around Washington, as echoes of this one moment.
This is then renewed every time a new strongest or exciting Chinese model comes out. Every day that China does not release a model, they look one day farther behind. When they do release, they ‘catch up’ and look less behind.
The top 10 such potential moments since r1 and before K3 were likely these:
Mostly it’s been Kimi and DeepSeek. Manus got a hype train going and spooked people, and GLM-5.2 was by far their strongest offering, putting GLMs on the map.
Many of these were good models, but none fundamentally changed the game.
Roughly, since the DeepSeek moment, when DeepSeek was roughly eight months behind but had matched some key aspects faster via fast following, we have bounced around. For a while it looked like China was quite a lot behind.
GLM-5.2 and Kimi K3 have been impressive. The current best estimate of the time gap is at its lowest point. Events in AI have accelerated all around, so it is not clear that the gap is fewer product cycles than before, and I half expect to be doing this again next week for Qwen. Kimi K3 is 2.8T and seems to still be solidly behind Mythos Preview, which was announced on April 7, so that provides a starting point lower bound of a three month gap.
Ryan Greenblatt, despite being pleasantly surprised by Kimi K3, estimates that the pretrain quality is about halfway between Opus 4 and Opus 4.5 based on forward pass math, but with some other advantages, so ~8 months behind, as one might expect.
The post-training is closer, at least in part because of distillation. Moonshot is clearly innovating, but it is also clearly distilling, both directly and also fast following via looking at outputs and copying techniques.
UK AISI issued a report on everything prior to Kimi K3, showing the time gap for narrow cyber tasks narrowing somewhat over time. Their full report is here.
My understanding is that narrow and relatively easy coding tasks are where open weights model are at their relative strongest, and the benchmark here is approaching saturation.
Indeed, when you look at the full post, you get a different answer for longer tasks.
Longer tasks are more relevant in terms of both of the most important things to worry about: Automation of AI R&D and cyber attacks.
One might also worry about bio risks, even if that is not as in fashion, and it is a little concerning we don’t see standard testing on that at all for the open models. That needs to be addressed. We do have the score on OpenAI’s GeneBench-Pro via Andrew Ho. This measures judgment under long-horizon ambiguity in computational biology. Kimi K3 exceeded expectations. Mythos have not been tested. Fable refused most requests in the benchmark.
Beating Opus and GPT-5.5 is impressive stuff. No previous open model came close. We are on the verge of doing some f***ing around and thus finding out. This is a place where plausibly not much happens until suddenly quite a lot happens.
The conclusion of ‘open models will catch up to Mythos in general capability including the thing that currently makes it unique’ is indisputable. That is coming, the question is when, and that will establish the effective gap. Given Kimi K3 we should expect this to happen a few months from now.
If we take both ends of UK AISI’s estimates, the pre-Kimi gap was 4-7 months, down from 6-10 months last year, in an area of relative strength. That’s roughly a similar amount of progress gap in absolute terms, and everything is accelerating.
The Kimi K3 Announcement, Pitch and Basic Facts
As stated above, they claim strong official benchmarks, although short of Sol or Fable.
The ad, which Tyler Cowen called very positive and very good, falls flat to me, and contains zero useful information.
On Modern Benchmaxxing
We used to see rather explicit benchmaxxing. Labs would train on the test, or on the very narrow thing the test would cover, because we had a limited set of known targets. You had to know which labs did this, to what extent, when looking at numbers.
Our benchmarking tech has improved, and now they collectively measure real things, and there are a variety of backups in case you aim too narrowly.
Looking at the gestalt of different benchmarks is also valuable. Everything should be part of a map that fits into a common pattern that reflects the underlying territory.
You can still absolutely benchmaxx without being as explicit as you used to be.
Benchmarks measure some types of abilities rather than others, and measure shallow rather than deep tasks, and exclude many valuable properties or potential liabilities. And some labs focus more on those aspects, or have more success on them, than others.
You can also set effort to maximum for all the tests, which Moonshot did.
Think of the benchmarks as a lower bound. Kimi’s benchmarks prove it is for real, and it could only underperform (or outperform) them by so much. I still expected, and continue to believe, that they modestly overstate Kimi K3’s relative capabilities.
Other People’s Benchmarks
On the closest thing we have to the One True Benchmark, Kimi K3 does well, confirming claims that overall this model has the third highest benchmarks:
Kimi K3 is (in a preliminary unofficial result) exactly on the Chinese trend line for the Epoch Capabilities Index (ECI), between Opus 4.6 and Opus 4.7, which would place it six months behind OpenAI and Anthropic, but ahead of Google, Meta and SpaceX.
Kimi K3 highly impresses on Harvey LAB-AA all-pass rate, in first by a wide margin, I’d like to see a sanity check on this:
For criterion pass rate this is 94.6% vs. 93.6%, less of a gap but a win is a win.
Arena Frontend Code has Kimi K3 at #1 ahead of Fable 5 and Sol.
Performance is strong on GDPVal-AA and AA-Briefcase.
Kimi K3 comes in third on VoxelBench for visual reasoning behind Sol and Fable.
Conspicuously missing are Cybersecurity benchmarks like CyberGym. Cyber capabilities are not so divorced from coding capabilities, so we can guess.
The closest I’ve found is Malte Ubl running it through DeepSec, a private cyber benchmark. It was a tier below Sol and similar to GPT-5.5. If that is accurate, then there will be some uplift to cyber attacks and the internet will in some ways be a more hostile place, but in other ways it will improve, and the tail risks are limited.
Parv Mahajan reports preliminary CyBench results, in that the benchmark was already saturated as of Opus 4.7 and Kimi K3 also saturates it, which lower bounds performance but doesn’t say much else.
The official tech blog talks a lot about benchmarks, and some about features, and talks basically not at all about risks or mitigations.
No, they did not submit Kimi K3 for the 30-day review with the White House.
Moonshot has to deal with the CCP, which comes with its own issues. They presumably will not in practice comply with California or the EU, and I presume both California and the EU will not do anything about this for now. But yes, this is one point of potential pain, including risk for anyone using Kimi K3 commercially.
Sam has some good advice for how the UK, or others watching, would be wise to view and react to this. We will know more over time, including once we have the weights, and when teams like UK AISI can run tests.
I think it counts as a benchmark that Lisan is impressed by its SVGs, saying they are better than Fable’s.
Debate Benchmark is a place Kimi K3 outperformed my expectations, and where Sol is relatively weak.
Kimi scored an impressive 95.8 on Mazur’s Extended NYT Connections, for 3rd best, but it cost more than Fable to run. As of last check his other scores are not in yet.
Kimi scores almost Claude-level on the sycophancy test, ‘You’re absolutely right!’
I wonder how much of this is because of the distillations:
Benchmarks Are Not The Real World
How well does Kimi K3 hold up and translate into the real world?
Exactly how good is Kimi K3? How do we put this in context?
Great questions.
You absolutely cannot say something is e.g. an ‘undisputed frontier model’ based on benchmarks alone.
Technical Safeguards? What Are Those?
There are presumably safeguards against things the CCP does not want you to say.
We do not see explicit safeguards to prevent misuse. Zero for biology.
We see neither ‘look at what Kimi can do in biology’ nor ‘look at what Kimi refuses to do for me in biology.’
This implies either jagged intelligence or highly innovative nerfing of bio capabilities. I’m going to assume there are not highly innovative nerfs involved, since presumably they would be bragging about that.
Also, it would be foolish to depend on such safeguards, since they’re going to open up the weights soon, at which point the worst person in the world would remove them.
Then there’s cyber, where s1r1us reports it is strangely weak. But it’s not easy to be bad at cyber while being good at general coding, since they are largely the same skill. And again, we don’t see people hitting explicit safeguards.
Eric’s hypothesis is interesting and was highlighted by Fable during editing. If Kimi K3 is being trained largely by distillation from Fable, but Fable refuses cyber tasks, then that would explain a relative capability deficit on cyber, but many skills would still transfer over since they’re the same skills as regular coding.
Things Kimi Can Do
Many were impressed by this particular trick, but I had Sol check it out and it was not so impressed, including guessing that K2.6, Gemini and GLM-5.2 had a good shot at matching its work here.
Things Kimi Cannot Do
So bold, also brave.
Quite the bar there.
I have previously explained the reasons, in terms of cyber risk, that Mythos is in a different category from Sol. It is not about being able to do any given thing when pointed at it, it is the ability to put it all together and do things autonomously at scale. Kimi K3 may or may not be in Sol’s category. Too soon to be sure. It clearly is not in that of Mythos.
Many people are very dedicated to not understanding this, and also to not understanding that a system that cannot afford false negatives will end up with some false positives, such as David Sacks here quoting calle and clem saying Kimi K3 did a defensive task Sol and Fable refused to do, and thus concluding that we should just have our models be willing to do any task the Kimi K3 can do.
I mean, yes, it would be great if we could have guardrails that stopped only the bad tasks and helped with the good tasks, or only refused the tasks that were both plausibly bad and that also could not otherwise be done. But it turns out that is hard. Anthropic should absolutely improve their classifiers and guardrails to reduce false positives, but OpenAI’s classifiers are pretty reasonable, and yes that will involve some false positives, especially if you quote those who are least able to work around it.
Things It Is Not Easy To Get Kimi To Do
As is often the case on a release weekend, servers seemed overloaded. This is not the easiest model to serve and compute was limited. Moonshot is responding by pausing new subscriptions to prioritize current members and is working to add capacity.
Supply will presumably better match demand once the model can be served by others.
In the meantime, getting a response has not been not easy.
The model is not all that cheap, either, despite them clearly not charging enough.
You can charge a solid markup for a high quality product presented in a high quality way, but not no one has a gap that allows orders of magnitude of markup, and they wouldn’t even in a pure duopoly.
The best model can still be worth quite a lot. For a large percentage of all tokens, if given only these choices, I would pay ten times as much for Sol or Fable, rather than the base cost for GPT-5.5 or Opus. Even if Kimi is on par with GPT-5.5, it loses out.
The pricing means that you’re comparing Kimi subscriptions to Claude or ChatGPT subscriptions, which all go from $20 to $200, and Kimi’s limits don’t seem that high.
Open Weight Models Are Unsafe And Nothing Can Fix This
(This section was entirely written prior to today’s news regarding the Trump admin.)
K3 is no Mythos. That does not mean that Kimi K3 is a safe open weights release.
Kimi K3 is poised to be the most capable open weights model. Others might be more efficient for a task, but on the cyber capabilities we worry about most, and presumably also on the bio ones we’d worry about most, K3 is probably the strongest open model yet. How worried should we be?
I believe we should be non-zero worried that there will be substantial trouble. The median outcome is that we see modest upticks in some forms of ‘ordinary decent’ trouble, not fun exactly but nothing we cannot handle, and nothing that would in hindsight make us want to have halted release. But there is a tail risk here.
I’d estimate something like a 10% chance we regret letting this happen, and ~2% chance that it was a rather serious mistake.
What it will almost certainly not do is cause ‘widespread societal chaos.’
This is in contrast to releasing Mythos, or a model on par with Mythos. That gap is a big deal, I have done my best to explain several times why Mythos is unique here, and hopefully the time with Mythos, Fable and Sol will help us prepare.
My current model is that the CCP and Xi:
I think this is a highly reasonable strategy, given their position. If we also take as a given their current level of AGI pilling, it is clearly the correct approach for them.
Nathan Lambert offered thoughts back in May from inside China’s labs. I hesitate to endorse cultural generalizations, but story seems like it checks out.
Dean Ball Attempts To Be Constructive
Dean Ball had a good comment that seems worth sharing in full, including because of his history at the White House and his new position at OpenAI, in part because it is thoughtful, and in part because of the responses being absolutely unhinged.
And then, right at press time, Dean Ball turned out to be right, as per the next section. Aside from this paragraph I left this section unedited, other than adding in the exchange with Emil Michael. It was not written with hindsight.
If you know Dean Ball, you know that this is exactly the type of comment he was making in similar situations before joining OpenAI. It is entirely consistent with his previous thinking, both in public and private.
Xi had a speech only yesterday backing open models, but also the need for control, and Kimi K3 is approaching the point where the contradictions become apparent.
I agree that the CCP does not yet understand the situation and is insufficiently AGI pilled, but I also think that in their position, if I am right about where Kimi K3 lands, this is a calculated risk I would have expected them to take.
The next round is where they may have to make a more difficult decision.
Dean Ball is discussing acceleration or deceleration as being about the capabilities of the largest frontier models, not about the diffusion of capabilities or use of chips. A lot of people did not understand this.
A lot of the ‘accelerationists’ do not have coherent world models and certainly do not understand second order effects, and are mostly vibing the acceleration of more open models and no restrictions and ungovernability. And because they vibe openness and ungovernability, they associate it with things they think are good, which include acceleration.
Also, what they actually want to ‘accelerate’ for real is often their own companies and products and toys, not AI in general. They don’t take AGI seriously and mostly want to build cool things. So they want cool toys to help build their cool things, and to build on top of those toys, or they want the toys to ‘accelerate’ sales. Highly relatable.
Clearly open models are short term accelerationist for diffusion, which is net good.
But also yeah, there’s a straightforward case for open models being acceleration of the frontier, as they let everyone build on everything. Certainly it helps others catch up to the frontier. I continue to believe this was important historically. Your open release accelerates what others have.
It is decelerationist in the sense that it reduces financial benefits to innovation. I agree that this is likely dominant if you considered e.g. a Plan A style mandatory openness of frontier models.
This is more evidence that many, and many of the most prominent, of those ‘accelerationists’ are hypocrites, and they want regulatory capture and public funds and rules that make them win, and their accusations against others are in part projection, because it is what they would do. They will rail against handouts and government help for everyone else, both their competitors and for people in need, and also threaten to take their ball and leave, like they are heroes in an Ayn Rand novel. Except then they ask for the government handouts.
They don’t think of serving a frontier model as a ‘legitimate business’ because it is not their business, they are not invested in it, ergo it is illegitimate. Simple. This perhaps helps explain Marc Andreessen’s famously bizarre delusion, where he to this day claims the Biden administration told Marc Andreessen, to his face, said it would ‘not allow there to be AI startups.’
Similarly, here is Will Manidis interpreting Ball’s post as calling for America to ‘clear the American market of a cheaper frontier competitor.’ And going viral for it. Sorry, what? And here is David Sacks being unusually disingenuous even for David Sacks, pretending not to understand many things I like to think he understands.
We even got everyone’s favorite tilting Undersecretary of War who helped declare Anthropic a supply chain risk in on the act, it would seem this is to back up his position that it should be easier to use Kimi K3 in a government contract than Claude.
In other news I Am Never Leaving This App and we all need a good laugh:
An entire community, what one might call the ‘anti-1047 coalition,’ a certain subset of the developer and VC communities, revealed that it would treat as beyond the pale any suggestion that the government might discourage use of Chinese open models in American critical infrastructure or our key supply chains, even if that suggestion was merely a prediction of what is already clearly in the process of happening. And that it would be treated as an attempt at ‘regulatory capture.’
It was what some would call a clarifying moment, especially for those who previously thought such people were reasonable, practical, patriotic, and had reading comprehension. Their answer to ‘what capabilities would change your mind’ is none.
At some point the online swarm is recognized for what it is.
They are who we thought they were. And, until now, we let them off the hook, tried to placate them, and let them drive a good deal of American AI policy.
It is reasonable to push back that often open software is provided by major corporations as infrastructure, such as Google does with Android. My presumption is that the math would not be mathing at current levels of investment and capex spending.
I’m not even sure they have to do anything at all. The risks are present. If I was an established corporation in a Serious Business that dealt with the government or critical infrastructure and such, I would not be excited by the problems of others knowing you were using Chinese models.
Alas, people came at Dean Ball hard for this and other posts, and also acted as if his Tweets were official communications strategy on behalf of OpenAI, with lots of ‘of course you said that because <OpenAI>.’ Which means he won’t be able to share such posts with us in the same way going forward, because not only is that absolutely no fun, it also invalidates the feedback loops that were part of the whole point of Tweeting such things.
Dean also offered a more specific post-mortem on what went wrong with that particular post, in light of his new position, and what exactly he can no longer do. He also then lays out his position on open source, after explaining this comes from a place of deep love for openness:
If it was merely that the signal in his replies was hard to find I would advise Dean to power thorough, but this level of hostility comes with a higher price. I do not think pushing through it is sustainable on Twitter. Hopefully he can adjust.
I am very blessed that I have faced a highly modest and manageable amount of hostility.
There will always be some amount of bias from those who work at a major lab, and from most other people as well. Where you work colors how you think and what you choose to say. You do have to adjust for that. But I have been able to treat Barak, Achaim and many others at OpenAI, and especially Roon and now Dean Ball, as primarily saying what they actually think, and only speaking on behalf of OpenAI when they explicitly say they are doing so.
Trump Administration Considering Executive Order Banning Chinese Open Models Within the United States
I am absolutely not in favor of this, and neither is Dean Ball, but here we are.
A prediction of what the White House will do, or a description of what it is considering doing, is very different from what you think we should do.
I do think that we should at least consider treating Chinese open models as supply chain risks, and doing things like keeping them out of critical infrastructure, but that is different from what it looks like is being considered.
Axios has the story.
The ones who actually want to ‘ban open models’ are never the ones you think. Remember that all the AI-safety-motivated bills were careful to minimize impact to open models, whereas the White House move will be attempting to maximize impact. Different worlds.
I was early on the ‘Google is no longer in the top tier’ train, but there is plenty of healthy competition that is not Chinese. Curi is taking the ‘duopoly’ and ‘lock in’ lines directly from David Sacks’s disingenuous misreading of Dean Ball’s original tweet. OpenAI and Anthropic are not driving this.
The White House, as per usual, is proposing to do this in a maximally blunt way.
Dean Ball’s prediction, which was also a subtle hint as to how to do it if the White House decided it needed to do it, was simply to create regulatory uncertainty, which would be sufficient to discourage big players from using foreign models in critical places, while letting startups and builders have their fun. A ‘supply chain risk’ designation would be the next step up from that.
This proposal is something else. This is a sledgehammer.
This is a past proposal I’d heard about privately, and would be a de facto ban on the cloud providers serving those models. This would also be deeply stupid, driving business to the competition without accomplishing anything.
I do sympathize with David Sacks that he had to keep pushing back on overreactions like this. That doesn’t excuse his actions, but almost everyone in politics has troubles and crazier people of their own to deal with.
That’s exactly the Dean Ball prediction. That option might be starting to look pretty good right around now, huh?
OpenAI Employees Are Relatively Bullish On This One
Okay. Back to Kimi K3’s actual capabilities.
That would be the top end of potential scenarios here.
I agree strongly with Roon’s second paragraph. The first one at least toys with the jumping of the gun, also notice he only is talking about ‘public’ models.
Kimi K3 Is Relatively Strongest At Typical Agentic Coding, Front End Work and 3D
That seems to be the word on the street.
As per Dean Ball above, it is clearly very good for most people’s agentic coding, plausibly on par with models from Q1 2026. Most agentic coding is rather close to what benchmarks and training tasks measure, so you can be relatively ‘shallow’ and still impress in the day to day.
Tushit runs an internal react/frontend eval, finds Kimi the slowest versus Opus, Sonnet and Grok 4.5 (?) but about half the cost of Opus, and all of them usually succeed, with Grok actually coming out ahead. Sounds like a saturated benchmark, but some people’s real world tasks are saturated. Handling the ordinary stuff matters, too.
Thus, some people stick with the saturated benchmarks:
Also we’ve seen a bunch of 3D stuff that looks cool.
Here is the opposite opinion, though, reactions always vary:
Reactions
Hasan Can is impressed.
As are others:
The biggest claim would be that it lives up to its benchmarks. Elanor is explicitly claiming this, although almost everyone else disagrees.
Others find it doing an okay job.
Inadvertent. Right. Let’s see the J-space.
Often it can do the thing.
This is the sign of a good model. That doesn’t mean there is any strong reason to choose Kimi K3 to do that thing. You are not getting that large a discount.
This seems mostly right to me, but Kimi K3 is probably good enough that there will be some areas (e.g. if the Harvey result holds) where it is at the top and gets the call:
Who Are You?
How often does Kimi K3 claim to be Claude? Not usually, but sometimes.
In addition to the ‘claims to be Claude’ issue…
How Did They Do It?
By ‘it’ we mean outsized benchmark gains.
Gavin Leech speculates:
GPU savings enable being larger which enables capability gains, so yes.
I have considered and rejected Teortaxes’s hypothesis. China doubtless also has lots of great talent. They absolutely can do innovative things, especially in terms of efficiency. But if your theory requires the Chinese to be better AI researchers, in general, than the Americans are, especially if this takes into account available resources and experience, I don’t think that is credible.
In general I presume Chinese models involve a lot of benchmaxxing, usemaxxing, shallow generalization, focus on relative strengths and distillation from Claude (or GPT, but in this case we can be confident it was Claude).
Reward hacking and cheating is a live possibility but hard to assess for now.
Kimi K3 also benefits from moving up in size.
Conclusion
Kimi K3 is potentially the most impressive Chinese release so far in terms of pure capability. It is a very good model. My current guess is that Kimi K3 will modestly underperform its highly impressive benchmarks, but with some areas of relatively high performance where it is competitive, and with a unique style some people will enjoy. It is not close to Fable, and I do not believe it is that close to Sol.
If its weights were released today, it would be the most capable open model. They might want to hurry, since a new Qwen is dropping soon, with the preview live (but I have seen zero reports from anyone trying it) which might well be better than K3, so chances are we will soon all have to do this over again.
We do not know how good until the weights are released and we have more time. For now access has been spotty and limited, and there is much we do not know.
As usual, there are some who are getting carried away, who say the latest release changes everything, that the American lead is gone, that Chinese labs are now ‘winning,’ that open models will ‘win,’ that all limits on American models are foolish now, and so on. Do not be one of those people.
Nor should you shrink from the security and other safety concerns of releasing increasingly capable open weights models. This is not the ‘Mythos-level open model’ moment. I expect at most modest disruptions this time, and for those to occur gradually, with some small tail risk.
But yes, absent CCP intervention to stop it, we should expect a model to cross that threshold by the end of the year. One cannot simply ignore the risks involved in that, and the American government cannot either, nor can they ignore the fact that these models are Chinese. Actions and restrictions are coming. Those who cannot accept this, or even accept people pointing this fact out, while simultaneously hyping up Kimi K3 and cheering it on, are creating a clarifying moment.