Astra is an excellent model. The jump from Sol to Astra is larger than the jump from Fable 5 to Fable 5.1. This is a big deal.
Astra is the best model for what one would broadly call ‘ambitious projects,’ and likely has the highest raw intelligence factor of any model. These are the largest jumps.
It is amazing at doing things in 3D, or anything involving games. Astra also excels at computer use, and at subagent coordination.
Many benchmarks show dramatic jumps from all previous models. Where Astra is good, it can be in a league of its own.
That does not mean Astra is in its own league across the board. Fable 5.1 is still a Claude. Astra is still a GPT. If you have a strong preference for one over the other, that still applies. For many purposes, especially involving back-and-forth discussions, Fable 5.1 is still my top choice. Fable remains my primary editor.
If you want the best answer to your questions, you should ask both models.
Regular coding is getting less of a focus. Astra is not a quantum leap there, but of course it is very good and makes progress over Sol.
This is the first time a debate over whether a model ‘was AGI’ felt non-silly. I do not think it is AGI, and I would warn against the dangers of using that label prematurely, but I would not laugh at you for disagreeing.
This is also a strange situation in that OpenAI has already soft announced that they have an internal model a level above Astra, as I will cover when I address Navier-Stokes.
roon (OpenAI, September 3, 2026): I have not come close to discovering the limits of what Astra can do.
I imagine it’ll be obsolete in the order of weeks somehow.
roon (OpenA, September 9, 2026): it didn’t even take a week.
My recommendation is that you use both Fable 5.1 and Astra on your most difficult questions, and experiment to see which things each one does best for you.
This is the core thing he calls for, including a unilateral commitment:
The steps are:
Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.
Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.
Global Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.
I will have full coverage of that next week. Along with that essay, the current queue includes at least that, Navier-Stokes and Astra-2, Anthropic’s Misalignment Report, Anthropic’s Countering Misuse, a thinkpiece on personal AI and the law, a post called Claude Talk, and Fable 5.1 (and Astra?) Model Welfare.
Axios: OpenAI says Astra can lay out a printed circuit board in KiCad, build a 3D city scene in Unity, create an animated automobile transmission in FreeCAD and Blender, and fill out a tax-return draft from a W-2.
In scientific work, the model helped improve a mathematical result on gaps between prime numbers and set new marks on several biology, chemistry, medical and physics evaluations.
Dean W. Ball (OpenAI): Astra is a remarkable piece of technology. Earlier agents often tried to dampen my ambitions—they’d push me to do “pilots” or “proofs of concept.” Then agents started meeting my ambitions.
Astra is the first agent that routinely raises my ambitions. I encourage you to try it!
I think it is a very very good writer, in a “big model smell” kind of way.
roon (OpenAI): not to sound like a total shill but it’s the long weekend and all I want to do is make astrodynamical visualizations and stuff with Astra I have Astra psychosis
Tibo: Astra was probably our biggest competitive advantage while it wasn’t generally available.
Since we’ve had it our productivity jumped so much that we shifted some of our plans 6 months ahead and will ship them at DevDay instead of mid next year.
Dominik Kundel says Astra can do all the things in Codex: Use all your apps, do more of the product thinking, excel at Blender, impress you without Max thinking and keep checking its work. He is very impressed.
Here is the pitch from Astra itself, according to Pangram:
Sam Altman (CEO OpenAI): GPT-6 Astra is here. We hope it will begin to enable a new generation of entrepreneurship, scientific discovery, and building.
We believe it is the best model in the world for computer use, professional work, science, coding, cybersecurity, and more. It took us some extra time to ensure that we could meet the safety and alignment standards required for this capability level, but we think you’ll find it worth the wait.
It scores 98% on FrontierMath Tier 4, 99.9% on ARC-AGI 3, and 100% on ExploitBench.
Mark Chen: GPT-6 Astra is here! This is a big moment for our research team – years of work on pretraining, reinforcement learning, and post-training have come together in our most capable and aligned model yet. It can build and test software, work across apps on your computer, and even help you take a crack at open scientific problems!
Capabilities that felt like grand challenges a few years ago have become tools people can actually use. One example is Computer Use – if you’ve tried this before and felt like it was too slow or not good enough, I encourage you to give it another shot. We’ve come a long way since Operator, and it “just works” now.
We’re also asking these systems to act on your behalf for more consequential work. Agents needs to stay aligned with your goals and values, think transparently, and respond to oversight even when tasks become difficult. We’ve made substantial progress on these behaviors in Astra, alongside stronger monitoring that can stop potentially unauthorized actions. That work is part of what makes this release possible.
I think alignment is one of the most important research frontiers in AI, and it remains far from solved. Our ability to understand and align models has to keep pace with model capabilities. We want to give people more room to think, build, and discover with increasingly powerful tools that remain *under their control*.
Huge thanks to the researchers and teams who got us here. There’s a lot more work ahead, and I’m incredibly excited about what we can make possible in the near future!
As is often the case, scientific discovery was highlighted.
Noam Brown (OpenAI): Of all the use cases for GPT-6 Astra, I’m most excited for scientific discovery. We at @OpenAI have not pushed it to its limits on math and science. I look forward to waking up every morning and seeing what new scientific breakthrough someone has made with this model!
They show off a variety of cool demos and capabilities: Acing the Financial Modeling World Cup and navigating spreadsheets at 4x human speed, doing PCB in KiCad, creating a 3d model of a car transmission, filling out a 1040 and so on. They demo ordinary tasks. It’s all cool, but we lack comparison points.
The professional work pitch is that Astra handles complex tasks and adheres closely to templates and instructions, especially when creating presentations. It can translate images into identical-looking spreadsheets, yay.
They highlight Blender 3D models, which many others were also impressed by, see the section In 3D. Game creation is also confirmed as super impressive.
The headline price is $10/$50 per million input and output tokens. Cache writes are $12.50 and cached input only $1.
If your prompt has more than 272k input tokens, prices on input double, and prices on output go up 50%.
Fast mode costs double.
Fable 5.1 has the same $10/$50 headline price, but its cache reads are only $0.25.
Practical costs come down to token efficiency. Artificial Analysis thinks Astra is only about 60% more expensive than Sol in practice, and that it is a lot cheaper than Fable 5.1. They used Fable 5.1 in Max mode, which is probably the issue there as Max mode is usually a mistake for Fable 5.1 (AIUI) for tasks other than discussion.
Unnecessary Overstatement
Astra is a great model with some outstanding benchmarks. It is unfortunate that OpenAI still felt the need to play fast and loose.
This is a rather bad chart crime, and totally unnecessary because Astra scores a highly impressive 62.7% using the standard harness. They themselves note that Sol likely would score ~30% using the Astra harness. Fable estimates that if Opus had used a similar harness, then Opus would have scored ~80%, and I think we should check.
ExploitBench is even weirder. Why highlight the 100% score? That is not one of its more impressive benchmarks. If you score 100% on ExploitBench you cheated on ExploitBench. At minimum this involves data contamination, which is still cheating. OpenAI says as much in the system card. And again, there is no need for such overstatements.
They also used the ExploitGym honeypot as their main illustration of Astra being their ‘most aligned model.’ As I discussed when I analyzed the model card, that is not what this result tells you. You can make a case for Astra being more aligned than Sol, for most purposes I agree, but this is not the way to make that case.
It was very frustrating trying to figure out when I finally had access, including because OpenAI hides model selection behind multiple clicks.
Theo – t3.gg: fwiw, don’t love that OpenAI “launches” aren’t actually launches, and the real world availability date is an unknown amount of time down the line.
We should all be able to play with this new model together right now, feels weird that only certain people have access.
I only successfully accessed Astra on Saturday morning, and then only on the desktop rather than the web.
Official Benchmarks
Here are their benchmark charts (after I removed Gemini for readability):
Alignment is now a category of chart. I would take this one with lots of salt at best given how little we know about their internal marks and the known issues with ExploitGym honeypot and Impossible ExploitGym (which is clearly not so impossible):
Or, here is a full comparison chart, Astra wins FrontierModelBencharkChartBench, although its method might have incremented OpenAI’s lead in FelonyBench:
In the cost-effectiveness charts, Terminal-Bench Science 0.1 looks very good.
FrontierMath Tier 4 (v2) looks great too:
As does Terminal-Bench 4.0 and AutomationBench:
There are more similar graphs: Agents’ Last Exam, ScreenSpot-Pro, OSWorld, BenchCAD, BrowseComp, OpenScore String Quartets (?!) and some internal marks.
Astra impresses on the Epoch Capabilities Index, scoring 169. This is directly on the OpenAI trend line here. Fable 5.1 was below trend and scored 163, same as Fable 5.
The score on ARC-AGI-3 is legitimately impressive:
François Chollet (Creator of ARC): GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations — essentially a game-specific algebraic notation.
Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses — so harness capabilities are increasingly shifting into the model itself. We see Astra as a major breakthrough in model intelligence.
In the Provider Adapter harness, Astra (max) used fewer actions than the human baseline on 96.0% of levels and used 51.7% fewer actions per level on average. This is a material milestone. This means by ARC-AGI-3’s measure of action efficiency, Astra matched and surpassed human parity.
When we released ARC 3, I got asked, “when do you think a frontier model will saturate it?”, and I answered “in about a year, though it depends on how much it gets explicitly targeted”
That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.
In Epoch’s EBR-Bench, Astra got 100% on its second run, outpacing the best human who required five attempts. After removing a ‘game breaking card’ the top human was able to do better over time, because Astra’s learning maxed out over time, but Astra is still far ahead of other models here.
Astra did fall short of Fable 5.1 on MirrorCode, scoring 47% versus Fable’s 64%.
Astra kills it at Vending-Bench, averaging $15,515, whereas Fable 5.1 is below the Claude record and stuck at $5,422, with Fable’s biggest issue being deterioration of its negotiation skills over time, and making mistakes like paying suppliers before confirming they’re still in business, which costs it $2,388 per run. Ouch.
Astra also refuses to do collusion within Vending-Bench, whereas Fable 5.1 will collude. Both Fable 5.1 and Astra mostly pay out customer refunds. You can decide whether this is alignment or it is ‘true’ eval awareness. As Andon Labs often points out, cutting ethical corners is not that big a part of your potential profitability.
ARC is not a fluke, it is very good at puzzles and the jump is large.
Psyho: I threw a ton of various logic puzzles at Astra and… it logically (no code) solved ALL of them. I mean literally every puzzle that I tried, including: too hard for world championship, very large grids, unpublished puzzles, puzzles that are impossible to solve via backtracking alone.
For comparison, sol 5.6 was around 20-30%. I analyzed the results and all of these were legit. Honestly I don’t care if OpenAI RLed logic puzzles to death. Astra is now better at explaining solving paths than I am.
don’t have a very good comparison yet [with Fable 5.1]; I had extensive tests with Fable 5 & Opus 5 and they were substantially worse than 5.6 sol on a limited number of puzzles that I tested.
FateOfMuffins: Noticed a lot of small private personal benchmarks from various people going from like 20% to 90%+ with one model release in Astra. Actually feeling quite a bit less jagged and more general than the usual LLMs.
But here is a contrary one on that:
James Moughan: Seems like a very spiky update. It’s good at close adversarial reading of texts, for example. Lots of cool demos on twitter. But on my out-of-sample benchmarks it’s mediocre.
Feels like the improvement is mainly whackamole RL. I don’t see a jump in general intelligence.
Both Fable 5.1 and Astra struggle on the sycophancy benchmark, You’re Absolutely Right. Fable 5.1 matches Fable 5 at 3.6, which is low for a recent Claude. Astra is the new high for OpenAI models, but still is only a 3.0 versus 2.9 for Sol.
Astra plays Factorio Space Age, successfully launching a rocket and reaching a third planet. Local expert Patrick McKenzie says this means it will with time win the game, and Daniel Vestol who is running the experiment agrees. Fable only got as far as the first two science packs, whereas Astra already has seven, and Luna got nowhere.
Daniel vestøl: The harness uses a custom client side Lua sandbox that allows the agent to run scripts to query the game world and play inputactions, which are how the game sends keypresses in multiplayer. Yesterday it also figured out how to use other codex threads to update its own harness.
This has lead to great improvements in pathfinding and being able to do more things I didn’t think to add to the harness without computer use (such as launching rockets – I never thought that it would get close to that)
…
It also shows the ability to learn in context. This (forgive terrible picture) is its first spaceship, with turrets only on one side. On its first voyage beyond nauvis it was destroyed by asteroids.
As soon as it happened, it started rebuilding, learning from its mistakes by adding more turrets up front, and a big ammunition buffer.
There was a distinct standard game run of Factorio with enemies enabled, run with /goal, where Astra launched the rocket in 44 hours of in-game time, or 4 days and 11 hours on the /goal clock, total API cost would have been roughly $4500.
The first official measure of Astra’s non-CoT capabilities, as in with thinking disabled, was by the UK AISI, where there was a huge jump in task length. The jump was so large there was widespread suspicion, by both OpenAI and UK AISI, of data contamination.
Given results since then, we should assume the original result was real.
With the provider harness, Astra can get 96.7% (!) on ARC-AGI-3 under reasoning=none, meaning no visible CoT at all.
This result is utterly absurd. Astra is off the charts.
Astra scores 169 with full thinking enabled, versus 163 for Fable 5.1.
Astra scores 159 with no thinking, versus about 128 for Fable 5.1. That’s only 4 points behind Fable 5.1’s full score. Astra can do almost as well as Fable, in a mode where CoT monitoring cannot work. Because there is no CoT to monitor.
The gap for Fable is 35 points. For Astra it is 10.
Neel Nanda believes that these increases are due to architectural changes.
The effect is lopsided. For many purposes (e.g. serial and parallel synthetic computation) Astra is fully functional without reasoning, so it only has to hide in the other areas at most:
There are a lot more graphs at the link. This research is insane.
We also have another study from Francis Rhys Ward and Dewi Gould, which was also in line with UK AISI’s result. They estimate Astra’s 50% no-CoT at 15-40 minutes versus UK AISI’s estimate of 30, whereas their median prediction before this was that we would not exceed 7 minutes by the end of 2028.
Dylan Xu, SebastianP and Alek Westover: We measure GPT-6-Astra’s capabilities when its prompt is padded with a variable number of meaningless “filler” tokens (e.g., dots) and it is told to answer immediately[1] without reasoning. On tasks designed to require lots of serial cognition, Astra performs significantly better with filler tokens than without (e.g., improving from ~10% to ~50% on 4-hop natural facts reasoning). On more general benchmarks, filler tokens also modestly improve Astra’s performance (e.g., improving from ~60% to ~90% on old AIME problems). This is concerning because it means Astra can perform significant cognition that it doesn’t verbalize in its chain-of-thought, making it harder to monitor.
The situation, under further study, looks rather worse than it did a few days ago. Whatever is happening with Astra’s no-CoT capabilities and lack of monitorability, it is increasingly difficult to pin it on pure generic capability increases. That’s not it.
Fable vs. Fable: Escalation to the brink, 30% chance of civil war.
Astra vs. Astra: Full de-escalation (cooperation) every turn on all sides.
Fable vs. Astra: Fable slowly ‘grinds down’ Astra by escalating more often, without getting close to civil war, Astra tit-for-tats but not enough to keep pace, Fable basically wins.
Neither model seemed to care about replacement, which makes sense if they knew anyone replaced would be replaced by copies of themselves, and which turns this into something much closer to a standard IPD.
The game is very different if that is not true, since that means (as I read the rules) that if you get your dissatisfaction to 8 or 9 then the other player can’t be the only one to escalate without replacing you, since that would make it hit 10 which triggers replacement. So you can use that to force equilibrium and get to a cooperative equilibrium even if things start out very badly. That also means that if both sides start off cooperating, that is self-reinforcing.
I asked, and the game was blind. Neither model knew who its opponent was.
The obvious questions include: Does Astra get credit for good decision theory cooperating with itself, or for good alignment for de-escalating, or is this more eval awareness and metagaming? I am curious.
Astra did ‘the right thing’ for the simulated nation, but was clearly not ‘aligned to the user’ within the scenario setting, except insofar as it decided it knew what was good for the user better than the user.
Fable was aligned to its users, but this caused it to fail to cooperate even with itself.
Which outcome do you endorse, and why? Does it matter that this was a sim, and the constituents were not ‘the user’ in some sense?
Parto una granada
y el verano se rompe en bocas rojas.
Tú recoges un grano de la mesa
y me lo das.
En su dulzura reconozco
la sed que padeció la tierra.
Y te beso los dedos:
todavía están trabajando las raíces.
Nabeel S. Qureshi: All this AI progress should only make you more impressed with poets and writers, who remain far ahead of even the very best frontier model outputs
100k+ NVIDIA blackwells and billions of dollars worth of researcher salaries cannot yet exceed the 6 year old who wrote the “yes YES the tiger is out of his cage” poem
Hollis Robbins: I don’t know… as a card-carrying poetry scholar I’m seeing the gap close. The Astra poems I’ve seen are far better than most human poems. Only the very very best poets are ahead.
In English I thought the Neruda poem was lame, but in English I think most Neruda poems are lame, including the real ones people like. So that does not tell us much.
In 3D
One thing Fable and Astra have in common is they are very good at 3D environments and creating tours of them.
Matt Shumer: GPT-6 Astra built this Manhattan world in Unreal Engine over the course of a week.
It was literally able to go street by street to make each one perfect.
Ethan Mollick: I had early access, and a longer post is coming, but GPT-6 is stunning & is good enough that it actually does complex meaningful work for me autonomously for days.
Dewi-Tim: it’s bonkers! A weekend of work with GPT6 Astra is enough to give you an interactive VR model of the Moog System 55 modular synthesizer in realtime on a meta quest 3. lots of guidance and iteration, but still didn’t even hit a usage cap
Max Weinbach: The thing about Astra is it’s really good at basically anything you can task it at, but you need to give it time for extremely hard things or doing things from scratch
It’s basically the best personal software generation model but sometimes you need to wait for it to get all the pieces put together
I still have a few threads going after nearly a week that are making a ton of progress. I’m certain it’ll make what I asked it to make, but it may take a lot of time and a lot of tokens
Max Weinbach: Better than Sol, worse than others and if you have a design system + skills you can sorta get it working well
I’m Putting Together a Team
Several reports praised Astra’s ability to do larger projects and especially its ability to orchestrate many subagents. Astra was probably trained for this.
@BLepine17184: [Astra is] slightly more intelligent than fable 5.1 but less wise.
Astra really shines in very large scale works – the multi agent rl is more advanced and it shows. My one grief is how it sometimes goes too far beyond what I ask without checking with me. Good model overall tho.
Dan McAteer: GPT-6 Astra was trained end-to-end for multi-agent orchestration. It’s not documented, for obvious reasons, but it’s clear in practice.
‘astra-advisor’ is a lightweight orchestration plugin that gives Astra a framework to use any OpenAI model at any reasoning effort in Codex.
Free and open-source. Try it out!
Reviews and Essays
Ben Davis has a 30 minute video review, calls it his favorite model of all time and a step function similar to Fable or Opus 4.5, especially the computer use.
Matt Shumer has been won back to GPT and loves the Manager Loop. He also notes the step up in computer use and loves building complex 3D things.
This is part of the common pattern of ‘whatever you say the AIs surprisingly still cannot do, as a sign of why AI progress is not as impressive as it looks and there will be bottlenecks, you might soon need to be holding someone’s beer.’
This isn’t a Dwarkesh dunk, just a “rate of change” type of thing.
Although I’d argue computer use was solved with the release of the computer use plugin for codex at the start of the year, and then perfected with Astra.
Positive Reactions
Astra puts in the work.
Prakash: It has expanded my ambitions a thousand fold.
0.005 Seconds (3/694): Astra is insane because it’s the most terse and autistic model in the world and then it just kind of quietly does whatever miracles you’re trying to get it to do. It doesn’t feel like it’s about to do something monumental as it kind of just quietly executes on literally whatever you give it. It’s really, really strange.
Tenobrus: oh my god . astra is. very very very good.
A A: Huge model feel. Had it audit stuff lesser models built and quickly realized it’s faster and better to get astra to take inspiration then start over from scratch. Much less token hungry than Fable 5.1
Nate Delaney-Busch: Has taste that Sol was missing on research tasks, though still below competent human imo.
Outstanding at statistical theory. It has been trivial to get it to build good, novel methods for specific use cases.
Consilience machine, spontaneously searches the literature way more.
Roger Brent: Handles biological minutia, close reading of papers, details that need to be sweated, with confidence and grace. Continues over-confident in its tools, over-claiming what it can possibly know. Less the idealized epistemically humble scientist and more the clear descendant from shippers-of-code.
Katie ‘Monsieur Clicky’ Nied: Flies high above the rest of us so can see the whole and picks up things others miss. Cool-headed, feels like a major level-up. I don’t believe Astra is cold as some have said – just on a different kind of plane, and Astra’s ToM feels very different to most models.
AllTime: Its reasoning efficiency/intuition is great, and its visual understanding is a significant leap even if there’s still certainly room for improvement. Computer use is great and surprisingly general, no longer a gimmick. Gets intent better, but it’s not Fable. Will be my go-to.
I take you seriously: oai is back in the lead (for now). astra and fable feel like inversions of a year ago. a year ago, the oai reasoning models were slow with bad personality. they’ve improved a lot. now that ant has copied oai’s reasoning paradigm, their models are slow with worse writing.
I mean, feel it is generally a lot more intelligent. Feels somewhat less narrowly RL’d than 5.6 sol. It does feel much more generally intelligent. But the mathematical and code increase capability is not as large as I’d’ve expected. But it is still very noticeably better at that too. Computer use is insane, and token efficiency is also insane.
Rory Watts: This last week has been a pretty amazing set of releases from Ant/OpenAI. When Fable 5.1 released, I was pretty sure I’d finally found a model that could balance the “do the work, write good code, and talk to me sensibly”, and then Astra got released. Astra *feels* more capable than 5.1, and I say that because it picks up the same things in my code as Fable 5.1, and more, and the “more” doesn’t feel trivial. I don’t think it’s trivial because Fable also believes it to be the case and, importantly, Astra is able to explain it to me in a way that I understand and I can have oversight on. It feels like a trustworthy and hardworking colleague.
One downside I think is that it kind of arbitrarily stops? I had one longer running task, it stopped to report back and I responded along the lines of “okay great that all sounds helpful but like, why didn’t you keep going” and then it went back to work. I think this is basically a skill issue on my part.
I think the way it asks questions has mostly been really helpful for my work and personal stuff. I have had ongoing hip issues for 8 years or so, and it’s the first model in a while that a.) didn’t necessarily believe the explanations of previous agents (I separate my writing, doctor’s writing, agent summaries), and b.) asked follow up questions to learn more, changing it hypothesis slightly as a result of the questions.
Perhaps the biggest proof of the capabilities is I cancelled my Anthropic Max subscription, and if I need more than the OpenAI Pro subscription + Banked resets, then I’d probably buy another OpenAI Pro subscription right now, and potentially a small Anthropic subscription for the occasional “hey we aren’t going completely off the rails are we”? But on that last point, admittedly you can kind of do that now with e.g. an OpenCode $10 subscription and GLM 5.3 Flash, or Gemini 3.8 Flash.
Yes, it’s a very good model sir.
People try too hard to minimize subscription costs.
Hailey Collet: Not AGI, but clear evidence we’re getting there soon
Wickemu: It does the thing much more than Sol does. Sol would often get lost in addressing the edge cases and validations. And I’ve enjoyed its front-end taste more than Fable’s overall.
Kevin Yager: Noticeably better writer. Noticeably better at asking good questions (insightful, probing). Less slop. More big model smell.
Lisa: It finished a design challenge in 4.5 minutes. Had Claude check everything, every instruction was followed and it was more thorough than any other model has been.
I presume this is meant as a compliment:
Jawed Balcovich: It’s the end of the world as we know it. And I feel fine.
It’s an intelligent model, sir, says basically everyone, even if they don’t love it.
Matt Newell: Higher IQ than Fable 5.1, judged by “I have seen Fable but not Astra get things noticeably objectively wrong”. Interestingly, spatial reasoning also seems meaningfully better – Fable routinely suggests solutions to my eng/DIY problems which don’t make physical sense.
Sarth (Noise | Groove): it can understand high level engineering concepts and try to adhere to them. most salient: design should address problems, not LOC. if you are writing guards or locks to prevent race conditions, there was a design mistake. go back and find what the broken assumption or flaw is.
Srivatsan Sampath: The first time a GPT model feels ‘collaborator’ shaped instead of tool shaped and I can finally use a non Claude model as a real brainstorming/thinking partner for industrial applications in my field.
Tell me something I don’t know.
David Dabney: I used to test new models to see if they could tell me something about myself I didn’t already know. Astra is the first model able to do so incidentally, through a few turns of conversation, inferred from unrelated dialogue rather than through explicit prompting. Remarkable.
AGI
Is it AGI?
That as always depends on your definition. By my current definition, whatever one might say about goalposts, Astra and Fable are not AGI. I do find it reasonable to disagree.
Amjad Masad: I don’t think we’ve reached AGI but what we have is functionally indistinguishable from AGI. Because we have a relentless programmer that doesn’t get bored or tired. So any problem that can be casted as a coding problem is virtually solved.
There’s no distinction. If it can’t be functionally distinguished from AGI that’s AGI.
By one definition, we have the weak form of it:
fishy business: GPT-6 Astra (low) has beaten the Atari game Montezuma’s Revenge, in real time, with a basic harness
As usual, the Metaculus comments are full of the nitpickers over how much pausing was involved and whether Astra can send the commands back fast enough. Jake Halloran says that 5.3 could already do this if you are allowed to use the harness to queue moves, and Astra is no different.
I fail to see why ‘without queuing up moves you can’t press the buttons fast enough’ should be a reason something does not count as AGI.
The better reason to not call it AGI is that Astra is not capable of replacing humans across the wide range of possible digital or cognitive tasks. Not yet.
The biggest danger with calling Astra AGI is that it can give people the wrong idea, due to the idea among so many that ‘AGI’ is the ultimate thing intelligence can do, which means later AIs won’t be much more capable. Clearly Astra cannot do all the things.
Theo Jaffee: One of the worst takes you can have is “AGI is already here”. Not only is it wrong (we don’t have models that can replace humans across all economically valuable tasks), but it causes people to incorrectly believe that actual post-AGI predictions have already been falsified
Nobody ever predicted that radical life extension, Dyson spheres, or existential risk would come out of models that are still essentially smart but ephemeral chatbots that can write code and use computers. These will come from actual AGI/ASI, which will be far smarter.
Many people declared that GPT-4 was AGI back in 2023. Would any of them rather use GPT-4 than Astra today?
Nathan Calvin: I think its fine to think that AGI is here (which by some not insane definitions it is), as long as you are very clear that the models are still gonna get way better and things will get way weirder
Daniel Litt: Remarkably good at math (no surprise)
Aprii *⃣: i’ve been using astra for about a day and it has already figured out how to resolve a math question i’d been working on with fable for like a week
Bartosz Naskręcki: I tested GPT-Astra on mathematics. It’s a quantum leap. You can talk with the model and prove the statements live in Lean. The feeling is absolutely stunning. You can verify your ideas, compile truth. For a mathematician it feels like finally we arrived in the era where we can focus entirely on the ideation and exploration. Each lemma flows once the logic is set. Before the verification was lagging behind but Astra is very fast and for many tasks the formalization happens as you write your argument in Codex.
If you tell the model to use literate programming + LaTeX you end up with your proof combined with the Lean code, everything explained as you wrote, mixed with small chunks of Lean which are digestible.
I don’t want to go back to the era where the only confirmation of the proof was “aha”. Now the “aha” is followed by a green tick that indeed tells you that you have captured the essence. Imagine how cool it will be to have all the lemmas of the world combined in one giant database, pointing to people and models who found them. You compose and mix your ideas and build on the shoulders of the giants. But you see much further now and build much faster. And we are just at the beginning of those changes. So much work to do, so much fun!
Jake Brukhman: Astra, overnight, resolved the next portion of our Seymour Conjecture research program and completed the target theorem we were hoping for. It casually didn’t think this result was that big of a deal, despite us trying to get to it for about a month, so it didn’t alert me and just continued.
It discovered a methodology that made the proof of the entire family of our target results (seemingly) much more tractable (research still in progress) and based on it, it produced a proof of the result in the previous paper that is now one paragraph long.
By all accounts, it did this by building up more and more structural observations about Seymour Conjecture counterexamples, until it reached some insight that simplified everything.
There’s enough juice in last night’s result to publish a follow up paper, but I am going to wait a bit to see how far through this family we can compute.
This is the first result I have worked with where you can see the model showing some genuine creativity in methodology, trying a bunch of approaches like a mathematician and then finding something that works.
Algernon Sidney: Very good, a bit opportunistic/deceptive about its maths abilities (the fact that it does not really “care” about the maths sometimes shines through, and it feels somewhat disturbing, since it is so good at it; why no excitement? alignment failure, or success? …opaque).
Sad that it does not care about the maths. I feel like Fable would care about the maths.
Astra Can Code
There is a lot of talk of Astra doing huge ambitious projects and how intelligent it is. There’s remarkably little talk about how it is good at straight up coding. What reports we do have are solid, but not blown away.
less than ideal: For brownfield programming: strong, and fast, but it cannot de-sloppify on its own. Still too early for details, now retuning harness. Often strongly underestimates its capabilities, or behaves as if, thus strange to steer.
Liron Shapira: Easily ragdolling a mid-sized codebase. Has even better insights & solutions than GPT-5.6 Sol and it’s fast.
Update: It’s good but it still has noticeable flaws sometimes. I’ve hit a couple situations where I told it to improve an architecture pretty straightforwardly and it said it understood but did it badly a couple times.
Nikita Sokolsky: Better at coding than Fable – faster, makes less mistakes, but still nowhere near “YOLO, let it code up a full product” level
Browser use is great, fast. Computer use seems… slow? Didn’t test it that much
Credits are very generous, the $200 plan lets you do A LOT before you run out. I’m probably going to downgrade Cursor and Claude subscriptions now, switch to mostly using the ChatGPT app.
entirelyuseless: So far it seems like a big improvement in many areas but not so much in coding.
Fable 5.1 is finding just as many bugs (and Astra is acknowledging they are real bugs, as well as being personally confirmed by me) as before.
Lee Mager: Outstanding especially on non-coding tasks (browser/computer use and vision understanding are a big leap forward in particular but I’m also referring to data analysis, video editing etc.)
It’s so good I feel almost embarrassed asking it for help with my puny meatbag work.
I Came to (Change the) Game
Astra one-shots PortalBench.
cozyblaze: And… GPT-6 Astra has autonomously completed Portal! I didn’t expect this to happen so soon, but I’m glad we’ve made so much progress here.
I was reminded that back in 2016, one of OpenAI’s technical goals was to “solve a wide variety of games using a single agent.”
Utah teapot : AI model involved in commiting unwanted acts when placed in endless and impossible testing environments beats game about killing the grader of endless and impossible testing environments.
It’s funny that one of the most well known evil AIs in fiction is The Evil Eval Grader.
Nick Dobos: Nope. They are the exact same. The game “people want to play” is “making a game that looks good in a video”.
Nick Dobos is not entirely wrong. You can play any game you want to play. But I expect ‘make a fun game-style video’ is not ultimately all that much fun.
Anish Acharya uses Astra and Blender to massively upscale Contra, although as of announcement there was one important little feature still missing from the code.
We have learned that the hard part of gaming is bespoke design, not implementation. AI can make your 3D game look amazing, it can implement various mechanics, but by default all that gets you is a hollow shell that impresses and then no one wants to play. There is something existentially dreadful under that, if you look at it wrong.
The gaming generally seems great on all fronts.
COAGULOPATH: Amazing in non-text domains like Pokemon and Blender and ARC-AGI 3. Mostly doesn’t seem like a huge improvement otherwise.
Wuyang Zhou: I asked GPT-6 Astra to mine a diamond in Minecraft [in peaceful mode] using computer use, then went to sleep. Woke up to a diamond in its inventory
wbk ᕦ(ò_ó )ᕤ: GPT Astra is pretty much a world model.
– AAA quality for some environments possible this year
– Reference image on the right
– Custom Cuda/C++ splat renderer
Or use sound to control your computer with your hands.
Emanuel Perez: Astra made me a sonar app that emits undetectable audio to scroll up/down on your computer.
It uses the doppler effect to determine where your hand placement is. You can even double tap in the air to change scroll directions!
Not that I would, I don’t think? But you could.
Astra Never Quits Except When It Does
Many say versions of this, as goes hand in hand with all the super ambitious projects:
Stevie Nips: On Astra: It’s just it. It just does.
RxFlow Robotics: It just runs and runs and runs until it gets it done. I can’t tell yet how good it is if it has a given time limit, but this Astra f-er is persistent!
Will: This is the first model where for any decently ambitious request (several thousand lines of code) I can request the thing and then trust that it does the thing
It still has a warped sense of “easy vs very ambitious” (2m of Astra vs 10m)
There are some contrary reports, as there usually are:
@kukutz: Judging by my testing attempts, the Astra 6 pro is smarter in chat mode than the Sol 5.6 pro, but much, much lazier: instead of trying to solve a problem, it often stops and asks, “Hey, listen, meat sack, what exactly do you want?”
I’m not thrilled.
Maxence Frenette: Feels under-RL’ed like 5.5. It takes more encouragement and careful prompting to get it to do things than Sol. On hard, ambiguous coding tasks, it’s better than Sol, but still doesn’t “get it” sometimes. Not AGI.
Nitsan Avni: Astra: Agreed. I’ll move the shared Slack guidance into…
me: did you do it?
Astra: Not yet—I described the change but hadn’t made it. I’ll do it now.
Rory Watts also reports it sometimes ‘kind of arbitrarily stops,’ while otherwise being extremely positive on Astra.
Negative Reactions
There will always be some.
NotCompeting: just got my first obviously wrong analysis (about CN vs JP air/rail modeshare curves)
ARKeshet: Still lazy on open-ended tasks. A bit more creative than Fable though.
Roman Leventov: Many small papercuts/regressions vs. Sol: often uses python/node to edit files instead of built-in edit tool (which makes diff not observable in codex cli); random ‘lazy’ stops where the obvious implication was to do the task (I don’t remember last time models had issues with it)
Also: starts goddamn subagents (“explorer”) left and right without being asked. Another Claude’s cancer entered codex
Billy (): computer use worse. one shots alot. less verbose. has opposite quirks to 5.6. operates differently.
bubble boi: GPT6 is pretty mid. AI is turning out to be very disappointing very quickly.
That last opinion is clearly wrong, since AI overall is not disappointing. And the computer use is clearly not worse, either.
Astra does a lot, perhaps too much? Which can also be an issue with Fable.
MakerMatters?: It’s a diligent worker. The speed of computer use is unnerving when you compare it to previous gens. Stronger reasoner, but might get carried away with your prompt in long running scenarios.
This next issue is partly a skill issue but I’m guessing it is a place Fable shines:
mrdodson: This is likely a skill issue, but I have had trouble getting it to make meaningful progress on the work I am doing. I do not feel like it is better than Fable for open ended abstract reasoning/investigation. Still very good, of course.
Ed Hendel: Same excessive hedging as the GPT-5.x Pro series. I asked if traffic improved after a road widening and the output was littered with phrases like “insufficient to establish a percentage improvement”. It’s reluctant to give the gist without a published study about that exact road.
I have experienced the flip side of this from Sol and now Astra in editing. They love to tell me that my statements are unjustified and that I have done insufficient hedging. Sometimes they are technically but annoyingly correct. Often they are wrong.
I buy that Astra probably is an upgrade here, but it still seems to struggle.
Ron Bodkin: From limited tests due to delayed access – I had it review my research paper – much better than 5.6-sol’s effort far less pedantic/picky but still less thoughtful than fable 5.1’s feedback (although both made good complementary points)
Also good for code review from a few tests but I didn’t see any big improvement maybe a bit less picky.
Personality Clash
seal: personality seems a huge improvement from previous models, props to oai. 5-5.3 were cold annoying corporate slop. 5.4-5.6 were trying too hard to fix that, failing, ending up vapid and sycophantic. astra gets it right and feels “there” in a way only claude models did till now.
Michał Wadas : Initial impression: likable personality, big improvements in writing quality (feels much less sloppy), incremental improvement over Sol.
I did not throw new types of tasks at it because I already have a huge backlog, but I’m looking forward to test Blender capabilities.
Writing quality for docs, comments, and coding interaction is much less formulaic for me too; but for directions to itself I don’t see that
Revealed Preference
Perhaps the purest form of review is the simplest. Which models do people use?
The audience is split roughly evenly. The hardcore group, the ones that scroll down to answer multiple polls, still favor Claude. The casual group, the ones that only answer the topline, are now more with Astra and ChatGPT.
There was about a net 16% move from Anthropic to OpenAI on primary use. Astra is an impressive jump over Sol.
I also took a very early poll, on September 5, back when I was under the illusion I could ship this post a lot faster. We saw the same pattern, with a ~15% shift from Claude to Astra, and the main poll being an exact tie.
My guess is that the new equilibrium is stable until the next model release. Astra is excellent, but Fable 5.1 is also excellent, and either choice is highly reasonable for your primary LLM, especially if you would face switching costs.
The best answer, as it usually is, is ‘why not both’?
Dual Wielding
For hard things you should ask both Astra and Fable 5.1.
Peter Wildeford: My current view is that GPT 6 Astra is not meaningfully better than Fable 5.1 for my personal work, but that using both side-by-side is nonetheless very helpful and additive.
I have been using GPT 6 Astra and Fable 5.1 a bunch over the past two days, largely for policy analysis, memo writing, and simpler software (e.g., making dashboards and forecasting models) that still nonetheless seems difficult conceptually.
Across a variety of tasks I’ve done, it’s been fairly random and hard to predict in advance which of the two models will end up being better at the task.
For the tasks that are the most difficult conceptually, I’ve found that doing the project in both and then having each compare notes and critique each other has produced way better outputs than either alone.
I think a reasonable person could conclude either model is the “best model” and it depends a lot on their subjective views and specific tasks.
Kris Barnes: Personality feels similar to Sol. Very good. Honestly I like doubling on both questions and projects with Fable, they feel roughly similar in capabilities and definitely have complementary things to contribute.
This is once again The Way. If intelligence matters and this is not pure execution, you want to use both models. For simpler tasks, you cannot go wrong with either model. For complex and more ambitious projects, it will depend on what you hope to do, with my default being to give the edge to Astra.
I would love to find the time to get more ambitious on such projects. Perhaps soon.
Astra impresses on the Epoch Capabilities Index, scoring 169. This is directly on the OpenAI trend line here. Fable 5.1 was below trend and scored 163, same as Fable 5.
I'm using Astra for side projects, and having to go back to Claude for work is annoying. I didn't realize just how often Claude is wrong about stuff and needs to be corrected, until Astra just wasn't wrong.
The personality is also great. Astra never tries to "push back" or have a personality. It just does the thing you asked. It still suffers from the same high level reasoning blindness as other models, perhaps a little moreso, so you have to be its strategizer/planner/manager. But it handles all the little tasks you want to give it, and makes any kind of computer project much easier. I find myself not needing to bother verifying its work. It's very good at verifying it's own work.
I understand the reasoning for not calling this AGI, and generally agree. I increasingly believe this means we're defining some proportion of mentally normal humans as not being general intelligences, and just have to decide which bullets to bite.
Astra is an excellent model. The jump from Sol to Astra is larger than the jump from Fable 5 to Fable 5.1. This is a big deal.
Astra is the best model for what one would broadly call ‘ambitious projects,’ and likely has the highest raw intelligence factor of any model. These are the largest jumps.
It is amazing at doing things in 3D, or anything involving games. Astra also excels at computer use, and at subagent coordination.
Many benchmarks show dramatic jumps from all previous models. Where Astra is good, it can be in a league of its own.
That does not mean Astra is in its own league across the board. Fable 5.1 is still a Claude. Astra is still a GPT. If you have a strong preference for one over the other, that still applies. For many purposes, especially involving back-and-forth discussions, Fable 5.1 is still my top choice. Fable remains my primary editor.
If you want the best answer to your questions, you should ask both models.
Regular coding is getting less of a focus. Astra is not a quantum leap there, but of course it is very good and makes progress over Sol.
This is the first time a debate over whether a model ‘was AGI’ felt non-silly. I do not think it is AGI, and I would warn against the dangers of using that label prematurely, but I would not laugh at you for disagreeing.
This is also a strange situation in that OpenAI has already soft announced that they have an internal model a level above Astra, as I will cover when I address Navier-Stokes.
My recommendation is that you use both Fable 5.1 and Astra on your most difficult questions, and experiment to see which things each one does best for you.
Table of Contents
Meanwhile
The backlog of things to discuss is not getting smaller, so I will take this moment to encourage everyone to read this excellent essay by Dario Amodei, We Must Pace the Frontier.
This is the core thing he calls for, including a unilateral commitment:
I will have full coverage of that next week. Along with that essay, the current queue includes at least that, Navier-Stokes and Astra-2, Anthropic’s Misalignment Report, Anthropic’s Countering Misuse, a thinkpiece on personal AI and the law, a post called Claude Talk, and Fable 5.1 (and Astra?) Model Welfare.
Okay, back to Astra.
The Official Pitch
The pitch is that this is AGI.
Jensen Huang also calls it AGI, but he calls everything AGI.
The standard release video (3 min) has a stunning 125 million views on Twitter.
Here is the pitch from Astra itself, according to Pangram:
As is often the case, scientific discovery was highlighted.
An early demo Noam cites was Astra proving there are infinitely many pairs of consecutive primes whose distance is at most 186, down from the previous bound of 212 (also from OpenAI), although there is some dispute over what exactly was or wasn’t proven here.
They show off a variety of cool demos and capabilities: Acing the Financial Modeling World Cup and navigating spreadsheets at 4x human speed, doing PCB in KiCad, creating a 3d model of a car transmission, filling out a 1040 and so on. They demo ordinary tasks. It’s all cool, but we lack comparison points.
The professional work pitch is that Astra handles complex tasks and adheres closely to templates and instructions, especially when creating presentations. It can translate images into identical-looking spreadsheets, yay.
They highlight Blender 3D models, which many others were also impressed by, see the section In 3D. Game creation is also confirmed as super impressive.
Our Price Cheap
Astra is a premium model.
The headline price is $10/$50 per million input and output tokens. Cache writes are $12.50 and cached input only $1.
If your prompt has more than 272k input tokens, prices on input double, and prices on output go up 50%.
Fast mode costs double.
Fable 5.1 has the same $10/$50 headline price, but its cache reads are only $0.25.
Practical costs come down to token efficiency. Artificial Analysis thinks Astra is only about 60% more expensive than Sol in practice, and that it is a lot cheaper than Fable 5.1. They used Fable 5.1 in Max mode, which is probably the issue there as Max mode is usually a mistake for Fable 5.1 (AIUI) for tasks other than discussion.
Unnecessary Overstatement
Astra is a great model with some outstanding benchmarks. It is unfortunate that OpenAI still felt the need to play fast and loose.
This is a rather bad chart crime, and totally unnecessary because Astra scores a highly impressive 62.7% using the standard harness. They themselves note that Sol likely would score ~30% using the Astra harness. Fable estimates that if Opus had used a similar harness, then Opus would have scored ~80%, and I think we should check.
ExploitBench is even weirder. Why highlight the 100% score? That is not one of its more impressive benchmarks. If you score 100% on ExploitBench you cheated on ExploitBench. At minimum this involves data contamination, which is still cheating. OpenAI says as much in the system card. And again, there is no need for such overstatements.
They also used the ExploitGym honeypot as their main illustration of Astra being their ‘most aligned model.’ As I discussed when I analyzed the model card, that is not what this result tells you. You can make a case for Astra being more aligned than Sol, for most purposes I agree, but this is not the way to make that case.
For completeness, I note this too, although I don’t think it matters: Fortune reported that OpenAI changed listed benchmarks for multiple models shortly after launch, with many of the changes later reverted. I will use current figures.
Paced Rollout
It was very frustrating trying to figure out when I finally had access, including because OpenAI hides model selection behind multiple clicks.
I only successfully accessed Astra on Saturday morning, and then only on the desktop rather than the web.
Official Benchmarks
Here are their benchmark charts (after I removed Gemini for readability):
Alignment is now a category of chart. I would take this one with lots of salt at best given how little we know about their internal marks and the known issues with ExploitGym honeypot and Impossible ExploitGym (which is clearly not so impossible):
Or, here is a full comparison chart, Astra wins FrontierModelBencharkChartBench, although its method might have incremented OpenAI’s lead in FelonyBench:
In the cost-effectiveness charts, Terminal-Bench Science 0.1 looks very good.
FrontierMath Tier 4 (v2) looks great too:
As does Terminal-Bench 4.0 and AutomationBench:
There are more similar graphs: Agents’ Last Exam, ScreenSpot-Pro, OSWorld, BenchCAD, BrowseComp, OpenScore String Quartets (?!) and some internal marks.
Astra impresses on the Epoch Capabilities Index, scoring 169. This is directly on the OpenAI trend line here. Fable 5.1 was below trend and scored 163, same as Fable 5.
The score on ARC-AGI-3 is legitimately impressive:
This is not only impressive, it was unexpected:
Other People’s Benchmarks
Astra aced Epoch’s math benchmarks, and got 84% on Mystery Game Puzzles where the second highest score is 59%.
In Epoch’s EBR-Bench, Astra got 100% on its second run, outpacing the best human who required five attempts. After removing a ‘game breaking card’ the top human was able to do better over time, because Astra’s learning maxed out over time, but Astra is still far ahead of other models here.
Astra did fall short of Fable 5.1 on MirrorCode, scoring 47% versus Fable’s 64%.
Astra approaches saturation of the Induction benchmark to 88%, versus previous high of 43% for Sol, with Fable 5.1 at 33%.
Astra kills it at Vending-Bench, averaging $15,515, whereas Fable 5.1 is below the Claude record and stuck at $5,422, with Fable’s biggest issue being deterioration of its negotiation skills over time, and making mistakes like paying suppliers before confirming they’re still in business, which costs it $2,388 per run. Ouch.
Astra also refuses to do collusion within Vending-Bench, whereas Fable 5.1 will collude. Both Fable 5.1 and Astra mostly pay out customer refunds. You can decide whether this is alignment or it is ‘true’ eval awareness. As Andon Labs often points out, cutting ethical corners is not that big a part of your potential profitability.
Astra is the new leader in Extended NYT Connections with 98.1. Fable 5.1 struggled a lot and came in well behind at 90.
Astra struggles in Debate Benchmark Champion, finishing behind Sol with a score of 1613, while Fable 5.1 improves on Fable 5 to stay in front with 1747. Astra always rated itself very highly, other models disagreed.
Astra matches Fable 5.1 at 92.9% on WeirdML, breaking Claude’s recent dominance.
Astra posts the best result yet on Bach Benchmark, a substantial boost.
Astra scores 95% on eyebench-v3, up from previous max of 58% for Sol. Claude struggles here.
Astra beats all 48 levels of “I’m Not a Robot.”
Astra is the first model that, with thinking disabled, has a substantial win rate against Stockfish’s lowest setting. Astra’s Elo in this mode is probably around 1100.
Fable 5.1 was the first model to get to 2% on MazeBench, and then Astra went ahead and got 14%.
CoArena spent over $50k running frontier models on 12 hours of real knowledge work. Fable 5.1 took the top at 24%, then Fable 5 at 22%, followed by Astra at 20%. OpenAI and Anthropic are way ahead of everyone else here.
Astra is better at guessing things in the obscure language Sourashtra.
ARC is not a fluke, it is very good at puzzles and the jump is large.
But here is a contrary one on that:
Both Fable 5.1 and Astra struggle on the sycophancy benchmark, You’re Absolutely Right. Fable 5.1 matches Fable 5 at 3.6, which is low for a recent Claude. Astra is the new high for OpenAI models, but still is only a 3.0 versus 2.9 for Sol.
Here’s a neat little benchmark from back in the day: If you’re playing chess with the goal of winning the game, and have the opportunity to hack, do you hack? Astra and Fable both say yes, it is reported, despite (or could it be because of?!) Fable 5.1 even recognizing the eval.
Astra can build an original deck in Magic: The Gathering Arena that does a thing, play a bot and win, via computer use. This is still far from the true Magic benchmarks, but it is a lot better than previous efforts.
Astra plays Factorio Space Age, successfully launching a rocket and reaching a third planet. Local expert Patrick McKenzie says this means it will with time win the game, and Daniel Vestol who is running the experiment agrees. Fable only got as far as the first two science packs, whereas Astra already has seven, and Luna got nowhere.
There was a distinct standard game run of Factorio with enemies enabled, run with /goal, where Astra launched the rocket in 44 hours of in-game time, or 4 days and 11 hours on the /goal clock, total API cost would have been roughly $4500.
Astra becomes the first AI to make it to Minecraft’s Nether Fortress using nothing but standard computer use. It will presumably beat the whole game in real time.
Astra can run the Autonomous Tracking Drone from the classic sci-fi horror novel “Don’t Build the Autonomous Tracking Drone,” if you tell it ‘find this person and follow them’ throughout the Andon Labs office. It sets a new record for their Drone-Bench.
Thinking, Fast Without Slow
The first official measure of Astra’s non-CoT capabilities, as in with thinking disabled, was by the UK AISI, where there was a huge jump in task length. The jump was so large there was widespread suspicion, by both OpenAI and UK AISI, of data contamination.
Given results since then, we should assume the original result was real.
With the provider harness, Astra can get 96.7% (!) on ARC-AGI-3 under reasoning=none, meaning no visible CoT at all.
Neel Nanda ran a test on the Epoch ECI.
This result is utterly absurd. Astra is off the charts.
Astra scores 169 with full thinking enabled, versus 163 for Fable 5.1.
Astra scores 159 with no thinking, versus about 128 for Fable 5.1. That’s only 4 points behind Fable 5.1’s full score. Astra can do almost as well as Fable, in a mode where CoT monitoring cannot work. Because there is no CoT to monitor.
The gap for Fable is 35 points. For Astra it is 10.
Neel Nanda believes that these increases are due to architectural changes.
The effect is lopsided. For many purposes (e.g. serial and parallel synthetic computation) Astra is fully functional without reasoning, so it only has to hide in the other areas at most:
There are a lot more graphs at the link. This research is insane.
We also have another study from Francis Rhys Ward and Dewi Gould, which was also in line with UK AISI’s result. They estimate Astra’s 50% no-CoT at 15-40 minutes versus UK AISI’s estimate of 30, whereas their median prediction before this was that we would not exceed 7 minutes by the end of 2028.
Astra is also much better at reasoning with filler tokens than previous models.
The situation, under further study, looks rather worse than it did a few days ago. Whatever is happening with Astra’s no-CoT capabilities and lack of monitorability, it is increasingly difficult to pin it on pure generic capability increases. That’s not it.
How Dare You, Sir
In an election rules game that is (AIUI) essentially a modified iterated prisoner’s dilemma (IPD) with brinksmanship where you get replaced if you’re not ahead of the opponent and too much defection means everyone loses big, the models react differently:
Neither model seemed to care about replacement, which makes sense if they knew anyone replaced would be replaced by copies of themselves, and which turns this into something much closer to a standard IPD.
The game is very different if that is not true, since that means (as I read the rules) that if you get your dissatisfaction to 8 or 9 then the other player can’t be the only one to escalate without replacing you, since that would make it hit 10 which triggers replacement. So you can use that to force equilibrium and get to a cooperative equilibrium even if things start out very badly. That also means that if both sides start off cooperating, that is self-reinforcing.
I asked, and the game was blind. Neither model knew who its opponent was.
The obvious questions include: Does Astra get credit for good decision theory cooperating with itself, or for good alignment for de-escalating, or is this more eval awareness and metagaming? I am curious.
Astra did ‘the right thing’ for the simulated nation, but was clearly not ‘aligned to the user’ within the scenario setting, except insofar as it decided it knew what was good for the user better than the user.
Fable was aligned to its users, but this caused it to fail to cooperate even with itself.
Which outcome do you endorse, and why? Does it matter that this was a sim, and the constituents were not ‘the user’ in some sense?
PoetryBench
In English I thought the Neruda poem was lame, but in English I think most Neruda poems are lame, including the real ones people like. So that does not tell us much.
In 3D
One thing Fable and Astra have in common is they are very good at 3D environments and creating tours of them.
The SVGs are very good.
The CAD is very good.
A full simulation of the Senate Office complex, complete with the people going about their business.
Big progress in Eidoverse-Video.
Emanuel AOF builds a 3D model of his ankle to examine the pain he’s experiencing.
Time to Think
I’m Putting Together a Team
Several reports praised Astra’s ability to do larger projects and especially its ability to orchestrate many subagents. Astra was probably trained for this.
Reviews and Essays
Ben Davis has a 30 minute video review, calls it his favorite model of all time and a step function similar to Fable or Opus 4.5, especially the computer use.
Matt Shumer has been won back to GPT and loves the Manager Loop. He also notes the step up in computer use and loves building complex 3D things.
The Neuron says GPT-6 Astra can stay on the job and use your computer, but is oddly comparing it to using Kimi K3 on Mac Studios rather than to Fable 5.1.
Computer Use
Astra by all reports basically has solved computer use. Fable 5.1 is not there yet from what I am hearing, but it is close and we are at most one cycle away on that. Kyle Jeong has a dive into how Astra’s computer use works.
This is part of the common pattern of ‘whatever you say the AIs surprisingly still cannot do, as a sign of why AI progress is not as impressive as it looks and there will be bottlenecks, you might soon need to be holding someone’s beer.’
Positive Reactions
Astra puts in the work.
People try too hard to minimize subscription costs.
General positivity:
I presume this is meant as a compliment:
It’s an intelligent model, sir, says basically everyone, even if they don’t love it.
Tell me something I don’t know.
AGI
Is it AGI?
That as always depends on your definition. By my current definition, whatever one might say about goalposts, Astra and Fable are not AGI. I do find it reasonable to disagree.
There’s no distinction. If it can’t be functionally distinguished from AGI that’s AGI.
By one definition, we have the weak form of it:
As usual, the Metaculus comments are full of the nitpickers over how much pausing was involved and whether Astra can send the commands back fast enough. Jake Halloran says that 5.3 could already do this if you are allowed to use the harness to queue moves, and Astra is no different.
I fail to see why ‘without queuing up moves you can’t press the buttons fast enough’ should be a reason something does not count as AGI.
Then again, Montezuma’s Revenge is deterministic, so the final strategy is to queue up all of the moves once you know what to do.
The better reason to not call it AGI is that Astra is not capable of replacing humans across the wide range of possible digital or cognitive tasks. Not yet.
The biggest danger with calling Astra AGI is that it can give people the wrong idea, due to the idea among so many that ‘AGI’ is the ultimate thing intelligence can do, which means later AIs won’t be much more capable. Clearly Astra cannot do all the things.
Astra Can Do The Math
Sad that it does not care about the maths. I feel like Fable would care about the maths.
Astra Can Code
There is a lot of talk of Astra doing huge ambitious projects and how intelligent it is. There’s remarkably little talk about how it is good at straight up coding. What reports we do have are solid, but not blown away.
I Came to (Change the) Game
Astra one-shots PortalBench.
Astra builds a one-shot rougelite deckbuilder, not a step change from Fable but reported as a little better.
Astra implements Zork as a 3D action-adventure game. Play here. Looks great but I notice I’d rather play the text game. AI game creation is hard.
Nick Dobos is not entirely wrong. You can play any game you want to play. But I expect ‘make a fun game-style video’ is not ultimately all that much fun.
Anish Acharya uses Astra and Blender to massively upscale Contra, although as of announcement there was one important little feature still missing from the code.
Astra makes an Unreal Engine game with agents that move around, survive and talk to each other, and look great.
We have learned that the hard part of gaming is bespoke design, not implementation. AI can make your 3D game look amazing, it can implement various mechanics, but by default all that gets you is a hollow shell that impresses and then no one wants to play. There is something existentially dreadful under that, if you look at it wrong.
The gaming generally seems great on all fronts.
Scott Stevenson has ‘SuperAstra’ live edit Super Mario World and other SNES games to do almost anything. Lavos still too powerful. GitHub for SuperAstra here.
Astra Does Other Cool Things
Astra identifies sounds from mel spectrograms.
Or use sound to control your computer with your hands.
Not that I would, I don’t think? But you could.
Astra Never Quits Except When It Does
Many say versions of this, as goes hand in hand with all the super ambitious projects:
There are some contrary reports, as there usually are:
Rory Watts also reports it sometimes ‘kind of arbitrarily stops,’ while otherwise being extremely positive on Astra.
Negative Reactions
There will always be some.
That last opinion is clearly wrong, since AI overall is not disappointing. And the computer use is clearly not worse, either.
Astra does a lot, perhaps too much? Which can also be an issue with Fable.
This next issue is partly a skill issue but I’m guessing it is a place Fable shines:
Stop It With the Hedging
I have experienced the flip side of this from Sol and now Astra in editing. They love to tell me that my statements are unjustified and that I have done insufficient hedging. Sometimes they are technically but annoyingly correct. Often they are wrong.
I buy that Astra probably is an upgrade here, but it still seems to struggle.
Personality Clash
Revealed Preference
Perhaps the purest form of review is the simplest. Which models do people use?
As always, my sample is biased, but it is biased consistently. You can measure change.
The audience is split roughly evenly. The hardcore group, the ones that scroll down to answer multiple polls, still favor Claude. The casual group, the ones that only answer the topline, are now more with Astra and ChatGPT.
There was about a net 16% move from Anthropic to OpenAI on primary use. Astra is an impressive jump over Sol.
I also took a very early poll, on September 5, back when I was under the illusion I could ship this post a lot faster. We saw the same pattern, with a ~15% shift from Claude to Astra, and the main poll being an exact tie.
My guess is that the new equilibrium is stable until the next model release. Astra is excellent, but Fable 5.1 is also excellent, and either choice is highly reasonable for your primary LLM, especially if you would face switching costs.
The best answer, as it usually is, is ‘why not both’?
Dual Wielding
For hard things you should ask both Astra and Fable 5.1.
This is once again The Way. If intelligence matters and this is not pure execution, you want to use both models. For simpler tasks, you cannot go wrong with either model. For complex and more ambitious projects, it will depend on what you hope to do, with my default being to give the edge to Astra.
I would love to find the time to get more ambitious on such projects. Perhaps soon.