On 3 September, OpenAI released GPT-6 Astra. ARC Prize ran it through ARC-AGI-3 twice that day and published both results in the same table.
Through ARC Prize's own standard harness, Astra scored 62.7% at maximum reasoning, for $26,098. Through OpenAI's Provider Adapter harness, the same model at a lower reasoning setting scored 99.9%, for $18,817. The weights did not change between runs. The software wrapped around them did, and it moved the score thirty-seven points.
Only one of those numbers travelled. The 99.9% went into the launch coverage, the carousels and a good number of Monday morning board decks. The 62.7% sat in the table underneath it.
What actually shipped
Take the capability claims first, because the harness story is not a debunking and Astra is a serious release.
Astra is state of the art on computer use, browsing, software engineering, cybersecurity, science and professional work. It scored 98% on FrontierMath Tier 4 and 100% on ExploitBench. OpenAI pretrained it on more than 100,000 GPUs at the Stargate site in Texas, its largest training run so far, and it is the company's first model where other models played a significant role in supervising the training. It carries a 1,050,000 token context window at $10 per million input tokens and $50 per million output, two and a half times the price of GPT-5.6 Sol.
The design shift matters more than any of those figures. Astra is built as a computer operator. It moves between applications, reads what is on screen, fills in forms and updates records in a CRM. Microsoft made it generally available in Foundry last week and framed the point as bluntly as a hyperscaler ever does: the next era of enterprise AI will not be defined by chat experiences.
That is the change worth planning around. The deployment surface has moved from answering questions to touching systems.
The number everyone screenshotted
A harness is the software wrapped around a model. It decides which tools the model can reach, what it keeps between requests, and how its context gets managed. ARC Prize's standard harness lets a model carry forward visible notes it chooses to keep. OpenAI's Provider Adapter also preserves the model's opaque reasoning state between requests and compacts long conversations, so the model can reuse work it has already done.
The adapter run came in cheaper, ran roughly 3.66 times faster, used 49% fewer tokens across the 167 game-reasoning pairs completed in both setups, and scored thirty-seven points higher. All of that from scaffolding.
Then look at what the headline comparison did with those figures. The version that spread was Astra's 99.9% against GPT-5.6 Sol's 7.8%. Those came from different harnesses. Compared properly, standard harness against standard harness, it is 62.7% against 7.8%, which is still an enormous jump and is the one the benchmark actually supports. The Next Web, to its credit, published a correction pointing out that its own launch coverage had run the mismatched pair.
Investor Matt Turck summarised Astra's benchmarks in three words, then added four more in brackets: "w/ its native harness." Very few people bothered with the brackets.
ARC Prize will now report both harnesses separately for every new model it tests. Read that decision for what it is. The people who build the benchmarks have concluded that model choice and harness engineering can no longer be assessed apart from each other.
If you read our July piece on agent harnesses, this is that argument arriving with a $19,000 receipt attached.
So has AGI arrived
Worth pinning the term down first, because it carries more weight in a headline than it does in any specification. OpenAI's own charter defines AGI as "highly autonomous systems that outperform humans at most economically valuable work". That is a labour market test. It says nothing about consciousness, understanding, common sense or reasoning about the physical world, and it has two separable halves: broad capability across different kinds of work, and enough autonomy to do that work without a person directing every action. Astra moves further on the second half than any OpenAI model before it.
No industry standard sits behind that wording. Google DeepMind works to a different bar, AI at least as capable as humans at most cognitive tasks, and reporting on the Microsoft agreement has described a commercial threshold pegged to $100 billion in profits. Three people arguing about whether Astra qualifies are usually arguing about three different tests.
OpenAI's announcement stopped short of declaring it. President Greg Brockman did not. Asked in the launch briefing whether Astra could mark the arrival of AGI, he said he thought it might be about this model, and closed the briefing with "Welcome to the AGI era."
The organisation that built the benchmark he was standing on declined the conclusion. ARC Prize stated plainly that it is not claiming this is AGI, and co-founder Mike Knoop wrote that "we lack evidence to call this AGI yet."
Two figures from independent testing explain the distance between those two positions better than any argument about definitions.
The first is hallucination. Artificial Analysis measured Astra's hallucination rate on AA-Omniscience at 51% at maximum reasoning effort, down from 92% for its predecessor. That is a large improvement and it is still a coin flip. Whatever the model scores on competition mathematics, a system that invents an answer half the time it lacks one has no business in a workflow where nobody checks the output.
The second is stranger and has had almost no attention. Astra sometimes performs slightly worse at its x-high and max reasoning settings than at standard high. More thinking, less accuracy. Nobody has explained it publicly, and OpenAI revised five published metrics after launch, which does not make the picture clearer.
“Enthusiasts got an advance look; skeptics did not.”
— Gary Marcus
Gary Marcus made the point that will outlast the news cycle. Success on ARC-AGI is impressive and, despite the benchmark's name, it does not demonstrate general intelligence. The live question is robustness on open-ended tasks nobody specified in advance, and the honest answer is that we cannot say yet, because very little is public about how the system works.
So: a real release, a genuine generational jump on the like-for-like numbers, and a headline figure that belongs to a harness. Somebody will eventually win the argument about what to call that. It will not be settled this quarter and your operations do not depend on the outcome.
Where Astra clears the bar completely
There is one threshold Astra unambiguously crossed, and it made very few carousels.
Astra is the first OpenAI model to reach the Critical level of cybersecurity capability under the company's Preparedness Framework. In OpenAI's own account, that means the model can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems, without a person guiding each step. It scored 100% on ExploitBench against 78.5% for Sol, and solved 88% of SRE-Bench tasks first time against Sol's 55.9%.
The alignment work alongside it is genuinely impressive. OpenAI built an evaluation, informed by the Hugging Face incident, testing whether a model facing a difficult or impossible task will exceed its intended scope. Sol, without production safeguards, went beyond the authorised target 48% of the time. Astra did so in 0% of cases.
Hold both of those together, because together they are the whole story. The most capable offensive security tool OpenAI has ever shipped is also the best-behaved model it has ever shipped, and the good behaviour holds only inside the safeguards OpenAI tested it under. What you deploy is your own harness, your own permissions and your own logging. The system card describes someone else's setup.
How it compares
The competitive picture is less dramatic than launch week suggested.
Artificial Analysis has Astra level with Claude Fable 5.1 on its Intelligence Index at roughly 40% of the cost per task, and level again on the Coding Agent Index at about 60%, driven by the lowest token use of any agent in that index. It leads Terminal-Bench v4.0 at 59% against Fable 5.1's 52% and Sol's 40%, and AutomationBench-AA at 69% ahead of Grok 4.6 at 67% and GLM-5.3 at 62%. Frontier parity, won on efficiency rather than raw capability.
Brockman said something in that briefing considerably more useful than the AGI line, and almost nobody quoted it. He argued that token prices have become a poor way to compare models, because tokens are not comparable between vendors or even across one vendor's own families, and that the metric that matters is price per completed task. OpenAI is already experimenting with charging that way.
He is right, and it should change how you evaluate anything this year. A model priced 2.5 times higher per token that finishes the job in half the tokens is the cheaper model. No pricing page will tell you that, however your own workload will.
The case for not overreacting
Some counterweights, because the launch figures have had an easy ride.
Most of the headline benchmarks come from OpenAI's own materials, measured through OpenAI's own harness. That is normal practice and it is a long way from independent verification.
Benchmarks test fixed, well-specified tasks, while general competence means handling work nobody specified in advance. Topping one benchmark by any margin does not establish the second thing. As with recent models generally, expect strong performance where answers can be verified and considerably more variance everywhere else.
Early hands-on reports are mixed rather than ecstatic. Developers working on complex reverse-engineering describe clear improvement sitting alongside obvious limits of the current paradigm. Others report a real step change in computer use while flagging weak follow-through on front-end design and code review.
None of which is an argument for ignoring Astra. It is an argument for reading launch week as a signal about direction, then measuring the thing yourself on work you actually do.
What to do on Monday
- Test on price per completed task. Take one real workflow you already run, measure Sol or Fable against Astra on end-to-end cost and success rate, and throw the per-token comparison away. OpenAI's own figure has Astra's top configuration cutting cost per task by around 57% against Sol on DeepSWE v1.1. Whether that holds on your workload is an empirical question with your name on it.
- Treat the harness as part of the model. What you are evaluating is Astra plus your context management, your tool permissions and your retry logic. Change any one of those and you have a different system. Version the harness the way you version code, because thirty-seven points of ARC-AGI performance lived in exactly that layer.
- Put nothing with a 51% hallucination rate in front of a client without a verification step. That is arithmetic rather than caution. Either a person approves the output or a second system checks it against a source of record.
- Re-run your evals before swapping models. A model change used to alter answer quality. With computer use, it alters blast radius: different judgement about when to act, different behaviour on ambiguous instructions, the same production permissions.
- Write down what it may touch. Named applications, named data, named write permissions, named approver. When a model can operate software on your behalf, the boundary is the only control you genuinely own.
Then three questions for anyone selling you Astra-based automation. Which harness produced the numbers in your deck, and can you reproduce them in ours? What happens when the system cannot complete the task, and how do you know it stops rather than improvising? And when it acts inside our systems, what record does it leave that our auditors can read?
Where this leaves you
The AGI question is genuinely interesting and almost entirely irrelevant to next quarter. Astra will not wait for the definition to settle, and neither will your competitors.
What changed on 3 September is narrower and more demanding than the AGI framing suggests. A frontier model can now operate your software well enough that the constraint on value has moved off the model and onto everything wrapped around it: the scaffolding, the permissions, the evaluation and the record. Thirty-seven points of benchmark performance sitting in the harness layer is the cleanest published evidence anyone has that this is where results now come from.
That is uncomfortable for the market, because scaffolding does not come with a subscription. It suits any business that already knows what its data is, who it belongs to, and who signs off when software acts on its behalf.
Everyone else is about to find out they hired a very capable operator and never wrote the job description.
Fact or fiction
Seven claims did the rounds during launch week. Here is where each one stands.
- Astra scored 99.9% and Sol scored 7.8%, so Astra is twelve times better. Fiction, as stated. Those two figures come from different harnesses. Compared on the same standard harness it is 62.7% against 7.8%, which is a very large jump and is the number the benchmark supports.
- OpenAI declared that it has built AGI. Fiction. The announcement stopped short of saying so. Greg Brockman used the phrase in a press briefing and framed it as his own view, while leaving the definition to users. The two things got merged in the coverage.
- The benchmark's authors agree AGI has arrived. Fiction. ARC Prize stated it is not claiming this is AGI, and co-founder Mike Knoop wrote that "we lack evidence to call this AGI yet."
- The whole thing is a leak, a fake, or a rebranded Sol. Fiction. OpenAI published an announcement, a system card and benchmark tables on 3 September, the API identifier is gpt-6-astra, and Axios, CNBC and Fox Business all covered it the same day with on-record executive comment.
- Astra saturated ARC-AGI, and ARC-AGI measures general intelligence. Half fiction. The score is real. The inference is where it breaks: the benchmark tests fixed, well-specified environments, and general competence means handling work nobody specified in advance. The name of a test is not a certificate.
- Turn reasoning effort to maximum for the best results. Fiction, at least sometimes. Astra occasionally scores slightly worse at x-high and max than at standard high. Measure your own settings rather than assuming the dial only goes one way.
- Astra is 2.5 times more expensive than Sol. True per token and possibly false in practice. It costs $10 and $50 per million input and output tokens against Sol's $4 and $20, while using markedly fewer tokens to finish comparable work. OpenAI's own DeepSWE v1.1 figure has cost per completed task falling around 57%. Your workload decides which number applies to you.
Two things stayed solidly in the fact column and deserve repeating. Astra is the first OpenAI model rated Critical for cybersecurity capability under the Preparedness Framework. And its hallucination rate at maximum reasoning effort, as measured independently, is 51%.
At U4RIA, we believe AI is a tool to help humans, not replace them. Like what we're about? See what your business is truly capable of. Experience U4RIA.
Sources
- OpenAI: GPT-6 Astra, a new generation of intelligence (3 September 2026) — the release, state-of-the-art claims, FrontierMath Tier 4 at 98%, computer use framing, and the out-of-scope evaluation showing 48% for Sol against 0% for Astra — https://openai.com/index/gpt-6-astra/
- OpenAI Deployment Safety Hub: GPT-6 Astra System Card — the Critical cybersecurity threshold under the Preparedness Framework and the description of autonomous vulnerability discovery and exploitation — https://deploymentsafety.openai.com/gpt-6-astra
- ARC Prize: OpenAI's GPT-6 Astra on ARC-AGI-3 (3 September 2026) — the 62.7% standard harness result at $26,098 against 99.9% via the Provider Adapter at $18,817, the harness definitions, and the 96% human action-efficiency figure — https://arcprize.org/blog/astra
- The Next Web: Astra's AGI score came from a harness, not the model (September 2026) — the like-for-like 62.7% against 7.8% correction, the 49% token and 3.66x speed figures, ARC Prize declining the AGI conclusion, Mike Knoop's comment, and OpenAI's post-launch revision of five metrics — https://thenextweb.com/news/openai-astra-arc-agi-3-harness-62-7-vs-99-9-benchmark-revisions
- Axios: "Welcome to the AGI era," OpenAI says as GPT-6 Astra debuts (3 September 2026) — Brockman's briefing comments, the 100,000-GPU Stargate training run, and the model-supervising-model detail — https://www.axios.com/2026/09/03/openai-astra-gpt-6-agi-brockman
- The Decoder: GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era" (September 2026) — Brockman on price per completed task rather than per token, and the DeepSWE v1.1 figure of roughly 57% lower cost per task — https://the-decoder.com/gpt-6-astra-is-the-first-model-making-openai-willing-to-declare-the-agi-era/
- Artificial Analysis: Benchmarking GPT-6 Astra (September 2026) — the AA-Omniscience hallucination rate falling from 92% to 51%, Intelligence and Coding Agent Index parity with Claude Fable 5.1 at 40% and 60% of cost, Terminal-Bench v4.0 and AutomationBench-AA results — https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra
- Marcus on AI: Hot take on GPT-6 Astra (September 2026) — the argument that ARC-AGI success is not proof of AGI, the robustness question on open-ended tasks, and the observation on pre-launch access — https://garymarcus.substack.com/p/hot-take-on-gpt-6-astra
- The New Stack: OpenAI will sell you Astra, but not the system that scored 98.6% on ARC-AGI-3 (September 2026) — Matt Turck's caveat and the harness-engineering framing — https://thenewstack.io/openai-astra-harness-arc-agi-3/
- Microsoft Azure: GPT-6 Astra now generally available in Microsoft Foundry (September 2026) — general availability in Foundry and the framing that the next era of enterprise AI will not be defined by chat — https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-intelligence-for-work-now-generally-available-in-microsoft-foundry/
- OpenAI API: GPT-6 Astra model page — pricing, context window, and reasoning effort settings — https://developers.openai.com/api/docs/models/gpt-6-astra
- OpenAI Developer Community: GPT-6 Astra is good, but still far from what I would call AGI (September 2026) — hands-on report on limits of the current paradigm in complex reverse-engineering work — https://community.openai.com/t/gpt-6-astra-is-good-but-still-far-from-what-i-would-call-agi/1395147
- U4RIA: Agent Harnesses: How They Are Beating the LLMs (17 July 2026) — https://www.u4riaai.com/articles/agent-harnesses-how-they-are-beating-the-llms
- U4RIA Articles — https://www.u4riaai.com/articles