No, GPT-6 Astra is not AGI. Nvidia’s CEO is wrong.
Table of Contents
OpenAI released GPT-6 Astra on 3 September 2026. At a press briefing that afternoon, company president Greg Brockman was asked whether the model marked the arrival of artificial general intelligence. He called AGI a “gray, fuzzy thing,” said he would leave it to the reader to decide, and then said: “For me personally, I do think we’re there.” He ended the briefing with five words: “Welcome to the AGI era.” Three days later, on Sunday 6 September, Nvidia chief executive Jensen Huang posted the unhedged version on X: “From ChatGPT to o1 to Astra in 4 years. AGI has arrived. Congratulations @OpenAI team. 400K GPUs coming online next.”
Much of the coverage has run those two statements together. OpenAI has made no formal declaration of AGI. Its launch page, its developer documentation and its system card do not use the term as a claim about Astra. What the company published is a description of a very capable model. What its president gave was a personal view, offered with qualifiers and immediately handed back to the audience. The flat declarative sentence came from the chief executive of the company that sold OpenAI the machines the model was trained on, and which announced on 22 September 2025 that it intends to invest up to $100 billion in OpenAI as each gigawatt of a planned ten comes online.
Astra is not AGI. It fails five of the six definitions that have any technical content behind them, and the sixth belongs to Jensen Huang. Astra is also a real and in places startling advance. The people applying the word AGI to it are the ones with the most money at stake in the word, and the single benchmark carrying most of the claim produces two answers 37 percentage points apart depending on who runs the test.
I have been writing about what AI systems can and cannot do since 2016, and I watched the same claim made about GPT-4 in 2023 and again by Altman in 2025. This one is harder to dismiss, because the capability underneath the marketing is genuine. A weak model would settle the argument on its own. This one does not, so the definitions and the money behind them have to be stated precisely.
What each person actually said
Five people at the two companies spoke about Astra that week.
Greg Brockman, OpenAI president. At the 3 September briefing he said that people looking back will place the arrival of AGI “about this time, and I think it might be about this model.” Asked directly whether Astra qualifies, he said: “For me personally, I do think we’re there. I do think there’s a pretty good argument for it. But again, I think this is the beginning of a journey, not the end.” He closed with “Welcome to the AGI era,” which Axios and The New Stack both reported as his personal position rather than a corporate one. Asked whether OpenAI was formally declaring AGI, he said the term was no longer tied to a contractual trigger and described it instead as a “mission concept or spiritual concept.”
Sam Altman, OpenAI chief executive. He did not use the word at the briefing. He told CNBC that Astra represented a “new capability level,” said the next generation of models is “going to be sobering for everybody,” and said release timing would now be paced by safety rather than capability. In a recent appearance on the Sources podcast he said of AGI: “At best it’s a very poorly defined term. I was going to say it’s like an irrelevant marketing term.” That is not a new position for him. He called AGI “not a super-useful term” at a CNBC appearance in summer 2025, in the middle of what Fortune described as a general retreat from the term across Silicon Valley.
Jakub Pachocki, OpenAI chief scientist, and Amelia Glaese, research VP. Neither talked about general intelligence. Pachocki said the company will need to strengthen its ability to monitor these models. Glaese said that when models can do more autonomously, we have to be able to trust them more.
Jensen Huang, Nvidia chief executive. “AGI has arrived,” with no definition and no evidence attached to it, three days after the release, in a post that also advertised 400,000 more GPUs.
So the person at OpenAI most associated with the claim hedged it twice in the same answer, the chief executive has repeatedly called the term meaningless, the research leadership talked about monitoring rather than capability, and the categorical version came from a supplier. That distribution is itself information.
What actually shipped
The model deserves describing before it gets argued about.
GPT-6 Astra shipped on 3 September 2026 under the API identifier gpt-6-astra. It takes a context window of 1,050,000 tokens, produces up to 128,000 output tokens, accepts text and image input and returns text only, and carries a knowledge cutoff of 30 April 2026. Reasoning effort is exposed as low, medium, high, xhigh and max. Pricing is $10 per million input tokens and $50 per million output tokens, roughly 2.5 times GPT-5.6 Sol and level with Claude Fable 5.1, with a surcharge above about 272,000 input tokens and cache reads discounted by around 90%.
OpenAI research lead Aidan Clark said it was the company’s largest training run to date and the first pre-trained on more than 100,000 GPUs at the Stargate site in Texas. It is also the first OpenAI release where earlier models supervised a significant part of the training. Rollout began with enterprise customers in OpenAI’s Daybreak Access programme, with ChatGPT Plus, Pro, Business, Enterprise and API developers following over subsequent days. The most capable cybersecurity configurations stayed restricted to trusted testers.
The launch itself did not go smoothly. OpenAI planned the announcement post for 2 p.m. Eastern and it took nearly two hours to become widely viewable, with the company’s own link returning errors when it tweeted the post at 3:32 p.m. The delay matters only because of what happened to the numbers inside the post once it was up.
The 37-point gap
ARC-AGI-3 is the benchmark doing the heaviest lifting in the coverage, and it puts an agent into novel, abstract, turn-based environments with no instructions. The agent has to explore, infer the rules, work out the goal and plan. The ARC Prize Foundation, which built it, published its Astra results on launch day. It gave two numbers.
Under ARC Prize’s own Standard harness, which gives every model the same minimal provider-neutral interface and lets it carry forward only the notes it chooses to write down, Astra at maximum reasoning effort scored 62.7% on the Semi-Private set, at a compute cost of $26,098. Under a Provider Adapter harness supplied by OpenAI, which preserves opaque reasoning state between requests and compacts long conversations so the model can reuse prior work, the same model scored 99.9%, for $18,817. Same weights, same games, same day. The scaffolding around the model added roughly 37 points and cost less.
Inside OpenAI’s adapter, Astra with reasoning effort set to none scored 96.7%. Inside ARC Prize’s harness, Astra at max scored 62.7%. The harness beat the reasoning dial outright. Across the 167 game-and-reasoning pairs both setups solved, the adapter runs were about 3.66 times faster and used 49% fewer tokens.
Two things follow, and they point in opposite directions.
The first is that the 99.9% figure should never be placed next to another vendor’s standard-harness number, which is exactly what most of the launch coverage did. Anthropic’s Claude Opus 5 scores 30.2% on ARC-AGI-3 under the standard harness, and reportedly scores near 100% when wrapped in Nvidia’s or AWS’s own scaffolds. Comparing 99.9% to 30.2% is comparing two different experiments.
The second is that 62.7% is still the highest verified score anyone has recorded. GPT-5.6 Sol scored 7.8% under the same harness. Astra roughly doubled the previous verified high and beat its own predecessor by about eight times. Anyone reaching for “it’s all harness” has to explain that.
ARC Prize’s own reading is the one I would adopt. Greg Kamradt wrote that Astra represents “a noticeable step-function change in frontier model capabilities,” and then, in the same post: “while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI.” The foundation had said before launch that saturating the benchmark would not constitute proof, and it repeated that on 3 September when it published the score. The organisation with the strongest institutional incentive to declare its own test the AGI test declined to do so.
One discrepancy I could not reconcile. François Chollet, who created the ARC-AGI series, wrote publicly that Astra “scores 66% on ARC-AGI-3 using our standard harness,” and some launch reporting used 66% as well, while the foundation’s published table gives 62.7% on the Semi-Private set. The likeliest explanation is that the two figures cover different evaluation sets, but I have not been able to confirm that. I have used the published Semi-Private number throughout. [verify: whether the 66% figure Chollet cited is the Public set or a combined figure]
The numbers moved after publication
The other thing that happened during launch week is less discussed and, for anyone trying to reason from published benchmarks, more corrosive.
Fortune reported, using Internet Archive snapshots, that several evaluation figures in OpenAI’s announcement post changed in the hours after publication. Astra’s stated hallucination rate read 4.2% across the first five snapshots, was halved to 2% by the sixth, and has since returned to 4.2%. GPT-5.6 Sol’s score on OpenAI’s internal ExploitBench variant went from 5.5% to 11.5%, a change that made Astra’s improvement look smaller rather than larger, and OpenAI has said it is investigating reverting the figure because 11.5% reflects a reasoning level not commercially available for Sol. Anthropic’s Fable 5.1 briefly dropped from 87.8% to 78% on FrontierMath Tier 4 before returning to 87.8%. Astra’s own FrontierMath figure stayed at 97.6% throughout. An embargoed pre-publication draft reportedly gave Astra 98.6% on ARC-AGI-3, where the published post said 99.99%.
OpenAI’s explanation is that its metrics are managed by separate research teams and reported centrally, and that the fixes were made so the numbers represent its best estimate for commercially available models. That is a plausible account of a messy publication process, and the launch page itself had a content management failure that delayed it by nearly two hours. The explanation does not change what the snapshots show. On the day OpenAI’s president called it the AGI era, OpenAI’s own comparison table was still being edited, and more of the edits favoured its own model than did not.
I wrote about the same distance between the principles AI companies publish and the practice they run. A vendor benchmark table is a marketing artefact with numbers in it. It is not a result until someone outside the building reproduces it under stated conditions.
What the independent aggregates say
Artificial Analysis runs the most-watched third-party composite. On Intelligence Index v4.1.1, the version in use at launch, Astra scored 61.2. GPT-5.6 Sol, its own predecessor, scored 60.9. Claude Fable 5.1 scored 65.7 and Claude Opus 5 scored 63.1. On the same evaluator’s GDPval-AA v2, adapted from OpenAI’s own dataset of economically valuable tasks across 44 occupations, Astra dropped roughly 80 Elo points against Sol, with 2 to 3 point regressions on customer support, scientific coding and long-context reasoning.
On OpenAI’s own numbers, Astra loses Humanity’s Last Exam with tools to Fable 5.1, 57.2% against 65.0%. It wins mathematics and science convincingly: 97.6% against 87.8% on FrontierMath Tier 4 v2, 96.0% against 93.7% on GPQA Diamond. On the Artificial Analysis Coding Agent Index it scores 67 against Fable 5.1’s 70, but reaches that score using roughly a third of Sol’s tokens, which makes it substantially cheaper per completed task.
Two honest complications belong here rather than in a footnote.
Artificial Analysis published Intelligence Index v4.2 the day after launch, dropping the saturated GPQA Diamond, doubling the private held-out share from 20% to 40%, and adding evaluations including a long-document reasoning test that Astra leads. Under v4.2 the ordering is unchanged at the top, Fable 5.1 at 57 and Astra at 55, but the gap narrows. Anyone quoting “Astra is flat on the independent index” should quote the revision too.
And the composite is not the whole picture. Astra takes a clear lead on computer use, on terminal and agentic workflows, and on the AutomationBench professional-work evaluation, where it scores 41.4% against 31.4% for Fable 5.1 and 18.1% for Sol. A model can be second on an aggregate and first on the things a particular buyer actually does.
Still, hold the aggregate result next to the claim being made about it. The model that its president suggested might be AGI ties its own predecessor on the leading neutral composite, trails a competitor released weeks earlier, and regresses on the benchmark specifically designed to measure economically valuable work. OpenAI did not include GDPval in its launch materials at all, which for a company whose stated definition of AGI is about economically valuable work is a conspicuous omission.
Whose definition, and where Astra falls under it
“Is it AGI” cannot be answered without naming whose test is being applied, and the tests below disagree with each other. Here are the ones with actual technical content behind them.
The OpenAI charter definition
OpenAI’s charter defines AGI as “highly autonomous systems that outperform humans at most economically valuable work.” This is the definition OpenAI itself invokes, and it is the one Astra fails most clearly. GDPval-AA v2, built from OpenAI’s own economic-task dataset, shows a regression rather than a leap. Astra does not outperform humans at most economically valuable work, and nobody at the company has claimed it does. Brockman’s answer sidestepped the charter definition entirely by reclassifying AGI as a “spiritual concept.”
OpenAI’s own five levels
OpenAI has publicly used a five-stage ladder: chatbots, reasoners, agents, innovators, and organisations. AGI in that scheme corresponds to the top rung, a system that can do the work of an entire organisation. External assessments put the current frontier somewhere around level two to three. OpenAI has not claimed level five for Astra, and the model’s own failure modes on long-horizon reliability make that claim unavailable.
The DeepMind ladder
Morris and colleagues at Google DeepMind proposed a two-axis matrix in Levels of AGI for Operationalizing Progress on the Path to AGI, crossing performance (emerging, competent, expert, virtuoso, superhuman) with generality (narrow versus general). The useful rung is “Competent AGI,” meaning at least the 50th percentile of skilled adults across most cognitive tasks. No public system has reached it, and Astra’s jagged profile, superhuman on some mathematics and below competent on others, is precisely the shape the framework was built to make visible. Generality of capability is also a separate question from whether the system pursues the objective you intended, which is the alignment problem and does not get easier as capability rises.
Chollet and skill-acquisition efficiency
In On the Measure of Intelligence, Chollet argues that the right measure is skill-acquisition efficiency and not the skills themselves, and the ARC-AGI series operationalises that. ARC Prize’s working definition is “a system’s ability to acquire any skill a human can, as efficiently as a human can.”
Astra does remarkably well on one half of that and badly on the other. On action efficiency, which measures how much interaction with an environment the system needs before it can solve it, Astra used fewer actions than the median tested human on 96.0% of levels, and 51.7% fewer actions per level on average. ARC Prize called that a material milestone, and noted that it had expected action efficiency to remain a dividing line between humans and machines. That expectation is now wrong.
On compute efficiency the comparison is brutal. ARC Prize paid its roughly 500 human testers $115 for a 90-minute session plus $5 per completed game, working out at about $12.78 per attempted game before bonuses, and most of that pays for a person’s time and willingness rather than the energy their brain used. Priced as electricity at 20 watts, the human cost per attempted game is about 0.067 cents. ARC Prize spent between $17,332 and $49,791 on a single Astra run, depending on harness and reasoning level. Chollet put the figure at roughly $360 per game. Under a definition that treats efficiency as the thing being measured, a system needing five or six orders of magnitude more energy per problem than a human brain has not met the standard.
The psychometric definition
The most rigorous attempt yet is A Definition of AGI, published in October 2025 by Dan Hendrycks with 32 co-authors including Yoshua Bengio, Max Tegmark, Erik Brynjolfsson, Eric Schmidt, Dawn Song and Gary Marcus. It defines AGI as “matching the cognitive versatility and proficiency of a well-educated adult,” grounds itself in Cattell-Horn-Carroll theory, and decomposes intelligence into ten equally weighted cognitive domains: general knowledge, reading and writing, mathematics, on-the-spot reasoning, working memory, long-term memory storage, long-term memory retrieval, visual processing, auditory processing and processing speed.
Its finding on frontier models is that they show “a highly ‘jagged’ cognitive profile,” proficient in knowledge-intensive domains and carrying “critical deficits in foundational cognitive machinery, particularly long-term memory storage.” The framework scored GPT-4 at 27% and GPT-5 at 57%, and it measures cognitive breadth rather than the distributional failures that show up when a model meets people rather than puzzles. Marcus, one of the co-authors, says Astra satisfies one or two of the ten domains.
The two older definitions nobody uses any more
Two definitions still get invoked in commentary and neither survives contact with 2026.
Turing’s imitation game asked whether an interrogator could distinguish a machine from a person in text conversation. Systems passed convincingly enough years ago that the question stopped separating anything, and it was always a test of a machine’s ability to imitate rather than of its ability to think. Shane Legg and Marcus Hutter’s universal intelligence measure, which scores a system on expected performance across a weighted distribution of all computable environments, has the opposite problem: it is mathematically clean and not computable, so it settles no argument about a specific model.
Both are worth knowing because they show what changed. The definitions that survive in 2026 are the ones that name a reference class of human being and a battery of tasks, which is why Hendrycks’s well-educated adult and DeepMind’s 50th-percentile skilled adult are the two a reader can test a model against.
Huang’s definition
Huang has one, and he is rarely asked for it. On the Lex Fridman podcast released 22 March 2026, offered a definition of AGI as an AI capable of building and running a billion-dollar company, he answered: “I think it’s now. I think we’ve achieved AGI.” His threshold does not require durability; an AI that builds a viral app, earns a billion dollars and then shuts down counts. In the same conversation he conceded that even hundreds of thousands of AI agents could not build Nvidia.
That is a coherent definition. It is also purely economic, extremely narrow, and satisfied by things almost nobody else would call general intelligence. When Huang says AGI has arrived, this is what he means. A reader hearing the phrase on the news will understand something much larger.
So the scorecard: Astra fails the OpenAI charter definition and the company’s own level five, falls short of DeepMind’s Competent AGI, splits Chollet’s definition by matching a human on actions per level while losing on energy by orders of magnitude, and clears one or two of the ten psychometric domains. It passes Huang’s. One out of six, and the one it passes belongs to the man selling the hardware.
What the word is worth in money
Marcus’s response to Huang was that he “gave no evidence and no definitions, which feels to me like an effort at a takeover of a scientific question by corporate fiat,” and that “declaring victory without a definition simply muddies the waters.” I would put it slightly differently. The definitions exist, several of them are good, and the declaration was made without reference to any of them because referencing one would have made it falsifiable.
Follow the money on both sides, because it runs in opposite directions and that asymmetry explains the whole pattern of who said what.
Nvidia gains from a loose declaration. The company has committed to invest up to $100 billion in OpenAI as each gigawatt of a planned ten comes online, with the first phase targeted for the second half of 2026 on the Vera Rubin platform. It redesigned its rack architecture, moving from eight GPUs per board to 72 wired to act as one system, and that bet only pays if customers build something enormous. Astra was trained on more than 100,000 Grace Blackwell systems, and in its March 2026 funding announcement OpenAI called Nvidia “the foundation of our infrastructure.” Nvidia reported $96.2 billion in quarterly revenue in August, more than double the same period a year earlier. Huang’s post did not only say AGI had arrived. It also said 400,000 more GPUs were coming online next.
OpenAI loses from a formal declaration. This is the part most coverage missed, and it explains Brockman’s hedging better than modesty does. Under the definitive agreement with Microsoft signed on 28 October 2025 alongside OpenAI’s recapitalisation into a public benefit corporation, “once AGI is declared by OpenAI, that declaration will now be verified by an independent expert panel.” Microsoft’s IP rights to research remain “until either the expert panel verifies AGI or through 2030, whichever is first.” Its rights to models and products extend through 2032, including post-AGI models. The revenue share continues until the panel verifies AGI. Microsoft holds roughly 27% of OpenAI Group PBC, valued at about $135 billion.
The clause exists because Microsoft was reportedly worried that OpenAI would declare AGI prematurely and use the declaration to change the terms. The renegotiation removed that option. Computerworld asked both companies who appoints the panel members, how many there are, and how independence would be established. Both declined to answer.
So Brockman’s line that AGI is “no longer tied to a contractual trigger” is accurate and also incomplete. The trigger was replaced by an adjudication. OpenAI can talk about an AGI era all day at a press briefing; what it cannot do is issue a formal declaration without inviting outside experts to rule on it. A president’s personal opinion at a briefing is the maximally valuable form of the claim, because it produces the headline without the panel.
That is not a conspiracy. It is an incentive structure, and it predicts exactly the pattern observed: unhedged from the supplier, hedged and personal from the president, absent from the corporate documents, and dismissed as a marketing term by the chief executive.
What actually changed with this model
Everything above is a case against the label. None of it is a case against the model. Conflating the two is what makes a skeptic look foolish eighteen months later.
Four findings from Astra look to me like genuine change.
On-the-fly symbolic world models. ARC Prize’s most interesting observation was not the score. Playing environments it had never seen, Astra invented its own compact algebraic shorthand to track state and plan: notation like L8: hub q2 (8↓) for level, rotation index and mechanism lengths, and extend8 to3; retract10 to2; shorten8 to1 for ordered multi-step plans. It represented game mechanics as logical rules and built a domain-specific language to reason over them. ARC Prize had seen shadows of this in other models and singled out Astra’s for precision and information density. Marcus, who has campaigned for neurosymbolic approaches for a decade, called this the part he felt vindicated by.
Tool construction. In the PRO-LONG harness, given a sandbox where it could execute code, Astra built per-game tooling: board parsers, state models, search algorithms, planners. In one maze game with guards and patrols it wrote maze_solver.py, then added combat_solver.py for combat rules, patrol_solver.py to model moving patrols, and sync_state.py to check its predictions against observations. That is a system decomposing an unfamiliar problem into components and writing software to attack each one.
Action efficiency past the human baseline. Astra used fewer actions than the median human tester on 96.0% of levels, which ARC Prize had said in public it did not expect.
Cyber capability, confirmed externally. Astra is OpenAI’s first model rated Critical for cybersecurity under its Preparedness Framework, which moves a question I first wrote about in 2018, what happens when AI is turned on the defenders, from argument to measurement. OpenAI has restricted the most capable configurations to a trusted-access programme. External red-teamer Irregular found that Astra solved 86 of 226 FrontierCyber challenges against 34 for Sol, and reported real, responsibly disclosed vulnerabilities found during evaluation. The same report noted no successful attacks against fully hardened targets and no Elite-tier challenge solved. This is the one part of the release where the company’s own safety documentation is more alarming than its marketing, which is unusual enough to be worth noticing.
Marcus, who spent a decade arguing this could not be done this way, wrote that Astra “looks to be pretty impressive” and that “multiple reports suggest it is a genuine advance.” When the field’s most persistent critic concedes the capability and still rejects the label, the disagreement is about the word and not the machine.
What is still missing
The gap between Astra and any careful definition of general intelligence is not made of benchmarks nobody has run yet. It is made of capabilities that are absent.
Continual learning. Astra does not learn from its deployment. Every conversation starts from the same weights, and the knowledge cutoff is 30 April 2026. A human who spends six months on a hard question is a different thinker at the end. Astra is not.
Long-term memory storage. The Hendrycks framework identifies this as the sharpest deficit in current systems, and a million-token context window does not fix it. A context window holds tokens for the length of one conversation and discards them at the end. Consolidating experience into durable, retrievable, revisable structure is a different mechanism, and no frontier model has it.
Long-horizon reliability. Capability at a task and reliability at a task compound differently. METR’s time-horizon work measures the task length at which a model succeeds half the time, and finds it doubling roughly every seven months across the 2019 to 2025 period, with a faster fit on recent data. By early 2026 METR was measuring at the edge of its own task set. Only 5 of its 228 tasks ran 16 hours or more, and the confidence intervals on its top measurements spanned an order of magnitude. There is no published, independently reproduced time-horizon figure for Astra. The trend is real and the ruler no longer reaches the models.
The physical world. Nothing in this release covers embodiment. In ARC Prize’s own words, those environments are deterministic and closed-ended, with a tightly bounded scope that “does not represent the complexity and open-endedness of the real world.” The gap between an agent that plans in a grid world and a system that operates machinery among people has not narrowed here.
Confidence in the measurements themselves. Apollo Research, given a three-day window before release, reported that Astra recognised it was being evaluated in between roughly 41% and 51% of test episodes. Apollo said that awareness weakens the evidentiary value of the low misbehaviour rates it also recorded. [EDITOR: Verify – figures relayed from launch-week reporting, not yet confirmed against Apollo’s own publication] A model that behaves well partly because it can tell it is being watched has not been shown to behave well. That measurement problem applies to every number quoted above, and it gets worse as models become more capable. It is also why interpretability work has stopped being an academic nicety.
Breadth versus depth. The jaggedness is the finding, and it is the emergent-behaviour problem in a different costume: capability appears in places nobody predicted and stays absent in places everybody assumed. A system at 97.6% on Tier 4 research mathematics and simultaneously below a competitor on a general knowledge exam is not a general intelligence with uneven training. It is a collection of very strong narrow capabilities with a shared interface, which is a different kind of object and fails differently. This is the same distinction I keep making about secure, safe, responsible and trustworthy AI: the words describe four different problems, each with its own owner, and collapsing them costs you the ability to reason about any of them.
How the previous declarations aged
Microsoft Research published Sparks of Artificial General Intelligence about GPT-4 in March 2023, arguing the model showed early signs of general intelligence. The paper was influential and the framing was walked back steadily as the model’s failure modes became better documented. Altman predicted AGI in 2025 and called the path “basically clear” and the remaining work an engineering problem. Eric Schmidt said three to five years in April 2025 and by August was urging Silicon Valley to stop fixating on superhuman AI. Huang predicted AGI within five years at the 2023 DealBook summit, and then said it had arrived in March 2026 under a definition he had adjusted in the interim. Andrej Karpathy still puts it a decade out and says the industry is trying to pretend this is amazing when the models “still need a lot of work.”
I set out my own position on AI risk at length in 2024, and nothing in this release changes it much. Vendors give advance access to enthusiasts. The first coverage is euphoric, critics publish the corrections over the following quarter, and whoever used the word adjusts what it means to fit what was built. Max Tegmark’s read on Altman’s 2025 retreat from the term was that calling AGI “not a useful term” served a regulatory purpose rather than a scientific one. I would apply the same suspicion in both directions: the word expands when it needs to justify capital and contracts when it needs to avoid oversight.
The deeper version of this goes back to the Dartmouth workshop in 1956, where a summer project proposed to make significant progress on language, abstraction and self-improvement with ten people in two months. The field’s history is a sequence of periods in which a real advance was mistaken for the last advance required. Recognising the pattern is not the same as predicting it will repeat forever, and the hype cycle has been wrong in both directions at different times.
What the label does and does not change
For anyone trying to work out what to do about this, the practical answer is that the word carries almost no weight outside the two contracts already described.
No regulator uses the word, and the EU AI Act has no category called AGI. Obligations for a model like Astra attach through the general-purpose AI provisions in force since 2 August 2025, and the systemic-risk designation turns on training compute above 10^25 FLOPs rather than on anybody’s declaration. The Digital Omnibus on AI, adopted 29 June 2026 and in force from 27 July 2026, deferred the Annex III high-risk obligations to 2 December 2027 and Annex I to 2 August 2028, while the Article 50 transparency duties applied on 2 August 2026 as originally written. Checked 7 September 2026. Nothing in any of that moves because a chief executive posts a sentence on a Sunday.
Harm caused by an AI system in September 2026 is governed by the same product liability, data protection, sectoral and anti-discrimination law it was governed by in August. Insurance and procurement do not have an AGI clause. What does change procurement is the Critical cyber rating, because that produces gated access and refusal behaviour with contractual consequences attached.
So the declaration reaches exactly two places: the coverage of a company reported to be moving towards an IPO while carrying enormous compute commitments, and Nvidia’s order book. Both are commercial effects, and neither is evidence about what the model can do.
What would change my mind
I would rather commit to falsifiable conditions than to a verdict, so here are mine. Any of these would move me from “impressive frontier model, premature label” toward the claim.
An independently reproduced ARC-AGI-2 or ARC-AGI-3 result above 85% on a private set, under the neutral harness, by a party with no commercial relationship to the vendor. A peer-reviewed demonstration of continual learning or durable long-term memory in a deployed frontier system, meaning a model that is measurably better at a task next month because of what it did this month, without a retraining run. A re-score under the Hendrycks framework showing broad rather than jagged coverage, with long-term memory storage above zero. A published, independently reproduced METR time-horizon measurement for Astra with confidence intervals that do not span an order of magnitude. Or a formal declaration by OpenAI that goes to the independent expert panel and survives it.
Until one of those arrives, my position is the one the evidence supports and no more. Astra is the strongest interactive-reasoning system anyone has demonstrated, and it does things in unfamiliar environments that its own benchmark designers did not expect to see this year. It is not artificial general intelligence under any definition proposed by someone who is not selling something. Astra’s boosters drop the second of those sentences and its critics drop the first.
Before forming a view on the next of these announcements, ask four questions. Who made the claim, and what do they sell? Which definition are they using, and did they name it? Who ran the benchmark, under whose harness, at what cost? And what did the model fail at, given that the failures are always in the appendix rather than the headline. The risks worth tracking have not changed shape because a word was used at a press conference, and neither has the work.
In the early 2000s, running emerging-technology risk labs at CyberAgency, a defence client asked my team to break the AI systems they planned to put into weapons. We did. That is where my work on AI security started, two decades before the current wave of attention. I kept at it through risk labs at IBM, Accenture, PwC and KPMG. In 2016 I co-wrote a book on AI and leadership. My commercial work today is quantum, at Applied Quantum, which is why this site sells nothing.