OpenAI Says "Welcome to the AGI Era." The Data Disagrees
OpenAI released GPT-6 Astra calling it the start of the "AGI era," but its own benchmark scores trail Claude on core reasoning tests.
Two words closed OpenAI's launch briefing for its newest model on Thursday, September 3: "Welcome to the AGI era." The company's president, Greg Brockman, said them deliberately, in front of reporters, about GPT-6 Astra, OpenAI's new flagship system. It is the most consequential sentence a frontier AI lab has said in public this year. It is also, by the company's own published benchmark data, not quite backed up by the numbers sitting one page below it.
What Actually Shipped
Astra replaces GPT-5.6 Sol as OpenAI's flagship model, arriving just over a year after GPT-5's original release and built, according to the company, on its largest training run to date, using more than 100,000 GPUs at its Stargate site in Texas. OpenAI is rolling it out in stages rather than all at once: enterprise customers in its Daybreak cybersecurity program received access first on launch day, with ChatGPT Plus, Pro, Business, and Enterprise subscribers following in what the company describes as "the coming days," alongside developer access through the OpenAI API, Amazon Bedrock, and Microsoft Foundry. Pricing lands at the premium end of the current frontier market, $10 per million input tokens and $50 per million output tokens on the standard API, roughly two and a half times the promotional rate OpenAI had charged for Sol.
The headline capability, and the one OpenAI's own materials lead with, is computer use. Brockman described Astra as able to do "anything a human can do with a computer," and OpenAI's demonstrations backed that framing with genuinely varied examples: the model formatting a legal contract, building a working 3D game, laying out a printed circuit board in KiCad, animating a car transmission in FreeCAD and Blender, and filling out a tax return draft directly from a W-2 form, all without step-by-step human direction at each stage.
The Sentence Doing More Work Than the Science
Brockman's framing was carefully hedged even as it made headlines. Asked directly whether Astra represented the arrival of AGI, he told reporters, "I think it might be about this model," before closing with the more quotable, less qualified line that immediately traveled everywhere: "Welcome to the AGI era." In a separate account of the same briefing, Brockman reportedly acknowledged that AGI "lacks a universally accepted definition," calling it "a gray, fuzzy thing" rather than a single, measurable threshold a model either crosses or does not.
That acknowledgment matters enormously, because it means OpenAI is not actually claiming Astra passed some agreed-upon technical bar for general intelligence. Notably, OpenAI's own official launch page does not use the word AGI at all. It calls Astra "the world's best computer use model" and its "most intelligent and aligned model," considerably narrower and more defensible language than the sentence its own president chose to end the press briefing with. A meaningful gap between what a company puts in its official documentation and what its president says out loud to reporters is itself worth noticing, particularly for a claim this consequential.
What Brockman Said About AGI Seven Years Ago
The clearest evidence that this framing deserves real scrutiny, rather than automatic acceptance, comes from Brockman's own past words. Kingy.ai's detailed analysis of the launch dug up Brockman's 2019 description of AGI: a system mastering fields at world-expert level across more domains than any human, comparing it to having a Marie Curie, an Alan Turing, and a Johann Sebastian Bach combined in one mind. Astra's own published results do not support that bar being met. On Humanity's Last Exam, a benchmark specifically designed to test the outer edge of expert-level reasoning across disciplines, Astra scored 57.2 percent, a result that trails several of Anthropic's Claude models on the same test, according to TechSpot's review of the released benchmark data.
That is a genuinely awkward number for a company simultaneously declaring the arrival of general intelligence. Either Brockman's own 2019 definition of AGI has quietly loosened considerably by 2026, from a multi-domain polymath standard to something closer to "a model capable of extended, autonomous work on ordinary computer tasks," or the "AGI era" framing is doing more rhetorical work than the underlying benchmark data can actually support. The charitable reading, one commentator noted, is that Brockman now sees AGI as describing a trajectory or historical era rather than a single completed checklist. The skeptical reading is a familiar pattern in AI marketing: when the model finally arrives, the definition of the milestone quietly expands to fit it.
The Cybersecurity Number That Isn't Rhetorical at All
Separate from the AGI framing, one part of Astra's release carries none of that ambiguity, because OpenAI confirmed it directly days before this launch: Astra is the first model in company history to be formally classified as reaching the "Critical" tier under OpenAI's own Preparedness Framework for cybersecurity risk, the designation reserved for a model capable of independently discovering and chaining together working exploits against hardened, real-world systems without step-by-step human guidance. Astra scored a perfect 100 percent on ExploitBench and discovered two genuine, previously unknown zero-day vulnerabilities during evaluation, which OpenAI has since disclosed to the relevant software maintainers.
That classification is precisely why the model's rollout to most users remains restricted well past launch day. OpenAI is limiting Astra's most advanced cybersecurity capabilities to vetted organizations through a program it calls Daybreak Blue, prioritizing entities responsible for protecting critical infrastructure, while broader access to those specific capabilities stays subject to tighter restrictions and ongoing monitoring. OpenAI also disclosed that the model went through the White House's voluntary vetting process for advanced AI systems before release, and Brockman said that review did not require any changes to Astra's existing safeguards, a detail worth noting precisely because it is the kind of claim regulators, not the company itself, are best positioned to verify independently over time.
The Competitive Backdrop This Launch Landed Inside
The timing of Astra's release is not incidental. It arrived just two days after Anthropic released Claude Fable 5.1, and Astra's own benchmark results sit in direct, public tension with Anthropic's models on at least the Humanity's Last Exam comparison already noted. That rivalry has been intensifying on every front recently, not just model capability. Anthropic posted its first-ever operating profit this month, roughly two years ahead of its own internal financial guidance, a milestone that gives the company genuine standing to compete on business fundamentals rather than purely on benchmark leaderboards. The two companies are also now competing directly for the same institutional customers in far more consequential arenas than consumer chatbots: OpenAI's models recently joined the Pentagon's GenAI.mil platform alongside xAI's Grok, while Anthropic's Claude remains conspicuously excluded following its own legal fight with the department over surveillance guardrails. Reading Astra's "AGI era" declaration purely as a scientific claim, rather than partly as a competitive one timed against a well-resourced rival's own major release, would be missing half of what is actually happening here.
Why This Framing Matters Beyond the Semantics
There is a genuine, practical reason the AGI question is not just an academic argument over terminology. OpenAI's own corporate structure has historically tied real financial and governance consequences to a formal AGI determination, a threshold that would trigger specific contractual changes in the company's relationship with Microsoft and its broader investor base. One analysis of the launch noted that Brockman himself acknowledged the company's old contractual AGI trigger "no longer exists" in its original form, describing AGI now more as a mission or aspirational concept than a hard technical or legal line. That is a meaningful admission buried inside an otherwise triumphant announcement: the specific, binding definition of AGI that used to carry real consequences for OpenAI's own governance has been quietly loosened at almost exactly the moment the company is publicly declaring the term's arrival.
What to Actually Watch From Here
Astra's genuine, well-documented capabilities, extended autonomous computer use, a perfect score on a serious exploit-discovery benchmark, and a training run built on more than 100,000 GPUs, are real and substantial advances regardless of how the AGI framing eventually shakes out. What remains unresolved is whether "AGI era" turns out to be a label historians settle on in retrospect, the way OpenAI is clearly betting it will, or whether it becomes another instance of a term stretched loose enough to fit whatever model a company happens to be launching that particular week. The most honest reading of Thursday's announcement is the one several independent analysts converged on directly: OpenAI has built a genuinely capable, genuinely risky model, and has chosen to wrap it in language its own benchmark data, and its own president's older definitions, do not fully support yet.
Written by
Mr. Aayush Bhatt
Software Engineer interested in how models work and where they fail.




