OpenAI didn’t hedge when it launched GPT-6 Astra on September 3, 2026. President Greg Brockman told the room to get ready for “the AGI era” and said he personally believes the company has crossed that line. That’s not a random influencer take — that’s the company’s president, on the record, at the launch briefing.
So naturally, everyone wants to know: is it true? Or is this just the latest round of AGI marketing?
Having dug through the benchmark tables, the independent testing, and what early users are actually reporting, the honest answer is: it’s complicated, and the complicated part is exactly where the story gets interesting.
The Numbers That Are Genuinely Hard to Argue With
Some of Astra’s results aren’t up for debate. They’re measurable, reproducible, and frankly impressive:
- 97.6% on FrontierMath Tier 4 — a benchmark made up of unpublished, expert-level math problems
- 100% on ExploitBench — a test of finding and chaining together security exploits
- 95.9% on BenchCAD — covering CAD-style reasoning and 3D/spatial tasks
- First OpenAI model to cross the company’s own “Critical” threshold for cybersecurity capability, meaning it can independently discover previously unknown vulnerabilities with minimal human steering
Those aren’t soft numbers. They’re the kind of scores that would have sounded far-fetched a year ago.
The Score Everyone’s Repeating — And Why It Deserves a Second Look
The number that spread fastest online is 98.6% (OpenAI’s own table lists it closer to 99.9%) on ARC-AGI-3, a benchmark specifically built to resist memorization. Unlike older AI tests, ARC-AGI-3 doesn’t use static puzzles — it drops models into hundreds of interactive game environments with zero instructions and no stated goal. When it launched in March 2026, the best AI model in the world scored under 1%. Humans, by contrast, solve nearly all of it without training.
So a model hitting 98%+ sounds like the finish line for AGI, right?
Here’s the catch: ARC Prize (the organization behind the benchmark) actually tested Astra two different ways.
- Under the standard harness — the same conditions every model on the public leaderboard is tested under — Astra scored 62.7%.
- The much higher number came from a setup that let Astra retain its own working notes across moves within a single game session, an advantage other models on the leaderboard weren’t given.
62.7% is still the best score any model has ever posted on ARC-AGI-3. That’s a real, legitimate achievement. But it’s a very different number from the one that went viral, and ARC Prize itself has been explicit that even a clean win here wouldn’t prove AGI, since these are still closed, rule-bound game worlds — not the messy, open-ended real world.
Read More Blog : Mini AI Desktop PCs: The New Work-From-Home Essential?
It Doesn’t Win Everything — And That Matters
If Astra were unambiguously AGI, you’d expect it to sweep the field. It doesn’t.
On Humanity’s Last Exam with tools, Astra scores 57.2%, while Claude Fable 5.1 scores 65.0% — a clear lead for Anthropic’s model. Astra also trails Claude Fable 5.1, Claude Opus 5, and Claude Fable 5 on the Artificial Analysis Intelligence Index, according to OpenAI’s own published comparison table.
The practical read from early testers seems to be a split personality: Astra looks like the stronger pick for computer operation, 3D/Blender-to-Unreal workflows, scientific research tasks, and cybersecurity, while Claude’s models still hold an edge in clean, mergeable code and frontend design judgment. One is shaping up to be the “operator,” the other the “craftsperson.”
What Real Users Are Saying (Not Just the Benchmarks)
Benchmark scores are one thing. Hands-on reports are another, and they’re arguably more telling.
Early testers with access ahead of launch have described major leaps in computer use, 3D modeling, data analysis, and running multiple coordinated AI agents at once — while also flagging real rough edges and quirks. One tester called their early access experience “a taste of AGI,” which is a notably different, more grounded claim than OpenAI’s own marketing line.
Specific workflow reports have also surfaced: modeling a house in Blender and turning it into a walkable Unreal Engine 5 scene, generating playable game prototypes from a single prompt, and running full scientific data-analysis loops — opening data, running the analysis, generating plots, and writing up what changed, largely unsupervised.
The Question People Actually Care About: Jobs
Benchmarks are interesting. Job security is what people are actually anxious about.
Astra is being pitched as a model that finishes tasks rather than just assisting with them — filling out forms, editing legal documents, booking reservations, and completing full workflows without a human at the keyboard. OpenAI’s own leadership has framed this as a coming “boom of entrepreneurship” rather than a wave of layoffs.
That optimism runs against a well-worn fear cycle. Back in 2025, warnings about AI wiping out a huge share of entry-level white-collar jobs within five years built into a real “jobs apocalypse” narrative — one that cooled off through early 2026 as actual layoff numbers failed to match the predictions. By this past spring, the share of executives expecting AI to meaningfully cut headcount had dropped sharply from earlier estimates.
Astra’s launch reignited that fear almost overnight, with lists of “at-risk” job categories circulating within a day of the announcement.
The honest middle ground: every previous “AI will take your job” warning was about a model that could draft or suggest — still requiring a human to actually execute. Astra is the first model explicitly built to close that gap and operate a browser or an application start to finish. That’s a real, technical step change. But OpenAI notably didn’t publish a score on GDPval, its own benchmark designed specifically to measure performance on real, paid occupational work — a conspicuous omission for a model being marketed as job-ready. Early testing shows strong but uneven, task-by-task results rather than the broad, reliable competence needed to fully hand over someone’s job.
So — Is It AGI or Not?
Here’s the fairest way to put it: GPT-6 Astra doesn’t prove AGI, but it makes dismissing the AGI conversation a lot harder than it used to be.
It has broad competence across math, science, cybersecurity, 3D/CAD work, and long, multi-step computer-use tasks. It uses tools. It runs agent swarms. It can operate software end to end rather than just talking about it. That’s genuinely AGI-ish behavior.
But AGI was never meant to be settled by a launch event, a single benchmark, or a viral demo video. It’s a threshold that keeps moving specifically because every time a system gets impressively capable, the definition shifts again. Astra might be the clearest example of that pattern yet — not a finish line, but the goalposts visibly moving in real time.
Frequently Asked Questions
Is GPT-6 Astra actually AGI?
Not by the clearest available evidence. It shows a genuine capability jump, but its most-cited benchmark score came from a non-standard testing setup, and it hasn’t demonstrated it beats humans at most economically valuable work — the more rigorous definition of AGI.
What is ARC-AGI-3, and why does it matter here?
It’s a benchmark built to resist memorization by testing models in interactive, unlabeled game environments rather than static puzzles. Astra’s viral 98%+ score came from a non-standard setup; under the same standard conditions every other model is tested with, it scored 62.7% — still a record, but a very different number.
Does GPT-6 Astra beat Claude’s models?
It depends on the task. Astra leads on 3D/CAD work, computer-use benchmarks, cybersecurity, and advanced math. Claude Fable 5.1 currently leads on Humanity’s Last Exam with tools and the Artificial Analysis Intelligence Index, and is reportedly still favored for clean, mergeable code.
Will GPT-6 Astra actually cause job losses?
It’s the first model built to complete full computer-based tasks rather than just assist with them, which makes older warnings more technically plausible than before. Whether that translates into real job losses will be determined by real-world deployment over months, not by a launch demo.