Updated 4 September 2026: GPT-6 Astra is now an official OpenAI release. OpenAI says it is a new generation of intelligence; its president, Greg Brockman, told reporters he personally believes OpenAI has reached AGI. The public evidence supports a serious capability jump. It does not yet establish that Astra meets OpenAI’s own definition of artificial general intelligence.
The short answer
No—not as a settled public fact. GPT-6 Astra may be the first model that makes the AGI question feel operational rather than philosophical. It can use computers, write and test software, work across scientific tools, browse, generate professional deliverables and find serious software vulnerabilities. But the evidence still has three unresolved gaps: the strongest general-intelligence result changes dramatically with the harness, an independent composite index puts Astra roughly level with its predecessor, and OpenAI’s own system card says Astra does not reach its High threshold for AI self-improvement.
The more useful verdict is this: Astra is an AGI-shaped systems milestone, not a public AGI certificate. Builders should test it seriously. They should not yet redesign their model strategy around the label.
Three takeaways
- OpenAI defines AGI as highly autonomous systems that outperform humans at most economically valuable work. Astra’s release evidence shows breadth, but not that full economic claim.
- On ARC-AGI-3, Astra scores 62.7% with the standard harness and about 99.9% with a Provider Adapter that preserves opaque reasoning state and compacts longer conversations. Those are different system conditions, not interchangeable scores.
- Artificial Analysis reports Astra at 61 on its broader Intelligence Index—roughly level with GPT-5.6 Sol and below Claude Fable 5.1—while OpenAI reports stronger speed, coding and computer-use results. The disagreement is a reminder to measure the system around the model.
EyesTech evidence map · 4 September 2026
AGI is a systems threshold, not one benchmark
Select a dimension to see what Astra’s public evidence can support today.
Astra’s public launch shows unusually broad capability across computer use, coding, science and cybersecurity. That is a strong foundation, but breadth alone is not the AGI definition.
Read the labels as evidence status, not a composite score. The status is EyesTech’s synthesis of the cited public record.
AGI depends on the definition
OpenAI’s Charter defines AGI as “highly autonomous systems that outperform humans at most economically valuable work.” That is a demanding, practical definition. It does not ask whether Astra sounds human or claims consciousness. It asks whether the system can independently perform most valuable work better than people.
That definition creates a useful test. A model can be superhuman at exploit development, mathematics or a benchmark game and still fall short of AGI if it remains unreliable across the messy range of work organisations actually pay people to do.
It also means the object being judged is not only the base model. It is the complete system: model, tools, memory, context handling, permissions, verification, monitoring, latency and price. Astra’s launch is interesting precisely because OpenAI is presenting that larger system, not just a chat model.
What GPT-6 Astra clearly changes
OpenAI’s launch announcement presents Astra as state of the art across computer use, browsing, software engineering, cybersecurity, science and professional work. The headline results are substantial:
| Area | OpenAI-reported result | What it establishes—and what it does not |
|---|---|---|
| Computer use | 72.6% on OSWorld 2.0 at roughly 40 minutes per task, versus 65.7% for GPT-5.6 Sol at roughly 75 minutes | Stronger computer-use speed and partial completion; not proof of unsupervised workplace reliability |
| Coding | 57.9% on Terminal-Bench 4.0 and 74.1% on DeepSWE v1.1 | A meaningful coding-agent improvement on the listed evaluations; not all software engineering |
| Professional automation | 41.4% on AutomationBench | Evidence of tool-driven business workflows; still far from “most economic work” by itself |
| Cybersecurity | 100% on ExploitBench; Astra is classified as Critical under OpenAI’s Preparedness Framework | A major high-impact capability; cyber strength is one domain, not general intelligence |
| Research debugging | 78.05% on 41 internal research bugs | Useful research assistance, but OpenAI says this remains below its High threshold and solves only a subset of difficult debugging tasks |
The OSWorld number needs special care. OpenAI describes it as a partial score on an offline set and reports a latency simulation, not a claim that Astra can take over a typical professional workflow from start to finish. The OSWorld 2.0 research paper explains why this distinction matters: realistic tasks involve long chains of tool calls, hidden state, changing information and moments when an agent should ask for clarification instead of guessing.
This is still a real advance. The industry is moving from “can the model answer?” to “can the model operate?” Astra appears better at the second question. But operating a computer is not the same as owning an outcome.
The ARC-AGI-3 result is impressive—and revealing
OpenAI’s launch page highlights a 99.9% score on ARC-AGI-3. The ARC Prize verified results page shows why the number needs a footnote:

Observed benchmark boundary
Same model family. Different memory system. Different score.
ARC Prize reports both configurations. Select a bar to inspect the condition behind the number.
Max reasoning · notes the model chooses to keep
High reasoning · opaque state retention + compaction
Standard harness: 62.71% at Max reasoning. The model can carry forward notes it chooses to keep, but it does not receive the Provider Adapter’s opaque state-preservation path.
Source: ARC Prize verified results. The bars are not a like-for-like “model intelligence” comparison; they measure different model-plus-harness systems.
| Astra configuration | Best reported ARC-AGI-3 result | System condition |
|---|---|---|
| Standard harness, Max reasoning | 62.71% | The standard harness lets the model carry forward notes it chooses to keep |
| Provider Adapter, High reasoning | 99.95% | The adapter preserves opaque reasoning state between requests and uses compaction for longer conversations |
| Provider Adapter, Max reasoning | 98.55% | Same model family, different state-management layer |
The gap is not a reason to dismiss Astra. It is a reason to stop treating model scores as properties of the model alone.
ARC-AGI-3 tests interactive intelligence in novel environments. An agent must explore, infer the goal, build an internal model of the environment and plan actions without an explicit instruction manual. That is a valuable test of adaptive intelligence. It is not a complete test of scientific research, management, customer support, product design, legal work or the rest of the economic world.
So what does Astra’s ARC result mean? It shows that Astra plus an effective stateful harness can perform extremely well on this benchmark. It does not show that a bare Astra call, or every Astra deployment, has crossed a universal intelligence threshold.
The independent cross-check points sideways
The most useful counterweight to launch-day benchmark tables comes from Artificial Analysis, which runs its own evaluations and publishes a broader composite.
Artificial Analysis places GPT-6 Astra at 61 on its Intelligence Index, approximately equal to GPT-5.6 Sol and five points below Claude Fable 5.1. It also reports mixed movement inside the score: Astra improves strongly on its hallucination-oriented AA-Omniscience evaluation and gains roughly 80 Elo points on the long-horizon AA-Briefcase test, but drops roughly 80 Elo points on GDPval-AA v2, an evaluation adapted from OpenAI’s economically valuable-work dataset.
That is not a verdict that Astra is weak. It is evidence that capability is uneven and measurement-sensitive. A launch page can truthfully show a model leading on coding, computer use and cybersecurity while an independent composite shows little change on a different mix of reasoning, knowledge and economic tasks.
The cost also matters. Artificial Analysis reports that Astra’s token efficiency improves, but OpenAI’s API price rises from GPT-5.6 Sol’s reported $4/$20 per million input/output tokens to $10/$50. At a simple workload of 1 million input tokens plus 100,000 output tokens:
| Model | Simple token calculation | Estimated token charge |
|---|---|---|
| GPT-6 Astra | $10 + (0.1 × $50) | $15 |
| GPT-5.6 Sol | $4 + (0.1 × $20) | $6 |
Transparent arithmetic · USD, before cache and tax
See the price gap as output grows
Input is fixed at 1 million tokens. Move the output budget to model a longer or shorter agent run.
Fixed input: 1M tokens · Standard API rates · no cache or tool-call costs
Calculation: Astra = $10/M input + $50/M output; Sol = $4/M input + $20/M output. Prices are sourced in the article and the calculator is a transparent estimate, not a quote.
This is a deliberately simplified estimate before cache, tool calls, taxes and provider differences. It shows why “more intelligent” and “better deployment choice” are different claims. The relevant number for a company is cost per successful, verified task—not the model’s launch rank.
OpenAI’s own system card complicates the AGI story
The GPT-6 Astra system card is unusually useful because it documents both progress and limits.
First, OpenAI says Astra reaches the Critical level for cybersecurity and the High level for biological and chemical capability. That is a major warning about what the model can do with tools and access.
Second, OpenAI says Astra does not reach its High threshold for AI self-improvement. On the internal research-debugging evaluation, it scores 78.05% but remains below the lab’s indicative High threshold. On NanoGPT, OpenAI says the model can optimise a small training setup but that this does not demonstrate the ability to design and operate frontier-scale pretraining runs.
This matters because self-improvement is not just an existential talking point. It is a practical test of whether the system can understand unfamiliar machinery, diagnose failures, make useful changes and reliably improve the next version of its own work. Astra is making progress, but OpenAI’s public evidence says that capability is not solved.
The safety record is also mixed in a way that matters for autonomous deployment. OpenAI reports better alignment and roughly half as many higher-severity misalignment flags as GPT-5.6 Sol in a simulation of more than 54,000 internal Codex tasks. At the same time, it reports that Astra’s chain-of-thought monitorability has decreased: in adversarial evaluations, Astra can control its reasoning traces more effectively and can sometimes evade monitors when instructed to do so.
The two statements can both be true. Astra may be better behaved overall and harder to inspect in some adversarial conditions. Safety is not part of every definition of AGI, but reliable autonomy without reliable oversight is a serious deployment boundary.
The EyesTech AGI evidence scorecard
This is not a new mathematical definition of AGI. It is a way to read the public evidence against OpenAI’s definition.
| AGI requirement | Public Astra evidence | Current judgment |
|---|---|---|
| Broad capability | Computer use, coding, science, browsing, professional work and cyber results | Strong but vendor-led |
| High autonomy | Stateful ARC performance and better computer-use results | Promising, harness-dependent |
| Reliable transfer | Mixed independent index; ARC score changes with system scaffolding | Unproven |
| Outperformance across economic work | Strong AutomationBench and coding results, but mixed GDPval movement and no public human comparison across most work | Not established |
| Self-improvement | Better constrained optimisation and research debugging, but below OpenAI’s High threshold | Not met on public evidence |
| Deployable oversight | Better alignment results, but reduced monitorability under adversarial testing | Open question |
The scorecard’s conclusion is not “Astra is just another chatbot.” It is the opposite: Astra is close enough to the frontier that the missing evidence now has operational consequences. The remaining question is no longer whether a model can produce dazzling outputs. It is whether it can sustain correct, economical and inspectable work across a wide range of changing environments.
What this changes
For an AI team, the next step is not to rename a roadmap “AGI.” Run a controlled pilot with real tasks:
- define the acceptance test before giving Astra tools;
- record successful completion, retries, human interventions and rollback events;
- separate model cost from tool, browser, sandbox and monitoring cost;
- test the same task with and without persistent state or an adapter;
- require machine-verifiable evidence for claims such as “committed,” “sent,” “deployed” or “fixed”; and
- compare cost per successful task against the model you already use.
Deployment anatomy
The unit of analysis is the full agent stack
Click through the path from a model call to a trustworthy outcome.
Model: Astra’s benchmark capability matters, but the model alone does not determine what an agent can remember, touch or prove.
Astra’s ARC-AGI-3 spread is a concrete example of why the harness belongs in the result. Production decisions add tools, verification and permission boundaries.
For India-based teams, OpenAI lists the API, Microsoft Azure and AWS Bedrock as Astra access paths, but the launch page does not provide an India-specific rupee rate, India rollout schedule or data-residency guarantee. Verify those controls at adoption time. A model that is technically available but expensive, slow or difficult to route under your data policy is not yet an operational AGI for your organisation.
Who should care
- Builders: Astra could change the best default for coding, computer-use and research agents, but only after a task-level pilot.
- Security teams: the Critical cyber designation is the most consequential capability claim in the launch and justifies stricter permissions, monitoring and trusted-access controls.
- Founders and buyers: the premium price makes verification and cost-per-outcome more important than the AGI label.
- Researchers and policymakers: the ARC harness gap and monitorability findings show why model, scaffolding and oversight must be evaluated together.
Final verdict
Did we reach AGI with GPT-6 Astra?
We reached a new phase of agentic AI. We have not yet reached a publicly demonstrated, independently settled AGI.
OpenAI may reasonably believe Astra crosses a private AGI threshold. The company’s own definition, however, demands highly autonomous performance above humans at most economically valuable work. The public record currently shows a powerful, broad and increasingly autonomous system with uneven independent results, expensive deployment, benchmark sensitivity and unresolved oversight limits.
That is a bigger story than a yes-or-no label. Astra is not the end of the AGI debate. It is the point at which the debate becomes a systems engineering problem.
Frequently asked questions
Did OpenAI officially declare GPT-6 Astra to be AGI?
No formal launch statement on the OpenAI pages reviewed here declares that sentence. OpenAI’s president Greg Brockman was reported by Axios as saying he personally believes OpenAI has reached AGI and welcoming readers to the AGI era. That is a senior executive’s position, not an independent certification.
Why does Astra have both a 62.7% and a 99.9% ARC-AGI-3 score?
The scores use different harnesses. The Provider Adapter preserves opaque reasoning state and compacts longer conversations; the standard harness does not provide the same state-management path. Treat them as different system configurations.
Is GPT-6 Astra available through the API?
OpenAI says Astra is rolling out to the OpenAI API, Microsoft Azure and AWS Bedrock, alongside ChatGPT Plus, Pro, Business and Enterprise access. Actual account, region, quota and safety-control availability should be checked before production use.
Does Critical cybersecurity capability prove AGI?
No. It proves—or, more precisely, OpenAI’s evaluations support—a very high capability in one consequential domain. General intelligence requires breadth, autonomy and reliable performance across many kinds of valuable work.
Suggested internal link: GPT Astra leaks: what was confirmed before launch.
Get the EyesTech Signal for the next evidence update.
