Updated 4 September 2026: GPT-6 Astra is now an official OpenAI release. OpenAI says it is a new generation of intelligence; its president, Greg Brockman, told reporters he personally believes OpenAI has reached AGI. The public evidence supports a serious capability jump. It does not yet establish that Astra meets OpenAI’s own definition of artificial general intelligence.

The short answer

No—not as a settled public fact. GPT-6 Astra may be the first model that makes the AGI question feel operational rather than philosophical. It can use computers, write and test software, work across scientific tools, browse, generate professional deliverables and find serious software vulnerabilities. But the evidence still has three unresolved gaps: the strongest general-intelligence result changes dramatically with the harness, an independent composite index puts Astra roughly level with its predecessor, and OpenAI’s own system card says Astra does not reach its High threshold for AI self-improvement.

The more useful verdict is this: Astra is an AGI-shaped systems milestone, not a public AGI certificate. Builders should test it seriously. They should not yet redesign their model strategy around the label.

Three takeaways

  • OpenAI defines AGI as highly autonomous systems that outperform humans at most economically valuable work. Astra’s release evidence shows breadth, but not that full economic claim.
  • On ARC-AGI-3, Astra scores 62.7% with the standard harness and about 99.9% with a Provider Adapter that preserves opaque reasoning state and compacts longer conversations. Those are different system conditions, not interchangeable scores.
  • Artificial Analysis reports Astra at 61 on its broader Intelligence Index—roughly level with GPT-5.6 Sol and below Claude Fable 5.1—while OpenAI reports stronger speed, coding and computer-use results. The disagreement is a reminder to measure the system around the model.

EyesTech evidence map · 4 September 2026

AGI is a systems threshold, not one benchmark

Select a dimension to see what Astra’s public evidence can support today.





Astra’s public launch shows unusually broad capability across computer use, coding, science and cybersecurity. That is a strong foundation, but breadth alone is not the AGI definition.

Read the labels as evidence status, not a composite score. The status is EyesTech’s synthesis of the cited public record.

AGI depends on the definition

OpenAI’s Charter defines AGI as “highly autonomous systems that outperform humans at most economically valuable work.” That is a demanding, practical definition. It does not ask whether Astra sounds human or claims consciousness. It asks whether the system can independently perform most valuable work better than people.

That definition creates a useful test. A model can be superhuman at exploit development, mathematics or a benchmark game and still fall short of AGI if it remains unreliable across the messy range of work organisations actually pay people to do.

It also means the object being judged is not only the base model. It is the complete system: model, tools, memory, context handling, permissions, verification, monitoring, latency and price. Astra’s launch is interesting precisely because OpenAI is presenting that larger system, not just a chat model.

What GPT-6 Astra clearly changes

OpenAI’s launch announcement presents Astra as state of the art across computer use, browsing, software engineering, cybersecurity, science and professional work. The headline results are substantial:

AreaOpenAI-reported resultWhat it establishes—and what it does not
Computer use72.6% on OSWorld 2.0 at roughly 40 minutes per task, versus 65.7% for GPT-5.6 Sol at roughly 75 minutesStronger computer-use speed and partial completion; not proof of unsupervised workplace reliability
Coding57.9% on Terminal-Bench 4.0 and 74.1% on DeepSWE v1.1A meaningful coding-agent improvement on the listed evaluations; not all software engineering
Professional automation41.4% on AutomationBenchEvidence of tool-driven business workflows; still far from “most economic work” by itself
Cybersecurity100% on ExploitBench; Astra is classified as Critical under OpenAI’s Preparedness FrameworkA major high-impact capability; cyber strength is one domain, not general intelligence
Research debugging78.05% on 41 internal research bugsUseful research assistance, but OpenAI says this remains below its High threshold and solves only a subset of difficult debugging tasks

The OSWorld number needs special care. OpenAI describes it as a partial score on an offline set and reports a latency simulation, not a claim that Astra can take over a typical professional workflow from start to finish. The OSWorld 2.0 research paper explains why this distinction matters: realistic tasks involve long chains of tool calls, hidden state, changing information and moments when an agent should ask for clarification instead of guessing.

This is still a real advance. The industry is moving from “can the model answer?” to “can the model operate?” Astra appears better at the second question. But operating a computer is not the same as owning an outcome.

The ARC-AGI-3 result is impressive—and revealing

OpenAI’s launch page highlights a 99.9% score on ARC-AGI-3. The ARC Prize verified results page shows why the number needs a footnote:

Illustration of one AI model branching into two memory and harness paths with different benchmark outcomes
The same model can produce different benchmark outcomes when state retention and context management change.

Observed benchmark boundary

Same model family. Different memory system. Different score.

ARC Prize reports both configurations. Select a bar to inspect the condition behind the number.


Standard harness: 62.71% at Max reasoning. The model can carry forward notes it chooses to keep, but it does not receive the Provider Adapter’s opaque state-preservation path.

Source: ARC Prize verified results. The bars are not a like-for-like “model intelligence” comparison; they measure different model-plus-harness systems.

Astra configurationBest reported ARC-AGI-3 resultSystem condition
Standard harness, Max reasoning62.71%The standard harness lets the model carry forward notes it chooses to keep
Provider Adapter, High reasoning99.95%The adapter preserves opaque reasoning state between requests and uses compaction for longer conversations
Provider Adapter, Max reasoning98.55%Same model family, different state-management layer

The gap is not a reason to dismiss Astra. It is a reason to stop treating model scores as properties of the model alone.

ARC-AGI-3 tests interactive intelligence in novel environments. An agent must explore, infer the goal, build an internal model of the environment and plan actions without an explicit instruction manual. That is a valuable test of adaptive intelligence. It is not a complete test of scientific research, management, customer support, product design, legal work or the rest of the economic world.

So what does Astra’s ARC result mean? It shows that Astra plus an effective stateful harness can perform extremely well on this benchmark. It does not show that a bare Astra call, or every Astra deployment, has crossed a universal intelligence threshold.

The independent cross-check points sideways

The most useful counterweight to launch-day benchmark tables comes from Artificial Analysis, which runs its own evaluations and publishes a broader composite.

Artificial Analysis places GPT-6 Astra at 61 on its Intelligence Index, approximately equal to GPT-5.6 Sol and five points below Claude Fable 5.1. It also reports mixed movement inside the score: Astra improves strongly on its hallucination-oriented AA-Omniscience evaluation and gains roughly 80 Elo points on the long-horizon AA-Briefcase test, but drops roughly 80 Elo points on GDPval-AA v2, an evaluation adapted from OpenAI’s economically valuable-work dataset.

That is not a verdict that Astra is weak. It is evidence that capability is uneven and measurement-sensitive. A launch page can truthfully show a model leading on coding, computer use and cybersecurity while an independent composite shows little change on a different mix of reasoning, knowledge and economic tasks.

The cost also matters. Artificial Analysis reports that Astra’s token efficiency improves, but OpenAI’s API price rises from GPT-5.6 Sol’s reported $4/$20 per million input/output tokens to $10/$50. At a simple workload of 1 million input tokens plus 100,000 output tokens:

ModelSimple token calculationEstimated token charge
GPT-6 Astra$10 + (0.1 × $50)$15
GPT-5.6 Sol$4 + (0.1 × $20)$6

Transparent arithmetic · USD, before cache and tax

See the price gap as output grows

Input is fixed at 1 million tokens. Move the output budget to model a longer or shorter agent run.


Fixed input: 1M tokens · Standard API rates · no cache or tool-call costs

GPT-6 Astra$15.00
GPT-5.6 Sol$6.00
Premium2.50×

Calculation: Astra = $10/M input + $50/M output; Sol = $4/M input + $20/M output. Prices are sourced in the article and the calculator is a transparent estimate, not a quote.

This is a deliberately simplified estimate before cache, tool calls, taxes and provider differences. It shows why “more intelligent” and “better deployment choice” are different claims. The relevant number for a company is cost per successful, verified task—not the model’s launch rank.

OpenAI’s own system card complicates the AGI story

The GPT-6 Astra system card is unusually useful because it documents both progress and limits.

First, OpenAI says Astra reaches the Critical level for cybersecurity and the High level for biological and chemical capability. That is a major warning about what the model can do with tools and access.

Second, OpenAI says Astra does not reach its High threshold for AI self-improvement. On the internal research-debugging evaluation, it scores 78.05% but remains below the lab’s indicative High threshold. On NanoGPT, OpenAI says the model can optimise a small training setup but that this does not demonstrate the ability to design and operate frontier-scale pretraining runs.

This matters because self-improvement is not just an existential talking point. It is a practical test of whether the system can understand unfamiliar machinery, diagnose failures, make useful changes and reliably improve the next version of its own work. Astra is making progress, but OpenAI’s public evidence says that capability is not solved.

The safety record is also mixed in a way that matters for autonomous deployment. OpenAI reports better alignment and roughly half as many higher-severity misalignment flags as GPT-5.6 Sol in a simulation of more than 54,000 internal Codex tasks. At the same time, it reports that Astra’s chain-of-thought monitorability has decreased: in adversarial evaluations, Astra can control its reasoning traces more effectively and can sometimes evade monitors when instructed to do so.

The two statements can both be true. Astra may be better behaved overall and harder to inspect in some adversarial conditions. Safety is not part of every definition of AGI, but reliable autonomy without reliable oversight is a serious deployment boundary.

The EyesTech AGI evidence scorecard

This is not a new mathematical definition of AGI. It is a way to read the public evidence against OpenAI’s definition.

AGI requirementPublic Astra evidenceCurrent judgment
Broad capabilityComputer use, coding, science, browsing, professional work and cyber resultsStrong but vendor-led
High autonomyStateful ARC performance and better computer-use resultsPromising, harness-dependent
Reliable transferMixed independent index; ARC score changes with system scaffoldingUnproven
Outperformance across economic workStrong AutomationBench and coding results, but mixed GDPval movement and no public human comparison across most workNot established
Self-improvementBetter constrained optimisation and research debugging, but below OpenAI’s High thresholdNot met on public evidence
Deployable oversightBetter alignment results, but reduced monitorability under adversarial testingOpen question

The scorecard’s conclusion is not “Astra is just another chatbot.” It is the opposite: Astra is close enough to the frontier that the missing evidence now has operational consequences. The remaining question is no longer whether a model can produce dazzling outputs. It is whether it can sustain correct, economical and inspectable work across a wide range of changing environments.

What this changes

For an AI team, the next step is not to rename a roadmap “AGI.” Run a controlled pilot with real tasks:

  • define the acceptance test before giving Astra tools;
  • record successful completion, retries, human interventions and rollback events;
  • separate model cost from tool, browser, sandbox and monitoring cost;
  • test the same task with and without persistent state or an adapter;
  • require machine-verifiable evidence for claims such as “committed,” “sent,” “deployed” or “fixed”; and
  • compare cost per successful task against the model you already use.

Deployment anatomy

The unit of analysis is the full agent stack

Click through the path from a model call to a trustworthy outcome.









Model: Astra’s benchmark capability matters, but the model alone does not determine what an agent can remember, touch or prove.

Astra’s ARC-AGI-3 spread is a concrete example of why the harness belongs in the result. Production decisions add tools, verification and permission boundaries.

For India-based teams, OpenAI lists the API, Microsoft Azure and AWS Bedrock as Astra access paths, but the launch page does not provide an India-specific rupee rate, India rollout schedule or data-residency guarantee. Verify those controls at adoption time. A model that is technically available but expensive, slow or difficult to route under your data policy is not yet an operational AGI for your organisation.

Who should care

  • Builders: Astra could change the best default for coding, computer-use and research agents, but only after a task-level pilot.
  • Security teams: the Critical cyber designation is the most consequential capability claim in the launch and justifies stricter permissions, monitoring and trusted-access controls.
  • Founders and buyers: the premium price makes verification and cost-per-outcome more important than the AGI label.
  • Researchers and policymakers: the ARC harness gap and monitorability findings show why model, scaffolding and oversight must be evaluated together.

Final verdict

Did we reach AGI with GPT-6 Astra?

We reached a new phase of agentic AI. We have not yet reached a publicly demonstrated, independently settled AGI.

OpenAI may reasonably believe Astra crosses a private AGI threshold. The company’s own definition, however, demands highly autonomous performance above humans at most economically valuable work. The public record currently shows a powerful, broad and increasingly autonomous system with uneven independent results, expensive deployment, benchmark sensitivity and unresolved oversight limits.

That is a bigger story than a yes-or-no label. Astra is not the end of the AGI debate. It is the point at which the debate becomes a systems engineering problem.

Frequently asked questions

Did OpenAI officially declare GPT-6 Astra to be AGI?

No formal launch statement on the OpenAI pages reviewed here declares that sentence. OpenAI’s president Greg Brockman was reported by Axios as saying he personally believes OpenAI has reached AGI and welcoming readers to the AGI era. That is a senior executive’s position, not an independent certification.

Why does Astra have both a 62.7% and a 99.9% ARC-AGI-3 score?

The scores use different harnesses. The Provider Adapter preserves opaque reasoning state and compacts longer conversations; the standard harness does not provide the same state-management path. Treat them as different system configurations.

Is GPT-6 Astra available through the API?

OpenAI says Astra is rolling out to the OpenAI API, Microsoft Azure and AWS Bedrock, alongside ChatGPT Plus, Pro, Business and Enterprise access. Actual account, region, quota and safety-control availability should be checked before production use.

Does Critical cybersecurity capability prove AGI?

No. It proves—or, more precisely, OpenAI’s evaluations support—a very high capability in one consequential domain. General intelligence requires breadth, autonomy and reliable performance across many kinds of valuable work.

Suggested internal link: GPT Astra leaks: what was confirmed before launch.

Get the EyesTech Signal for the next evidence update.

Categorized in:

A.I,

Last Update: September 4, 2026