Google announced Gemini 4 Argon on September 30, with introductory API pricing of $2 per million input tokens and $10 per million output tokens. Initial access is restricted to trusted cyber defenders. Its benchmark results support a serious contender for professional work, while leaving substantial questions about deployment reliability. Google’s announcement

For readers deciding whether to subscribe, integrate or switch models, three details matter: wider availability is still pending, the introductory rates will double, and each benchmark score needs its own interpretation.

Who can access Gemini 4 Argon?

The initial rollout runs through Google’s Fairwind Program. Google identifies paid API customers and Google AI Ultra subscribers as the starting audiences for its subsequent broader release, without giving a firm public date. The announcement does not establish immediate access for Ultra subscribers or a timetable for Google AI Pro. Google’s rollout statement

Buying a subscription on that basis means paying before the promised capability is available to your account. Check actual eligibility and the model selector before treating Argon as a reason to upgrade.

Existing Gemini availability is a separate question. EyesTech’s Gemini 3.8 Flash access guide covers that model’s API and Antigravity routes; those details do not establish Argon access.

Pricing doubles after the introductory period

Google gives two sets of token rates and a 95% discount on cached input. Its launch footnote does not specify when the introductory period ends. Google’s pricing announcement and footnote

Usage, per million tokensIntroductory rateAfter introductory pricing
Fresh input$2$4
Output$10$20

At the introductory input rate, the stated cache discount implies $0.10 per million cached input tokens. That is arithmetic, subject to the eventual service’s caching rules and any separate charges.

Consider a job consuming 100,000 fresh input tokens and 20,000 billed output tokens:

  • Input: 0.1 × $2 = $0.20.
  • Output: 0.02 × $10 = $0.20.
  • Total: $0.40 at introductory rates; $0.80 afterward.

For 10,000 identical jobs, that becomes $4,000 versus $8,000 before tools, retries or other charges. These are worked examples, not measured Argon workloads.

Caching helps most when repeated input dominates spending. If 90,000 of those input tokens qualify for the discount and 10,000 remain fresh, introductory input cost falls to $0.029. Output still costs $0.20, making the total $0.229—a reduction of about 43%, despite the 95% discount on the cached portion.

The useful purchasing metric is cost per accepted result. Divide spending across all attempts by the number of outputs that meet your acceptance criteria. Add review time when comparing the actual cost of completing work.

The benchmark profile changes with the task

Selected rows from Google’s published comparison show why one overall victory claim would be misleading:

BenchmarkArgonGPT-6 AstraClaude Opus 5.5
Vals Index68.9%63.1%67.0%
DeepSWE v1.177.9%74.1%74.2%
FrontierSWE v255.0%65.5%62.3%
Terminal-Bench 4.057.4%58.2%66.4%
Harvey’s Legal Agent Benchmark19.6%5.4%3.8%
OSWorld 2.0, offline subset, partial score69.2%72.6%Not reported

Argon’s coding results change direction across these tests. A team choosing a repository agent should examine which tasks resemble its own work before using the DeepSWE lead to justify migration.

The legal row raises another question: how much absolute performance is sufficient? A large lead over competing systems can coexist with a low score. The result needs the benchmark’s grading definition before anyone translates it into a claim about dependable legal work.

Likewise, a missing comparator is an evidence gap. It supplies no basis for assigning a zero or declaring a win.

Google’s full Gemini 4 Argon benchmark comparison

Google’s complete launch comparison, including benchmarks where Argon trails. Source: Google DeepMind. Published results; EyesTech did not rerun these tests.

The methods explain what the scores measure

Google’s evaluation document describes Argon testing at the highest thinking settings unless otherwise noted. The table combines Google’s own runs, evaluator leaderboards and providers’ reported results.

Several qualifications affect interpretation:

  • Video inputs differ. LVBench uses one frame per second for Gemini, versus fixed budgets of 800 frames for Astra, 300 for Fable 5.1 and 600 for Opus 5.5 because of API limits.
  • Verification time differs. Argon’s Terminal-Bench Science run uses a sixfold verifier timeout to address verification timeouts.
  • Computer-use grading differs. OSWorld reports partial credit on the offline subset; Argon’s result is the maximum across three runs. Anthropic’s combined online/offline results are omitted.

These conditions are documented. They define the comparison being made. Unequal video sampling can affect how much evidence a system receives; a changed verifier timeout can affect which completed work gets scored. Neither establishes misconduct, and neither should disappear from a capability verdict.

A benchmark also evaluates the surrounding agent setup: tools, context handling, budgets and verification. Its outcome cannot automatically be attributed to model weights alone.

Google’s DeepSWE v1.1 comparison chart

DeepSWE v1.1 results as presented by Google. The methods document identifies different sources for the comparison runs.

Independent evaluation adds evidence, with its own limits

Vals AI’s Argon profile reports 68.90% on its Vals Index, ranking first among 41 evaluated models. Its wider results also place Claude Sonnet 5.5 above Argon on Vibe Code Bench: 92.39% versus 91.91%.

That is useful evidence beyond the selected competitors in Google’s launch table. It shows how the apparent winner can change when another model enters the comparison. A narrow numerical gap still needs uncertainty and task relevance before it becomes a purchasing decision.

The meaning of “accuracy” also changes by benchmark. Vals’ Finance Agent v2 methodology uses weighted partial credit, with critical checks that can invalidate an answer. It separately reports an All-Pass metric requiring every check to succeed.

Consequently, a 65.4% primary finance score cannot be read as 65.4% of jobs completed perfectly. For a workflow where every material number must be correct, strict acceptance criteria are closer to the business requirement than a partial-credit average.

Google’s Vals Finance Agent v2 comparison chart

Finance Agent v2 primary scores as shown in Google’s announcement. These should be read using the benchmark’s partial-credit definition.

A million-token output limit is a ceiling

Google announces an output ceiling of one million tokens. Vals lists a 262,144-token maximum for its evaluation configuration. These describe different settings; neither guarantees that every eventual access route will expose the same limit. Google’s announcement, Vals’ configuration

If a run consumed one million billed output tokens, output alone would cost $10 at introductory rates or $20 afterward. A large ceiling gives difficult tasks more room, but it also allows more spending before a result arrives.

Measure whether longer runs produce more accepted work. Token capacity by itself supplies no answer about correctness, speed or economic value.

What would justify switching?

Argon has earned a place in a professional-work evaluation. A deployment decision needs results on the work you actually intend to assign it.

Once access is available, freeze a representative task set and compare Argon with your current model. Keep inputs, permissions and acceptance criteria consistent. Record total tokens, elapsed time, failed attempts and human corrections. Set spending and time limits in advance, and rerun the cost calculation using the post-introductory rates.

The evidence that would strengthen the case is straightforward: more accepted results at an affordable total cost, reproduced across tasks and repeated runs on the version customers can actually use.

Until then, the published benchmarks support a promising candidate. Access, workload-specific reliability and the full bill remain the conditions that determine whether it deserves your production traffic.

Sources checked October 1, 2026. EyesTech has not performed hands-on Argon testing.

Last Update: October 1, 2026