GPT-6.1 Sol’s standard API rates are $2 per million ordinary input tokens, $0.10 for cached reads, $2.50 for cache writes and $10 for output. The low read price is only one component of a request bill. A reusable prefix must first be written, and savings depend on how often it is reused.

These rates were checked September 30. The worked examples below use standard processing and requests below the long-context surcharge threshold. They are calculations, not measured bills or latency benchmarks.

Ten uses change the arithmetic

Take a stable 100,000-token prefix used in ten requests. With caching disabled, its ordinary input cost is 100,000 × 10 × $2 ÷ 1,000,000 = $2.00.

If the first use writes the entire prefix and nine subsequent requests fully read it from cache, the write costs $0.25 and the reads cost $0.09. The prefix total is $0.34, an 83% reduction against that ordinary-input comparison. This assumes every later request gets the full hit, with no expiry or rewritten prefix.

For two uses, the equivalent comparison is $0.40 without caching and $0.26 for one write plus one read. For a prefix used only once, a $0.25 write costs more than the $0.20 ordinary input charge. Reuse is what pays for the initial premium.

Output can dominate the total

Now add 2,000 newly generated output tokens to each of the ten requests. At $10 per million, those outputs cost $0.20 in total. The simplified combined bill becomes $0.54 with the assumed cache behavior, versus $2.20 using ordinary input: a reduction of approximately 75.5%.

That example contains no dynamic input after the prefix, tool charges, retries or processing premiums. Add those before treating it as an application budget. If the output becomes much longer while the cached prefix stays the same, the read discount applies to a smaller share of the complete bill.

Misses change the break-even point

The perfect-hit example is useful, but a production prefix may be rewritten or expire. Keep the same ten requests and 100,000-token prefix. Assume the first request writes it; among the next nine, a fraction h gets a full read, while every miss writes the full prefix again. Ordinary-input comparison remains $2.00.

Under those assumptions, prefix cost is $0.25 + 9 × [$0.25 × (1 − h) + $0.01 × h] = $2.50 − $2.16h. Caching wins when h exceeds 0.50 ÷ 2.16, or about 23.15% of the nine later requests. This fraction describes expected or aggregated hit behavior; one ten-request sequence has an integer number of hits.

Ten-request example: first write, then variable full-prefix hits
Later hit fractionPrefix costVersus $2 input
0%$2.50$0.50 more
25%$1.96$0.04 less
50%$1.42$0.58 less
75%$0.88$1.12 less
100%$0.34$1.66 less

At zero hits, ten writes cost $2.50, above the $2.00 ordinary-input comparison. At a 25% later-request hit rate, the saving is only four cents before any other cost. The read price alone makes these two workloads look similar; the writes reveal the difference.

For a long-running workload in which each prefix token is either read at $0.10 or written at $2.50 per million, the analogous average-cost comparison is 2.50 × (1 − h) + 0.10 × h < 2.00. Its break-even is about 20.83%. The ten-request threshold is higher because it includes a guaranteed initial write. Both examples assume full-prefix hits or full rewrites; partial hits and mixed uncached material require the actual token counts.

A cheap cached prefix can still be the wrong context

A stable rubric followed by a new document is a sensible boundary. But if a product’s rules change, keeping the old rubric solely to preserve hits can make answers incorrect. Version stable instructions when their meaning changes, then include the resulting writes in the budget. Cache eligibility and evidence freshness are different requirements.

EyesTech’s context-engineering guide discusses provenance, freshness and selecting information for the task. Those requirements still apply to cheap tokens. Retrieving a smaller, current document set can be more useful than repeatedly reading a large stale prefix, even if it lowers the apparent hit rate.

Compare cost per accepted result

For an application trial, record the complete bill across successful and failed attempts, then divide it by the number of results that meet a fixed acceptance criterion. Preserve the rejected outputs and the reason for rejection. That exposes a prompt that saves input cost but creates more retries or manual correction.

EyesTech’s earlier Claude Fable cost analysis uses the same completed-work perspective. Its model-specific numbers are not transferred into this calculation. The reusable method is to account for the whole task, including the attempts that did not produce usable work.

Separate cache reads, cache writes, ordinary input, output and other charges in the log. Add elapsed time and acceptance status. A rollout can then be judged on whether it lowers the observed cost of accepted work at the required quality, rather than whether its cached-token counter is large.

Choose the boundary around reusable content

OpenAI documents a 1,024-visible-token minimum prefix for GPT-5.6 and later. Explicit-only caching requires a selected breakpoint; without one, it creates no cache writes. Content after the last explicit breakpoint is ordinary input. The current setting provides at least 30 minutes of eligibility after the latest write or reuse; the entry may be retained longer.

A fixed grading rubric followed by a new document is a useful example. The rubric can form the reusable prefix; the document may be used only once. Writing both on every request can add cost for material that never receives a later read. Keep the content useful rather than padding a prompt solely to meet a minimum.

Measure writes and reads separately

The usage response distinguishes cached tokens and cache-write tokens. Log both alongside total input and output. A high read rate is encouraging, but it does not expose how much changing content was repeatedly written.

The model page also specifies different prices for long contexts, processing modes and regional processing. Recalculate using the settings actually used by the application. The practical optimization target is the cost of a successful task, including its failed attempts, rather than the cheapest token category on the price card.

Last Update: September 30, 2026