Anthropic released Claude Haiku 5.5 on October 7, 2026, introducing a 1,000,000-token context window, 128,000-token maximum output, and a dual-tier pricing model: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, surging fivefold to $0.50 input and $2.50 output beyond that threshold.
The 100,000-Token Pricing Discontinuity
On October 7, 2026, Anthropic announced Claude Haiku 5.5 under the model identifier claude-haiku-5-5. The release expands the Haiku tier’s context window from legacy spans to a full 1,000,000 tokens (1M tokens) with up to 128,000 tokens of generation per call. For developers who previously benchmarked Claude Opus 5.5 for complex architecture, Haiku 5.5 represents Anthropic’s bid for high-throughput, low-latency execution.
However, the headline rate of $0.10 per million input tokens (MTok) and $0.50 per million output tokens carries a hard structural boundary. The moment an incoming request reaches 100,001 tokens, Anthropic transitions billing to an elevated tier:
- Base Tier (≤ 100,000 tokens): $0.10 / MTok input, $0.50 / MTok output.
- Cliff Tier (> 100,000 tokens): $0.50 / MTok input, $2.50 / MTok output.
- Batch Processing: 50% discount on both tiers ($0.05 / $0.25 sub-100k; $0.25 / $1.25 post-100k).
This creates an immediate 400% cost inflation—an exact 5.0× multiplier across both prompt ingestion and token generation. For engineering teams deploying Haiku as an autonomous subagent orchestrator, this step-function change fundamentally alters multi-turn agent unit economics.
| Pricing Dimension | Sub-100k Context Window | Post-100k Context Window | Tier Jump Multiplier |
|---|---|---|---|
| Input Tokens (per MTok) | $0.10 | $0.50 | 5.0× (+400%) |
| Output Tokens (per MTok) | $0.50 | $2.50 | 5.0× (+400%) |
| Prompt Cache Write (5m TTL) | $0.125 | $0.625 | 5.0× (+400%) |
| Prompt Cache Read (per MTok) | $0.010 | $0.050 | 5.0× (+400%) |
| Batch API Input (per MTok) | $0.050 | $0.250 | 5.0× (+400%) |
| Batch API Output (per MTok) | $0.250 | $1.250 | 5.0× (+400%) |
Agentic Loop Telemetry: OSWorld 2.1 and Terminal Benchmarks
The commercial relevance of Haiku 5.5 stems from its measurable synthetic capability uplift. In developer discussions on Hacker News, infrastructure engineers emphasized that while frontier tiers like Claude Fable 5.1 remain the standard for high-ambiguity system design, Haiku 5.5 is optimized for high-frequency deterministic tasks: terminal execution, tool-call dispatch, and intermediate log compaction.
On Anthropic’s published evaluations, Haiku 5.5 demonstrates a substantial performance leap over its predecessor, Haiku 4.5:
| Benchmark Evaluation | Claude Haiku 4.5 | Claude Haiku 5.5 | Operational Lift |
|---|---|---|---|
| OSWorld 2.1 (Computer Use) | 15.7% | 72.4% | +56.7% (4.61×) |
| Terminal-Bench 4.0 | 0.0% | 39.2% | +39.2% (New Capability) |
| GDPval-AA Elo (Artificial Analysis) | 735 Elo | 1,620 Elo | +885 Elo (+120.4%) |
Haiku 5.5 is also the first model in Anthropic’s compact tier to incorporate “Adaptive Thinking,” introducing a granular effort configuration parameter (low, medium, high, max). This allows developers to tune token deliberation time and cost on a per-request basis without changing endpoints. However, because thinking tokens count toward output billing, high-effort reasoning on prompts exceeding 100k tokens is billed at $2.50 / MTok rather than $0.50 / MTok.
The Worked Math: A 40-Step Repository Subagent
To evaluate the financial impact of the 100k threshold, consider a standard developer workflow: an autonomous triage subagent navigating a medium-sized enterprise repository to resolve an integration bug.
Modeled Scenario Assumptions
- Base Context: 70,000 tokens of system prompt, repository AST, and dependency manifests.
- Per-Turn Accumulation: Each turn appends 1,500 tokens of tool definitions, bash stdout, and git diffs.
- Execution Length: 40 iterative turns.
- Output per Turn: 500 generated tokens (actions, JSON payloads, or brief verification traces).
- Prompt Caching: Cold write on Turn 1; 100% cache hit on the static base context (70k tokens) across subsequent turns.
Under this trajectory, the agent crosses the 100,000-token pricing cliff on Turn 21 (70,000 base + 21 × 1,500 = 101,500 tokens). From that moment forward, every single turn incurs the elevated $0.50 / $2.50 rate card.
Turns 1–20 (Sub-100k Tier)
- • Cached Base Reads (19 × 70k @ $0.01): $0.0133
- • Dynamic Ingestion (315k @ $0.10): $0.0315
- • Initial Cache Write (70k @ $0.125): $0.0088
- • Output Generation (10k @ $0.50): $0.0050
- Subtotal: $0.0586
Turns 21–40 (Cliff Tier — >100k)
- • Cached Base Reads (20 × 70k @ $0.05): $0.0700
- • Dynamic Ingestion (765k @ $0.50): $0.3825
- • Output Generation (10k @ $2.50): $0.0250
- Subtotal: $0.4775
The total cost for the unmanaged 40-step run is $0.5361. Despite executing exactly half of the loop’s actions, Turns 21–40 generate 89.1% of the total task expenditure. The 5× surge in cache read rates ($0.01 to $0.05 per MTok) combines with uncached dynamic prompt ingestion at $0.50 per MTok to create rapid cost acceleration.
Context Compaction Economics: Why Pruning at 90k Saves 72%

To eliminate this unmanaged expenditure, the orchestrator can inject an automated context compaction pass at Turn 19, just before crossing the 100,000-token boundary.
- Compaction Trigger: At Turn 19 (98,500 total prompt tokens), the orchestrator triggers an inline summary prompt. The 28,500 tokens of accumulated tool outputs are condensed into an 8,000-token structured state machine snapshot.
- Compaction Execution Cost: Generating the 2,000-token state snapshot costs 2,000 tokens @ $0.50 / MTok = $0.0010.
- New Prompt Baseline: 70,000 static base + 8,000 state snapshot = 78,000 tokens.
- Remaining Execution: Turns 20 through 40 accumulate fresh diffs starting from 78,000 tokens, finishing Turn 40 at 108,000 tokens. The agent spends only 5 turns above the cliff instead of 20.
Task Cost Comparison: Unmanaged vs. Actively Compacted
Net Efficiency Delta: Spending $0.0010 on an inline summary pass eliminates $0.3880 in cliff surcharges, reducing total execution cost by 72.4% per successful task.
Architectural Guardrails for Anthropic API Consumers
To keep autonomous agent swarms operating inside the base rate card, engineering teams should incorporate three concrete patterns into their API gateway architecture:
1. Strict Prefix Invariant Segmentation
Anthropic’s prompt caching functions via sequential prefix matching. If dynamic tool returns or variable scratchpad tokens are inserted before static documentation, cache misses cascade and prompt blocks re-evaluate at full uncached rates. Structure requests with static system prompts and tool schemas at the head of the array, followed by cache breakpoints (ephemeral), before appending rolling history.
2. High-Watermark Eviction Thresholds
Establish an explicit eviction watermark at 90,000 tokens in the orchestration middleware. When usage.input_tokens reaches 90k, trigger an asynchronous compaction subagent to distill execution logs into a declarative state snapshot, evicting raw shell outputs before the subsequent turn crosses the 100k boundary.
3. Effort Throttling on Extended Contexts
Because thinking tokens count toward output billing, running Haiku 5.5 at max adaptive thinking on contexts exceeding 100k tokens incurs generation charges of $2.50 per million tokens. Cap the effort parameter at low or medium once context expands beyond 85k tokens, or route complex single-turn reasoning tasks to Sonnet with compressed context.
FinOps Verdict: The High-Volume Threshold
Claude Haiku 5.5 is not an unconditional price reduction across all workloads; it is a pricing fork. Below 100,000 tokens, it delivers 72.4% OSWorld automation at $0.10 per million input tokens—making it the most economical agent executor in Anthropic’s catalog. Above 100,000 tokens, its cost profile matches baseline rates of larger frontier models.
For systems architects and FinOps leads, the defining operational metric of a Haiku 5.5 deployment is not raw throughput—it is the percentage of requests maintained beneath the 100,000-token boundary.
