All pricing in this analysis was verified directly against provider documentation in September 2026: AWS EC2 Pricing (p5 instance family, us-east-1), CoreWeave Rate Card (on-demand 8×H100 HGX nodes), RunPod GPU Cloud Pricing (Secure Cloud tier, post-August 2026 price revision), and AWS Data Transfer Pricing (internet egress, standard regions). Historical H100 spot pricing sourced from SiliconData.com market reports and SemiAnalysis GPU cost tracking (2023–2026). TCO model methodology: one 8×H100 node, 720 hours per monthnth continuous operation, 20TB scratch storage, 15TB monthly outbound egress, standard enterprise support tiers. Next pricing review: September 21, 2026.

I track GPU rental pricing for a living. Every week I pull rate cards from six cloud providers, run them through a TCO model, and send a digest to engineering leads deciding where to put their next training run. And I want to be direct with you about what I’m seeing in September 2026: the H100 market has bifurcated. Neo-clouds collapsed from $8.50 per GPU-hour in 2023 down to $3.99 per GPU-hour on RunPod today — a 53% drop in three years. AWS’s p5.48xlarge list price sits at $55.04 per node-hour ($6.88 per GPU-hour) — and once you model egress taxes, FSx Lustre storage, NAT Gateway transit, and mandatory support tiers, the real all-in rate climbs past $9.85 per GPU-hour. That spread is not a pricing anomaly. It is a deliberate infrastructure tax, and most engineering teams are still paying it without questioning it.

Neo-cloud H100 on-demand pricing reached $3.99 per GPU-hour (RunPod Secure Cloud) in September 2026 — down 53% from the 2023 peak of $8.50–$10.00 per GPU-hour. CoreWeave on-demand is $6.16 per GPU-hour ($49.24/8-GPU node); with a 1-year reserved contract (up to 60% discount), effective rates reach approximately $2.46 per GPU-hour. AWS p5.48xlarge lists at $6.88 per GPU-hour ($55.04 per node) — but with egress ($0.09 per GB), FSx Lustre storage, NAT Gateway overhead, and enterprise support surcharges, all-in effective rates exceed $9.85 per GPU-hour. Key architectural trade-off: CoreWeave’s NVIDIA Quantum-2 InfiniBand delivers 50%+ MFU on large clusters vs. RunPod RoCEv2’s ~34% MFU, making provider choice a topology decision, not just a billing comparison.
- The 70% Price Collapse: Four Structural Catalysts
- September 2026 Verified Pricing: RunPod vs CoreWeave vs AWS
- The AWS Tax: Egress, FSx, NAT & Support Anatomy
- 30-Day TCO Model: One 8×H100 Node, Fully Loaded
- Interconnect Forensics: InfiniBand vs EFAv2 vs RoCEv2
- Spot Eviction Economics & the Checkpointing Penalty
- Break-Even Math: When to Commit vs Rent
- The Tri-Tier Procurement Architecture
- GPU MSA Negotiation: 4 Non-Negotiable Clauses
- FAQ
The 70% Price Collapse: Four Structural Catalysts
This is not a market correction. It is a structural regime change, and the four catalysts that drove it are still accelerating. Understanding them is not academic — it determines your negotiating leverage when signing the next cloud MSA.
Source: SiliconData.com, SemiAnalysis
Source: ValueAdd VC GPU tracker
Source: Presenc.ai market data
Source: RunPod pricing page, Sep 2026
Catalyst 1: The Blackwell Migration. NVIDIA’s B200 and GB200 NVL72 rack-scale systems began volume delivery in late 2025. OpenAI, Anthropic, Meta, and xAI — the tier-1 buyers who were paying $8+ per hour for H100 access — migrated their frontier training runs to Blackwell. That mass departure offloaded vast H100 & H200 cluster inventories into secondary rental pools overnight.
Catalyst 2: Inference Efficiency Compounding. In 2023, serving a 70B parameter model required two to four H100s in BF16 precision. By 2026, FP8 E4M3 quantization, AWQ W4A16, and speculative decoding halved or quartered VRAM and compute requirements for equivalent token throughput. Inference demand splintered across more providers and smaller instances. The H100 moat dissolved. For context on how quantization methods are reshaping inference economics, see our analysis of DeepSeek’s MLA architecture slashing KV cache by 93% — the same architectural thinking is compressing GPU-hour demand per token across all providers.
Catalyst 3: Neo-Cloud Debt Covenant Pressure. CoreWeave, Lambda, Crusoe, and RunPod raised tens of billions in asset-backed debt facilities collateralized by physical GPU inventories. Those debt covenants require continuous cash flow to service principal and interest. When uncontracted utilization drops below ~75%, neo-clouds are financially compelled to slash spot rates toward marginal operating cost — roughly $0.55–$0.75 per GPU-hour in electricity and colocation cooling. The headline says “cheap GPUs.” The mechanism is “debt service urgency.”
Catalyst 4: Secondary Market Saturation. Enterprises that over-provisioned private clusters in 2023–2024 began renting idle cycles on multi-tenant hosting platforms, creating an active secondary spot exchange. Hundreds of new GPU cloud providers entered the market between 2023 and 2026, breaking the supply monopoly previously held by hyperscalers. The market is eating itself clean. The full 2026 AI Inference & Hardware Economics statistics breakdown shows this saturation effect across all chip generations.
September 2026 Verified Pricing: RunPod vs CoreWeave vs AWS
These are verified rates pulled from provider documentation in September 2026. I’m flagging this explicitly because pricing from Q1 2025 dossiers circulates on Slack channels and in procurement decks months after rate adjustments. Use these figures for your current budget cycle.
The AWS Tax: Egress, FSx, NAT & Support Anatomy
I’ve reviewed dozens of AWS invoices from AI startups who came to us asking why their cloud bill was “twice what they budgeted.” Every time, the answer is the same four line items. Let me break each one down so you can model it before signing anything.
AWS list compute: $6.88 per GPU-hour. Add egress ($0.09 per GB verified, aws.amazon.com/ec2/pricing), FSx Lustre storage, NAT Gateway processing, and enterprise support tier surcharge — all-in effective rate exceeds $9.85 per GPU-hour for a realistic 15TB per monthnth egress workflow. Neo-clouds: zero egress or $0.01 per GB max.
Verified: AWS charges $0.09 per GB for the first 10TB of outbound internet transfer per month (aws.amazon.com/ec2/pricing, September 2026). A 20TB model checkpoint flush — standard for 70B parameter full-weight fine-tuning — costs $1,843 per save. Four checkpoint saves per week: $7,372 per monthnth in pure egress overhead. RunPod: free inbound, nominal outbound. CoreWeave enterprise: $0.01–$0.015 per GB negotiated rate.
Multi-node distributed training on AWS requires Amazon FSx for Lustre for shared POSIX storage. AWS now also offers an Intelligent-Tiering storage class starting below $0.005 per GB-month for elastically-provisioned workloads — but provisioned SSD tiers (which most ML teams require for deterministic throughput) carry significantly higher costs depending on the throughput level (MB/s-per-TiB) you need. Local NVMe scratch on p5 instances (8× 3.84TB NVMe SSDs) does not survive instance termination. CoreWeave native GPFS & BeeGFS: $0.08–$0.12 per GB-month. RunPod: $0.07 per GB-month.
Multi-node clusters spanning AWS Availability Zones incur cross-AZ transfer fees of $0.01 per GB in each direction. Private-subnet clusters require NAT Gateways at $0.045 per hour + $0.045 per GB processed. Moving 30TB of pre-training shards through a NAT Gateway adds $1,350 in gateway transit fees alone. CoreWeave and RunPod: zero intra-cluster networking fees — flat fabric, no artificial per-GB VPC tollbooths.
Production AI workloads require AWS Business or Enterprise Support to access technical account managers and sub-15-minute response SLAs. AWS Enterprise Support charges a sliding scale: 10% of the first $150k monthly spend, 7% of the next $350k. A 4× p5.48xlarge cluster running $220k per monthnth in compute triggers a mandatory ~$17,000+ per monthnth support levy. Neo-clouds bundle equivalent Slack & Teams engineering support directly into annual commit contracts at no surcharge.
30-Day TCO Model: One 8×H100 Node, Fully Loaded
This is the model I run for every client who comes to me with an AWS bill they can’t explain. One 8×H100 SXM5 node, 720 hours continuous operation (30 days), 20TB scratch storage, 15TB monthly outbound egress, standard enterprise support needs. The numbers will surprise you if you’ve been quoting from compute-only rate cards.
* 1-yr AWS Savings Plan rate estimated based on historical ~42% discount from OD; CoreWeave 1-yr reserved estimated at 60% off OD per public statements. All figures are approximations for planning purposes. Verify current rates with provider directly before signing commitments.
AWS on-demand delivers an effective all-in rate exceeding $9.85 per GPU-hour, versus CoreWeave 1-year reserved at approximately $2.80 per GPU-hour fully loaded — a 3.5× premium. Even with a 1-year AWS Savings Plan, hyperscaler infrastructure remains 2.2× more expensive than a dedicated CoreWeave contract when egress and storage are modeled. This gap is why the largest enterprise AI teams are multi-cloud: AWS for compliance-gated production APIs, neo-clouds for training and batch jobs.
Interconnect Forensics: InfiniBand vs EFAv2 vs RoCEv2
The number that most procurement teams ignore — and that destroys their training economics — is Model FLOPs Utilization (MFU). MFU measures the fraction of your theoretical GPU FLOP budget that produces useful gradient updates. The rest is wasted in inter-node synchronization, NCCL timeouts, and PFC pause storms. The interconnect fabric is the primary determinant of MFU at multi-node scale.
P = GPU count in collective, Sparams = gradient volume (bytes), Binter = bidirectional inter-node bandwidth, Lhop = per-hop latency. CoreWeave NVIDIA Quantum-2 InfiniBand: sub-1.5 μs hop latency, SHARP in-network all-reduce. AWS EFAv2 SRD: 5–8 μs. RunPod RoCEv2 under PFC pause storms: 120ms NCCL timeouts documented in empirical 32-GPU Llama 3.1 70B tests.
The practical implication: RunPod at $3.99 per hour with ~34% MFU has a higher effective cost per useful FLOP than CoreWeave 1-year reserved at ~$2.80 per hour with 50%+ MFU. Cheap headline compute rates do not survive FLOP-per-dollar accounting on multi-node training workloads. This is the calculation I push every client to run before committing to a spot-only strategy for large models. For context on how architecture-level efficiency gains interact with infrastructure cost, our GRPO vs PPO VRAM analysis shows how algorithmic choices halve cluster node counts before you even pick a provider.
Spot Eviction Economics & the Checkpointing Penalty
Where Woverhead is the fractional compute wasted to eviction recovery (mid-step eviction rollback + Tsave + Treload+compile). Empirical RunPod Community spot: eviction rate 8.4% daily avg, surging to 22.1% during peak NA hours (14:00–21:00 UTC). 140GB FP16 checkpoint flush at 1Gbps network volume: ~19 minutes — longer than the 30–45 second termination warning. Net wasted compute: ~14.2% of billed hours. Effective spot rate: $3.99 ÷ (1 − 0.142) ≈ $4.65 per GPU-hour effective, not the $3.99 list.
Termination notice: 30–45 seconds webhook
140GB checkpoint flush: ~19 minutes at 1Gbps
Net wasted compute overhead: ~14.2%
Effective rate: ~$4.65 per GPU-hour (not $3.99)
Termination notice: 2-minute EC2 IMDS signal
Capacity Blocks for ML: zero eviction risk
CB price premium: 15–25% over OD
Better graceful-shutdown window than RunPod
Capacity via predictable on-demand or reserved
Cluster uptime ETTR: 97–98%
InfiniBand — no PFC pause storm risk
Best for: multi-week sustained training runs
Break-Even Math: When to Commit vs Rent
If your actual utilization fraction U > Ube, a reserved or committed contract is cheaper. All break-even calculations below use September 2026 verified pricing.
Rule: If your AWS cluster runs more than 14 hours per day, commit to a 1-year Savings Plan. Below that, on-demand is cheaper.
Rule: If you need H100 capacity for more than 8.6 hours per day, a CoreWeave 1-year reserved cluster running 24/7 costs less per node-hour than AWS on-demand, even ignoring AWS egress and storage surcharges.
Rule: After accounting for ~14.2% wasted compute from evictions, Community Cloud spot saves roughly 14–27% vs Secure Cloud on-demand — not the 37%+ implied by list prices alone. Factor in checkpoint engineering overhead before committing to spot-only strategies.
The Tri-Tier Procurement Architecture
This is the architecture I recommend to every engineering leader I work with who is spending more than $50k per monthnth on GPU compute. It is not a theoretical framework — it is the actual topology I have seen deployed at Series B through pre-IPO AI companies that successfully defended their gross margin through the 2024–2026 pricing transition. It also aligns with the economics documented in our American cloud markup teardown.

GPU MSA Negotiation: 4 Non-Negotiable Clauses
I have reviewed GPU Master Services Agreements for teams at Series A through late-stage. The contracts that end up causing disputes share a common profile: they guarantee host availability, not silicon health; they lock compute into a single generation; they do not cap egress; and they give spot workloads 30 seconds to checkpoint before termination. Here are the four clauses that change all of that.
Standard SLAs guarantee host ping availability (99.9%), not silicon health. Your MSA must mandate: billing suspends immediately if NVLink or PCIe interconnect drops below 850 GB/s, double-bit ECC error counts cross threshold, or GPU clocks throttle below 1,590 MHz due to datacenter thermal strain.
Eliminate rigid single-generation lock-in. Require contractually guaranteed migration rights: committed H100 spend must be portable to Blackwell B200 or Rubin allocations at prevailing market price-per-FLOP equivalents. CoreWeave’s “Flex Reservations” model is a template for this kind of workload-flexible capacity commitment.
When negotiating multi-year hyperscaler commitments, demand an Egress Waiver Amendment capping outbound transfer to S3 or Cloudflare R2 at $0.015 per GB or zero. AWS has granted this to enterprise customers with documented AI workload profiles at >$500k ARR. Present egress spend projections and migration intent to neo-cloud as negotiating leverage.
For spot and preemptible commitments, require a minimum 120-second ACPI shutdown notice via POSIX SIGTERM signals. This guarantees in-flight PyTorch state checkpointing to distributed storage before pod termination — the difference between a recoverable training run and a multi-hour rollback.
FAQ
The H100 price crash is real — but it is unevenly distributed. Neo-clouds absorbed the deflation. AWS absorbed it into new line items. The gap between CoreWeave’s ~$2.80 per GPU-hour all-in (1-year reserved) and AWS’s $9.85+ per GPU-hour all-in is not market inefficiency. It is a deliberate infrastructure margin architecture built on egress taxation, FSx Lustre storage mandates, NAT Gateway tollbooths, and support tier percentages. The procurement teams winning right now are the ones who have decomposed the bill, built a multi-cloud topology that routes each workload class to its cost-optimal provider, and negotiated egress waivers into every hyperscaler MSA from day one. The ones still paying a single-vendor AWS rate for training workloads are funding their competitors’ gross margin.
