The GPU Price Crash of 2025 and Why It Reversed
Image generated by AI

The GPU Price Crash of 2025 and Why It Reversed

I have been tracking GPU rental prices since 2023. The story everyone told back then was simple: demand outstrips supply, prices only go up, buy now or regret it. That story was wrong for two years. And now it is wrong again, but in the other direction.

So here is what actually happened - with numbers, not narratives.

The crash: 2023 to late 2025

In early 2023, an H100 rented for roughly $8/hr. Every company with a machine learning team had placed massive GPU orders simultaneously. Nvidia could not ship fast enough. Prices were absurd.

Then the supply caught up. All those orders arrived. Cloud providers, GPU brokers, and sovereign AI clusters all came online within the same 12-month window. The market went from scarcity to glut almost overnight.

By late 2025, one-year contract rates for H100s had fallen to $1.70/hr - a 75-80% crash from peak. Spot prices on marketplaces like Vast.ai dipped below $1.85/hr. Lambda Labs was listing on-demand H100s at $2.99-$3.99/hr while AWS still charged $6.88/hr for the same card.

The oversupply was real. Brokers who had pre-purchased thousands of GPUs were renting them at a loss just to cover power bills.

H100 pricing across providers (mid-2025 to mid-2026)

Provider Type Mid-2025 ($/hr) Mid-2026 ($/hr) Change
AWS (p5.48xlarge) On-demand $6.88 $6.98 +1.5%
Lambda Labs On-demand $2.99 $3.99 +33%
CoreWeave 1-year contract $1.70 ~$2.40 +41%
Vast.ai Spot $1.85 $2.50-3.20 +35-73%
Nvidia DGX Cloud On-demand ~$3.70 ~$4.40 +19%

Spot pricing saves 50-70% over on-demand across all providers. That spread has held steady through both the crash and the reversal. I covered the mechanics of that spread in Spot GPU Pricing Explained.

The reversal: 2026

Three things happened at once.

Nvidia raised the floor. In early 2026, Nvidia announced roughly a 20% price increase on H100 rental rates through its DGX Cloud and partner programs. That reset the baseline for everyone downstream.

Blackwell demand ate the surplus. New model architectures - reasoning models, multi-modal systems, long-context inference - need more compute per query. The University of Rhode Island measured GPT-5 class queries at 18.35 Wh average versus 2.12 Wh for GPT-4. That is 8.6x more energy per query, which translates directly to more GPU-hours consumed per dollar of API revenue.

Training runs got bigger again. Frontier labs resumed large pre-training runs on Blackwell clusters. That pulled supply out of the rental market. Blackwell spot rates jumped 48% in 60 days - from $2.75/hr to $4.08/hr.

One-year H100 contracts climbed roughly 40% off their late-2025 lows. The crash was over.

API vs self-hosting: where the break-even sits now

The API pricing story is its own crash. GPT-4 class inference dropped 99% in three years - from $30/$60 per million tokens in 2023 to roughly $2/$12 per million tokens in 2026. I broke down the full cost comparison in The Real Cost of GPT-4 vs Open Source.

So when does renting your own GPUs make sense? The rough break-even sits at about $20K/month in API spend, or 100M+ tokens per month. Below that, APIs win on operational simplicity. Above that, self-hosting on spot GPUs starts saving real money - if you can handle the interruptions. I wrote about that tradeoff in Can You Run Production LLM on Spot GPUs?.

The market underneath all of this is enormous and growing. SNS Insider projects the AI inference infrastructure market at $22.80B in 2025, growing to $229.95B by 2035 at a 26.02% CAGR. That is a lot of GPU-hours being bought and sold.

What I actually do with this data

I built Smart Inference partly because I needed to track these prices myself. It watches spot rates across providers and routes inference jobs to whatever is cheapest at that moment. For batch workloads and anything that can tolerate a few seconds of latency, it works well.

Honest limitation: it does not help with reserved contracts. If you are at the scale where one-year commitments make sense, you need to negotiate those directly. And for teams spending under $5K/month on inference, the API route is simpler. Smart Inference is useful in the middle - teams spending $5K-50K/month who want GPU-tier pricing without GPU-tier ops burden.

What comes next

The 2023-2025 crash was a supply overshoot. The 2026 reversal is demand catching up. Neither was permanent.

Blackwell supply will ramp through late 2026 and into 2027. When it does, expect another pricing correction downward - probably not as steep as the H100 crash, because Nvidia learned from the oversupply cycle and is managing allocation more tightly.

The pattern is clear: GPU pricing follows semiconductor cycles, not software pricing logic. Plan for 18-month swings of 30-50% in either direction. Lock in contracts when prices are low. Use spot when they are high. And do not assume today’s price is tomorrow’s price.


Methodology and sources

All pricing data was collected between January 2025 and July 2026 from publicly listed rates on provider websites and pricing pages. Specific sources:

  • AWS pricing: aws.amazon.com/ec2/pricing/on-demand - p5.48xlarge (8x H100) instance, per-GPU hourly rate derived from instance price
  • Lambda Labs pricing: lambdalabs.com/service/gpu-cloud - listed on-demand H100 rates
  • Vast.ai spot pricing: vast.ai/gpu-market - real-time marketplace; figures represent observed low-end spot clearing prices
  • Nvidia DGX Cloud: nvidia.com/dgx-cloud - published rates and partner announcements
  • Blackwell spot rate movement (48% in 60 days): Aggregated from GPU marketplace listings, January-March 2026
  • Nvidia H100 rental price increase (~20%): Nvidia partner communications and DGX Cloud pricing updates, Q1 2026
  • GPT-4 class API pricing decline (99%): OpenAI published pricing pages, 2023 vs 2026; corroborated by Anthropic and Google DeepMind competitive pricing
  • Energy per query (18.35 Wh vs 2.12 Wh): University of Rhode Island, AI energy consumption research, 2025-2026
  • AI Inference Infrastructure market ($22.80B to $229.95B, 26.02% CAGR): SNS Insider - AI Inference Infrastructure Market Report, 2025
  • Self-hosting break-even ($20K/month / 100M+ tokens): Derived from published cloud GPU rates vs API token pricing at scale, cross-referenced with operational cost estimates from multiple infrastructure teams

Pricing snapshots reflect publicly available rates at time of collection. Contract rates vary by commitment length, volume, and negotiation. Spot prices are volatile by nature. All figures are USD.