Live · 32 GPUs · 5 providers

AIVory  ·  GPU Price Index

The cheapest GPU cloud, one table.

Every GPU we track, every provider we compare, cheapest live spot price first. Real numbers from the same feed that powers our 32 individual GPU pages — nothing hand-typed, nothing hidden behind a form.

Renting a GPU by the hour sounds simple until you try to price-shop it. Every cloud publishes its own rate card, spot prices move by the minute, and "cheap" on one provider's page means nothing next to another's. So instead of checking five tabs, we built one.

This page pulls the same live spot data that feeds our GPU marketplace and each GPU's own landing page, lines every card up side by side, and shows the single cheapest offer we're currently tracking for each one — consumer cards like the RTX 3090 through datacenter flagships like the B300. Click through to any GPU's own page for full specs, use cases, and FAQs.

32
GPU models compared

From the RTX 3090 to the B300 — every card with its own landing page on this site.

70
Live spot offers tracked

Individual offers in the current snapshot, across every connected provider.

5
Providers aggregated

RunPod, Vast.ai, AWS Spot, Azure Spot and Crusoe Cloud, checked side by side.

The comparison

Every GPU, cheapest offer first.

Click a column to sort. Click a GPU to see its full pricing history and specs.

Live spot pricing
Details
RTX 3090Cheapest 24 GB Ampere $0.10/hr vast-ai View pricing →
RTX A2000 6 GB Ampere $0.12/hr runpod View pricing →
RTX 3070 8 GB Ampere $0.13/hr runpod View pricing →
RTX 3080 12 GB Ampere $0.17/hr runpod View pricing →
RTX A4000 16 GB Ampere $0.17/hr runpod View pricing →
RTX A5000 24 GB Ampere $0.17/hr runpod View pricing →
RTX 4000 Ada 20 GB Ada Lovelace $0.18/hr runpod View pricing →
T4 16 GB Turing $0.18/hr azure-spot View pricing →
RTX 4070 Ti 12 GB Ada Lovelace $0.19/hr runpod View pricing →
RTX A4500 20 GB Ampere $0.19/hr runpod View pricing →
V100 32 GB Volta $0.19/hr runpod View pricing →
RTX 4080 16 GB Ada Lovelace $0.27/hr runpod View pricing →
RTX 2000 Ada 16 GB Ada Lovelace $0.28/hr runpod View pricing →
RTX 4090 24 GB Ada Lovelace $0.29/hr vast-ai View pricing →
A6000 48 GB Ampere $0.33/hr runpod View pricing →
A40 48 GB Ampere $0.35/hr runpod View pricing →
L4 24 GB Ada Lovelace $0.37/hr vast-ai View pricing →
RTX 5080 16 GB Blackwell $0.39/hr runpod View pricing →
RTX 5090 32 GB Blackwell $0.39/hr vast-ai View pricing →
RTX 5000 Ada 32 GB Ada Lovelace $0.50/hr runpod View pricing →
L40 48 GB Ada Lovelace $0.51/hr vast-ai View pricing →
L40S 48 GB Ada Lovelace $0.57/hr vast-ai View pricing →
RTX 6000 Ada 48 GB Ada Lovelace $0.74/hr runpod View pricing →
RTX PRO 4500 32 GB Blackwell $0.76/hr runpod View pricing →
A100 40GB 80 GB Ampere $1.14/hr crusoe-cloud View pricing →
A10G 24 GB Ampere $1.20/hr aws-spot View pricing →
RTX PRO 6000 96 GB Blackwell $2.23/hr runpod View pricing →
MI300X 192 GB CDNA 3 $2.35/hr runpod View pricing →
H100 80 GB Hopper $2.35/hr runpod View pricing →
H200 141 GB Hopper $4.71/hr runpod View pricing →
B200 180 GB Blackwell $6.48/hr runpod View pricing →
B300 288 GB Blackwell $8.19/hr runpod View pricing →

Prices shown are per-GPU, per-hour, spot rate — billed per second after a 15-minute minimum. They move by the hour and sometimes by the minute; the row you're looking at may already have a cheaper offer by the time you click through.

Methodology

How this list is built.

The numbers come from the same pricing feed that runs AIVory's Smart Inference GPU marketplace: live offers pulled from RunPod, Vast.ai, AWS Spot, Azure Spot and Crusoe Cloud. For every GPU model, we take the single cheapest live offer across every one of its variants (an "A100" row, for example, covers both the 40GB and 80GB cards) and show that price, that provider, and a link to the full page for that GPU.

The table above is rendered server-side from a snapshot of that feed taken when this page was last built, so it loads instantly and stays fully readable to search engines and any browser with JavaScript off. Once the page is open, the same script that powers our live GPU board tries to refresh the price cells from the current feed in the background — if that succeeds you're looking at the current second's prices; if it can't reach the feed, you're looking at the snapshot, and the status badge above the table says so.

Two honest limits. First, this is not every GPU cloud on Earth — it's the providers AIVory aggregates today, not Lambda, CoreWeave, Paperspace or the dozens of smaller resellers we don't yet connect to. Second, we are not a neutral third party: this table sits on AIVory's own site, and every GPU links to a page where you can rent that card through us. The prices themselves are real and traceable to the same feed powering our marketplace — we'd rather you catch us wrong on a number than trust us on our word.

Reading the table

VRAM and architecture, in practice.

VRAM is the number that decides what actually fits. A quantized 7-13B model runs comfortably in 16-24 GB (T4, RTX 4090, L4). A 70B model in FP16 needs 140+ GB, so it wants a multi-GPU H100 or A100 setup, or one of the newer 80-192 GB cards like the H200 or MI300X. Anything past 200B parameters — DeepSeek V3.2, Qwen 3 235B — is a sharded, multi-card job regardless of which GPU you pick.

Architecture tells you the generation, and roughly the price-to-performance band. Turing (T4) and Ampere (A100, A10G, RTX 30-series, A-series Quadro cards) are the value tier: older, cheaper, still fine for inference at moderate throughput. Ada Lovelace (RTX 40-series, L4, L40, L40S) sits in the middle: faster tensor cores, still consumer-card pricing on the RTX side. Hopper (H100, H200) and Blackwell (B200, B300, RTX 50-series, RTX PRO 6000) are the current frontier — the only realistic choice for training or serving the largest open models at production throughput, and priced accordingly.

If you'd rather not manage a GPU at all — no instance to babysit, no driver setup, no idle-timeout to configure — Smart Inference routes your API calls to the cheapest working provider automatically, including spot GPUs when a model needs one, behind one OpenAI-compatible endpoint.

FAQ

Common questions.

What is a spot GPU?

Spare GPU capacity a cloud provider sells cheaply because it isn't currently reserved by another customer. The trade-off is that the instance can be reclaimed with little notice if the provider needs the capacity back. For batch training, fine-tuning, and fault-tolerant inference, spot pricing typically runs 50 to 80 percent below the same card's on-demand rate.

Why do prices change?

Spot pricing works like an auction: it floats with supply and demand for that specific GPU, in that specific region, at that specific moment. When more capacity sits idle, prices drop. When a provider's supply tightens, or when a burst of training jobs starts up, prices climb. That is also why the same GPU model can show wildly different prices across providers at the same time — each provider's pool of idle capacity is separate.

How is this list built?

From a live feed of spot offers across RunPod, Vast.ai, AWS Spot, Azure Spot and Crusoe Cloud — the same feed behind AIVory's GPU marketplace and every individual GPU page on this site. For each GPU model we show the single cheapest live offer among its variants. The table is server-rendered from a snapshot at build time, then refreshed in your browser from the live feed where possible. See the methodology section above for the full explanation, including what this list doesn't cover.

What is the cheapest GPU cloud right now?

It changes by the hour, which is the entire point of this page — check the table above instead of trusting a fixed answer. As a rule of thumb, consumer cards on peer-to-peer marketplaces like Vast.ai tend to undercut datacenter-only providers on older Ampere and Ada cards, while RunPod is frequently the cheapest source for newer Hopper and Blackwell datacenter GPUs.

What happens if my spot instance gets reclaimed?

The provider ends your session, usually with a short warning window. If you're renting directly, you restart on the next cheapest available offer yourself. If you're using Smart Inference, the router can fail over automatically to another provider so a reclaimed instance doesn't take your API down with it.

Which GPU should I pick for running an LLM?

Match VRAM to model size first, price second. A quantized model under 13B fits an RTX 4090, L4, or T4. A 70B model wants at least one H100 or A100 80GB, more if you're running it unquantized. Anything above 200B parameters needs a multi-GPU cluster regardless of which card you choose. See the "reading the table" section above for the full breakdown by architecture.

Smart Inference

Don't want to shop for a GPU at all?

Renting a card yourself is the hands-on path. If you'd rather send API calls and let routing handle the rest, Smart Inference picks the cheapest working provider per request — including spot GPUs when a model needs one — behind one OpenAI-compatible endpoint.

Cheap GPUs. Real prices. No guessing.

One table, updated live, linked to every GPU's own page.