AIVory  ·  GPU Marketplace

Live spot pricing

NVIDIA V100 — data center power at pocket-change prices.

The original tensor core GPU that launched the deep learning revolution, now available at rock-bottom spot prices. At $0.19/hr the V100 is the cheapest data center-grade GPU you can rent — with 16 or 32 GB of HBM2 and ECC memory, it runs 7B–13B models reliably for pennies.

At a glance

V100 specifications.

Key hardware specs that determine what workloads this GPU handles.

32GB
VRAM

HBM2 memory

900 GB/s
Memory Bandwidth

peak throughput

300W
TDP

thermal design power

Volta
Architecture

NVIDIA GPU architecture

Spot pricing

V100: live hourly rates.

Every provider offering this GPU on the spot market, sorted cheapest first.

Loading spot prices…

Prices in USD per GPU-hour · spot instances · sorted cheapest first

Use cases

What the V100 is built for.

  1. Self-hosting 7B–13B models at the lowest cost

    If your workload runs on Mistral Nemo 12B, Gemma 2 9B, or any sub-13B model, the V100 is the cheapest data center GPU to run it on. At $0.19/hr you pay less than $140/month for 24/7 inference — cheaper than most API subscriptions, with full control over your data and no rate limits.

  2. Development, testing, and prototyping on real hardware

    Before committing to H100 or A100 clusters for production, prototype your inference pipeline on a V100. The tensor cores, ECC memory, and NVLink support mirror the data center experience at a fraction of the cost. If your model runs well on a V100, it will run better on everything above it.

  3. High-volume batch processing where latency doesn't matter

    Embedding generation, document classification, sentiment analysis, and other batch workloads don't need cutting-edge throughput — they need cheap GPU-hours. Spin up 10 V100s for $1.90/hr total and chew through millions of records overnight. The cost per processed item is unbeatable.

FAQ

Common questions.

V100 16GB vs 32GB — which should I pick?

The 16GB variant handles any model under ~14 GB (7B in FP16, 13B in INT8). The 32GB variant is needed for 13B models in FP16 or 27B models in INT4. If you're running Mistral Nemo 12B or Gemma 2 9B in FP16, you want the 32GB. For smaller quantized models, save money with the 16GB.

Why is the V100 so much cheaper than newer GPUs?

The V100 shipped in 2017 and has been superseded by Ampere (A100), Hopper (H100/H200), and Ada Lovelace (L40S, RTX 4090). Cloud providers have large fleets of depreciated V100s, and demand has shifted to newer hardware. The oversupply drives spot prices to $0.19/hr — but the GPU itself is still a 300W data center card with ECC memory, NVLink, and first-generation tensor cores.

Is the V100 still fast enough for production inference?

For sub-13B models, yes. The V100's FP16 tensor cores deliver ~125 TFLOPS, which is plenty for serving Mistral Nemo 12B or Gemma 2 9B at reasonable latency. You won't match an A100's throughput, but at one-sixth the price, the V100 often wins on cost-per-token. The bottleneck is VRAM, not compute — if your model fits, the V100 handles it.

Can I run Llama 3.3 70B on V100s?

Technically yes, with 4–5 V100 32GB cards and tensor parallelism. Practically, this is a bad idea — the V100's PCIe 3.0 interconnect is too slow for multi-GPU inference at acceptable latency. Use a single A100 80GB or H200 instead. The V100 is best for models that fit on a single card.

Rent a V100. Right now.

Spot pricing, per-second billing, no commitment.

Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.