HBM2 memory
AIVory · GPU Marketplace
Live spot pricingNVIDIA V100 — data center power at pocket-change prices.
The original tensor core GPU that launched the deep learning revolution, now available at rock-bottom spot prices. At $0.19/hr the V100 is the cheapest data center-grade GPU you can rent — with 16 or 32 GB of HBM2 and ECC memory, it runs 7B–13B models reliably for pennies.
At a glance
V100 specifications.
Key hardware specs that determine what workloads this GPU handles.
peak throughput
thermal design power
NVIDIA GPU architecture
Spot pricing
V100: live hourly rates.
Every provider offering this GPU on the spot market, sorted cheapest first.
Prices in USD per GPU-hour · spot instances · sorted cheapest first
Recommended models
AI models that run well on V100.
Tested model-GPU pairings with notes on why each is a good fit.
Use cases
What the V100 is built for.
-
Self-hosting 7B–13B models at the lowest cost
If your workload runs on Mistral Nemo 12B, Gemma 2 9B, or any sub-13B model, the V100 is the cheapest data center GPU to run it on. At $0.19/hr you pay less than $140/month for 24/7 inference — cheaper than most API subscriptions, with full control over your data and no rate limits.
-
Development, testing, and prototyping on real hardware
Before committing to H100 or A100 clusters for production, prototype your inference pipeline on a V100. The tensor cores, ECC memory, and NVLink support mirror the data center experience at a fraction of the cost. If your model runs well on a V100, it will run better on everything above it.
-
High-volume batch processing where latency doesn't matter
Embedding generation, document classification, sentiment analysis, and other batch workloads don't need cutting-edge throughput — they need cheap GPU-hours. Spin up 10 V100s for $1.90/hr total and chew through millions of records overnight. The cost per processed item is unbeatable.
FAQ
Common questions.
V100 16GB vs 32GB — which should I pick?
The 16GB variant handles any model under ~14 GB (7B in FP16, 13B in INT8). The 32GB variant is needed for 13B models in FP16 or 27B models in INT4. If you're running Mistral Nemo 12B or Gemma 2 9B in FP16, you want the 32GB. For smaller quantized models, save money with the 16GB.
Why is the V100 so much cheaper than newer GPUs?
The V100 shipped in 2017 and has been superseded by Ampere (A100), Hopper (H100/H200), and Ada Lovelace (L40S, RTX 4090). Cloud providers have large fleets of depreciated V100s, and demand has shifted to newer hardware. The oversupply drives spot prices to $0.19/hr — but the GPU itself is still a 300W data center card with ECC memory, NVLink, and first-generation tensor cores.
Is the V100 still fast enough for production inference?
For sub-13B models, yes. The V100's FP16 tensor cores deliver ~125 TFLOPS, which is plenty for serving Mistral Nemo 12B or Gemma 2 9B at reasonable latency. You won't match an A100's throughput, but at one-sixth the price, the V100 often wins on cost-per-token. The bottleneck is VRAM, not compute — if your model fits, the V100 handles it.
Can I run Llama 3.3 70B on V100s?
Technically yes, with 4–5 V100 32GB cards and tensor parallelism. Practically, this is a bad idea — the V100's PCIe 3.0 interconnect is too slow for multi-GPU inference at acceptable latency. Use a single A100 80GB or H200 instead. The V100 is best for models that fit on a single card.
Rent a V100. Right now.
Spot pricing, per-second billing, no commitment.
Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.