HBM3 memory
AIVory · GPU Marketplace
Live spot pricingNVIDIA H100 — the gold standard for AI compute.
The enterprise flagship GPU that defines modern AI infrastructure. Fastest tensor core throughput, 80 GB HBM3, and Hopper architecture deliver unmatched performance for large-model training and high-throughput production inference. Spot prices range from $2.35/hr to $22/hr depending on provider.
At a glance
H100 specifications.
Key hardware specs that determine what workloads this GPU handles.
peak throughput
thermal design power
NVIDIA GPU architecture
Spot pricing
H100: live hourly rates.
Every provider offering this GPU on the spot market, sorted cheapest first.
Prices in USD per GPU-hour · spot instances · sorted cheapest first
Recommended models
AI models that run well on H100.
Tested model-GPU pairings with notes on why each is a good fit.
Use cases
What the H100 is built for.
-
Training and fine-tuning 70B+ models
The H100's FP8 tensor cores and 3.35 TB/s bandwidth make it the only practical choice for full fine-tuning or LoRA training on models above 70B parameters. A cluster of 8 H100s can fine-tune Llama 3.3 70B in hours, not days.
-
High-throughput production serving at 1000+ req/s
When your inference workload needs to sustain thousands of concurrent requests with tight latency SLAs, the H100's transformer engine delivers 2-3x the throughput of the previous-generation A100 on the same model, making it the cost-effective choice at scale despite the higher hourly rate.
-
Multi-GPU clusters for frontier model deployment
Models above 200B parameters require sharding across multiple GPUs. The H100's fourth-generation NVLink provides 900 GB/s bidirectional bandwidth between cards, minimizing the overhead of tensor parallelism and making 8-GPU clusters perform nearly linearly.
FAQ
Common questions.
What is the difference between H100 SXM, PCIe, and NVL variants?
SXM is the full-power variant (700W) with NVLink mesh for multi-GPU clusters — used in DGX systems. PCIe fits standard server slots at 350W with lower interconnect bandwidth. NVL pairs two H100 dies on a single board with 188 GB combined VRAM. For AI training, SXM is the standard; PCIe works for single-GPU inference; NVL suits large-context inference workloads.
How much faster is the H100 than the A100 for LLM inference?
On transformer workloads, the H100 delivers 2-3x the inference throughput of an A100 at comparable batch sizes. The gap comes from the FP8 transformer engine, higher memory bandwidth (3.35 vs 2.0 TB/s), and the Hopper architecture's improved tensor cores. For latency-sensitive serving, you can serve the same traffic with fewer H100s than A100s.
Why do H100 spot prices vary by 10x across providers?
Provider pricing depends on supply, region, and contract structure. RunPod and Vast.ai access wholesale and mining-adjacent capacity, so they can offer $2-4/hr. AWS and Azure charge $18-22/hr on spot because their H100 supply serves enterprise customers with higher reliability expectations. The GPU is identical — the price reflects the provider's overhead and SLA.
What is the minimum cluster size for running a 70B model on H100?
Llama 3.3 70B in FP16 requires ~140 GB of VRAM, so you need at least 2x H100 80GB cards. With 4-bit quantization (GPTQ/AWQ), it fits on a single H100. For production serving at high throughput, 4x H100 with tensor parallelism is the common configuration. Training requires 8x H100 minimum for reasonable iteration speed.
Rent a H100. Right now.
Spot pricing, per-second billing, no commitment.
Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.