AIVory  ·  GPU Marketplace

Live spot pricing

NVIDIA H100 — the gold standard for AI compute.

The enterprise flagship GPU that defines modern AI infrastructure. Fastest tensor core throughput, 80 GB HBM3, and Hopper architecture deliver unmatched performance for large-model training and high-throughput production inference. Spot prices range from $2.35/hr to $22/hr depending on provider.

At a glance

H100 specifications.

Key hardware specs that determine what workloads this GPU handles.

80GB
VRAM

HBM3 memory

3.35 TB/s
Memory Bandwidth

peak throughput

700W
TDP

thermal design power

Hopper
Architecture

NVIDIA GPU architecture

Spot pricing

H100: live hourly rates.

Every provider offering this GPU on the spot market, sorted cheapest first.

Loading spot prices…

Prices in USD per GPU-hour · spot instances · sorted cheapest first

Use cases

What the H100 is built for.

  1. Training and fine-tuning 70B+ models

    The H100's FP8 tensor cores and 3.35 TB/s bandwidth make it the only practical choice for full fine-tuning or LoRA training on models above 70B parameters. A cluster of 8 H100s can fine-tune Llama 3.3 70B in hours, not days.

  2. High-throughput production serving at 1000+ req/s

    When your inference workload needs to sustain thousands of concurrent requests with tight latency SLAs, the H100's transformer engine delivers 2-3x the throughput of the previous-generation A100 on the same model, making it the cost-effective choice at scale despite the higher hourly rate.

  3. Multi-GPU clusters for frontier model deployment

    Models above 200B parameters require sharding across multiple GPUs. The H100's fourth-generation NVLink provides 900 GB/s bidirectional bandwidth between cards, minimizing the overhead of tensor parallelism and making 8-GPU clusters perform nearly linearly.

FAQ

Common questions.

What is the difference between H100 SXM, PCIe, and NVL variants?

SXM is the full-power variant (700W) with NVLink mesh for multi-GPU clusters — used in DGX systems. PCIe fits standard server slots at 350W with lower interconnect bandwidth. NVL pairs two H100 dies on a single board with 188 GB combined VRAM. For AI training, SXM is the standard; PCIe works for single-GPU inference; NVL suits large-context inference workloads.

How much faster is the H100 than the A100 for LLM inference?

On transformer workloads, the H100 delivers 2-3x the inference throughput of an A100 at comparable batch sizes. The gap comes from the FP8 transformer engine, higher memory bandwidth (3.35 vs 2.0 TB/s), and the Hopper architecture's improved tensor cores. For latency-sensitive serving, you can serve the same traffic with fewer H100s than A100s.

Why do H100 spot prices vary by 10x across providers?

Provider pricing depends on supply, region, and contract structure. RunPod and Vast.ai access wholesale and mining-adjacent capacity, so they can offer $2-4/hr. AWS and Azure charge $18-22/hr on spot because their H100 supply serves enterprise customers with higher reliability expectations. The GPU is identical — the price reflects the provider's overhead and SLA.

What is the minimum cluster size for running a 70B model on H100?

Llama 3.3 70B in FP16 requires ~140 GB of VRAM, so you need at least 2x H100 80GB cards. With 4-bit quantization (GPTQ/AWQ), it fits on a single H100. For production serving at high throughput, 4x H100 with tensor parallelism is the common configuration. Training requires 8x H100 minimum for reasonable iteration speed.

Rent a H100. Right now.

Spot pricing, per-second billing, no commitment.

Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.