AIVory  ·  GPU Marketplace

Live spot pricing

NVIDIA H200 — 141 GB of raw memory.

The memory monster. With 141 GB of HBM3e and 4.8 TB/s bandwidth, the H200 is the first GPU where Llama 3.3 70B runs in full FP16 on a single card — no sharding, no quantization, no multi-GPU overhead. When your bottleneck is VRAM, the H200 eliminates it.

At a glance

H200 specifications.

Key hardware specs that determine what workloads this GPU handles.

141GB
VRAM

HBM3e memory

4.8 TB/s
Memory Bandwidth

peak throughput

700W
TDP

thermal design power

Hopper
Architecture

NVIDIA GPU architecture

Spot pricing

H200: live hourly rates.

Every provider offering this GPU on the spot market, sorted cheapest first.

Loading spot prices…

Prices in USD per GPU-hour · spot instances · sorted cheapest first

Use cases

What the H200 is built for.

  1. Single-GPU FP16 inference for 70B dense models

    The H200's defining advantage is eliminating multi-GPU setups for 70B models. Llama 3.3 70B in FP16 needs ~140 GB — it fits on one H200 with room for KV cache. This removes tensor parallelism overhead, simplifies deployment, halves your failure surface, and cuts operational complexity by an order of magnitude.

  2. Reduced cluster size for 200B+ MoE models

    When each GPU holds 141 GB instead of 80 GB, you need fewer cards for the same model. A DeepSeek V3.2 deployment that requires 8x H100 can run on 5-6x H200, reducing NVLink traffic, simplifying orchestration, and often lowering total cost despite the higher per-GPU rate.

  3. Large context window workloads

    Long-context inference is memory-bound — the KV cache for a 128K context on a 70B model can consume 40+ GB alone. The H200's 141 GB accommodates both the model weights and massive KV caches, while the 4.8 TB/s bandwidth keeps attention computation fast even at extreme sequence lengths.

FAQ

Common questions.

Why does 141 GB of VRAM matter for AI inference?

Model weights, KV cache, and activation memory all compete for VRAM. A 70B model in FP16 needs ~140 GB just for weights. On an H100 (80 GB), you must either quantize or shard across two GPUs. The H200's 141 GB removes that trade-off — run full precision on one card, with headroom for batch processing and long contexts.

Is the H200's premium over the H100 worth paying?

For memory-bound workloads, yes. If your model fits on an H100 with room to spare (e.g., 13B or 30B models), the H200's extra VRAM goes unused and you overpay. But if you are running 70B models, processing 128K+ contexts, or deploying large MoE architectures, the H200 saves money by reducing GPU count and eliminating sharding overhead.

When will H200 spot availability improve?

H200 started shipping in volume in late 2025. As of mid-2026, spot availability is growing but still trails the A100 and H100. RunPod and a few specialised providers offer H200 spot instances. Expect broader availability and lower prices through the second half of 2026 as more units enter the secondary market.

H200 for training vs inference — where does it shine?

The H200's advantage is memory capacity, not raw compute over the H100. For training, it lets you fit larger micro-batches and reduce gradient accumulation steps, modestly improving training throughput. For inference, the impact is dramatic: single-GPU serving of 70B models eliminates all multi-GPU coordination, cutting latency and operational complexity in half.

Rent a H200. Right now.

Spot pricing, per-second billing, no commitment.

Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.