AIVory  ·  GPU Marketplace

Live spot pricing

NVIDIA RTX PRO 6000 — 96 GB of Blackwell in a workstation card.

The first professional GPU to cross the 80 GB barrier without using HBM. 96 GB of GDDR7 and Blackwell tensor cores make the RTX PRO 6000 the only workstation card that runs Llama 3.3 70B in full FP16 on a single GPU. At $2.23/hr on spot, it undercuts the A100 80GB while offering 20% more VRAM and a newer architecture.

At a glance

RTX PRO 6000 specifications.

Key hardware specs that determine what workloads this GPU handles.

96GB
VRAM

GDDR7 memory

1.8 TB/s
Memory Bandwidth

peak throughput

350W
TDP

thermal design power

Blackwell
Architecture

NVIDIA GPU architecture

Spot pricing

RTX PRO 6000: live hourly rates.

Every provider offering this GPU on the spot market, sorted cheapest first.

Loading spot prices…

Prices in USD per GPU-hour · spot instances · sorted cheapest first

Use cases

What the RTX PRO 6000 is built for.

  1. Running 70B models on a single professional GPU

    Before the RTX PRO 6000, running Llama 3.3 70B on one card meant either an A100 80GB (tight on VRAM, $1.14/hr), an H100 80GB (expensive at $2.35/hr), or the MI300X (AMD ecosystem). The RTX PRO 6000 at 96 GB and $2.23/hr offers more memory than the A100, a lower price than the H100, and NVIDIA's CUDA ecosystem — a unique combination.

  2. Workstation-based development on frontier models

    AI researchers who need to iterate on 70B+ models locally — running inference, debugging prompts, testing quantization — can install the RTX PRO 6000 in their workstation instead of renting cloud GPU instances. The 350W TDP and standard PCIe form factor fit in high-end desktop chassis with adequate cooling.

  3. Secure on-premises inference for sensitive data

    Organizations that can't send data to cloud providers — defense contractors, healthcare systems, financial institutions — need on-premises GPU infrastructure. The RTX PRO 6000's 96 GB VRAM runs 70B models locally, and NVIDIA's professional driver certification meets enterprise security requirements. Spot cloud instances serve as a development environment before deploying to on-prem hardware.

FAQ

Common questions.

RTX PRO 6000 vs A100 80GB — which is better for 70B models?

The RTX PRO 6000 has 96 GB GDDR7 at 1.8 TB/s; the A100 has 80 GB HBM2e at 2.0 TB/s. The A100's HBM bandwidth gives it an edge in raw token generation speed, but the RTX PRO 6000's extra 16 GB VRAM provides more room for KV cache and higher concurrency. Pricing favors the RTX PRO 6000 ($2.23 vs $1.14-$3.10/hr depending on variant). For memory-bound workloads with large context windows, the RTX PRO 6000's extra VRAM wins.

Why GDDR7 instead of HBM?

GDDR7 is significantly cheaper to manufacture than HBM3/HBM3e, which allows NVIDIA to offer 96 GB on a professional card at workstation pricing. HBM provides higher bandwidth per pin, but GDDR7 closes the gap — the RTX PRO 6000's 1.8 TB/s is comparable to the A100's 2.0 TB/s HBM2e. The tradeoff is power: GDDR7 draws more watts per GB/s than HBM, hence the 350W TDP.

Does the RTX PRO 6000 support multi-GPU NVLink?

Yes. The RTX PRO 6000 supports NVLink for two-card configurations, creating a unified 192 GB memory pool with high-bandwidth interconnect. This enables running 200B+ dense models or 70B models in FP16 with generous batching headroom across two cards.

Is the RTX PRO 6000 the same chip as the consumer RTX 5090?

They share the Blackwell architecture family but are different GPUs. The RTX PRO 6000 has 3x the VRAM (96 vs 32 GB), uses GDDR7 with ECC, includes professional driver certification, and supports NVLink. The RTX 5090 has higher bandwidth per GB (1.79 TB/s for 32 GB) and costs far less ($0.39/hr). They target completely different workloads: the RTX 5090 is for consumer AI enthusiasts, the RTX PRO 6000 is for enterprise deployment.

Rent a RTX PRO 6000. Right now.

Spot pricing, per-second billing, no commitment.

Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.