HBM3 memory
AIVory · GPU Marketplace
Live spot pricingAMD MI300X — 192 GB of HBM3 without the NVIDIA tax.
AMD's answer to the H100 packs 192 GB of HBM3 and 5.3 TB/s memory bandwidth on a single accelerator. At $2.35/hr on spot — the same price as an H100 with 2.4x the VRAM — the MI300X is the most memory-per-dollar you can rent for large-model inference. ROCm and vLLM support make it a drop-in option for teams willing to leave CUDA behind.
At a glance
MI300X specifications.
Key hardware specs that determine what workloads this GPU handles.
peak throughput
thermal design power
NVIDIA GPU architecture
Spot pricing
MI300X: live hourly rates.
Every provider offering this GPU on the spot market, sorted cheapest first.
Prices in USD per GPU-hour · spot instances · sorted cheapest first
Recommended models
AI models that run well on MI300X.
Tested model-GPU pairings with notes on why each is a good fit.
Use cases
What the MI300X is built for.
-
Running 70B models in full precision on a single card
The MI300X's 192 GB HBM3 holds Llama 3.3 70B or Qwen 2.5 72B in full FP16 on one GPU — no quantization, no multi-card sharding. At $2.35/hr, that's the same hourly rate as an H100 SXM that would need quantization or a second card for the same model. If your inference stack runs on vLLM, the switch is straightforward.
-
Cost-optimized inference for memory-bound workloads
When your bottleneck is VRAM, not compute — long-context serving, large batch inference, or multi-model hosting — the MI300X offers 192 GB at the same price point as 80 GB H100s. You get 2.4x the memory per dollar, which directly translates to more concurrent requests or larger models without upgrading to multi-GPU.
-
Vendor diversification for production AI infrastructure
Running your entire inference fleet on NVIDIA creates supply-chain risk. The MI300X gives teams a second-source option with competitive performance. vLLM, PyTorch, and JAX all support ROCm, so workloads that don't depend on CUDA-specific libraries can run on MI300X without code changes.
FAQ
Common questions.
How does the MI300X compare to the H100 for LLM inference?
The MI300X has 2.4x the VRAM (192 vs 80 GB) and 1.6x the memory bandwidth (5.3 vs 3.35 TB/s) of an H100 SXM. On memory-bound workloads like large-model inference, the MI300X matches or exceeds H100 throughput. On compute-bound workloads with small models and high batch sizes, the H100's tensor cores and software maturity give it an edge. The MI300X's strongest advantage is fitting 70B+ models on a single card without quantization.
Is the ROCm software stack mature enough for production?
For inference: yes. vLLM's ROCm backend has been production-grade since mid-2025, and major cloud providers run MI300X inference at scale. PyTorch and JAX support ROCm natively. For training: it depends on your framework. Megatron-LM and DeepSpeed have ROCm forks, but CUDA-only training libraries may need porting. If your workload is inference on vLLM, the software gap is negligible.
Why is the MI300X priced the same as the H100 despite having more VRAM?
AMD prices the MI300X competitively to win market share from NVIDIA. The $2.35/hr spot rate reflects AMD's strategy of offering more memory at parity pricing. For cloud providers, MI300X supply is growing as AMD ramps production, and the lower demand relative to H100 keeps spot prices from spiking. This is good for users — you get more hardware for the same money.
Can I train models on the MI300X or is it inference-only?
The MI300X supports both training and inference. Its CDNA 3 architecture includes matrix cores optimized for FP16/BF16/FP8 training. Multi-GPU MI300X clusters connected via Infinity Fabric can train 70B+ models. The practical limitation is software: if your training pipeline relies on CUDA-specific kernels (FlashAttention custom builds, NCCL), you may need to switch to ROCm equivalents (Triton, RCCL).
Rent a MI300X. Right now.
Spot pricing, per-second billing, no commitment.
Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.