HBM2e memory
AIVory · GPU Marketplace
Live spot pricingNVIDIA A100 — the proven workhorse.
The most widely available data center GPU in the cloud, with two VRAM variants and the most mature software ecosystem in AI. Best-in-class documentation, battle-tested tooling, and consistent availability across every major provider make the A100 the pragmatic backbone of production ML.
At a glance
A100 40GB specifications.
Key hardware specs that determine what workloads this GPU handles.
peak throughput
thermal design power
NVIDIA GPU architecture
Spot pricing
A100 40GB: live hourly rates.
Every provider offering this GPU on the spot market, sorted cheapest first.
Prices in USD per GPU-hour · spot instances · sorted cheapest first
Recommended models
AI models that run well on A100 40GB.
Tested model-GPU pairings with notes on why each is a good fit.
Use cases
What the A100 40GB is built for.
-
Running 70B dense models without quantization
The A100 80GB is the sweet spot for teams that need full FP16 precision on 70B-class models without the complexity of multi-GPU setups. Load Llama 3.3 70B or Qwen 2.5 72B on a single card with enough headroom for KV cache and batch processing — no quantization artifacts, no tensor parallelism overhead.
-
Multi-GPU inference clusters with proven NVLink
The A100's third-generation NVLink provides 600 GB/s bidirectional bandwidth between GPUs. Unlike consumer cards that rely on PCIe for inter-GPU communication, A100 clusters scale predictably. Every major ML framework has been optimized for A100 NVLink topology over the past four years.
-
Cost-efficient batch processing with established tooling
When you need to process millions of tokens overnight or run evaluation suites across hundreds of prompts, the A100 offers the best cost-per-token at the data center tier. Its widespread availability means spot capacity is consistently available, and every inference framework from vLLM to TGI is thoroughly tested on A100 hardware.
FAQ
Common questions.
A100 40GB vs 80GB — which variant should I choose?
Choose 40GB if your model fits in 35GB or less after quantization (most 7B-13B models in FP16, or 30B models in INT8). Choose 80GB when you need to run 70B models in FP16 or want maximum KV cache for long-context workloads. The 80GB variant costs 15-30% more per hour but avoids the headaches of model sharding.
When is the A100 still the right choice over the H100?
When cost matters more than peak throughput. The A100 delivers roughly half the LLM inference performance of an H100 at roughly a third of the spot price. For batch workloads without strict latency requirements, two A100s can match one H100's throughput at a lower total cost. The A100 also has better spot availability across more providers and regions.
Why is A100 availability consistently the highest across cloud providers?
The A100 shipped in 2020 and has been the default data center GPU for four years. Providers bought massive quantities, and as customers upgrade to H100/H200, the freed-up A100 capacity enters the spot market. This oversupply keeps prices competitive and availability high — you rarely face stock-outs on A100 spot.
What are the software compatibility advantages of the A100?
Every major ML framework, inference server, and optimization tool has been developed and tested on A100 first. CUDA 11.x+ support is rock-solid, FP16 and INT8 tensor cores are thoroughly documented, and community resources (guides, benchmarks, debugging tips) are vastly more abundant for A100 than for any other GPU. If your stack is new or experimental, A100 minimizes integration surprises.
Rent a A100 40GB. Right now.
Spot pricing, per-second billing, no commitment.
Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.