AIVory  ·  GPU Marketplace

Live spot pricing

NVIDIA A100 — the proven workhorse.

The most widely available data center GPU in the cloud, with two VRAM variants and the most mature software ecosystem in AI. Best-in-class documentation, battle-tested tooling, and consistent availability across every major provider make the A100 the pragmatic backbone of production ML.

At a glance

A100 40GB specifications.

Key hardware specs that determine what workloads this GPU handles.

80GB
VRAM

HBM2e memory

2.0 TB/s
Memory Bandwidth

peak throughput

400W
TDP

thermal design power

Ampere
Architecture

NVIDIA GPU architecture

Spot pricing

A100 40GB: live hourly rates.

Every provider offering this GPU on the spot market, sorted cheapest first.

Loading spot prices…

Prices in USD per GPU-hour · spot instances · sorted cheapest first

Use cases

What the A100 40GB is built for.

  1. Running 70B dense models without quantization

    The A100 80GB is the sweet spot for teams that need full FP16 precision on 70B-class models without the complexity of multi-GPU setups. Load Llama 3.3 70B or Qwen 2.5 72B on a single card with enough headroom for KV cache and batch processing — no quantization artifacts, no tensor parallelism overhead.

  2. Multi-GPU inference clusters with proven NVLink

    The A100's third-generation NVLink provides 600 GB/s bidirectional bandwidth between GPUs. Unlike consumer cards that rely on PCIe for inter-GPU communication, A100 clusters scale predictably. Every major ML framework has been optimized for A100 NVLink topology over the past four years.

  3. Cost-efficient batch processing with established tooling

    When you need to process millions of tokens overnight or run evaluation suites across hundreds of prompts, the A100 offers the best cost-per-token at the data center tier. Its widespread availability means spot capacity is consistently available, and every inference framework from vLLM to TGI is thoroughly tested on A100 hardware.

FAQ

Common questions.

A100 40GB vs 80GB — which variant should I choose?

Choose 40GB if your model fits in 35GB or less after quantization (most 7B-13B models in FP16, or 30B models in INT8). Choose 80GB when you need to run 70B models in FP16 or want maximum KV cache for long-context workloads. The 80GB variant costs 15-30% more per hour but avoids the headaches of model sharding.

When is the A100 still the right choice over the H100?

When cost matters more than peak throughput. The A100 delivers roughly half the LLM inference performance of an H100 at roughly a third of the spot price. For batch workloads without strict latency requirements, two A100s can match one H100's throughput at a lower total cost. The A100 also has better spot availability across more providers and regions.

Why is A100 availability consistently the highest across cloud providers?

The A100 shipped in 2020 and has been the default data center GPU for four years. Providers bought massive quantities, and as customers upgrade to H100/H200, the freed-up A100 capacity enters the spot market. This oversupply keeps prices competitive and availability high — you rarely face stock-outs on A100 spot.

What are the software compatibility advantages of the A100?

Every major ML framework, inference server, and optimization tool has been developed and tested on A100 first. CUDA 11.x+ support is rock-solid, FP16 and INT8 tensor cores are thoroughly documented, and community resources (guides, benchmarks, debugging tips) are vastly more abundant for A100 than for any other GPU. If your stack is new or experimental, A100 minimizes integration surprises.

Rent a A100 40GB. Right now.

Spot pricing, per-second billing, no commitment.

Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.