AIVory  ·  GPU Marketplace

Live spot pricing

NVIDIA A40 — 48 GB with ECC for the price of a coffee.

The A40 packs 48 GB of ECC-protected GDDR6 and Ampere tensor cores into a passively-cooled data center form factor. At $0.35/hr on spot, it's the cheapest 48 GB GPU available — undercutting even the A6000 — and the ECC memory means bit-flip errors won't silently corrupt your model outputs during long-running inference sessions.

At a glance

A40 specifications.

Key hardware specs that determine what workloads this GPU handles.

48GB
VRAM

GDDR6 memory

696 GB/s
Memory Bandwidth

peak throughput

300W
TDP

thermal design power

Ampere
Architecture

NVIDIA GPU architecture

Spot pricing

A40: live hourly rates.

Every provider offering this GPU on the spot market, sorted cheapest first.

Loading spot prices…

Prices in USD per GPU-hour · spot instances · sorted cheapest first

Use cases

What the A40 is built for.

  1. Budget-friendly 48 GB inference for MoE and mid-range models

    The A40 at $0.35/hr is the cheapest 48 GB GPU on the spot market — cheaper than the A6000 ($0.33/hr with less ECC protection) and far cheaper than the L40 ($0.51/hr) or L40S ($0.57/hr). For workloads where you need 48 GB of VRAM but don't need the absolute fastest throughput, the A40 delivers the most memory per dollar.

  2. ECC-protected inference for regulated and financial workloads

    The A40's ECC memory corrects single-bit errors in real time, preventing silent data corruption during inference. This matters for financial modeling, healthcare AI, and any workload where a wrong output has legal or monetary consequences. Non-ECC GPUs like the RTX 4090 can produce bit flips under sustained load — the A40 eliminates that risk.

  3. Batch processing and offline analysis on large models

    Document summarization, classification, and embedding generation at scale don't need sub-100ms latency — they need cheap GPU-hours with enough VRAM. The A40 handles 30B models in FP16 or 70B in INT4 at $0.35/hr. Process a million documents overnight for under $3.

FAQ

Common questions.

A40 vs A6000 — which 48 GB Ampere GPU should I pick?

The A40 ($0.35/hr) and A6000 ($0.33/hr) are priced similarly with the same 48 GB VRAM and Ampere architecture. Key differences: the A40 has ECC memory (error correction) and is passively cooled for data centers; the A6000 has higher bandwidth (768 vs 696 GB/s) and display outputs for workstation use. For inference reliability, the A40's ECC is the tiebreaker. For raw throughput, the A6000's extra bandwidth gives it a slight edge.

How does the A40 compare to the newer L40 and L40S?

The L40/L40S use Ada Lovelace architecture (one generation newer) with FP8 support and higher tensor throughput. They cost more ($0.51-$0.57/hr) but deliver roughly 40% better inference performance per GPU-hour. If your workload is latency-sensitive, the L40S wins on speed. If it's throughput-at-minimum-cost, the A40 wins on price.

Does the A40 support multi-GPU configurations?

Yes. A40 servers support up to 4 or 8 GPUs connected via PCIe 4.0. NVLink is not available on the A40 (unlike the A100), so inter-GPU bandwidth is limited to PCIe speeds (~32 GB/s per direction). This is fine for independent model instances but too slow for tensor-parallel inference on models sharded across cards. Use multi-A40 for multi-model hosting, not multi-GPU sharding.

Is the A40 suitable for model training?

For fine-tuning models up to ~15B parameters in BF16, the A40 is usable. LoRA and QLoRA fine-tuning of 30B-70B models works within the 48 GB envelope. Full pre-training at scale should use A100 or H100 for the HBM bandwidth advantage (2.0+ TB/s vs 696 GB/s). The A40 is a training-capable inference card, not a training-first card.

Rent a A40. Right now.

Spot pricing, per-second billing, no commitment.

Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.