GDDR6 memory
AIVory · GPU Marketplace
Live spot pricingNVIDIA A40 — 48 GB with ECC for the price of a coffee.
The A40 packs 48 GB of ECC-protected GDDR6 and Ampere tensor cores into a passively-cooled data center form factor. At $0.35/hr on spot, it's the cheapest 48 GB GPU available — undercutting even the A6000 — and the ECC memory means bit-flip errors won't silently corrupt your model outputs during long-running inference sessions.
At a glance
A40 specifications.
Key hardware specs that determine what workloads this GPU handles.
peak throughput
thermal design power
NVIDIA GPU architecture
Spot pricing
A40: live hourly rates.
Every provider offering this GPU on the spot market, sorted cheapest first.
Prices in USD per GPU-hour · spot instances · sorted cheapest first
Recommended models
AI models that run well on A40.
Tested model-GPU pairings with notes on why each is a good fit.
Use cases
What the A40 is built for.
-
Budget-friendly 48 GB inference for MoE and mid-range models
The A40 at $0.35/hr is the cheapest 48 GB GPU on the spot market — cheaper than the A6000 ($0.33/hr with less ECC protection) and far cheaper than the L40 ($0.51/hr) or L40S ($0.57/hr). For workloads where you need 48 GB of VRAM but don't need the absolute fastest throughput, the A40 delivers the most memory per dollar.
-
ECC-protected inference for regulated and financial workloads
The A40's ECC memory corrects single-bit errors in real time, preventing silent data corruption during inference. This matters for financial modeling, healthcare AI, and any workload where a wrong output has legal or monetary consequences. Non-ECC GPUs like the RTX 4090 can produce bit flips under sustained load — the A40 eliminates that risk.
-
Batch processing and offline analysis on large models
Document summarization, classification, and embedding generation at scale don't need sub-100ms latency — they need cheap GPU-hours with enough VRAM. The A40 handles 30B models in FP16 or 70B in INT4 at $0.35/hr. Process a million documents overnight for under $3.
FAQ
Common questions.
A40 vs A6000 — which 48 GB Ampere GPU should I pick?
The A40 ($0.35/hr) and A6000 ($0.33/hr) are priced similarly with the same 48 GB VRAM and Ampere architecture. Key differences: the A40 has ECC memory (error correction) and is passively cooled for data centers; the A6000 has higher bandwidth (768 vs 696 GB/s) and display outputs for workstation use. For inference reliability, the A40's ECC is the tiebreaker. For raw throughput, the A6000's extra bandwidth gives it a slight edge.
How does the A40 compare to the newer L40 and L40S?
The L40/L40S use Ada Lovelace architecture (one generation newer) with FP8 support and higher tensor throughput. They cost more ($0.51-$0.57/hr) but deliver roughly 40% better inference performance per GPU-hour. If your workload is latency-sensitive, the L40S wins on speed. If it's throughput-at-minimum-cost, the A40 wins on price.
Does the A40 support multi-GPU configurations?
Yes. A40 servers support up to 4 or 8 GPUs connected via PCIe 4.0. NVLink is not available on the A40 (unlike the A100), so inter-GPU bandwidth is limited to PCIe speeds (~32 GB/s per direction). This is fine for independent model instances but too slow for tensor-parallel inference on models sharded across cards. Use multi-A40 for multi-model hosting, not multi-GPU sharding.
Is the A40 suitable for model training?
For fine-tuning models up to ~15B parameters in BF16, the A40 is usable. LoRA and QLoRA fine-tuning of 30B-70B models works within the 48 GB envelope. Full pre-training at scale should use A100 or H100 for the HBM bandwidth advantage (2.0+ TB/s vs 696 GB/s). The A40 is a training-capable inference card, not a training-first card.
Rent a A40. Right now.
Spot pricing, per-second billing, no commitment.
Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.