AIVory  ·  GPU Marketplace

Live spot pricing

NVIDIA RTX A4000 — 16 GB professional power in a single slot.

A single-slot, 140-watt professional card with 16 GB of GDDR6 and Ampere tensor cores. The RTX A4000 at $0.17/hr is one of the cheapest Ampere GPUs on the spot market — bringing professional driver stability and ECC support to the same price tier as the legacy T4, with significantly better compute performance.

At a glance

RTX A4000 specifications.

Key hardware specs that determine what workloads this GPU handles.

16GB
VRAM

GDDR6 memory

448 GB/s
Memory Bandwidth

peak throughput

140W
TDP

thermal design power

Ampere
Architecture

NVIDIA GPU architecture

Spot pricing

RTX A4000: live hourly rates.

Every provider offering this GPU on the spot market, sorted cheapest first.

Loading spot prices…

Prices in USD per GPU-hour · spot instances · sorted cheapest first

Use cases

What the RTX A4000 is built for.

  1. Cheapest Ampere GPU for 7B model production serving

    The RTX A4000 at $0.17/hr matches the RTX A5000 as the cheapest Ampere professional card and undercuts the A10G ($1.20/hr) by 7x. For 7B model inference with professional drivers — where you need Ampere tensor cores but don't need 24 GB VRAM — the A4000 is the most cost-effective option available.

  2. High-density single-slot deployment

    The A4000's single-slot, 140W form factor allows packing 4-8 GPUs into a standard server chassis. Each card serves an independent 7B model endpoint, creating a multi-model inference server at $0.68-$1.36/hr total. No other Ampere professional card offers this density — the A5000 and A6000 are dual-slot.

  3. Professional driver environments for regulated inference

    Financial services, healthcare, and government AI deployments often mandate NVIDIA professional drivers for GPU compute. The RTX A4000 at $0.17/hr is the cheapest entry point for driver-certified inference — 7x less expensive than the A10G and comparable to consumer cards that lack certification.

FAQ

Common questions.

RTX A4000 vs T4 — both 16 GB, which to use?

The RTX A4000 (Ampere, $0.17/hr) and T4 (Turing, $0.18/hr) are priced nearly identically. The A4000 delivers ~2x the tensor throughput thanks to Ampere architecture, supports BF16 that the T4 lacks, and has higher bandwidth (448 vs 320 GB/s). The T4 uses less power (70W vs 140W) and has wider cloud availability. For inference performance at the same price, the A4000 is the better card.

Can the RTX A4000 handle 13B models?

In INT4 quantization, yes — 13B INT4 needs ~7 GB, leaving 9 GB for KV cache. In INT8, 13B needs ~13 GB, leaving only 3 GB — workable for low-concurrency use but tight for production. In FP16, 13B needs ~26 GB, which exceeds the 16 GB VRAM. For comfortable 13B serving, consider the 20 GB RTX 4000 Ada or 24 GB A5000.

Is the RTX A4000 the same as the consumer RTX 4000?

No. The RTX A4000 is an Ampere-generation professional card (GA104 chip). The 'RTX 4000' branding was later reused for the Ada Lovelace professional lineup (RTX 4000 Ada, AD104 chip). They are different GPUs from different generations. The RTX A4000 is Ampere; the RTX 4000 Ada is Ada Lovelace with higher performance and FP8 support.

What's the realistic throughput for Mistral 7B on the A4000?

With vLLM or TGI serving Mistral 7B in FP16, expect 30-40 tokens/second per request on the A4000. With INT8 quantization, throughput rises to 50-65 tokens/second. These numbers assume single-user latency — batched serving with multiple concurrent users increases total throughput at the cost of per-user latency. For interactive chat, the A4000 handles 5-10 concurrent conversations at acceptable latency.

Rent a RTX A4000. Right now.

Spot pricing, per-second billing, no commitment.

Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.