AIVory  ·  GPU Marketplace

Live spot pricing

NVIDIA A10G — the AWS inference default with Ampere under the hood.

The A10G is the GPU behind AWS G5 instances — the most widely deployed inference accelerator in the world's largest cloud. 24 GB of GDDR6, Ampere tensor cores, and 150W TDP make it a balanced choice for serving 7B-14B models. Spot pricing starts at $1.20/hr, with wide availability across AWS regions and other cloud providers.

At a glance

A10G specifications.

Key hardware specs that determine what workloads this GPU handles.

24GB
VRAM

GDDR6 memory

600 GB/s
Memory Bandwidth

peak throughput

150W
TDP

thermal design power

Ampere
Architecture

NVIDIA GPU architecture

Spot pricing

A10G: live hourly rates.

Every provider offering this GPU on the spot market, sorted cheapest first.

Loading spot prices…

Prices in USD per GPU-hour · spot instances · sorted cheapest first

Use cases

What the A10G is built for.

  1. AWS-native inference for teams already on Amazon

    If your infrastructure runs on AWS, the A10G is the path of least resistance. G5 instances with A10G GPUs are available in every major AWS region, integrate natively with SageMaker endpoints, ECS, and EKS, and spot availability is excellent. At $1.20/hr on spot, you get a production-grade inference GPU without leaving the AWS ecosystem.

  2. Serving 13B-14B code and instruction models

    The 24 GB VRAM sweet spot handles models like Code Llama 13B, Qwen 2.5 14B, and Phi-3 Medium 14B in FP16 or INT8. For development teams running self-hosted code completion, the A10G provides the right balance of memory capacity and cost — enough VRAM for 14B models, enough bandwidth for interactive latency.

  3. Multi-model serving with shared GPU resources

    Deploy multiple smaller models (3B-7B) on a single A10G using frameworks like vLLM with model sharding. The 24 GB VRAM can hold two 7B models simultaneously or four 3B models, with GPU multiplexing handling request routing. This consolidation cuts per-model costs below $0.30/hr.

FAQ

Common questions.

A10G vs RTX 3090 — why is the A10G more expensive?

Both have 24 GB VRAM, but the A10G costs $1.20/hr vs $0.10/hr for the RTX 3090. The price difference comes from the provider ecosystem, not the hardware. A10G pricing is set by AWS and major clouds with enterprise SLAs, while RTX 3090 spot instances come from GPU-specialized providers. For raw inference performance, they're comparable — the RTX 3090 actually has higher bandwidth (936 vs 600 GB/s). If your workload doesn't need AWS-specific integrations, the RTX 3090 is a better value.

Can the A10G run 70B models?

Not on a single card. 70B in INT4 needs ~35 GB, which exceeds the 24 GB VRAM. You'd need multi-GPU A10G instances (AWS offers up to 4x A10G in g5.12xlarge), but inter-GPU communication over PCIe is slow. For 70B models, a single A100 80GB or L40S 48 GB is a better fit.

Why does A10G spot pricing vary from $1.20 to $6.64?

The wide spread reflects the difference between GPU-specialized providers (Vast.ai, RunPod at the low end) and major cloud providers (AWS, GCP at the high end). AWS G5 spot pricing in popular regions (us-east-1, us-west-2) tends to be $1.20-$2.00/hr, while on-demand is $4-6/hr. The $6.64 represents peak-demand spot pricing in constrained regions.

Is the A10G being replaced by the L4?

In Google Cloud, yes — the L4 is GCP's recommended inference GPU and has effectively replaced the T4. In AWS, the A10G remains the standard for G5 instances, with no announced successor. The L4 offers better performance-per-watt (72W vs 150W) but less memory bandwidth (300 vs 600 GB/s). They coexist across providers rather than one replacing the other.

Rent a A10G. Right now.

Spot pricing, per-second billing, no commitment.

Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.