AIVory  ·  GPU Marketplace

Live spot pricing

NVIDIA RTX A2000 — the cheapest GPU on the spot market, period.

At $0.12/hr, the RTX A2000 is the absolute lowest-cost GPU you can rent for tensor-core-accelerated compute. Six gigabytes of GDDR6 and Ampere architecture in a compact, 70-watt single-slot form factor. VRAM limits you to 2B-3B models, but for embedding generation, classification, and tiny model inference, nothing else comes close on price.

At a glance

RTX A2000 specifications.

Key hardware specs that determine what workloads this GPU handles.

6GB
VRAM

GDDR6 memory

288 GB/s
Memory Bandwidth

peak throughput

70W
TDP

thermal design power

Ampere
Architecture

NVIDIA GPU architecture

Spot pricing

RTX A2000: live hourly rates.

Every provider offering this GPU on the spot market, sorted cheapest first.

Loading spot prices…

Prices in USD per GPU-hour · spot instances · sorted cheapest first

Use cases

What the RTX A2000 is built for.

  1. Text embedding and vector indexing at the lowest GPU cost

    Embedding models like E5-small, MiniLM, and all-MiniLM-L6 use 0.2-0.5 GB of VRAM. The RTX A2000's 6 GB handles large batch embedding with 5+ GB free for input buffers. At $0.12/hr, indexing millions of documents for vector search costs $2.88/day — less than a cup of coffee for enterprise-scale RAG infrastructure.

  2. Classification and NER with compact transformer models

    Named entity recognition, sentiment analysis, and text classification with DistilBERT, DeBERTa-v3-small, or custom fine-tuned models require 0.3-1 GB of VRAM. The RTX A2000 processes these at Ampere tensor-core speeds for $0.12/hr, making it viable to run GPU-accelerated NLP pipelines that would otherwise fall back to CPU for cost reasons.

  3. GPU compute testing in CI/CD pipelines

    Automated tests that validate CUDA compatibility, model loading, and inference pipeline correctness need the cheapest available GPU. The RTX A2000 at $0.12/hr is perfect for CI/CD GPU runners — run your test suite on real tensor core hardware at $2.88/day, catching GPU-specific bugs that CPU-only CI misses.

FAQ

Common questions.

Can the RTX A2000 run any useful LLM?

Yes, but only very small ones. Gemma 2 2B in FP16 (4 GB) and TinyLlama 1.1B (2.2 GB) both fit. Llama 3.2 3B in FP16 needs 6 GB — it fits but with zero headroom for KV cache, making it impractical for multi-turn conversations. For any model above 3B, you need more VRAM. The RTX A2000 excels at embedding, classification, and sub-2B inference.

RTX A2000 6 GB vs 12 GB — which am I getting?

NVIDIA released both a 6 GB and 12 GB variant. Spot listings should specify the VRAM size. The 12 GB variant opens up 3B-7B models in INT8, which dramatically expands the card's usefulness. At $0.12/hr, either variant is a bargain — but check the listing carefully before renting.

Is the RTX A2000 faster than a CPU for inference?

For transformer models, significantly. Even a high-end 64-core CPU running llama.cpp serves Gemma 2 2B at 10-15 tokens/second. The RTX A2000's Ampere tensor cores deliver 30-50+ tokens/second on the same model. The GPU advantage is smaller than with larger cards (fewer tensor cores, lower bandwidth), but the price-performance ratio still favors the GPU for any workload that fits in 6 GB.

Why is the RTX A2000 cheaper than every other GPU?

Three factors: tiny chip (GA106), low VRAM (6 GB GDDR6), and market positioning. NVIDIA designed the A2000 as the entry-level professional card for CAD and light compute. On spot markets, these cards come from workstations being retired or upgraded — providers price them at $0.12/hr because demand for 6 GB GPUs is low. This makes them a hidden gem for workloads that don't need more VRAM.

Rent a RTX A2000. Right now.

Spot pricing, per-second billing, no commitment.

Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.