GDDR6 memory
AIVory · GPU Marketplace
Live spot pricingNVIDIA RTX A2000 — the cheapest GPU on the spot market, period.
At $0.12/hr, the RTX A2000 is the absolute lowest-cost GPU you can rent for tensor-core-accelerated compute. Six gigabytes of GDDR6 and Ampere architecture in a compact, 70-watt single-slot form factor. VRAM limits you to 2B-3B models, but for embedding generation, classification, and tiny model inference, nothing else comes close on price.
At a glance
RTX A2000 specifications.
Key hardware specs that determine what workloads this GPU handles.
peak throughput
thermal design power
NVIDIA GPU architecture
Spot pricing
RTX A2000: live hourly rates.
Every provider offering this GPU on the spot market, sorted cheapest first.
Prices in USD per GPU-hour · spot instances · sorted cheapest first
Recommended models
AI models that run well on RTX A2000.
Tested model-GPU pairings with notes on why each is a good fit.
Use cases
What the RTX A2000 is built for.
-
Text embedding and vector indexing at the lowest GPU cost
Embedding models like E5-small, MiniLM, and all-MiniLM-L6 use 0.2-0.5 GB of VRAM. The RTX A2000's 6 GB handles large batch embedding with 5+ GB free for input buffers. At $0.12/hr, indexing millions of documents for vector search costs $2.88/day — less than a cup of coffee for enterprise-scale RAG infrastructure.
-
Classification and NER with compact transformer models
Named entity recognition, sentiment analysis, and text classification with DistilBERT, DeBERTa-v3-small, or custom fine-tuned models require 0.3-1 GB of VRAM. The RTX A2000 processes these at Ampere tensor-core speeds for $0.12/hr, making it viable to run GPU-accelerated NLP pipelines that would otherwise fall back to CPU for cost reasons.
-
GPU compute testing in CI/CD pipelines
Automated tests that validate CUDA compatibility, model loading, and inference pipeline correctness need the cheapest available GPU. The RTX A2000 at $0.12/hr is perfect for CI/CD GPU runners — run your test suite on real tensor core hardware at $2.88/day, catching GPU-specific bugs that CPU-only CI misses.
FAQ
Common questions.
Can the RTX A2000 run any useful LLM?
Yes, but only very small ones. Gemma 2 2B in FP16 (4 GB) and TinyLlama 1.1B (2.2 GB) both fit. Llama 3.2 3B in FP16 needs 6 GB — it fits but with zero headroom for KV cache, making it impractical for multi-turn conversations. For any model above 3B, you need more VRAM. The RTX A2000 excels at embedding, classification, and sub-2B inference.
RTX A2000 6 GB vs 12 GB — which am I getting?
NVIDIA released both a 6 GB and 12 GB variant. Spot listings should specify the VRAM size. The 12 GB variant opens up 3B-7B models in INT8, which dramatically expands the card's usefulness. At $0.12/hr, either variant is a bargain — but check the listing carefully before renting.
Is the RTX A2000 faster than a CPU for inference?
For transformer models, significantly. Even a high-end 64-core CPU running llama.cpp serves Gemma 2 2B at 10-15 tokens/second. The RTX A2000's Ampere tensor cores deliver 30-50+ tokens/second on the same model. The GPU advantage is smaller than with larger cards (fewer tensor cores, lower bandwidth), but the price-performance ratio still favors the GPU for any workload that fits in 6 GB.
Why is the RTX A2000 cheaper than every other GPU?
Three factors: tiny chip (GA106), low VRAM (6 GB GDDR6), and market positioning. NVIDIA designed the A2000 as the entry-level professional card for CAD and light compute. On spot markets, these cards come from workstations being retired or upgraded — providers price them at $0.12/hr because demand for 6 GB GPUs is low. This makes them a hidden gem for workloads that don't need more VRAM.
Rent a RTX A2000. Right now.
Spot pricing, per-second billing, no commitment.
Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.