AIVory  ·  GPU Marketplace

Live spot pricing

NVIDIA RTX 3080 — Ampere inference on a shoestring.

The RTX 3080 delivers Ampere tensor cores with 10-12 GB of GDDR6X at $0.17/hr — matching the cheapest professional cards at a fraction of the usual data center cost. VRAM is the bottleneck: 10-12 GB limits you to 7B models in FP16 or 12B in INT4, but for those workloads the 760 GB/s bandwidth keeps token generation fast.

At a glance

RTX 3080 specifications.

Key hardware specs that determine what workloads this GPU handles.

12GB
VRAM

GDDR6X memory

760 GB/s
Memory Bandwidth

peak throughput

320W
TDP

thermal design power

Ampere
Architecture

NVIDIA GPU architecture

Spot pricing

RTX 3080: live hourly rates.

Every provider offering this GPU on the spot market, sorted cheapest first.

Loading spot prices…

Prices in USD per GPU-hour · spot instances · sorted cheapest first

Use cases

What the RTX 3080 is built for.

  1. Sub-$0.20/hr inference for 3B-7B models

    The RTX 3080 at $0.17/hr is tied with the RTX A4000 as the cheapest Ampere GPU on the spot market. For 3B-7B models in INT8, the 760 GB/s bandwidth generates tokens faster than both the A4000 (448 GB/s) and T4 (320 GB/s). If your model fits in 10-12 GB, the RTX 3080 offers the best throughput-per-dollar in the Ampere generation.

  2. Batch embedding generation at scale

    Text embedding models like E5, BGE, and Jina typically use 1-2 GB of VRAM, leaving 8-10 GB on the RTX 3080 for large batch processing. At $0.17/hr, generating embeddings for millions of documents is extraordinarily affordable — spin up 10 RTX 3080 instances for $1.70/hr total and index an entire document corpus overnight.

  3. Student and researcher experimentation on real GPUs

    At $0.17/hr, the RTX 3080 costs $4/day for 24-hour access to Ampere tensor cores. Students learning ML inference, researchers prototyping quantization strategies, and hobbyists experimenting with local LLMs can access real GPU hardware at coffee-money pricing. The 10-12 GB VRAM forces you to learn quantization early — a useful skill.

FAQ

Common questions.

RTX 3080 10 GB vs 12 GB — which do I get?

The 12 GB variant adds 2 GB of VRAM and slightly more CUDA cores. For 7B models in FP16 (~14 GB), neither variant fits — you need INT8 or INT4. For 7B in INT8 (~7 GB), both work but the 12 GB variant gives more KV cache room. For 3B-4B models in FP16, both work comfortably. Spot listings don't always specify the variant — check the VRAM reported in the instance details.

RTX 3080 vs RTX 3090 — worth paying extra?

The RTX 3090 costs $0.10/hr with 24 GB — actually cheaper than the RTX 3080's $0.17/hr. This pricing inversion exists because the 3090 is more commonly available on spot markets. If both are available, the RTX 3090 is strictly better: more VRAM, higher bandwidth (936 vs 760 GB/s), and lower spot pricing. The RTX 3080 only makes sense when RTX 3090 supply is exhausted.

Is the RTX 3080 worth using when the RTX 3090 is cheaper?

In most cases, no — grab the 3090 if it's available and cheaper. The RTX 3080 exists on spot markets as surplus gaming hardware. Its value is availability: when 3090 instances are fully claimed, the RTX 3080 is your next-cheapest Ampere option. Think of it as overflow capacity, not a first choice.

Can the RTX 3080 run 12B-13B models?

In INT4 quantization, a 13B model needs ~7 GB — fits on both variants with room for KV cache. In INT8, it needs ~13 GB — only fits on the 12 GB variant with zero headroom (impractical). In FP16, it needs ~26 GB — far too large. For 12B-13B inference on the RTX 3080, INT4 is the only practical option, and output quality is noticeably lower than INT8 or FP16.

Rent a RTX 3080. Right now.

Spot pricing, per-second billing, no commitment.

Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.