GDDR6 memory
AIVory · GPU Marketplace
Live spot pricingNVIDIA T4 — the original cloud inference GPU, now dirt cheap.
The T4 defined what cloud inference looks like. Launched in 2018 as the first GPU designed specifically for data center inference, its 16 GB of GDDR6 and 70-watt TDP made it the default choice across AWS, GCP, and Azure. At $0.18/hr on spot, the T4 remains the absolute cheapest way to get tensor-core-accelerated inference in the cloud.
At a glance
T4 specifications.
Key hardware specs that determine what workloads this GPU handles.
peak throughput
thermal design power
NVIDIA GPU architecture
Spot pricing
T4: live hourly rates.
Every provider offering this GPU on the spot market, sorted cheapest first.
Prices in USD per GPU-hour · spot instances · sorted cheapest first
Recommended models
AI models that run well on T4.
Tested model-GPU pairings with notes on why each is a good fit.
Use cases
What the T4 is built for.
-
Lowest-cost self-hosted inference for 3B-7B models
If your model is under 7B parameters and you want the absolute cheapest GPU-accelerated inference, the T4 at $0.18/hr is hard to beat. A month of 24/7 serving costs ~$130 — less than most API subscriptions, with no rate limits and full control over your data. The T4's Turing tensor cores handle FP16 and INT8 matmuls, so quantized 7B models run at respectable throughput.
-
Embedding generation and vector database indexing
Generating embeddings with models like E5, BGE, or Jina Embeddings doesn't need the latest hardware — it needs cheap, reliable GPU hours. Spin up T4 spot instances at $0.18/hr and churn through millions of documents. The 16 GB VRAM handles large embedding models, and the low power draw means providers keep plenty of T4 supply available.
-
Model prototyping before scaling to production hardware
Test your inference pipeline, measure latency baselines, and validate model serving configurations on the cheapest available GPU. If your model works on a T4, it works everywhere. Scale up to L4 or A100 for production throughput once the pipeline is proven. The T4 is your $0.18/hr sandbox.
FAQ
Common questions.
T4 vs V100 — which is better for inference?
Different strengths. The V100 has 32 GB HBM2 with 900 GB/s bandwidth — better for models between 16-32 GB. The T4 has 16 GB GDDR6 at 320 GB/s — slower bandwidth but native INT8 tensor cores that the V100 lacks. For sub-7B models in INT8, the T4 matches or beats the V100 on throughput despite the lower bandwidth. For larger models or FP16 workloads, the V100 wins. Pricing is similar ($0.18 vs $0.19/hr).
Is the T4 too old for modern LLM workloads?
For small models, no. The T4's Turing architecture supports the tensor operations that modern LLM frameworks need. vLLM, TGI, and Triton all run on the T4. The limitations are practical: 16 GB VRAM caps you at 7B models in FP16 (13B in INT4), and the 320 GB/s bandwidth limits token generation speed. For production serving of 7B models at moderate traffic, the T4 still delivers.
Why is the T4 only available from a few providers?
Most major cloud providers have moved their T4 fleets to reserved and on-demand pricing tiers, where enterprise customers use them for legacy workloads. Spot availability is concentrated on GPU-specialized providers like Vast.ai, RunPod, and Lambda, where T4s are plentiful from depreciated data center hardware. The limited spot supply keeps the price pinned at $0.18/hr — enough providers to be reliable, not so many that you're spoiled for choice.
Can I use the T4 for training?
For fine-tuning small models (under 3B), yes. LoRA fine-tuning of 7B models is possible with gradient checkpointing and INT8 mixed precision. Full fine-tuning is limited to models under ~5B parameters due to the 16 GB VRAM constraint. For any serious training, the A100 or H100 is a better investment — the T4 was designed for inference, and that's where it shines.
Rent a T4. Right now.
Spot pricing, per-second billing, no commitment.
Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.