AIVory  ·  GPU Marketplace

Live spot pricing

NVIDIA L4 — the cloud-native inference card that sips power.

Designed from the ground up for always-on inference at scale. The L4 packs 24 GB of GDDR6 and Ada Lovelace tensor cores into a 72-watt envelope — roughly one-fifth the power draw of an A100. At $0.37/hr on spot, it slots between the legacy T4 and the premium L40S as the sweet spot for serving 7B-13B models in production.

At a glance

L4 specifications.

Key hardware specs that determine what workloads this GPU handles.

24GB
VRAM

GDDR6 memory

300 GB/s
Memory Bandwidth

peak throughput

72W
TDP

thermal design power

Ada Lovelace
Architecture

NVIDIA GPU architecture

Spot pricing

L4: live hourly rates.

Every provider offering this GPU on the spot market, sorted cheapest first.

Loading spot prices…

Prices in USD per GPU-hour · spot instances · sorted cheapest first

Use cases

What the L4 is built for.

  1. Always-on inference endpoints with minimal power overhead

    The L4's 72W TDP is less than a desktop light bulb. For teams running inference 24/7, power cost is a real factor — an L4 serving Gemma 2 9B costs $0.37/hr in compute and pennies in electricity, while an A100 doing the same job draws 400W and costs $1.14/hr. The L4 wins on total cost of ownership for workloads under 13B parameters.

  2. Google Cloud inference deployments and Vertex AI

    The L4 is Google Cloud's default inference accelerator, making it the natural choice for teams already on GCP. Spot L4 instances integrate with Vertex AI endpoints, Cloud Run GPU, and GKE autopilot. If your stack is GCP-native, the L4 is the path of least resistance for GPU inference.

  3. Video and image generation with Stable Diffusion

    Beyond LLMs, the L4's Ada Lovelace architecture excels at image generation workloads. Stable Diffusion XL runs at 15+ images/second on the L4, and the 24 GB VRAM handles high-resolution outputs (2048x2048) without tiling. At $0.37/hr, batch image generation is remarkably affordable.

FAQ

Common questions.

L4 vs T4 — why pay more for the L4?

The L4 costs $0.37/hr vs $0.18/hr for the T4, but delivers roughly 3x the inference throughput on transformer models. Ada Lovelace tensor cores are two generations ahead of Turing, and the L4 supports INT8/FP8 natively while the T4 is limited to INT8/FP16. If you're serving enough traffic that a T4 runs at capacity, two T4s ($0.36/hr) deliver less throughput than one L4 ($0.37/hr).

Can the L4 run 70B models?

Not on a single card. Llama 3.3 70B in INT4 needs ~35 GB, which exceeds the L4's 24 GB. For 70B inference, look at the A100 80GB, L40S (48 GB), or H100. The L4 is optimized for 3B-13B models, where it offers the best performance-per-watt in NVIDIA's lineup.

Why does the L4 spot price range from $0.37 to $10.52?

The wide price spread reflects different providers and regions. The $0.37 rate comes from GPU-specialized providers like Vast.ai and RunPod, while the $10.52 rate is from major cloud providers (AWS, GCP) where spot pricing is closer to on-demand. Always sort by price and check availability — the cheap end of the range is where the value is.

Is the L4 just a cut-down L40S?

They share the Ada Lovelace architecture, but the L4 is a different chip (AD104 vs AD102). The L4 has fewer CUDA cores, half the VRAM (24 vs 48 GB), and one-third the power draw (72W vs 350W). Think of the L4 as the inference-focused derivative and the L40S as the full-power workstation/data center variant. For models under 24 GB, the L4 is the more cost-effective choice.

Rent a L4. Right now.

Spot pricing, per-second billing, no commitment.

Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.