AIVory  ·  GPU Marketplace

Live spot pricing

NVIDIA L40 — 48 GB of Ada muscle at half the L40S price.

The original L40 delivers the same 48 GB VRAM and Ada Lovelace architecture as the L40S, without the inference-specific optimizations — and at a lower price point. Starting at $0.51/hr on spot, the L40 runs 30B-70B quantized models comfortably and handles mixed workloads that combine inference with graphics or video encoding.

At a glance

L40 specifications.

Key hardware specs that determine what workloads this GPU handles.

48GB
VRAM

GDDR6X memory

864 GB/s
Memory Bandwidth

peak throughput

300W
TDP

thermal design power

Ada Lovelace
Architecture

NVIDIA GPU architecture

Spot pricing

L40: live hourly rates.

Every provider offering this GPU on the spot market, sorted cheapest first.

Loading spot prices…

Prices in USD per GPU-hour · spot instances · sorted cheapest first

Use cases

What the L40 is built for.

  1. Serving 27B-31B models in full precision

    The L40's 48 GB VRAM holds Qwen 3.5 27B or Gemma 4 31B in FP16 without quantization. At $0.51/hr, this is the cheapest way to serve a 30B-class model at full precision — the L40S costs $0.57/hr for the same VRAM, and the A100 80GB starts at $1.14/hr with overkill capacity.

  2. Mixed inference and rendering pipelines

    Unlike the inference-focused L40S, the original L40 retains full graphics capabilities including ray tracing and NVENC/NVDEC. Teams running AI-powered video processing, 3D rendering with neural denoising, or real-time graphics applications benefit from a single GPU that handles both compute and display output.

  3. Staging environment for production A100 workloads

    If your production fleet runs on A100 80GB, the L40's 48 GB VRAM and similar throughput characteristics make it an affordable staging proxy. Deploy quantized versions of your production models on L40 spot instances to validate serving infrastructure, load balancing, and client integrations before committing to A100 on-demand pricing.

FAQ

Common questions.

What is the difference between the L40 and L40S?

Same Ada Lovelace chip (AD102), same 48 GB VRAM, same 864 GB/s bandwidth. The L40S adds hardware-level optimizations for transformer inference (enhanced INT8/FP8 tensor core scheduling) and disables the display output and some graphics features. In practice, the L40S delivers 10-15% higher LLM inference throughput. The L40 is cheaper ($0.51 vs $0.57/hr) and retains full graphics capabilities, making it better for mixed workloads.

Can the L40 run Llama 3.3 70B?

In INT4 quantization (GPTQ or AWQ), yes. The quantized model needs ~35 GB, leaving 13 GB for KV cache. For production serving at high throughput, this is tight — consider the A100 80GB instead. But for development, testing, and low-traffic endpoints, the L40 handles 70B INT4 at $0.51/hr.

Why choose the L40 over the cheaper A6000?

Both have 48 GB VRAM, but the L40's Ada Lovelace architecture delivers roughly 50% higher tensor throughput than the A6000's Ampere architecture. The L40 also supports FP8 and has better INT8 performance. At $0.51/hr vs $0.33/hr for the A6000, you pay ~55% more for ~50% more speed — a wash on cost-per-token, but the L40 serves requests at lower latency.

Is the L40 good for training or just inference?

The L40 works for fine-tuning models up to ~20B parameters. LoRA fine-tuning of Llama 3.3 70B in INT4 is possible with careful memory management. For full fine-tuning of large models, the 48 GB VRAM limits you to ~15B parameters in BF16 with optimizer states. The L40 is primarily an inference and mixed-workload card — for serious training, the A100 80GB or H100 offers better memory and bandwidth.

Rent a L40. Right now.

Spot pricing, per-second billing, no commitment.

Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.