GDDR6X memory
AIVory · GPU Marketplace
Live spot pricingNVIDIA L40S — purpose-built for inference.
48 GB of GDDR6X on the Ada Lovelace architecture at half the power draw of an A100. The L40S handles every quantized model from 7B to 70B and was designed from the ground up for always-on inference workloads where reliability and efficiency beat raw training throughput.
At a glance
L40S specifications.
Key hardware specs that determine what workloads this GPU handles.
peak throughput
thermal design power
NVIDIA GPU architecture
Spot pricing
L40S: live hourly rates.
Every provider offering this GPU on the spot market, sorted cheapest first.
Prices in USD per GPU-hour · spot instances · sorted cheapest first
Recommended models
AI models that run well on L40S.
Tested model-GPU pairings with notes on why each is a good fit.
Use cases
What the L40S is built for.
-
Production inference for quantized 24B-70B models
The L40S sits in the sweet spot between consumer GPUs (24 GB) and the A100 (80 GB). Its 48 GB comfortably runs INT4/INT8 quantized versions of models up to 72B parameters, while the Ada Lovelace tensor cores deliver competitive inference latency. For teams that need more VRAM than an RTX 4090 but not the full expense of an A100, the L40S is the answer.
-
Power-efficient always-on deployments
At 350W TDP versus the A100's 400W and the H100's 700W, the L40S draws less power for every token generated. In a colocation or on-premises setup where you pay per kilowatt-hour, this translates directly into lower operating costs. For inference services that run 24/7, the power savings compound into meaningful budget reductions over months.
-
Mixed inference and video encoding pipelines
The L40S inherits Ada Lovelace's NVENC/NVDEC engines alongside its tensor cores. Workloads that combine LLM inference with real-time video processing — think AI-powered video analytics, captioning pipelines, or interactive media generation — can run both tasks on a single L40S instead of splitting across GPU types.
FAQ
Common questions.
L40S vs A100 — when is the L40S the better choice?
Choose the L40S when your workload is inference-only and your model fits in 48 GB. The L40S typically costs 40-60% less per hour than an A100 on spot markets, draws less power, and delivers competitive inference throughput on Ada Lovelace tensor cores. Choose the A100 when you need 80 GB for larger models, when you need NVLink for multi-GPU scaling, or when your workflow includes training.
Why does the L40S use GDDR6X instead of HBM?
HBM (High Bandwidth Memory) is expensive and primarily benefits training workloads where raw memory bandwidth is the bottleneck. Inference workloads are often compute-bound or latency-bound rather than bandwidth-bound, especially at typical batch sizes. GDDR6X provides enough bandwidth for inference at dramatically lower manufacturing cost, which is why L40S spot prices are a fraction of A100 prices.
Which popular models fit in the L40S's 48 GB VRAM?
In FP16: all models up to ~24B (Mistral Small, Gemma 2 27B). In INT8: models up to ~45B (Gemma 4 31B, CodeLlama 34B). In INT4/GPTQ: models up to ~72B (Qwen 2.5 72B, Llama 3.3 70B). The MoE model Mixtral 8x7B (46.7B total) also fits comfortably. Anything above 72B requires multi-GPU or an 80GB+ card.
Can the L40S handle multi-modal inference workloads?
Yes. The L40S's Ada Lovelace architecture includes hardware acceleration for both tensor operations and media encoding. Vision-language models like LLaVA, multi-modal Gemma, and Qwen-VL run well on the L40S. The dedicated NVENC/NVDEC engines handle video preprocessing without competing for tensor core resources, making it efficient for pipelines that mix text, image, and video inference.
Rent a L40S. Right now.
Spot pricing, per-second billing, no commitment.
Browse the live marketplace, pick your GPU, deploy in one click. Credits from $10.