Price floor Every GPU at least 30% below the market median — re-checked against the marketplace weekly.

See the proof

Datacenter · Ada Lovelace · launched 2023 · prices checked 2026-09-03

Rent NVIDIA L40S — 48 GB, $0.466/hr on-demand

  • VRAM 48 GBGDDR6
  • FP16 tensor 362TFLOPS
  • PowerScore 255RTX 3090 = 100
  • Configs 1–8×PCIe 4.0
  • Online now 59 8 regions

Billed per second, price locked at deploy. Storage $0.08/GB/mo · bandwidth $0.01/GB — the whole fee schedule. Next weekly market re-check: 2026-09-10.

NVIDIA L40S 48 GB cloud GPU for rent

Datacenter · Ada Lovelace architecture

The L40S is NVIDIA's universal datacenter card for Ada: 48 GB GDDR6, 18,176 CUDA cores, FP8 support and passive cooling at 350 W. It is the go-to for 24/7 inference endpoints, SDXL/Flux serving and 32B-class LLMs — server-grade reliability at a consumer-adjacent price.

A server-class card built for sustained 24/7 load: passive cooling in proper chassis, ECC memory, and drivers validated for compute. The sweet spot for inference fleets and fine-tuning jobs that need stability more than headline FLOPS.

NVIDIA L40S specs: VRAM, TFLOPS, bandwidth

GPU modelNVIDIA L40S ArchitectureAda Lovelace (2023)
VRAM48 GB GDDR6 Memory bandwidth864 GB/s
FP16 tensor perf.362 TFLOPS FP32 perf.91.6 TFLOPS
CUDA cores18,176 TDP350 W
PowerScore (RTX 3090 = 100)255 PCIe generationGen 4.0
Multi-GPU1× – 8× Max instance storage8,000 GB NVMe
Network up to5,000 Mbps CUDA12.4 – 13.0

Bandwidth, CUDA cores, TDP and FP32 are public NVIDIA figures; FP16 tensor is the dense (non-sparsity) number. Machine-level values come from live inventory.

L40S price per hour: on-demand, interruptible, reserved

One public rule sets every price on this page: the marketplace median for the L40S ($0.67/hr, snapshot 2026-09-03) × 0.70, rounded down — so on-demand is $0.466, 30% below market. Interruptible halves it; a 3-month reservation takes another 35% off.

ModePer GPU-hour Per day (24 h)Per month (730 h)What you get
On-demand$0.466 $11.18$340 Guaranteed capacity, price locked at deploy, stop anytime
Interruptible$0.233 $5.59$170 Flat −50%; may pause under capacity pressure, disk kept, auto-requeue
Reserved (3 months)$0.302 $7.25$220 −35% on on-demand, rate locked for the term, capacity held

Per GPU: an 8× machine costs exactly 8× — no multi-GPU premium. Estimate a full month with storage and bandwidth in the GPU cost calculator.

What you can run on a L40S (48 GB VRAM)

With 48 GB of GDDR6, a single card holds a ~14B-parameter LLM in FP16 or up to ~72B parameters quantized to 4-bit, with room for KV-cache at practical context lengths. Scale to 8× GPUs on one machine for bigger models or bigger batches — the per-GPU price stays $0.466.

L40S availability by region

59 × L40S across 10 machines, live from inventory:

  • GB London
  • DE Frankfurt
  • US Ashburn, VA
  • US Chicago, IL
  • AU Sydney
  • US Los Angeles, CA
  • SG Singapore
  • US Seattle, WA

L40S vs alternatives: price per TFLOP

GPUVRAMFP16 On-demand$ / TFLOP-hr
L40S this card 48 GB362 $0.466 $1.29‰
A100 PCIE 80 GB312 $0.662 $2.12‰
L40 48 GB181 $0.234 $1.29‰
L4 24 GB121 $0.225 $1.86‰
A10 24 GB125 $0.168 $1.34‰

‰ = dollars per 1,000 TFLOP-hours of FP16 — a rough value-for-compute yardstick across cards.

Renting a L40S: frequently asked questions

How much does it cost to rent an NVIDIA L40S per hour?

$0.466 per GPU-hour on-demand — a fixed price set at least 30% below the current market median of $0.67. Interruptible capacity costs $0.233/hr and a 3-month reservation $0.302/hr. Around $340/month if you keep one running non-stop, billed per second.

What can a L40S with 48 GB VRAM run?

In LLM terms, roughly a 14B-parameter model in FP16 or up to ~72B parameters 4-bit quantized on a single card, with context headroom. Multi-GPU instances (up to 8× on current inventory) multiply that; diffusion and rendering workloads fit comfortably at this VRAM class.

Is the L40S available to rent right now?

Yes — 59 GPUs across 10 machines in 8 regions are listed as we render this page. Configurations go from 1× to 8×. Deploy from the console and it is running in about 30 seconds.

How do I deploy a L40S?

Create an account (email + password, no card, no KYC), top up in crypto, open the console, filter by L40S, pick a machine and a template such as PyTorch, vLLM or ComfyUI. The same deploy is one command with the CLI: powergpu launch --gpu l40s --template pytorch.

Why is the L40S cheaper here than on GPU marketplaces?

We price from the public marketplace median and fix our on-demand rate at least 30% below it, rounded down. The price is re-checked weekly (next check 2026-09-10) and published — no auctions, no per-host roulette, no bidding.

Deploy your first GPU in under a minute

Top up in crypto, benchmark us against your current provider. Per-second billing, fixed prices ≥ 30% below market — cancel by just stopping the instance.