Compare · prices checked 2026-09-03
L40S vs A100 PCIE: specs, price per hour, which to rent
NVIDIA L40S (48 GB, $0.466/hr) against NVIDIA A100 PCIE (80 GB, $0.662/hr): public specs side by side, live fixed prices from our sheet, what fits in each card's VRAM, and a verdict written for real workloads — both are rentable right now.
L40S vs A100 PCIE specifications
Public NVIDIA figures (dense, non-sparsity). The last column is L40S relative to A100 PCIE.
| Spec | L40S | A100 PCIE | Difference |
|---|---|---|---|
| Architecture | Ada Lovelace (2023) | Ampere (2021) | — |
| VRAM | 48 GB GDDR6 | 80 GB HBM2e | −40% |
| Memory bandwidth | 864 GB/s | 1,935 GB/s | −55% |
| FP16 tensor (dense) | 362 TFLOPS | 312 TFLOPS | +16% |
| FP32 | 91.6 TFLOPS | 19.5 TFLOPS | +370% |
| CUDA cores | 18,176 | 6,912 | +163% |
| TDP | 350 W | 300 W | +17% |
| PCIe · NVLink | Gen 4.0 · no NVLink | Gen 4.0 · no NVLink | — |
| PowerScore (RTX 3090 = 100) | 255 | 220 | +16% |
| Max GPUs per machine | 8× | 8× | — |
L40S vs A100 PCIE price per hour
Fixed rates from our sheet — every on-demand price is the marketplace median × 0.70, rounded down. Monthly = 730 hours.
| Rate | L40S | A100 PCIE | Cheaper |
|---|---|---|---|
| On-demand, per GPU-hour | $0.466 | $0.662 | L40S (−30%) |
| Interruptible, per GPU-hour | $0.233 | $0.331 | L40S |
| Reserved (3 mo), per GPU-hour | $0.302 | $0.430 | L40S |
| On-demand, per month | $340 | $483 | L40S |
| Market median (reference) | $0.67 | $0.95 | — |
| $ per 1,000 FP16 TFLOP-hours | $1.29 | $2.12 | L40S (better value) |
| $ per GB of VRAM per hour | $0.0097 | $0.0083 | A100 PCIE (better value) |
Try a full month with storage and bandwidth in the GPU cost calculator.
What fits in VRAM: 48 GB vs 80 GB
| Workload | L40S | A100 PCIE |
|---|---|---|
| Largest LLM in FP16, one card | ~14B | ~32B |
| Largest LLM at 4-bit, one card | ~72B | ~123B |
| Flux dev (FP8, ~17 GB) | fits | fits |
| Wan 2.x 14B video (offloaded) | yes | yes |
| 70B 4-bit LLM on one card | yes | yes |
Rules of thumb: ~2.4 GB per billion parameters in FP16 all-in, ~0.62 GB in 4-bit. Full tables in the VRAM guide.
Verdict: which should you rent?
The L40S (48 GB GDDR6, Ada, FP8) is faster for FP16/FP8 inference and image generation; the A100 PCIe (80 GB HBM2e) has more memory, more bandwidth and FP64. Pick the L40S for diffusion and ≤32B serving, the A100 for 70B models and training.
- Cheaper per hour: L40S ($0.466 vs $0.662, −30%).
- More VRAM: A100 PCIE (80 GB vs 48 GB).
- More FP16 throughput: L40S (about 1.2×).
- Best value per TFLOP-hour: L40S.
- Best value per GB of VRAM: A100 PCIE.
- Multi-GPU: L40S over PCIe · A100 PCIE over PCIe.
Related comparisons
- H100 PCIE vs A100 PCIE $2.147 vs $0.662 per hour Compare
- A100 SXM4 vs A100 PCIE $0.583 vs $0.662 per hour Compare
- RTX 4090 vs A100 PCIE $0.262 vs $0.662 per hour Compare
- RTX 4090 vs L40S $0.262 vs $0.466 per hour Compare
- L40S vs H100 PCIE $0.466 vs $2.147 per hour Compare
- RTX 5090 vs L40S $0.318 vs $0.466 per hour Compare
L40S vs A100 PCIE: FAQ
Deeper reading: H100 vs H200 vs B200, RTX 4090 vs RTX 5090, how cloud GPU pricing works.
Is the L40S faster than the A100 PCIE?
On dense FP16 tensor throughput the L40S leads by about 1.2× (362 vs 312 TFLOPS). Memory bandwidth matters as much for inference: L40S 864 GB/s vs A100 PCIE 1,935 GB/s.
Which is cheaper to rent, the L40S or the A100 PCIE?
The L40S: $0.466/hr on-demand versus $0.662/hr — 30% less. Interruptible rates are $0.233 (L40S) and $0.331 (A100 PCIE). Per TFLOP-hour the better value is the L40S.
Which has more VRAM and what does that change?
The A100 PCIE has 80 GB versus 48 GB. In LLM terms that is roughly a 32B FP16 model (or ~123B in 4-bit) on one card against 14B FP16 (~72B 4-bit). If the model does not fit, speed is irrelevant.
Can I rent both on PowerGPU right now?
Yes — 59 × L40S and 55 × A100 PCIE are online as this page renders, deployable in about 30 seconds, billed per second, paid in crypto with no KYC.