Side by side
| RTX 4090 | RTX 5090 | |
|---|---|---|
| Architecture | Ada Lovelace | Blackwell |
| VRAM | 24 GB GDDR6X | 32 GB GDDR7 |
| Memory bandwidth | 1.0 TB/s | 1.79 TB/s |
| FP16 tensor (dense) | 330 TF | 419 TF |
| Rent, on-demand | $0.262/hr | $0.318/hr |
| Rent, interruptible | $0.131/hr | $0.159/hr |
The premium is $0.056/hr (21%) today. The question is always: does the job use the extra 8 GB or the extra bandwidth? If not, the 4090's price wins.
LLM inference math
- 8B class (fits both) — the 5090's bandwidth pushes ~1.6–1.8× the tokens/s; per token that beats its 1.21× price. Win: 5090, narrowly.
- 14B class — FP16 needs ~31 GB: 5090 serves it native, 4090 must quantize. Win: 5090.
- 32B 4-bit (~21 GB weights) — loads on both, but KV-cache drowns the 4090 at real concurrency. Win: 5090, decisively.
- Batch embeddings / Whisper — compute-bound small models: rent whichever is cheaper per hour that day; it is usually the 4090.
Image & video generation
SDXL at 1024px is not VRAM-bound: the 4090's ~25–30% speed deficit is smaller than its 18% price advantage — more images per dollar on the 4090. Flux dev flips it: FP16 weights + text encoders brush against 24 GB, and every offload event stalls the 4090 while the 5090 keeps everything resident. Video models (Wan, LTX) are 5090-or-bigger territory — see the video playbook.
Fine-tuning
QLoRA ceilings: ~13B on 24 GB, ~34B on 32 GB. If your target model is Qwen 32B, the 5090 is the cheapest single-card trainer on the sheet; at 8B, the 4090 (or even a 3090 at $0.104/hr) does the same epochs for less. Walkthrough with live consumption numbers: QLoRA guide.
The verdict table
| Job | Rent | Why |
|---|---|---|
| SDXL batches | 4090 | images/$ wins |
| Flux heavy workflows | 5090 | stays resident in 32 GB |
| 7–8B serving | 5090 | bandwidth → tokens/$ |
| 32B 4-bit serving | 5090 | only one with cache headroom |
| 8–13B QLoRA | 4090 | same result, lower rate |
| 34B QLoRA | 5090 | single-card ceiling |
| Video generation | 5090+ | 24 GB is below entry |
Undecided? Rent both for an hour — $0.580 total on-demand — and benchmark your actual workload. That experiment costs less than this article took to read.
Put the numbers to work
Every price in this guide is our live rate — fixed, ≥30% under the market median, billed per second. Deploy the exact setup above from the console in about 30 seconds, paid in crypto, no card and no KYC.


