Use case · image generation
Image generation GPUs: 6,246 images per dollar
ComfyUI and Forge boot in ~30 seconds on cards the community actually tunes for: RTX 4090 at $0.262/hr. Generate interactively, batch overnight on interruptible at half price, and stop paying the second the queue empties.
The image cards, ranked by images-per-dollar
| Tier | GPU | VRAM | On-demand | Interruptible | Why this card | Action |
|---|---|---|---|---|---|---|
| Good | RTX 3090 | 24 GB | $0.104 | $0.052 | 24 GB at the lowest price that runs everything — the batch-farm workhorse. | Deploy |
| Better | RTX 4090 | 24 GB | $0.262 | $0.131 | The community default: fastest ecosystem support, ~2.2 s Flux dev images. | Deploy |
| Best | RTX 5090 | 32 GB | $0.318 | $0.159 | 32 GB GDDR7 headroom for full-precision Flux, video nodes and monster upscale chains. | Deploy |
Serving images to users? The 48 GB L40S on a serverless endpoint handles concurrent workflows with server-grade cooling.
Interactive vs batch: two bills
| An evening of prompting | 3 h interactive ComfyUI on a 4090, on-demand | $0.79 |
|---|---|---|
| A 10,000-image batch | ~6.1 h on a 4090, interruptible queue | $0.80 |
| A style LoRA | ~45 min Kohya training run | $0.10 |
| Model library, always warm | 120 GB volume (checkpoints, LoRAs, VAEs) | $9.60/mo |
Set up once, generate forever
- ComfyUI template ships with Manager — install nodes from the browser.
- Models on a volume — 40 GB of checkpoints load in seconds on every new instance.
- Forge for A1111 muscle memory — same extensions, lower VRAM, faster attention.
- API mode — every ComfyUI workflow exports as JSON and runs headless for pipelines.
Moving pictures? The video generation playbook continues from here.
Stable Diffusion & Flux GPUs: FAQ
What does an AI image actually cost to generate?
Flux dev at ~2.2 s/image on an RTX 4090 works out to about $0.00016 per image on-demand — roughly 6,246 images per dollar. Batch on interruptible and it doubles. SDXL is 2–3× faster still.
Which GPU for ComfyUI with Flux?
Flux dev FP8 wants ~17 GB, so 24 GB cards are the entry point: the RTX 4090 at $0.262/hr is the community default. Full-precision Flux, heavy upscaler chains or video nodes benefit from the 32 GB RTX 5090.
How do I train a style LoRA?
The Kohya template: 20–40 images, captions, ~30–60 minutes on a 4090 — a couple of dollars end to end. Keep datasets and outputs on a volume so trainer instances stay disposable.