Guides · 16 articles · numbers included
Cloud GPU guides where the prices are never stale
Every figure in these articles is rendered live from our price sheet — the same data the console bills from. Written by the team that runs the fleet, updated with every weekly market re-check.
All guides 16
Costs, hardware choices and hands-on walkthroughs — every number on these pages is today's sheet.
Costs & pricingCloud GPU pricing explained (2026): on-demand vs spot vs reservedWhat actually drives cloud GPU prices in 2026, how marketplaces and hyperscalers differ, and when each billing mode saves you money.
Choosing hardwareLLM VRAM requirements: how much GPU memory for 7B–405B modelsComplete sizing tables for running and training LLMs — FP16, INT8 and 4-bit, with KV-cache math and the cheapest GPU that fits each model.
Choosing hardwareH100 vs H200 vs B200 (2026): specs, price per hour, which to rentMemory, bandwidth and real rental economics of NVIDIA's three datacenter flagships — and when the older card is the better deal.
Choosing hardwareRTX 5090 vs RTX 4090 for AI (2026): benchmarks, VRAM, rental costSpecs, VRAM, real throughput differences and cost per run for inference, fine-tuning and image generation on both consumer flagships.
Hands-on walkthroughHow to fine-tune Llama 3.1 8B with QLoRA on a single GPUA complete, copy-pasteable walkthrough: dataset to merged weights in about an hour on a single 24 GB card, with axolotl.
Hands-on walkthroughHow to deploy vLLM on a cloud GPU: an OpenAI-compatible endpointDeploy vLLM on a rented GPU, pick the right card for your model size, benchmark tokens per second and put a price on every million tokens.
Costs & pricingCheapest cloud GPU for Stable Diffusion & Flux in 2026 ($/image)What SDXL and Flux actually need, images-per-dollar on eight rentable GPUs, and where paying more per hour costs less per image.
Hands-on walkthroughBlender cloud rendering on GPUs: setup, cost per frame, pitfallsRender Cycles scenes on rented RTX hardware: headless setup, per-frame cost math, and the mistakes that quietly triple a render bill.
Costs & pricingHow much does it cost to rent an H100? Per-hour math for 2026H100 SXM, PCIe and NVL rental prices per hour and per month, what marketplaces and hyperscalers charge, the rent-vs-buy break-even, and five ways to pay less.
Choosing hardwareBest cloud GPU for LLM inference in 2026, by model sizeFrom 8B to 405B: the cheapest rentable card that holds each model, indicative tokens per second with vLLM, and the cost per million tokens on today's sheet.
Hands-on walkthroughHow to run ComfyUI on a cloud GPU: setup, models, cost per imageDeploy ComfyUI on a rented RTX 4090 or 5090 in 30 seconds, keep checkpoints on a volume, run workflows headless through the API, and know what each image costs.
Hands-on walkthroughHow to rent a GPU with crypto and no KYC (2026)Renting cloud GPUs with USDT, Bitcoin or Monero and no identity check: how top-ups work, which coin to use, what data is kept, and the step-by-step deploy.
Choosing hardwareA100 vs H100 for fine-tuning (2026): cost per run, not per hourSame 80 GB, 2.7× the hourly price: when an H100 finishes a LoRA, QLoRA or full fine-tune fast enough to beat the A100 on total cost — live prices, break-even rule.
Costs & pricingCheapest cloud GPU for Ollama (2026): 8B to 70B models, by the hourWhich rented card runs each Ollama model size at Q4, indicative tokens per second per card, and what an always-on private assistant costs per month.
Choosing hardwareWan 2.x video generation: which GPU, minutes per clip, cost per clipVRAM needs for Wan 2.1 and 2.2 (1.3B, 5B, 14B), realistic minutes per five-second clip on RTX 4090, 5090, L40S, H100 and H200, and the price of a batch night.
Costs & pricingBest GPU for Whisper transcription (2026): speed and cost per audio hourfaster-whisper large-v3 throughput on T4, L4, RTX 3060, RTX 4090 and H100 — and the only number that matters: cents per hour of audio transcribed. Looking for reference material instead?
The documentation covers the platform itself; the use-case playbooks pre-match GPUs to workloads; the API reference has curl you can run before signing up.