Datacenter flagship · Blackwell · launched 2024 · prices checked 2026-09-03
Rent NVIDIA B200 — 192 GB, $4.204/hr on-demand
- VRAM 192 GBHBM3e
- FP16 tensor 2,250TFLOPS
- PowerScore 1585RTX 3090 = 100
- Configs 1–8×NVLink
- Online now 53 4 regions
Billed per second, price locked at deploy. Storage $0.08/GB/mo · bandwidth $0.01/GB — the whole fee schedule. Next weekly market re-check: 2026-09-10.
Datacenter flagship · Blackwell architecture
The B200 is the Blackwell workhorse: 192 GB HBM3e, 8 TB/s and a second-generation Transformer Engine with FP4. Against an H100 it trains roughly 2–3× faster per GPU and serves large models several times faster, so a B200 hour frequently costs less per token than an H100 hour despite the higher rate.
This is training-grade silicon: HBM3e memory feeding tensor cores at multi-TB/s, NVLink for scaling past one card, and the reliability profile of Tier-III datacenter hosts. Teams rent it for pre-training, long fine-tunes and high-throughput inference where batch size is money.
NVIDIA B200 specs: VRAM, TFLOPS, bandwidth
| GPU model | NVIDIA B200 | Architecture | Blackwell (2024) |
|---|---|---|---|
| VRAM | 192 GB HBM3e | Memory bandwidth | 8,000 GB/s |
| FP16 tensor perf. | 2,250 TFLOPS | FP32 perf. | — |
| CUDA cores | — | TDP | 1000 W |
| PowerScore (RTX 3090 = 100) | 1585 | PCIe generation | Gen 5.0 |
| Multi-GPU | 1× – 8× · NVLink | Max instance storage | 16,000 GB NVMe |
| Network up to | 10,000 Mbps | CUDA | 12.4 – 13.0 |
Bandwidth, CUDA cores, TDP and FP32 are public NVIDIA figures; FP16 tensor is the dense (non-sparsity) number. Machine-level values come from live inventory.
B200 price per hour: on-demand, interruptible, reserved
One public rule sets every price on this page: the marketplace median for the B200 ($6.01/hr, snapshot 2026-09-03) × 0.70, rounded down — so on-demand is $4.204, 30% below market. Interruptible halves it; a 3-month reservation takes another 35% off.
| Mode | Per GPU-hour | Per day (24 h) | Per month (730 h) | What you get |
|---|---|---|---|---|
| On-demand | $4.204 | $100.90 | $3,069 | Guaranteed capacity, price locked at deploy, stop anytime |
| Interruptible | $2.102 | $50.45 | $1,534 | Flat −50%; may pause under capacity pressure, disk kept, auto-requeue |
| Reserved (3 months) | $2.732 | $65.57 | $1,994 | −35% on on-demand, rate locked for the term, capacity held |
Per GPU: an 8× machine costs exactly 8× — no multi-GPU premium. Estimate a full month with storage and bandwidth in the GPU cost calculator.
What you can run on a B200 (192 GB VRAM)
With 192 GB of HBM3e, a single card holds a ~72B-parameter LLM in FP16 or up to ~235B parameters quantized to 4-bit, with room for KV-cache at practical context lengths. Scale to 8× GPUs on one machine with NVLink for bigger models or bigger batches — the per-GPU price stays $4.204.
- Recommended for llm training — H100 SXM from 30%+ under market
- One-click template: vLLM on a B200
- One-click template: PyTorch NGC on a B200
- One-click template: Axolotl — Fine Tuning on a B200
- Sizing help: LLM VRAM requirements guide
B200 availability by region
53 × B200 across 9 machines, live from inventory:
Ashburn, VA
Tokyo
Dallas, TX
Singapore
B200 vs alternatives: price per TFLOP
| GPU | VRAM | FP16 | On-demand | $ / TFLOP-hr |
|---|---|---|---|---|
| B200 this card | 192 GB | 2,250 | $4.204 | $1.87‰ |
| H200 | 141 GB | 990 | $3.058 | $3.09‰ |
| H100 SXM | 80 GB | 990 | $1.587 | $1.60‰ |
| A100 SXM4 | 80 GB | 312 | $0.583 | $1.87‰ |
| B300 | 288 GB | 2,800 | $7.875 | $2.81‰ |
‰ = dollars per 1,000 TFLOP-hours of FP16 — a rough value-for-compute yardstick across cards.
- H200 vs B200 — specs, price per hour, which to rent
- H100 SXM vs B200 — specs, price per hour, which to rent
- B200 vs B300 — specs, price per hour, which to rent
- All GPU comparisons
Renting a B200: frequently asked questions
How much does it cost to rent an NVIDIA B200 per hour?
$4.204 per GPU-hour on-demand — a fixed price set at least 30% below the current market median of $6.01. Interruptible capacity costs $2.102/hr and a 3-month reservation $2.732/hr. Around $3,069/month if you keep one running non-stop, billed per second.
What can a B200 with 192 GB VRAM run?
In LLM terms, roughly a 72B-parameter model in FP16 or up to ~235B parameters 4-bit quantized on a single card, with context headroom. Multi-GPU instances (up to 8× on current inventory) multiply that; diffusion and rendering workloads fit comfortably at this VRAM class.
Is the B200 available to rent right now?
Yes — 53 GPUs across 9 machines in 4 regions are listed as we render this page. Configurations go from 1× to 8× with NVLink on multi-GPU chassis. Deploy from the console and it is running in about 30 seconds.
How do I deploy a B200?
Create an account (email + password, no card, no KYC), top up in crypto, open the console, filter by B200, pick a machine and a template such as PyTorch, vLLM or ComfyUI. The same deploy is one command with the CLI: powergpu launch --gpu b200 --template pytorch.
Why is the B200 cheaper here than on GPU marketplaces?
We price from the public marketplace median and fix our on-demand rate at least 30% below it, rounded down. The price is re-checked weekly (next check 2026-09-10) and published — no auctions, no per-host roulette, no bidding.