Compare · prices checked 2026-09-03
H200 vs B200: specs, price per hour, which to rent
NVIDIA H200 (141 GB, $3.058/hr) against NVIDIA B200 (192 GB, $4.204/hr): public specs side by side, live fixed prices from our sheet, what fits in each card's VRAM, and a verdict written for real workloads — both are rentable right now.
H200 vs B200 specifications
Public NVIDIA figures (dense, non-sparsity). The last column is H200 relative to B200.
| Spec | H200 | B200 | Difference |
|---|---|---|---|
| Architecture | Hopper (2023) | Blackwell (2024) | — |
| VRAM | 141 GB HBM3e | 192 GB HBM3e | −27% |
| Memory bandwidth | 4,800 GB/s | 8,000 GB/s | −40% |
| FP16 tensor (dense) | 990 TFLOPS | 2,250 TFLOPS | −56% |
| FP32 | 67.0 TFLOPS | — | — |
| CUDA cores | 16,896 | — | — |
| TDP | 700 W | 1000 W | −30% |
| PCIe · NVLink | Gen 5.0 · no NVLink | Gen 5.0 · NVLink | — |
| PowerScore (RTX 3090 = 100) | 697 | 1585 | −56% |
| Max GPUs per machine | 8× | 8× | — |
H200 vs B200 price per hour
Fixed rates from our sheet — every on-demand price is the marketplace median × 0.70, rounded down. Monthly = 730 hours.
| Rate | H200 | B200 | Cheaper |
|---|---|---|---|
| On-demand, per GPU-hour | $3.058 | $4.204 | H200 (−27%) |
| Interruptible, per GPU-hour | $1.529 | $2.102 | H200 |
| Reserved (3 mo), per GPU-hour | $1.987 | $2.732 | H200 |
| On-demand, per month | $2,232 | $3,069 | H200 |
| Market median (reference) | $4.37 | $6.01 | — |
| $ per 1,000 FP16 TFLOP-hours | $3.09 | $1.87 | B200 (better value) |
| $ per GB of VRAM per hour | $0.0217 | $0.0219 | H200 (better value) |
Try a full month with storage and bandwidth in the GPU cost calculator.
What fits in VRAM: 141 GB vs 192 GB
| Workload | H200 | B200 |
|---|---|---|
| Largest LLM in FP16, one card | ~49B | ~72B |
| Largest LLM at 4-bit, one card | ~141B | ~235B |
| Flux dev (FP8, ~17 GB) | fits | fits |
| Wan 2.x 14B video (offloaded) | yes | yes |
| 70B 4-bit LLM on one card | yes | yes |
Rules of thumb: ~2.4 GB per billion parameters in FP16 all-in, ~0.62 GB in 4-bit. Full tables in the VRAM guide.
Verdict: which should you rent?
The B200 roughly doubles FP16 throughput and bandwidth over the H200 and adds FP4; the H200 stays cheaper per hour with 141 GB. For throughput-critical training and FP4 inference the B200 usually costs less per token; for memory capacity at a lower rate, the H200.
- Cheaper per hour: H200 ($3.058 vs $4.204, −27%).
- More VRAM: B200 (192 GB vs 141 GB).
- More FP16 throughput: B200 (about 2.3×).
- Best value per TFLOP-hour: B200.
- Best value per GB of VRAM: H200.
- Multi-GPU: H200 over PCIe · B200 with NVLink.
H200 vs B200: FAQ
Deeper reading: H100 vs H200 vs B200, RTX 4090 vs RTX 5090, how cloud GPU pricing works.
Is the B200 faster than the H200?
On dense FP16 tensor throughput the B200 leads by about 2.3× (2,250 vs 990 TFLOPS). Memory bandwidth matters as much for inference: H200 4,800 GB/s vs B200 8,000 GB/s.
Which is cheaper to rent, the H200 or the B200?
The H200: $3.058/hr on-demand versus $4.204/hr — 27% less. Interruptible rates are $1.529 (H200) and $2.102 (B200). Per TFLOP-hour the better value is the B200.
Which has more VRAM and what does that change?
The B200 has 192 GB versus 141 GB. In LLM terms that is roughly a 72B FP16 model (or ~235B in 4-bit) on one card against 49B FP16 (~141B 4-bit). If the model does not fit, speed is irrelevant.
Can I rent both on PowerGPU right now?
Yes — 41 × H200 and 53 × B200 are online as this page renders, deployable in about 30 seconds, billed per second, paid in crypto with no KYC.