Price floor Every GPU at least 30% below the market median — re-checked against the marketplace weekly.

See the proof

Template · LLM serving & chat · CUDA 13

Run vLLM on a cloud GPU, in 30 seconds

Production LLM serving with PagedAttention — an OpenAI-compatible endpoint from any HF model. vLLM is the production standard for serving open LLMs: PagedAttention, continuous batching, tensor parallelism and an OpenAI-compatible API. The template exposes port 8000 with TLS, reads MODEL from an environment variable and caches weights on the instance disk or a mounted volume. From $0.159/hr on an interruptible RTX 5090.

Running in ~30 s

Pre-cached image, ports mapped, TLS terminated — the stack is working before you finish reading this.

Template is free

You pay the GPU price only: vLLM on a RTX 5090 is $0.318/hr on-demand, $0.159/hr interruptible, billed per second.

Volumes for state

Models, datasets and outputs live on a $0.08/GB/mo volume; the instance stays disposable.

Private by default

Dedicated GPU, encrypted disk, crypto payments, no KYC, no stored IPs — Jupyter and SSH behind your own credentials.

Best GPUs for vLLM

Three price points that run this template well — Good, Better, Best. Every model in the catalogue works; these are the value picks.

TierGPUVRAMOn-demandInterruptibleWhy this cardAction
GoodRTX 509032 GB$0.318$0.15932 GB GDDR7: the lowest $/token for 7B–14B FP16 and 32B 4-bit models.Deploy
BetterL40S48 GB$0.466$0.23348 GB and datacenter cooling for 24/7 endpoints up to 32B.Deploy
BestH100 PCIE80 GB$2.147$1.07380 GB for quantized 70B on one card, FP8 for maximum throughput.Deploy

Need more VRAM? The full catalogue lists all 76 models with live availability; the VRAM guide sizes models to cards.

Deploy vLLM from the console, CLI or API

Pick the template in the console deploy bar, or script it:

  • Console — filter by GPU, choose vLLM in the template picker, set disk and env, deploy.
  • CLIpip install powergpu, then the command on the right. CLI reference.
  • APIPOST /v1/instances with "template": "vllm". REST reference.
  • Own image — any OCI reference works too; we inject the NVIDIA runtime. Template docs.
deploy — vllm
$ powergpu launch --gpu rtx-5090 --template vllm \
    --disk 100 --volume models:/workspace/models
 instance i-7a41c0e2 running (27.9s)
# vLLM · RTX 5090 · $0.318/hr · per second
# https://i-7a41c0e2.powergpu.io:8888 (TLS)
$ powergpu stop i-7a41c0e2   # billing ends this second

What is inside the vLLM template

Imagepowergpu/vllm
CUDACUDA 13
Accessalso builds for ARM hosts · SSH shell · JupyterLab on a mapped port
CategoryLLM serving & chat
StorageInstance NVMe disk (sized at deploy) + optional network volumes
BillingGPU price only, per second — no template fee, no setup fee

Other llm serving & chat templates

vLLM on a cloud GPU: FAQ

Environment variables, ports and custom images are covered in the template docs.

How much does it cost to run vLLM on a cloud GPU?

Only the GPU price — the template is free. From $0.159/hr on an interruptible RTX 5090, $0.318/hr on-demand on a RTX 5090. Billing is per second, so an hour of tinkering costs an hour, not a day. Storage is $0.08/GB/month.

How long does vLLM take to start?

About 30 seconds from the deploy click: the image is pre-cached on hosts, ports are mapped and TLS is terminated for you. Restarting a stopped instance is faster, and your disk is exactly as you left it.

Can I keep my models and outputs between sessions?

Yes — attach a volume at deploy. Everything on the volume survives instance destruction and mounts on the next instance in the region in seconds, at $0.08/GB/month. The instance disk itself survives stop/start but not destroy.

Is the endpoint really OpenAI-compatible?

Yes — point any OpenAI SDK at https://<instance>.powergpu.io:8000/v1 with the API key you set in VLLM_API_KEY. Chat, completions, embeddings and streaming all work unchanged.

Deploy your first GPU in under a minute

Top up in crypto, benchmark us against your current provider. Per-second billing, fixed prices ≥ 30% below market — cancel by just stopping the instance.