Running in ~30 s
Pre-cached image, ports mapped, TLS terminated — the stack is working before you finish reading this.
Products
Developers & company
GPUs · on-demand per hour
Template · LLM serving & chat · CUDA 12.9
A full chat UI over Ollama — conversations, RAG and model management in the browser. Open WebUI on top of Ollama is a private ChatGPT-style interface: multi-user chat, document RAG, model switching and prompt libraries, all served from your own GPU behind TLS. Teams rent it as an internal assistant that never sends data to a third party. From $0.131/hr on an interruptible RTX 4090.
Pre-cached image, ports mapped, TLS terminated — the stack is working before you finish reading this.
You pay the GPU price only: Open WebUI (Ollama) on a RTX 4090 is $0.262/hr on-demand, $0.131/hr interruptible, billed per second.
Models, datasets and outputs live on a $0.08/GB/mo volume; the instance stays disposable.
Dedicated GPU, encrypted disk, crypto payments, no KYC, no stored IPs — Jupyter and SSH behind your own credentials.
Three price points that run this template well — Good, Better, Best. Every model in the catalogue works; these are the value picks.
| Tier | GPU | VRAM | On-demand | Interruptible | Why this card | Action |
|---|---|---|---|---|---|---|
| Good | RTX 4090 | 24 GB | $0.262 | $0.131 | 24 GB serves a 14B assistant to a small team. | Deploy |
| Better | RTX 5090 | 32 GB | $0.318 | $0.159 | 32 GB for 32B 4-bit models with more concurrent users. | Deploy |
| Best | L40S | 48 GB | $0.466 | $0.233 | 48 GB and passive cooling for an always-on team deployment. | Deploy |
Need more VRAM? The full catalogue lists all 76 models with live availability; the VRAM guide sizes models to cards.
Pick the template in the console deploy bar, or script it:
$ powergpu launch --gpu rtx-4090 --template open-webui-ollama \
--disk 100 --volume models:/workspace/models
✓ instance i-7a41c0e2 running (27.9s)
# Open WebUI (Ollama) · RTX 4090 · $0.262/hr · per second
# https://i-7a41c0e2.powergpu.io:8888 (TLS)
$ powergpu stop i-7a41c0e2 # billing ends this second
| Image | powergpu/openwebui |
|---|---|
| CUDA | CUDA 12.9 |
| Access | SSH shell · JupyterLab on a mapped port |
| Category | LLM serving & chat |
| Storage | Instance NVMe disk (sized at deploy) + optional network volumes |
| Billing | GPU price only, per second — no template fee, no setup fee |
Environment variables, ports and custom images are covered in the template docs.
Only the GPU price — the template is free. From $0.131/hr on an interruptible RTX 4090, $0.262/hr on-demand on a RTX 4090. Billing is per second, so an hour of tinkering costs an hour, not a day. Storage is $0.08/GB/month.
About 30 seconds from the deploy click: the image is pre-cached on hosts, ports are mapped and TLS is terminated for you. Restarting a stopped instance is faster, and your disk is exactly as you left it.
Yes — attach a volume at deploy. Everything on the volume survives instance destruction and mounts on the next instance in the region in seconds, at $0.08/GB/month. The instance disk itself survives stop/start but not destroy.
Yes — Open WebUI has accounts and roles, and Ollama queues concurrent requests. For heavy concurrency, pair it with the vLLM template as the backend instead.
Top up in crypto, benchmark us against your current provider. Per-second billing, fixed prices ≥ 30% below market — cancel by just stopping the instance.