Salad
Distributed inference cloud — RTX 3090 $0.09/h, RTX 4090 $0.16/h
- Cheapest consumer GPUs — RTX 3090 from $0.09/h
- Massive horizontal scale (1000+ nodes)
Cheapest GPU clouds · July 2026
Rent GPU compute from $0.02/h. 11 budget GPU clouds ranked by raw price — with the trade-offs spelled out.
If your priority is squeezing maximum compute out of every dollar, four GPU clouds dominate the budget tier in 2026: RunPod (best value), Vast.ai, TensorDock, and Hyperstack. Hyperscalers (AWS, GCP, Azure) are systematically 3–5× more expensive for raw GPU compute and only make sense if you need their proprietary ML services.
The cheapest GPU clouds use one or more of these tactics:
Reality check: the cheapest tier requires fault-tolerant code (checkpointing, retry logic). For always-on production inference, add 50–80% to the sticker price for "Secure" or "On-Demand" tiers.
Spiky inference? The idle hours are the real cost. A $0.16/h instance left running around the clock is $3.84/day whether or not requests arrive. If your endpoint is bursty rather than constant, a serverless GPU that bills per second and scales to zero drops the idle hours entirely: you pay for active compute, not wall-clock time. The trade-off is a cold-start penalty on the first request after scale-down. Serverless stops paying off once utilization is high enough (very roughly a third of the day) that a dedicated cheap instance is simply cheaper to leave on.
Why no AWS, Azure or GCP here: all three do have sub-$1 GPU instances — AWS g4dn.xlarge and Azure NC4as_T4_v3 are both $0.526/h, GCP's g2-standard-4 is $0.71/h. But the first two are 2018-era T4 cards, and a T4 at $0.53 buys less throughput than a RunPod RTX A5000 at $0.16. Match the GPU class instead and the hyperscalers land 3–5× higher: an 8× H100 node is $55.04/h on AWS p5 and $88.49/h on GCP a3-highgpu-8g, against $1.99/h per H100 on RunPod Community.
The sticker price is not the thing that bites. Take a job needing 20 GPU-hours of real compute, checkpointed every 30 minutes, on an interruptible instance that gets reclaimed roughly every 3 hours. Each preemption throws away up to a half-hour of work and costs another 10–15 minutes to land a new node and reload weights.
| Tier | Rate | Billed hours | Total | Wall-clock |
|---|---|---|---|---|
| Interruptible / community | $0.18/h | ~22.3 h | ~$4.00 | ~23.3 h |
| On-demand / secure | $0.35/h | 20 h | $7.00 | 20 h |
That works out to about 7 preemptions, 1.7 hours of recomputed work, and 17% more wall-clock. Interruptible still wins on money, and usually will. It loses on everything else, and you only get that price if you already wrote checkpoint-and-resume logic that works. The $3 you save is real, but it is not worth an engineering afternoon unless you are running that job repeatedly or at much larger scale.
The break-even is roughly this: use the cheap tier for anything batch, repeated, or restartable. Pay for the secure tier when a failed run costs you more than the price gap, which is nearly always true for a demo the next morning or an inference endpoint with users on it.
| Provider | Starting Price | Top GPUs | Highlights | Rating | CTA |
|---|---|---|---|---|---|
| Salad | from $0.02/h | RTX 3090, RTX 4090, RTX 3080 ≤24GB |
| ★★★★☆ | View pricing |
| Vast.ai Editor's Choice | from $0.03/h | RTX 3090, RTX 4090, A100 ≤80GB |
| ★★★★☆ | View pricing |
| TensorDock | from $0.10/h | RTX 4090, RTX 3090, A100 80GB ≤80GB |
| ★★★★☆ | View pricing |
| Hyperstack | from $0.15/h | RTX A4000, RTX A6000, L40 ≤80GB |
| ★★★★☆ | View pricing |
| RunPod Editor's Choice | from $0.16/h | RTX A5000, RTX 3090, RTX 4090 ≤80GB |
| ★★★★★ | View pricing |
| Massed Compute | from $0.35/h | RTX A6000, A40, A100 80GB ≤80GB |
| ★★★★☆ | View pricing |
| OVH GPU | from €0.36/h | T4, V100, A100 ≤80GB |
| ★★★★☆ | View pricing |
| Jarvis Labs | from $0.41/h | A30, L4, A100 40GB ≤80GB |
| ★★★★☆ | View pricing |
| Paperspace | from $0.45/h | A100, A6000, RTX 4000 ≤80GB |
| ★★★★☆ | View pricing |
| Lambda Labs Editor's Choice | from $0.69/h | Quadro RTX 6000, A100 40GB, A100 80GB ≤80GB |
| ★★★★★ | View pricing |
| Scaleway | from €0.79/h | L4, L40S, H100 ≤80GB |
| ★★★★☆ | View pricing |
Distributed inference cloud — RTX 3090 $0.09/h, RTX 4090 $0.16/h
Cheapest GPU cloud — peer-to-peer marketplace for budget training
Marketplace GPU cloud — RTX A4000 from $0.10/h, RTX 4090 $0.35/h, H100 SXM5 $2.25/h
Global GPU cloud specialist — RTX A4000 from $0.15/h, plus H100, H200 and B200
Best value GPU cloud — huge selection, community + secure cloud
Workstation-grade GPUs for AI/ML/VFX — A100 from $1.79/h
For most workloads, RunPod is the best-value choice: RTX A5000 Community Cloud from $0.16/h with reliable infrastructure, persistent volumes and Serverless endpoints. If you only chase the absolute floor price: Salad starts at $0.02/h, though that tier is GTX 10-series hardware; its RTX 3090 is $0.09/h (stateless inference only, no training) and Vast.ai marketplace instances start at $0.02/h on a Tesla V100 32GB (interruptible — hosts can reclaim hardware anytime). The right pick depends on whether you need reliability and persistent state; for 9 out of 10 users that means RunPod.
No, not the marketplace/community tiers. Use them for: batch training with checkpoints, hobby projects, hyperparameter sweeps, batch inference. For production APIs, use RunPod Secure ($0.27/h+) or Lambda Labs ($0.69/h+) — still cheap, but on dedicated hardware. Hetzner is EU-sovereign and dedicated but no longer budget: its cheapest GPU server is €1.42/h.
AWS bundles its GPU compute with proprietary services (SageMaker, IAM, VPC, support tiers) and prices for enterprise customers who value the ecosystem. For pure compute, you pay 3-5× more. Specialist clouds skip this overhead. Use AWS only when you need its ecosystem.
Vast.ai 4090 community at $0.34/h or RunPod Community 4090 at $0.39/h. Both fit Llama 3 8B QLoRA in 24GB. Total run cost for a typical fine-tune (~12 hours): $4-5. Compare to AWS at $3.06/h = $37 for the same job.
Persistent storage ($0.10–0.20/GB/month), egress data transfer ($0.05-0.12/GB), static IPs ($3-10/month), and idle time charges (some providers bill for stopped pods retaining storage). RunPod and Vast.ai are the most transparent; hyperscalers have the worst hidden cost reputation.
Get an email when GPU prices drop or availability changes at your preferred provider.
No spam. Unsubscribe any time.