Salad
Distributed inference cloud — RTX 3090 $0.09/h, RTX 4090 $0.16/h
- Cheapest consumer GPUs — RTX 3090 from $0.09/h
- Massive horizontal scale (1000+ nodes)
Serverless GPU · July 2026
Pay only when your model runs — no idle bills. 6 serverless GPU providers ranked by cold start, pricing, and reliability.
Serverless GPU inference means paying only when your model is actually running — no idle-time bills, no cluster management. Six providers dominate serverless GPU in 2026: RunPod Serverless (best overall), Together AI (per-token LLM APIs), Salad Cloud (distributed), Modal (developer-friendly), Nebius (Kubernetes-native), and CoreWeave (enterprise scale).
There are three flavors of "serverless GPU":
The right pick depends on traffic pattern: spiky low-QPS → Container serverless (RunPod). Steady high-QPS → dedicated GPU. LLM inference at scale → Per-token API (Together AI) until you exceed ~$5k/month, then self-host.
| Provider | Starting Price | Top GPUs | Highlights | Rating | CTA |
|---|---|---|---|---|---|
| Salad | from $0.02/h | RTX 3090, RTX 4090, RTX 3080 ≤24GB |
| ★★★★☆ | View pricing |
| RunPod Editor's Choice | from $0.16/h | RTX A5000, RTX 3090, RTX 4090 ≤80GB |
| ★★★★★ | View pricing |
| Lambda Labs Editor's Choice | from $0.69/h | Quadro RTX 6000, A100 40GB, A100 80GB ≤80GB |
| ★★★★★ | View pricing |
| Nebius Editor's Choice | from $1.55/h | H100, H200, B200 ≤192GB |
| ★★★★★ | View pricing |
| Together AI | from $3.99/h | H100, H200, A100 80GB ≤141GB |
| ★★★★☆ | View pricing |
| CoreWeave | from $6.50/h | L40S, H100 SXM, A100 SXM ≤80GB |
| ★★★★☆ | View pricing |
Distributed inference cloud — RTX 3090 $0.09/h, RTX 4090 $0.16/h
Best value GPU cloud — huge selection, community + secure cloud
On-demand H100 clusters — developer-favourite for serious ML
EU-sovereign AI cloud from the Netherlands — full GDPR compliance, H100 to B200
Inference-first GPU cloud — H100/H200 with optimized serving stacks
Enterprise H100 clusters — Kubernetes-native GPU cloud
A serverless GPU is a cloud compute model where you pay per-second (or per-request) for GPU inference instead of renting a persistent GPU instance. The provider auto-scales workers up when requests arrive and down to zero when idle. Best providers in 2026: RunPod Serverless for containers, Together AI for LLM APIs, Salad for distributed inference.
RunPod Serverless charges per compute unit — roughly $0.0002–$0.0006 per second on RTX 4090, or ~$0.50/hour equivalent while running. Together AI charges $0.10–$5 per million tokens depending on model size. Salad averages $0.02–$0.05/hour equivalent on consumer GPUs. Compare to $0.30–$2.50/hour for a persistent RTX 4090/H100 — serverless wins when your utilization is under ~40%.
Cold starts on RunPod Serverless are 2–8 seconds for smaller models (7B LLMs, SD 1.5) and 10–30 seconds for larger ones (70B LLMs, FLUX). You can pin warm workers at extra cost to eliminate cold starts. Per-token APIs (Together AI, Anyscale) have zero cold start because models are always loaded.
RunPod Serverless is best for custom container inference at low cost. Modal has the smoothest developer UX (Python decorators, no Dockerfile needed) but is 2× more expensive. Together AI is best for chat/completion APIs with popular open-source models. Pick RunPod for cost + flexibility, Modal for fast prototyping, Together AI for pre-hosted LLM inference.
No, serverless GPU platforms are optimized for inference (short, stateless workloads). For training or fine-tuning, use persistent instances: RunPod Pods, Lambda Labs on-demand, or Vast.ai. Training runs need consistent GPU memory across hours/days — the container startup overhead of serverless kills training economics.
AWS Lambda does not support GPU. Amazon SageMaker Serverless Inference supports CPU only. The closest AWS equivalent is SageMaker Real-Time Inference with auto-scaling, but it bills by instance-hour, not per-second — so it is not truly serverless. Use RunPod Serverless or Modal on AWS-adjacent workloads.
Get an email when GPU prices drop or availability changes at your preferred provider.
No spam. Unsubscribe any time.