Modal
US · gpu cloud
Serverless GPU platform where you deploy your own inference code.
- Category
- gpu cloud
- Free tier
- Yes
- OpenAI API
- Custom
- Adds on top
- None - billed for compute, not tokens
Margin over the underlying model price
What Modal charges
Live prices for widely-hosted open-weight models, sampled from the shared catalogue. Compare the same row against other hosts on each model page.
| Model | Input / 1M | Output / 1M | Context | Throughput |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | $0.300 | $1.20 | 1.0M | - |
| GLM 5.3 Flash | $0.450 | $1.50 | 1.0M | - |
| GLM 5.3 | $1.40 | $4.40 | 1.0M | - |
| Qwen3.8 2.4T A95B | $2.00 | $6.00 | 1M | - |
| Kimi K3 | $3.00 | $15.00 | 1.0M | - |
Modal alternatives
Other services in the same category. These are genuine substitutes - services in a different category solve a different problem.
CoreWeave
Large-scale GPU cloud used to host frontier training and inference.
Crusoe
GPU cloud powered by otherwise-stranded and low-carbon energy.
Hyperbolic
Pivoted from serverless inference to GPU rental during 2026; the dedicated inference pricing page is gone.
Lambda
GPU cloud. Its serverless Inference API is being wound down, so treat it as GPU rental rather than a token-billed endpoint.
RunPod
Rentable GPUs and serverless workers for self-managed model serving.
Frequently asked
Is Modal OpenAI-compatible?
No. Modal uses its own API shape, so you will need its SDK or a translation layer such as LiteLLM.
Does Modal have a free tier?
Yes - Modal offers free usage, though limits and eligible models change frequently. Check the pricing page before relying on it.
How does Modal charge?
Per second of GPU time Margin over the underlying model price: none - billed for compute, not tokens.