Skip to content
LLMs
Fireworks logo

Fireworks

US · serverless

Low-latency serverless inference with its own optimised runtime and on-demand GPU deployments.

Free tierOpenAI-compatible
Category
serverless
Free tier
Yes
OpenAI API
Compatible
Adds on top
None - sets its own prices

Margin over the underlying model price

What Fireworks charges

Live prices for widely-hosted open-weight models, sampled from the shared catalogue. Compare the same row against other hosts on each model page.

Models served by Fireworks
ModelInput / 1MOutput / 1MContextThroughput
GLM 5.3 Flash$0.150$0.5001.0M-
DeepSeek V4.1 Flash$0.220$0.6601.0M-
DeepSeek V4 Flash Vision Exp$0.220$0.6601.0M-
DeepSeek V4 Flash 0731$0.220$0.6601.0M-
Muse Glimmer 30B$0.350$1.50131K-
DeepSeek V4 Pro 0813$1.32$3.961.0M-
GLM 5.3$1.40$4.401.0M-
Kimi K3$3.00$15.001.0M-

Fireworks alternatives

Other services in the same category. These are genuine substitutes - services in a different category solve a different problem.

Frequently asked

Is Fireworks OpenAI-compatible?

Yes. Fireworks accepts the OpenAI chat-completions request shape, so most SDKs work by changing the base URL and API key.

Does Fireworks have a free tier?

Yes - Fireworks offers free usage, though limits and eligible models change frequently. Check the pricing page before relying on it.

How does Fireworks charge?

Per token, plus per-GPU-hour for dedicated Margin over the underlying model price: none - sets its own prices.