Skip to content
LLMs
inclusionAI logo

inclusionAI: Ling 3.0 Flash

by inclusionAI

inclusionAI: Ling 3.0 Flash is a large language model from inclusionAI. It costs $0.021 per million input tokens and $0.063 per million output tokens. Its context window is 262K tokens.

ReasoningTool callingStructured outputPrompt cachingOpen weights
Input / 1M tokens
$0.021
Output / 1M tokens
$0.063
Cached input / 1M
$0.0042

On repeated prefixes

Context window
262K

tokens

Who serves it cheapest

2 hosts serve Ling 3.0 Flash. Same weights, same API - the price difference is pure margin and routing.

Providers serving inclusionAI: Ling 3.0 Flash, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
NovitaCheapest$0.021$0.063262K-100.0%
DeepInfrabf16$0.060$0.180131K-98.1%

The spread between Novita and DeepInfra is 2.9× for identical weights. Quantization and context limits differ, so check both columns before switching.

Benchmarks

Independent scores published alongside the catalogue.

Coding index
50.6
Agentic index
21.0

About Ling 3.0 Flash

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

Specifications

inclusionAI: Ling 3.0 Flash specifications
Model IDinclusionai/ling-3.0-flash
ProviderinclusionAI
Context window262K tokens
Max output33K tokens
Input modalitiestext
Output modalitiestext
Knowledge cutoff-
Open weightsYes - inclusionAI/Ling-3.0-flash
ReleasedJuly 23, 2026

Cheaper alternatives

Models that cost less than Ling 3.0 Flash while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does inclusionAI: Ling 3.0 Flash cost?

$0.021 per million input tokens and $0.063 per million output tokens. Cached input reads cost $0.0042 per million tokens.

What is the context window of inclusionAI: Ling 3.0 Flash?

262K tokens, with up to 33K tokens of output per request.

Which provider serves inclusionAI: Ling 3.0 Flash cheapest?

Novita at $0.021 per million input tokens - 2.9× cheaper than DeepInfra, the most expensive of the 2 hosts serving it.