Thinking Machines: Inkling
by Thinking MachinesThinking Machines: Inkling is a large language model from Thinking Machines. It costs $1.00 per million input tokens and $4.05 per million output tokens. Its context window is 1.0M tokens.
- Input / 1M tokens
- $1.00
- Output / 1M tokens
- $4.05
- Cached input / 1M
- $0.170
- Context window
- 1.0M
On repeated prefixes
tokens
Who serves it cheapest
3 hosts serve Inkling. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfraCheapestfp8 | $0.950 | $4.05 | 524K | - | 93.3% |
| BaseTenfp8 | $1.00 | $4.05 | 1.0M | - | 97.4% |
| Together | $1.00 | $4.05 | 524K | - | 97.4% |
The spread between DeepInfra and Together is about the same for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 25.5
- Coding index
- 52.1
- Agentic index
- 24.3
About Inkling
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Specifications
| Model ID | thinkingmachines/inkling |
|---|---|
| Provider | Thinking Machines |
| Context window | 1.0M tokens |
| Max output | 472K tokens |
| Input modalities | text, image, audio |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - thinkingmachines/Inkling |
| Released | July 17, 2026 |
Cheaper alternatives
Models that cost less than Inkling while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Thinking Machines: Inkling cost?
$1.00 per million input tokens and $4.05 per million output tokens. Cached input reads cost $0.170 per million tokens.
What is the context window of Thinking Machines: Inkling?
1.0M tokens, with up to 472K tokens of output per request.
Which provider serves Thinking Machines: Inkling cheapest?
DeepInfra at $0.950 per million input tokens - about the same cheaper than Together, the most expensive of the 3 hosts serving it.