Thinking Machines: Inkling Small
by Thinking MachinesThinking Machines: Inkling Small is a large language model from Thinking Machines. It costs $0.450 per million input tokens and $1.20 per million output tokens. Its context window is 1.0M tokens.
- Input / 1M tokens
- $0.450
- Output / 1M tokens
- $1.20
- Cached input / 1M
- $0.100
- Context window
- 1.0M
On repeated prefixes
tokens
Who serves it cheapest
3 hosts serve Inkling Small. Same weights, same API - the price difference is pure margin and routing.
| Provider | Input / 1M | Output / 1M | Context | Throughput | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfraCheapestfp8 | $0.450 | $1.20 | 524K | - | 49.3% |
| BaseTenfp8 | $0.500 | $1.20 | 1.0M | - | 100.0% |
| Together | $0.500 | $1.20 | 524K | - | 99.5% |
The spread between DeepInfra and Together is 1.1× for identical weights. Quantization and context limits differ, so check both columns before switching.
Benchmarks
Independent scores published alongside the catalogue.
- Intelligence index
- 26.1
- Coding index
- 52.9
- Agentic index
- 25.0
About Inkling Small
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Specifications
| Model ID | thinkingmachines/inkling-small |
|---|---|
| Provider | Thinking Machines |
| Context window | 1.0M tokens |
| Max output | 262K tokens |
| Input modalities | text, image, audio |
| Output modalities | text |
| Knowledge cutoff | - |
| Open weights | Yes - thinkingmachines/Inkling-Small |
| Released | July 30, 2026 |
Cheaper alternatives
Models that cost less than Inkling Small while keeping at least half its context window and every input modality it supports.
Frequently asked
How much does Thinking Machines: Inkling Small cost?
$0.450 per million input tokens and $1.20 per million output tokens. Cached input reads cost $0.100 per million tokens.
What is the context window of Thinking Machines: Inkling Small?
1.0M tokens, with up to 262K tokens of output per request.
Which provider serves Thinking Machines: Inkling Small cheapest?
DeepInfra at $0.450 per million input tokens - 1.1× cheaper than Together, the most expensive of the 3 hosts serving it.