Skip to content
LLMs
Thinking Machines logo

Thinking Machines: Inkling

by Thinking Machines

Thinking Machines: Inkling is a large language model from Thinking Machines. It costs $1.00 per million input tokens and $4.05 per million output tokens. Its context window is 1.0M tokens.

Free tier availableReasoningTool callingStructured outputPrompt cachingBatch tierOpen weightsimage inputaudio input
Input / 1M tokens
$1.00
Output / 1M tokens
$4.05
Cached input / 1M
$0.170

On repeated prefixes

Context window
1.0M

tokens

Who serves it cheapest

3 hosts serve Inkling. Same weights, same API - the price difference is pure margin and routing.

Providers serving Thinking Machines: Inkling, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
DeepInfraCheapestfp8$0.950$4.05524K-93.3%
BaseTenfp8$1.00$4.051.0M-97.4%
Together$1.00$4.05524K-97.4%

The spread between DeepInfra and Together is about the same for identical weights. Quantization and context limits differ, so check both columns before switching.

Benchmarks

Independent scores published alongside the catalogue.

Intelligence index
25.5
Coding index
52.1
Agentic index
24.3

About Inkling

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Specifications

Thinking Machines: Inkling specifications
Model IDthinkingmachines/inkling
ProviderThinking Machines
Context window1.0M tokens
Max output472K tokens
Input modalitiestext, image, audio
Output modalitiestext
Knowledge cutoff-
Open weightsYes - thinkingmachines/Inkling
ReleasedJuly 17, 2026

Cheaper alternatives

Models that cost less than Inkling while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does Thinking Machines: Inkling cost?

$1.00 per million input tokens and $4.05 per million output tokens. Cached input reads cost $0.170 per million tokens.

What is the context window of Thinking Machines: Inkling?

1.0M tokens, with up to 472K tokens of output per request.

Which provider serves Thinking Machines: Inkling cheapest?

DeepInfra at $0.950 per million input tokens - about the same cheaper than Together, the most expensive of the 3 hosts serving it.