Skip to content
LLMs
Inference.net logo

Inference.net: Schematron V2 Turbo

by Inference.net

Inference.net: Schematron V2 Turbo is a large language model from Inference.net. It costs $0.030 per million input tokens and $0.150 per million output tokens. Its context window is 128K tokens.

Structured outputPrompt cachingOpen weights
Input / 1M tokens
$0.030
Output / 1M tokens
$0.150
Cached input / 1M
$0.030

On repeated prefixes

Context window
128K

tokens

Who serves it cheapest

1 hosts serve Schematron V2 Turbo. Same weights, same API - the price difference is pure margin and routing.

Providers serving Inference.net: Schematron V2 Turbo, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
InferenceNetCheapest$0.030$0.150128K-100.0%

About Schematron V2 Turbo

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...

Specifications

Inference.net: Schematron V2 Turbo specifications
Model IDinference-net/schematron-v2-turbo
ProviderInference.net
Context window128K tokens
Max output8K tokens
Input modalitiestext
Output modalitiestext
Knowledge cutoff-
Open weightsYes - inference-net/schematron-v2-granite-4.0-h-micro
ReleasedSeptember 12, 2026

Cheaper alternatives

Models that cost less than Schematron V2 Turbo while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does Inference.net: Schematron V2 Turbo cost?

$0.030 per million input tokens and $0.150 per million output tokens. Cached input reads cost $0.030 per million tokens.

What is the context window of Inference.net: Schematron V2 Turbo?

128K tokens, with up to 8K tokens of output per request.