Skip to content
LLMs
Inference.net logo

Inference.net: Schematron V2 Small

by Inference.net

Inference.net: Schematron V2 Small is a large language model from Inference.net. It costs $0.050 per million input tokens and $0.230 per million output tokens. Its context window is 128K tokens.

Structured outputPrompt cachingOpen weights
Input / 1M tokens
$0.050
Output / 1M tokens
$0.230
Cached input / 1M
$0.050

On repeated prefixes

Context window
128K

tokens

Who serves it cheapest

1 hosts serve Schematron V2 Small. Same weights, same API - the price difference is pure margin and routing.

Providers serving Inference.net: Schematron V2 Small, cheapest first
ProviderInput / 1MOutput / 1MContextThroughputUptime 24h
InferenceNetCheapest$0.050$0.230128K-100.0%

About Schematron V2 Small

Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. Extraction instructions must be supplied through a JSON schema...

Specifications

Inference.net: Schematron V2 Small specifications
Model IDinference-net/schematron-v2-small
ProviderInference.net
Context window128K tokens
Max output4K tokens
Input modalitiestext
Output modalitiestext
Knowledge cutoff-
Open weightsYes - inference-net/schematron-v2-llama-3.2-3b
ReleasedSeptember 12, 2026

Cheaper alternatives

Models that cost less than Schematron V2 Small while keeping at least half its context window and every input modality it supports.

Frequently asked

How much does Inference.net: Schematron V2 Small cost?

$0.050 per million input tokens and $0.230 per million output tokens. Cached input reads cost $0.050 per million tokens.

What is the context window of Inference.net: Schematron V2 Small?

128K tokens, with up to 4K tokens of output per request.