Cheapest models with 128k context.
prices as of · re-ranked daily · 273 qualifying models
128k tokens is roughly a 300-page book in one call, which is the point at which you stop chopping documents into pieces. 273 models hold that much and are sold by 3 providers or more, from $0.08. Prices are per million tokens, three parts input to one part output, at each provider's standard rate.
Today's pick is Llama 3.2 3B Instruct at Inference, $0.02 in and $0.02 out per million tokens, read on 16 Sept 2026.
- 1
InferenceLlama 3.2 3B Instruct
131k context, sold by 17 providers, $0.02 / $0.02 per M tokens at InferenceHosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.
$0.08/M tokens
- 2
Nous PortalMistral Nemo
128k context, sold by 25 providers, $0.02 / $0.03 per M tokens at Nous PortalResells OpenRouter's catalog at 0.80× the listed price, funded by a subscription whose credits carry a 10% bonus.
$0.08/M tokens
- 3
Merge GatewayGLM-5.3-Flash
1M context, sold by 76 providers, $0.01 / $0.05 per M tokens at Merge GatewayHosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.
$0.10/M tokens
- 4
Kilo GatewayLlama 3.1 8B Instruct
131k context, sold by 24 providers, $0.02 / $0.04 per M tokens at Kilo GatewayRoutes to a host serving the open weights; precision not disclosed.
$0.10/M tokens
- 5
NanoGPTMercury 2.5
260k context, sold by 8 providers, $0.04 / $0.004 per M tokens at NanoGPTThe same price Inception charges.
$0.12/M tokens
- 6
OpenRouterLing 3.0 Flash
262k context, sold by 12 providers, $0.02 / $0.06 per M tokens at OpenRouterThe same price OpenRouter charges.
$0.13/M tokens
- 7
Eden AIQwen3.7 Flash
1M context, sold by 15 providers, $0.02 / $0.08 per M tokens at Eden AIAlibaba's China-region price list, which is lower than the international one. A region, not a deal.
$0.14/M tokens
- 8
Kilo Gatewaygpt-oss-20b
131k context, sold by 43 providers, $0.02 / $0.10 per M tokens at Kilo GatewayRoutes to a host serving the open weights; precision not disclosed.
$0.16/M tokens
- 9
Azure Cognitive ServicesMinistral 3B (latest)
128k context, sold by 14 providers, $0.04 / $0.04 per M tokens at Azure Cognitive ServicesThe same price Mistral AI charges.
$0.16/M tokens
- 10
NanoGPTNemotron 3.5 Lightning
262k context, sold by 12 providers, $0.05 / $0.01 per M tokens at NanoGPTRoutes to a host serving the open weights; precision not disclosed.
$0.16/M tokens
Frequently asked
- What is the cheapest of the models with 128k context right now?
- Llama 3.2 3B Instruct at Inference, $0.02 in and $0.02 out per million tokens, read on 16 Sept 2026.
- Which lab shows up most in this ranking?
- Meta trained 2 of the top 10 models here.
- How is this ranked, and how current is it?
- Every row is the cheapest price for that model that we can account for: the maker's own, a price below it with a published reason, or a price at or above it. A price below the maker's own with nothing to explain it is shown on the model's page and ranked nowhere. Prices are read daily from 212 providers' public catalogs and blended three parts input to one part output per million tokens. Last refresh: 16 Sept 2026.