Cheapest models with 128k context.

prices as of · re-ranked daily · 273 qualifying models

128k tokens is roughly a 300-page book in one call, which is the point at which you stop chopping documents into pieces. 273 models hold that much and are sold by 3 providers or more, from $0.08. Prices are per million tokens, three parts input to one part output, at each provider's standard rate.

Today's pick is Llama 3.2 3B Instruct at Inference, $0.02 in and $0.02 out per million tokens, read on 16 Sept 2026.

Frequently asked

What is the cheapest of the models with 128k context right now?
Llama 3.2 3B Instruct at Inference, $0.02 in and $0.02 out per million tokens, read on 16 Sept 2026.
Which lab shows up most in this ranking?
Meta trained 2 of the top 10 models here.
How is this ranked, and how current is it?
Every row is the cheapest price for that model that we can account for: the maker's own, a price below it with a published reason, or a price at or above it. A price below the maker's own with nothing to explain it is shown on the model's page and ranked nowhere. Prices are read daily from 212 providers' public catalogs and blended three parts input to one part output per million tokens. Last refresh: 16 Sept 2026.
Related rankingsCheapest LLM APICheapest open-weight modelsCheapest models with tool callingBrowse all 337 models