Best AI model for summarising documents.
prices as of · re-ranked daily · 286 qualifying models
Reports, contracts and transcripts go in whole and come back as a paragraph. The bill is almost all input tokens, so the input price is the number that decides this job. We count only models with 128k of context or more, and 286 of the 337 we track qualify. The cheapest of them costs $0.012 per 1,000 pages at Merge Gateway. They are ranked by what the work costs, not by the price per million tokens, because a model that is cheap to prompt and dear to answer looks different once you know the shape of the job.
The cheapest model that can do this job is GLM-5.3-Flash at Merge Gateway: $0.012 per 1,000 pages, read on 16 Sept 2026.
Cheapest that qualifies
GLM-5.3-Flash
$0.012per 1,000 pages
$0.01 / $0.05 per M tokens · $0.06 at list
lowest cost per 1,000 pages
Cheapest from the maker
Command R7B Arabic
$0.0315per 1,000 pages
$0.04 / $0.15 per M tokens · $0.0315 at list
bought from the lab that trained it
Most context
DeepSeek V4.1 Flash
$0.0304per 1,000 pages
$0.04 / $0.08 per M tokens · $0.126 at list
the longest single call on this list
What actually matters here
- 128k context or more
- A 300-page report fits in one call at 128k tokens. Below that you split the document up and pay to re-read the overlap on every piece.
- Input price decides it
- Ten tokens go in for every one that comes out, so a model with a cheap input price and a dear output price is the right shape for this work.
- Reasoning is optional
- Summarising is not a puzzle. A reasoning model bills its thinking as output tokens, which is money spent on the half of the job that is already cheap.
Cheapest 10 for summarising documents
- 1
Merge GatewayGLM-5.3-Flash
$0.012 per 1,000 pages · $0.01 / $0.05 per M tokens · 1M contextHosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.
✓ 1M context✓ tool calling✓ structured output
$0.10/M tokens
- 2
Nous PortalMistral Nemo
$0.0126 per 1,000 pages · $0.02 / $0.03 per M tokens · 128k contextResells OpenRouter's catalog at 0.80× the listed price, funded by a subscription whose credits carry a 10% bonus.
✓ 128k context✓ tool calling✓ open weights
$0.08/M tokens
- 3
InferenceLlama 3.2 3B Instruct
$0.0132 per 1,000 pages · $0.02 / $0.02 per M tokens · 131k contextHosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.
✓ 131k context✓ structured output✓ open weights
$0.08/M tokens
- 4
Kilo GatewayLlama 3.1 8B Instruct
$0.0144 per 1,000 pages · $0.02 / $0.04 per M tokens · 131k contextRoutes to a host serving the open weights; precision not disclosed.
✓ 131k context✓ tool calling✓ structured output
$0.10/M tokens
- 5
OpenRouterLing 3.0 Flash
$0.0164 per 1,000 pages · $0.02 / $0.06 per M tokens · 262k contextThe same price OpenRouter charges.
✓ 262k context✓ tool calling✓ structured output
$0.13/M tokens
- 6
Eden AIQwen3.7 Flash
$0.0168 per 1,000 pages · $0.02 / $0.08 per M tokens · 1M contextAlibaba's China-region price list, which is lower than the international one. A region, not a deal.
✓ 1M context✓ tool calling✓ structured output
$0.14/M tokens
- 7
Kilo Gatewaygpt-oss-20b
$0.018 per 1,000 pages · $0.02 / $0.10 per M tokens · 131k contextRoutes to a host serving the open weights; precision not disclosed.
✓ 131k context✓ tool calling✓ structured output
$0.16/M tokens
- 8
NanoGPTMercury 2.5
$0.0242 per 1,000 pages · $0.04 / $0.004 per M tokens · 260k contextThe same price Inception charges.
✓ 260k context✓ tool calling✓ structured output
$0.12/M tokens
- 9
Alibaba Cloud (China)Qwen Flash
$0.0262 per 1,000 pages · $0.02 / $0.22 per M tokens · 1M contextAlibaba's China-region Model Studio, which prices the Qwen models below the international list. A region, not a deal.
✓ 1M context✓ tool calling
$0.28/M tokens
- 10
Azure Cognitive ServicesMinistral 3B (latest)
$0.0264 per 1,000 pages · $0.04 / $0.04 per M tokens · 128k contextThe same price Mistral AI charges.
✓ 128k context✓ tool calling✓ open weights
$0.16/M tokens
A page is about 600 tokens in and 60 out. A cost per 1,000 pages is that multiplied out at the row's own input and output rate. The price beside it is the standard price, three parts input to one part output per million tokens, so the two numbers answer different questions: $0.10 per million tokens is what the model costs, and $0.012 per 1,000 pages is what the work costs.
Frequently asked
- What is the cheapest model for summarising documents?
- GLM-5.3-Flash at Merge Gateway, $0.012 per 1,000 pages on $0.01 / $0.05 per million tokens in and out, read on 16 Sept 2026.
- How is the cost per unit worked out?
- A page is about 600 tokens in and 60 out. Multiply that by 1,000 and price it at each row's own input and output rate. Nothing else is counted: no cache discount, no batch rate, no free tier.
- Why do only 286 models qualify?
- The job needs 128k of context or more. Models that fall short are still in the index, they just cannot do this job as described.
- How current are these prices?
- Prices are read daily from 212 providers' public catalogs and this ranking recomputes with them. Last refresh: 16 Sept 2026.