Best AI model for answering support questions.
prices as of · re-ranked daily · 280 qualifying models
A customer asks something, the model looks it up in your help centre and writes the reply. Short in, longer out, and it has to be able to call your search. We count only models with 32k of context or more and tool calling, and 280 of the 337 we track qualify. The cheapest of them costs $0.0092 per 1,000 replies at NanoGPT. They are ranked by what the work costs, not by the price per million tokens, because a model that is cheap to prompt and dear to answer looks different once you know the shape of the job.
The cheapest model that can do this job is Mercury 2.5 at NanoGPT: $0.0092 per 1,000 replies, read on 16 Sept 2026.
Cheapest that qualifies
Mercury 2.5
$0.0092per 1,000 replies
$0.04 / $0.004 per M tokens · $0.053 at list
lowest cost per 1,000 replies
Cheapest open weights
Mistral Nemo
$0.0126per 1,000 replies
$0.02 / $0.03 per M tokens · $0.075 at list
weights you could host yourself instead
Cheapest from the maker
Ministral 8B (latest)
$0.05per 1,000 replies
$0.10 / $0.10 per M tokens · $0.05 at list
bought from the lab that trained it
What actually matters here
- Tool calling required
- The answer lives in your help centre, not in the model. It has to be able to call a search function and write from what comes back.
- Output price decides it
- A reply is longer than the question, so this is one of the few jobs where the output price is the one to watch.
- 32k context is enough
- A question, a few retrieved articles and the reply fit comfortably. Paying for a million tokens of context you never fill is waste.
Cheapest 10 for answering support questions
- 1
NanoGPTMercury 2.5
$0.0092 per 1,000 replies · $0.04 / $0.004 per M tokens · 260k contextThe same price Inception charges.
✓ 260k context✓ tool calling✓ structured output
$0.12/M tokens
- 2
Nous PortalMistral Nemo
$0.0126 per 1,000 replies · $0.02 / $0.03 per M tokens · 128k contextResells OpenRouter's catalog at 0.80× the listed price, funded by a subscription whose credits carry a 10% bonus.
✓ 128k context✓ tool calling✓ open weights
$0.08/M tokens
- 3
NanoGPTNemotron 3.5 Lightning
$0.013 per 1,000 replies · $0.05 / $0.01 per M tokens · 262k contextRoutes to a host serving the open weights; precision not disclosed.
✓ 262k context✓ tool calling✓ structured output
$0.16/M tokens
- 4
Kilo GatewayLlama 3.1 8B Instruct
$0.016 per 1,000 replies · $0.02 / $0.04 per M tokens · 131k contextRoutes to a host serving the open weights; precision not disclosed.
✓ 131k context✓ tool calling✓ structured output
$0.10/M tokens
- 5
Merge GatewayGLM-5.3-Flash
$0.018 per 1,000 replies · $0.01 / $0.05 per M tokens · 1M contextHosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.
✓ 1M context✓ tool calling✓ structured output
$0.10/M tokens
- 6
Azure Cognitive ServicesMinistral 3B (latest)
$0.02 per 1,000 replies · $0.04 / $0.04 per M tokens · 128k contextThe same price Mistral AI charges.
✓ 128k context✓ tool calling✓ open weights
$0.16/M tokens
- 7
NanoGPTMuse Spark 1.3 Contributor
$0.0206 per 1,000 replies · $0.10 / $0.002 per M tokens · 1M contextThe same price Meta charges.
✓ 1M context✓ tool calling✓ structured output
$0.30/M tokens
- 8
NanoGPTMuse Spark 1.2 Contributor
$0.0206 per 1,000 replies · $0.10 / $0.002 per M tokens · 1M contextThe same price Meta charges.
✓ 1M context✓ tool calling✓ structured output
$0.30/M tokens
- 9
OpenRouterLing 3.0 Flash
$0.0231 per 1,000 replies · $0.02 / $0.06 per M tokens · 262k contextThe same price OpenRouter charges.
✓ 262k context✓ tool calling✓ structured output
$0.13/M tokens
- 10
SiliconFlow (China)Qwen2.5 7B Instruct
$0.025 per 1,000 replies · $0.05 / $0.05 per M tokens · 131k contextSiliconFlow's China-region catalog, priced in the domestic market.
✓ 131k context✓ tool calling✓ open weights
$0.20/M tokens
A support reply is about 200 tokens in and 300 out. A cost per 1,000 replies is that multiplied out at the row's own input and output rate. The price beside it is the standard price, three parts input to one part output per million tokens, so the two numbers answer different questions: $0.12 per million tokens is what the model costs, and $0.0092 per 1,000 replies is what the work costs.
Frequently asked
- What is the cheapest model for answering support questions?
- Mercury 2.5 at NanoGPT, $0.0092 per 1,000 replies on $0.04 / $0.004 per million tokens in and out, read on 16 Sept 2026.
- How is the cost per unit worked out?
- A support reply is about 200 tokens in and 300 out. Multiply that by 1,000 and price it at each row's own input and output rate. Nothing else is counted: no cache discount, no batch rate, no free tier.
- Why do only 280 models qualify?
- The job needs 32k of context or more, tool calling. Models that fall short are still in the index, they just cannot do this job as described.
- How current are these prices?
- Prices are read daily from 212 providers' public catalogs and this ranking recomputes with them. Last refresh: 16 Sept 2026.