Best AI model for a chat assistant.
prices as of · re-ranked daily · 280 qualifying models
An assistant in your product, holding a conversation and calling your API when it needs to. Every turn re-reads the whole conversation, so the input bill grows with the chat. We count only models with 32k of context or more and tool calling, and 280 of the 337 we track qualify. The cheapest of them costs $0.069 per 1,000 conversations at Nous Portal. They are ranked by what the work costs, not by the price per million tokens, because a model that is cheap to prompt and dear to answer looks different once you know the shape of the job.
The cheapest model that can do this job is Mistral Nemo at Nous Portal: $0.069 per 1,000 conversations, read on 16 Sept 2026.
Cheapest that qualifies
Mistral Nemo
$0.069per 1,000 conversations
$0.02 / $0.03 per M tokens · $0.525 at list
lowest cost per 1,000 conversations
Cheapest from the maker
Command R7B Arabic
$0.188per 1,000 conversations
$0.04 / $0.15 per M tokens · $0.188 at list
bought from the lab that trained it
Most context
DeepSeek V4.1 Flash
$0.169per 1,000 conversations
$0.04 / $0.08 per M tokens · $0.75 at list
the longest single call on this list
What actually matters here
- Tool calling required
- An assistant that can only talk is a search box with extra steps. Tool calling is what lets it check an order or book a slot.
- Input grows with the chat
- Each turn resends the conversation so far, so a ten-turn chat reads its own history nine times. Ask your provider about cached input before you scale it.
- Latency is not on this page
- Price is measured here; speed is not. For a live assistant, test the two or three cheapest rows for how fast the first token arrives.
Cheapest 10 for a chat assistant
- 1
Nous PortalMistral Nemo
$0.069 per 1,000 conversations · $0.02 / $0.03 per M tokens · 128k contextResells OpenRouter's catalog at 0.80× the listed price, funded by a subscription whose credits carry a 10% bonus.
✓ 128k context✓ tool calling✓ open weights
$0.08/M tokens
- 2
Merge GatewayGLM-5.3-Flash
$0.07 per 1,000 conversations · $0.01 / $0.05 per M tokens · 1M contextHosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.
✓ 1M context✓ tool calling✓ structured output
$0.10/M tokens
- 3
Kilo GatewayLlama 3.1 8B Instruct
$0.08 per 1,000 conversations · $0.02 / $0.04 per M tokens · 131k contextRoutes to a host serving the open weights; precision not disclosed.
✓ 131k context✓ tool calling✓ structured output
$0.10/M tokens
- 4
OpenRouterLing 3.0 Flash
$0.0945 per 1,000 conversations · $0.02 / $0.06 per M tokens · 262k contextThe same price OpenRouter charges.
✓ 262k context✓ tool calling✓ structured output
$0.13/M tokens
- 5
Eden AIQwen3.7 Flash
$0.101 per 1,000 conversations · $0.02 / $0.08 per M tokens · 1M contextAlibaba's China-region price list, which is lower than the international one. A region, not a deal.
✓ 1M context✓ tool calling✓ structured output
$0.14/M tokens
- 6
Kilo Gatewaygpt-oss-20b
$0.11 per 1,000 conversations · $0.02 / $0.10 per M tokens · 131k contextRoutes to a host serving the open weights; precision not disclosed.
✓ 131k context✓ tool calling✓ structured output
$0.16/M tokens
- 7
NanoGPTMercury 2.5
$0.122 per 1,000 conversations · $0.04 / $0.004 per M tokens · 260k contextThe same price Inception charges.
✓ 260k context✓ tool calling✓ structured output
$0.12/M tokens
- 8
Azure Cognitive ServicesMinistral 3B (latest)
$0.14 per 1,000 conversations · $0.04 / $0.04 per M tokens · 128k contextThe same price Mistral AI charges.
✓ 128k context✓ tool calling✓ open weights
$0.16/M tokens
- 9
NanoGPTNemotron 3.5 Lightning
$0.155 per 1,000 conversations · $0.05 / $0.01 per M tokens · 262k contextRoutes to a host serving the open weights; precision not disclosed.
✓ 262k context✓ tool calling✓ structured output
$0.16/M tokens
- 10
LLM Gatewaygpt-oss-120b
$0.166 per 1,000 conversations · $0.03 / $0.14 per M tokens · 131k contextRoutes to a host serving the open weights; precision not disclosed.
✓ 131k context✓ tool calling✓ structured output
$0.24/M tokens
A conversation of about ten turns is 3,000 tokens in and 500 out in total. A cost per 1,000 conversations is that multiplied out at the row's own input and output rate. The price beside it is the standard price, three parts input to one part output per million tokens, so the two numbers answer different questions: $0.08 per million tokens is what the model costs, and $0.069 per 1,000 conversations is what the work costs.
Frequently asked
- What is the cheapest model for a chat assistant?
- Mistral Nemo at Nous Portal, $0.069 per 1,000 conversations on $0.02 / $0.03 per million tokens in and out, read on 16 Sept 2026.
- How is the cost per unit worked out?
- A conversation of about ten turns is 3,000 tokens in and 500 out in total. Multiply that by 1,000 and price it at each row's own input and output rate. Nothing else is counted: no cache discount, no batch rate, no free tier.
- Why do only 280 models qualify?
- The job needs 32k of context or more, tool calling. Models that fall short are still in the index, they just cannot do this job as described.
- How current are these prices?
- Prices are read daily from 212 providers' public catalogs and this ranking recomputes with them. Last refresh: 16 Sept 2026.