Best AI model for a chat assistant.

prices as of · re-ranked daily · 280 qualifying models

An assistant in your product, holding a conversation and calling your API when it needs to. Every turn re-reads the whole conversation, so the input bill grows with the chat. We count only models with 32k of context or more and tool calling, and 280 of the 337 we track qualify. The cheapest of them costs $0.069 per 1,000 conversations at Nous Portal. They are ranked by what the work costs, not by the price per million tokens, because a model that is cheap to prompt and dear to answer looks different once you know the shape of the job.

The cheapest model that can do this job is Mistral Nemo at Nous Portal: $0.069 per 1,000 conversations, read on 16 Sept 2026.

Cheapest that qualifies

Mistral Nemo

Nous Portal

$0.069per 1,000 conversations

$0.02 / $0.03 per M tokens · $0.525 at list

lowest cost per 1,000 conversations

Cheapest from the maker

Command R7B Arabic

Cohere

$0.188per 1,000 conversations

$0.04 / $0.15 per M tokens · $0.188 at list

bought from the lab that trained it

Most context

DeepSeek V4.1 Flash

TrustedRouter

$0.169per 1,000 conversations

$0.04 / $0.08 per M tokens · $0.75 at list

the longest single call on this list

What actually matters here

Tool calling required
An assistant that can only talk is a search box with extra steps. Tool calling is what lets it check an order or book a slot.
Input grows with the chat
Each turn resends the conversation so far, so a ten-turn chat reads its own history nine times. Ask your provider about cached input before you scale it.
Latency is not on this page
Price is measured here; speed is not. For a live assistant, test the two or three cheapest rows for how fast the first token arrives.

Cheapest 10 for a chat assistant

A conversation of about ten turns is 3,000 tokens in and 500 out in total. A cost per 1,000 conversations is that multiplied out at the row's own input and output rate. The price beside it is the standard price, three parts input to one part output per million tokens, so the two numbers answer different questions: $0.08 per million tokens is what the model costs, and $0.069 per 1,000 conversations is what the work costs.

Frequently asked

What is the cheapest model for a chat assistant?
Mistral Nemo at Nous Portal, $0.069 per 1,000 conversations on $0.02 / $0.03 per million tokens in and out, read on 16 Sept 2026.
How is the cost per unit worked out?
A conversation of about ten turns is 3,000 tokens in and 500 out in total. Multiply that by 1,000 and price it at each row's own input and output rate. Nothing else is counted: no cache discount, no batch rate, no free tier.
Why do only 280 models qualify?
The job needs 32k of context or more, tool calling. Models that fall short are still in the index, they just cannot do this job as described.
How current are these prices?
Prices are read daily from 212 providers' public catalogs and this ranking recomputes with them. Last refresh: 16 Sept 2026.
Related jobsBest AI model for answering support questionsBest AI model for writing and reviewing codeCheapest models with tool callingCheapest LLM API