Devstral Small vs Step 3.5 Flash 2603.

model vs model · prices as of

Devstral Small and Step 3.5 Flash 2603 split the 5 rounds this table counts: the two list price rows, the cheapest price we can explain, the standard price and context. The weighting you pick decides the rest.

Mistral AI

Devstral Small

$0.24/M tokens

cheapest we can explain, at NanoGPT · $0.0396 per 1,000 pages

128k context · tool calling · open weights

NanoGPT
StepFun

Step 3.5 Flash 2603

$0.57/M tokens

cheapest we can explain, at Vercel AI Gateway · $0.072 per 1,000 pages

256k context · tool calling · open weights

List price in (lower wins)

$0.10/M

$0.10/M

List price out (lower wins)

$0.30/M

$0.30/M

Cheapest we can explain (lower wins)

$0.24/M at NanoGPT

$0.57/M at Vercel AI Gateway

Standard price (lower wins)

$0.60/M

$0.60/M

Cost per 1,000 pages (lower wins)

$0.0396

$0.072

Context

128k tokens

256k tokens

Open weights

yes

yes

Tool calling

yes

yes

Structured output

not stated

not stated

Reads images

no

no

Reasoning

no

yes

Released

10 Jul 2025

2 Apr 2026

Providers selling it

9

16

Price we measure against

the maker's own

the maker's own

Trains on your prompts (no is better)

not stated

not stated

Providers selling both · 3

  • NanoGPT$0.24 / $0.60
  • Eden AI$0.60 / $0.57
  • LLMTR$0.60 / $0.60

Devstral Small then Step 3.5 Flash 2603, per million tokens

Go further

Add a third or fourth model, or switch the weighting, in the tool.

What moved in the last seven days

  1. Step 3.5 Flash 2603Nous Portal$0.08 / $0.24 $0.10 / $0.3025% dearer ·

Frequently asked

Which is cheaper, Devstral Small or Step 3.5 Flash 2603?
Devstral Small, at $0.24 per million tokens against $0.57 for Step 3.5 Flash 2603, three parts input to one part output, read on 16 Sept 2026.
Where is each one cheapest?
Devstral Small: $0.24 at NanoGPT. Step 3.5 Flash 2603: $0.57 at Vercel AI Gateway.
Can I buy both from one provider?
Yes. 3 providers sell both: NanoGPT, Eden AI, LLMTR. The cheapest of them for the two together is NanoGPT.
Which holds more context?
Step 3.5 Flash 2603, at 256k tokens against 128k.

Prices exclude tax and are per million tokens, read 16 Sept 2026 from each provider's own catalog. The standard price blends three parts input to one part output; a page is 600 tokens in and 60 out. Bars are scaled to the larger value in each row.

Related head-to-headsDevstral Small vs Gemini Embedding 2Devstral Small vs Mistral NemoDevstral Small vs Mistral Small 3.2Devstral Small vs Pixtral 12BDevstral Small vs Voxtral Small (latest)Devstral Small vs Voxtral Small 24B 2507Morecompare any two modelsevery model we track