R1 Distill Llama 70B vs Devstral Medium.
model vs model · prices as of
R1 Distill Llama 70B and Devstral Medium split the 5 rounds this table counts: the two list price rows, the cheapest price we can explain, the standard price and context. The weighting you pick decides the rest.
R1 Distill Llama 70B
$1.20/M tokens
cheapest we can explain, at DeepInfra · $0.156 per 1,000 pages
8k context · open weights
Devstral Medium
$3.20/M tokens
Mistral AI list price, nothing below it is accounted for · $0.36 per 1,000 pages
128k context · tool calling · open weights
List price in (lower wins)
$0.80/M
$0.40/M
List price out (lower wins)
$0.80/M
$2/M
Cheapest we can explain (lower wins)
$1.20/M at DeepInfra
none below list
Standard price (lower wins)
$3.20/M
$3.20/M
Cost per 1,000 pages (lower wins)
$0.156
$0.36
Context
8k tokens
128k tokens
Open weights
yes
yes
Tool calling
no
yes
Structured output
no
not stated
Reads images
no
no
Reasoning
yes
no
Released
23 Jan 2025
10 Jul 2025
Providers selling it
18
4
Price we measure against
OpenRouter
the maker's own
Trains on your prompts (no is better)
not stated
no provider to check
Providers selling both · 1
Eden AI$2.90 / $3.20
R1 Distill Llama 70B then Devstral Medium, per million tokens
Go further
Add a third or fourth model, or switch the weighting, in the tool.
Frequently asked
- Which is cheaper, R1 Distill Llama 70B or Devstral Medium?
- R1 Distill Llama 70B, at $1.20 per million tokens against $3.20 for Devstral Medium, three parts input to one part output, read on 16 Sept 2026.
- Where is each one cheapest?
- R1 Distill Llama 70B: $1.20 at DeepInfra. Devstral Medium: nobody goes below the maker's own price with a reason we can point at, so $3.20 is the number to beat.
- Can I buy both from one provider?
- Yes. One provider sells both: Eden AI. The cheapest of them for the two together is Eden AI.
- Which holds more context?
- Devstral Medium, at 128k tokens against 8k.
Prices exclude tax and are per million tokens, read 16 Sept 2026 from each provider's own catalog. The standard price blends three parts input to one part output; a page is 600 tokens in and 60 out. Bars are scaled to the larger value in each row.