GPT-4.1 mini vs Hermes 3 70B Instruct.
model vs model · prices as of
Hermes 3 70B Instruct takes 3 of 5 rounds against GPT-4.1 mini. It is also the cheaper of the two to run, at $0.66 per million tokens against $2.55.
GPT-4.1 mini
$2.55/M tokens
cheapest we can explain, at Poe · $0.305 per 1,000 pages
1M context · tool calling · reads images
Hermes 3 70B Instruct
$0.66/M tokens
cheapest we can explain, at Hyperbolic · $0.09 per 1,000 pages
131k context · open weights
List price in (lower wins)
$0.40/M
$0.70/M
List price out (lower wins)
$1.60/M
$0.70/M
Cheapest we can explain (lower wins)
$2.55/M at Poe
$0.66/M at Hyperbolic
Standard price (lower wins)
$2.80/M
$2.80/M
Cost per 1,000 pages (lower wins)
$0.305
$0.09
Context
1M tokens
131k tokens
Open weights
no
yes
Tool calling
yes
no
Structured output
yes
yes
Reads images
yes
no
Reasoning
no
no
Released
14 Apr 2025
18 Aug 2024
Providers selling it
32
7
Price we measure against
the maker's own
OpenRouter
Trains on your prompts (no is better)
not stated
not stated
Providers selling both · 4
NanoGPT$2.80 / $1.63
Kilo Gateway$2.80 / $2.80
OpenRouter$2.80 / $2.80
Eden AI$3.08 / $2.80
GPT-4.1 mini then Hermes 3 70B Instruct, per million tokens
Go further
Add a third or fourth model, or switch the weighting, in the tool.
What moved in the last seven days
- GPT-4.1 mini
Nous Portal$0.32 / $1.28 $0.40 / $1.6025% dearer ·
Frequently asked
- Which is cheaper, GPT-4.1 mini or Hermes 3 70B Instruct?
- Hermes 3 70B Instruct, at $0.66 per million tokens against $2.55 for GPT-4.1 mini, three parts input to one part output, read on 16 Sept 2026.
- Where is each one cheapest?
- GPT-4.1 mini: $2.55 at Poe. Hermes 3 70B Instruct: $0.66 at Hyperbolic.
- Can I buy both from one provider?
- Yes. 4 providers sell both: NanoGPT, Kilo Gateway, OpenRouter, Eden AI. The cheapest of them for the two together is NanoGPT.
- Which holds more context?
- GPT-4.1 mini, at 1M tokens against 131k.
Prices exclude tax and are per million tokens, read 16 Sept 2026 from each provider's own catalog. The standard price blends three parts input to one part output; a page is 600 tokens in and 60 out. Bars are scaled to the larger value in each row.