Gemini 3.7 Flash vs Hermes 4 405B.
model vs model · prices as of
Gemini 3.7 Flash and Hermes 4 405B split the 5 rounds this table counts: the two list price rows, the cheapest price we can explain, the standard price and context. The weighting you pick decides the rest.
Gemini 3.7 Flash
$6/M tokens
Google list price, nothing below it is accounted for · $0.675 per 1,000 pages
1M context · tool calling · reads images
Hermes 4 405B
$2.10/M tokens
cheapest we can explain, at NanoGPT · $0.252 per 1,000 pages
131k context · open weights
List price in (lower wins)
$0.75/M
$1/M
List price out (lower wins)
$3.75/M
$3/M
Cheapest we can explain (lower wins)
none below list
$2.10/M at NanoGPT
Standard price (lower wins)
$6/M
$6/M
Cost per 1,000 pages (lower wins)
$0.675
$0.252
Context
1M tokens
131k tokens
Open weights
no
yes
Tool calling
yes
no
Structured output
yes
yes
Reads images
yes
no
Reasoning
yes
yes
Released
13 Aug 2026
26 Aug 2025
Providers selling it
31
8
Price we measure against
the maker's own
OpenRouter
Trains on your prompts (no is better)
no provider to check
not stated
Providers selling both · 7
NanoGPT$6 / $2.10
Eden AI$6 / $6
Kilo Gateway$6 / $6
OpenRouter$6 / $6
Requesty$6 / $6
Cortecs$6.21 / $6.19
TrustedRouter$6.33 / $6.33
Gemini 3.7 Flash then Hermes 4 405B, per million tokens
Go further
Add a third or fourth model, or switch the weighting, in the tool.
What moved in the last seven days
- Gemini 3.7 Flash
Nous Portal$0.60 / $3 $0.75 / $3.7525% dearer · - Gemini 3.7 Flash
Requesty$0.82 / $4.13 $0.75 / $3.759% cheaper ·
Frequently asked
- Which is cheaper, Gemini 3.7 Flash or Hermes 4 405B?
- Hermes 4 405B, at $2.10 per million tokens against $6 for Gemini 3.7 Flash, three parts input to one part output, read on 16 Sept 2026.
- Where is each one cheapest?
- Gemini 3.7 Flash: nobody goes below the maker's own price with a reason we can point at, so $6 is the number to beat. Hermes 4 405B: $2.10 at NanoGPT.
- Can I buy both from one provider?
- Yes. 7 providers sell both: NanoGPT, Eden AI, Kilo Gateway, OpenRouter, Requesty, Cortecs, and more. The cheapest of them for the two together is NanoGPT.
- Which holds more context?
- Gemini 3.7 Flash, at 1M tokens against 131k.
Prices exclude tax and are per million tokens, read 16 Sept 2026 from each provider's own catalog. The standard price blends three parts input to one part output; a page is 600 tokens in and 60 out. Bars are scaled to the larger value in each row.