GLM-5.3-Flash vs Llama 3.2 3B Instruct.
model vs model · prices as of
GLM-5.3-Flash takes 3 of 5 rounds against Llama 3.2 3B Instruct. Llama 3.2 3B Instruct counters on price: $0.08 per million tokens against $0.10, for 131k context instead of 1M.
GLM-5.3-Flash
$0.10/M tokens
cheapest we can explain, at Merge Gateway · $0.012 per 1,000 pages
1M context · tool calling · reads images · open weights
Llama 3.2 3B Instruct
$0.08/M tokens
cheapest we can explain, at Inference · $0.0132 per 1,000 pages
131k context · open weights
List price in (lower wins)
$0.07/M
$0.05/M
List price out (lower wins)
$0.25/M
$0.33/M
Cheapest we can explain (lower wins)
$0.10/M at Merge Gateway
$0.08/M at Inference
Standard price (lower wins)
$0.47/M
$0.48/M
Cost per 1,000 pages (lower wins)
$0.012
$0.0132
Context
1M tokens
131k tokens
Open weights
yes
yes
Tool calling
yes
no
Structured output
yes
yes
Reads images
yes
no
Reasoning
yes
no
Released
26 Aug 2026
25 Sept 2024
Providers selling it
76
17
Price we measure against
the maker's own
OpenRouter
Trains on your prompts (no is better)
not stated
not stated
Providers selling both · 13
NanoGPT$0.24 / $0.14
LLM Gateway$0.51 / $0.14
Novita$0.84 / $0.14
TrustedRouter$0.47 / $0.51
DeepInfra$0.95 / $0.08
Eden AI$0.95 / $0.08
OpenRouter$0.57 / $0.48
Together AI$0.95 / $0.24
GLM-5.3-Flash then Llama 3.2 3B Instruct, per million tokens · showing the 8 cheapest of 13
Go further
Add a third or fourth model, or switch the weighting, in the tool.
What moved in the last seven days
- GLM-5.3-Flash
Cortecs$0.21 / $0.52 $0.10 / $0.3641% cheaper · - GLM-5.3-Flash
GMI Cloud$0.11 / $0.38 $0.10 / $0.357% cheaper · - GLM-5.3-Flash
io.net Intelligence$0.21 / $0.70 $0.21 / $0.690.9% cheaper · - GLM-5.3-Flash
Kenari$0.15 / $0.50 $0.20 / $0.6026% dearer · - GLM-5.3-Flash
NextBit$0.15 / $0.50 $0.25 / $0.9074% dearer · - GLM-5.3-Flash
Nous Portal$0.06 / $0.20 $0.07 / $0.2525% dearer · - GLM-5.3-Flash
OpenRouter$0.15 / $0.50 $0.09 / $0.3040% cheaper · - GLM-5.3-Flash
Reka$0.15 / $0.50 $0.23 / $0.7550% dearer · - GLM-5.3-Flash
Requesty$0.20 / $0.60 $0.15 / $0.5021% cheaper · - GLM-5.3-Flash
StreamLake$0.11 / $0.37 $0.10 / $0.357% cheaper · - GLM-5.3-Flash
Vancine$0.06 / $0.20 $0.12 / $0.40100% dearer · - Llama 3.2 3B Instruct
TrustedRouter$0.03 / $0.05 $0.05 / $0.35248% dearer · - GLM-5.3-Flash
Modal$0.15 / $0.50 $0.45 / $1.50200% dearer ·
Frequently asked
- Which is cheaper, GLM-5.3-Flash or Llama 3.2 3B Instruct?
- Llama 3.2 3B Instruct, at $0.08 per million tokens against $0.10 for GLM-5.3-Flash, three parts input to one part output, read on 16 Sept 2026.
- Where is each one cheapest?
- GLM-5.3-Flash: $0.10 at Merge Gateway. Llama 3.2 3B Instruct: $0.08 at Inference.
- Can I buy both from one provider?
- Yes. 13 providers sell both: NanoGPT, LLM Gateway, Novita, TrustedRouter, DeepInfra, Eden AI, and more. The cheapest of them for the two together is NanoGPT.
- Which holds more context?
- GLM-5.3-Flash, at 1M tokens against 131k.
Prices exclude tax and are per million tokens, read 16 Sept 2026 from each provider's own catalog. The standard price blends three parts input to one part output; a page is 600 tokens in and 60 out. Bars are scaled to the larger value in each row.