GLM-5.3-Flash vs Llama 3.2 3B Instruct.

model vs model · prices as of

GLM-5.3-Flash takes 3 of 5 rounds against Llama 3.2 3B Instruct. Llama 3.2 3B Instruct counters on price: $0.08 per million tokens against $0.10, for 131k context instead of 1M.

Z.aiwins 3 rounds

GLM-5.3-Flash

$0.10/M tokens

cheapest we can explain, at Merge Gateway · $0.012 per 1,000 pages

1M context · tool calling · reads images · open weights

Meta

Llama 3.2 3B Instruct

$0.08/M tokens

cheapest we can explain, at Inference · $0.0132 per 1,000 pages

131k context · open weights

List price in (lower wins)

$0.07/M

$0.05/M

List price out (lower wins)

$0.25/M

$0.33/M

Cheapest we can explain (lower wins)

$0.10/M at Merge Gateway

$0.08/M at Inference

Standard price (lower wins)

$0.47/M

$0.48/M

Cost per 1,000 pages (lower wins)

$0.012

$0.0132

Context

1M tokens

131k tokens

Open weights

yes

yes

Tool calling

yes

no

Structured output

yes

yes

Reads images

yes

no

Reasoning

yes

no

Released

26 Aug 2026

25 Sept 2024

Providers selling it

76

17

Price we measure against

the maker's own

OpenRouter

Trains on your prompts (no is better)

not stated

not stated

Providers selling both · 13

  • NanoGPT$0.24 / $0.14
  • LLM Gateway$0.51 / $0.14
  • Novita$0.84 / $0.14
  • TrustedRouter$0.47 / $0.51
  • DeepInfra$0.95 / $0.08
  • Eden AI$0.95 / $0.08
  • OpenRouter$0.57 / $0.48
  • Together AI$0.95 / $0.24

GLM-5.3-Flash then Llama 3.2 3B Instruct, per million tokens · showing the 8 cheapest of 13

Go further

Add a third or fourth model, or switch the weighting, in the tool.

What moved in the last seven days

  1. GLM-5.3-FlashCortecs$0.21 / $0.52 $0.10 / $0.3641% cheaper ·
  2. GLM-5.3-FlashGMI Cloud$0.11 / $0.38 $0.10 / $0.357% cheaper ·
  3. GLM-5.3-Flashio.net Intelligence$0.21 / $0.70 $0.21 / $0.690.9% cheaper ·
  4. GLM-5.3-FlashKenari$0.15 / $0.50 $0.20 / $0.6026% dearer ·
  5. GLM-5.3-FlashNextBit$0.15 / $0.50 $0.25 / $0.9074% dearer ·
  6. GLM-5.3-FlashNous Portal$0.06 / $0.20 $0.07 / $0.2525% dearer ·
  7. GLM-5.3-FlashOpenRouter$0.15 / $0.50 $0.09 / $0.3040% cheaper ·
  8. GLM-5.3-FlashReka$0.15 / $0.50 $0.23 / $0.7550% dearer ·
  9. GLM-5.3-FlashRequesty$0.20 / $0.60 $0.15 / $0.5021% cheaper ·
  10. GLM-5.3-FlashStreamLake$0.11 / $0.37 $0.10 / $0.357% cheaper ·
  11. GLM-5.3-FlashVancine$0.06 / $0.20 $0.12 / $0.40100% dearer ·
  12. Llama 3.2 3B InstructTrustedRouter$0.03 / $0.05 $0.05 / $0.35248% dearer ·
  13. GLM-5.3-FlashModal$0.15 / $0.50 $0.45 / $1.50200% dearer ·

Frequently asked

Which is cheaper, GLM-5.3-Flash or Llama 3.2 3B Instruct?
Llama 3.2 3B Instruct, at $0.08 per million tokens against $0.10 for GLM-5.3-Flash, three parts input to one part output, read on 16 Sept 2026.
Where is each one cheapest?
GLM-5.3-Flash: $0.10 at Merge Gateway. Llama 3.2 3B Instruct: $0.08 at Inference.
Can I buy both from one provider?
Yes. 13 providers sell both: NanoGPT, LLM Gateway, Novita, TrustedRouter, DeepInfra, Eden AI, and more. The cheapest of them for the two together is NanoGPT.
Which holds more context?
GLM-5.3-Flash, at 1M tokens against 131k.

Prices exclude tax and are per million tokens, read 16 Sept 2026 from each provider's own catalog. The standard price blends three parts input to one part output; a page is 600 tokens in and 60 out. Bars are scaled to the larger value in each row.

Related head-to-headsGLM-5.3-Flash vs Qwen-Omni TurboLlama 3.2 3B Instruct vs Qwen-Omni TurboMorecompare any two modelsevery model we track