DeepInfra model pricing.

prices as of · 106 models · 180 prices

Serves open-weight models on its own GPUs, often at fp8 or fp4, and publishes the precision per model.

DeepInfra is a host selling 106 models, from $0.08 per million tokens for Llama 3.2 3B Instruct, read on 16 Sept 2026.

DeepInfra prices

Kind of provider

Host

based in US

Where it runs

not stated

no regions published

How you pay

not stated

no payment terms published

Free tier

none published

nothing given away

Where DeepInfra's prices land, and why

DeepInfra publishes 180 prices we can compare, and 84 of them are below the list price for that model. 79 of those have a published reason behind them, so they rank like any other price. The other 3 sit below list on models whose weights are closed, with nothing published to account for the gap, so they are listed here in full and ranked nowhere on the site. Nobody here has read its terms yet, so treat what it says about itself as its own claim.

DeepInfra does not say whether it trains on what you send it. It does not say how long it keeps requests.

not yet checked: terms · privacy policy ↗ · status page ↗

Models to start with

all 180 prices ↓

Everything DeepInfra sells, against the list price

Cheapest first inside each group. The share is this row against the price the lab that trained the model charges.

Cheaper, and we know why

79 prices

Below list, with a published reason: hosting the open weights, a regional price list, credits, a stated discount.

  • Llama 3.2 3B Instruct

    Meta · 131k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.02 / $0.02in / out per M

    0.17×of list price

  • Mistral Nemo

    Mistral AI · 128k context

    Serves the open weights at fp8, a lower precision than the maker's, so not quite the same product.

    $0.02 / $0.03in / out per M

    0.14×of list price

  • Llama 3.1 8B Instruct

    Meta · 131k context

    Serves the open weights at fp8, a lower precision than the maker's, so not quite the same product.

    $0.02 / $0.04in / out per M

    0.43×of list price

  • Llama 3.1 8B Instruct

    Meta · 131k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.02 / $0.05in / out per M

    0.48×of list price

  • Qwen2.5 7B Instruct

    Alibaba Cloud · 131k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.04 / $0.10in / out per M

    0.18×of list price

  • Mistral Small 3.2 24B

    Mistral AI · 256k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.07 / $0.20in / out per M

    0.80×of list price

  • Mistral Small 3.2 24B

    Mistral AI · 256k context

    Serves the open weights at fp8, a lower precision than the maker's, so not quite the same product.

    $0.07 / $0.20in / out per M

    0.80×of list price

  • DeepSeek V4 Flash 0731

    DeepSeek · 1M context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.09 / $0.18in / out per M

    0.43×of list price

  • DeepSeek V4 Flash 0731

    DeepSeek · 1M context

    Serves the open weights at fp8, a lower precision than the maker's, so not quite the same product.

    $0.09 / $0.18in / out per M

    0.43×of list price

  • Qwen3 32B

    Alibaba Cloud · 131k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.08 / $0.28in / out per M

    0.11×of list price

  • Qwen3 32B

    Alibaba Cloud · 131k context

    Serves the open weights at fp8, a lower precision than the maker's, so not quite the same product.

    $0.08 / $0.28in / out per M

    0.11×of list price

  • Gemma 4 26B A4B IT

    Google · 262k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.07 / $0.34in / out per M

    0.52×of list price

  • Gemma 4 26B A4B IT

    Google · 262k context

    Serves the open weights at fp8, a lower precision than the maker's, so not quite the same product.

    $0.07 / $0.34in / out per M

    0.52×of list price

  • Qwen3 14B

    Alibaba Cloud · 131k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.12 / $0.24in / out per M

    0.24×of list price

  • Qwen3 14B

    Alibaba Cloud · 131k context

    Serves the open weights at fp8, a lower precision than the maker's, so not quite the same product.

    $0.12 / $0.24in / out per M

    0.24×of list price

  • Nemotron 3 Super

    NVIDIA · 262k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision bf16.

    $0.09 / $0.40in / out per M

    0.47×of list price

  • Qwen3 235B-A22B

    Alibaba Cloud · 131k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.09 / $0.55in / out per M

    0.17×of list price

  • Qwen3 235B-A22B

    Alibaba Cloud · 131k context

    Serves the open weights at fp8, a lower precision than the maker's, so not quite the same product.

    $0.09 / $0.55in / out per M

    0.17×of list price

  • Qwen3 30B A3B Instruct 2507

    Alibaba Cloud · 262k context

    8% under the list price, an ordinary reseller margin.

    $0.12 / $0.50in / out per M

    0.95×of list price

  • Qwen3 30B A3B Instruct 2507

    Alibaba Cloud · 262k context

    8% under the list price, an ordinary reseller margin.

    $0.12 / $0.50in / out per M

    0.95×of list price

  • Qwen3-VL 30B-A3B

    Alibaba Cloud · 131k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.15 / $0.60in / out per M

    0.75×of list price

  • Qwen3-VL 30B-A3B

    Alibaba Cloud · 131k context

    Serves the open weights at fp8, a lower precision than the maker's, so not quite the same product.

    $0.15 / $0.60in / out per M

    0.75×of list price

  • DeepSeek V3.2

    DeepSeek · 164k context

    7% under the list price, an ordinary reseller margin.

    $0.26 / $0.38in / out per M

    0.94×of list price

  • DeepSeek V3.2

    DeepSeek · 164k context

    7% under the list price, an ordinary reseller margin.

    $0.26 / $0.38in / out per M

    0.94×of list price

  • R1 Distill Llama 70B

    DeepSeek · 8k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.20 / $0.60in / out per M

    0.38×of list price

  • Qwen3.6 35B-A3B

    Alibaba Cloud · 262k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.10 / $0.95in / out per M

    0.56×of list price

  • Qwen3.6 35B-A3B

    Alibaba Cloud · 262k context

    Serves the open weights at fp8, a lower precision than the maker's, so not quite the same product.

    $0.10 / $0.95in / out per M

    0.56×of list price

  • Qwen3-Next 80B-A3B Instruct

    Alibaba Cloud · 131k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.09 / $1.10in / out per M

    0.39×of list price

  • Qwen3-Next 80B-A3B Instruct

    Alibaba Cloud · 131k context

    Serves the open weights at fp8, a lower precision than the maker's, so not quite the same product.

    $0.09 / $1.10in / out per M

    0.39×of list price

  • Qwen3.5 35B-A3B

    Alibaba Cloud · 262k context

    Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.

    $0.14 / $1in / out per M

    0.52×of list price

showing the 30 cheapest of 79 · every row in the open dataset

Same price as the maker

98 prices

The same price the maker charges, sold by someone else.

  • gpt-oss-20b

    OpenAI · 131k context

    The same price OpenRouter charges.

    $0.03 / $0.14in / out per M

    1.05×of list price

  • gpt-oss-20b

    OpenAI · 131k context

    The same price OpenRouter charges.

    $0.03 / $0.14in / out per M

    1.05×of list price

  • gpt-oss-20b

    OpenAI · 131k context

    The same price OpenRouter charges.

    $0.03 / $0.14in / out per M

    1.05×of list price

  • Mistral Small 3

    Mistral AI · 33k context

    The same price OpenRouter charges.

    $0.05 / $0.08in / out per M

    1.00×of list price

  • Mistral Small 3

    Mistral AI · 33k context

    The same price OpenRouter charges.

    $0.05 / $0.08in / out per M

    1.00×of list price

  • Gemma 3 4B

    Google · 131k context

    The same price OpenRouter charges.

    $0.05 / $0.10in / out per M

    1.00×of list price

  • Gemma 3 4B

    Google · 131k context

    The same price OpenRouter charges.

    $0.05 / $0.10in / out per M

    1.00×of list price

  • Gemma 3 4B

    Google · 131k context

    The same price OpenRouter charges.

    $0.05 / $0.10in / out per M

    1.00×of list price

  • gpt-oss-120b

    OpenAI · 131k context

    The same price OpenRouter charges.

    $0.04 / $0.17in / out per M

    1.00×of list price

  • gpt-oss-120b

    OpenAI · 131k context

    The same price OpenRouter charges.

    $0.04 / $0.17in / out per M

    1.00×of list price

  • gpt-oss-120b

    OpenAI · 131k context

    The same price OpenRouter charges.

    $0.04 / $0.17in / out per M

    1.00×of list price

  • Gemma 3 12B

    Google · 131k context

    The same price OpenRouter charges.

    $0.05 / $0.15in / out per M

    1.00×of list price

  • Gemma 3 12B

    Google · 131k context

    The same price OpenRouter charges.

    $0.05 / $0.15in / out per M

    1.00×of list price

  • Gemma 3 12B

    Google · 131k context

    The same price OpenRouter charges.

    $0.05 / $0.15in / out per M

    1.00×of list price

  • Nemotron 3 Nano 30B A3B

    NVIDIA · 262k context

    The same price OpenRouter charges.

    $0.05 / $0.20in / out per M

    1.00×of list price

  • Nemotron 3 Nano 30B A3B

    NVIDIA · 262k context

    The same price OpenRouter charges.

    $0.05 / $0.20in / out per M

    1.00×of list price

  • Phi 4

    Microsoft · 16k context

    The same price OpenRouter charges.

    $0.07 / $0.14in / out per M

    1.00×of list price

  • Phi 4

    Microsoft · 16k context

    The same price OpenRouter charges.

    $0.07 / $0.14in / out per M

    1.00×of list price

  • Phi 4

    Microsoft · 16k context

    The same price OpenRouter charges.

    $0.07 / $0.14in / out per M

    1.00×of list price

  • Ling 3.0 Flash

    InclusionAI · 262k context

    186% above OpenRouter's price.

    $0.06 / $0.18in / out per M

    2.86×of list price

  • Ling 3.0 Flash

    InclusionAI · 262k context

    186% above OpenRouter's price.

    $0.06 / $0.18in / out per M

    2.86×of list price

  • Ling 3.0 Flash Fin

    InclusionAI · 262k context

    The same price OpenRouter charges.

    $0.06 / $0.18in / out per M

    1.00×of list price

  • Ling 3.0 Flash Fin

    InclusionAI · 262k context

    The same price OpenRouter charges.

    $0.06 / $0.18in / out per M

    1.00×of list price

  • Ling 3.0 Flash VL

    InclusionAI · 131k context

    The same price OpenRouter charges.

    $0.06 / $0.18in / out per M

    1.00×of list price

  • Ling 3.0 Flash VL

    InclusionAI · 131k context

    The same price OpenRouter charges.

    $0.06 / $0.18in / out per M

    1.00×of list price

  • Gemma 3 27B

    Google · 131k context

    The same price OpenRouter charges.

    $0.08 / $0.16in / out per M

    0.58×of list price

  • Gemma 3 27B

    Google · 131k context

    The same price OpenRouter charges.

    $0.08 / $0.16in / out per M

    0.58×of list price

  • Nemotron 3.5 Lightning

    NVIDIA · 262k context

    The same price OpenRouter charges.

    $0.08 / $0.20in / out per M

    1.00×of list price

  • Qwen3.5-9B

    Alibaba Cloud · 262k context

    The same price OpenRouter charges.

    $0.10 / $0.15in / out per M

    1.00×of list price

  • Qwen3.5-9B

    Alibaba Cloud · 262k context

    The same price OpenRouter charges.

    $0.10 / $0.15in / out per M

    1.00×of list price

showing the 30 cheapest of 98 · every row in the open dataset

Cheaper, and nobody says why

3 prices · never ranked

Below list on a closed model with nothing to account for it. Shown, never ranked, never a pick.

  • Mistral Nemo

    Mistral AI · 128k context

    0.13× the list price, from one source and far below what hosting the weights costs. Listed until a second source agrees.

    $0.02 / $0.03in / out per M

    0.14×of list price

  • Qwen3.8 Flash

    Alibaba Cloud · 1M context

    0.75× Alibaba Cloud's list price on a closed model, and nobody says why. Listed, never ranked.

    $0.11 / $0.38in / out per M

    0.78×of list price

  • Qwen3.8 Max

    Alibaba Cloud · 1M context

    0.82× Alibaba Cloud's list price on a closed model, and nobody says why. Listed, never ranked.

    $1.65 / $4.95in / out per M

    0.83×of list price

Frequently asked

What is DeepInfra?
Serves open-weight models on its own GPUs, often at fp8 or fp4, and publishes the precision per model. It sells 106 of the models we track, and trades as Deep Infra, Inc., based in US.
Is DeepInfra cheaper than buying from the maker?
84 of its 180 prices sit below the list price for the model, and 79 of those have a published reason behind them, so they can lead a ranking. 3 are listed and never ranked.
Does DeepInfra train on what you send it?
DeepInfra does not say whether it trains on what you send it. It does not say how long it keeps requests.
How current are these DeepInfra prices?
They come from its own public catalog, last read 16 Sept 2026. Prices exclude tax.