Vultr Serverless Inference model pricing.
prices as of · 8 models · 8 prices
Serves open-weight models on its own GPUs.
Vultr Serverless Inference is a cloud platform selling 8 models, from $1.90 per million tokens for DeepSeek V4 Flash 0731, read on 16 Sept 2026.
Kind of provider
Cloud platform
based in US
Where it runs
not stated
no regions published
How you pay
not stated
no payment terms published
Free tier
none published
nothing given away
Where Vultr Serverless Inference's prices land, and why
Vultr Serverless Inference publishes 8 prices we can compare, and 4 of them are below the list price for that model. 4 of those have a published reason behind them, so they rank like any other price. Nobody here has read its terms yet, so treat what it says about itself as its own claim.
Vultr Serverless Inference does not say whether it trains on what you send it. It does not say how long it keeps requests.
not yet checked: terms
Models to start with
all 8 prices ↓- DeepSeek V4 Flash 0731DeepSeek · 1M context · 97 providers sell it$0.30 / $181% above listbuy at Vultr Serverless Inference ↗
- GLM-5.2Z.ai · 1M context · 94 providers sell it$0.85 / $3.1034% under listbuy at Vultr Serverless Inference ↗
- Kimi K2.6Moonshot AI · 262k context · 82 providers sell it$0.30 / $1.2069% under listbuy at Vultr Serverless Inference ↗
- DeepSeek V3.2DeepSeek · 164k context · 54 providers sell it$0.55 / $1.65166% above listbuy at Vultr Serverless Inference ↗
- MiniMax-M2.7MiniMax · 205k context · 49 providers sell it$0.30 / $1.20same as listbuy at Vultr Serverless Inference ↗
Everything Vultr Serverless Inference sells, against the list price
Cheapest first inside each group. The share is this row against the price the lab that trained the model charges.
Cheaper, and we know why
4 pricesBelow list, with a published reason: hosting the open weights, a regional price list, credits, a stated discount.
- Kimi K2.6
Moonshot AI · 262k context
Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.
$0.30 / $1.20in / out per M
0.31×of list price
- Qwen3.5 397B-A17B
Alibaba Cloud · 262k context
Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.
$0.30 / $2in / out per M
0.54×of list price
- Qwen3.6 27B
Alibaba Cloud · 262k context
Hosts the open weights on its own servers, so the price is its own, not a resale. Precision not disclosed.
$0.30 / $2in / out per M
0.54×of list price
- GLM-5.2
Z.ai · 1M context
Serves the open weights at fp8, a lower precision than the maker's, so not quite the same product.
$0.85 / $3.10in / out per M
0.66×of list price
Same price as the maker
4 pricesThe same price the maker charges, sold by someone else.
$0.30 / $1in / out per M
1.81×of list price
$0.30 / $1.20in / out per M
1.00×of list price
$0.55 / $1.65in / out per M
2.66×of list price
$0.55 / $1.65in / out per M
1.52×of list price
Frequently asked
- What is Vultr Serverless Inference?
- Serves open-weight models on its own GPUs. It sells 8 of the models we track, and trades as The Constant Company, LLC, based in US.
- Is Vultr Serverless Inference cheaper than buying from the maker?
- 4 of its 8 prices sit below the list price for the model, and 4 of those have a published reason behind them, so they can lead a ranking.
- Does Vultr Serverless Inference train on what you send it?
- Vultr Serverless Inference does not say whether it trains on what you send it. It does not say how long it keeps requests.
- How current are these Vultr Serverless Inference prices?
- They come from the public catalogs we read, last read 16 Sept 2026. Prices exclude tax.