Run Qwen3 1.7B on a VPS.
prices as of · Contabo as of · 303 of 524 plans fit, disk included · re-ranked daily
Small model with a thinking mode; strong for its size on structured output. On a CPU-only VPS it needs about 3.7 GB of RAM: 1.3 GB of Q4_K_M weights, 922 MB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 303 of the 524 plans in our index fit and include a disk; the cheapest comfortable pick is HOSTKEY's vm.mini · 4 vCPU · 6 GB at $5.93/mo, streaming an estimated 7.8–16 tok/s.
The cheapest VPS that runs Qwen3 1.7B is HOSTKEY vm.mini at $5.93/mo, 6 GB of RAM against the 3.7 GB the model needs, checked 16 Sept 2026.
Best VPS plans for Qwen3 1.7B
ranked by estimated tokens/s per dollar, comfortable fits first
- 1runs comfortably~7.8–16 tok/s$5.93/moView at HOSTKEY ↗
- 2runs comfortably~12–23 tok/s$8.65/mochecked 10 days agoView at Contabo ↗
- 3runs comfortably~7.8–16 tok/s$6.35/mochecked 10 days agoView at Contabo ↗
- 4runs comfortably~23–47 tok/s$20.78/moView at netcup ↗
- 5runs comfortably~7.8–16 tok/s$7.58/moView at HOSTKEY ↗
- 6runs comfortably~12–23 tok/s$11.70/moView at HOSTKEY ↗
Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbors are, so treat these as order-of-magnitude estimates.
RAM needed · CPU inference
3.7 GB
- weights · Q4_K_M
- 1.3 GB
- KV cache · 8K context
- 922 MB
- OS + runtime headroom
- 1.5 GB
Other quants: Q8_0 weights 2.2 GB (near-lossless, about half of F16); F16 4.1 GB. Comfortable = 20% headroom over the total.
Model card
- Size
- 2.03B parameters
- Context
- 32K tokens
- Kind
- chat
- Released
- 2025-04
- License
- Apache-2.0
Facts fetched from Hugging Face on 7 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 3,414,459 downloads.
Or buy hardware · 14 reference machines fit
all machines →- Raspberry Pi 5 (16 GB)16 GB · 17 GB/s · ~3.6–5.6 tok/s$305*= 51 mo of VPS
- GeForce RTX 5060 Ti 16 GB (card only)16 GB · 448 GB/s · ~100+ tok/s$800*= 10+ yrs of VPS
- Mac mini (M6, 16 GB)16 GB · 153 GB/s · ~65–100 tok/s$899= 10+ yrs of VPS
- GeForce RTX 5070 Ti 16 GB (card only)16 GB · 896 GB/s · ~100+ tok/s$1,200*= 10+ yrs of VPS
- Mac mini (M6, 32 GB)32 GB · 153 GB/s · ~65–100 tok/s$1,299= 10+ yrs of VPS
- GeForce RTX 5080 16 GB (card only)16 GB · 960 GB/s · ~100+ tok/s$1,600*= 10+ yrs of VPS
Buy · per month
$8.78
$8.47 hardware + $0.31 power
Rent · per month
$5.93
HOSTKEY vm.mini · 4 vCPU · 6 GB
Break-even
The Raspberry Pi 5 (16 GB) pays for itself after 54 months of replacing the VPS, and it streams an estimated 3.6–5.6 tok/s against the VPS's CPU-only pace.
* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.
Frequently asked
- How much RAM does Qwen3 1.7B need?
- About 3.7 GB for CPU inference at Q4_K_M: 1.3 GB of weights, 922 MB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (112 KB per token for this model).
- What is the cheapest VPS that can run Qwen3 1.7B?
- HOSTKEY vm.mini · 4 vCPU · 6 GB (6 GB RAM, 4 vCPU) at $5.93/mo excl. VAT runs it comfortably as of 16 Sept 2026. OVHcloud VPS-1 2027 · 2 vCPU / 4 GB · EU at $5.18/mo is a tight fit.
- How fast will Qwen3 1.7B run on a VPS without a GPU?
- Roughly 7.8–16 tok/s on the top pick. Token generation is limited by memory bandwidth, because each new token reads all of the weights once. More vCPUs help, dedicated ones most, but a GPU or an Apple silicon machine is 10 to 50 times faster.