Run Qwen3 0.6B on a VPS.

prices as of · Contabo as of · 303 of 524 plans fit, disk included · re-ranked daily

Sub-gigabyte model for rerankers, intent detection and edge devices. On a CPU-only VPS it needs about 2.9 GB of RAM: 512 MB of Q4_K_M weights, 922 MB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 303 of the 524 plans in our index fit and include a disk; the cheapest comfortable pick is OVHcloud's VPS-1 2027 · 2 vCPU / 4 GB · EU at $5.18/mo, streaming an estimated 11–21 tok/s.

The cheapest VPS that runs Qwen3 0.6B is OVHcloud VPS-1 2027 at $5.18/mo, 4 GB of RAM against the 2.9 GB the model needs, checked 16 Sept 2026.

Best VPS plans for Qwen3 0.6B

ranked by estimated tokens/s per dollar, comfortable fits first

  1. 1
    HOSTKEYvm.mini · 4 vCPU · 6 GB

    4 vCPU shared · 6 GB RAM · 120 GB SSD

    runs comfortably~21–42 tok/s$5.93/moView at HOSTKEY
  2. 2
    ContaboCloud VPS 6 · 6 vCPU · 12 GB

    6 vCPU shared · 12 GB RAM · 200 GB SSD

    runs comfortably~32–63 tok/s$8.65/mochecked 10 days agoView at Contabo
  3. 3
    ContaboCloud VPS 4 · 4 vCPU · 8 GB

    4 vCPU shared · 8 GB RAM · 100 GB SSD

    runs comfortably~21–42 tok/s$6.35/mochecked 10 days agoView at Contabo
  4. 4
    netcupRS 2000 G12 · 8 dedicated cores · 16 GB

    8 vCPU dedicated · 16 GB RAM · 512 GB NVMe

    runs comfortably~63–127 tok/s$20.78/moView at netcup
  5. 5
    HOSTKEYvm.v2-mini · 4 vCPU · 8 GB

    4 vCPU shared · 8 GB RAM · 120 GB NVMe

    runs comfortably~21–42 tok/s$7.58/moView at HOSTKEY
  6. 6
    HOSTKEYvm.medium · 6 vCPU · 12 GB

    6 vCPU shared · 12 GB RAM · 160 GB SSD

    runs comfortably~32–63 tok/s$11.70/moView at HOSTKEY

Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbors are, so treat these as order-of-magnitude estimates.

RAM needed · CPU inference

2.9 GB

weights · Q4_K_M
512 MB
KV cache · 8K context
922 MB
OS + runtime headroom
1.5 GB

Other quants: Q8_0 weights 819 MB (near-lossless, about half of F16); F16 1.5 GB. Comfortable = 20% headroom over the total.

Model card

Size
0.75B parameters
Context
32K tokens
Kind
chat
Released
2025-04
License
Apache-2.0

Facts fetched from Hugging Face on 7 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 21,327,590 downloads.

ollama run qwen3:0.6bmodel card ↗Alibaba

Or buy hardware · 14 reference machines fit

all machines →
  • Raspberry Pi 5 (16 GB)16 GB · 17 GB/s · ~9.8–15 tok/s$305*= 59 mo of VPS
  • GeForce RTX 5060 Ti 16 GB (card only)16 GB · 448 GB/s · ~100+ tok/s$800*= 10+ yrs of VPS
  • Mac mini (M6, 16 GB)16 GB · 153 GB/s · ~100+ tok/s$899= 10+ yrs of VPS
  • GeForce RTX 5070 Ti 16 GB (card only)16 GB · 896 GB/s · ~100+ tok/s$1,200*= 10+ yrs of VPS
  • Mac mini (M6, 32 GB)32 GB · 153 GB/s · ~100+ tok/s$1,299= 10+ yrs of VPS
  • GeForce RTX 5080 16 GB (card only)16 GB · 960 GB/s · ~100+ tok/s$1,600*= 10+ yrs of VPS

Buy · per month

$8.78

$8.47 hardware + $0.31 power

Rent · per month

$5.18

OVHcloud VPS-1 2027 · 2 vCPU / 4 GB · EU

Break-even

The Raspberry Pi 5 (16 GB) pays for itself after 63 months of replacing the VPS, and it streams an estimated 9.8–15 tok/s against the VPS's CPU-only pace.

* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.

Frequently asked

How much RAM does Qwen3 0.6B need?
About 2.9 GB for CPU inference at Q4_K_M: 512 MB of weights, 922 MB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (112 KB per token for this model).
What is the cheapest VPS that can run Qwen3 0.6B?
OVHcloud VPS-1 2027 · 2 vCPU / 4 GB · EU (4 GB RAM, 2 vCPU) at $5.18/mo excl. VAT runs it comfortably as of 16 Sept 2026.
How fast will Qwen3 0.6B run on a VPS without a GPU?
Roughly 21–42 tok/s on the top pick. Token generation is limited by memory bandwidth, because each new token reads all of the weights once. More vCPUs help, dedicated ones most, but a GPU or an Apple silicon machine is 10 to 50 times faster.
Same lineupQwen3 1.7BQwen3 4BQwen3 8BQwen3 14BQwen3 32BQwen3 30B-A3B (MoE)Qwen2.5-Coder 7BQwen2.5-Coder 32BSame size classLlama 3.2 1BLlama 3.2 3BGemma 3 1BGemma 3 4BVPS for AI agents (API-hosted models) →