Run Qwen3 30B-A3B (MoE) on a VPS.

prices as of · Contabo as of · 81 of 524 plans fit, disk included · re-ranked daily

The CPU darling: 30B of knowledge, 3B active per token, so it streams like a small model on a 24–32 GB VPS. On a CPU-only VPS it needs about 20.9 GB of RAM: 18.6 GB of Q4_K_M weights, 819 MB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 81 of the 524 plans in our index fit and include a disk; the cheapest comfortable pick is netcup's VPS 4000 G12 · 12 vCore · 32 GB at $31.43/mo, streaming an estimated 6–12 tok/s.

The cheapest VPS that runs Qwen3 30B-A3B (MoE) is netcup VPS 4000 G12 at $31.43/mo, 32 GB of RAM against the 20.9 GB the model needs, checked 16 Sept 2026.

Best VPS plans for Qwen3 30B-A3B (MoE)

ranked by estimated tokens/s per dollar, comfortable fits first

  1. 1
    netcupRS 4000 G12 · 12 dedicated cores · 32 GB

    12 vCPU dedicated · 32 GB RAM · 1024 GB NVMe

    runs comfortably~9–18 tok/s$38.71/moView at netcup
  2. 2
    ContaboCloud VPS 12 · 12 vCPU · 48 GB

    12 vCPU shared · 48 GB RAM · 400 GB SSD

    runs comfortably~6–12 tok/s$28.85/mochecked 10 days agoView at Contabo
  3. 3
    netcupVPS 4000 G12 · 12 vCore · 32 GB

    12 vCPU shared · 32 GB RAM · 1024 GB NVMe

    runs comfortably~6–12 tok/s$31.43/moView at netcup
  4. 4
    HetznerCX53 · 16 vCPU · 32 GB

    16 vCPU shared · 32 GB RAM · 320 GB NVMe

    runs comfortably~6–12 tok/s$34.03/monot available nowView at Hetzner
  5. 5
    HOSTKEYvm.v2-heavy · 8 vCPU · 32 GB

    8 vCPU shared · 32 GB RAM · 240 GB NVMe

    runs comfortably~4.8–9.6 tok/s$32.31/moView at HOSTKEY
  6. 6
    ContaboCloud VPS 16 · 16 vCPU · 64 GB

    16 vCPU shared · 64 GB RAM · 500 GB SSD

    runs comfortably~6–12 tok/s$42.69/mochecked 10 days agoView at Contabo

Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated, ×0.5 for MoE routing overhead), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbors are, so treat these as order-of-magnitude estimates.

RAM needed · CPU inference

20.9 GB

weights · Q4_K_M
18.6 GB
KV cache · 8K context
819 MB
OS + runtime headroom
1.5 GB

Other quants: Q8_0 weights 32.5 GB (near-lossless, about half of F16); F16 61.1 GB. Comfortable = 20% headroom over the total.

Model card

Size
31B total · 3.3B active
Context
40K tokens
Kind
chat
Released
2025-04
License
Apache-2.0

Facts fetched from Hugging Face on 7 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 2,104,492 downloads.

ollama run qwen3:30b-a3bmodel card ↗Alibaba

Or buy hardware · 8 reference machines fit

all machines →
  • Mac mini (M6, 32 GB)32 GB · 153 GB/s · ~20–31 tok/s$1,299= 41 mo of VPS
  • Framework Desktop (Ryzen AI Max+ 395, 64 GB)64 GB · 256 GB/s · ~34–52 tok/s$1,659*= 53 mo of VPS
  • Mac mini (M5 Pro, 64 GB)64 GB · 307 GB/s · ~40–62 tok/s$2,299= 73 mo of VPS
  • Mac Studio (M5 Max, 36 GB)36 GB · 460 GB/s · ~60–93 tok/s$2,499= 80 mo of VPS
  • Framework Desktop (Ryzen AI Max+ 395, 128 GB)128 GB · 256 GB/s · ~34–52 tok/s$3,149*= 100 mo of VPS
  • GeForce RTX 5090 32 GB (card only)32 GB · 1792 GB/s · ~100+ tok/s$4,300*= 10+ yrs of VPS

Buy · per month

$37.45

$36.08 hardware + $1.36 power

Rent · per month

$31.43

netcup VPS 4000 G12 · 12 vCore · 32 GB

Break-even

The Mac mini (M6, 32 GB) pays for itself after 43 months of replacing the VPS, and it streams an estimated 20–31 tok/s against the VPS's CPU-only pace.

* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.

Frequently asked

How much RAM does Qwen3 30B-A3B (MoE) need?
About 20.9 GB for CPU inference at Q4_K_M: 18.6 GB of weights, 819 MB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (96 KB per token for this model).
What is the cheapest VPS that can run Qwen3 30B-A3B (MoE)?
netcup VPS 4000 G12 · 12 vCore · 32 GB (32 GB RAM, 12 vCPU) at $31.43/mo excl. VAT runs it comfortably as of 16 Sept 2026. OVHcloud VPS-4 2027 · 8 vCPU / 24 GB · EU at $27.11/mo is a tight fit.
How fast will Qwen3 30B-A3B (MoE) run on a VPS without a GPU?
Roughly 9–18 tok/s on the top pick. Token generation is limited by memory bandwidth, because each new token reads the active experts' weights once. More vCPUs help, dedicated ones most, but a GPU or an Apple silicon machine is 10 to 50 times faster.
Same lineupQwen3 0.6BQwen3 1.7BQwen3 4BQwen3 8BQwen3 14BQwen3 32BQwen2.5-Coder 7BQwen2.5-Coder 32BSame size classLlama 4 Scout (109B MoE)gpt-oss-20b (MoE)gpt-oss-120b (MoE)VPS for AI agents (API-hosted models) →