Run DeepSeek-R1 Distill Qwen 14B on a VPS.

prices as of · Contabo as of · 178 of 524 plans fit, disk included · re-ranked daily

Balanced reasoning distill for math and code on a 16 GB plan. On a CPU-only VPS it needs about 12 GB of RAM: 9 GB of Q4_K_M weights, 1.5 GB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 178 of the 524 plans in our index fit and include a disk; the cheapest comfortable pick is HOSTKEY's vm.v2-medium · 8 vCPU · 16 GB at $16.15/mo, streaming an estimated 2.1–4.3 tok/s.

The cheapest VPS that runs DeepSeek-R1 Distill Qwen 14B is HOSTKEY vm.v2-medium at $16.15/mo, 16 GB of RAM against the 12 GB the model needs, checked 16 Sept 2026.

Best VPS plans for DeepSeek-R1 Distill Qwen 14B

ranked by estimated tokens/s per dollar, comfortable fits first

  1. 1
    netcupRS 2000 G12 · 8 dedicated cores · 16 GB

    8 vCPU dedicated · 16 GB RAM · 512 GB NVMe

    runs comfortably~3.2–6.4 tok/s$20.78/moView at netcup
  2. 2
    ContaboCloud VPS 8 · 8 vCPU · 24 GB

    8 vCPU shared · 24 GB RAM · 300 GB SSD

    runs comfortably~2.1–4.3 tok/s$16.15/mochecked 10 days agoView at Contabo
  3. 3
    HOSTKEYvm.v2-medium · 8 vCPU · 16 GB

    8 vCPU shared · 16 GB RAM · 160 GB NVMe

    runs comfortably~2.1–4.3 tok/s$16.15/moView at HOSTKEY
  4. 4
    HetznerCX43 · 8 vCPU · 16 GB

    8 vCPU shared · 16 GB RAM · 160 GB NVMe

    runs comfortably~2.1–4.3 tok/s$18.45/monot available nowView at Hetzner
  5. 5
    HOSTKEYvm.v3-medium · 8 vCPU · 16 GB

    8 vCPU shared · 16 GB RAM · 160 GB NVMe

    runs comfortably~2.1–4.3 tok/s$18.46/moView at HOSTKEY
  6. 6
    netcupVPS 2000 G12 · 8 vCore · 16 GB

    8 vCPU shared · 16 GB RAM · 512 GB NVMe

    runs comfortably~2.1–4.3 tok/s$18.67/moView at netcup

Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbors are, so treat these as order-of-magnitude estimates.

RAM needed · CPU inference

12 GB

weights · Q4_K_M
9 GB
KV cache · 8K context
1.5 GB
OS + runtime headroom
1.5 GB

Other quants: Q8_0 weights 15.7 GB (near-lossless, about half of F16); F16 29.6 GB. Comfortable = 20% headroom over the total.

Model card

Size
15B parameters
Context
128K tokens
Kind
reasoning
Released
2025-01
License
MIT

Facts fetched from Hugging Face on 7 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 362,613 downloads.

ollama run deepseek-r1:14bmodel card ↗DeepSeek

Or buy hardware · 14 reference machines fit

all machines →
  • Raspberry Pi 5 (16 GB)16 GB · 17 GB/s · ~0.5–0.8 tok/s$305*= 19 mo of VPS
  • GeForce RTX 5060 Ti 16 GB (card only)16 GB · 448 GB/s · ~26–40 tok/s$800*= 50 mo of VPS
  • Mac mini (M6, 16 GB)16 GB · 153 GB/s · ~9–14 tok/s · tight$899= 56 mo of VPS
  • GeForce RTX 5070 Ti 16 GB (card only)16 GB · 896 GB/s · ~53–81 tok/s$1,200*= 74 mo of VPS
  • Mac mini (M6, 32 GB)32 GB · 153 GB/s · ~9–14 tok/s$1,299= 80 mo of VPS
  • GeForce RTX 5080 16 GB (card only)16 GB · 960 GB/s · ~56–86 tok/s$1,600*= 99 mo of VPS

Buy · per month

$8.78

$8.47 hardware + $0.31 power

Rent · per month

$16.15

HOSTKEY vm.v2-medium · 8 vCPU · 16 GB

Break-even

The Raspberry Pi 5 (16 GB) pays for itself after 19 months of replacing the VPS, and it streams an estimated 0.5–0.8 tok/s against the VPS's CPU-only pace.

* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.

Frequently asked

How much RAM does DeepSeek-R1 Distill Qwen 14B need?
About 12 GB for CPU inference at Q4_K_M: 9 GB of weights, 1.5 GB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (192 KB per token for this model).
What is the cheapest VPS that can run DeepSeek-R1 Distill Qwen 14B?
HOSTKEY vm.v2-medium · 8 vCPU · 16 GB (16 GB RAM, 8 vCPU) at $16.15/mo excl. VAT runs it comfortably as of 16 Sept 2026. HOSTKEY vm.medium · 6 vCPU · 12 GB at $11.70/mo is a tight fit.
How fast will DeepSeek-R1 Distill Qwen 14B run on a VPS without a GPU?
Roughly 3.2–6.4 tok/s on the top pick. Token generation is limited by memory bandwidth, because each new token reads all of the weights once. More vCPUs help, dedicated ones most, but a GPU or an Apple silicon machine is 10 to 50 times faster.
Same lineupDeepSeek-R1-0528 Qwen3 8BDeepSeek-R1 Distill Qwen 32BDeepSeek-R1 Distill Llama 70BSame size classQwen3 14BGemma 3 12BMistral Nemo 12BPhi-4 14BVPS for AI agents (API-hosted models) →