Run DeepSeek-R1 Distill Llama 70B on a VPS.
prices as of · Contabo as of · 22 of 524 plans fit, disk included · re-ranked daily
Top-tier open reasoning; a 64 GB dedicated box is the entry ticket. On a CPU-only VPS it needs about 46.5 GB of RAM: 42.5 GB of Q4_K_M weights, 2.5 GB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 22 of the 524 plans in our index fit and include a disk; the cheapest comfortable pick is netcup's VPS 8000 G12 · 16 vCore · 64 GB at $46.49/mo, streaming an estimated 0.6–1.1 tok/s.
The cheapest VPS that runs DeepSeek-R1 Distill Llama 70B is netcup VPS 8000 G12 at $46.49/mo, 64 GB of RAM against the 46.5 GB the model needs, checked 16 Sept 2026.
Best VPS plans for DeepSeek-R1 Distill Llama 70B
ranked by estimated tokens/s per dollar, comfortable fits first
- 1runs comfortably~0.8–1.7 tok/s$69.20/moView at netcup ↗
- 2runs comfortably~0.6–1.1 tok/s$42.69/mochecked 10 days agoView at Contabo ↗
- 3runs comfortably~0.6–1.1 tok/s$46.49/moView at netcup ↗
- 4runs comfortably~0.8–1.7 tok/s$102.70/moView at Cherry Servers ↗
- 5runs comfortably~0.8–1.7 tok/s$125.78/moView at Cherry Servers ↗
- 6runs comfortably~0.6–1.1 tok/s$91.16/mochecked 10 days agoView at Contabo ↗
Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbors are, so treat these as order-of-magnitude estimates.
RAM needed · CPU inference
46.5 GB
- weights · Q4_K_M
- 42.5 GB
- KV cache · 8K context
- 2.5 GB
- OS + runtime headroom
- 1.5 GB
Other quants: Q8_0 weights 75 GB (near-lossless, about half of F16); F16 141.1 GB. Comfortable = 20% headroom over the total.
Model card
- Size
- 71B parameters
- Context
- 128K tokens
- Kind
- reasoning
- Released
- 2025-01
- License
- MIT
Facts fetched from Hugging Face on 7 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 81,732 downloads.
Or rent it as an API · 18 providers sell R1 Distill Llama 70B
every price for R1 Distill Llama 70B →DeepSeek list price
$3.20/M tokens
Cheapest price we can explain
$1.20/M tokens
DeepInfra, read 16 Sept 2026
What the VPS price buys
39M tokens/mo
at $46.49/mo
Token prices are per million, three parts input to one part output, and exclude tax. A rented box costs the same whether or not you use it and runs whatever else you put on it; an API charges for what you send and nothing when you stop.
Or buy hardware · 5 reference machines fit
all machines →- Framework Desktop (Ryzen AI Max+ 395, 64 GB)64 GB · 256 GB/s · ~3.1–4.8 tok/s · tight$1,659*= 36 mo of VPS
- Mac mini (M5 Pro, 64 GB)64 GB · 307 GB/s · ~3.8–5.8 tok/s · tight$2,299= 50 mo of VPS
- Framework Desktop (Ryzen AI Max+ 395, 128 GB)128 GB · 256 GB/s · ~3.1–4.8 tok/s$3,149*= 68 mo of VPS
- NVIDIA DGX Spark (128 GB)128 GB · 273 GB/s · ~3.4–5.1 tok/s$4,699*= 101 mo of VPS
- Mac Studio (M5 Ultra, 96 GB)96 GB · 1200 GB/s · ~15–23 tok/s$5,499= 118 mo of VPS
Buy · per month
$50.75
$46.08 hardware + $4.67 power
Rent · per month
$46.49
netcup VPS 8000 G12 · 16 vCore · 64 GB
Break-even
The Framework Desktop (Ryzen AI Max+ 395, 64 GB) pays for itself after 40 months of replacing the VPS, and it streams an estimated 3.1–4.8 tok/s against the VPS's CPU-only pace.
* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.
Frequently asked
- How much RAM does DeepSeek-R1 Distill Llama 70B need?
- About 46.5 GB for CPU inference at Q4_K_M: 42.5 GB of weights, 2.5 GB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (320 KB per token for this model).
- What is the cheapest VPS that can run DeepSeek-R1 Distill Llama 70B?
- netcup VPS 8000 G12 · 16 vCore · 64 GB (64 GB RAM, 16 vCPU) at $46.49/mo excl. VAT runs it comfortably as of 16 Sept 2026.
- How fast will DeepSeek-R1 Distill Llama 70B run on a VPS without a GPU?
- Roughly 0.8–1.7 tok/s on the top pick. Token generation is limited by memory bandwidth, because each new token reads all of the weights once. More vCPUs help, dedicated ones most, but a GPU or an Apple silicon machine is 10 to 50 times faster.