Run Llama 3.3 70B on a VPS.

prices as of · Contabo as of · 22 of 524 plans fit, disk included · re-ranked daily

Near-frontier quality from a dense 70B; needs a 64 GB box and patience on CPU. On a CPU-only VPS it needs about 46.5 GB of RAM: 42.5 GB of Q4_K_M weights, 2.5 GB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 22 of the 524 plans in our index fit and include a disk; the cheapest comfortable pick is netcup's VPS 8000 G12 · 16 vCore · 64 GB at $46.49/mo, streaming an estimated 0.6–1.1 tok/s.

The cheapest VPS that runs Llama 3.3 70B is netcup VPS 8000 G12 at $46.49/mo, 64 GB of RAM against the 46.5 GB the model needs, checked 16 Sept 2026.

Best VPS plans for Llama 3.3 70B

ranked by estimated tokens/s per dollar, comfortable fits first

  1. 1
    netcupRS 8000 G12 · 16 dedicated cores · 64 GB

    16 vCPU dedicated · 64 GB RAM · 2048 GB NVMe

    runs comfortably~0.8–1.7 tok/s$69.20/moView at netcup
  2. 2
    ContaboCloud VPS 16 · 16 vCPU · 64 GB

    16 vCPU shared · 64 GB RAM · 500 GB SSD

    runs comfortably~0.6–1.1 tok/s$42.69/mochecked 10 days agoView at Contabo
  3. 3
    netcupVPS 8000 G12 · 16 vCore · 64 GB

    16 vCPU shared · 64 GB RAM · 2048 GB NVMe

    runs comfortably~0.6–1.1 tok/s$46.49/moView at netcup
  4. 4
    Cherry ServersCloud ARM VDS 12 · 12 vCPU · 56 GB

    12 vCPU dedicated · 56 GB RAM · 300 GB NVMe

    runs comfortably~0.8–1.7 tok/s$102.70/moView at Cherry Servers
  5. 5
    Cherry ServersCloud ARM VDS 16 · 16 vCPU · 72 GB

    16 vCPU dedicated · 72 GB RAM · 400 GB NVMe

    runs comfortably~0.8–1.7 tok/s$125.78/moView at Cherry Servers
  6. 6
    ContaboCloud VPS Plus 16 · 16 vCPU · 64 GB

    16 vCPU shared · 64 GB RAM · 750 GB NVMe

    runs comfortably~0.6–1.1 tok/s$91.16/mochecked 10 days agoView at Contabo

Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbors are, so treat these as order-of-magnitude estimates.

RAM needed · CPU inference

46.5 GB

weights · Q4_K_M
42.5 GB
KV cache · 8K context
2.5 GB
OS + runtime headroom
1.5 GB

Other quants: Q8_0 weights 75 GB (near-lossless, about half of F16); F16 141.1 GB. Comfortable = 20% headroom over the total.

Model card

Size
71B parameters
Context
128K tokens
Kind
chat
Released
2024-12
License
Llama 3.3 Community License

Facts fetched from Hugging Face on 7 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 844,566 downloads.

ollama run llama3.3:70bmodel card ↗Meta

Or rent it as an API · 46 providers sell Llama-3.3-70B-Instruct

every price for Llama-3.3-70B-Instruct

Meta list price

$0.62/M tokens

Cheapest price we can explain

$0.38/M tokens

NanoGPT, read 16 Sept 2026

What the VPS price buys

122M tokens/mo

at $46.49/mo

Token prices are per million, three parts input to one part output, and exclude tax. A rented box costs the same whether or not you use it and runs whatever else you put on it; an API charges for what you send and nothing when you stop.

Or buy hardware · 5 reference machines fit

all machines →
  • Framework Desktop (Ryzen AI Max+ 395, 64 GB)64 GB · 256 GB/s · ~3.1–4.8 tok/s · tight$1,659*= 36 mo of VPS
  • Mac mini (M5 Pro, 64 GB)64 GB · 307 GB/s · ~3.8–5.8 tok/s · tight$2,299= 50 mo of VPS
  • Framework Desktop (Ryzen AI Max+ 395, 128 GB)128 GB · 256 GB/s · ~3.1–4.8 tok/s$3,149*= 68 mo of VPS
  • NVIDIA DGX Spark (128 GB)128 GB · 273 GB/s · ~3.4–5.1 tok/s$4,699*= 101 mo of VPS
  • Mac Studio (M5 Ultra, 96 GB)96 GB · 1200 GB/s · ~15–23 tok/s$5,499= 118 mo of VPS

Buy · per month

$50.75

$46.08 hardware + $4.67 power

Rent · per month

$46.49

netcup VPS 8000 G12 · 16 vCore · 64 GB

Break-even

The Framework Desktop (Ryzen AI Max+ 395, 64 GB) pays for itself after 40 months of replacing the VPS, and it streams an estimated 3.1–4.8 tok/s against the VPS's CPU-only pace.

* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.

Frequently asked

How much RAM does Llama 3.3 70B need?
About 46.5 GB for CPU inference at Q4_K_M: 42.5 GB of weights, 2.5 GB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (320 KB per token for this model).
What is the cheapest VPS that can run Llama 3.3 70B?
netcup VPS 8000 G12 · 16 vCore · 64 GB (64 GB RAM, 16 vCPU) at $46.49/mo excl. VAT runs it comfortably as of 16 Sept 2026.
How fast will Llama 3.3 70B run on a VPS without a GPU?
Roughly 0.8–1.7 tok/s on the top pick. Token generation is limited by memory bandwidth, because each new token reads all of the weights once. More vCPUs help, dedicated ones most, but a GPU or an Apple silicon machine is 10 to 50 times faster.
Same lineupLlama 3.2 1BLlama 3.2 3BLlama 3.1 8BLlama 4 Scout (109B MoE)Same size classDeepSeek-R1 Distill Llama 70BVPS for AI agents (API-hosted models) →