Run Llama 4 Scout (109B MoE) on a VPS.
prices as of · Contabo as of · 3 of 524 plans fit, disk included · re-ranked daily
Mixture-of-experts: 109B of weights to hold in RAM but only 17B active per token, so it runs surprisingly fast once it fits. On a CPU-only VPS it needs about 68.4 GB of RAM: 65.4 GB of Q4_K_M weights, 1.5 GB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 3 of the 524 plans in our index fit and include a disk; the cheapest comfortable pick is Linode's Linode 90GB at $240/mo, streaming an estimated 0.7–1.4 tok/s.
The cheapest VPS that runs Llama 4 Scout (109B MoE) is Linode Linode 90GB at $240/mo, 90 GB of RAM against the 68.4 GB the model needs, checked 16 Sept 2026.
Best VPS plans for Llama 4 Scout (109B MoE)
ranked by estimated tokens/s per dollar, comfortable fits first
- 1runs comfortably~0.7–1.4 tok/s$240/moView at Linode ↗
- 2runs comfortably~1.2–2.3 tok/s$114.24/mochecked 10 days agoView at Contabo ↗
- 3tight fit~1.7–3.5 tok/s$125.78/moView at Cherry Servers ↗
Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated, ×0.5 for MoE routing overhead), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbors are, so treat these as order-of-magnitude estimates.
RAM needed · CPU inference
68.4 GB
- weights · Q4_K_M
- 65.4 GB
- KV cache · 8K context
- 1.5 GB
- OS + runtime headroom
- 1.5 GB
Other quants: Q8_0 weights 114.5 GB (near-lossless, about half of F16); F16 1.8 GB. Comfortable = 20% headroom over the total.
Model card
- Size
- 109B total · 17B active
- Context
- 10.24M tokens
- Kind
- chat · vision
- Released
- 2025-04
- License
- Llama 4 Community License
Facts fetched from Hugging Face on 7 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 162,521 downloads.
Or buy hardware · 3 reference machines fit
all machines →- Framework Desktop (Ryzen AI Max+ 395, 128 GB)128 GB · 256 GB/s · ~6.5–10 tok/s$3,149*= 13 mo of VPS
- NVIDIA DGX Spark (128 GB)128 GB · 273 GB/s · ~7–11 tok/s$4,699*= 20 mo of VPS
- Mac Studio (M5 Ultra, 96 GB)96 GB · 1200 GB/s · ~31–47 tok/s · tight$5,499= 23 mo of VPS
Buy · per month
$92.14
$87.47 hardware + $4.67 power
Rent · per month
$240
Linode Linode 90GB
Break-even
The Framework Desktop (Ryzen AI Max+ 395, 128 GB) pays for itself after 13 months of replacing the VPS, and it streams an estimated 6.5–10 tok/s against the VPS's CPU-only pace.
* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.
Frequently asked
- How much RAM does Llama 4 Scout (109B MoE) need?
- About 68.4 GB for CPU inference at Q4_K_M: 65.4 GB of weights, 1.5 GB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (192 KB per token for this model).
- What is the cheapest VPS that can run Llama 4 Scout (109B MoE)?
- Linode Linode 90GB (90 GB RAM, 4 vCPU) at $240/mo excl. VAT runs it comfortably as of 16 Sept 2026. Cherry Servers Cloud ARM VDS 16 · 16 vCPU · 72 GB at $125.78/mo is a tight fit.
- How fast will Llama 4 Scout (109B MoE) run on a VPS without a GPU?
- Roughly 0.7–1.4 tok/s on the top pick. Token generation is limited by memory bandwidth, because each new token reads the active experts' weights once. More vCPUs help, dedicated ones most, but a GPU or an Apple silicon machine is 10 to 50 times faster.