Run Llama 4 Scout (109B MoE) on a VPS.

prices as of · Contabo as of · 3 of 524 plans fit, disk included · re-ranked daily

Mixture-of-experts: 109B of weights to hold in RAM but only 17B active per token, so it runs surprisingly fast once it fits. On a CPU-only VPS it needs about 68.4 GB of RAM: 65.4 GB of Q4_K_M weights, 1.5 GB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 3 of the 524 plans in our index fit and include a disk; the cheapest comfortable pick is Linode's Linode 90GB at $240/mo, streaming an estimated 0.7–1.4 tok/s.

The cheapest VPS that runs Llama 4 Scout (109B MoE) is Linode Linode 90GB at $240/mo, 90 GB of RAM against the 68.4 GB the model needs, checked 16 Sept 2026.

Best VPS plans for Llama 4 Scout (109B MoE)

ranked by estimated tokens/s per dollar, comfortable fits first

  1. 1
    LinodeLinode 90GB

    4 vCPU dedicated · 90 GB RAM · 90 GB SSD

    runs comfortably~0.7–1.4 tok/s$240/moView at Linode
  2. 2
    ContaboCloud VPS Plus 18 · 18 vCPU · 96 GB

    18 vCPU shared · 96 GB RAM · 900 GB NVMe

    runs comfortably~1.2–2.3 tok/s$114.24/mochecked 10 days agoView at Contabo
  3. 3
    Cherry ServersCloud ARM VDS 16 · 16 vCPU · 72 GB

    16 vCPU dedicated · 72 GB RAM · 400 GB NVMe

    tight fit~1.7–3.5 tok/s$125.78/moView at Cherry Servers

Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated, ×0.5 for MoE routing overhead), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbors are, so treat these as order-of-magnitude estimates.

RAM needed · CPU inference

68.4 GB

weights · Q4_K_M
65.4 GB
KV cache · 8K context
1.5 GB
OS + runtime headroom
1.5 GB

Other quants: Q8_0 weights 114.5 GB (near-lossless, about half of F16); F16 1.8 GB. Comfortable = 20% headroom over the total.

Model card

Size
109B total · 17B active
Context
10.24M tokens
Kind
chat · vision
Released
2025-04
License
Llama 4 Community License

Facts fetched from Hugging Face on 7 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 162,521 downloads.

ollama run llama4:scoutmodel card ↗Meta

Or buy hardware · 3 reference machines fit

all machines →
  • Framework Desktop (Ryzen AI Max+ 395, 128 GB)128 GB · 256 GB/s · ~6.5–10 tok/s$3,149*= 13 mo of VPS
  • NVIDIA DGX Spark (128 GB)128 GB · 273 GB/s · ~7–11 tok/s$4,699*= 20 mo of VPS
  • Mac Studio (M5 Ultra, 96 GB)96 GB · 1200 GB/s · ~31–47 tok/s · tight$5,499= 23 mo of VPS

Buy · per month

$92.14

$87.47 hardware + $4.67 power

Rent · per month

$240

Linode Linode 90GB

Break-even

The Framework Desktop (Ryzen AI Max+ 395, 128 GB) pays for itself after 13 months of replacing the VPS, and it streams an estimated 6.5–10 tok/s against the VPS's CPU-only pace.

* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.

Frequently asked

How much RAM does Llama 4 Scout (109B MoE) need?
About 68.4 GB for CPU inference at Q4_K_M: 65.4 GB of weights, 1.5 GB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (192 KB per token for this model).
What is the cheapest VPS that can run Llama 4 Scout (109B MoE)?
Linode Linode 90GB (90 GB RAM, 4 vCPU) at $240/mo excl. VAT runs it comfortably as of 16 Sept 2026. Cherry Servers Cloud ARM VDS 16 · 16 vCPU · 72 GB at $125.78/mo is a tight fit.
How fast will Llama 4 Scout (109B MoE) run on a VPS without a GPU?
Roughly 0.7–1.4 tok/s on the top pick. Token generation is limited by memory bandwidth, because each new token reads the active experts' weights once. More vCPUs help, dedicated ones most, but a GPU or an Apple silicon machine is 10 to 50 times faster.
Same lineupLlama 3.2 1BLlama 3.2 3BLlama 3.1 8BLlama 3.3 70BSame size classQwen3 30B-A3B (MoE)gpt-oss-20b (MoE)gpt-oss-120b (MoE)VPS for AI agents (API-hosted models) →