Run Llama 3.2 1B on a VPS.
prices as of · Contabo as of · 303 of 524 plans fit, disk included · re-ranked daily
Tiny instruct model for classification, routing and on-device assistants; runs on the smallest box in the index. On a CPU-only VPS it needs about 2.6 GB of RAM: 819 MB of Q4_K_M weights, 307 MB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 303 of the 524 plans in our index fit and include a disk; the cheapest comfortable pick is OVHcloud's VPS-1 2027 · 2 vCPU / 4 GB · EU at $5.18/mo, streaming an estimated 6.4–13 tok/s.
The cheapest VPS that runs Llama 3.2 1B is OVHcloud VPS-1 2027 at $5.18/mo, 4 GB of RAM against the 2.6 GB the model needs, checked 16 Sept 2026.
Best VPS plans for Llama 3.2 1B
ranked by estimated tokens/s per dollar, comfortable fits first
- 1runs comfortably~13–26 tok/s$5.93/moView at HOSTKEY ↗
- 2runs comfortably~19–38 tok/s$8.65/mochecked 10 days agoView at Contabo ↗
- 3runs comfortably~13–26 tok/s$6.35/mochecked 10 days agoView at Contabo ↗
- 4runs comfortably~38–77 tok/s$20.78/moView at netcup ↗
- 5runs comfortably~13–26 tok/s$7.58/moView at HOSTKEY ↗
- 6runs comfortably~19–38 tok/s$11.70/moView at HOSTKEY ↗
Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbors are, so treat these as order-of-magnitude estimates.
RAM needed · CPU inference
2.6 GB
- weights · Q4_K_M
- 819 MB
- KV cache · 8K context
- 307 MB
- OS + runtime headroom
- 1.5 GB
Other quants: Q8_0 weights 1.3 GB (near-lossless, about half of F16); F16 2.5 GB. Comfortable = 20% headroom over the total.
Model card
- Size
- 1.24B parameters
- Context
- 128K tokens
- Kind
- chat
- Released
- 2024-09
- License
- Llama 3.2 Community License
Facts fetched from Hugging Face on 7 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 6,010,185 downloads.
Or rent it as an API · 11 providers sell Llama 3.2 1B Instruct
every price for Llama 3.2 1B Instruct →Meta list price
$0.28/M tokens
Cheapest price we can explain
$0.04/M tokens
Inference, read 16 Sept 2026
What the VPS price buys
130M tokens/mo
at $5.18/mo
Token prices are per million, three parts input to one part output, and exclude tax. A rented box costs the same whether or not you use it and runs whatever else you put on it; an API charges for what you send and nothing when you stop.
Or buy hardware · 14 reference machines fit
all machines →- Raspberry Pi 5 (16 GB)16 GB · 17 GB/s · ~5.9–9.1 tok/s$305*= 59 mo of VPS
- GeForce RTX 5060 Ti 16 GB (card only)16 GB · 448 GB/s · ~100+ tok/s$800*= 10+ yrs of VPS
- Mac mini (M6, 16 GB)16 GB · 153 GB/s · ~100+ tok/s$899= 10+ yrs of VPS
- GeForce RTX 5070 Ti 16 GB (card only)16 GB · 896 GB/s · ~100+ tok/s$1,200*= 10+ yrs of VPS
- Mac mini (M6, 32 GB)32 GB · 153 GB/s · ~100+ tok/s$1,299= 10+ yrs of VPS
- GeForce RTX 5080 16 GB (card only)16 GB · 960 GB/s · ~100+ tok/s$1,600*= 10+ yrs of VPS
Buy · per month
$8.78
$8.47 hardware + $0.31 power
Rent · per month
$5.18
OVHcloud VPS-1 2027 · 2 vCPU / 4 GB · EU
Break-even
The Raspberry Pi 5 (16 GB) pays for itself after 63 months of replacing the VPS, and it streams an estimated 5.9–9.1 tok/s against the VPS's CPU-only pace.
* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.
Frequently asked
- How much RAM does Llama 3.2 1B need?
- About 2.6 GB for CPU inference at Q4_K_M: 819 MB of weights, 307 MB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (32 KB per token for this model).
- What is the cheapest VPS that can run Llama 3.2 1B?
- OVHcloud VPS-1 2027 · 2 vCPU / 4 GB · EU (4 GB RAM, 2 vCPU) at $5.18/mo excl. VAT runs it comfortably as of 16 Sept 2026.
- How fast will Llama 3.2 1B run on a VPS without a GPU?
- Roughly 13–26 tok/s on the top pick. Token generation is limited by memory bandwidth, because each new token reads all of the weights once. More vCPUs help, dedicated ones most, but a GPU or an Apple silicon machine is 10 to 50 times faster.