Run BGE-M3 on a VPS.
prices as of · 303 of 524 plans fit, disk included · re-ranked daily
Multilingual dense + sparse embeddings for hybrid search. On a CPU-only VPS it needs about 2.6 GB of RAM: 1.1 GB of weights in the model's native format and 1.5 GB for the OS and runtime. 303 of the 524 plans in our index fit and include a disk; the cheapest comfortable pick is OVHcloud's VPS-1 2027 · 2 vCPU / 4 GB · EU at $5.18/mo.
The cheapest VPS that runs BGE-M3 is OVHcloud VPS-1 2027 at $5.18/mo, 4 GB of RAM against the 2.6 GB the model needs, checked 16 Sept 2026.
Best VPS plans for BGE-M3
ranked by price among plans that fit
- 1runs comfortably$5.18/moView at OVHcloud ↗
- 2runs comfortably$5.35/moView at OVHcloud ↗
- 3runs comfortably$5.73/moView at netcup ↗
- 4runs comfortably$5.93/moView at HOSTKEY ↗
- 5runs comfortably$6.27/moView at HOSTKEY ↗
- 6runs comfortably$6.33/monot available nowView at Hetzner ↗
RAM needed · CPU inference
2.6 GB
- weights · native format
- 1.1 GB
- OS + runtime headroom
- 1.5 GB
Model card
- Size
- 0.568B parameters
- Context
- 8K tokens
- Kind
- embedding
- Released
- 2024-01
- License
- MIT
Facts fetched from Hugging Face on 7 Sept 2026: sizes from bits-per-weight · 37,730,552 downloads.
Frequently asked
- How much RAM does BGE-M3 need?
- About 2.6 GB for CPU inference at Q4_K_M: 1.1 GB of weights, 0 MB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (0 KB per token for this model).
- What is the cheapest VPS that can run BGE-M3?
- OVHcloud VPS-1 2027 · 2 vCPU / 4 GB · EU (4 GB RAM, 2 vCPU) at $5.18/mo excl. VAT runs it comfortably as of 16 Sept 2026.
- How fast will BGE-M3 run on a VPS without a GPU?
- Embedding models are small and run well on CPUs. Throughput scales with vCPUs rather than RAM.