Cheapest GPU to run a 70B model.
prices as of · re-ranked daily · 98 qualifying machines
A 70-billion-parameter model has to fit in GPU memory before it answers anything, which rules out most of this market and decides the price of the rest. We count only machines with 80 GB on one board, or GPUs adding up to 160 GB, and 98 of the 156 rentals we read qualify. The cheapest of them is $1.19 per hour served at A100 PCIe · Community Cloud, which rents for $1.19/GPU-hr. They are ranked by what the work costs rather than by the hourly rate, because a faster card at twice the price can still be the cheaper way to buy the job.
The cheapest machine that can do this job is A100 PCIe · Community Cloud at RunPod: $1.19 per hour served, read 17 Sept 2026.
Cheapest that fits
A100 PCIe · Community Cloud
$1.19per hour served
$1.19/GPU-hr
lowest cost per hour served
Cheapest interruptible
1A100.22V · 1x A100 SXM4 80GB
$1.71per hour served
$1.71/GPU-hr · $0.855 on spot
same machine, taken back at any moment
Most memory for the money
4A6000.40V · 4x RTX A6000 48GB
$2.37per hour served
$0.592/GPU-hr · $1.18 on spot
the largest instance in the first twenty-five
What actually matters here
- Memory before anything else
- A model that does not fit does not run slowly, it does not run. Memory is the first filter and speed is the second, which is why this page counts gigabytes before it counts anything else.
- 80 GB holds it in 8-bit
- One H100, H200, A100 80GB or MI300X takes the whole model at 8-bit weights, with room for a short context. It is the cheapest shape that serves at all.
- 160 GB holds it in 16-bit
- Two 80 GB boards, or one 192 GB board, run the model at full precision. Where the GPUs are in one instance they talk over the node's own links rather than over the network.
- What the context costs
- Every conversation in flight takes memory of its own on top of the weights, so a board that fits the model exactly serves one request at a time.
Cheapest 10 per hour served
- 1
RunPodA100 PCIe · Community Cloud
$1.19 per hour served · 1x A100 80GB PCIe · 80 GB · 8 vCPU · 117 GB RAM · $1.19/hr for the machineCommunity Cloud: vetted third-party hosts, cheaper and less consistent · storage billed separately · billed per second
✓ 80 GB per GPU
$1.19/GPU-hr
- 2
RunPodA100 SXM · Community Cloud
$1.39 per hour served · 1x A100 80GB SXM · 80 GB · 16 vCPU · 125 GB RAM · $1.39/hr for the machineCommunity Cloud: vetted third-party hosts, cheaper and less consistent · storage billed separately · billed per second
✓ 80 GB per GPU
$1.39/GPU-hr
- 3
RunPodA100 PCIe · Secure Cloud
$1.59 per hour served · 1x A100 80GB PCIe · 80 GB · 8 vCPU · 117 GB RAM · $1.59/hr for the machineSecure Cloud: RunPod's own datacentres · storage billed separately · billed per second
✓ 80 GB per GPU
$1.59/GPU-hr
- 4
RunPodA100 SXM · Secure Cloud
$1.59 per hour served · 1x A100 80GB SXM · 80 GB · 16 vCPU · 125 GB RAM · $1.59/hr for the machineSecure Cloud: RunPod's own datacentres · storage billed separately · billed per second
✓ 80 GB per GPU
$1.59/GPU-hr
- 5
RunPodRTX Pro 6000 · Community Cloud
$1.69 per hour served · 1x RTX PRO 6000 · 96 GB · 16 vCPU · 188 GB RAM · $1.69/hr for the machineCommunity Cloud: vetted third-party hosts, cheaper and less consistent · storage billed separately · billed per second
✓ 96 GB per GPU
$1.69/GPU-hr
- 6
Verda (DataCrunch)1A100.22V · 1x A100 SXM4 80GB
$1.71 per hour served · 1x A100 80GB SXM · 80 GB · 22 vCPU · 120 GB RAM · $1.71/hr for the machine · $0.855 at its spot rateDedicated Hardware Instance · storage billed separately
✓ 80 GB per GPU✓ spot rate published
$1.71/GPU-hr
- 7
Verda (DataCrunch)1RTXPRO6000.30V · 1x RTX PRO 6000 96GB
$1.84 per hour served · 1x RTX PRO 6000 · 96 GB · 30 vCPU · 90 GB RAM · $1.84/hr for the machine · $0.921 at its spot rateDedicated Hardware Instance · storage billed separately
✓ 96 GB per GPU✓ spot rate published
$1.84/GPU-hr
- 8
Verda (DataCrunch)1RTXPRO6000.30V.CC · 1x RTX PRO 6000 CC 96GB
$1.88 per hour served · 1x RTX PRO 6000 · 96 GB · 30 vCPU · 90 GB RAM · $1.88/hr for the machine · $0.939 at its spot rateDedicated Hardware Instance · storage billed separately
✓ 96 GB per GPU✓ spot rate published
$1.88/GPU-hr
- 9
RunPodH100 PCIe · Community Cloud
$1.99 per hour served · 1x H100 PCIe · 80 GB · 16 vCPU · 188 GB RAM · $1.99/hr for the machineCommunity Cloud: vetted third-party hosts, cheaper and less consistent · storage billed separately · billed per second
✓ 80 GB per GPU
$1.99/GPU-hr
- 10
Vultrvbm-72c-480gb-gh200-gpu
$1.99 per hour served · 1x GH200 · 96 GB · 72 vCPU · 480 GB RAM · $1.99/hr for the machineBare metal · $2009.28/mo if kept a whole month
✓ 96 GB per GPU
$1.99/GPU-hr
A 70-billion-parameter model is about 70 GB of 8-bit weights and about 140 GB at 16-bit, before anything is left over for the context window. We count a single board of 80 GB or more, which holds the 8-bit weights, and any instance whose GPUs add up to 160 GB or more, which holds the 16-bit ones. The price is what the whole instance bills for an hour, because that is what you rent. The price beside each row is the on-demand rate for one GPU for one hour, so the two numbers answer different questions: $1.19/GPU-hr is what the GPU costs, and $1.19 per hour served is what the work costs.
Frequently asked
- What is the cheapest GPU to run a 70b model?
- A100 PCIe · Community Cloud at RunPod, $1.19 per hour served. That is 1x A100 80GB PCIe with 80 GB of memory each, renting for $1.19/GPU-hr. Read on 17 Sept 2026.
- How is the cost worked out?
- A 70-billion-parameter model is about 70 GB of 8-bit weights and about 140 GB at 16-bit, before anything is left over for the context window. We count a single board of 80 GB or more, which holds the 8-bit weights, and any instance whose GPUs add up to 160 GB or more, which holds the 16-bit ones. The price is what the whole instance bills for an hour, because that is what you rent.
- Why do only 98 machines qualify?
- The job needs 80 GB on one board, or GPUs adding up to 160 GB. The other 58 rentals we read are still in the index; they just cannot do this job as described. Serverless rates are left out entirely, because they buy a container that runs when it is called rather than a machine you hold.
- Can I pay less on spot?
- 50 of the 98 qualifying machines publish an interruptible rate, and the column beside each row shows what the same work would cost on it if the run were never preempted. It can be, so treat the spot figure as a floor rather than a quote.
- How current are these prices?
- Prices are read daily from 10 providers' own public pages and keyless APIs, and this ranking recomputes with them. Last refresh: 17 Sept 2026.