Cheapest GPU to run a 70B model.

prices as of · re-ranked daily · 98 qualifying machines

A 70-billion-parameter model has to fit in GPU memory before it answers anything, which rules out most of this market and decides the price of the rest. We count only machines with 80 GB on one board, or GPUs adding up to 160 GB, and 98 of the 156 rentals we read qualify. The cheapest of them is $1.19 per hour served at A100 PCIe · Community Cloud, which rents for $1.19/GPU-hr. They are ranked by what the work costs rather than by the hourly rate, because a faster card at twice the price can still be the cheaper way to buy the job.

The cheapest machine that can do this job is A100 PCIe · Community Cloud at RunPod: $1.19 per hour served, read 17 Sept 2026.

Cheapest that fits

A100 PCIe · Community Cloud

RunPod

$1.19per hour served

$1.19/GPU-hr

lowest cost per hour served

Cheapest interruptible

1A100.22V · 1x A100 SXM4 80GB

Verda (DataCrunch)

$1.71per hour served

$1.71/GPU-hr · $0.855 on spot

same machine, taken back at any moment

Most memory for the money

4A6000.40V · 4x RTX A6000 48GB

Verda (DataCrunch)

$2.37per hour served

$0.592/GPU-hr · $1.18 on spot

the largest instance in the first twenty-five

What actually matters here

Memory before anything else
A model that does not fit does not run slowly, it does not run. Memory is the first filter and speed is the second, which is why this page counts gigabytes before it counts anything else.
80 GB holds it in 8-bit
One H100, H200, A100 80GB or MI300X takes the whole model at 8-bit weights, with room for a short context. It is the cheapest shape that serves at all.
160 GB holds it in 16-bit
Two 80 GB boards, or one 192 GB board, run the model at full precision. Where the GPUs are in one instance they talk over the node's own links rather than over the network.
What the context costs
Every conversation in flight takes memory of its own on top of the weights, so a board that fits the model exactly serves one request at a time.

Cheapest 10 per hour served

A 70-billion-parameter model is about 70 GB of 8-bit weights and about 140 GB at 16-bit, before anything is left over for the context window. We count a single board of 80 GB or more, which holds the 8-bit weights, and any instance whose GPUs add up to 160 GB or more, which holds the 16-bit ones. The price is what the whole instance bills for an hour, because that is what you rent. The price beside each row is the on-demand rate for one GPU for one hour, so the two numbers answer different questions: $1.19/GPU-hr is what the GPU costs, and $1.19 per hour served is what the work costs.

Frequently asked

What is the cheapest GPU to run a 70b model?
A100 PCIe · Community Cloud at RunPod, $1.19 per hour served. That is 1x A100 80GB PCIe with 80 GB of memory each, renting for $1.19/GPU-hr. Read on 17 Sept 2026.
How is the cost worked out?
A 70-billion-parameter model is about 70 GB of 8-bit weights and about 140 GB at 16-bit, before anything is left over for the context window. We count a single board of 80 GB or more, which holds the 8-bit weights, and any instance whose GPUs add up to 160 GB or more, which holds the 16-bit ones. The price is what the whole instance bills for an hour, because that is what you rent.
Why do only 98 machines qualify?
The job needs 80 GB on one board, or GPUs adding up to 160 GB. The other 58 rentals we read are still in the index; they just cannot do this job as described. Serverless rates are left out entirely, because they buy a container that runs when it is called rather than a machine you hold.
Can I pay less on spot?
50 of the 98 qualifying machines publish an interruptible rate, and the column beside each row shows what the same work would cost on it if the run were never preempted. It can be, so treat the spot figure as a floor rather than a quote.
How current are these prices?
Prices are read daily from 10 providers' own public pages and keyless APIs, and this ranking recomputes with them. Last refresh: 17 Sept 2026.
Related jobsCheapest GPU node to train for a weekCheapest GPU to fine-tune a 7B modelCheapest H100 rentalCheapest GPU memory by the gigabyte