FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Guides

Sizing GPUs for a 70B model: from parameters to KV cache and quantisation

When a client asks how many GPUs they need to run a model in their own datacentre, you should have the answer before you leave the meeting, not after a failed deployment.

In brief

  • Weights are only the start: 70B parameters at FP16 already take 140 GB, before KV cache and overhead.
  • KV cache grows with context length and the number of simultaneous users, so those are the numbers you must get from the client.
  • INT4 can take a 70B model from two 80 GB GPUs down to one, but measure it again with get_memory_footprint.
ShareLinkedInFacebookX
GraphicFive steps to estimate VRAM for serving an LLM
  1. 1Weight memoryParameters × bytes per parameter: 70B at FP16 is 140 GB, at INT4 35 GB
  2. 2KV cache per token2 × layers × KV heads × head_dim × bytes per element
  3. 3Multiply by real load× context length (including the response) × concurrent requests
  4. 4Add 10–20% headroomFor activations, CUDA context and the framework
  5. 5Compare with client VRAMIf short, quantise, cut context or add GPUs, then measure again

Weights are only the first step; KV cache, driven by context length and user count, usually decides the number of GPUs.

Graphic: FDE Times

Second meeting with a banking client. Their infrastructure team will not let data leave the building, so the model must run on-prem, and the head of IT asks directly: “To run a 70B model, how many cards do we need to buy?” If you answer “let me check and get back to you”, you lose a week.

If you answer wrongly, the client spends money on hardware and the model still will not run.

When a model has to run on a client’s infrastructure, you should be able to turn a model name into gigabytes, and gigabytes into a GPU count, right there at the table. The arithmetic is not hard. The hard part is knowing which components to count, and which ones the client will not mention unprompted.

This guide moves from the concept of parameters, through a complete hand calculation for a 70B model, to code you can use to check it. It ends with the mistakes that throw estimates off by a factor of two.

What are parameters, and why do they determine the GPU count?

IBM defines an LLM’s parameters as the settings that control the model’s output and behaviour, with two main types: weights and biases. Weights are numerical values that express how much importance the model assigns to a particular input. Taken together, a model can have billions of parameters.

For anyone doing deployment, each parameter is a number that has to sit in GPU memory while the model runs. So the first formula is simple: weight memory = number of parameters × bytes per parameter. The “70B” in the model name is the parameter count; what remains is knowing how many bytes each parameter occupies.

That is where quantisation comes in. IBM describes it as a way of simplifying all the mathematics inside the model, making it smaller and more efficient. Instead of storing each weight in 2 bytes (FP16/BF16), you store it in 1 byte (INT8) or half a byte (INT4).

Precision Bytes per parameter Weights for a 70B model
FP16 / BF16 2 140 GB
INT8 1 70 GB
INT4 0.5 35 GB

Hugging Face’s Transformers documentation states that 8-bit quantisation with bitsandbytes halves memory compared with 16-bit, and 4-bit compresses it further. On weights alone, 140 GB at FP16 already exceeds a single 80 GB GPU.

That is why an analysis on Machine Learning at Scale notes that INT4 takes a 70B model from needing two 80 GB GPUs to fitting on one.

Weights are only half the story

Stop at the table above and you will tell the client “INT4, one 80 GB card is enough”, and you may be wrong. When serving requests, the model also holds a KV cache: the key and value vectors for every token already processed, so it does not have to recompute them at each new generation step.

The per-token formula is: kv_bytes_per_token = 2 × n_layers × n_kv_heads × head_dim × bytes_per_element. The factor of 2 accounts for both keys and values. The other values are in the model’s config file; read them directly from there.

KV cache grows with the context window, meaning all the text the model can refer to while generating a response. Claude’s documentation stresses that the context window includes the response. So a request with a 6,000-token prompt and a 2,000-token answer occupies KV cache for 8,000 tokens, not 6,000.

A worked example

Suppose the bank’s 70B model has the following hypothetical configuration: 80 layers, 8 KV heads, head_dim of 128, and KV cache stored at 2 bytes per element. These numbers are for illustration only; in practice, take them from the model’s config.json.

KV per token = 2 × 80 × 8 × 128 × 2 = 327,680 bytes, roughly 0.33 MB. The client says typical context is 8,192 tokens and they want to serve 10 users at once. One request costs 327,680 × 8,192 ≈ 2.68 GB; ten requests cost about 26.8 GB.

Now add it up. At INT4: 35 GB of weights + 26.8 GB of KV cache = 61.8 GB. Machine Learning at Scale recommends reserving an extra 10–20% on top of weights and KV cache for activations, CUDA context and the framework; taking 20% to be safe gives about 74.2 GB. It fits on one 80 GB GPU, but only just.

At FP16 the picture changes completely: 140 + 26.8 = 166.8 GB, plus 20% comes to about 200 GB. Two 80 GB GPUs can hold the weights but not this load; you need three cards, or fewer concurrent users. Same model, different precision, a difference of two cards.

Then the client adds: “We want to feed in whole long contracts, around 32 thousand tokens.” Say the context is now 32,768 tokens.

One request takes 327,680 × 32,768 ≈ 10.7 GB of KV cache; ten requests take about 107 GB. Add 35 GB of INT4 weights and 20% headroom, and the total rises to about 171 GB: three 80 GB GPUs instead of one. The single-card plan collapsed over one question about document length.

Long context costs more elsewhere too

Memory is not the only cost. IBM notes that the compute required for attention grows with the square of sequence length, so long requests consume disproportionate resources. Claude’s documentation also warns that longer context is not automatically better, because of context rot.

This is the moment to revisit the requirement with the client. Rather than buying more cards to cram an entire 32,768-token contract into the prompt, you can propose splitting documents into chunks and including only the relevant sections. Careful sizing often produces a leaner design, not just a shopping list of cards.

Don’t trust the numbers on paper: measure

Once you have done the arithmetic, verify it on real hardware. The Transformers documentation lets you load a model in 8-bit with device_map=“auto” to use whatever GPUs are available, and provides get_memory_footprint to measure the loaded model. If you do not have an 80 GB card yet, test the formula on a small model first:

from transformers import AutoModelForCausalLM, BitsAndBytesConfig

def weight_gb(params, bytes_per_param):
    return params * bytes_per_param / 1e9

def kv_gb(n_layers, n_kv_heads, head_dim, bytes_el, context, concurrency):
    per_token = 2 * n_layers * n_kv_heads * head_dim * bytes_el
    return per_token * context * concurrency / 1e9

# Estimate for the 70B INT4 case in the example above, plus 20% headroom
total = (weight_gb(70e9, 0.5) + kv_gb(80, 8, 128, 2, 8192, 10)) * 1.2
print(f"Estimate for 70B INT4 case: {total:.1f} GB")

# Verify the weight formula on a small model, loaded in 8-bit (1 byte/parameter)
model_id = "your-small-model-name"  # replace with any small model on Hugging Face
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    quantization_config=BitsAndBytesConfig(load_in_8bit=True),
    device_map="auto",
)
print(f"Calculated: {weight_gb(model.num_parameters(), 1):.2f} GB")
print(f"Measured:   {model.get_memory_footprint() / 1e9:.2f} GB")

In the code above, the comments estimate the 70B INT4 case with 20% headroom, then check the weight formula on a small 8-bit model (1 byte per parameter); replace model_id with any small model on Hugging Face. The last two lines print the hand-calculated figure and the measured one.

get_memory_footprint reports only the loaded model, not the KV cache under load. So the final step is always to run a test at the concurrency and context length the client actually needs, and watch memory.

Five mistakes that skew the estimate

The most common mistake is counting only the weights, as in the example above: 35 GB looks comfortable until ten users submit long documents at once. The second is forgetting that the response also sits in the context, and so sizing KV cache from prompt length alone.

The third is confusing parameters with bytes, assuming “70B means 70 GB” without asking about precision. The fourth is skipping the 10–20% headroom, and then watching the model crash out of memory in the middle of the demo.

The fifth is subtler: treating quantisation as free. Quantisation simplifies the mathematics inside the model, so you need to run the client’s own eval suite on the INT4 version before promising FP16-level quality. Show the client the results and let them decide on the trade-off.

Bringing this skill to your CV and interviews

When reading FDE job descriptions, look for phrases such as “on-prem deployment”, “air-gapped”, “GPU sizing” or “inference optimization”: that is where this skill gets paid. On a CV, one concrete line carries more weight than any adjective, for instance describing how you estimated VRAM for a model, chose INT4 after measuring quality, and reduced the number of GPUs to buy.

In a case interview, if asked “the client wants to run model X; how many GPUs do they need?”, do not rush to a number.

Ask five questions back: what GPUs the client has, how much VRAM each card has, how many concurrent users, how long the context is, and what level of quality is acceptable.

Those five questions show you understand what the final number depends on.

Next time a client asks “how many cards?”, the right answer starts with “let me ask you five questions” and ends with a calculation both sides can check.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
5 sources
Read next on the roadmap · Stage 5: DeploymentWhen the client's source team quietly changes the schema: writing data contracts that keep pipelines intactAt some point the client's source team will rename a field without warning. Whether your pipeline notices straight away depends on the data contract you wrote in week one.