Model Memory & GPU Sizing Calculators

Whether a model fits, and on what: weight memory at each precision, KV-cache growth with context length and batch size, the memory a quantisation step actually saves, and what the GPU hours cost.

4 calculators in this category

Which one do you need?

LLM GPU VRAM Requirement Calculator
Size the GPU memory a language model needs from its parameter count, quantisation, context length and batch size, and see how many GPUs that takes.
KV Cache Size Calculator
Compute the key-value cache a transformer holds at your context length and batch size, the term that dominates long-context serving, and what fits.
Model Quantisation Memory Savings Calculator
Compare model footprint across precisions and get the compression ratio, memory freed and effective bits per weight before you run the quantisation.
GPU Hours Cost Calculator
Price a block of GPU work from the hourly rate, GPU count and run length, then compare on-demand against spot once interruption overhead is counted.

More ai, llm & machine learning engineering categories

Related categories in other subjects

Where this work overlaps other trades and disciplines.