LLM Inference Cost & Serving Calculators

What it costs to actually run a language model in production: per-token API pricing split across input and output, cost per user per month, prompt-cache savings, throughput in tokens per second, and the point at which self-hosting beats paying per token.

5 calculators in this category

Which one do you need?

LLM API Token Cost Calculator
Estimate LLM API spend from input and output token counts and per-million pricing, including prompt-cache discounts, per request, per month and per year.
LLM Cost Per User Per Month Calculator
Work out what one active user costs you in LLM tokens each month, what gross margin your price leaves, and the price a target margin requires.
Self-Hosted LLM vs API Break-Even Calculator
Find the monthly token volume where renting GPUs beats paying per token, including the utilisation assumption that usually decides the answer.
Prompt Caching Savings Calculator
Work out what caching a long system prompt actually saves once the write premium, the discounted read rate and the number of cache hits are counted.
LLM Throughput & Latency Calculator (Tokens Per Second)
Estimate tokens per second, time to first token and total response time from GPU memory bandwidth and model size, with batching and efficiency.

More ai, llm & machine learning engineering categories

Related categories in other subjects

Where this work overlaps other trades and disciplines.