Throughput Calculator

· LIVE

A weight-bandwidth ceiling, not measured throughput. KV memory uses a fixed example: 80 layers, 64 KV heads, head dimension 128, and BF16 cache. Model size changes weights only.

Model size70B params
Batch size8
Sequence length4096 tokens
Precision
GPU
Estimate
Weight-read ceiling (tokens/sec)
383
Example memory
155.9 GB
Example fits on 1 GPU?
needs sharding
Param memory (fp8)70.0 GB
KV cache @ batch=8, seq=409685.9 GB
HBM bandwidth (H100 SXM)3350 GB/s
⚠ The ceiling ignores KV-cache traffic and compute limits. Precision describes weight storage, not native GPU support. Actual numbers depend on kernel quality, continuous batching, speculative decoding, and a dozen other things this tool doesn't model.

What it does

Shows an ideal weight-read ceiling: GPU memory bandwidth divided by weight bytes, multiplied by batch size. It is an upper bound for a simplified decode step, not a throughput prediction.

The memory example assumes 80 layers, 64 KV heads, a head dimension of 128, and a BF16 cache. These dimensions stay fixed when you change the parameter count. Use the Training Memory Calculator for architecture-specific memory estimates.

Limitations

  • The ceiling omits KV-cache traffic, compute limits, communication, and kernel overhead. Longer sequences increase the displayed memory but do not change this weight-only ceiling.
  • Precision sets weight storage size. It does not establish that a GPU supports native arithmetic at that precision or that suitable kernels exist.
  • Memory fit reserves 15% headroom. A configuration that needs sharding cannot achieve the displayed ceiling on a single GPU; multi-GPU performance is not modeled.
  • Use measured benchmarks for deployment decisions.

See How To Scale Your Model: Transformer Inference for the full bandwidth and compute model.