Throughput Calculator
· LIVEA weight-bandwidth ceiling, not measured throughput. KV memory uses a fixed example: 80 layers, 64 KV heads, head dimension 128, and BF16 cache. Model size changes weights only.
What it does
Shows an ideal weight-read ceiling: GPU memory bandwidth divided by weight bytes, multiplied by batch size. It is an upper bound for a simplified decode step, not a throughput prediction.
The memory example assumes 80 layers, 64 KV heads, a head dimension of 128, and a BF16 cache. These dimensions stay fixed when you change the parameter count. Use the Training Memory Calculator for architecture-specific memory estimates.
Limitations
- The ceiling omits KV-cache traffic, compute limits, communication, and kernel overhead. Longer sequences increase the displayed memory but do not change this weight-only ceiling.
- Precision sets weight storage size. It does not establish that a GPU supports native arithmetic at that precision or that suitable kernels exist.
- Memory fit reserves 15% headroom. A configuration that needs sharding cannot achieve the displayed ceiling on a single GPU; multi-GPU performance is not modeled.
- Use measured benchmarks for deployment decisions.
See How To Scale Your Model: Transformer Inference for the full bandwidth and compute model.