Playground

Tools

Explore ML systems in real-time with live, interactive tools

Throughput Calculator

· LIVE

A weight-bandwidth ceiling, not measured throughput. KV memory uses a fixed example: 80 layers, 64 KV heads, head dimension 128, and BF16 cache. Model size changes weights only.

Model size70B params
Batch size8
Sequence length4096 tokens
Precision
GPU
Estimate
Weight-read ceiling (tokens/sec)
383
Example memory
155.9 GB
Example fits on 1 GPU?
needs sharding
Param memory (fp8)70.0 GB
KV cache @ batch=8, seq=409685.9 GB
HBM bandwidth (H100 SXM)3350 GB/s
⚠ The ceiling ignores KV-cache traffic and compute limits. Precision describes weight storage, not native GPU support. Actual numbers depend on kernel quality, continuous batching, speculative decoding, and a dozen other things this tool doesn't model.
Browse all

All tools & demos.

Around the web

Canonical references.