Playground
Tools
Explore ML systems in real-time with live, interactive tools
Throughput Calculator
· LIVEA weight-bandwidth ceiling, not measured throughput. KV memory uses a fixed example: 80 layers, 64 KV heads, head dimension 128, and BF16 cache. Model size changes weights only.
All tools & demos.
Throughput Calculator
Explore a weight-bandwidth ceiling and a fixed KV-cache memory example across precisions, batch sizes, and GPUs.
Training Compute Calculator
Turn a model size and token budget into training FLOPs, GPU-hours, wall-clock days, and cost, using the 6ND rule and real GPU peak numbers.
Training Memory Calculator
Per-GPU VRAM breakdown for training and inference — params, gradients, optimizer state, activations, KV cache, with ZeRO/FSDP sharding.
Attention Visualizer
Explore a synthetic attention map illustrating causal masking and recurring patterns.
Tokenizer explorer
Paste any text and see how popular tokenizers split it, side by side, with token counts and cost per provider.
Attention pattern atlas
A gallery of attention head signatures from popular open models — induction heads, name movers, sink tokens, retrieval circuits — annotated and searchable.
Eval Harness Playground
Run a small, fixed set of evals against an OpenAI-compatible endpoint and get a scorecard for quality, latency, and cost.
Kernel Benchmark
Side-by-side timings for Triton, CUDA, and PyTorch implementations of the same kernel — attention, layernorm, GEMM, custom — across shapes and dtypes.
Model Card Generator
Generate a structured model card from a checkpoint and an evaluation log — covers intended use, training data summary, evals, limitations, and ethics.
Canonical references.
- Tokenization
The Tokenizer Playground
Tokenize the same text with GPT-4, Claude, LLaMA, Mistral, Gemma, and more — switch tokenizers instantly, or load any Hugging Face tokenizer. Runs in the browser.
Xenova · Hugging Face - Tokenization
Tiktokenizer
OpenAI-focused tokenizer playground. Visualize cl100k, o200k, and legacy encodings with per-token highlights.
dqbd - Memory & VRAM
LLM Model VRAM Calculator
Widely-referenced inference VRAM estimator for popular open-source models with quantization and context-length sliders.
NyxKrage · Hugging Face - Memory & VRAM
APXML VRAM Calculator
Inference and fine-tuning VRAM calculator covering Nvidia GPUs and Apple Silicon. Good for picking hardware for a target model.
APXML - Architecture
LLM Visualization
A 3D, animated walk through the entire forward pass of a nano-GPT model, layer by layer. The clearest mental model of how a transformer works.
Brendan Bycroft - Training & Scaling
Chinchilla Scaling Calculator
Enter a model size, get the Chinchilla-optimal training-token count per Hoffmann et al. 2022 — with an interactive params-vs-tokens chart.
Nathan Godey