Modelling
Transformer inference sizing. Will it fit, and how fast will it decode. First-order roofline estimates.
How to use this tab
This tab answers two questions about running an AI model on GPUs: will it fit? and how fast will it be? You choose a model and some hardware, and the app works out the answer instantly. It's a smart estimate, not a real benchmark, so treat the numbers as a close guide.
Try it in three steps
- On the left, pick a model (the AI) and a GPU (the chip that runs it). Every control has an button that explains it.
- Set how many GPUs you have and how you split the model across them, then how much work you send it (batch size and how long the questions and answers are).
- Watch the pictures and numbers on the right change as you drag the sliders. Nothing is saved or sent anywhere — experiment freely.
What the pictures mean
- Memory box — everything that must fit in the GPU's memory: the model's weights, the KV cache (its memory of the conversation), and working space. If the box overflows, it won't fit.
- Pipe — how hard the GPU is reading its memory. This is usually the speed limit when writing an answer.
- Chip (die) — how busy the GPU's calculators are. This is the speed limit when reading a long question.
- Interconnect — the cables and network the GPUs use to talk to each other when a model is shared across many of them.
A model always has one thing slowing it down most — the bottleneck. It's either waiting on memory, waiting on math, or waiting to talk to other GPUs. The tool tells you which, so you know what to fix.
Supported model types
- Language models — the familiar chatbots and text models (Llama, Qwen, DeepSeek, GLM, Nemotron, and more). Dense and mixture-of-experts (MoE), including latent-attention (MLA) and hybrid Mamba designs. They read a question and write an answer one token at a time.
- Diffusion models — image generators (Stable Diffusion XL, FLUX.1, SD 3.5, PixArt-Σ, DiT-XL) and video generators (CogVideoX). They start from noise and clean it up over many steps. Sized by images (or clips) per second and time per image, not tokens. Video adds a frames control.
- Autoregressive image (VAR) — image generators built like a language model, predicting an image coarse-to-fine one scale at a time (VAR-d30). Modelled as a transformer: set Output tokens to the image's token count (about 680 for 256²).
- Vision-language (VLM) — LLMs that also read images (Qwen2.5-VL, Pixtral, LLaVA-OneVision). Each image adds a few hundred to a few thousand visual patch tokens to the prompt. Sized like an LLM; images inflate prefill and KV memory.
- Vision-language-action (VLA) — robotics policies (pi-zero, SmolVLA). A VLM backbone with a flow-matching action expert that produces a chunk of continuous actions per observation. Sized by robots driven at the control frequency (typically 50 Hz).
- Speech (ASR) — Whisper family. Encoder-decoder: 30-second audio window through the encoder, autoregressive text decode with cross-attention. Reports real-time factor (RTF), audio-sec/sec, and encoder-vs-decoder time. Sized by concurrent real-time streams and minimum RTF.
- Embeddings & rerankers — text encoders that turn a document into a vector (BGE-M3, Jina, E5-Mistral) or score a query-doc pair (BGE reranker). A single bidirectional forward pass, no generation. Sized by docs per second and time per batch. Encoders love batching.
- World models (JEPA) — self-supervised video models (V-JEPA and V-JEPA 2). They watch a clip and turn it into an embedding in a single pass, no words in or out. Sized by clips per second. Because a clip is thousands of space-time patches, attention cost grows with the square of the patch count. The -AC variant is an action-conditioned world model: it also rolls a predictor forward over a planning horizon.
Pick the type from the grouped model menu on the left. The controls and result cards change to match the type you chose.
Word list
- Token — a chunk of text, roughly a short word or word-piece. Models read and write in tokens.
- Weights — the model's learned knowledge, stored as billions of numbers. Bigger models have more and need more memory.
- GPU — the specialized chip that does the heavy math. HBM is its fast on-chip memory.
- KV cache — short notes the model keeps about the conversation so it doesn't re-read everything for each new word. It grows with the length of the chat.
- Batch — how many requests are handled together at once.
- Prefill vs decode — prefill is reading the whole question (math-heavy); decode is writing the answer one word at a time (memory-heavy).
- Parallelism (TP / PP / EP) — different ways to split one model across several GPUs when it's too big or too slow for one.
- Weight format / quantization — how precisely each number is stored. Fewer bits saves memory and can run faster, with a small accuracy cost.
Want the opposite — you know your users and speed goals and want to know how many GPUs to buy? Use the Workload tab.
Sweet spot ≈ batch 96: 4.7k tok/s aggregate at 49 tok/s per user. Beyond it, throughput is within 10% of peak while per-user speed keeps falling. Estimate, single replica of this config.
Default rates are the median on-demand price across GPU cloud providers (GetDeploying GPU price index, 2026-10-07). Providers vary 2-3x around it, so enter your own quote when you have one. Committed and spot apply a typical discount.
First-order: GPU draw = idle 30% of TDP + linear to TDP with utilization. Host overhead 30% of GPU. PUE 1.20. Grid intensities are directional estimates anchored to public 2024 grid-mix data.
Wall-clock ≈ weights ÷ min(source, node PCIe). Sharded assumes a well-parallelised loader (e.g. FSDP with parallel range reads). This is where Gen5 PCIe hosts (H100 and newer) pay off vs Gen4 (A100, L4, L40S): ~2× faster once the source can keep up.
First-order roofline (MBU 0.8, MFU 0.7) with 1.5 GB runtime overhead per GPU. MoE/EP sharding, activation memory and KV allocation waste are approximate; projected models are estimates from the current generation.