◢◤ GENAI_CALC_
SYS_TIME UTC+0

Modelling

Transformer inference sizing. Will it fit, and how fast will it decode. First-order roofline estimates.

How to use this tab

This tab answers two questions about running an AI model on GPUs: will it fit? and how fast will it be? You choose a model and some hardware, and the app works out the answer instantly. It's a smart estimate, not a real benchmark, so treat the numbers as a close guide.

Try it in three steps

  1. On the left, pick a model (the AI) and a GPU (the chip that runs it). Every control has an button that explains it.
  2. Set how many GPUs you have and how you split the model across them, then how much work you send it (batch size and how long the questions and answers are).
  3. Watch the pictures and numbers on the right change as you drag the sliders. Nothing is saved or sent anywhere — experiment freely.

What the pictures mean

  • Memory box — everything that must fit in the GPU's memory: the model's weights, the KV cache (its memory of the conversation), and working space. If the box overflows, it won't fit.
  • Pipe — how hard the GPU is reading its memory. This is usually the speed limit when writing an answer.
  • Chip (die) — how busy the GPU's calculators are. This is the speed limit when reading a long question.
  • Interconnect — the cables and network the GPUs use to talk to each other when a model is shared across many of them.

A model always has one thing slowing it down most — the bottleneck. It's either waiting on memory, waiting on math, or waiting to talk to other GPUs. The tool tells you which, so you know what to fix.

Supported model types

  • Language models — the familiar chatbots and text models (Llama, Qwen, DeepSeek, GLM, Nemotron, and more). Dense and mixture-of-experts (MoE), including latent-attention (MLA) and hybrid Mamba designs. They read a question and write an answer one token at a time.
  • Diffusion models — image generators (Stable Diffusion XL, FLUX.1, SD 3.5, PixArt-Σ, DiT-XL) and video generators (CogVideoX). They start from noise and clean it up over many steps. Sized by images (or clips) per second and time per image, not tokens. Video adds a frames control.
  • Autoregressive image (VAR) — image generators built like a language model, predicting an image coarse-to-fine one scale at a time (VAR-d30). Modelled as a transformer: set Output tokens to the image's token count (about 680 for 256²).
  • Vision-language (VLM) — LLMs that also read images (Qwen2.5-VL, Pixtral, LLaVA-OneVision). Each image adds a few hundred to a few thousand visual patch tokens to the prompt. Sized like an LLM; images inflate prefill and KV memory.
  • Vision-language-action (VLA) — robotics policies (pi-zero, SmolVLA). A VLM backbone with a flow-matching action expert that produces a chunk of continuous actions per observation. Sized by robots driven at the control frequency (typically 50 Hz).
  • Speech (ASR) — Whisper family. Encoder-decoder: 30-second audio window through the encoder, autoregressive text decode with cross-attention. Reports real-time factor (RTF), audio-sec/sec, and encoder-vs-decoder time. Sized by concurrent real-time streams and minimum RTF.
  • Embeddings & rerankers — text encoders that turn a document into a vector (BGE-M3, Jina, E5-Mistral) or score a query-doc pair (BGE reranker). A single bidirectional forward pass, no generation. Sized by docs per second and time per batch. Encoders love batching.
  • World models (JEPA) — self-supervised video models (V-JEPA and V-JEPA 2). They watch a clip and turn it into an embedding in a single pass, no words in or out. Sized by clips per second. Because a clip is thousands of space-time patches, attention cost grows with the square of the patch count. The -AC variant is an action-conditioned world model: it also rolls a predictor forward over a planning horizon.

Pick the type from the grouped model menu on the left. The controls and result cards change to match the type you chose.

Word list

  • Token — a chunk of text, roughly a short word or word-piece. Models read and write in tokens.
  • Weights — the model's learned knowledge, stored as billions of numbers. Bigger models have more and need more memory.
  • GPU — the specialized chip that does the heavy math. HBM is its fast on-chip memory.
  • KV cache — short notes the model keeps about the conversation so it doesn't re-read everything for each new word. It grows with the length of the chat.
  • Batch — how many requests are handled together at once.
  • Prefill vs decode — prefill is reading the whole question (math-heavy); decode is writing the answer one word at a time (memory-heavy).
  • Parallelism (TP / PP / EP) — different ways to split one model across several GPUs when it's too big or too slow for one.
  • Weight format / quantization — how precisely each number is stored. Fewer bits saves memory and can run faster, with a small accuracy cost.

Want the opposite — you know your users and speed goals and want to know how many GPUs to buy? Use the Workload tab.

Fits on device
Yes
50.9 GB free / GPU
Max concurrent seqs
191
8k ctx · paged KV
Decode throughput
2.7k tok/s
aggregate, 8 GPU
Per sequence
84.6 tok/s
single user decode
Time to first token
3.16 s
prefill · memory-bound
Full request
51.59 s
TTFT + 4096 tok decode
HBM Memory 29.1 GB / 80.0 GB
Weights 16.4 GB
KV cache 10.2 GB
HBM bandwidth 2419 / 3350 GB/s
72% of 3.35 TB/s
GPU compute die block size ∝ die area
Tensor cores 5%
CUDA / vector cores 26%
L2 cache 72%
HBM I/O (mem controllers)◄ bottleneck 72%
NVLink I/O 1%
Scheduler · uncore structural
48 / 990 TFLOPS
HBM I/O tracks memory bandwidth, L2 tracks the hotter of memory/compute (so it stays busy in prefill), link tracks interconnect; scheduler shown structural
Interconnect
Cluster topology 1 node × 8 GPU · TP 8 · DP 1 · Single node
Node 0
g0
g1
g2
g3
g4
g5
g6
g7
TP all-reduce over NVLink · 12.4 GB/s
Batch sweep — throughput vs per-user speed
sweet spot5.2k01300182464128
Aggregate throughput Per-user speed
x: batch size (requests in flight)
At batch 32
2.7k tok/s
Per user
84.6 tok/s
TTFT
3159 ms
Cost / 1M tok
$2.85

Sweet spot ≈ batch 96: 4.7k tok/s aggregate at 49 tok/s per user. Beyond it, throughput is within 10% of peak while per-user speed keeps falling. Estimate, single replica of this config.

Cost estimate
Estimates only. Rates vary 2-3x between providers, and committed contracts or private pricing can change the bill a lot. Get a real quote before you decide.
Purchasing
GPUs billed
8 ×
$3.47/GPU-hr
Cluster / hour
$27.76
On-demand
Cluster / day
$666
24 × hourly
Cost / 1M tokens
$2.85
output tokens

Default rates are the median on-demand price across GPU cloud providers (GetDeploying GPU price index, 2026-10-07). Providers vary 2-3x around it, so enter your own quote when you have one. Committed and spot apply a typical discount.

Power & energy
Estimates only. First-order approximations, not audited figures. Grid intensities are directional, and TDP-based power under-counts real host draw. For anything you report externally, use measured figures from your provider or your own meters.
Emissions method
Power draw
7.04 kW
564 W / GPU · PUE 1.20
Energy / day
169 kWh
61.6 MWh / year
CO₂e / year
29.3 t
475 g/kWh · location-based
CO₂e / 1M tokens
343.1 g
output tokens

First-order: GPU draw = idle 30% of TDP + linear to TDP with utilization. Host overhead 30% of GPU. PUE 1.20. Grid intensities are directional estimates anchored to public 2024 grid-mix data.

Cold-start / model load
Cold start
1.4 s
source-bound
Per-GPU bytes
17.6 GB
1/8 of the model
PCIe (node)
504 GB/s
63 per GPU × 8
Source
100 GB/s
Parallel file system (Lustre-class)

Wall-clock ≈ weights ÷ min(source, node PCIe). Sharded assumes a well-parallelised loader (e.g. FSDP with parallel range reads). This is where Gen5 PCIe hosts (H100 and newer) pay off vs Gen4 (A100, L4, L40S): ~2× faster once the source can keep up.

First-order roofline (MBU 0.8, MFU 0.7) with 1.5 GB runtime overhead per GPU. MoE/EP sharding, activation memory and KV allocation waste are approximate; projected models are estimates from the current generation.

runs in your browser prices as of 2026-10-07 first-order estimates // not a benchmark