◢◤ GENAI_CALC_
SYS_TIME UTC+0

Workload Sizing

Describe the demand and per-user SLOs. The smallest cluster that meets them is sized live.

How to use this tab

This tab answers "how much hardware do I need to buy?". It is the reverse of the Modelling tab. There you pick the hardware and see how it runs. Here you say how busy your service will be and the promises you want to keep, and the app finds the smallest cluster of GPUs that can do it. Every box has an i button — tap it to learn what that box means.

  1. Pick your model (the AI brain) and a GPU to try (the chip that runs it). You can upload a model's config.json to add your own.
  2. Describe the demand: how many people use it at once (concurrency), and whether that number is the busiest moment or a calmer average. The app always plans for the busiest moment.
  3. Set the average length of questions (input) and answers (output), in tokens (word-pieces).
  4. Set your promises (SLOs): how long a user waits for the first word (TTFT), and how fast words then stream (throughput).

The app then shows the smallest cluster that both fits the model in memory and keeps every promise, plus how other GPUs would compare. The numbers are quick physics-based estimates, not a real benchmark.

Supported model types

  • Language models — chatbots and text models (Llama, Qwen, DeepSeek, GLM, Nemotron). Sized by concurrent users and per-user speed promises (TTFT and throughput).
  • Diffusion models — image generators (SDXL, FLUX.1, SD 3.5, PixArt-Σ, DiT-XL) and video (CogVideoX). Sized by target images or clips per second and a max time per image/clip.
  • Vision-language (VLM) — LLMs that also read images (Qwen2.5-VL, Pixtral, LLaVA-OneVision). Sized like an LLM plus an "images per request" knob that adds visual-patch tokens.
  • Vision-language-action (VLA) — robotics policies (pi-zero, SmolVLA). Sized by target robots × control frequency and a max per-chunk latency.
  • Speech (ASR) — Whisper family (large-v3, large-v3-turbo, medium). Sized by concurrent real-time streams and minimum RTF.
  • Embeddings & rerankers — text encoders (BGE-M3, Jina v3, E5-Mistral) and rerankers (BGE reranker). Sized by target docs (or query-doc pairs) per second and a max time per batch.
  • World models (JEPA) — video models (V-JEPA, V-JEPA 2, and the action-conditioned V-JEPA 2-AC) that turn a clip into an embedding. Sized by target clips per second and a max time per clip.
  • Autoregressive image (VAR) — image generators built like a language model (VAR-d30). Sized like a language model; set output tokens to the image's token count.

Pick the type from the grouped model menu. The demand and SLO controls change to match: users and tokens for language, images or clips per second for the others.

Word list
  • Token — a word-piece. Models read and write tokens, not whole words. Roughly ¾ of a word each.
  • Concurrency — how many requests are being answered at the same moment.
  • P90 — the busy level that only the busiest 10% of moments go above. Higher than the average, lower than the all-time peak.
  • TTFT — "time to first token": how long a user waits before the first word appears.
  • Throughput — how many tokens per second the answer streams at.
  • SLO — a service promise you commit to keeping (like "under 1.5 seconds to first word").
  • TP / PP — two ways to split one model across several GPUs: TP splits each layer's math, PP puts different layers on different GPUs.
Cluster
8 GPU
1 node × 8 · H100 SXM 80GB
Parallelism
TP8 · PP1
1 replicas · 8 GPU/replica
Batch / replica
512
serves 512 concurrent
Cluster throughput
16.5k tok/s
aggregate generation
Per-user TTFT
49 ms
target ≤ 1500 ms
Per-user throughput
32.2 tok/s
target ≥ 30 tok/s
Full request
15.95 s
TTFT + 512 tok
Fits / GPU
Yes
memory-bound decode
Cluster topology 1 node × 8 GPU · TP 8 · DP 1 · Single node
Node 0
g0
g1
g2
g3
g4
g5
g6
g7
TP all-reduce over NVLink · 75.6 GB/s
HBM Memory 73.9 GB / 80.0 GB
Weights 16.4 GB
KV cache 51.0 GB
GPU compute die block size ∝ die area
Tensor cores 29%
CUDA / vector cores 32%
L2 cache 70%
HBM I/O (mem controllers)◄ bottleneck 70%
NVLink I/O 8%
Scheduler · uncore structural
Cost estimate
Estimates only. Rates vary 2-3x between providers, and committed contracts or private pricing can change the bill a lot. Get a real quote before you decide.
Purchasing
GPUs billed
8 ×
$3.47/GPU-hr
Cluster / hour
$27.76
On-demand
Cluster / day
$666
24 × hourly
Cost / 1M tokens
$0.468
output tokens

Default rates are the median on-demand price across GPU cloud providers (GetDeploying GPU price index, 2026-10-07). Providers vary 2-3x around it, so enter your own quote when you have one. Committed and spot apply a typical discount.

Power & energy
Estimates only. First-order approximations, not audited figures. Grid intensities are directional, and TDP-based power under-counts real host draw. For anything you report externally, use measured figures from your provider or your own meters.
Emissions method
Power draw
6.88 kW
551 W / GPU · PUE 1.20
Energy / day
165 kWh
60.3 MWh / year
CO₂e / year
28.6 t
475 g/kWh · location-based
CO₂e / 1M tokens
55.0 g
output tokens

First-order: GPU draw = idle 30% of TDP + linear to TDP with utilization. Host overhead 30% of GPU. PUE 1.20. Grid intensities are directional estimates anchored to public 2024 grid-mix data.

Cold-start / model load
Cold start
1.4 s
source-bound
Per-GPU bytes
17.6 GB
1/8 of the model
PCIe (node)
504 GB/s
63 per GPU × 8
Source
100 GB/s
Parallel file system (Lustre-class)

Wall-clock ≈ weights ÷ min(source, node PCIe). Sharded assumes a well-parallelised loader (e.g. FSDP with parallel range reads). This is where Gen5 PCIe hosts (H100 and newer) pay off vs Gen4 (A100, L4, L40S): ~2× faster once the source can keep up.

GPU comparison — cluster to meet this workload · click a row to use that GPU
Optimise for
GPUGPUsNodesTP·PP·DP$/hour$/1M tokResult
★B300 (Blackwell Ultra) 288GB4 1 2·1·2 $35.28 $0.553 meets SLOs
B200 (Blackwell) 192GB4 1 2·1·2 $34.20 $0.536 meets SLOs
H200 SXM 141GB8 1 4·1·2 $43.20 $0.519 meets SLOs
H100 SXM 80GB8 1 8·1·1 $27.76 $0.468 meets SLOs
A100 SXM 80GB16 2 8·1·2 $29.28 $0.490 meets SLOs
A100 SXM 40GB32 4 8·1·4 $58.56 $0.867 meets SLOs
RTX PRO 6000 Blackwell 96GB32 4 8·1·4 $68.80 $1.05 meets SLOs
RTX 4090 24GB88 11 8·1·11 $39.60 $0.693 meets SLOs
RTX PRO 4500 Blackwell 32GB128 16 8·1·16 — — meets SLOs
L40S 48GB176 22 8·1·22 $276 $4.69 meets SLOs
L4 24GB— min throughput 30 tok/s/user exceeds the hardware ceiling (best 13 tok/s even at batch 1, full tensor-parallel) — a per-user decode limit, not a cluster-size cap
A10G 24GB— min throughput 30 tok/s/user exceeds the hardware ceiling (best 26 tok/s even at batch 1, full tensor-parallel) — a per-user decode limit, not a cluster-size cap
T4 16GB— min throughput 30 tok/s/user exceeds the hardware ceiling (best 14 tok/s even at batch 1, full tensor-parallel) — a per-user decode limit, not a cluster-size cap

First-order sizing: TTFT is a single request's prefill; per-user throughput is the decode rate at the sized batch. Cluster = TP·PP model-parallel groups replicated (DP) until the provisioned concurrency is served. Estimates, not a benchmark.

runs in your browser prices as of 2026-10-07 first-order estimates // not a benchmark