Training / Fine-tuning Calculator
Will it fit to train, and how long will it take? First-order estimates for full fine-tuning, LoRA and QLoRA.
How to use this tab
This tab sizes training and fine-tuning, not serving. The big difference from the other tabs is memory: to train a model you must hold not just its weights, but also its gradients, the optimizer's running state, and the activations from the forward pass. With normal AdamW that is about 16 bytes for every parameter — so a model that runs on one GPU often needs six to eight times the memory to train.
Full vs LoRA vs QLoRA
- Full fine-tune — trains every weight. Most accurate, most memory: gradients and optimizer state for the whole model.
- LoRA — freezes the model and trains tiny added "adapter" weights. Only the adapters need gradients and optimizer state, so the memory drops enormously.
- QLoRA — LoRA on top of a model squeezed to 4-bit. This is how a 70B model fine-tunes on a single 80 GB GPU.
The controls
- ZeRO / sharding — split the model states across your data-parallel GPUs so each holds a slice. Higher stages shard more (optimizer → +gradients → +weights).
- Activation checkpointing — throw away activations and recompute them in the backward pass. Saves a lot of memory for ~30% more compute.
- Micro-batch × grad-accum — the effective batch is micro-batch × grad-accum × replicas. Grad-accum lets a small micro-batch reach a large effective batch.
Numbers are first-order estimates, not a benchmark. v1 covers transformer language and visual-AR models; diffusion and JEPA training will follow. To size serving instead, use the Modelling or Workload tabs.
Global batch 0M tok = 1 × 16 accum × 1 replicas × 4k.
Default rates are the median on-demand price across GPU cloud providers (GetDeploying GPU price index, 2026-10-07). Providers vary 2-3x around it, so enter your own quote when you have one. Committed and spot apply a typical discount.
First-order: GPU draw = idle 30% of TDP + linear to TDP with utilization. Host overhead 30% of GPU. PUE 1.20. Grid intensities are directional estimates anchored to public 2024 grid-mix data.
Wall-clock ≈ weights ÷ min(source, node PCIe). Sharded assumes a well-parallelised loader (e.g. FSDP with parallel range reads). This is where Gen5 PCIe hosts (H100 and newer) pay off vs Gen4 (A100, L4, L40S): ~2× faster once the source can keep up.
First-order training roofline (MFU 0.45) with 2.0 GB overhead/GPU. Model states use the mixed-precision AdamW convention (weight 2 + grad 2 + optimizer 12 B/param). Activation memory follows the Korthikanti et al. per-layer estimate. Pipeline bubble and exact comm overlap are approximate; estimates, not a benchmark.