◢◤ GENAI_CALC_
SYS_TIME UTC+0

Training / Fine-tuning Calculator

Will it fit to train, and how long will it take? First-order estimates for full fine-tuning, LoRA and QLoRA.

How to use this tab

This tab sizes training and fine-tuning, not serving. The big difference from the other tabs is memory: to train a model you must hold not just its weights, but also its gradients, the optimizer's running state, and the activations from the forward pass. With normal AdamW that is about 16 bytes for every parameter — so a model that runs on one GPU often needs six to eight times the memory to train.

Full vs LoRA vs QLoRA

  • Full fine-tune — trains every weight. Most accurate, most memory: gradients and optimizer state for the whole model.
  • LoRA — freezes the model and trains tiny added "adapter" weights. Only the adapters need gradients and optimizer state, so the memory drops enormously.
  • QLoRA — LoRA on top of a model squeezed to 4-bit. This is how a 70B model fine-tunes on a single 80 GB GPU.

The controls

  • ZeRO / sharding — split the model states across your data-parallel GPUs so each holds a slice. Higher stages shard more (optimizer → +gradients → +weights).
  • Activation checkpointing — throw away activations and recompute them in the backward pass. Saves a lot of memory for ~30% more compute.
  • Micro-batch × grad-accum — the effective batch is micro-batch × grad-accum × replicas. Grad-accum lets a small micro-batch reach a large effective batch.

Numbers are first-order estimates, not a benchmark. v1 covers transformer language and visual-AR models; diffusion and JEPA training will follow. To size serving instead, use the Modelling or Workload tabs.

Fits to train
Yes
59.2 GB free / GPU
Trainable params
225M
0.32% of 71B
Training throughput
7.2k tok/s
1 replicas · 8 GPU each
Time to train
4.8 days
1B × 3 ep · compute-bound
Step time
9.09 s
16 micro-steps + sync
MFU
45 %
445 TFLOP/s / GPU
Per-GPU memory to train
HBM Memory 20.8 GB / 80.0 GB
Base weights (frozen) 16.5 GB
Weights (frozen) 16.5 GB
Gradients 54 MB
Optimizer states 323 MB
Activations 2.9 GB
Overhead 1.0 GB
Total / GPU 20.8 GB / 80.0 GB

Global batch 0M tok = 1 × 16 accum × 1 replicas × 4k.

Cost estimate
Estimates only. Rates vary 2-3x between providers, and committed contracts or private pricing can change the bill a lot. Get a real quote before you decide.
Purchasing
GPUs billed
8 ×
$3.47/GPU-hr
Cluster / hour
$27.76
On-demand
Cluster / day
$666
24 × hourly
Total for the run
$3,209
4.8 days × $27.76/hr

Default rates are the median on-demand price across GPU cloud providers (GetDeploying GPU price index, 2026-10-07). Providers vary 2-3x around it, so enter your own quote when you have one. Committed and spot apply a typical discount.

Power & energy
Estimates only. First-order approximations, not audited figures. Grid intensities are directional, and TDP-based power under-counts real host draw. For anything you report externally, use measured figures from your provider or your own meters.
Emissions method
Power draw
7.82 kW
627 W / GPU · PUE 1.20
Energy / day
188 kWh
68.5 MWh / year
CO₂e / year
32.5 t
475 g/kWh · location-based
CO₂e for the run
429 kg
115.6 h × 7.82 kW

First-order: GPU draw = idle 30% of TDP + linear to TDP with utilization. Host overhead 30% of GPU. PUE 1.20. Grid intensities are directional estimates anchored to public 2024 grid-mix data.

Cold-start / model load
Cold start
1.4 s
source-bound
Per-GPU bytes
17.6 GB
1/8 of the model
PCIe (node)
504 GB/s
63 per GPU × 8
Source
100 GB/s
Parallel file system (Lustre-class)

Wall-clock ≈ weights ÷ min(source, node PCIe). Sharded assumes a well-parallelised loader (e.g. FSDP with parallel range reads). This is where Gen5 PCIe hosts (H100 and newer) pay off vs Gen4 (A100, L4, L40S): ~2× faster once the source can keep up.

First-order training roofline (MFU 0.45) with 2.0 GB overhead/GPU. Model states use the mixed-precision AdamW convention (weight 2 + grad 2 + optimizer 12 B/param). Activation memory follows the Korthikanti et al. per-layer estimate. Pipeline bubble and exact comm overlap are approximate; estimates, not a benchmark.

runs in your browser prices as of 2026-10-07 first-order estimates // not a benchmark