Self-host vs API
When does owning GPUs beat paying per token? Configure the exact cluster (same controls as Modelling), set your API price, find the crossover.
How to use this tab
This tab answers "should I run this model myself, or just call an API?". Self-hosting means renting GPUs and running the model yourself — you pay for the whole cluster around the clock, busy or not. An API charges per token — only for what you use, but each token costs a bit more.
- Configure the cluster with the full Modelling controls — model, GPU, GPU count, parallelism, quantization, and the request shape (input/output tokens). The forward model computes the cluster's throughput (the same engine as the Modelling tab).
- Set the duty cycle: what fraction of the month that cluster is actually busy. Flat-out all day is very different from spiking at lunch and idle overnight.
- Type the API price for the same model (input and output $/1M). Seeded from the model's real market price — replace with your quote.
The app shows which is cheaper today, the break-even duty cycle, and a chart of both. The API is a straight line up from zero; self-hosting is a staircase — one cluster serves only so much, so cost jumps a step as you add servers.
This month, at 40% duty cycle
Self-host is billed 24/7 for the cluster you configured. API volume is derived from the cluster's peak throughput scaled by the duty cycle, so both sides serve the same traffic.