ML//training//training cost

Training cost is the money and resources spent to produce a model's weights, and it is really three different magnitudes that public figures routinely mix: the compute of one final training run, the total cost of the research and development that made that run possible, and the cost of serving the finished model to users. Keeping them apart is what lets an engineer read a headline about a cheap model, or budget a project of their own, without comparing a single run with someone else's whole programme.


Training cost is the money and resources spent to produce a model's weights, and it is really three different magnitudes that public figures routinely mix: the compute of one final training run, the total cost of the research and development that made that run possible, and the cost of serving the finished model to users. Keeping them apart is what lets an engineer read a headline about a cheap model, or budget a project of their own, without comparing a single run with someone else's whole programme.

The first magnitude is the one that is easy to compute: accelerator hours times a price per hour. A run on 2,000 GPUs for 60 days is about 2.9 million GPU-hours, and at 2 dollars an hour that is about 6 million dollars. DeepSeek's V3 report gave exactly this kind of figure (about 2.8 million H800 hours, about 5.6 million dollars at an assumed rental price) and stated that it excluded earlier research and ablations; read as the cost of the model, it misled many.

Compute of a run covers the final pre-training pass and, often reported apart, the post-training stages. It scales with parameters times tokens (scaling laws) and falls every year with better methods (algorithmic progress) and better hardware.

Total development cost adds the failed and exploratory runs (often several times the final one), the people, the data collection, filtering and labelling, the evaluation and red teaming, and the hardware owned or reserved for all of it (CapEx). It is the number a founder actually has to raise.

Serving cost is paid per request for as long as the model is used: accelerators, memory for the KV cache, energy and operations. For a popular model it overtakes the training bill, which is why the split between training and inference matters (inference economics).

Compare like with like.

A run's compute, a lab's development budget and a product's serving bill differ by orders of magnitude; a claim about one says almost nothing about the others, and architecture choices move them separately (a mixture of experts cuts compute per token, a smaller KV cache cuts serving memory).

For a project of one's own the same accounting runs in reverse: write the scope document first, estimate the tokens each feature consumes, pick the smallest model that meets the quality bar, and only then decide whether to train, fine-tune with LoRA or call an API.