ML//inference//test-time compute

Test-time compute is the computation a trained model spends on one request beyond a single pass per output token, and it is the knob that lets a system buy a better answer with more time and money instead of a bigger model. It covers every way of working harder on the problem in front of it: writing a longer chain of reasoning (extended thinking), sampling many answers and keeping the most frequent, searching a tree of partial solutions with an evaluator, or calling tools and checking the results before answering.


Test-time compute is the computation a trained model spends on one request beyond a single pass per output token, and it is the knob that lets a system buy a better answer with more time and money instead of a bigger model. It covers every way of working harder on the problem in front of it: writing a longer chain of reasoning (extended thinking), sampling many answers and keeping the most frequent, searching a tree of partial solutions with an evaluator, or calling tools and checking the results before answering.

The idea is older than language models. AlphaGo combined a network that had learned good moves with Monte Carlo tree search that spent seconds on the position at hand, and AlphaCode generated huge numbers of candidate programs and kept those that passed the example tests. Reasoning models (o1, o3) showed accuracy rising with thinking time on mathematics and code; OpenAI has not published how they search internally, so describing o3 as MCTS is an analogy for search and evaluation, and no more.

More reasoning at inference improves the answer; it does not improve the model.

The weights are the same after the request as before it, so the next request starts from scratch: solving a problem after an hour of thinking teaches the model nothing, unless someone turns that experience into training data (continual learning).

It pays most where an answer can be checked. A verifier (unit tests, a compiler, a proof checker, a constraint check) lets the system generate many candidates and keep a correct one; without a reliable judge, extra samples only add confident alternatives. Open tasks whose quality is not a 0 or 1 (a business decision, a medical explanation) are where scoring is hard and search helps least (reward function).

It costs latency and tokens on every call, which is why it is a design choice per task (inference economics). A planner that runs overnight can think for minutes; the attitude loop of a drone, closing every few milliseconds, cannot think at all, and the learning there happened offline.

The system is more than the model. The same weights look far more capable inside a harness that gives them time, tools and a checker, which is why calling a reasoning system just an LLM is like calling AlphaGo just a convolutional network: true of the component, misleading about the system (agentic system).