ML//benchmark//time horizon

The time horizon of an AI system is the length of task, measured by how long a skilled human takes to do it, that the system completes with a given success rate (usually 50%), and it is METR's measure of how far an agent can work on its own. A model with a one-hour horizon succeeds half the time on tasks that take an expert about an hour.


The time horizon of an AI system is the length of task, measured by how long a skilled human takes to do it, that the system completes with a given success rate (usually 50%), and it is METR's measure of how far an agent can work on its own. A model with a one-hour horizon succeeds half the time on tasks that take an expert about an hour.

The measure turns capability into a unit people can picture and plot over time. METR's measurements show the horizon doubling about every seven months over 2019 to 2025 and faster since 2024, all on software tasks, where success can be checked automatically.

It is measured on code, the domain with the cheapest checks, so its pace says little about domains where checking requires the physical world.

It saturates at the long end. A suite with only a handful of very long tasks cannot place a model whose horizon exceeds them, so the newest models have intervals that span several times their central value.

A doubling time is a slope over a chosen window, and different windows give different rates.