ML//AI progress//jagged frontier

The jagged frontier is the irregular boundary between the tasks an AI system does well and those it does badly, in which tasks of similar apparent difficulty can fall on opposite sides, and it is used to decide which work to hand to a model and which to keep. Picture capability as a mountain range over domains (translation, vision, chemistry, driving, social interaction, art): some peaks are high, some valleys deep, and the height of one peak says little about its neighbour. A model can explain advanced chemistry and then misjudge which of two stacked boxes will fall.


The jagged frontier is the irregular boundary between the tasks an AI system does well and those it does badly, in which tasks of similar apparent difficulty can fall on opposite sides, and it is used to decide which work to hand to a model and which to keep. Picture capability as a mountain range over domains (translation, vision, chemistry, driving, social interaction, art): some peaks are high, some valleys deep, and the height of one peak says little about its neighbour. A model can explain advanced chemistry and then misjudge which of two stacked boxes will fall.

The name comes from a 2023 field experiment by researchers from Harvard Business School and other universities with Boston Consulting Group, in which 758 consultants worked on realistic tasks with or without GPT-4 (Dell'Acqua et al., 2023). On tasks inside the model's frontier, those with the model finished 12.2 % more tasks, 25.1 % faster and at more than 40 % higher rated quality; on a task chosen to lie just outside it, they were 19 percentage points less likely to reach the right answer than colleagues without it, because the model's output looked as convincing there as anywhere else.

The danger lies on the frontier, where the model is wrong and sounds right.

A general reputation for capability tells nothing about one particular task, so delegation has to be decided task by task, with a check on the output wherever an error is costly.

It is measured task by task. One aggregate score hides the shape; per-task benchmarks and the time horizon of agents show where the peaks and valleys are. An engineer deploying a model in a plant evaluates it on that plant's documents and questions, never on its leaderboard rank.

It moves unevenly too. Progress raises some peaks quickly (mathematics and code, where answers can be checked and so trained against) and others slowly (open judgment, physical interaction), which connects it to the difficulty of writing a reward for open tasks (reward function) and to Moravec's paradox for the bodily skills.

It is an argument about the shape of capability, never about its height. The frontier can advance everywhere and stay jagged, which is why a system that keeps a person or a verifier at the right points outperforms one that trusts the model uniformly.

It belongs to the wider study of AI progress, which asks how fast the whole range rises.