ML//benchmark

A benchmark is a standardized task distribution plus a scoring procedure used to compare systems. In ML, examples include MMLU and HumanEval. The score compresses many observations into something comparable, which is both its value and its danger.


A benchmark is a standardized task distribution plus a scoring procedure used to compare systems. In ML, examples include MMLU and HumanEval. The score compresses many observations into something comparable, which is both its value and its danger.

Contamination, test-format overfitting, cherry-picking and harness differences can all make a clean number answer a dirtier question than readers assume. Benchmarks remain useful when the task, environment, uncertainty and failure modes travel with the score.

A leaderboard is a map. Optimization pressure has a habit of moving the roads toward whatever the map happens to measure.