ML//benchmark//BIG-Bench

- BIG-Bench is a collaborative benchmark containing more than 200 tasks across reasoning, language, knowledge, mathematics, and social understanding.


BIG-Bench is a collaborative benchmark containing more than 200 tasks across reasoning, language, knowledge, mathematics, and social understanding.

BIG-Bench Hard (BBH): the subset where models initially struggled (multi-step reasoning focus).

Its breadth is the point and the problem: an aggregate score can hide catastrophic failure on a narrow task. It is more useful as a capability microscope than as a single-number leaderboard.

evaluation : : Broad suites reveal uneven capability profiles that headline averages compress away