ML//benchmark//BIG-Bench
- BIG-Bench is a collaborative benchmark containing more than 200 tasks across reasoning, language, knowledge, mathematics, and social understanding.
BIG-Bench is a collaborative benchmark containing more than 200 tasks across reasoning, language, knowledge, mathematics, and social understanding.
BIG-Bench Hard (BBH): the subset where models initially struggled (multi-step reasoning focus).
Its breadth is the point and the problem: an aggregate score can hide catastrophic failure on a narrow task. It is more useful as a capability microscope than as a single-number leaderboard.
evaluation : : Broad suites reveal uneven capability profiles that headline averages compress away