ML//benchmark//multi-hop reasoning

Multi-hop reasoning is answering a question that needs two or more facts chained together, none of which answers it alone, and multi-hop benchmarks are the test sets built to check whether a model, or a retrieval system, can find and connect those facts instead of matching one passage. Given *drug X inhibits protein Y* and *protein Y takes part in process Z*, the question *which process could drug X affect?* needs both hops; a system that retrieves only the first sentence, or that pattern-matches *drug X* to the nearest familiar answer, fails.


Multi-hop reasoning is answering a question that needs two or more facts chained together, none of which answers it alone, and multi-hop benchmarks are the test sets built to check whether a model, or a retrieval system, can find and connect those facts instead of matching one passage. Given drug X inhibits protein Y and protein Y takes part in process Z, the question which process could drug X affect? needs both hops; a system that retrieves only the first sentence, or that pattern-matches drug X to the nearest familiar answer, fails.

It is the question shape of real engineering diagnosis. Asking which production lines are exposed if supplier S stops? chains supplier to part, part to machine, machine to line; why did pump P-101 trip? chains a bearing temperature to a lubrication schedule to a missed work order. A RAG system that pulls the three most similar passages often finds the first link and misses the rest, because the second fact shares no words with the question.

Each hop multiplies the chances of failure.

If a system gets each link right 90 % of the time, a three-hop chain comes out right about 73 % of the time, which is why multi-hop questions separate systems that look equal on single-fact ones.

Building such a test set is often synthetic: start from a structure of known relations (entities and links, as in a knowledge graph), compose chains of two or three hops, and write questions whose answer is the end of the chain. The structure gives the ground truth for free (synthetic data).

A good multi-hop item cannot be shortcut: if one of the facts alone gives away the answer, or the answer is a famous pair, the benchmark measures memory instead of chaining. Distractor passages with similar words are added on purpose.

Systems answer multi-hop questions better when they decompose them: retrieve, read, ask the next sub-question, retrieve again (chain of thought, iterative retrieval, GraphRAG). Scoring the intermediate hops, not only the final answer, shows where the chain broke.

A model that gives the right final answer through a wrong chain is a known failure; for an audit or a safety case the chain itself has to be checked.