benchmarks.wiki / Public workspace
BIG-Bench Hard
Reasoning
BIG-Bench Hard is the set of 23 BIG-Bench tasks on which models of the time scored below the average human rater. It is best known for showing how much chain-of-thought prompting changes the result.
bbh
Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.
Reference links
At a glance
What a problem looks like
Varies by task: boolean expressions, date understanding, object counting, logical deduction and others.
How it is scored
Each task ships its own target; nothing is imported here.
Why it is hard
These are exactly the tasks that resisted scale, and several of them need an explicit multi-step procedure rather than a single judgement.
Cite
@article{suzgun2022bbh,
title={Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them},
author={Mirac Suzgun and Nathan Scales and Nathanael Schärli and Sebastian Gehrmann and Yi Tay and Hyung Won Chung and Aakanksha Chowdhery and Quoc V. Le and Ed H. Chi and Denny Zhou and Jason Wei},
journal={arXiv preprint arXiv:2210.09261},
year={2022}
}Discussion
Contribute
Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.
Public problem summary
Catalogue metadata only; public problem statistics are not available.
- Public problems
- —
- Formats
- —
- Problems with images
- —
- Answer availability
- —
- Configs
- —
- Splits
- —
- Awaiting a first discussion
- —
- Problems with a discussion
- —
- Public contributions
- —