benchmarks.wiki / Public workspace

BIG-Bench Hard

Reasoning

BIG-Bench Hard is the set of 23 BIG-Bench tasks on which models of the time scored below the average human rater. It is best known for showing how much chain-of-thought prompting changes the result.

bbh

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

At a glance

What a problem looks like

Varies by task: boolean expressions, date understanding, object counting, logical deduction and others.

How it is scored

Each task ships its own target; nothing is imported here.

Why it is hard

These are exactly the tasks that resisted scale, and several of them need an explicit multi-step procedure rather than a single judgement.

Cite

@article{suzgun2022bbh,
  title={Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them},
  author={Mirac Suzgun and Nathan Scales and Nathanael Schärli and Sebastian Gehrmann and Yi Tay and Hyung Won Chung and Aakanksha Chowdhery and Quoc V. Le and Ed H. Chi and Denny Zhou and Jason Wei},
  journal={arXiv preprint arXiv:2210.09261},
  year={2022}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion
Problems with a discussion
Public contributions