benchmarks.wiki / Public workspace
WinoGrande
Reasoning
WinoGrande scales the Winograd Schema Challenge to 44,000 problems in which a pronoun must be resolved between two sentences that differ by one word. It was built through a procedure that strips dataset-side bias.
winogrande
Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.
Reference links
At a glance
What a problem looks like
A sentence with a blank and two candidate fillers.
How it is scored
The source ships the correct option; nothing is imported here.
Metric: accuracy, reported across training-set sizes from 160 to 40,000 examples
Why it is hard
The two sentences differ by a single word, so every cue except the intended commonsense one is held constant.
Cite
@article{sakaguchi2019winogrande,
title={WinoGrande: An Adversarial Winograd Schema Challenge at Scale},
author={Keisuke Sakaguchi and Ronan Le Bras and Chandra Bhagavatula and Yejin Choi},
journal={arXiv preprint arXiv:1907.10641},
year={2019}
}Discussion
Contribute
Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.
Public problem summary
Catalogue metadata only; public problem statistics are not available.
- Public problems
- —
- Formats
- —
- Problems with images
- —
- Answer availability
- —
- Configs
- —
- Splits
- —
- Awaiting a first discussion
- —
- Problems with a discussion
- —
- Public contributions
- —