benchmarks.wiki / Public workspace

WinoGrande

Reasoning

WinoGrande scales the Winograd Schema Challenge to 44,000 problems in which a pronoun must be resolved between two sentences that differ by one word. It was built through a procedure that strips dataset-side bias.

winogrande

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

At a glance

What a problem looks like

A sentence with a blank and two candidate fillers.

How it is scored

The source ships the correct option; nothing is imported here.

Metric: accuracy, reported across training-set sizes from 160 to 40,000 examples

Why it is hard

The two sentences differ by a single word, so every cue except the intended commonsense one is held constant.

Cite

@article{sakaguchi2019winogrande,
  title={WinoGrande: An Adversarial Winograd Schema Challenge at Scale},
  author={Keisuke Sakaguchi and Ronan Le Bras and Chandra Bhagavatula and Yejin Choi},
  journal={arXiv preprint arXiv:1907.10641},
  year={2019}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion
Problems with a discussion
Public contributions