benchmarks.wiki / Public workspace

HellaSwag

Reasoning

HellaSwag asks which of four continuations naturally follows a sentence or a video caption. The wrong options are produced by adversarial filtering, so items that are nearly trivial for people remain hard for models.

hellaswag

Full text200 tasksProblems imported

Official source

Full text: eligible public problems include text, images and discussion.

At a glance

What a problem looks like

An activity label and a context sentence, with the four candidate continuations as a list. Which one is correct is not part of the prompt.

Read an actual problem

How it is scored

The source ships the index of the correct continuation; this catalogue keeps it in restricted storage and never renders it.

Metric: accuracy over the 10,042 validation items

Why it is hard

The distractors were selected specifically because models rank them highly, so the gap it measures is between surface plausibility and actual commonsense.

Cite

@article{zellers2019hellaswag,
  title={HellaSwag: Can a Machine Really Finish Your Sentence?},
  author={Rowan Zellers and Ari Holtzman and Yonatan Bisk and Ali Farhadi and Yejin Choi},
  journal={arXiv preprint arXiv:1905.07830},
  year={2019}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Tasks

Whole public benchmark.

Public problems
200
Formats
text (200)
Problems with images
0
Answer availability
Answer published by the source (200)
Configs
default (200)
Splits
validation (200)
Awaiting a first discussion
200
Problems with a discussion
0
Public contributions
0

Provenance and terms

Pinned source revisions, newest first:

  • 218ec52e09a7e7462a5400043bb9a69a41d06b76 · acquired 2026-09-12 · release 7fb6311c6df2
  • 218ec52e09a7e7462a5400043bb9a69a41d06b76 · acquired 2026-09-12 · release da7444a35989