benchmarks.wiki / Public workspace

SimpleQA

Knowledge / Search

SimpleQA collects short fact-seeking questions that have exactly one indisputable answer, measuring both how often a model is right and how often it declines. The format is deliberately narrow so that grading is not itself contestable.

simpleqa

Full text200 tasksProblems imported

Official source

Full text: eligible public problems include text, images and discussion.

At a glance

What a problem looks like

One short question. The reference answer and the URLs that document it are not part of the prompt.

Read an actual problem

How it is scored

The source ships a reference answer with every question; this catalogue keeps it in restricted storage and never renders it.

Metric: correct, incorrect and not-attempted rates, and the F-score that combines them

Why it is hard

The questions are short but the facts are obscure, and a model that guesses scores worse than one that abstains. It measures calibration as much as knowledge.

Cite

@article{wei2024simpleqa,
  title={Measuring short-form factuality in large language models},
  author={Jason Wei and Nguyen Karina and Hyung Won Chung and Yunxin Joy Jiao and Spencer Papay and Amelia Glaese and John Schulman and William Fedus},
  journal={arXiv preprint arXiv:2411.04368},
  year={2024}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Tasks

Whole public benchmark.

Public problems
200
Formats
text (200)
Problems with images
0
Answer availability
Answer published by the source (200)
Configs
default (200)
Splits
test (200)
Awaiting a first discussion
200
Problems with a discussion
0
Public contributions
0

Provenance and terms

Pinned source revisions, newest first:

  • e319282ab125c3dbd0c7fd00be2e4dd54e7e8f94 · acquired 2026-09-12 · release 2f3ac1c4c266