benchmarks.wiki / Public workspace

MMLU

Reasoning / Knowledge

MMLU asks four-way multiple-choice questions across 57 subjects, from elementary material to professional-level examinations. Since 2020 it has been the most widely cited single reference point for general capability.

mmlu

Full text200 tasksProblems imported

Official source

Full text: eligible public problems include text, images and discussion.

At a glance

What a problem looks like

A question with four options, plus the subject it belongs to. The options are imported as a list; the index of the correct one is not.

Read an actual problem

How it is scored

The source ships the answer index with every problem; this catalogue keeps it in restricted storage and never renders it.

Metric: accuracy (fraction of questions answered correctly, usually reported as a macro average over the 57 subjects)

Why it is hard

No single subject is hard; the breadth is. A model must hold professional-level material in law, medicine and mathematics at once, and the four-way format leaves little room to recover from a gap in knowledge.

Cite

@article{hendrycks2020mmlu,
  title={Measuring Massive Multitask Language Understanding},
  author={Dan Hendrycks and Collin Burns and Steven Basart and Andy Zou and Mantas Mazeika and Dawn Song and Jacob Steinhardt},
  journal={arXiv preprint arXiv:2009.03300},
  year={2020}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Tasks

Whole public benchmark.

Public problems
200
Formats
text (200)
Problems with images
0
Answer availability
Answer published by the source (200)
Configs
all (200)
Splits
test (200)
Awaiting a first discussion
200
Problems with a discussion
0
Public contributions
0

Provenance and terms

Pinned source revisions, newest first:

  • c30699e8356da336a370243923dbaf21066bb9fe · acquired 2026-09-12 · release 8cfe6e6995d0