benchmarks.wiki / Public workspace

MBPP

Code

Mostly Basic Python Problems asks for a short Python function from a one-sentence description, checked by assertions. The sanitized split is the hand-verified subset.

mbpp

Full text200 tasksProblems imported

Official source

Full text: eligible public problems include text, images and discussion.

At a glance

What a problem looks like

A one-sentence task description in prompt. The reference code and the assertion list are kept out of the prompt.

Read an actual problem

How it is scored

For the 200 imported problems, every column outside the prompt is separated into restricted storage, but the converter does not interpret it, so this catalogue does not claim a usable answer key exists.

Why it is hard

The description is short enough to be ambiguous, so the model has to infer the intended signature and behaviour from the tests it cannot see.

Cite

@article{austin2021program,
  title={Program Synthesis with Large Language Models},
  author={Austin, Jacob and Odena, Augustus and Nye, Maxwell and Bosma, Maarten and Michalewski, Henryk and Dohan, David and Jiang, Ellen and Cai, Carrie and Terry, Michael and Le, Quoc and others},
  journal={arXiv preprint arXiv:2108.07732},
  year={2021}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Tasks

Whole public benchmark.

Public problems
200
Formats
text (200)
Problems with images
0
Answer availability
Answer availability not recorded (200)
Configs
sanitized (200)
Splits
test (200)
Awaiting a first discussion
200
Problems with a discussion
0
Public contributions
0

Provenance and terms

Pinned source revisions, newest first:

  • 4bb6404fdc6cacfda99d4ac4205087b89d32030c · acquired 2026-09-11 · release a9c6c8747255