benchmarks.wiki / Public workspace

MBPP+

Code

MBPP+ takes the 378 problems of the sanitized MBPP set and multiplies their tests by roughly 35 times. It removes the same test-weakness inflation that HumanEval+ addresses.

mbpp-plus

Full text200 tasksProblems imported

Official source

Full text: eligible public problems include text, images and discussion.

At a glance

What a problem looks like

A short natural-language programming task. The reference program and the expanded assertions are not part of the prompt.

Read an actual problem

How it is scored

Correctness is decided by executing the candidate against the expanded assertions, so the reference program the source ships is not a graded answer; it stays in restricted storage.

Metric: pass@1 on the expanded tests

Why it is hard

Each task is a few lines of Python, but the specification is one sentence long, so most of the difficulty is in reading what is actually being asked for.

Cite

@article{liu2023mbppplus,
  title={Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation},
  author={Jiawei Liu and Chunqiu Steven Xia and Yuyao Wang and Lingming Zhang},
  journal={arXiv preprint arXiv:2305.01210},
  year={2023}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Tasks

Whole public benchmark.

Public problems
200
Formats
code (200)
Problems with images
0
Answer availability
Answer published by the source (200)
Configs
default (200)
Splits
test (200)
Awaiting a first discussion
200
Problems with a discussion
0
Public contributions
0

Provenance and terms

Pinned source revisions, newest first:

  • b2d74c91837c3f2a20c1299ae98133cbe7cfa077 · acquired 2026-09-12 · release 10a7a2abecad
  • b2d74c91837c3f2a20c1299ae98133cbe7cfa077 · acquired 2026-09-12 · release 532e203491c3