benchmarks.wiki / Public workspace

SWE-bench Verified

Code / Agents

SWE-bench Verified is the 500-issue subset of SWE-bench that human annotators confirmed to be solvable and correctly specified. Each task is a real GitHub issue in a Python repository.

swe-bench-verified

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

At a glance

What a problem looks like

A real GitHub issue with the repository, base commit and the tests that must go from failing to passing.

How it is scored

The dataset declares no licence, and a run needs the repository checkout and test harness rather than the issue text alone. The catalogue links to the official distribution.

Why it is hard

A patch must make specific tests pass without breaking others, inside a real codebase the model has not seen.

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion
Problems with a discussion
Public contributions