benchmarks.wiki / Public workspace
SWE-bench Verified
Code / Agents
SWE-bench Verified is the 500-issue subset of SWE-bench that human annotators confirmed to be solvable and correctly specified. Each task is a real GitHub issue in a Python repository.
swe-bench-verified
Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.
Reference links
At a glance
What a problem looks like
A real GitHub issue with the repository, base commit and the tests that must go from failing to passing.
How it is scored
The dataset declares no licence, and a run needs the repository checkout and test harness rather than the issue text alone. The catalogue links to the official distribution.
Why it is hard
A patch must make specific tests pass without breaking others, inside a real codebase the model has not seen.
Discussion
Contribute
Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.
Public problem summary
Catalogue metadata only; public problem statistics are not available.
- Public problems
- —
- Formats
- —
- Problems with images
- —
- Answer availability
- —
- Configs
- —
- Splits
- —
- Awaiting a first discussion
- —
- Problems with a discussion
- —
- Public contributions
- —