Benchmark AI / Public workspace

SWE-Bench Pro

Code / Agents

SWE-Bench Pro evaluates agents on long-horizon software engineering tasks in real repositories. The public card describes 731 tasks containing issue descriptions, repository identifiers, and base commits.

Japanese introduction

実際のソフトウェアリポジトリに対する修正課題で、長い工程を要する開発能力を評価します。公開データカードの731課題には問題文、対象リポジトリ、修正開始点のcommitなどが含まれます。

swe-bench-pro

Full text200 tasksProblems imported

Official source

Full text: eligible public problems include text, images and discussion.

Tasks

Whole public benchmark.

Public problems
200
Formats
text (200)
Problems with images
0
Answer availability
Answer published by the source (200)
Configs
Default (200)
Splits
test (200)
Awaiting a first discussion
200

Discussion