Reasoning
ARC-AGI-2
arc-agi-2
ARC-AGI-2 evaluates the ability to infer unfamiliar rules from a few examples and apply them to new grids. Training and evaluation are separate; this import covers the 120 public evaluation tasks.
Benchmark AI / Public workspace
Choose a field to explore public problems and discussions.
Fields and counts refer to this catalog page.
Reasoning
arc-agi-2
ARC-AGI-2 evaluates the ability to infer unfamiliar rules from a few examples and apply them to new grids. Training and evaluation are separate; this import covers the 120 public evaluation tasks.
Reasoning / Agents
arc-agi-3
ARC-AGI-3 is an interactive benchmark for exploring unfamiliar game environments, learning their rules, and planning actions. The official toolkit provides an environment interface, with game versions and trial records t…
Reasoning
bbeh
BIG-Bench Extra Hard evaluates language-model reasoning through challenging task categories. It offers full and mini distributions; this import contains 200 Boolean-expression questions from the full distribution.
Agents / Search
browsecomp
BrowseComp evaluates web browsing and the ability to identify answers from multiple clues. This catalog supports introductions and discussion of search methods without republishing decrypted questions or answers.
Agents / Search
browsecomp-plus
BrowseComp-Plus supports evaluation of retrieval methods using BrowseComp tasks and a search corpus. This catalog links to the official distribution without republishing plaintext questions or corpus content.
Science
critpt
CritPt evaluates scientific understanding and multi-step reasoning and computation on research-level physics problems. Its public dataset contains 70 challenges with problem descriptions and code templates.
Mathematics
frontiermath
FrontierMath evaluates advanced mathematical reasoning using problems written by experts. This catalog links to the official introduction and public examples while excluding private evaluation problems.
Science
frontierscience
FrontierScience evaluates the ability to solve expert-level scientific tasks. Its public data separates olympiad and research problems so that competition and research tasks can be examined independently.
Agents / Multimodal
gaia
GAIA evaluates general AI assistants on tasks involving information gathering and tool use. Some tasks include attachments; this catalog provides an introduction in view of access and resharing restrictions.
Code
hlce
Humanity's Last Code Exam evaluates problem solving and code generation using ICPC World Finals and IOI problems. The project describes 235 problems; this import selects its public ICPC dataset of 146 problems.
Reasoning / Multimodal
hle
Humanity's Last Exam evaluates AI knowledge and reasoning across academic disciplines. It includes text and image questions; this catalog provides an introduction and official links in accordance with the publisher's req…
Code
livecodebench
LiveCodeBench continuously collects new competitive-programming problems to evaluate coding capabilities. Its initial release_v1 contains 400 problems, with source platform, contest date, and dataset version tracked expl…
Reasoning / Long context
longbench-v2
LongBench v2 evaluates deep understanding and reasoning over long contexts through multiple-choice questions. Its official description lists 503 questions spanning tasks such as single-document and multi-document QA and …
Multimodal
mmmu-pro
MMMU-Pro evaluates multimodal understanding across academic disciplines using images and text. The official card lists 1,730 questions per setting, separating four-option and ten-option standard settings from the vision …
Mathematics / Science / Multimodal
olympiadbench
OlympiadBench evaluates scientific reasoning on Olympiad-level mathematics and physics problems. Its official description lists 8,476 English and Chinese problems with separate text-only and multimodal settings.
Mathematics
omni-math
Omni-MATH evaluates mathematical reasoning on Olympiad-level problems. Its official dataset contains 4,428 problems accompanied by domain and difficulty information.
Science / Code
scicode
SciCode evaluates the ability to solve scientific research problems through code. Problems are decomposed into subproblems; this dev import preserves the relationships between 15 parent problems and 50 subproblems.
Code / Agents
swe-bench-pro
SWE-Bench Pro evaluates agents on long-horizon software engineering tasks in real repositories. The public card describes 731 tasks containing issue descriptions, repository identifiers, and base commits.
Agents / Code
terminal-bench-2-1
Terminal-Bench 2.1 evaluates agents performing tasks in terminal environments. Each task supplies instructions and environment configuration, and version 2.1 is tracked separately from 2.0.
Reasoning
zebralogicbench
ZebraLogicBench evaluates reasoning on logic puzzles that require satisfying clues to determine attribute assignments. The official card lists 1,000 grid-mode puzzles and a separate multiple-choice setting.