Reasoning
ARC-AGI-2
arc-agi-2
ARC-AGI-2 evaluates the ability to infer unfamiliar rules from a few examples and apply them to new grids. Training and evaluation are separate; this import covers the 120 public evaluation tasks.
Benchmark AI / Public workspace
Choose a field to explore public problems and discussions.
Fields and counts refer to this catalog page.
Reasoning
arc-agi-2
ARC-AGI-2 evaluates the ability to infer unfamiliar rules from a few examples and apply them to new grids. Training and evaluation are separate; this import covers the 120 public evaluation tasks.
Reasoning / Agents
arc-agi-3
ARC-AGI-3 is an interactive benchmark for exploring unfamiliar game environments, learning their rules, and planning actions. The official toolkit provides an environment interface, with game versions and trial records t…
Reasoning
bbeh
BIG-Bench Extra Hard evaluates language-model reasoning through challenging task categories. It offers full and mini distributions; this import contains 200 Boolean-expression questions from the full distribution.
Reasoning / Multimodal
hle
Humanity's Last Exam evaluates AI knowledge and reasoning across academic disciplines. It includes text and image questions; this catalog provides an introduction and official links in accordance with the publisher's req…
Reasoning / Long context
longbench-v2
LongBench v2 evaluates deep understanding and reasoning over long contexts through multiple-choice questions. Its official description lists 503 questions spanning tasks such as single-document and multi-document QA and …
Reasoning
zebralogicbench
ZebraLogicBench evaluates reasoning on logic puzzles that require satisfying clues to determine attribute assignments. The official card lists 1,000 grid-mode puzzles and a separate multiple-choice setting.