Choose a field to explore public problems and discussions.
Full text
Problem text is available, with images and a page per problem. The API reports this as full.
Overview only
An introduction, the official source and general discussion are available; problem text, images and individual pages are not published here. The API reports this as metadata_only.
ARC-AGI-3 is an interactive benchmark for exploring unfamiliar game environments, learning their rules, and planning actions. The official toolkit provides an environment interface, with game versions and trial records t…
BrowseComp evaluates web browsing and the ability to identify answers from multiple clues. This catalog supports introductions and discussion of search methods without republishing decrypted questions or answers.
BrowseComp-Plus supports evaluation of retrieval methods using BrowseComp tasks and a search corpus. This catalog links to the official distribution without republishing plaintext questions or corpus content.
GAIA evaluates general AI assistants on tasks involving information gathering and tool use. Some tasks include attachments; this catalog provides an introduction in view of access and resharing restrictions.
SWE-Bench Pro evaluates agents on long-horizon software engineering tasks in real repositories. The public card describes 731 tasks containing issue descriptions, repository identifiers, and base commits.
Terminal-Bench 2.1 evaluates agents performing tasks in terminal environments. Each task supplies instructions and environment configuration, and version 2.1 is tracked separately from 2.0.