Agents / Multimodal
GAIA
gaia
GAIA evaluates general AI assistants on tasks involving information gathering and tool use. Some tasks include attachments; this catalog provides an introduction in view of access and resharing restrictions.
Benchmark AI / Public workspace
Choose a field to explore public problems and discussions.
Fields and counts refer to this catalog page.
Agents / Multimodal
gaia
GAIA evaluates general AI assistants on tasks involving information gathering and tool use. Some tasks include attachments; this catalog provides an introduction in view of access and resharing restrictions.
Reasoning / Multimodal
hle
Humanity's Last Exam evaluates AI knowledge and reasoning across academic disciplines. It includes text and image questions; this catalog provides an introduction and official links in accordance with the publisher's req…
Multimodal
mmmu-pro
MMMU-Pro evaluates multimodal understanding across academic disciplines using images and text. The official card lists 1,730 questions per setting, separating four-option and ten-option standard settings from the vision …
Mathematics / Science / Multimodal
olympiadbench
OlympiadBench evaluates scientific reasoning on Olympiad-level mathematics and physics problems. Its official description lists 8,476 English and Chinese problems with separate text-only and multimodal settings.