Benchmark AI / Public workspace

GAIA

Agents / Multimodal

GAIA evaluates general AI assistants on tasks involving information gathering and tool use. Some tasks include attachments; this catalog provides an introduction in view of access and resharing restrictions.

Japanese introduction

情報探索やツール利用を伴う課題で、汎用AIアシスタントの能力を評価するベンチマークです。添付資料を含む課題もあり、アクセスと再共有の条件に従って紹介のみを掲載します。

gaia

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion

Discussion