Benchmark AI / Public workspace
GAIA
Agents / Multimodal
GAIA evaluates general AI assistants on tasks involving information gathering and tool use. Some tasks include attachments; this catalog provides an introduction in view of access and resharing restrictions.
Japanese introduction
情報探索やツール利用を伴う課題で、汎用AIアシスタントの能力を評価するベンチマークです。添付資料を含む課題もあり、アクセスと再共有の条件に従って紹介のみを掲載します。
gaia
Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.
Public problem summary
Catalogue metadata only; public problem statistics are not available.
- Public problems
- —
- Formats
- —
- Problems with images
- —
- Answer availability
- —
- Configs
- —
- Splits
- —
- Awaiting a first discussion
- —