Benchmark AI / Public workspace

Humanity's Last Exam

Reasoning / Multimodal

Humanity's Last Exam evaluates AI knowledge and reasoning across academic disciplines. It includes text and image questions; this catalog provides an introduction and official links in accordance with the publisher's request.

Japanese introduction

幅広い学問分野の専門的な問題を通じて、AIの知識と推論能力を評価するベンチマークです。テキストと画像を含む問題を扱い、このカタログでは公開者の要請に従って紹介と公式リンクを提供します。

hle

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion

Discussion