Benchmark AI / Public workspace

LongBench v2

Reasoning / Long context

LongBench v2 evaluates deep understanding and reasoning over long contexts through multiple-choice questions. Its official description lists 503 questions spanning tasks such as single-document and multi-document QA and code-repository understanding.

Japanese introduction

長い資料の深い理解と推論を、多肢選択問題で評価するベンチマークです。公式紹介では503問を収録し、単一・複数文書の質問応答やコードリポジトリ理解などを扱います。

longbench-v2

Full text200 tasksProblems imported

Official source

Full text: eligible public problems include text, images and discussion.

Tasks

Whole public benchmark.

Public problems
200
Formats
text (200)
Problems with images
0
Answer availability
Answer published by the source (200)
Configs
Default (200)
Splits
train (200)
Awaiting a first discussion
200

Discussion