Benchmark AI / Public workspace

BrowseComp

Agents / Search

BrowseComp evaluates web browsing and the ability to identify answers from multiple clues. This catalog supports introductions and discussion of search methods without republishing decrypted questions or answers.

Japanese introduction

ウェブ上の情報を探し、複数の手掛かりから答えを特定する能力を評価するベンチマークです。問題と答えの平文再公開を初期対象とせず、紹介と検索手法の議論を提供します。

browsecomp

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion

Discussion