benchmarks.wiki / Public workspace

AssistantBench

Agents / Search

AssistantBench collects realistic, time-consuming tasks that can only be answered by working through the live web, written so that each has one determinable answer. It measures whether a web agent is actually useful.

assistantbench

Full text181 tasksProblems imported

Official source

Full text: eligible public problems include text, images and discussion.

At a glance

What a problem looks like

One task statement. The evaluation subset, the answer, the gold source URL and the explanation are not part of the prompt.

Read an actual problem

How it is scored

The test split ships no answer at all — the source withholds it — so this catalogue has nothing to store or render for it.

Metric: accuracy and answer rate over the held-out test set, scored by the official submission service

Why it is hard

The answer is not on any single page. It has to be assembled across sites, and the search itself is most of the work.

Cite

@article{yoran2024assistantbench,
  title={AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?},
  author={Ori Yoran and Samuel Joseph Amouyal and Chaitanya Malaviya and Ben Bogin and Ofir Press and Jonathan Berant},
  journal={arXiv preprint arXiv:2407.15711},
  year={2024}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Tasks

Whole public benchmark.

Public problems
181
Formats
text (181)
Problems with images
0
Answer availability
Answer withheld by the source (181)
Configs
default (181)
Splits
test (181)
Awaiting a first discussion
181
Problems with a discussion
0
Public contributions
0

Provenance and terms

Pinned source revisions, newest first:

  • 482cbbc0400f6d048438c4021727f21a10cbff49 · acquired 2026-09-12 · release 75ac7aefce1d
  • 482cbbc0400f6d048438c4021727f21a10cbff49 · acquired 2026-09-12 · release 02c88c74db5a