benchmarks.wiki / Public workspace

WebArena

Agents / Search

WebArena gives an agent self-hosted clones of a shopping site, a forum, a wiki and a code host. It scores whether a natural-language goal was achieved on the live sites.

webarena

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

At a glance

What a problem looks like

A natural-language goal on a self-hosted web application, with a programmatic check of the resulting state.

How it is scored

Tasks run against self-hosted web applications and are graded by the resulting state, so there is no static problem text to republish.

Why it is hard

Long horizons on real interfaces: navigation, forms and search all have to work before the goal state can be reached.

Cite

@article{zhou2023webarena,
  title={WebArena: A Realistic Web Environment for Building Autonomous Agents},
  author={Zhou, Shuyan and Xu, Frank F and Zhu, Hao and Zhou, Xuhui and Lo, Robert and Sridhar, Abishek and Cheng, Xianyi and Bisk, Yonatan and Fried, Daniel and Alon, Uri and others},
  journal={arXiv preprint arXiv:2307.13854},
  year={2023}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion
Problems with a discussion
Public contributions