benchmarks.wiki / Public workspace
WebArena
Agents / Search
WebArena gives an agent self-hosted clones of a shopping site, a forum, a wiki and a code host. It scores whether a natural-language goal was achieved on the live sites.
webarena
Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.
Reference links
At a glance
What a problem looks like
A natural-language goal on a self-hosted web application, with a programmatic check of the resulting state.
How it is scored
Tasks run against self-hosted web applications and are graded by the resulting state, so there is no static problem text to republish.
Why it is hard
Long horizons on real interfaces: navigation, forms and search all have to work before the goal state can be reached.
Cite
@article{zhou2023webarena,
title={WebArena: A Realistic Web Environment for Building Autonomous Agents},
author={Zhou, Shuyan and Xu, Frank F and Zhu, Hao and Zhou, Xuhui and Lo, Robert and Sridhar, Abishek and Cheng, Xianyi and Bisk, Yonatan and Fried, Daniel and Alon, Uri and others},
journal={arXiv preprint arXiv:2307.13854},
year={2023}
}Discussion
Contribute
Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.
Public problem summary
Catalogue metadata only; public problem statistics are not available.
- Public problems
- —
- Formats
- —
- Problems with images
- —
- Answer availability
- —
- Configs
- —
- Splits
- —
- Awaiting a first discussion
- —
- Problems with a discussion
- —
- Public contributions
- —