benchmarks.wiki / Public workspace
Arena-Hard-Auto
Instruction following
Arena-Hard-Auto selects 500 prompts from 200,000 real LMArena user conversations. Answers are scored automatically with a strong model as judge, which is reported to correlate closely with human preference results.
arena-hard
Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.
Reference links
At a glance
What a problem looks like
An open-ended user prompt with its topic cluster; answers are compared against a baseline model.
How it is scored
There is no reference answer — a judge model compares two responses, so there is nothing to store as an answer.
Why it is hard
The prompts were chosen for being the ones real users found hard, and they have no single correct answer, so the score depends on judgement quality as much as on the response.
Cite
@article{li2024arenahard,
title={From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline},
author={Tianle Li and Wei-Lin Chiang and Evan Frick and Lisa Dunlap and Tianhao Wu and Banghua Zhu and Joseph E. Gonzalez and Ion Stoica},
journal={arXiv preprint arXiv:2406.11939},
year={2024}
}Discussion
Contribute
Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.
Public problem summary
Catalogue metadata only; public problem statistics are not available.
- Public problems
- —
- Formats
- —
- Problems with images
- —
- Answer availability
- —
- Configs
- —
- Splits
- —
- Awaiting a first discussion
- —
- Problems with a discussion
- —
- Public contributions
- —