benchmarks.wiki / Public workspace

Arena-Hard-Auto

Instruction following

Arena-Hard-Auto selects 500 prompts from 200,000 real LMArena user conversations. Answers are scored automatically with a strong model as judge, which is reported to correlate closely with human preference results.

arena-hard

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

At a glance

What a problem looks like

An open-ended user prompt with its topic cluster; answers are compared against a baseline model.

How it is scored

There is no reference answer — a judge model compares two responses, so there is nothing to store as an answer.

Why it is hard

The prompts were chosen for being the ones real users found hard, and they have no single correct answer, so the score depends on judgement quality as much as on the response.

Cite

@article{li2024arenahard,
  title={From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline},
  author={Tianle Li and Wei-Lin Chiang and Evan Frick and Lisa Dunlap and Tianhao Wu and Banghua Zhu and Joseph E. Gonzalez and Ion Stoica},
  journal={arXiv preprint arXiv:2406.11939},
  year={2024}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion
Problems with a discussion
Public contributions