benchmarks.wiki / Public workspace
MT-Bench
Instruction following
MT-Bench is a set of 80 two-turn questions across eight categories, published with the study that established LLM-as-a-judge evaluation. It looks at multi-turn consistency and instruction following.
mt-bench
Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.
Reference links
At a glance
What a problem looks like
Two successive user turns in one category, the second depending on the answer to the first.
How it is scored
Scored by a judge model on a 1–10 scale; there is no reference answer to store for most categories.
Metric: average judge score over the eight categories, reported per turn
Why it is hard
The second turn depends on the first, so an answer that was locally fine but unhelpful for what follows is penalised where a single-turn benchmark would not notice.
Cite
@article{zheng2023mtbench,
title={Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena},
author={Lianmin Zheng and Wei-Lin Chiang and Ying Sheng and Siyuan Zhuang and Zhanghao Wu and Yonghao Zhuang and Zi Lin and Zhuohan Li and Dacheng Li and Eric P. Xing and Hao Zhang and Joseph E. Gonzalez and Ion Stoica},
journal={arXiv preprint arXiv:2306.05685},
year={2023}
}Discussion
Contribute
Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.
Public problem summary
Catalogue metadata only; public problem statistics are not available.
- Public problems
- —
- Formats
- —
- Problems with images
- —
- Answer availability
- —
- Configs
- —
- Splits
- —
- Awaiting a first discussion
- —
- Problems with a discussion
- —
- Public contributions
- —