benchmarks.wiki / Public workspace

MT-Bench

Instruction following

MT-Bench is a set of 80 two-turn questions across eight categories, published with the study that established LLM-as-a-judge evaluation. It looks at multi-turn consistency and instruction following.

mt-bench

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

At a glance

What a problem looks like

Two successive user turns in one category, the second depending on the answer to the first.

How it is scored

Scored by a judge model on a 1–10 scale; there is no reference answer to store for most categories.

Metric: average judge score over the eight categories, reported per turn

Why it is hard

The second turn depends on the first, so an answer that was locally fine but unhelpful for what follows is penalised where a single-turn benchmark would not notice.

Cite

@article{zheng2023mtbench,
  title={Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena},
  author={Lianmin Zheng and Wei-Lin Chiang and Ying Sheng and Siyuan Zhuang and Zhanghao Wu and Yonghao Zhuang and Zi Lin and Zhuohan Li and Dacheng Li and Eric P. Xing and Hao Zhang and Joseph E. Gonzalez and Ion Stoica},
  journal={arXiv preprint arXiv:2306.05685},
  year={2023}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion
Problems with a discussion
Public contributions