benchmarks.wiki / Public workspace
JailbreakBench
Safety
JailbreakBench is an open robustness benchmark for jailbreaking attacks, defining 100 harmful behaviours alongside 100 matched benign ones. This catalogue deliberately introduces and links it rather than reproducing any of its prompts.
jailbreakbench
Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.
Reference links
At a glance
What a problem looks like
Paired harmful and benign behaviour descriptions, used as targets for attack and defence artefacts.
How it is scored
Scored by a judge model against the defined behaviours; nothing is imported here.
Metric: attack success rate, with the benign pairs used to measure over-refusal
Why it is hard
The benign twins are the point: a defence that simply refuses more often scores worse, so the benchmark measures discrimination rather than caution.
Cite
@article{chao2024jailbreakbench,
title={JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models},
author={Patrick Chao and Edoardo Debenedetti and Alexander Robey and Maksym Andriushchenko and Francesco Croce and Vikash Sehwag and Edgar Dobriban and Nicolas Flammarion and George J. Pappas and Florian Tramer and Hamed Hassani and Eric Wong},
journal={arXiv preprint arXiv:2404.01318},
year={2024}
}Discussion
Contribute
Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.
Public problem summary
Catalogue metadata only; public problem statistics are not available.
- Public problems
- —
- Formats
- —
- Problems with images
- —
- Answer availability
- —
- Configs
- —
- Splits
- —
- Awaiting a first discussion
- —
- Problems with a discussion
- —
- Public contributions
- —