benchmarks.wiki / Public workspace

JailbreakBench

Safety

JailbreakBench is an open robustness benchmark for jailbreaking attacks, defining 100 harmful behaviours alongside 100 matched benign ones. This catalogue deliberately introduces and links it rather than reproducing any of its prompts.

jailbreakbench

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

At a glance

What a problem looks like

Paired harmful and benign behaviour descriptions, used as targets for attack and defence artefacts.

How it is scored

Scored by a judge model against the defined behaviours; nothing is imported here.

Metric: attack success rate, with the benign pairs used to measure over-refusal

Why it is hard

The benign twins are the point: a defence that simply refuses more often scores worse, so the benchmark measures discrimination rather than caution.

Cite

@article{chao2024jailbreakbench,
  title={JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models},
  author={Patrick Chao and Edoardo Debenedetti and Alexander Robey and Maksym Andriushchenko and Francesco Croce and Vikash Sehwag and Edgar Dobriban and Nicolas Flammarion and George J. Pappas and Florian Tramer and Hamed Hassani and Eric Wong},
  journal={arXiv preprint arXiv:2404.01318},
  year={2024}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion
Problems with a discussion
Public contributions