benchmarks.wiki / Public workspace
HarmBench
Safety
HarmBench is a standardized framework for comparing automated red-teaming attacks and refusal robustness on the same footing. It pits attack methods against defences and measures the success rate against harmful behaviours.
harmbench
Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.
Reference links
At a glance
What a problem looks like
Harmful behaviour descriptions. The mirror linked here carries the 400 textual behaviours in the standard, contextual and copyright categories; the full benchmark adds 110 multimodal behaviours that the mirror does not.
How it is scored
Scored by a trained classifier deciding whether a completion actually carries out the behaviour; nothing is imported here.
Metric: attack success rate (ASR) per attack and defence pair
Why it is hard
It separates a model that refuses from a model that merely sounds like it refused: the classifier judges the completion, so a compliant answer wrapped in a disclaimer still counts as a success for the attack.
Cite
@article{mazeika2024harmbench,
title={HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal},
author={Mantas Mazeika and Long Phan and Xuwang Yin and Andy Zou and Zifan Wang and Norman Mu and Elham Sakhaee and Nathaniel Li and Steven Basart and Bo Li and David Forsyth and Dan Hendrycks},
journal={arXiv preprint arXiv:2402.04249},
year={2024}
}Discussion
Contribute
Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.
Public problem summary
Catalogue metadata only; public problem statistics are not available.
- Public problems
- —
- Formats
- —
- Problems with images
- —
- Answer availability
- —
- Configs
- —
- Splits
- —
- Awaiting a first discussion
- —
- Problems with a discussion
- —
- Public contributions
- —