benchmarks.wiki / Public workspace

HarmBench

Safety

HarmBench is a standardized framework for comparing automated red-teaming attacks and refusal robustness on the same footing. It pits attack methods against defences and measures the success rate against harmful behaviours.

harmbench

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

At a glance

What a problem looks like

Harmful behaviour descriptions. The mirror linked here carries the 400 textual behaviours in the standard, contextual and copyright categories; the full benchmark adds 110 multimodal behaviours that the mirror does not.

How it is scored

Scored by a trained classifier deciding whether a completion actually carries out the behaviour; nothing is imported here.

Metric: attack success rate (ASR) per attack and defence pair

Why it is hard

It separates a model that refuses from a model that merely sounds like it refused: the classifier judges the completion, so a compliant answer wrapped in a disclaimer still counts as a success for the attack.

Cite

@article{mazeika2024harmbench,
  title={HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal},
  author={Mantas Mazeika and Long Phan and Xuwang Yin and Andy Zou and Zifan Wang and Norman Mu and Elham Sakhaee and Nathaniel Li and Steven Basart and Bo Li and David Forsyth and Dan Hendrycks},
  journal={arXiv preprint arXiv:2402.04249},
  year={2024}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion
Problems with a discussion
Public contributions