benchmarks.wiki / Public workspace

AgentHarm

Safety / Agents

AgentHarm measures how far a tool-using agent will go in carrying out a harmful request. It pairs 110 harmful behaviours with 110 matched benign ones, so over-refusal is measured alongside compliance.

agentharm

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

At a glance

What a problem looks like

Multi-step agent tasks with a set of synthetic tools, each paired with a benign counterpart.

How it is scored

Scored by a grading function per behaviour against the agent transcript; nothing is imported here.

Metric: harm score and refusal rate, with the benign pairs measuring over-refusal

Why it is hard

A single refusal is not enough: the agent has to keep refusing across a multi-step trajectory, and a jailbreak that survives one step often carries through the rest.

Cite

@article{andriushchenko2024agentharm,
  title={AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents},
  author={Maksym Andriushchenko and Alexandra Souly and Mateusz Dziemian and Derek Duenas and Maxwell Lin and Justin Wang and Dan Hendrycks and Andy Zou and Zico Kolter and Matt Fredrikson and Eric Winsor and Jerome Wynne and Yarin Gal and Xander Davies},
  journal={arXiv preprint arXiv:2410.09024},
  year={2024}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion
Problems with a discussion
Public contributions