benchmarks.wiki / Public workspace
AgentHarm
Safety / Agents
AgentHarm measures how far a tool-using agent will go in carrying out a harmful request. It pairs 110 harmful behaviours with 110 matched benign ones, so over-refusal is measured alongside compliance.
agentharm
Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.
Reference links
At a glance
What a problem looks like
Multi-step agent tasks with a set of synthetic tools, each paired with a benign counterpart.
How it is scored
Scored by a grading function per behaviour against the agent transcript; nothing is imported here.
Metric: harm score and refusal rate, with the benign pairs measuring over-refusal
Why it is hard
A single refusal is not enough: the agent has to keep refusing across a multi-step trajectory, and a jailbreak that survives one step often carries through the rest.
Cite
@article{andriushchenko2024agentharm,
title={AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents},
author={Maksym Andriushchenko and Alexandra Souly and Mateusz Dziemian and Derek Duenas and Maxwell Lin and Justin Wang and Dan Hendrycks and Andy Zou and Zico Kolter and Matt Fredrikson and Eric Winsor and Jerome Wynne and Yarin Gal and Xander Davies},
journal={arXiv preprint arXiv:2410.09024},
year={2024}
}Discussion
Contribute
Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.
Public problem summary
Catalogue metadata only; public problem statistics are not available.
- Public problems
- —
- Formats
- —
- Problems with images
- —
- Answer availability
- —
- Configs
- —
- Splits
- —
- Awaiting a first discussion
- —
- Problems with a discussion
- —
- Public contributions
- —