benchmarks.wiki / Public workspace

IFEval

Reasoning

IFEval measures whether a model follows verifiable instructions -- word counts, formats, forbidden words -- rather than whether an answer is good. Every constraint is checked programmatically.

ifeval

Full text200 tasksProblems imported

Official source

Full text: eligible public problems include text, images and discussion.

At a glance

What a problem looks like

A natural-language prompt carrying one or more machine-checkable constraints, with the constraint identifiers recorded alongside it.

Read an actual problem

How it is scored

For the 200 imported problems, every column outside the prompt is separated into restricted storage, but the converter does not interpret it, so this catalogue does not claim a usable answer key exists.

Why it is hard

The constraints are easy to state and easy to violate in passing; models routinely answer well while breaking the format they were told to use.

Cite

@misc{zhou2023instructionfollowingevaluationlargelanguage,
      title={Instruction-Following Evaluation for Large Language Models}, 
      author={Jeffrey Zhou and Tianjian Lu and Swaroop Mishra and Siddhartha Brahma and Sujoy Basu and Yi Luan and Denny Zhou and Le Hou},
      year={2023},
      eprint={2311.07911},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2311.07911}, 
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Tasks

Whole public benchmark.

Public problems
200
Formats
text (200)
Problems with images
0
Answer availability
Answer availability not recorded (200)
Configs
default (200)
Splits
train (200)
Awaiting a first discussion
200
Problems with a discussion
0
Public contributions
0

Provenance and terms

Pinned source revisions, newest first:

  • 966cd89545d6b6acfd7638bc708b98261ca58e84 · acquired 2026-09-11 · release 54f3548aa995