benchmarks.wiki / Public workspace

RULER

Reasoning / Long context

RULER is a synthetic benchmark for the context length a model can actually use. It generates retrieval, multi-hop tracing, aggregation and question-answering tasks at any length, exposing the gap against the advertised window.

ruler

Overview onlyNot imported on this site

Official source

Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.

At a glance

What a problem looks like

Synthetic long documents with planted information to retrieve, trace, or aggregate.

How it is scored

Generated with its own answer key; nothing is imported here.

Metric: accuracy per task at each context length, and the effective length at which performance falls below a threshold

Why it is hard

It separates the advertised context window from the usable one. Most models degrade well before the limit they claim, and the point of failure depends on the task, not just the length.

Cite

@article{hsieh2024ruler,
  title={RULER: What's the Real Context Size of Your Long-Context Language Models?},
  author={Cheng-Ping Hsieh and Simeng Sun and Samuel Kriman and Shantanu Acharya and Dima Rekesh and Fei Jia and Yang Zhang and Boris Ginsburg},
  journal={arXiv preprint arXiv:2404.06654},
  year={2024}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Public problem summary

Catalogue metadata only; public problem statistics are not available.

Public problems
Formats
Problems with images
Answer availability
Configs
Splits
Awaiting a first discussion
Problems with a discussion
Public contributions