benchmarks.wiki / Public workspace
RULER
Reasoning / Long context
RULER is a synthetic benchmark for the context length a model can actually use. It generates retrieval, multi-hop tracing, aggregation and question-answering tasks at any length, exposing the gap against the advertised window.
ruler
Overview only: problem text has not been published on this site; the catalogue records the benchmark introduction, category and official source.
Reference links
At a glance
What a problem looks like
Synthetic long documents with planted information to retrieve, trace, or aggregate.
How it is scored
Generated with its own answer key; nothing is imported here.
Metric: accuracy per task at each context length, and the effective length at which performance falls below a threshold
Why it is hard
It separates the advertised context window from the usable one. Most models degrade well before the limit they claim, and the point of failure depends on the task, not just the length.
Cite
@article{hsieh2024ruler,
title={RULER: What's the Real Context Size of Your Long-Context Language Models?},
author={Cheng-Ping Hsieh and Simeng Sun and Samuel Kriman and Shantanu Acharya and Dima Rekesh and Fei Jia and Yang Zhang and Boris Ginsburg},
journal={arXiv preprint arXiv:2404.06654},
year={2024}
}Discussion
Contribute
Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.
Public problem summary
Catalogue metadata only; public problem statistics are not available.
- Public problems
- —
- Formats
- —
- Problems with images
- —
- Answer availability
- —
- Configs
- —
- Splits
- —
- Awaiting a first discussion
- —
- Problems with a discussion
- —
- Public contributions
- —