benchmarks.wiki / Public workspace

MuSR

Reasoning

MuSR poses multistep soft reasoning problems over algorithmically generated narratives. It has three domains — murder mysteries, object placements and team allocation — and this catalogue imports the murder-mystery set.

musr

Full text200 tasksProblems imported

Official source

Full text: eligible public problems include text, images and discussion.

At a glance

What a problem looks like

A narrative of several hundred words, a question about it, and the candidate answers as a list. The correct index is not part of the prompt.

Read an actual problem

How it is scored

The source ships the answer index and the chosen option; this catalogue keeps both in restricted storage and never renders them.

Metric: accuracy over the 250 problems in each domain

Why it is hard

The clues are spread through a long narrative and none of them decides the answer alone. The chain has to be assembled from commonsense inferences that the text never states.

Cite

@article{sprague2023musr,
  title={MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning},
  author={Zayne Sprague and Xi Ye and Kaj Bostrom and Swarat Chaudhuri and Greg Durrett},
  journal={arXiv preprint arXiv:2310.16049},
  year={2023}
}

Discussion

Contribute

Share a reproduction, report an evaluation issue, or suggest a correction to this benchmark overview.

Tasks

Whole public benchmark.

Public problems
200
Formats
text (200)
Problems with images
0
Answer availability
Answer published by the source (200)
Configs
default (200)
Splits
murder_mysteries (200)
Awaiting a first discussion
200
Problems with a discussion
0
Public contributions
0

Provenance and terms

Pinned source revisions, newest first:

  • 7c365b439a222150f317764d4f16ae6c96d7d94a · acquired 2026-09-12 · release 2a1b3a07a10a