Benchmark AI / Public workspace

Terminal-Bench 2.1

Agents / Code

Terminal-Bench 2.1 evaluates agents performing tasks in terminal environments. Each task supplies instructions and environment configuration, and version 2.1 is tracked separately from 2.0.

Japanese introduction

ターミナル環境で作業を遂行するエージェントの能力を評価するベンチマークです。各課題に作業指示と環境設定があり、2.0とは別の版として扱います。

terminal-bench-2-1

Full text89 tasksProblems imported

Official source

Full text: eligible public problems include text, images and discussion.

Tasks

Whole public benchmark.

Public problems
89
Formats
text (89)
Problems with images
0
Answer availability
Answer availability not recorded (89)
Configs
Default (89)
Splits
tasks (89)
Awaiting a first discussion
89

Discussion