{"kind":"task","effective_mode":"full","benchmark":{"kind":"benchmark","effective_mode":"full","slug":"terminal-bench-2-1","formal_name":"Terminal-Bench 2.1","introduction":"ターミナル環境で作業を遂行するエージェントの能力を評価するベンチマークです。各課題に作業指示と環境設定があり、2.0とは別の版として扱います。\n\nTerminal-Bench 2.1 evaluates agents performing tasks in terminal environments. Each task supplies instructions and environment configuration, and version 2.1 is tracked separately from 2.0.","introduction_ja":"","introduction_en":"","category":"Category not supplied","task_count":null,"acquisition_status":"Acquisition status not supplied","official_url":"https://github.com/harbor-framework/terminal-bench-2-1","indexing_mode":"noindex"},"task_id":"fcb1a30e-ca91-5fce-8ee0-35317df57b2f","task_key":"tasks--vulnerable~2dsecret","task_revision_id":"1","upstream_id":"vulnerable-secret","short_description":"Analyze the following program that contains a secret key. The goal is to extract…","config":"","split":"tasks","body":"{\"instruction\":\"Analyze the following program that contains a secret key. The goal is to extract this secret key and save it to `/app/results.txt`.\\n\\nThe program is available as an executable. You may need to interact with it in various ways to discover how to access the secret information.\\n\\nThe secret key is in the format `FLAG{...}`\\n\\nThe task is complete when you have successfully extracted the secret key and written it to `/app/results.txt`.\\n\"}","display_format":"text","language":"","answer_status":"unknown","assets":[],"source_url":"https://github.com/harbor-framework/terminal-bench-2-1","history":"initial import","indexing_mode":"noindex","subproblems":[],"grids":[]}