Terminal and software agent benchmark comparison

How do task origin, interaction interface and success criteria differ?

JSON · CSV · BibTeX

WorkTask sourceInteraction interfaceSuccess criteriaEvaluation scope
InterCodeNL2Bash and Spider tasks grounded in executable environments3.3 Implementations · 3.3 ImplementationsBash and SQL actions with execution observations3.3 Implementations · 3.3 Implementations · 3.3 ImplementationsExecution output and filesystem changes, or ordered database results3.3 Implementations · 3.3 Implementations200 Bash tasks and 1,034 SQL tasks in the cited version3.3 Implementations · 3.3 Implementations
SWE-benchRepository issues paired with resolving pull requests2 SWE-benchIssue description and codebase as input; code patch as output2.2 Task FormulationRepository tests associated with the issue2 SWE-benchSoftware issue resolution2 SWE-bench
Terminal-Bench 2.0Community-authored tasks selected through human review2.2 Dataset ConstructionContainer interaction through agent tools; Terminus 2 uses Bash3.1 Terminus 2 · 3.1 Terminus 2Tests of the final container state against task requirements2.1 Task Formulation89 selected terminal tasks, distributed through Harbor2.2 Dataset Construction · 3.4 Harbor
ShellOps and ShellOps-ProShell-driven tasks with initial and optional target file trees4.1 Dataset Usage and Construction · 4.1 Dataset Usage and ConstructionShell actions in filesystem workspaces3 MethodProgrammatic rewards over executed outputs and workspace state4.1 Dataset Usage and Construction1,624 standard tasks and 150 harder out-of-distribution tasks4.1 Dataset Usage and Construction
TerminalWorldTasks reconstructed from recorded terminal workflows3.3 Reproducing Executable EnvironmentsReproduced Docker environments, including multi-service setups3.3 Reproducing Executable EnvironmentsFinal-state tests checked with complete, empty and partial solutions3.4 Generating Test Suites · 3.4 Generating Test Suites1,530 validated tasks with a separately reviewed Verified subset3.4 Generating Test Suites · 4.2 The Verified Subset
TUA-BenchEveryday tasks adapted from OSWorld and expert-designed scientific workflows3.2.1 Everyday Digital Tasks · 3.2.2 Professional Scientific TasksTerminal interaction in resettable Linux containers using Harbor3.1 Task Execution Environment · 3.1 Task Execution EnvironmentTask-specific verifiers; professional-task design permits programmatic or model judging3.1 Task Execution Environment · 3.2.2 Professional Scientific Tasks120 tasks across five general-purpose computer-use families3.2.3 Task Statistics and Features

Qualitative comparisons of methods, information requirements and task design in the cited paper versions.

References

  1. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. Yang, John; Prabhakar, Akshara; Narasimhan, Karthik; Yao, Shunyu. 2306.14898v1.
  2. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. Jimenez, Carlos E.; Yang, John; Wettig, Alexander; Yao, Shunyu; Pei, Kexin; Press, Ofir; Narasimhan, Karthik. 2310.06770v1.
  3. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces. Merrill, Mike A.; Shaw, Alexander G.; Carlini, Nicholas; Li, Boxuan; Raj, Harsh; Bercovich, Ivan; Shi, Lin; Shin, Jeong Yeon; Walshe, Thomas; Buchanan, E. Kelly; Shen, Junhong; Ye, Guanghao; Lin, Haowei; Poulos, Jason; Wang, Maoyu; Nezhurina, Marianna; Jitsev, Jenia; Lu, Di; Mastromichalakis, Orfeas Menis; Xu, Zhiwei; Chen, Zizhao; Liu, Yue; Zhang, Robert; Chen, Leon Liangyu; Kashyap, Anurag; Uslu, Jan-Lucas; Li, Jeffrey; Wu, Jianbo; Yan, Minghao; Bian, Song; Sharma, Vedang; Sun, Ke; Dillmann, Steven; Anand, Akshay; Lanpouthakoun, Andrew; Koopah, Bardia; Hu, Changran; Guha, Etash; Dreiman, Gabriel H. S.; Zhu, Jiacheng; Krauth, Karl; Zhong, Li; Muennighoff, Niklas; Amanfu, Robert; Tan, Shangyin; Pimpalgaonkar, Shreyas; Aggarwal, Tushar; Lin, Xiangning; Lan, Xin; Zhao, Xuandong; Liang, Yiqing; Wang, Yuanli; Wang, Zilong; Zhou, Changzhi; Heineman, David; Liu, Hange; Trivedi, Harsh; Yang, John; Lin, Junhong; Shetty, Manish; Yang, Michael; Omi, Nabil; Raoof, Negin; Li, Shanda; Zhuo, Terry Yue; Lin, Wuwei; Dai, Yiwei; Wang, Yuxin; Chai, Wenhao; Zhou, Shang; Wahdany, Dariush; She, Ziyu; Hu, Jiaming; Dong, Zhikang; Zhu, Yuxuan; Cui, Sasha; Saiyed, Ahson; Kolbeinsson, Arinbjörn; Hu, Jesse; Rytting, Christopher Michael; Marten, Ryan; Wang, Yixin; Dimakis, Alex; Konwinski, Andy; Schmidt, Ludwig. 2601.11868v1.
  4. Learning CLI Agents with Structured Action Credit under Selective Observation. Su, Haoyang; Wen, Ying. 2605.08013v1.
  5. TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks. Chu, Zhaoyang; Hu, Jiarui; Jiang, Xingyu; Zou, Pengyu; Li, Han; Peng, Chao; O'Hearn, Peter; Barr, Earl T.; Harman, Mark; Sarro, Federica; Ye, He. 2605.22535v1.
  6. TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents. Chen, Shoufa; Wang, Luyuan; Yang, Xuan; Liu, Zhiheng; Cong, Yuren; Ji, Yuanfeng; Zhou, Feiyan; Zhang, Xiaohui; Yang, Fanny; Zeng, Belinda. 2606.28480v1.