This may take a moment. We're preparing the latest results.
This may take a moment. We're preparing the latest results.
Terminal-Bench is a collection of tasks and an evaluation harness to help agent makers quantify their agents' terminal mastery. This is the Terminal-Bench 2.1 dataset (89 verified tasks) run on the Harbor harness (terminal-bench/terminal-bench-2-1, agent: terminus-2).
Compare Models against this Benchmark