Skip to main content
We just open sourced a tiny GPT-style cognitive core built in pure Rust.See our repository

Terminal-Bench 2.1

One harness to unlock the potential of all models

Compare Ante runs across models on the same Terminal-Bench 2.1 task set, using consistent parameters and verified benchmark results.

Best Accuracy82.7%
Leading Model
DeepSeek V4 Flash 0731max
Task Set89 tasks
Trials368 passed / 445 trials
Benchmark principlesRead more
We benchmark what we ship.Every eval uses a pinned public . No eval-only branches or benchmark-specific prompts.
The runs are auditable.Every result links its raw Harbor run, so anyone can inspect the trials behind the number.
The constraints are official.All trials follow the official Terminal-Bench parameters: 89 tasks, 5 trials per task, strict timeouts, and hardware limits.
Model org
All model orgsTB 2.1 · Same parameters: 89 tasks · 5 trials/task · Updated Aug 9, 2026
#ModelSame-modelAgentSource

Terminal-Bench Reference

Verified Public Leaderboard

Same parameters: 89 tasks · 5 trials/task · Updated Jul 17, 2026 · Source: Terminal-Bench official verified rows

For how different models perform on TB 2.1, see Vals AI's Terminal-Bench 2.1 benchmark.

17 official rows
#AgentModelAccuracyRun date