Pass IndexThe State of AISign in

terminal-bench@2.1

terminal-bench@2.1 is run by Terminal-Bench. It has ranked 17 models, of which the catalogue holds 13, scoring from 58.7 to 83.8 on Accuracy.

measured by Terminal-Bench · the board itself

PlaceModelMetricScore
1stof 17Claude Fable 5Accuracy (xhigh)83.8
2ndof 17GPT-5.5Accuracy (xhigh)83.1
3rdof 17Claude Fable 5Accuracy (high)80.4
4thof 17Grok 4.5Accuracy (high)79.3
5thof 17Claude Opus 4.8Accuracy (high)78.9
6thof 17GPT-5.6 TerraAccuracy (max)78.4
8thof 17Muse Spark 1.1Accuracy (xhigh)76.2
9thof 17GPT-5.6 LunaAccuracy (max)75.7
10thof 17Claude Sonnet 5Accuracy (high)74.6
11thof 17gemini-3-pro-previewAccuracy73.9
12thof 17Claude Opus 4.7Accuracy68.9
14thof 17Gemini 3.1 Pro PreviewAccuracy65.8
17thof 17GLM 5.1Accuracy58.7

Read from the board on 2026-08-25