Mystery game puzzles — Epoch AI
Mystery game puzzles — Epoch AI is run by Epoch AI. It has ranked 127 models, of which the catalogue holds 64, scoring from 2 to 84.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 127 | GPT 6 Astra | mean_score (max) | 84 |
| 2ndof 127 | Claude Opus 5 | mean_score (max) | 59 |
| 3rdof 127 | Claude Fable 5.1 | mean_score (max) | 58 |
| 4thof 127 | GPT-5.6 Sol | mean_score (max) | 58 |
| 5thof 127 | GPT-5.5 | mean_score (xhigh) | 56 |
| 6thof 127 | Claude Fable 5 | mean_score (max) | 52 |
| 8thof 127 | Gemini 3.8 Flash | mean_score (high) | 47 |
| 9thof 127 | DeepSeek V4 Pro 0813 | mean_score (max) | 43 |
| 10thof 127 | Qwen 3.8 Max | mean_score (xhigh) | 38 |
| 11thof 127 | Gemini 3.7 Flash | mean_score (high) | 37 |
| 13thof 127 | GPT-5.4 | mean_score (xhigh) | 37 |
| 14thof 127 | Claude Opus 4.8 | mean_score (max) | 36 |
| 15thof 127 | Claude Sonnet 5 | mean_score (max) | 35 |
| 16thof 127 | GPT-5.6 Terra | mean_score (max) | 35 |
| 17thof 127 | Grok 4.6 | mean_score (xhigh) | 34 |
| 18thof 127 | DeepSeek V4 Flash (0731) | mean_score (max) | 34 |
| 19thof 127 | Gemini 3.1 Pro Preview | mean_score (high) | 34 |
| 20thof 127 | GLM-5.3 | mean_score (max) | 33 |
| 23rdof 127 | Qwen3.7 Max | mean_score (max) | 32 |
| 24thof 127 | Gemini 3.5 Flash | mean_score (high) | 32 |
| 26thof 127 | Gemini 3.6 Flash | mean_score (high) | 30 |
| 27thof 127 | o3 | mean_score (high) | 29 |
| 32ndof 127 | Claude Opus 4.7 | mean_score (max) | 28 |
| 35thof 127 | Gemini 3 Flash | mean_score (low) | 26 |
| 36thof 127 | Kimi K3 | mean_score (max) | 26 |
| 39thof 127 | Claude Opus 4.6 | mean_score (max) | 25 |
| 41stof 127 | GPT-5.2 | mean_score (high) | 23 |
| 43rdof 127 | GPT-5 | mean_score (high) | 23 |
| 45thof 127 | Qwen3.6 35B A3B | mean_score (none) | 22 |
| 46thof 127 | Claude Opus 4.5 | mean_score | 22 |
| 47thof 127 | GPT-5.6 Luna | mean_score (max) | 21 |
| 48thof 127 | Claude Opus 4.1 | mean_score | 21 |
| 50thof 127 | Qwen3.5-Flash | mean_score | 20 |
| 51stof 127 | Nemotron 3 Ultra | mean_score | 20 |
| 55thof 127 | GPT-5.1 | mean_score (low) | 19 |
| 56thof 127 | GLM 5.2 | mean_score (low) | 19 |
| 57thof 127 | Gemini 3.5 Flash-Lite | mean_score (low) | 19 |
| 58thof 127 | Qwen3.5 397B A17B | mean_score (none) | 18 |
| 59thof 127 | Qwen3.6-Flash | mean_score (none) | 18 |
| 63rdof 127 | Kimi K2.6 | mean_score | 18 |
| 64thof 127 | Qwen 3.7 Plus | mean_score | 17 |
| 66thof 127 | Qwen3.5-122B-A10B | mean_score (none) | 17 |
| 67thof 127 | Qwen 3.5 Plus | mean_score (none) | 17 |
| 70thof 127 | DeepSeek V4 Pro | mean_score (none) | 17 |
| 71stof 127 | Claude Sonnet 4.5 | mean_score | 17 |
| 76thof 127 | Claude Sonnet 4.6 | mean_score (low) | 16 |
| 79thof 127 | Qwen3.7 Flash | mean_score | 15 |
| 92ndof 127 | GPT-4 (0613) | mean_score | 12 |
| 95thof 127 | Qwen3.6 Plus | mean_score (none) | 12 |
| 96thof 127 | GPT-4o-mini (2024-07-18) | mean_score | 12 |
| 98thof 127 | GPT-5.4 mini | mean_score (none) | 11 |
| 101stof 127 | GPT-5 mini | mean_score (minimal) | 10 |
| 102ndof 127 | GPT-5 nano | mean_score (medium) | 9 |
| 103rdof 127 | Qwen3 235B A22B Thinking 2507 | mean_score | 9 |
| 105thof 127 | GPT-5.4 nano | mean_score (none) | 9 |
| 106thof 127 | GLM 5.3 Flash | mean_score (max) | 8 |
| 108thof 127 | MiniMax M3 | mean_score (none) | 8 |
| 111thof 127 | o3-mini | mean_score (high) | 7 |
| 112thof 127 | Qwen3.6 27B | mean_score (none) | 7 |
| 114thof 127 | GPT-4.1 mini | mean_score | 7 |
| 117thof 127 | Inkling Small | mean_score (xhigh) | 6 |
| 118thof 127 | o4-mini | mean_score (high) | 5 |
| 125thof 127 | GPT-3.5 Turbo | mean_score | 3 |
| 126thof 127 | GPT OSS 120B | mean_score (medium) | 2 |
Read from the board on 2026-09-16