Trivia qa — Epoch AI
Trivia qa — Epoch AI is run by Epoch AI. It has ranked 115 models, of which the catalogue holds 21, scoring from 0.452 to 0.876 on EM.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 115 | Llama-2-70b-hf | EM | 0.876 |
| 2ndof 115 | Claude 2.0 | EM | 0.875 |
| 8thof 115 | GPT-3.5 Turbo (1106) | EM | 0.858 |
| 11thof 115 | GPT-4 (0613) | EM | 0.848 |
| 18thof 115 | DeepSeek-V3 | EM | 0.829 |
| 19thof 115 | Llama 3.1 405B | EM | 0.827 |
| 28thof 115 | DeepSeek-V2 | EM | 0.8 |
| 29thof 115 | falcon-40b | EM | 0.799 |
| 30thof 115 | Llama-2-13b | EM | 0.796 |
| 39thof 115 | llama-13b | EM | 0.779 |
| 46thof 115 | Mistral-7B-v0.1 | EM | 0.752 |
| 49thof 115 | Phi-3-medium-128k-instruct | EM | 0.739 |
| 50thof 115 | Llama-2-7b | EM | 0.737 |
| 56thof 115 | gemma-7b | EM | 0.723 |
| 66thof 115 | llama-7b | EM | 0.71 |
| 78thof 115 | Llama 3 8B Instruct | EM | 0.677 |
| 82ndof 115 | falcon-7b | EM | 0.646 |
| 88thof 115 | Phi-3-mini-4k-instruct | EM | 0.64 |
| 99thof 115 | Phi-3-small-8k-instruct | EM | 0.581 |
| 111thof 115 | gemma-2b | EM | 0.532 |
| 114thof 115 | phi-2 | EM | 0.452 |
Read from the board on 2026-09-16