Bbh — Epoch AI
Bbh — Epoch AI is run by Epoch AI. It has ranked 88 models, of which the catalogue holds 31, scoring from 0.288 to 0.875 on Average.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 2ndof 88 | DeepSeek-V3 | Average | 0.875 |
| 4thof 88 | Llama 3.1 405B | Average | 0.829 |
| 7thof 88 | Phi-3-medium-128k-instruct | Average | 0.814 |
| 10thof 88 | Phi-3-small-8k-instruct | Average | 0.791 |
| 11thof 88 | DeepSeek-V2 | Average | 0.788 |
| 13thof 88 | GPT-4 (0613) | Average | 0.751 |
| 14thof 88 | Phi-3-mini-4k-instruct | Average | 0.717 |
| 15thof 88 | Yi-34B-Chat | Average | 0.717 |
| 16thof 88 | StableBeluga2 | Average | 0.693 |
| 17thof 88 | Llama-2-70b-hf | Average | 0.649 |
| 18thof 88 | GPT-3.5 Turbo (older v0613) | Average | 0.616 |
| 19thof 88 | phi-2 | Average | 0.594 |
| 21stof 88 | Llama-2-70b-chat | Average | 0.585 |
| 23rdof 88 | Llama-2-13b-chat | Average | 0.582 |
| 24thof 88 | Mistral-7B-v0.1 | Average | 0.561 |
| 25thof 88 | gemma-7b | Average | 0.551 |
| 27thof 88 | Qwen-14B-Chat | Average | 0.55 |
| 28thof 88 | Yi-34B | Average | 0.543 |
| 29thof 88 | Qwen-14B | Average | 0.534 |
| 40thof 88 | Baichuan2-13B-Chat | Average | 0.472 |
| 42ndof 88 | Llama-2-13b | Average | 0.47 |
| 43rdof 88 | Mistral 7B | Average | 0.45 |
| 50thof 88 | Baichuan-13B-Base | Average | 0.43 |
| 51stof 88 | Yi-6B | Average | 0.428 |
| 62ndof 88 | Llama-2-7b | Average | 0.392 |
| 65thof 88 | llama-13b | Average | 0.379 |
| 67thof 88 | falcon-40b | Average | 0.371 |
| 72ndof 88 | gemma-2b | Average | 0.352 |
| 77thof 88 | llama-7b | Average | 0.335 |
| 81stof 88 | Baichuan-7B | Average | 0.325 |
| 85thof 88 | falcon-7b | Average | 0.288 |
Read from the board on 2026-09-16