Pass IndexThe State of AISign in

Gsm8k — Epoch AI

Gsm8k — Epoch AI is run by Epoch AI. It has ranked 170 models, of which the catalogue holds 48, scoring from 0.018 to 0.945 on EM.

measured by Epoch AI · the board itself

PlaceModelMetricScore
1stof 170DeepSeek-Coder-V2-InstructEM0.945
2ndof 170Qwen2.5-Coder-14B-InstructEM0.942
3rdof 170Qwen2.5 Coder 32B InstructEM0.93
5thof 170GPT-4o-mini (2024-07-18)EM0.913
6thof 170Qwen 2.5 Coder 32BEM0.911
7thof 170GPT-4 (0613)EM0.9
8thof 170Phi-3.5-MoE-instructEM0.887
10thof 170DeepSeek-Coder-V2-Lite-InstructEM0.876
11thof 170Qwen2.5-Coder-7B-InstructEM0.867
13thof 170Phi-3.5-mini-instructEM0.862
15thof 170gemma-2-9bEM0.849
17thof 170Qwen2.5-Coder-7BEM0.839
18thof 170Llama 3.1 8B InstructEM0.824
21stof 170Qwen2.5-Coder-3B-InstructEM0.807
23rdof 170Yi-34B-ChatEM0.76
27thof 170Llama-2-70b-hfEM0.696
28thof 170StableBeluga2EM0.696
29thof 170Yi-34BEM0.672
34thof 170Qwen-14BEM0.613
36thof 170Qwen-14B-ChatEM0.612
39thof 170Llama-2-70b-chatEM0.587
40thof 170GPT-3.5 Turbo (older v0613)EM0.578
41stof 170starcoder2-15bEM0.577
48thof 170Mistral-7B-v0.1EM0.544
49thof 170falcon-180BEM0.544
51stof 170falcon-11BEM0.538
56thof 170Mistral 7BEM0.517
68thof 170gemma-7bEM0.464
71stof 170Baichuan2-13B-ChatEM0.457
85thof 170Llama-2-13b-chatEM0.369
88thof 170Mistral-7B-Instruct-v0.2EM0.354
91stof 170Qwen2.5-Coder-0.5BEM0.345
93rdof 170Llama-2-13bEM0.343
94thof 170falcon-40b-instructEM0.338
96thof 170starcoder2-7bEM0.327
97thof 170Yi-6BEM0.325
109thof 170Baichuan-13B-BaseEM0.268
110thof 170falcon-40bEM0.25
116thof 170vicuna-13b-v1.3EM0.226
118thof 170starcoder2-3bEM0.216
122ndof 170llama-13bEM0.205
128thof 170gemma-2bEM0.177
129thof 170Llama-2-7bEM0.167
142ndof 170llama-7bEM0.11
147thof 170BloomEM0.095
148thof 170Baichuan-7BEM0.092
154thof 170falcon-7bEM0.068
164thof 170opt-66bEM0.018

Read from the board on 2026-09-16