Berkeley Function-Calling Leaderboard (BFCL) V4
Berkeley Function-Calling Leaderboard (BFCL) V4 is run by UC Berkeley (Gorilla / Shishir Patil et al.). It has ranked 109 models, of which the catalogue holds 25, scoring from 51.4 to 77.47 on Overall Acc.
measured by UC Berkeley (Gorilla / Shishir Patil et al.) · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 109 | Claude Opus 4.5 | Overall Acc (fc) | 77.47 |
| 2ndof 109 | Claude Sonnet 4.5 | Overall Acc (fc) | 73.24 |
| 3rdof 109 | gemini-3-pro-preview | Overall Acc (prompt) | 72.51 |
| 4thof 109 | GLM-4.6 | Overall Acc (fc thinking) | 72.38 |
| 5thof 109 | Grok 4.1 Fast Reasoning | Overall Acc (fc) | 69.57 |
| 6thof 109 | Claude Haiku 4.5 | Overall Acc (fc) | 68.7 |
| 7thof 109 | gemini-3-pro-preview | Overall Acc (fc) | 68.14 |
| 8thof 109 | o3 | Overall Acc (prompt) | 63.05 |
| 9thof 109 | Grok 4 | Overall Acc (prompt) | 62.97 |
| 10thof 109 | Grok 4 | Overall Acc (fc) | 61.38 |
| 11thof 109 | Kimi K2 0711 | Overall Acc | 59.06 |
| 12thof 109 | Grok 4.1 Fast Non-Reasoning | Overall Acc (fc) | 58.29 |
| 13thof 109 | Command A | Overall Acc (fc, reasoning) | 57.06 |
| 14thof 109 | DeepSeek V3.2 Exp | Overall Acc (prompt + thinking) | 56.73 |
| 15thof 109 | Gemini 2.5 Flash | Overall Acc (fc) | 56.24 |
| 16thof 109 | GPT-5.2 | Overall Acc (fc) | 55.87 |
| 17thof 109 | GPT-5 mini | Overall Acc (fc) | 55.46 |
| 18thof 109 | xLAM-2-32b-fc-r | Overall Acc (fc) | 54.66 |
| 19thof 109 | DeepSeek V3.2 Exp | Overall Acc (fc) | 54.12 |
| 20thof 109 | GPT-4.1 | Overall Acc (fc) | 53.96 |
| 21stof 109 | o4-mini | Overall Acc (fc) | 53.24 |
| 22ndof 109 | Llama-xLAM-2-70b-fc-r | Overall Acc (fc) | 53.07 |
| 23rdof 109 | Qwen3-235B-A22B-Instruct-2507 | Overall Acc | 52.15 |
| 24thof 109 | GPT-5 nano | Overall Acc (fc) | 51.45 |
| 25thof 109 | Nanbeige4-3B-Thinking-2511 | Overall Acc (fc) | 51.4 |
Read from the board on 2026-08-25