Pass IndexThe State of AISign in

Berkeley Function-Calling Leaderboard (BFCL) V4

Berkeley Function-Calling Leaderboard (BFCL) V4 is run by UC Berkeley (Gorilla / Shishir Patil et al.). It has ranked 109 models, of which the catalogue holds 25, scoring from 51.4 to 77.47 on Overall Acc.

measured by UC Berkeley (Gorilla / Shishir Patil et al.) · the board itself

PlaceModelMetricScore
1stof 109Claude Opus 4.5Overall Acc (fc)77.47
2ndof 109Claude Sonnet 4.5Overall Acc (fc)73.24
3rdof 109gemini-3-pro-previewOverall Acc (prompt)72.51
4thof 109GLM-4.6Overall Acc (fc thinking)72.38
5thof 109Grok 4.1 Fast ReasoningOverall Acc (fc)69.57
6thof 109Claude Haiku 4.5Overall Acc (fc)68.7
7thof 109gemini-3-pro-previewOverall Acc (fc)68.14
8thof 109o3Overall Acc (prompt)63.05
9thof 109Grok 4Overall Acc (prompt)62.97
10thof 109Grok 4Overall Acc (fc)61.38
11thof 109Kimi K2 0711Overall Acc59.06
12thof 109Grok 4.1 Fast Non-ReasoningOverall Acc (fc)58.29
13thof 109Command AOverall Acc (fc, reasoning)57.06
14thof 109DeepSeek V3.2 ExpOverall Acc (prompt + thinking)56.73
15thof 109Gemini 2.5 FlashOverall Acc (fc)56.24
16thof 109GPT-5.2Overall Acc (fc)55.87
17thof 109GPT-5 miniOverall Acc (fc)55.46
18thof 109xLAM-2-32b-fc-rOverall Acc (fc)54.66
19thof 109DeepSeek V3.2 ExpOverall Acc (fc)54.12
20thof 109GPT-4.1Overall Acc (fc)53.96
21stof 109o4-miniOverall Acc (fc)53.24
22ndof 109Llama-xLAM-2-70b-fc-rOverall Acc (fc)53.07
23rdof 109Qwen3-235B-A22B-Instruct-2507Overall Acc52.15
24thof 109GPT-5 nanoOverall Acc (fc)51.45
25thof 109Nanbeige4-3B-Thinking-2511Overall Acc (fc)51.4

Read from the board on 2026-08-25