Benchmarks / Berkeley Function Calling Leaderboard (BFCL) V4
Reported by Berkeley Function Calling Leaderboard (BFCL) V4
Berkeley Function Calling Leaderboard (BFCL) V4
The same task on functions and questions contributed by real users.
- Results dated
- 16 Dec 2025
- Models
- 109
- Unit
- % correct
- Licence
- Apache License 2.0
| # | Model | BFCL: single-turn calls (user-contributed) % correct, higher is better |
|---|---|---|
| 1 | bitagent-bounty-8bBittensor |
93.1%
|
| 2 | gemini-3-pro-previewGoogle |
83.1%
|
| 3 | Qwen3 32BAlibaba · qwen3-32b |
82.0%
|
| 3 | Qwen3 32BAlibaba · qwen3-32b |
82.0%
|
| 5 | mistral-large-2411Mistral |
81.9%
|
| 6 | gemini-3-pro-previewGoogle |
81.7%
|
| 7 | Claude Sonnet 4.5Anthropic · claude-sonnet-4.5 |
81.1%
|
| 8 | GLM 4.6Z.ai · glm-4.6 |
80.9%
|
| 9 | Nova 2 LiteAmazon · nova-2-lite-v1 |
80.8%
|
| 10 | arch-agent-32bKatanemo |
80.7%
|
| 11 | Qwen3 8BAlibaba · qwen3-8b |
80.5%
|
| 12 | Qwen3 8BAlibaba · qwen3-8b |
80.1%
|
| 13 | Qwen3 14BAlibaba · qwen3-14b |
80.0%
|
| 14 | Claude Opus 4.5Anthropic · claude-opus-4.5 |
79.8%
|
| 15 | nanbeige4-3b-thinking-2511Nanbeige |
79.4%
|
| 16 | Qwen3 14BAlibaba · qwen3-14b |
79.3%
|
| 17 | Mistral Small 3.2 24BMistral · mistral-small-3.2-24b-instruct |
79.0%
|
| 18 | GPT-4.1OpenAI · gpt-4.1 |
78.9%
|
| 19 | Qwen3 235B A22B Instruct 2507Alibaba · qwen3-235b-a22b-2507 |
78.7%
|
| 19 | Claude Haiku 4.5Anthropic · claude-haiku-4.5 |
78.7%
|
| 19 | Kimi K2 0711Moonshot AI · kimi-k2 |
78.7%
|
| 22 | command-a-reasoningCohere |
78.6%
|
| 23 | Nova Pro 1.0Amazon · nova-pro-v1 |
78.5%
|
| 23 | Command ACohere · command-a |
78.5%
|
| 25 | grok-4-1-fast-reasoningxAI |
78.5%
|
| 26 | Qwen3 30B A3B Instruct 2507Alibaba · qwen3-30b-a3b-instruct-2507 |
78.4%
|
| 27 | Gemini 2.5 FlashGoogle · gemini-2.5-flash |
78.2%
|
| 28 | Qwen3 30B A3B Instruct 2507Alibaba · qwen3-30b-a3b-instruct-2507 |
77.9%
|
| 28 | grok-4-1-fast-non-reasoningxAI |
77.9%
|
| 30 | palmyra-x-004Writer |
77.9%
|
| 31 | toolace-2-8bHuawei Noah And Ustc |
77.4%
|
| 32 | Mistral Small 3.2 24BMistral · mistral-small-3.2-24b-instruct |
77.3%
|
| 33 | Llama 3.3 70B InstructMeta · llama-3.3-70b-instruct |
76.6%
|
| 34 | qwen3-4b-instruct-2507Alibaba |
76.4%
|
| 35 | Claude Opus 4.5Anthropic · claude-opus-4.5 |
76.0%
|
| 35 | DeepSeek V3.2 ExpDeepSeek · deepseek-v3.2-exp |
76.0%
|
| 37 | grok-4-0709xAI |
75.6%
|
| 38 | xlam-2-32b-fc-rSalesforce |
75.5%
|
| 39 | falcon3-10b-instructTII |
75.4%
|
| 40 | GPT-4.1 MiniOpenAI · gpt-4.1-mini |
74.8%
|
| 41 | qwen3-4b-instruct-2507Alibaba |
74.7%
|
| 41 | Llama 4 ScoutMeta · llama-4-scout |
74.7%
|
| 43 | qwen3-1.7bAlibaba |
74.6%
|
| 44 | Gemma 3 27BGoogle · gemma-3-27b-it |
74.5%
|
| 45 | Gemini 2.5 FlashGoogle · gemini-2.5-flash |
74.4%
|
| 46 | Gemma 3 12BGoogle · gemma-3-12b-it |
74.2%
|
| 47 | Mistral NemoMistral · mistral-nemo |
74.0%
|
| 48 | Mistral NemoMistral · mistral-nemo |
73.8%
|
| 49 | llama-4-maverick-17b-128e-instruct-fp8Meta |
73.7%
|
| 50 | o3OpenAI |
73.2%
|
| 51 | arch-agent-3bKatanemo |
72.9%
|
| 52 | grok-4-0709xAI |
72.5%
|
| 53 | xlam-2-70b-fc-rSalesforce |
72.2%
|
| 54 | Llama 3.1 8B InstructMeta · llama-3.1-8b-instruct |
70.8%
|
| 54 | o4 MiniOpenAI · o4-mini |
70.8%
|
| 56 | GPT-5 NanoOpenAI · gpt-5-nano |
70.7%
|
| 57 | hammer2.1-3bMadeagents |
70.5%
|
| 58 | GPT-5.2OpenAI · gpt-5.2 |
70.4%
|
| 59 | nanbeige3.5-pro-thinkingNanbeige |
70.0%
|
| 59 | GPT-4.1OpenAI · gpt-4.1 |
70.0%
|
| 61 | hammer2.1-1.5bMadeagents |
69.5%
|
| 61 | hammer2.1-7bMadeagents |
69.5%
|
| 63 | Command R7B (12-2024)Cohere · command-r7b-12-2024 |
69.1%
|
| 64 | Qwen3 235B A22B Instruct 2507Alibaba · qwen3-235b-a22b-2507 |
68.9%
|
| 65 | GPT-4.1 MiniOpenAI · gpt-4.1-mini |
68.8%
|
| 66 | falcon3-7b-instructTII |
68.3%
|
| 67 | mistral-large-2411Mistral |
68.1%
|
| 68 | Mistral Medium 3Mistral · mistral-medium-3 |
68.0%
|
| 68 | xlam-2-8b-fc-rSalesforce |
68.0%
|
| 70 | bielik-11b-v2.3-instructSpeakleash And Ack Cyfronet Agh |
67.8%
|
| 71 | arch-agent-1.5bKatanemo |
67.7%
|
| 72 | coalm-70bUiuc Oumi |
67.3%
|
| 73 | GPT-5.2OpenAI · gpt-5.2 |
67.1%
|
| 74 | coalm-8bUiuc Oumi |
66.8%
|
| 75 | Nova Micro 1.0Amazon · nova-micro-v1 |
66.3%
|
| 76 | o3OpenAI |
66.2%
|
| 77 | o4 MiniOpenAI · o4-mini |
66.1%
|
| 78 | Mistral Medium 3Mistral · mistral-medium-3 |
66.0%
|
| 79 | Gemini 2.5 Flash LiteGoogle · gemini-2.5-flash-lite |
65.8%
|
| 80 | minicpm3-4bOpenbmb |
65.2%
|
| 81 | xlam-2-3b-fc-rSalesforce |
62.9%
|
| 82 | GPT-5 MiniOpenAI · gpt-5-mini |
62.5%
|
| 83 | Gemma 3 4BGoogle · gemma-3-4b-it |
60.8%
|
| 84 | GPT-4.1 NanoOpenAI · gpt-4.1-nano |
60.8%
|
| 85 | Phi 4Microsoft · phi-4 |
60.7%
|
| 86 | granite-3.1-8b-instructIBM |
60.3%
|
| 86 | granite-3.2-8b-instructIBM |
60.3%
|
| 88 | GPT-5 NanoOpenAI · gpt-5-nano |
59.4%
|
| 89 | granite-20b-functioncallingIBM |
58.7%
|
| 90 | GPT-5 MiniOpenAI · gpt-5-mini |
58.6%
|
| 91 | Llama 3.2 3B InstructMeta · llama-3.2-3b-instruct |
58.3%
|
| 92 | qwen3-0.6bAlibaba |
56.6%
|
| 93 | xlam-2-1b-fc-rSalesforce |
55.1%
|
| 94 | Gemini 2.5 Flash LiteGoogle · gemini-2.5-flash-lite |
54.9%
|
| 95 | hammer2.1-0.5bMadeagents |
54.6%
|
| 96 | falcon3-3b-instructTII |
54.5%
|
| 97 | DeepSeek V3.2 ExpDeepSeek · deepseek-v3.2-exp |
53.7%
|
| 98 | Claude Haiku 4.5Anthropic · claude-haiku-4.5 |
52.5%
|
| 99 | GPT-4.1 NanoOpenAI · gpt-4.1-nano |
50.3%
|
| 100 | rzn-tPhronetic Ai |
49.7%
|
| 101 | qwen3-0.6bAlibaba |
49.4%
|
| 102 | Claude Sonnet 4.5Anthropic · claude-sonnet-4.5 |
46.6%
|
| 103 | granite-4.0-350mIBM |
46.1%
|
| 104 | minicpm3-4bOpenbmb |
43.1%
|
| 105 | gemma-3-1b-itGoogle |
11.8%
|
| 106 | Llama 3.2 1B InstructMeta · llama-3.2-1b-instruct |
11.8%
|
| 107 | falcon3-1b-instructTII |
2.9%
|
| 108 | ministral-8b-2410Mistral |
0.0%
|
| 108 | llama-3.1-nemotron-ultra-253b-v1NVIDIA |
0.0%
|
Results as published by Berkeley Function Calling Leaderboard (BFCL) V4; we do not re-run them.
What it measures
The same task on functions and questions contributed by real users.
What it does not measure
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
Berkeley Function Calling Leaderboard (BFCL) V4 by the UC Berkeley Gorilla team, Apache License 2.0.