Models / Qwen3 30B A3B Instruct 2507

Alibaba

Qwen3 30B A3B Instruct 2507

25 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Alibaba
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

1,397
Business, management and finance · rank 122 of 402
Unit
Arena rating, higher is better
Range
1,388 to 1,407
Sample
4105 votes
Configuration
Qwen3 30B A3B Instruct 2507
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,319
Creative writing · rank 182 of 407
Unit
Arena rating, higher is better
Range
1,308 to 1,330
Sample
2932 votes
Configuration
Qwen3 30B A3B Instruct 2507
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,395
Expert prompts · rank 141 of 359
Unit
Arena rating, higher is better
Range
1,378 to 1,412
Sample
1149 votes
Configuration
Qwen3 30B A3B Instruct 2507
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,366
Instruction following · rank 165 of 409
Unit
Arena rating, higher is better
Range
1,359 to 1,374
Sample
5922 votes
Configuration
Qwen3 30B A3B Instruct 2507
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,383
Overall · rank 165 of 409
Unit
Arena rating, higher is better
Range
1,378 to 1,388
Sample
23179 votes
Configuration
Qwen3 30B A3B Instruct 2507
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,342
Writing, literature and language · rank 178 of 408
Unit
Arena rating, higher is better
Range
1,333 to 1,350
Sample
5022 votes
Configuration
Qwen3 30B A3B Instruct 2507
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
79.9%
Irrelevance detection · rank 59 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
74.8%
Irrelevance detection · rank 72 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
17.6%
Memory · rank 51 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
9.7%
Memory · rank 73 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
30.0%
Multi-turn tasks · rank 40 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
23.5%
Multi-turn tasks · rank 48 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
41.4%
Overall accuracy · rank 41 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
36.7%
Overall accuracy · rank 53 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
81.2%
Relevance detection · rank 41 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
93.8%
Relevance detection · rank 8 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
85.8%
Single-turn calls (curated) · rank 37 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
88.9%
Single-turn calls (curated) · rank 9 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
77.9%
Single-turn calls (user-contributed) · rank 28 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
78.4%
Single-turn calls (user-contributed) · rank 26 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
22.5%
Web search · rank 40 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
17.5%
Web search · rank 46 of 109
Unit
% correct, higher is better
Configuration
qwen3-30b-a3b-instruct-2507
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
Reported by UGI Leaderboard
13.0%
Requested-length error · rank 103 of 370
Unit
% off the requested word count, lower is better
Configuration
Qwen3 30B A3B Instruct 2507
Measured
9 Oct 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.31
Style adherence · rank 302 of 370
Unit
score from 0 to 1, higher is better
Configuration
Qwen3 30B A3B Instruct 2507
Measured
9 Oct 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
27.9
Writing score · rank 285 of 370
Unit
score out of 100, higher is better
Configuration
Qwen3 30B A3B Instruct 2507
Measured
9 Oct 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Qwen3 30B A3B Instruct 2507 with