Models / Mistral Small 3.2 24B

Mistral

Mistral Small 3.2 24B

25 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Mistral
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

1,363
Business, management and finance · rank 173 of 402
Unit
Arena rating, higher is better
Range
1,352 to 1,374
Sample
2919 votes
Configuration
Mistral Small 3.2 24B
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,323
Creative writing · rank 175 of 407
Unit
Arena rating, higher is better
Range
1,310 to 1,336
Sample
2125 votes
Configuration
Mistral Small 3.2 24B
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,335
Expert prompts · rank 198 of 359
Unit
Arena rating, higher is better
Range
1,315 to 1,354
Sample
837 votes
Configuration
Mistral Small 3.2 24B
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,338
Instruction following · rank 194 of 409
Unit
Arena rating, higher is better
Range
1,329 to 1,347
Sample
4351 votes
Configuration
Mistral Small 3.2 24B
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,357
Overall · rank 193 of 409
Unit
Arena rating, higher is better
Range
1,351 to 1,362
Sample
17352 votes
Configuration
Mistral Small 3.2 24B
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,329
Writing, literature and language · rank 191 of 408
Unit
Arena rating, higher is better
Range
1,319 to 1,339
Sample
3730 votes
Configuration
Mistral Small 3.2 24B
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
65.7%
Irrelevance detection · rank 88 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
87.9%
Irrelevance detection · rank 17 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
15.1%
Memory · rank 56 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
18.1%
Memory · rank 50 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
14.8%
Multi-turn tasks · rank 58 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
11.5%
Multi-turn tasks · rank 63 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
32.4%
Overall accuracy · rank 59 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
37.1%
Overall accuracy · rank 51 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
93.8%
Relevance detection · rank 8 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
87.5%
Relevance detection · rank 27 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
89.7%
Single-turn calls (curated) · rank 4 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
73.6%
Single-turn calls (curated) · rank 81 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
79.0%
Single-turn calls (user-contributed) · rank 17 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
77.3%
Single-turn calls (user-contributed) · rank 32 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
7.5%
Web search · rank 60 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
31.0%
Web search · rank 34 of 109
Unit
% correct, higher is better
Configuration
mistral-small-3.2-24b-instruct
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
Reported by UGI Leaderboard
17.0%
Requested-length error · rank 140 of 370
Unit
% off the requested word count, lower is better
Configuration
Mistral Small 3.2 24B
Measured
6 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.35
Style adherence · rank 184 of 370
Unit
score from 0 to 1, higher is better
Configuration
Mistral Small 3.2 24B
Measured
6 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
36.3
Writing score · rank 233 of 370
Unit
score out of 100, higher is better
Configuration
Mistral Small 3.2 24B
Measured
6 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Mistral Small 3.2 24B with