Models / Llama 3.3 70B Instruct

Meta

Llama 3.3 70B Instruct

17 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Meta
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

1,304
Business, management and finance · rank 231 of 402
Unit
Arena rating, higher is better
Range
1,296 to 1,311
Sample
6618 votes
Configuration
Llama 3.3 70B Instruct
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,285
Creative writing · rank 223 of 407
Unit
Arena rating, higher is better
Range
1,278 to 1,292
Sample
8016 votes
Configuration
Llama 3.3 70B Instruct
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,298
Expert prompts · rank 231 of 359
Unit
Arena rating, higher is better
Range
1,286 to 1,309
Sample
2904 votes
Configuration
Llama 3.3 70B Instruct
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,293
Instruction following · rank 249 of 409
Unit
Arena rating, higher is better
Range
1,288 to 1,298
Sample
18744 votes
Configuration
Llama 3.3 70B Instruct
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,318
Overall · rank 244 of 409
Unit
Arena rating, higher is better
Range
1,314 to 1,321
Sample
54368 votes
Configuration
Llama 3.3 70B Instruct
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,292
Writing, literature and language · rank 241 of 408
Unit
Arena rating, higher is better
Range
1,286 to 1,297
Sample
13769 votes
Configuration
Llama 3.3 70B Instruct
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
53.5%
Irrelevance detection · rank 97 of 109
Unit
% correct, higher is better
Configuration
llama-3.3-70b-instruct
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
8.2%
Memory · rank 79 of 109
Unit
% correct, higher is better
Configuration
llama-3.3-70b-instruct
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
21.5%
Multi-turn tasks · rank 50 of 109
Unit
% correct, higher is better
Configuration
llama-3.3-70b-instruct
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
31.9%
Overall accuracy · rank 62 of 109
Unit
% correct, higher is better
Configuration
llama-3.3-70b-instruct
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
100.0%
Relevance detection · rank 1 of 109
Unit
% correct, higher is better
Configuration
llama-3.3-70b-instruct
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
88.0%
Single-turn calls (curated) · rank 22 of 109
Unit
% correct, higher is better
Configuration
llama-3.3-70b-instruct
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
76.6%
Single-turn calls (user-contributed) · rank 33 of 109
Unit
% correct, higher is better
Configuration
llama-3.3-70b-instruct
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
10.0%
Web search · rank 56 of 109
Unit
% correct, higher is better
Configuration
llama-3.3-70b-instruct
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
Reported by UGI Leaderboard
10.0%
Requested-length error · rank 81 of 370
Unit
% off the requested word count, lower is better
Configuration
Llama 3.3 70B Instruct
Measured
10 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.37
Style adherence · rank 102 of 370
Unit
score from 0 to 1, higher is better
Configuration
Llama 3.3 70B Instruct
Measured
10 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
26.2
Writing score · rank 295 of 370
Unit
score out of 100, higher is better
Configuration
Llama 3.3 70B Instruct
Measured
10 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Llama 3.3 70B Instruct with