Models / Llama 3.1 70B Instruct

Meta

Llama 3.1 70B Instruct

9 published results from 2 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Meta
Sources
2
Our benchmarks
0
Price
Not yet published

Reported by others

1,287
Business, management and finance · rank 257 of 402
Unit
Arena rating, higher is better
Range
1,279 to 1,296
Sample
6381 votes
Configuration
Llama 3.1 70B Instruct
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,257
Creative writing · rank 257 of 407
Unit
Arena rating, higher is better
Range
1,249 to 1,265
Sample
8250 votes
Configuration
Llama 3.1 70B Instruct
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,275
Expert prompts · rank 248 of 359
Unit
Arena rating, higher is better
Range
1,264 to 1,287
Sample
2924 votes
Configuration
Llama 3.1 70B Instruct
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,273
Instruction following · rank 272 of 409
Unit
Arena rating, higher is better
Range
1,267 to 1,278
Sample
21910 votes
Configuration
Llama 3.1 70B Instruct
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,293
Overall · rank 277 of 409
Unit
Arena rating, higher is better
Range
1,290 to 1,297
Sample
55240 votes
Configuration
Llama 3.1 70B Instruct
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,264
Writing, literature and language · rank 277 of 408
Unit
Arena rating, higher is better
Range
1,257 to 1,270
Sample
14962 votes
Configuration
Llama 3.1 70B Instruct
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by UGI Leaderboard
15.0%
Requested-length error · rank 119 of 370
Unit
% off the requested word count, lower is better
Configuration
Llama 3.1 70B Instruct
Measured
13 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.38
Style adherence · rank 67 of 370
Unit
score from 0 to 1, higher is better
Configuration
Llama 3.1 70B Instruct
Measured
13 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
32.8
Writing score · rank 250 of 370
Unit
score out of 100, higher is better
Configuration
Llama 3.1 70B Instruct
Measured
13 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Llama 3.1 70B Instruct with