Models / Kimi K2 0905

Moonshot AI

Kimi K2 0905

12 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Moonshot AI
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

1,403
Overall · rank 119 of 177
Unit
Arena rating, higher is better
Range
1,389 to 1,418
Sample
1447 votes
Configuration
Kimi K2 0905
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,404
Business, management and finance · rank 110 of 402
Unit
Arena rating, higher is better
Range
1,391 to 1,417
Sample
2041 votes
Configuration
Kimi K2 0905
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,382
Creative writing · rank 94 of 407
Unit
Arena rating, higher is better
Range
1,366 to 1,397
Sample
1475 votes
Configuration
Kimi K2 0905
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,417
Expert prompts · rank 106 of 359
Unit
Arena rating, higher is better
Range
1,393 to 1,441
Sample
571 votes
Configuration
Kimi K2 0905
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,390
Instruction following · rank 125 of 409
Unit
Arena rating, higher is better
Range
1,379 to 1,401
Sample
2870 votes
Configuration
Kimi K2 0905
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,419
Overall · rank 110 of 409
Unit
Arena rating, higher is better
Range
1,412 to 1,425
Sample
11537 votes
Configuration
Kimi K2 0905
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,382
Writing, literature and language · rank 115 of 408
Unit
Arena rating, higher is better
Range
1,370 to 1,394
Sample
2564 votes
Configuration
Kimi K2 0905
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by UGI Leaderboard
22.0%
Requested-length error · rank 202 of 370
Unit
% off the requested word count, lower is better
Configuration
Kimi K2 0905
Measured
15 Oct 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.31
Style adherence · rank 315 of 370
Unit
score from 0 to 1, higher is better
Configuration
Kimi K2 0905
Measured
15 Oct 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
54.4
Writing score · rank 130 of 370
Unit
score out of 100, higher is better
Configuration
Kimi K2 0905
Measured
15 Oct 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
98.6%
Answer rate · rank 76 of 108
Unit
% of documents, higher is better
Configuration
Kimi K2 0905
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
17.9%
Hallucination rate · rank 97 of 108
Unit
% of summaries, lower is better
Configuration
Kimi K2 0905
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare Kimi K2 0905 with