Models / Kimi K2.5

Moonshot AI

Kimi K2.5

29 published results from 4 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Moonshot AI
Sources
4
Our benchmarks
0
Price
Not yet published

Reported by others

1,436
Overall · rank 31 of 44
Unit
Arena rating, higher is better
Range
1,429 to 1,443
Sample
21569 votes
Configuration
Kimi K2.5
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,429
Overall · rank 88 of 177
Unit
Arena rating, higher is better
Range
1,423 to 1,435
Sample
8236 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,445
Overall · rank 60 of 177
Unit
Arena rating, higher is better
Range
1,442 to 1,448
Sample
74273 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,426
Business, management and finance · rank 73 of 402
Unit
Arena rating, higher is better
Range
1,411 to 1,441
Sample
1541 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,441
Business, management and finance · rank 64 of 402
Unit
Arena rating, higher is better
Range
1,435 to 1,446
Sample
14405 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,390
Creative writing · rank 83 of 407
Unit
Arena rating, higher is better
Range
1,373 to 1,406
Sample
1319 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,423
Creative writing · rank 54 of 407
Unit
Arena rating, higher is better
Range
1,417 to 1,430
Sample
12225 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,451
Expert prompts · rank 60 of 359
Unit
Arena rating, higher is better
Range
1,428 to 1,474
Sample
630 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,483
Expert prompts · rank 35 of 359
Unit
Arena rating, higher is better
Range
1,475 to 1,490
Sample
7011 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,432
Instruction following · rank 58 of 409
Unit
Arena rating, higher is better
Range
1,420 to 1,444
Sample
2337 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,439
Instruction following · rank 57 of 409
Unit
Arena rating, higher is better
Range
1,434 to 1,444
Sample
24198 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,430
Overall · rank 89 of 409
Unit
Arena rating, higher is better
Range
1,424 to 1,437
Sample
8358 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,450
Overall · rank 60 of 409
Unit
Arena rating, higher is better
Range
1,447 to 1,454
Sample
74638 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,404
Writing, literature and language · rank 83 of 408
Unit
Arena rating, higher is better
Range
1,391 to 1,417
Sample
1939 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,428
Writing, literature and language · rank 67 of 408
Unit
Arena rating, higher is better
Range
1,423 to 1,433
Sample
17888 votes
Configuration
Kimi K2.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by OpenHands Index
49.2%
Average · rank 18 of 34
Unit
% resolved, higher is better
Configuration
Kimi K2.5 (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
32.8%
Front end (SWE-Bench Multimodal) · rank 18 of 34
Unit
% resolved, higher is better
Configuration
Kimi K2.5 (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
18.8%
Greenfield (Commit0) · rank 21 of 34
Unit
% resolved, higher is better
Configuration
Kimi K2.5 (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
63.6%
Information gathering (GAIA) · rank 17 of 34
Unit
% resolved, higher is better
Configuration
Kimi K2.5 (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
68.8%
Issue resolution (SWE-Bench) · rank 27 of 34
Unit
% resolved, higher is better
Configuration
Kimi K2.5 (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
61.9%
Testing (SWT-Bench) · rank 22 of 34
Unit
% resolved, higher is better
Configuration
Kimi K2.5 (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by UGI Leaderboard
24.0%
Requested-length error · rank 217 of 370
Unit
% off the requested word count, lower is better
Configuration
Kimi K2.5
Measured
28 Jan 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
10.0%
Requested-length error · rank 81 of 370
Unit
% off the requested word count, lower is better
Configuration
Kimi K2.5
Measured
28 Jan 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.34
Style adherence · rank 215 of 370
Unit
score from 0 to 1, higher is better
Configuration
Kimi K2.5
Measured
28 Jan 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.34
Style adherence · rank 229 of 370
Unit
score from 0 to 1, higher is better
Configuration
Kimi K2.5
Measured
28 Jan 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
61.5
Writing score · rank 90 of 370
Unit
score out of 100, higher is better
Configuration
Kimi K2.5
Measured
28 Jan 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
62.4
Writing score · rank 87 of 370
Unit
score out of 100, higher is better
Configuration
Kimi K2.5
Measured
28 Jan 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
92.2%
Answer rate · rank 101 of 108
Unit
% of documents, higher is better
Configuration
Kimi K2.5
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
14.2%
Hallucination rate · rank 90 of 108
Unit
% of summaries, lower is better
Configuration
Kimi K2.5
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare Kimi K2.5 with