Models / Kimi K3

Moonshot AI

Kimi K3

35 published results from 5 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Moonshot AI
Sources
5
Our benchmarks
3
Price
Not yet published

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
0
Ladder Elo · rank 12 of 16
Unit
ladder Elo, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
6.8 s
Median move time · rank 9 of 16
Unit
milliseconds, lower is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · CatalogBench
95.8%
Content quality · rank 12 of 18
Unit
% of checks, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Failures
invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.6%
Failed outputs · rank 15 of 18
Unit
% of products, lower is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Failures
invalid JSON: 1 of 168 attempts
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
1.8%
Publish-ready listings · rank 12 of 18
Unit
% of products, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Failures
invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Reliably publish-ready · rank 10 of 18
Unit
% of products, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Failures
invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
7.18
Unsupported claims · rank 12 of 18
Unit
claims per product, lower is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Failures
invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
99.4%
Channel compliance · rank 9 of 18
Unit
% of products, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
97.0%
Conflicts caught · rank 16 of 18
Unit
% of conflicts, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
96.0%
Content quality · rank 13 of 18
Unit
% of checks, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
$0.0469
Cost per product · rank 17 of 18
Unit
US dollars, lower is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
94.8%
Decision accuracy · rank 9 of 18
Unit
% of decisions, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
94.4%
Field accuracy · rank 3 of 18
Unit
% of missing fields, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
5.7%
Invented values · rank 17 of 18
Unit
% of filled values, lower is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
47.6%
Publish-ready listings · rank 9 of 18
Unit
% of products, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
17.9%
Reliably publish-ready · rank 11 of 18
Unit
% of products, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
1.39
Unsupported claims · rank 9 of 18
Unit
claims per product, lower is better
Configuration
Kimi K3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · ROASBench
$1.01
Cost of a run · rank 14 of 17
Unit
US dollars, lower is better
Configuration
Kimi K3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
0
Months over budget · rank 1 of 17
Unit
months of 12, lower is better
Configuration
Kimi K3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
44.1
Overall score · rank 8 of 17
Unit
score out of 100, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
£0.50
Return on ad spend · rank 8 of 17
Unit
profit per £1 spent, higher is better
Configuration
Kimi K3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Reported by others

0.09
Confirmed task success · rank 4 of 46
Unit
IPS effect estimate, higher is better
Range
0.08 to 0.10
Sample
97021 observations
Configuration
Kimi K3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.07
Praise over complaint · rank 9 of 46
Unit
IPS effect estimate, higher is better
Range
0.06 to 0.09
Sample
37953 observations
Configuration
Kimi K3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.01
Steerability · rank 20 of 46
Unit
IPS effect estimate, higher is better
Range
-0.02 to 0.01
Sample
110549 observations
Configuration
Kimi K3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
10631569 observations
Configuration
Kimi K3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,472
Overall · rank 10 of 177
Unit
Arena rating, higher is better
Range
1,468 to 1,476
Sample
26132 votes
Configuration
Kimi K3 (max)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,489
Business, management and finance · rank 2 of 402
Unit
Arena rating, higher is better
Range
1,480 to 1,498
Sample
5149 votes
Configuration
Kimi K3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,458
Creative writing · rank 12 of 407
Unit
Arena rating, higher is better
Range
1,449 to 1,467
Sample
5249 votes
Configuration
Kimi K3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,533
Expert prompts · rank 1 of 359
Unit
Arena rating, higher is better
Range
1,521 to 1,545
Sample
2457 votes
Configuration
Kimi K3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,486
Instruction following · rank 6 of 409
Unit
Arena rating, higher is better
Range
1,479 to 1,493
Sample
8851 votes
Configuration
Kimi K3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,488
Overall · rank 7 of 409
Unit
Arena rating, higher is better
Range
1,483 to 1,493
Sample
26400 votes
Configuration
Kimi K3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,475
Writing, literature and language · rank 8 of 408
Unit
Arena rating, higher is better
Range
1,467 to 1,483
Sample
6938 votes
Configuration
Kimi K3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by tau2-bench
17.5%
Consistency (pass^4) · rank 12 of 21
Unit
% of tasks, higher is better
Configuration
kimi-k3
Measured
24 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
37.1%
Task success (pass^1) · rank 11 of 21
Unit
% of tasks, higher is better
Configuration
kimi-k3
Measured
24 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.

Compare Kimi K3 with