Confirm Action

Are you sure you want to proceed?

Model intelligence profile

Kimi K2.5 benchmark results and model details

This family profile brings together 3 published benchmark records for Moonshot AI Kimi K2.5. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.

Published benchmark evidence available · Business Skills V3 not yet authorized

Provider

Moonshot AI

Base model family

Release date

2026-01-27

Exact-family source only

Published benchmarks

3

Native scales kept separate

Artificial Analysis

Reference available

As of 2026-07-16

Benchmark-level evidence

Published Kimi K2.5 benchmark results

These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.

How sources are reviewed →

Spring Prompt benchmark

PredictTheWeek pilot

7.0%

Mean prediction score

Can LLMs anticipate next week’s Guardian agenda from last week’s coverage?

This benchmark is not on a weekly live cadence yet. The table and charts reflect a single multi-model comparison on one forecast window—a pilot snapshot, not an updating leaderboard. The evaluation setup (clustering, prompts, automated judge, and scoring checks) is still being refined; reported scores and details may change as we improve the pipeline.

1 published configuration

Kimi K2.5

7.0%

Accuracy (partial support or better)
13.0%
Outcome-line coverage
4.9%
Pilot composite
7.7%
Source: Spring Prompt Recorded 2 Apr 2026 16–22 Mar 2026 → 23–29 Mar 2026 Source record →

Spring Prompt benchmark

ROASBench

13.10

Average ROASBench score

A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.

1 published configuration

MoonshotAI: Kimi K2.5

13.10

Average contribution profit
$-292,423
Average ROAS
0.97×
Completed runs
2
Score variability
±2.02
Source: Spring Prompt Recorded 25 Jul 2026 Published ROASBench cache · 12-month simulation Source record →

source-native benchmark

Repository Issue Resolution

68.8%

OpenHands SWE-Bench resolved

Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.

Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.

1 published configuration

OpenHands v1.8.3 + Kimi-K2.5

68.8%

95% task-bootstrap interval
Unavailable
Task evidence
not-eligible-for-task-bootstrap
Source: OpenHands Index Recorded 30 Jun 2026 OpenHands · SWE-Bench 2026.06.30-3015ac6 Source record →

Independent operational reference

Artificial Analysis details

Representative configuration: Kimi K2.5 (Reasoning). Lifecycle: deprecated; released 2026-01-27.

Input price

$0.60 / 1M tokens

Output price

$3.00 / 1M tokens

Median output speed

49.8 tok/s

Operational reference only. Price and throughput are not model-quality scores and never affect Spring Prompt comparisons.

Early-user signal

Initial community opinions

Anecdotal · never scored

A paraphrased editorial synthesis of 4 Reddit discussions created in the 14 days after launch (27 Jan 2026 to 10 Feb 2026). Anecdotal context only — it never affects scores or rankings.

The search-indexed launch-window sample treated Kimi K2.5 as a capable open-weight option for coding, vision, and agent workflows, particularly when instructions were explicit. Recurring limitations included uneven non-coding style, occasional hallucination reports in particular hosts, and hardware demands that placed meaningful self-hosting outside ordinary consumer setups.

Independent early-test signal

Early technical field tests

X · anecdotal · never scored

An editorial paraphrase of 2 launch-window field tests from 2 independent authors on X, including 2 reports with a described method or inspectable artifact. The window runs from 27 Jan 2026 up to 10 Feb 2026, and exact model identity was editor reviewed.

Early Kimi K2.5 field reports highlighted multimodal tool use and parallel-agent workflows. One tester exercised more than one hundred interleaved image and tool calls, while another used the model in Kimi Code for parallel research and software work and found that software tasks benefited from more deliberate system prompts and sub-agent coordination. The evidence suggests high orchestration potential with host and configuration sensitivity.

These reports are selectively surfaced and are not a representative sample. They never affect benchmark scores, rankings, winners, or comparison outcomes.

Business Skills V3 · proposed

How Spring Prompt plans to test Kimi K2.5

The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.

Proposed reasoning
Provider-default reasoning
Provider revision
Exact provider revision will be resolved and frozen only after preflight and execution approval
Tools and service
No tools · Provider-default service tier
Planning configuration
kimi-k2.5
Generation controls
Temperature 0 proposed

Future first-party coverage

15 proposed Business Skills V3 task areas

These links describe evaluation contracts, not published Kimi K2.5 results.

Model comparisons

Compare Kimi K2.5 side by side

Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.

Publication safeguards

What must pass before a V3 result appears

  1. 1First-party task-local comparisons and eligible external evidence must both be present.
  2. 2Model and provider configuration identity must match the reviewed release exactly.
  3. 3Coverage, reliability, judge diagnostics, and sealed stability checks must pass.
  4. 4Uncertainty and missing evidence remain visible when results are published.

Stable family URL

Evidence can grow without changing the page

New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.

More from Moonshot AI

Other owned model profiles