Provider
Moonshot AI
Base model family
Model intelligence profile
This family profile brings together 7 published benchmark records for Moonshot AI Kimi K2.7 Code. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.
Provider
Moonshot AI
Base model family
Release date
Not yet verified
Exact-family source only
Published benchmarks
7
Native scales kept separate
Artificial Analysis
No exact match
As of 2026-07-16
Benchmark-level evidence
These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.
source-native benchmark
42
A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.
Kimi K2.7 Code
42
source-native benchmark
34%
Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.
Kimi K2.7 Code
34%
source-native benchmark
90%
The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.
Kimi K2.7 Code
90%
source-native benchmark
33%
A broad expert-level benchmark of difficult academic reasoning and knowledge questions.
Kimi K2.7 Code
33%
source-native benchmark
63%
An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.
Kimi K2.7 Code
63%
source-native benchmark
67%
A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.
Kimi K2.7 Code
67%
source-native benchmark
90%
A legacy agentic tool-use benchmark for conversational work in a telecom environment.
Kimi K2.7 Code
90%
Independent operational reference
No exact, reviewed Artificial Analysis operational record is attached to this family. We do not inherit price, release, or throughput values from a nearby model name.
Early-user signal
No editor-reviewed community-opinions paragraph is active for this model. An active paragraph requires at least two retained Reddit discussions from the minimum 14-day launch window—implemented exactly as release date through day 14, end-exclusive—plus paraphrase-only review and a sealed publication artifact. Community reports stay separate from every benchmark score and model comparison.
Independent early-test signal
An editorial paraphrase of 2 launch-window field tests from 2 independent authors on X, including 2 reports with a described method or inspectable artifact. The window runs from 12 Jun 2026 up to 26 Jun 2026, and exact model identity was editor reviewed.
Early Kimi K2.7 Code evidence covered two distinct evaluations: a real-world engineering benchmark based on whether maintainers would merge generated changes, and a game-oriented reasoning and cost test. Both showed useful coding or reasoning performance, but their task designs and success measures differ substantially. The findings support practical potential without implying uniform behavior across repositories or agents.
These reports are selectively surfaced and are not a representative sample. They never affect benchmark scores, rankings, winners, or comparison outcomes.
Business Skills V3 · proposed
The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.
Future first-party coverage
These links describe evaluation contracts, not published Kimi K2.7 Code results.
Model comparisons
Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.
Publication safeguards
Stable family URL
New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.
More from Moonshot AI