OpenAI
GPT-5.6 Sol
- Benchmark records
- 9
- Snapshot
- 2026-07-30
- Release date
- 2026-07-09
Independent model comparison
Compare the models using only evidence that can be identified and attributed. Exact benchmark matches appear first; independent facts and user reports remain separate and do not create an overall ranking.
Exact comparison coverage
7 benchmark records share an exact reviewed comparison key.
Evidence strength
Editorial review pending
Substantial matched evidence exists, but this pair has not passed the pair-native publication gate.
Models at a glance
OpenAI
Moonshot AI
| Detail | GPT-5.6 Sol | Kimi K2.7 Code |
|---|---|---|
| Provider |
OpenAI |
Moonshot AI |
| Release date Shown only where an exact reviewed release date is available. |
2026-07-09 |
Not verified |
| Input price Artificial Analysis operational reference; see the attributed details below. |
$5.00 / 1M tokens |
Not verified |
| Output price Artificial Analysis operational reference; see the attributed details below. |
$30.00 / 1M tokens |
Not verified |
| Median output speed Artificial Analysis operational reference; see the attributed details below. |
55.9 tok/s |
Not verified |
| Exact benchmark coverage Only records carrying the same reviewed comparison identity are counted. |
7 shared records |
7 shared records |
Price and speed appear only when an exact reviewed operational reference is attached to that model identity.
Primary comparison evidence
A shared benchmark name is not enough. A result appears side by side only when both model records carry the same reviewed comparison key. Duplicate or unversioned records are omitted rather than guessed into alignment.
A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.
GPT-5.6 Sol
56
Artificial Analysis Intelligence Index
Configuration: GPT-5.6 Sol (high)
Artificial Analysis · as of 30 Jul 2026 · Intelligence Index v4.1 · 2026-07-30
Kimi K2.7 Code
42
Artificial Analysis Intelligence Index
Configuration: Kimi K2.7 Code
Artificial Analysis · as of 30 Jul 2026 · Intelligence Index v4.1 · 2026-07-30
Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.
GPT-5.6 Sol
56%
GDPval-AA v2
Configuration: GPT-5.6 Sol (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Kimi K2.7 Code
34%
GDPval-AA v2
Configuration: Kimi K2.7 Code
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.
GPT-5.6 Sol
93%
GPQA Diamond
Configuration: GPT-5.6 Sol (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Kimi K2.7 Code
90%
GPQA Diamond
Configuration: Kimi K2.7 Code
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
A broad expert-level benchmark of difficult academic reasoning and knowledge questions.
GPT-5.6 Sol
44%
Humanity's Last Exam
Configuration: GPT-5.6 Sol (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Kimi K2.7 Code
33%
Humanity's Last Exam
Configuration: Kimi K2.7 Code
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.
GPT-5.6 Sol
69%
IFBench
Configuration: GPT-5.6 Sol (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Kimi K2.7 Code
63%
IFBench
Configuration: Kimi K2.7 Code
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.
GPT-5.6 Sol
87%
Terminal-Bench v2.1
Configuration: GPT-5.6 Sol (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Kimi K2.7 Code
67%
Terminal-Bench v2.1
Configuration: Kimi K2.7 Code
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
A legacy agentic tool-use benchmark for conversational work in a telecom environment.
GPT-5.6 Sol
83%
τ²-Bench Telecom
Configuration: GPT-5.6 Sol (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Kimi K2.7 Code
90%
τ²-Bench Telecom
Configuration: Kimi K2.7 Code
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Secondary context · not head-to-head
The sources below evaluated or discussed each model independently. Putting them in adjacent columns makes them easier to inspect, but it does not make their metrics, samples, or observations directly comparable.
Attributed third-party facts
These facts retain their source, date, protocol, and model identity. They are contextual model-level signals—not a Spring Prompt overall score.
Input price
$5.00 / 1M tokens
Representative configuration on the cited source page.
Artificial Analysis · as of 16 Jul 2026
Output price
$30.00 / 1M tokens
Representative configuration on the cited source page.
Artificial Analysis · as of 16 Jul 2026
Median output speed
55.9 tok/s
Median output tokens received per second after generation begins; this excludes time to first token and is not end-to-end response latency.
Artificial Analysis · as of 16 Jul 2026
No reviewed Artificial Analysis record is attached to this model snapshot.
Anecdotal model-level reports
Recurring themes from each model's declared observation window. They are not a representative survey and are not direct comparisons between these models.
Initial community opinions
The search-indexed launch-window sample highlighted GPT-5.6 Sol's autonomous code review, broad bug discovery, project comprehension, and promising frontend output. Counter-reports focused on heavy quota use, slow execution, overengineering, guardrail interference, rollout friction, and occasional basic mistakes. Strong initial results therefore sat alongside unresolved cost and consistency concerns.
Source venue: Reddit · window: 2026-07-09 to 2026-07-23 · observed items: 4
No editor-reviewed Reddit opinions paragraph is active for this model snapshot.
Independent early-test reports
Editor-paraphrased reports from independent authors in each model's fixed 14-day launch window. These selectively surfaced anecdotes are not direct head-to-head tests unless they also appear in the dedicated direct-report section above.
No editor-reviewed X field-test paragraph is active for this model snapshot.
Early technical field tests
Early Kimi K2.7 Code evidence covered two distinct evaluations: a real-world engineering benchmark based on whether maintainers would merge generated changes, and a game-oriented reasoning and cost test. Both showed useful coding or reasoning performance, but their task designs and success measures differ substantially. The findings support practical potential without implying uniform behavior across repositories or agents.
Source venue: X · window: 2026-06-12 to 2026-06-26 · 2 tests · 2 independent authors · 2 with a method or artifact · identity reviewed
X field tests remain separate from Reddit community opinions, Artificial Analysis facts, and every scored benchmark.
Spring Prompt does not name an overall winner from unrelated benchmark scores, third-party metrics, Reddit opinions, or X field-test reports.
Different benchmarks measure different constructs and may use different populations, prompts, harnesses, model revisions, tools, and score scales. This page does not average those values or infer a winner from source coverage. Test both models on your own production work before making a consequential choice.