OpenAI
GPT-5.5
- Benchmark records
- 4
- Snapshot
- 2026-07-20
- Release date
- 2026-04-23
Independent model comparison
Compare the models using only evidence that can be identified and attributed. Exact benchmark matches appear first; independent facts and user reports remain separate and do not create an overall ranking.
Exact comparison coverage
1 benchmark record share an exact reviewed comparison key.
Evidence strength
Limited exact coverage
The page is useful for inspection but does not yet contain enough pair-native evidence for Search publication.
Models at a glance
OpenAI
OpenAI
| Detail | GPT-5.5 | GPT-5.6 Sol |
|---|---|---|
| Provider |
OpenAI |
OpenAI |
| Release date Shown only where an exact reviewed release date is available. |
2026-04-23 |
2026-07-09 |
| Input price Artificial Analysis operational reference; see the attributed details below. |
$5.00 / 1M tokens |
$5.00 / 1M tokens |
| Output price Artificial Analysis operational reference; see the attributed details below. |
$30.00 / 1M tokens |
$30.00 / 1M tokens |
| Median output speed Artificial Analysis operational reference; see the attributed details below. |
66.3 tok/s |
55.9 tok/s |
| Exact benchmark coverage Only records carrying the same reviewed comparison identity are counted. |
1 shared record |
1 shared record |
Price and speed appear only when an exact reviewed operational reference is attached to that model identity.
Primary comparison evidence
A shared benchmark name is not enough. A result appears side by side only when both model records carry the same reviewed comparison key. Duplicate or unversioned records are omitted rather than guessed into alignment.
A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.
GPT-5.5
25.68
Average ROASBench score
Configuration: OpenAI: GPT-5.5
Spring Prompt · Reviewed public-catalogue SQLite projection · 12-month simulation
GPT-5.6 Sol
4 configurations
Average ROASBench score
Configuration: 4 configurations compared separately
Spring Prompt · Reviewed public-catalogue SQLite projection · 12-month simulation
Secondary context · not head-to-head
The sources below evaluated or discussed each model independently. Putting them in adjacent columns makes them easier to inspect, but it does not make their metrics, samples, or observations directly comparable.
Attributed third-party facts
These facts retain their source, date, protocol, and model identity. They are contextual model-level signals—not a Spring Prompt overall score.
Input price
$5.00 / 1M tokens
Representative configuration on the cited source page.
Artificial Analysis · as of 16 Jul 2026
Output price
$30.00 / 1M tokens
Representative configuration on the cited source page.
Artificial Analysis · as of 16 Jul 2026
Median output speed
66.3 tok/s
Median output tokens received per second after generation begins; this excludes time to first token and is not end-to-end response latency.
Artificial Analysis · as of 16 Jul 2026
Input price
$5.00 / 1M tokens
Representative configuration on the cited source page.
Artificial Analysis · as of 16 Jul 2026
Output price
$30.00 / 1M tokens
Representative configuration on the cited source page.
Artificial Analysis · as of 16 Jul 2026
Median output speed
55.9 tok/s
Median output tokens received per second after generation begins; this excludes time to first token and is not end-to-end response latency.
Artificial Analysis · as of 16 Jul 2026
Anecdotal model-level reports
Recurring themes from each model's declared observation window. They are not a representative survey and are not direct comparisons between these models.
Initial community opinions
In the search-indexed launch-window sample, GPT-5.5 was frequently praised for speed, focus, complex debugging, planning, and more natural chat responses. The main caveats were heavier usage at high effort, uneven frontend work, rollout or integration differences, and dependence on the surrounding Codex harness. Reports suggested a practical productivity gain for many workflows but not a uniform change across coding specialties.
Source venue: Reddit · window: 2026-04-23 to 2026-05-07 · observed items: 4
Initial community opinions
The search-indexed launch-window sample highlighted GPT-5.6 Sol's autonomous code review, broad bug discovery, project comprehension, and promising frontend output. Counter-reports focused on heavy quota use, slow execution, overengineering, guardrail interference, rollout friction, and occasional basic mistakes. Strong initial results therefore sat alongside unresolved cost and consistency concerns.
Source venue: Reddit · window: 2026-07-09 to 2026-07-23 · observed items: 4
Spring Prompt does not name an overall winner from unrelated benchmark scores, third-party metrics, Reddit opinions, or X field-test reports.
Different benchmarks measure different constructs and may use different populations, prompts, harnesses, model revisions, tools, and score scales. This page does not average those values or infer a winner from source coverage. Test both models on your own production work before making a consequential choice.