Provider
OpenAI
Base model family
Model intelligence profile
This family profile brings together 7 published benchmark records for OpenAI GPT-5.4. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.
Provider
OpenAI
Base model family
Release date
2026-03-05
Exact-family source only
Published benchmarks
7
Native scales kept separate
Artificial Analysis
Reference available
As of 2026-07-16
Benchmark-level evidence
These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.
Spring Prompt benchmark
26.5%
Mean prediction score
Can LLMs anticipate next week’s Guardian agenda from last week’s coverage?
This benchmark is not on a weekly live cadence yet. The table and charts reflect a single multi-model comparison on one forecast window—a pilot snapshot, not an updating leaderboard. The evaluation setup (clustering, prompts, automated judge, and scoring checks) is still being refined; reported scores and details may change as we improve the pipeline.
GPT-5.4
26.5%
Spring Prompt benchmark
18.39
Average ROASBench score
A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.
OpenAI: GPT-5.4
18.39
source-native benchmark
2 configurations
Coding Agent Index v1.2 point score
Artificial Analysis Coding Agent Index v1.2 point scores and task-specific operational measurements for exact agent and model configurations.
These are exact agent-plus-model configurations, not bare-model scores.
Codex - GPT-5.4 (medium)
41.1
Cursor CLI - GPT-5.4 (medium)
40.1
source-native benchmark
75.6%
OpenHands SWE-Bench resolved
Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.
Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.
OpenHands v1.18.1 + GPT-5.4
75.6%
source-native benchmark
55.7%
Overall pass@1
Source-native STATE-Bench completion, UX and cost results for stateful enterprise workflows.
A harnessed-model workflow result.
GPT-5.4 · high
55.7%
source-native benchmark
58.6%
Overall pass@1
Source-native STATE-Bench completion, UX and cost results for stateful enterprise workflows.
A harnessed-model workflow result.
GPT-5.4 · high
58.6%
source-native benchmark
87.02%
Direct benchmark score
Directly measured ability to return accurate values in the requested structured schema across the benchmark's evaluated text, image and audio modalities.
Point order reproduces the source's direct Overall score. Rank ranges come from overlap of marginal record-cluster bootstrap intervals for Overall; they are not simultaneous confidence intervals for rank.
GPT-5.4
87.02%
Independent operational reference
Representative configuration: GPT-5.4 (xhigh). Lifecycle: deprecated; released 2026-03-05.
Input price
$2.50 / 1M tokens
Output price
$15.00 / 1M tokens
Median output speed
150.3 tok/s
Operational reference only. Price and throughput are not model-quality scores and never affect Spring Prompt comparisons.
Early-user signal
A paraphrased editorial synthesis of 4 Reddit discussions created in the 14 days after launch (5 Mar 2026 to 19 Mar 2026). Anecdotal context only — it never affects scores or rankings.
In the search-indexed launch-window sample, GPT-5.4 drew positive reports for complex coding, self-correction, writing, and customization, though the perceived change depended on the surface and task. Some users retained older Codex or GPT variants for smaller work because they found them faster, cheaper, or more predictable, and the unified model naming caused early confusion.
Independent early-test signal
An editorial paraphrase of 2 launch-window field tests from 2 independent authors on X, including 2 reports with a described method or inspectable artifact. The window runs from 5 Mar 2026 up to 19 Mar 2026, and exact model identity was editor reviewed.
Early GPT-5.4 evidence showed a marked improvement on a professional-agent benchmark, while a concrete distillation workflow exposed a different tradeoff: fast data generation but disappointing teacher-data quality and rapid quota consumption. The combined picture is stronger agent capability without a universal workflow conclusion; output quality, quotas, and task framing still matter.
These reports are selectively surfaced and are not a representative sample. They never affect benchmark scores, rankings, winners, or comparison outcomes.
Business Skills V3 · proposed
The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.
Future first-party coverage
These links describe evaluation contracts, not published GPT-5.4 results.
Model comparisons
Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.
Publication safeguards
Stable family URL
New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.
More from OpenAI