Confirm Action

Are you sure you want to proceed?

Model intelligence profile

GPT-5.4 Mini benchmark results and model details

This family profile brings together 4 published benchmark records for OpenAI GPT-5.4 Mini. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.

Published benchmark evidence available · Business Skills V3 not yet authorized

Provider

OpenAI

Base model family

Release date

Not yet verified

Exact-family source only

Published benchmarks

4

Native scales kept separate

Artificial Analysis

No exact match

As of 2026-07-16

Benchmark-level evidence

Published GPT-5.4 Mini benchmark results

These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.

How sources are reviewed →

Spring Prompt benchmark

BulletBench

147

60-second bullet Elo

Chess decision quality under a shared 60-second game clock and a 10+1 per-move clock.

Ladder Elo: anchored to a Stockfish skill ladder (random mover = 400). Internally consistent ordering, not FIDE-calibrated.

1 published configuration

GPT-5.4 mini (low)

147

Bullet 95% interval
2–267
10+1 Elo
0
Bullet games
96
Bullet median move time
2.60s
Average bullet game cost
$0.0322
Source: Spring Prompt Recorded 3 Aug 2026 Ladder v1 · controls 10+1 and 60 Source record →

Spring Prompt benchmark

PredictTheWeek pilot

22.5%

Mean prediction score

Can LLMs anticipate next week’s Guardian agenda from last week’s coverage?

This benchmark is not on a weekly live cadence yet. The table and charts reflect a single multi-model comparison on one forecast window—a pilot snapshot, not an updating leaderboard. The evaluation setup (clustering, prompts, automated judge, and scoring checks) is still being refined; reported scores and details may change as we improve the pipeline.

1 published configuration

GPT-5.4 mini

22.5%

Accuracy (partial support or better)
29.0%
Outcome-line coverage
8.2%
Pilot composite
15.5%
Source: Spring Prompt Recorded 2 Apr 2026 16–22 Mar 2026 → 23–29 Mar 2026 Source record →

Spring Prompt benchmark

ROASBench

11.83

Average ROASBench score

A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.

1 published configuration

OpenAI: GPT-5.4 Mini

11.83

Average contribution profit
$-353,629
Average ROAS
0.87×
Completed runs
1
Score variability
±0.00
Source: Spring Prompt Recorded 25 Jul 2026 Published ROASBench cache · 12-month simulation Source record →

source-native benchmark

Structured Output Reliability

84.70%

Direct benchmark score

Directly measured ability to return accurate values in the requested structured schema across the benchmark's evaluated text, image and audio modalities.

Point order reproduces the source's direct Overall score. Rank ranges come from overlap of marginal record-cluster bootstrap intervals for Overall; they are not simultaneous confidence intervals for rank.

1 published configuration

GPT-5.4-Mini

84.70%

95% interval
84.20–85.18
Evaluated modality coverage
100%
Source: Structured Output Benchmark (SOB) Recorded 17 Jul 2026 Sob-v1@da785a8521c8954283b2989d01e54d80c4e023c6:upstream-provider-configs:temperature-0-where-supported:max-output-2048:reasoning-disabled-or-minimum-where-required:official-modality-weights Source record →

Independent operational reference

Artificial Analysis details

No exact, reviewed Artificial Analysis operational record is attached to this family. We do not inherit price, release, or throughput values from a nearby model name.

Early-user signal

Initial community opinions

Anecdotal · never scored

No editor-reviewed community-opinions paragraph is active for this model. An active paragraph requires at least two retained Reddit discussions from the minimum 14-day launch window—implemented exactly as release date through day 14, end-exclusive—plus paraphrase-only review and a sealed publication artifact. Community reports stay separate from every benchmark score and model comparison.

Independent early-test signal

Early technical field tests

X · anecdotal · never scored

No independent field-test write-up has cleared review for this model yet — insufficient independent tests.

Business Skills V3 · proposed

How Spring Prompt plans to test GPT-5.4 Mini

The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.

Proposed reasoning
Provider-default reasoning
Provider revision
Exact provider revision will be resolved and frozen only after preflight and execution approval
Tools and service
No tools · Provider-default service tier
Planning configuration
gpt-5.4-mini
Generation controls
Provider-managed reasoning; no temperature override proposed

Future first-party coverage

15 proposed Business Skills V3 task areas

These links describe evaluation contracts, not published GPT-5.4 Mini results.

Model comparisons

Compare GPT-5.4 Mini side by side

Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.

Publication safeguards

What must pass before a V3 result appears

  1. 1First-party task-local comparisons and eligible external evidence must both be present.
  2. 2Model and provider configuration identity must match the reviewed release exactly.
  3. 3Coverage, reliability, judge diagnostics, and sealed stability checks must pass.
  4. 4Uncertainty and missing evidence remain visible when results are published.

Stable family URL

Evidence can grow without changing the page

New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.

More from OpenAI

Other owned model profiles