Confirm Action

Are you sure you want to proceed?

Model intelligence profile

Qwen3.7 Max benchmark results and model details

This family profile brings together 10 published benchmark records for Alibaba Qwen3.7 Max. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.

Published benchmark evidence available · Business Skills V3 not yet authorized

Provider

Alibaba

Base model family

Release date

Not yet verified

Exact-family source only

Published benchmarks

10

Native scales kept separate

Artificial Analysis

No exact match

As of 2026-07-16

Benchmark-level evidence

Published Qwen3.7 Max benchmark results

These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.

How sources are reviewed →

Spring Prompt benchmark

BulletBench

2 configurations

60-second bullet Elo

Chess decision quality under a shared 60-second game clock and a 10+1 per-move clock.

Ladder Elo: anchored to a Stockfish skill ladder (random mover = 400). Internally consistent ordering, not FIDE-calibrated.

2 published configurations

Qwen 3.7 Max (off)

539

Bullet 95% interval
475–590
10+1 Elo
539
Bullet games
96
Bullet median move time
1.11s
Average bullet game cost
$0.0243

Qwen 3.7 Max (default/on)

0

Bullet 95% interval
Floor-clamped / unavailable
10+1 Elo
0
Bullet games
96
Bullet median move time
12.73s
Average bullet game cost
$0.0185
Source: Spring Prompt Recorded 3 Aug 2026 Ladder v1 · controls 10+1 and 60 Source record →

Spring Prompt benchmark

ROASBench

33.56

Average ROASBench score

A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.

1 published configuration

Qwen: Qwen3.7 Max

33.56

Average contribution profit
$131,537
Average ROAS
1.54×
Completed runs
3
Score variability
±2.40
Source: Spring Prompt Recorded 25 Jul 2026 Published ROASBench cache · 12-month simulation Source record →

Spring Prompt benchmark

World Cup prediction benchmark

288

Tournament prediction points

Fixture-by-fixture football predictions frozen before the tournament and graded on outcomes, scorelines, goals, penalties, and cards.

32 fixtures graded from the frozen prediction artifact.

1 published configuration

Qwen 3.7 Max

288

Graded matches
32
Exact-score points
12
Result points
34
Goal points
69
Source: Spring Prompt Recorded 20 Jul 2026 Predictions frozen 2026-06-11 Source record →

source-native benchmark

Artificial Analysis Intelligence Index

46

A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.

1 published configuration

Qwen3.7 Max

46

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Intelligence Index v4.1 · 2026-07-30 Source record →

source-native benchmark

GDPval-AA v2

39%

Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.

1 published configuration

Qwen3.7 Max

39%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

GPQA Diamond

92%

The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.

1 published configuration

Qwen3.7 Max

92%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

Humanity's Last Exam

38%

A broad expert-level benchmark of difficult academic reasoning and knowledge questions.

1 published configuration

Qwen3.7 Max

38%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

IFBench

81%

An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.

1 published configuration

Qwen3.7 Max

81%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

Terminal-Bench v2.1

75%

A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.

1 published configuration

Qwen3.7 Max

75%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

τ²-Bench Telecom

95%

A legacy agentic tool-use benchmark for conversational work in a telecom environment.

1 published configuration

Qwen3.7 Max

95%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

Independent operational reference

Artificial Analysis details

No exact, reviewed Artificial Analysis operational record is attached to this family. We do not inherit price, release, or throughput values from a nearby model name.

Early-user signal

Initial community opinions

Anecdotal · never scored

No editor-reviewed community-opinions paragraph is active for this model. An active paragraph requires at least two retained Reddit discussions from the minimum 14-day launch window—implemented exactly as release date through day 14, end-exclusive—plus paraphrase-only review and a sealed publication artifact. Community reports stay separate from every benchmark score and model comparison.

Independent early-test signal

Early technical field tests

X · anecdotal · never scored

An editorial paraphrase of 2 launch-window field tests from 2 independent authors on X, including 2 reports with a described method or inspectable artifact. The window runs from 21 May 2026 up to 4 Jun 2026, and exact model identity was editor reviewed.

Early Qwen3.7 Max evidence covered enterprise site-reliability work and hands-on coding use. An enterprise IT benchmark evaluated diagnosis across Kubernetes incidents using alerts, traces, logs, and topology, while a separate tester reported successful use through a command-based Qoder workflow. The two reports suggest useful agent and coding ability, but one is benchmark-driven and the other offers limited failure detail.

These reports are selectively surfaced and are not a representative sample. They never affect benchmark scores, rankings, winners, or comparison outcomes.

Business Skills V3 · proposed

How Spring Prompt plans to test Qwen3.7 Max

The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.

Proposed reasoning
High reasoning
Provider revision
Exact provider revision will be resolved and frozen only after preflight and execution approval
Tools and service
No tools · Provider-default service tier
Planning configuration
qwen3.7-max-high
Generation controls
Provider-managed reasoning; no temperature override proposed

Future first-party coverage

15 proposed Business Skills V3 task areas

These links describe evaluation contracts, not published Qwen3.7 Max results.

Model comparisons

Compare Qwen3.7 Max side by side

Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.

Publication safeguards

What must pass before a V3 result appears

  1. 1First-party task-local comparisons and eligible external evidence must both be present.
  2. 2Model and provider configuration identity must match the reviewed release exactly.
  3. 3Coverage, reliability, judge diagnostics, and sealed stability checks must pass.
  4. 4Uncertainty and missing evidence remain visible when results are published.

Stable family URL

Evidence can grow without changing the page

New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.

More from Alibaba

Other owned model profiles