Confirm Action

Are you sure you want to proceed?

External benchmarks

Understand the evaluation, not just the score

Benchmark results describe a deployed evaluation system: model, harness, tools, budget, and scoring protocol. These reviewed explainers make that contract visible before comparing numbers.

Published benchmark explainers

Artificial Analysis

AA-LCR

Artificial Analysis's long-context benchmark for answering questions that require reasoning across multiple documents.

Leader Kimi K3 · Published configuration

75%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

AA-Omniscience Accuracy

The share of all AA-Omniscience questions answered correctly, including questions where a model abstains in the denominator.

Leader Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)

61%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

AA-Omniscience Index

A factual-reliability score that rewards correct answers, penalizes hallucinated answers, and leaves abstentions neutral.

Leader Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)

40

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

AA-Omniscience Non-Hallucination Rate

One minus the AA-Omniscience hallucination rate, where incorrect answers are divided by incorrect, partial, and not-attempted outcomes.

Leader MiniCPM5-1B (Non-reasoning)

99%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

APEX-Agents-AA

Artificial Analysis's Stirrup-based implementation of long-horizon professional-services tasks across realistic workplace tools.

Leader Gemini 3.5 Flash (high)

47%

Read the explainer → Public leaderboard snapshot · 2026-07-30

ARC Prize Foundation

ARC-AGI-3

An interactive abstract-reasoning benchmark that evaluates exploration, adaptation, planning, and action efficiency.

Leader Claude Opus 5 · High

30.16%

Read the explainer → ARC-AGI-3 (2026)

Artificial Analysis

Artificial Analysis Intelligence Index

A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.

Leader Claude Opus 5 · Adaptive Reasoning, Max Effort

61

Read the explainer → Intelligence Index v4.1 · 2026-07-30

Artificial Analysis

CritPt

A research-level physics benchmark using composite reasoning challenges and executable or symbolic answer formats.

Leader GPT-5.6 Sol · max

32%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

GDPval-AA v2

Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.

Leader Claude Opus 5 · Adaptive Reasoning, Max Effort

68%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

GPQA Diamond

The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.

Leader GPT-5.6 Sol · max

94%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

Humanity's Last Exam

A broad expert-level benchmark of difficult academic reasoning and knowledge questions.

Leader Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)

53%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

IFBench

An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.

Leader Grok 4.3 (medium)

83%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

ITBench-AA

Artificial Analysis's implementation of Kubernetes incident root-cause analysis from offline SRE snapshots.

Leader GPT-5.6 Sol · max

56%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

MMMU-Pro

A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.

Leader Claude Opus 5 · Adaptive Reasoning, Max Effort

85%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

SciCode

A scientific-programming benchmark requiring Python solutions to research-oriented computational problems.

Leader Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)

60%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

Terminal-Bench Hard

A legacy terminal-use evaluation of agentic coding, system administration, data processing, and related command-line tasks.

Leader GPT-5.6 Sol · max

66%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

Terminal-Bench v2.1

A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.

Leader GPT-5.6 Sol (xhigh)

90%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

τ²-Bench Telecom

A legacy agentic tool-use benchmark for conversational work in a telecom environment.

Leader GLM-5.2 (max)

99%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Artificial Analysis

τ³-Banking

An agentic banking customer-support benchmark for policy and knowledge retrieval, reasoning, dialogue, and chained tool calls.

Leader Kimi K3 · Published configuration

33%

Read the explainer → Public leaderboard snapshot · 2026-07-30

Where to go next