Artificial Analysis
Artificial Analysis's long-context benchmark for answering questions that require reasoning across multiple documents.
Leader
Kimi K3 · Published configuration
75%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
The share of all AA-Omniscience questions answered correctly, including questions where a model abstains in the denominator.
Leader
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
61%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
A factual-reliability score that rewards correct answers, penalizes hallucinated answers, and leaves abstentions neutral.
Leader
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
40
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
One minus the AA-Omniscience hallucination rate, where incorrect answers are divided by incorrect, partial, and not-attempted outcomes.
Leader
MiniCPM5-1B (Non-reasoning)
99%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
Artificial Analysis's Stirrup-based implementation of long-horizon professional-services tasks across realistic workplace tools.
Leader
Gemini 3.5 Flash (high)
47%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
ARC Prize Foundation
An interactive abstract-reasoning benchmark that evaluates exploration, adaptation, planning, and action efficiency.
Leader
Claude Opus 5 · High
30.16%
Read the explainer →
ARC-AGI-3 (2026)
Artificial Analysis
A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.
Leader
Claude Opus 5 · Adaptive Reasoning, Max Effort
61
Read the explainer →
Intelligence Index v4.1 · 2026-07-30
Artificial Analysis
A research-level physics benchmark using composite reasoning challenges and executable or symbolic answer formats.
Leader
GPT-5.6 Sol · max
32%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.
Leader
Claude Opus 5 · Adaptive Reasoning, Max Effort
68%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.
Leader
GPT-5.6 Sol · max
94%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
A broad expert-level benchmark of difficult academic reasoning and knowledge questions.
Leader
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
53%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.
Leader
Grok 4.3 (medium)
83%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
Artificial Analysis's implementation of Kubernetes incident root-cause analysis from offline SRE snapshots.
Leader
GPT-5.6 Sol · max
56%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.
Leader
Claude Opus 5 · Adaptive Reasoning, Max Effort
85%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
A scientific-programming benchmark requiring Python solutions to research-oriented computational problems.
Leader
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
60%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
A legacy terminal-use evaluation of agentic coding, system administration, data processing, and related command-line tasks.
Leader
GPT-5.6 Sol · max
66%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.
Leader
GPT-5.6 Sol (xhigh)
90%
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
A legacy agentic tool-use benchmark for conversational work in a telecom environment.
Read the explainer →
Public leaderboard snapshot · 2026-07-30
Artificial Analysis
An agentic banking customer-support benchmark for policy and knowledge retrieval, reasoning, dialogue, and chained tool calls.
Leader
Kimi K3 · Published configuration
33%
Read the explainer →
Public leaderboard snapshot · 2026-07-30