AI model benchmarks
What each benchmark measures, who measured it, and how today's models do. For one score across all of them, see the overall leaderboard; for launches and pace, models over time.
Measured by Spring Prompt
- BulletBenchWhen thinking time comes off the clock, which models are quick enough to still make good decisions?
- CatalogBenchCan a model turn a product feed and photos into a listing that is ready to publish, without making things up?
- DeckBenchCan a model turn an analysis into a deck you could present without fixing it?
- ROASBenchGiven a year of ad spend decisions, which models grow revenue and which ones burn the budget?
Reported by others
- APEX-AgentsShare of investment banking, consulting and corporate law tasks an agent completes in a simulated workplace with files and apps, graded against expert criteria (one attempt).
- Arena (formerly LMArena)Head-to-head human preference across all text prompts.
- Artificial AnalysisArtificial Analysis's blend of its reasoning, knowledge, coding and agent evaluations.
- Berkeley Function Calling Leaderboard (BFCL) V4BFCL's own weighted average across its test categories.
- GDP.pdfRubric score on 100 real professional tasks that need reasoning over PDFs (finance, legal, insurance, engineering, HR and more).
- Microsoft STATE-BenchShare of customer support tasks completed correctly, averaged over repeated runs.
- OpenHands IndexThe equally weighted average of the five category scores below.
- Remote Labor IndexShare of 240 real paid freelance projects (design, software, data, architecture and more) an agent delivers at a standard a client would accept.
- SimpleQA Verified (Epoch AI)Share of short factual questions answered correctly without search, as run by Epoch AI.
- tau2-benchShare of airline bookings and changes tasks completed within policy, averaged over trials.
- UGI LeaderboardUGI's blend of intelligence, style, repetition and length adherence in writing, tuned to average human preference.
- Vectara Hallucination LeaderboardShare of document summaries that contain something the document does not support.
- Vending-Bench 2Bank balance after an agent runs a simulated vending business for a year: ordering stock, setting prices and dealing with suppliers.
Quick answers
Which benchmarks does Spring Prompt run itself?
BulletBench, CatalogBench, DeckBench and ROASBench. We design and run these on business work, such as product listings, decks and ad budgets, and publish every model's output with the scores.
Where do the other results come from?
13 leaderboards run by other teams, including APEX-Agents, Arena (formerly LMArena) and Artificial Analysis. Each page links to the original and shows the dates of its results.
Which AI benchmark should I look at for my task?
Pick the one closest to the work: for customer support agents, tau2-bench and Microsoft STATE-Bench; for agents doing office and professional work, APEX-Agents and Remote Labor Index; for coding agents, OpenHands Index; for tool and function calling, Berkeley Function Calling Leaderboard (BFCL) V4; for made-up facts in summaries, Vectara Hallucination Leaderboard. For one score across all of them, see the overall leaderboard.