AI model benchmarks

What each benchmark measures, who measured it, and how today's models do. For one score across all of them, see the overall leaderboard; for launches and pace, models over time.

Measured by Spring Prompt

Reported by others

Quick answers

Which benchmarks does Spring Prompt run itself?

BulletBench, CatalogBench, DeckBench and ROASBench. We design and run these on business work, such as product listings, decks and ad budgets, and publish every model's output with the scores.

Where do the other results come from?

13 leaderboards run by other teams, including APEX-Agents, Arena (formerly LMArena) and Artificial Analysis. Each page links to the original and shows the dates of its results.

Which AI benchmark should I look at for my task?

Pick the one closest to the work: for customer support agents, tau2-bench and Microsoft STATE-Bench; for agents doing office and professional work, APEX-Agents and Remote Labor Index; for coding agents, OpenHands Index; for tool and function calling, Berkeley Function Calling Leaderboard (BFCL) V4; for made-up facts in summaries, Vectara Hallucination Leaderboard. For one score across all of them, see the overall leaderboard.