Models / Gemini 2.5 Pro

Google

Gemini 2.5 Pro

16 published results from 4 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Google
Sources
4
Our benchmarks
0
Price
Not yet published

Reported by others

1,431
Overall · rank 34 of 44
Unit
Arena rating, higher is better
Range
1,424 to 1,437
Sample
25110 votes
Configuration
Gemini 2.5 Pro
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,142
Overall · rank 27 of 34
Unit
Arena rating, higher is better
Range
1,137 to 1,147
Sample
83404 votes
Configuration
Gemini 2.5 Pro
Measured
24 Aug 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,462
Overall · rank 28 of 177
Unit
Arena rating, higher is better
Range
1,459 to 1,465
Sample
69736 votes
Configuration
Gemini 2.5 Pro
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,437
Business, management and finance · rank 69 of 402
Unit
Arena rating, higher is better
Range
1,433 to 1,442
Sample
22770 votes
Configuration
Gemini 2.5 Pro
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,443
Creative writing · rank 26 of 407
Unit
Arena rating, higher is better
Range
1,438 to 1,448
Sample
17590 votes
Configuration
Gemini 2.5 Pro
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,458
Expert prompts · rank 76 of 359
Unit
Arena rating, higher is better
Range
1,451 to 1,465
Sample
8260 votes
Configuration
Gemini 2.5 Pro
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,438
Instruction following · rank 61 of 409
Unit
Arena rating, higher is better
Range
1,434 to 1,442
Sample
35043 votes
Configuration
Gemini 2.5 Pro
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,446
Overall · rank 72 of 409
Unit
Arena rating, higher is better
Range
1,443 to 1,448
Sample
124887 votes
Configuration
Gemini 2.5 Pro
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,445
Writing, literature and language · rank 39 of 408
Unit
Arena rating, higher is better
Range
1,440 to 1,449
Sample
28118 votes
Configuration
Gemini 2.5 Pro
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by tau2-bench
1.0%
Consistency (pass^4) · rank 21 of 21
Unit
% of tasks, higher is better
Configuration
gemini-2.5-pro
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
13.7%
Task success (pass^1) · rank 20 of 21
Unit
% of tasks, higher is better
Configuration
gemini-2.5-pro
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by UGI Leaderboard
19.0%
Requested-length error · rank 163 of 370
Unit
% off the requested word count, lower is better
Configuration
Gemini 2.5 Pro
Measured
9 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.40
Style adherence · rank 29 of 370
Unit
score from 0 to 1, higher is better
Configuration
Gemini 2.5 Pro
Measured
9 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
65.2
Writing score · rank 63 of 370
Unit
score out of 100, higher is better
Configuration
Gemini 2.5 Pro
Measured
9 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
99.1%
Answer rate · rank 64 of 108
Unit
% of documents, higher is better
Configuration
Gemini 2.5 Pro
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
7.0%
Hallucination rate · rank 32 of 108
Unit
% of summaries, lower is better
Configuration
Gemini 2.5 Pro
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare Gemini 2.5 Pro with