Models / gemini-3-pro

Google

gemini-3-pro

23 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Google
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

1,451
Overall · rank 18 of 44
Unit
Arena rating, higher is better
Range
1,442 to 1,460
Sample
10769 votes
Configuration
gemini-3-pro
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,201
Overall · rank 5 of 34
Unit
Arena rating, higher is better
Range
1,196 to 1,207
Sample
37024 votes
Configuration
gemini-3-pro
Measured
24 Aug 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,481
Overall · rank 4 of 177
Unit
Arena rating, higher is better
Range
1,478 to 1,485
Sample
40987 votes
Configuration
gemini-3-pro
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,474
Business, management and finance · rank 11 of 402
Unit
Arena rating, higher is better
Range
1,467 to 1,481
Sample
7727 votes
Configuration
gemini-3-pro
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,484
Creative writing · rank 3 of 407
Unit
Arena rating, higher is better
Range
1,476 to 1,492
Sample
6510 votes
Configuration
gemini-3-pro
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,498
Expert prompts · rank 14 of 359
Unit
Arena rating, higher is better
Range
1,487 to 1,510
Sample
2814 votes
Configuration
gemini-3-pro
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,473
Instruction following · rank 15 of 409
Unit
Arena rating, higher is better
Range
1,466 to 1,479
Sample
11653 votes
Configuration
gemini-3-pro
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,485
Overall · rank 8 of 409
Unit
Arena rating, higher is better
Range
1,482 to 1,489
Sample
41921 votes
Configuration
gemini-3-pro
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,481
Writing, literature and language · rank 5 of 408
Unit
Arena rating, higher is better
Range
1,474 to 1,488
Sample
9527 votes
Configuration
gemini-3-pro
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by OpenHands Index
49.0%
Average · rank 19 of 34
Unit
% resolved, higher is better
Configuration
gemini-3-pro
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
36.8%
Front end (SWE-Bench Multimodal) · rank 11 of 34
Unit
% resolved, higher is better
Configuration
gemini-3-pro
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
25.0%
Greenfield (Commit0) · rank 14 of 34
Unit
% resolved, higher is better
Configuration
gemini-3-pro
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
44.2%
Information gathering (GAIA) · rank 25 of 34
Unit
% resolved, higher is better
Configuration
gemini-3-pro
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
70.6%
Issue resolution (SWE-Bench) · rank 25 of 34
Unit
% resolved, higher is better
Configuration
gemini-3-pro
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
68.6%
Testing (SWT-Bench) · rank 17 of 34
Unit
% resolved, higher is better
Configuration
gemini-3-pro
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by tau2-bench
66.0%
Consistency (pass^4) · rank 6 of 8
Unit
% of tasks, higher is better
Configuration
gemini-3-pro
Measured
2 Mar 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
80.5%
Task success (pass^1) · rank 6 of 8
Unit
% of tasks, higher is better
Configuration
gemini-3-pro
Measured
2 Mar 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
4.1%
Consistency (pass^4) · rank 3 of 3
Unit
% of tasks, higher is better
Configuration
gemini-3-pro
Measured
2 Mar 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
18.0%
Task success (pass^1) · rank 3 of 3
Unit
% of tasks, higher is better
Configuration
gemini-3-pro
Measured
2 Mar 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
47.4%
Consistency (pass^4) · rank 5 of 8
Unit
% of tasks, higher is better
Configuration
gemini-3-pro
Measured
2 Mar 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
75.9%
Task success (pass^1) · rank 5 of 8
Unit
% of tasks, higher is better
Configuration
gemini-3-pro
Measured
2 Mar 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
74.6%
Consistency (pass^4) · rank 3 of 8
Unit
% of tasks, higher is better
Configuration
gemini-3-pro
Measured
2 Mar 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
91.0%
Task success (pass^1) · rank 4 of 8
Unit
% of tasks, higher is better
Configuration
gemini-3-pro
Measured
2 Mar 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.

Compare gemini-3-pro with