1,439
Overall · rank 28 of 44
- Unit
- Arena rating, higher is better
- Range
- 1,430 to 1,449
- Sample
- 7158 votes
- Configuration
- gemini-3-flash
- Measured
- 13 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,191
Overall · rank 8 of 34
- Unit
- Arena rating, higher is better
- Range
- 1,187 to 1,196
- Sample
- 149334 votes
- Configuration
- gemini-3-flash
- Measured
- 24 Aug 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,468
Overall · rank 22 of 177
- Unit
- Arena rating, higher is better
- Range
- 1,464 to 1,472
- Sample
- 30632 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,451
Overall · rank 47 of 177
- Unit
- Arena rating, higher is better
- Range
- 1,449 to 1,454
- Sample
- 89978 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,469
Business, management and finance · rank 18 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,461 to 1,477
- Sample
- 5763 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,455
Business, management and finance · rank 40 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,449 to 1,460
- Sample
- 18138 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,457
Creative writing · rank 12 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,447 to 1,466
- Sample
- 4825 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,447
Creative writing · rank 24 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,441 to 1,453
- Sample
- 14997 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,496
Expert prompts · rank 14 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,484 to 1,509
- Sample
- 2177 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,463
Expert prompts · rank 66 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,456 to 1,470
- Sample
- 8554 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,458
Instruction following · rank 31 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,451 to 1,465
- Sample
- 8494 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,443
Instruction following · rank 55 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,438 to 1,448
- Sample
- 29816 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,473
Overall · rank 22 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,468 to 1,477
- Sample
- 31283 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,458
Overall · rank 49 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,455 to 1,461
- Sample
- 90749 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,461
Writing, literature and language · rank 15 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,453 to 1,469
- Sample
- 7058 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,448
Writing, literature and language · rank 30 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,443 to 1,453
- Sample
- 21848 votes
- Configuration
- gemini-3-flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
49.0%
Average · rank 20 of 34
- Unit
- % resolved, higher is better
- Configuration
- gemini-3-flash
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
22.1%
Front end (SWE-Bench Multimodal) · rank 30 of 34
- Unit
- % resolved, higher is better
- Configuration
- gemini-3-flash
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
18.8%
Greenfield (Commit0) · rank 21 of 34
- Unit
- % resolved, higher is better
- Configuration
- gemini-3-flash
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
58.8%
Information gathering (GAIA) · rank 19 of 34
- Unit
- % resolved, higher is better
- Configuration
- gemini-3-flash
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
74.6%
Issue resolution (SWE-Bench) · rank 13 of 34
- Unit
- % resolved, higher is better
- Configuration
- gemini-3-flash
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
70.7%
Testing (SWT-Bench) · rank 11 of 34
- Unit
- % resolved, higher is better
- Configuration
- gemini-3-flash
- Measured
- 30 Jun 2026
- Not shown
- Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
68.0%
Consistency (pass^4) · rank 4 of 8
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-flash
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
82.5%
Task success (pass^1) · rank 3 of 8
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-flash
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
7.2%
Consistency (pass^4) · rank 2 of 3
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-flash
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
27.3%
Task success (pass^1) · rank 1 of 3
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-flash
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
51.8%
Consistency (pass^4) · rank 2 of 8
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-flash
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
76.8%
Task success (pass^1) · rank 4 of 8
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-flash
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
70.2%
Consistency (pass^4) · rank 5 of 8
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-flash
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
91.2%
Task success (pass^1) · rank 3 of 8
- Unit
- % of tasks, higher is better
- Configuration
- gemini-3-flash
- Measured
- 2 Mar 2026
- Not shown
- Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.