Models / Grok 4.6

xAI

Grok 4.6

15 published results from 2 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
xAI
Sources
2
Our benchmarks
0
Price
Not yet published

Reported by others

-0.03
Confirmed task success · rank 23 of 46
Unit
IPS effect estimate, higher is better
Range
-0.06 to -0.01
Sample
19134 observations
Configuration
Grok 4.6
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.00
Praise over complaint · rank 15 of 46
Unit
IPS effect estimate, higher is better
Range
-0.03 to 0.03
Sample
7955 observations
Configuration
Grok 4.6
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.03
Steerability · rank 6 of 46
Unit
IPS effect estimate, higher is better
Range
0.01 to 0.05
Sample
23132 observations
Configuration
Grok 4.6
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
2685682 observations
Configuration
Grok 4.6
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,461
Overall · rank 12 of 44
Unit
Arena rating, higher is better
Range
1,449 to 1,473
Sample
2258 votes
Configuration
Grok 4.6 (high)
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,449
Overall · rank 48 of 177
Unit
Arena rating, higher is better
Range
1,445 to 1,454
Sample
21035 votes
Configuration
Grok 4.6 (high)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,447
Business, management and finance · rank 47 of 402
Unit
Arena rating, higher is better
Range
1,438 to 1,457
Sample
4164 votes
Configuration
Grok 4.6 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,445
Creative writing · rank 24 of 407
Unit
Arena rating, higher is better
Range
1,435 to 1,454
Sample
4559 votes
Configuration
Grok 4.6 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,490
Expert prompts · rank 21 of 359
Unit
Arena rating, higher is better
Range
1,477 to 1,502
Sample
2489 votes
Configuration
Grok 4.6 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,451
Instruction following · rank 38 of 409
Unit
Arena rating, higher is better
Range
1,443 to 1,458
Sample
7687 votes
Configuration
Grok 4.6 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,453
Overall · rank 53 of 409
Unit
Arena rating, higher is better
Range
1,448 to 1,459
Sample
21380 votes
Configuration
Grok 4.6 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,446
Writing, literature and language · rank 30 of 408
Unit
Arena rating, higher is better
Range
1,437 to 1,454
Sample
5980 votes
Configuration
Grok 4.6 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by UGI Leaderboard
25.0%
Requested-length error · rank 228 of 370
Unit
% off the requested word count, lower is better
Configuration
Grok 4.6
Measured
3 Sep 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.36
Style adherence · rank 140 of 370
Unit
score from 0 to 1, higher is better
Configuration
Grok 4.6
Measured
3 Sep 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
63.3
Writing score · rank 80 of 370
Unit
score out of 100, higher is better
Configuration
Grok 4.6
Measured
3 Sep 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Grok 4.6 with