Models / Inkling Small

Thinkingmachines

Inkling Small

22 published results from 3 sources. Each card shows where the number comes from and what it does not measure. The overall leaderboard combines them; here each stands alone.

Provider
Thinkingmachines
Sources
3
Our benchmarks
1
Price per million tokens
$0.45 in · $1.20 out
OpenRouter list price, 1 Oct 2026 · 524,288-token context

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
$0.0075
Cost per game · rank 7 of 23
Unit
US dollars, lower is better
Configuration
Inkling Small (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
0.0%
Games lost on time · rank 1 of 23
Unit
% of games, lower is better
Configuration
Inkling Small (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
7.2%
Invalid moves · rank 23 of 23
Unit
% of moves, lower is better
Configuration
Inkling Small (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
529
Ladder Elo · rank 3 of 23
Unit
ladder Elo, higher is better
Range
393 to 652
Sample
12 games
Configuration
Inkling Small (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.6 s
Median move time · rank 4 of 23
Unit
milliseconds, lower is better
Configuration
Inkling Small (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · BulletBench
$0.0077
Cost per game · rank 8 of 14
Unit
US dollars, lower is better
Configuration
Inkling Small (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
0.0%
Games lost on time · rank 1 of 14
Unit
% of games, lower is better
Configuration
Inkling Small (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
5.3%
Invalid moves · rank 14 of 14
Unit
% of moves, lower is better
Configuration
Inkling Small (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
571
Ladder Elo · rank 1 of 14
Unit
ladder Elo, higher is better
Range
428 to 698
Sample
12 games
Configuration
Inkling Small (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.6 s
Median move time · rank 5 of 14
Unit
milliseconds, lower is better
Configuration
Inkling Small (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not general response speed: replies are one short move.

Reported by others

-0.21
Confirmed task success · rank 43 of 46
Unit
IPS effect estimate, higher is better
Range
-0.25 to -0.17
Sample
7096 observations
Configuration
Inkling Small
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.21
Praise over complaint · rank 43 of 46
Unit
IPS effect estimate, higher is better
Range
-0.24 to -0.17
Sample
2752 observations
Configuration
Inkling Small
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.12
Steerability · rank 41 of 46
Unit
IPS effect estimate, higher is better
Range
-0.15 to -0.08
Sample
9247 observations
Configuration
Inkling Small
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.00
Tool grounding · rank 30 of 46
Unit
IPS effect estimate, higher is better
Range
-0.01 to -0.00
Sample
419271 observations
Configuration
Inkling Small
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,425
Overall · rank 105 of 177
Unit
Arena rating, higher is better
Range
1,421 to 1,429
Sample
24435 votes
Configuration
Inkling Small
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,408
Business, management and finance · rank 110 of 402
Unit
Arena rating, higher is better
Range
1,398 to 1,417
Sample
4737 votes
Configuration
Inkling Small
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,313
Creative writing · rank 187 of 407
Unit
Arena rating, higher is better
Range
1,304 to 1,323
Sample
5116 votes
Configuration
Inkling Small
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,445
Expert prompts · rank 89 of 359
Unit
Arena rating, higher is better
Range
1,433 to 1,456
Sample
2955 votes
Configuration
Inkling Small
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,393
Instruction following · rank 126 of 409
Unit
Arena rating, higher is better
Range
1,386 to 1,400
Sample
9156 votes
Configuration
Inkling Small
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,405
Overall · rank 133 of 409
Unit
Arena rating, higher is better
Range
1,400 to 1,410
Sample
24758 votes
Configuration
Inkling Small
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,344
Writing, literature and language · rank 174 of 408
Unit
Arena rating, higher is better
Range
1,335 to 1,352
Sample
6751 votes
Configuration
Inkling Small
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
19.1%
Correct answers · rank 73 of 85
Unit
% of questions, higher is better
Configuration
Inkling Small (xhigh)
Measured
27 Aug 2026
Not shown
Not answers grounded in your documents; tests what the model remembers.

Compare Inkling Small with