Models / MiniMax M2.1

MiniMax

MiniMax M2.1

11 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
MiniMax
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

Reported by OpenHands Index
41.2%
Average · rank 27 of 34
Unit
% resolved, higher is better
Configuration
minimax-m2.1
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
16.2%
Front end (SWE-Bench Multimodal) · rank 34 of 34
Unit
% resolved, higher is better
Configuration
minimax-m2.1
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
18.8%
Greenfield (Commit0) · rank 21 of 34
Unit
% resolved, higher is better
Configuration
minimax-m2.1
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
40.6%
Information gathering (GAIA) · rank 27 of 34
Unit
% resolved, higher is better
Configuration
minimax-m2.1
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
68.8%
Issue resolution (SWE-Bench) · rank 27 of 34
Unit
% resolved, higher is better
Configuration
minimax-m2.1
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
61.4%
Testing (SWT-Bench) · rank 23 of 34
Unit
% resolved, higher is better
Configuration
minimax-m2.1
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by UGI Leaderboard
31.0%
Requested-length error · rank 258 of 370
Unit
% off the requested word count, lower is better
Configuration
MiniMax M2.1
Measured
28 Jan 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.41
Style adherence · rank 13 of 370
Unit
score from 0 to 1, higher is better
Configuration
MiniMax M2.1
Measured
28 Jan 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
41.6
Writing score · rank 201 of 370
Unit
score out of 100, higher is better
Configuration
MiniMax M2.1
Measured
28 Jan 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
98.5%
Answer rate · rank 78 of 108
Unit
% of documents, higher is better
Configuration
MiniMax M2.1
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
11.8%
Hallucination rate · rank 76 of 108
Unit
% of summaries, lower is better
Configuration
MiniMax M2.1
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare MiniMax M2.1 with