MiniMax-M3.1-Flash-Preview
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
MiniMax-M3.1-Flash-Preview · max
- Backend & testing
- 75
- Frontend & interaction
- 57
- Knowledge & reasoning
- 75
- Evaluation time
- 54m 7s
- Reference cost
- —
Totals for all three evaluations at this configuration; cost estimated from API usage.
Overall rank 19 / 33 models · Best tested effort
Backend tested · 9/28/2026
Published summary scores use the same effort. Axes may be evaluated on different dates; a backend date refers only to backend testing.
View task coverage and scoring →What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
- Overall score
- — — → 69.6
- Reference cost
- — — → —
- Evaluation time
- — — → 54m 7s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| low | — | 30 | — | — | 3m 21s | — | |
| medium | — | 34 | — | — | 4m 28s | — | |
| high | — | 65 | — | 75 | 28m 29s | — | |
| xhigh | — | 50 | 15 | — | 17m 55s | — | |
| max | 69.6 | 75 | 57 | 75 | 54m 7s | — |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Data snapshot published 10/1/2026