Qwen 3.8 Max
One configuration tested. Explore its results and compare with other models.
Highest overall score58.3
Qwen 3.8 Max · default
- Backend & testing
- 64
- Frontend & interaction
- 69
- Knowledge & reasoning
- 40
- Evaluation time
- 133m 3s
- Reference cost
- $3.41Partial cost
Totals for all three evaluations at this configuration; cost estimated from API usage.
Joint overall rank 27 / 33 models · Best tested effort
Backend tested · 9/10/2026
Published summary scores use the same effort. Axes may be evaluated on different dates; a backend date refers only to backend testing.
View task coverage and scoring →Tested configuration
Scores, evaluation time and reference cost for this configuration.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| default | 58.3 | 64 | 69 | 40 | 133m 3s | $3.41Partial cost |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Data snapshot published 10/1/2026