glm-5.3
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
glm-5.3 · max
- Backend & testing
- 65
- Frontend & interaction
- 78
- Knowledge & reasoning
- 70
- Evaluation time
- 157m 1s
- Reference cost
- $1.83
Totals for all three evaluations at this configuration; cost estimated from API usage.
Overall rank 18 / 33 models · Best tested effort
Backend tested · 10/1/2026
Published summary scores use the same effort. Axes may be evaluated on different dates; a backend date refers only to backend testing.
View task coverage and scoring →What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
- Overall score
- +8.4 pts 62.0 → 70.4
- Reference cost
- −$0.43 $2.26 → $1.83
- Evaluation time
- +130.9% 68m 0s → 157m 1s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| high | 62.0 | 56 | 82 | 50 | 68m 0s | $2.26 | |
| max | 70.4 | 65 | 78 | 70 | 157m 1s | $1.83 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Data snapshot published 10/1/2026