glm-5.3
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score70.0
glm-5.3 · max
- Backend & testing
- 64
- Frontend & interaction
- 78
- Knowledge & reasoning
- 70
What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
high max
- Overall score
- +4.4 pts 65.6 → 70.0
- Reference cost
- −$0.22 $2.13 → $1.91
- Evaluation time
- +148.6% 63m 2s → 156m 42s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| high | 65.6 | 65 | 82 | 50 | 63m 2s | $2.13 | |
| max | 70.0 | 64 | 78 | 70 | 156m 42s | $1.91 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026