glm-5.3-flash
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score51.7
glm-5.3-flash · max
- Backend & testing
- 43
- Frontend & interaction
- 75
- Knowledge & reasoning
- 40
What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
high max
- Overall score
- +1.7 pts 50.0 → 51.7
- Reference cost
- — — → —
- Evaluation time
- +7.4% 90m 3s → 96m 41s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| high | 50.0 | 53 | 76 | 20 | 90m 3s | $0.084Partial cost | |
| max | 51.7 | 43 | 75 | 40 | 96m 41s | $0.074Partial cost |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026