- Published
- Sep 16, 2026, 06:28 AM UTC
- Question pack
- coding-fast-v4.12
- Tested
- 2 measured configurations
- Batch
- evaluation-497e806be8b8470345169c1aad6752d74a0c1ce9e195a9eb7fcc02b628701c77
Measured configurations
| Configuration | Rank | Score | Elapsed | Reference cost | Route |
|---|---|---|---|---|---|
| GLM-5.3 / High | #25 | 62/100 | 38m 6s | $0.46 | Custom endpoint |
| GLM-5.3 / Max | #26 | 60/100 | 35m 52s | $0.41 | Custom endpoint |
Five-question profile
| Question | GLM-5.3 / High | GLM-5.3 / Max |
|---|---|---|
| Black-box Regression Audit | 10/20 | 10/20 |
| Retry Planner Counterexamples | 13/20 | 11/20 |
| CI Adversarial Audit | 14/20 | 13/20 |
| Transaction Regression Design | 15/20 | 15/20 |
| Cache Propagation Certificate | 10/20 | 11/20 |
Read High and Max as route-specific evidence
When both efforts appear in the same batch, their score and elapsed time are directly inspectable under the recorded protocol. A higher effort label still does not guarantee a higher score in every batch.
A small gap should be checked against the five-question profile and repeated comparable batches before it changes a default configuration.