- Published
- Aug 26, 2026, 11:00 PM UTC
- Question pack
- coding-fast-v4.10
- Coverage
- 2 measured configurations
- Batch
- snapshot-2026-08-26T23-00-00Z-r3
Measured configurations
| Configuration | Rank | Score | Elapsed | Reference cost | Route |
|---|---|---|---|---|---|
| GLM-5.3 / High | #8 | 68/100 | 23m 35s | $0.42 | Custom endpoint |
| GLM-5.3 / Max | #12 | 66/100 | 25m 6s | $0.44 | Custom endpoint |
Five-question profile
| Question | GLM-5.3 / High | GLM-5.3 / Max |
|---|---|---|
| Black-box Regression Audit | 13/20 | 10/20 |
| Retry Planner Counterexamples | 16/20 | 14/20 |
| CI Adversarial Audit | 12/20 | 14/20 |
| Transaction Regression Design | 15/20 | 15/20 |
| Cache Regression Test Design | 12/20 | 13/20 |
Read High and Max as route-specific evidence
When both efforts appear in the same batch, their score and elapsed time are directly inspectable under the recorded protocol. A higher effort label still does not guarantee a higher score in every batch.
A small gap should be checked against the five-question profile and repeated comparable batches before it changes a default configuration.