- Published
- Aug 26, 2026, 11:00 PM UTC
- Question pack
- coding-fast-v4.10
- Coverage
- 2 measured configurations
- Batch
- snapshot-2026-08-26T23-00-00Z-r3
Same-batch evidence
| Measure | Kimi K3 / High | GLM-5.3 / High |
|---|---|---|
| Rank | #21 | #8 |
| Score | 61/100 | 68/100 |
| Elapsed | 27m 47s | 23m 35s |
| Reference cost | $0.44 | $0.42 |
| Route | Custom endpoint | Custom endpoint |
| Black-box Regression Audit | 9/20 | 13/20 |
| Retry Planner Counterexamples | 13/20 | 16/20 |
| CI Adversarial Audit | 12/20 | 12/20 |
| Transaction Regression Design | 14/20 | 15/20 |
| Cache Regression Test Design | 13/20 | 12/20 |
The decision this comparison can support
Do the two configurations fail on the same capabilities, or does a similar total hide different operational risks?
Start with the score and question profile. Use elapsed time only after both configurations clear your quality floor, and leave cost undecided when either side is not covered.