- Published
- Sep 16, 2026, 06:28 AM UTC
- Question pack
- coding-fast-v4.12
- Tested
- 2 measured configurations
- Batch
- evaluation-497e806be8b8470345169c1aad6752d74a0c1ce9e195a9eb7fcc02b628701c77
Results from this batch
| Measure | Kimi K3 / High | GLM-5.3 / High |
|---|---|---|
| Rank | #21 | #25 |
| Score | 65/100 | 62/100 |
| Elapsed | 37m 56s | 38m 6s |
| Reference cost | $0.89 | $0.46 |
| Route | Custom endpoint | Custom endpoint |
| Black-box Regression Audit | 16/20 | 10/20 |
| Retry Planner Counterexamples | 15/20 | 13/20 |
| CI Adversarial Audit | 14/20 | 14/20 |
| Transaction Regression Design | 12/20 | 15/20 |
| Cache Propagation Certificate | 8/20 | 10/20 |
How to choose from these results
When the totals are close, which model does better on the skills your work needs?
Start with the score and question profile. Use elapsed time only after both configurations clear your quality floor, and leave cost undecided when either side is not covered.