- Published
- Sep 16, 2026, 06:28 AM UTC
- Question pack
- coding-fast-v4.12
- Tested
- 1 measured configuration
- Batch
- evaluation-497e806be8b8470345169c1aad6752d74a0c1ce9e195a9eb7fcc02b628701c77
Measured configurations
| Configuration | Rank | Score | Elapsed | Reference cost | Route |
|---|---|---|---|---|---|
| Kimi K3 / High | #21 | 65/100 | 37m 56s | $0.89 | Custom endpoint |
Five-question profile
| Question | Kimi K3 / High |
|---|---|
| Black-box Regression Audit | 16/20 |
| Retry Planner Counterexamples | 15/20 |
| CI Adversarial Audit | 14/20 |
| Transaction Regression Design | 12/20 |
| Cache Propagation Certificate | 8/20 |
What one measured configuration can answer
A single Kimi K3 configuration can be compared with other entries in the same batch. It cannot show whether another K3 effort or route would behave the same way.
When the total score is close to another model, open the five-question profile. Similar totals can come from different strengths and weaknesses.