Close-score comparison

Kimi K3 vs GLM-5.3

When totals are close, the five-question profile and repeated batches matter more than declaring a winner from one point.

Current batch

GLM-5.3 / High leads the selected batch by 7 points. GLM-5.3 / High completed 4m 11s sooner. Both configurations have reference-cost coverage.

Selected evidence
Published
Aug 26, 2026, 11:00 PM UTC
Question pack
coding-fast-v4.10
Coverage
2 measured configurations
Batch
snapshot-2026-08-26T23-00-00Z-r3

Same-batch evidence

MeasureKimi K3 / HighGLM-5.3 / High
Rank#21#8
Score61/10068/100
Elapsed27m 47s23m 35s
Reference cost$0.44$0.42
RouteCustom endpointCustom endpoint
Black-box Regression Audit9/2013/20
Retry Planner Counterexamples13/2016/20
CI Adversarial Audit12/2012/20
Transaction Regression Design14/2015/20
Cache Regression Test Design13/2012/20

The decision this comparison can support

Do the two configurations fail on the same capabilities, or does a similar total hide different operational risks?

Start with the score and question profile. Use elapsed time only after both configurations clear your quality floor, and leave cost undecided when either side is not covered.