Close-score comparison

Kimi K3 vs GLM-5.3

When totals are close, the five-question profile and repeated batches matter more than declaring a winner from one point.

Current batch

Kimi K3 / High leads the selected batch by 3 points. Kimi K3 / High completed 9s sooner. Reference costs are available for both configurations.

Evaluation details
Published
Sep 16, 2026, 06:28 AM UTC
Question pack
coding-fast-v4.12
Tested
2 measured configurations
Batch
evaluation-497e806be8b8470345169c1aad6752d74a0c1ce9e195a9eb7fcc02b628701c77

Results from this batch

MeasureKimi K3 / HighGLM-5.3 / High
Rank#21#25
Score65/10062/100
Elapsed37m 56s38m 6s
Reference cost$0.89$0.46
RouteCustom endpointCustom endpoint
Black-box Regression Audit16/2010/20
Retry Planner Counterexamples15/2013/20
CI Adversarial Audit14/2014/20
Transaction Regression Design12/2015/20
Cache Propagation Certificate8/2010/20

How to choose from these results

When the totals are close, which model does better on the skills your work needs?

Start with the score and question profile. Use elapsed time only after both configurations clear your quality floor, and leave cost undecided when either side is not covered.