Same-batch comparison

Qwen 3.8 vs Kimi K3

Compare the highest-scoring Qwen 3.8 and Kimi K3 effort levels from the same test batch.

Current batch

Kimi K3 / High leads the selected batch by 5 points. Kimi K3 / High completed 9m 52s sooner. Reference costs are available for both configurations.

Evaluation details
Published
Sep 16, 2026, 06:28 AM UTC
Question pack
coding-fast-v4.12
Tested
2 measured configurations
Batch
evaluation-497e806be8b8470345169c1aad6752d74a0c1ce9e195a9eb7fcc02b628701c77

Results from this batch

MeasureQwen 3.8 / MaxKimi K3 / High
Rank#28#21
Score60/10065/100
Elapsed47m 48s37m 56s
Reference cost$1.07$0.89
RouteCustom endpointCustom endpoint
Black-box Regression Audit12/2016/20
Retry Planner Counterexamples17/2015/20
CI Adversarial Audit8/2014/20
Transaction Regression Design14/2012/20
Cache Propagation Certificate9/208/20

How to choose from these results

Does the measured quality difference remain useful after you account for route reliability, elapsed time, and your own task profile?

Start with the score and question profile. Use elapsed time only after both configurations clear your quality floor, and leave cost undecided when either side is not covered.