Same-batch comparison

Qwen 3.8 vs Kimi K3

A same-batch comparison of the current measured configurations, with their route and missing-cost boundaries kept visible.

Current batch

Qwen 3.8 / Max leads the selected batch by 5 points. Qwen 3.8 / Max completed 6m 27s sooner. Both configurations have reference-cost coverage.

Selected evidence
Published
Aug 26, 2026, 11:00 PM UTC
Question pack
coding-fast-v4.10
Coverage
2 measured configurations
Batch
snapshot-2026-08-26T23-00-00Z-r3

Same-batch evidence

MeasureQwen 3.8 / MaxKimi K3 / High
Rank#10#21
Score66/10061/100
Elapsed21m 19s27m 47s
Reference cost$0.89$0.44
RouteCustom endpointCustom endpoint
Black-box Regression Audit10/209/20
Retry Planner Counterexamples17/2013/20
CI Adversarial Audit12/2012/20
Transaction Regression Design13/2014/20
Cache Regression Test Design14/2013/20

The decision this comparison can support

Does the measured quality difference remain useful after you account for route reliability, elapsed time, and your own task profile?

Start with the score and question profile. Use elapsed time only after both configurations clear your quality floor, and leave cost undecided when either side is not covered.