Configuration comparison

GPT-5.6 Sol vs Qwen 3.8

A same-batch comparison of the highest-scoring measured configuration from each family, with route and missing-cost boundaries left visible.

Current batch

GPT-5.6 Sol / Max leads the selected batch by 19 points. Qwen 3.8 / Max completed 23m 39s sooner. Both configurations have reference-cost coverage.

Selected evidence
Published
Aug 26, 2026, 11:00 PM UTC
Question pack
coding-fast-v4.10
Coverage
2 measured configurations
Batch
snapshot-2026-08-26T23-00-00Z-r3

Same-batch evidence

MeasureGPT-5.6 Sol / MaxQwen 3.8 / Max
Rank#1#10
Score85/10066/100
Elapsed44m 58s21m 19s
Reference cost$1.82$0.89
RouteOfficial loginCustom endpoint
Black-box Regression Audit14/2010/20
Retry Planner Counterexamples20/2017/20
CI Adversarial Audit18/2012/20
Transaction Regression Design17/2013/20
Cache Regression Test Design16/2014/20

The decision this comparison can support

Does the measured quality difference survive the route, elapsed-time, and cost-coverage trade-off for your work?

Start with the score and question profile. Use elapsed time only after both configurations clear your quality floor, and leave cost undecided when either side is not covered.