Configuration comparison

GPT-5.6 Sol vs Qwen 3.8

Compare the highest-scoring GPT-5.6 Sol and Qwen 3.8 effort levels from the same test batch.

Current batch

GPT-5.6 Sol / Max leads the selected batch by 24 points. GPT-5.6 Sol / Max completed 13m 25s sooner. Reference costs are available for both configurations.

Evaluation details
Published
Sep 16, 2026, 06:28 AM UTC
Question pack
coding-fast-v4.12
Tested
2 measured configurations
Batch
evaluation-497e806be8b8470345169c1aad6752d74a0c1ce9e195a9eb7fcc02b628701c77

Results from this batch

MeasureGPT-5.6 Sol / MaxQwen 3.8 / Max
Rank#6#28
Score84/10060/100
Elapsed34m 23s47m 48s
Reference cost$1.57$1.07
RouteCustom endpointCustom endpoint
Black-box Regression Audit14/2012/20
Retry Planner Counterexamples18/2017/20
CI Adversarial Audit15/208/20
Transaction Regression Design17/2014/20
Cache Propagation Certificate20/209/20

How to choose from these results

Does the higher-scoring option also fit your time and cost budget?

Start with the score and question profile. Use elapsed time only after both configurations clear your quality floor, and leave cost undecided when either side is not covered.