- Published
- Aug 26, 2026, 11:00 PM UTC
- Question pack
- coding-fast-v4.10
- Coverage
- 2 measured configurations
- Batch
- snapshot-2026-08-26T23-00-00Z-r3
Same-batch evidence
| Measure | GPT-5.6 Sol / Max | Qwen 3.8 / Max |
|---|---|---|
| Rank | #1 | #10 |
| Score | 85/100 | 66/100 |
| Elapsed | 44m 58s | 21m 19s |
| Reference cost | $1.82 | $0.89 |
| Route | Official login | Custom endpoint |
| Black-box Regression Audit | 14/20 | 10/20 |
| Retry Planner Counterexamples | 20/20 | 17/20 |
| CI Adversarial Audit | 18/20 | 12/20 |
| Transaction Regression Design | 17/20 | 13/20 |
| Cache Regression Test Design | 16/20 | 14/20 |
The decision this comparison can support
Does the measured quality difference survive the route, elapsed-time, and cost-coverage trade-off for your work?
Start with the score and question profile. Use elapsed time only after both configurations clear your quality floor, and leave cost undecided when either side is not covered.