- Published
- Sep 16, 2026, 06:28 AM UTC
- Question pack
- coding-fast-v4.12
- Tested
- 2 measured configurations
- Batch
- evaluation-497e806be8b8470345169c1aad6752d74a0c1ce9e195a9eb7fcc02b628701c77
Results from this batch
| Measure | GPT-5.6 Sol / Max | Qwen 3.8 / Max |
|---|---|---|
| Rank | #6 | #28 |
| Score | 84/100 | 60/100 |
| Elapsed | 34m 23s | 47m 48s |
| Reference cost | $1.57 | $1.07 |
| Route | Custom endpoint | Custom endpoint |
| Black-box Regression Audit | 14/20 | 12/20 |
| Retry Planner Counterexamples | 18/20 | 17/20 |
| CI Adversarial Audit | 15/20 | 8/20 |
| Transaction Regression Design | 17/20 | 14/20 |
| Cache Propagation Certificate | 20/20 | 9/20 |
How to choose from these results
Does the higher-scoring option also fit your time and cost budget?
Start with the score and question profile. Use elapsed time only after both configurations clear your quality floor, and leave cost undecided when either side is not covered.