- Published
- Sep 16, 2026, 06:28 AM UTC
- Question pack
- coding-fast-v4.12
- Tested
- 5 measured configurations
- Batch
- evaluation-497e806be8b8470345169c1aad6752d74a0c1ce9e195a9eb7fcc02b628701c77
Measured configurations
| Configuration | Rank | Score | Elapsed | Reference cost | Route |
|---|---|---|---|---|---|
| GPT-5.6 Sol / Max | #6 | 84/100 | 34m 23s | $1.57 | Custom endpoint |
| GPT-5.6 Sol / XHigh | #9 | 75/100 | 35m 6s | $1.22 | Custom endpoint |
| GPT-5.6 Sol / High | #11 | 73/100 | 13m 36s | $0.66 | Custom endpoint |
| GPT-5.6 Sol / Medium | #18 | 66/100 | 28m 4s | $0.74 | Custom endpoint |
| GPT-5.6 Sol / Low | #19 | 65/100 | 10m 43s | $0.53 | Custom endpoint |
Five-question profile
| Question | GPT-5.6 Sol / Max | GPT-5.6 Sol / XHigh | GPT-5.6 Sol / High | GPT-5.6 Sol / Medium | GPT-5.6 Sol / Low |
|---|---|---|---|---|---|
| Black-box Regression Audit | 14/20 | 14/20 | 16/20 | 12/20 | 16/20 |
| Retry Planner Counterexamples | 18/20 | 19/20 | 17/20 | 16/20 | 15/20 |
| CI Adversarial Audit | 15/20 | 12/20 | 9/20 | 13/20 | 12/20 |
| Transaction Regression Design | 17/20 | 17/20 | 19/20 | 13/20 | 13/20 |
| Cache Propagation Certificate | 20/20 | 13/20 | 12/20 | 12/20 | 9/20 |
What the effort spread can tell you
This family is measured at several effort levels through the recorded official-login route. The page keeps every effort separate so a strong Max result does not get attributed to Low, or vice versa.
Use the highest score and fastest completion as two different signals. If they point to different configurations, the choice depends on your quality floor rather than the family name.