gpt-5.6-sol
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score77.6
gpt-5.6-sol · max
- Backend & testing
- 83
- Frontend & interaction
- 73
- Knowledge & reasoning
- 75
What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
xhigh max
- Overall score
- +8.2 pts 69.4 → 77.6
- Reference cost
- +$0.69 $2.56 → $3.25
- Evaluation time
- +22.5% 71m 19s → 87m 21s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| low | 67.4 | 65 | 78 | 60 | 41m 49s | $1.53 | |
| medium | 73.9 | 70 | 83 | 70 | 59m 12s | $1.72 | |
| high | 75.7 | 73 | 90 | 65 | 58m 24s | $2.05 | |
| xhigh | 69.4 | 76 | 65 | 65 | 71m 19s | $2.56 | |
| max | 77.6 | 83 | 73 | 75 | 87m 21s | $3.25 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026