gpt-5.6-terra
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score70.5
gpt-5.6-terra · max
- Backend & testing
- 78
- Frontend & interaction
- 76
- Knowledge & reasoning
- 55
What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
xhigh max
- Overall score
- +13.6 pts 56.9 → 70.5
- Reference cost
- +$0.40 $1.54 → $1.94
- Evaluation time
- +97.7% 67m 17s → 133m 2s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| low | 34.9 | 49 | 21 | 30 | 43m 42s | $0.71 | |
| medium | 48.4 | 46 | 50 | 50 | 46m 2s | $0.75 | |
| high | 58.0 | 61 | 67 | 45 | 53m 5s | $1.03 | |
| xhigh | 56.9 | 62 | 72 | 35 | 67m 17s | $1.54 | |
| max | 70.5 | 78 | 76 | 55 | 133m 2s | $1.94 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026