gpt-5.6-luna
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score55.3
gpt-5.6-luna · max
- Backend & testing
- 61
- Frontend & interaction
- 78
- Knowledge & reasoning
- 25
What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
xhigh max
- Overall score
- +1.3 pts 54.0 → 55.3
- Reference cost
- +$0.18 $0.16 → $0.35
- Evaluation time
- +14.4% 75m 23s → 86m 14s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| low | 18.9 | 30 | 18 | 5 | 52m 31s | $0.067 | |
| medium | 30.2 | 47 | 13 | 25 | 18m 22s | $0.084 | |
| high | 38.1 | 45 | 52 | 15 | 32m 18s | $0.14 | |
| xhigh | 54.0 | 54 | 83 | 25 | 75m 23s | $0.16 | |
| max | 55.3 | 61 | 78 | 25 | 86m 14s | $0.35 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026