claude-opus-5
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score81.0
claude-opus-5 · max
- Backend & testing
- 66
- Frontend & interaction
- 87
- Knowledge & reasoning
- 95
What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
xhigh max
- Overall score
- +1.4 pts 79.6 → 81.0
- Reference cost
- +$2.32 $5.68 → $7.99
- Evaluation time
- +14.6% 80m 26s → 92m 12s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| high | 78.5 | 65 | 85 | 90 | 66m 37s | $4.06 | |
| xhigh | 79.6 | 70 | 87 | 85 | 80m 26s | $5.68 | |
| max | 81.0 | 66 | 87 | 95 | 92m 12s | $7.99 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026