claude-opus-5-5
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score89.1
claude-opus-5-5 · max
- Backend & testing
- 81
- Frontend & interaction
- 89
- Knowledge & reasoning
- 100
What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
xhigh max
- Overall score
- +2.1 pts 87.0 → 89.1
- Reference cost
- −$2.63 $6.29 → $3.67
- Evaluation time
- −44.6% 61m 15s → 33m 57s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| xhigh | 87.0 | 81 | 87 | 95 | 61m 15s | $6.29 | |
| max | 89.1 | 81 | 89 | 100 | 33m 57s | $3.67 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026