claude-opus-4-8
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score75.9
claude-opus-4-8 · max
- Backend & testing
- 60
- Frontend & interaction
- 88
- Knowledge & reasoning
- 85
What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
xhigh max
- Overall score
- +0.4 pts 75.5 → 75.9
- Reference cost
- +$7.16 $9.18 → $16.33
- Evaluation time
- +4.6% 102m 10s → 106m 52s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| high | 60.9 | 42 | 72 | 75 | 84m 26s | $4.40 | |
| xhigh | 75.5 | 56 | 87 | 90 | 102m 10s | $9.18 | |
| max | 75.9 | 60 | 88 | 85 | 106m 52s | $16.33 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026