claude-sonnet-5-5
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score85.8
claude-sonnet-5-5 · xhigh
- Backend & testing
- 84
- Frontend & interaction
- 89
- Knowledge & reasoning
- 85
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| xhigh | 85.8 | 84 | 89 | 85 | 48m 32s | $2.80 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026