ModelDial · Model results

claude-opus-4-8

Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.

Highest overall score75.9

claude-opus-4-8 · max

Backend & testing
60
Frontend & interaction
88
Knowledge & reasoning
85

What changes with more effort

Compare scores, cost and time at neighboring tested efforts.

xhigh max
Overall score
+0.4 pts
75.5 75.9
Reference cost
+$7.16
$9.18 $16.33
Evaluation time
+4.6%
102m 10s 106m 52s
View full comparison →

Results by reasoning effort

Select two or three configurations to compare scores and costs.

Swipe to see all results →

Model / effortOverallBackend & testingFrontend & interactionKnowledge & reasoningEvaluation timeReference costCompare
high60.942727584m 26s$4.40
xhigh75.5568790102m 10s$9.18
max75.9608885106m 52s$16.33

How to read these results

The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.

Evaluation details · Published 9/30/2026