ModelDial · Model results

gpt-5.6-sol

Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.

Highest overall score77.6

gpt-5.6-sol · max

Backend & testing
83
Frontend & interaction
73
Knowledge & reasoning
75

What changes with more effort

Compare scores, cost and time at neighboring tested efforts.

xhigh max
Overall score
+8.2 pts
69.4 77.6
Reference cost
+$0.69
$2.56 $3.25
Evaluation time
+22.5%
71m 19s 87m 21s
View full comparison →

Results by reasoning effort

Select two or three configurations to compare scores and costs.

Swipe to see all results →

Model / effortOverallBackend & testingFrontend & interactionKnowledge & reasoningEvaluation timeReference costCompare
low67.465786041m 49s$1.53
medium73.970837059m 12s$1.72
high75.773906558m 24s$2.05
xhigh69.476656571m 19s$2.56
max77.683737587m 21s$3.25

How to read these results

The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.

Evaluation details · Published 9/30/2026