ModelDial · Model results

gpt-6-astra

Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.

Highest overall score95.3

gpt-6-astra · xhigh

Backend & testing
95
Frontend & interaction
96
Knowledge & reasoning
95

What changes with more effort

Compare scores, cost and time at neighboring tested efforts.

high xhigh
Overall score
+4.2 pts
91.1 95.3
Reference cost
+$2.61
$4.40 $7.01
Evaluation time
+5.1%
79m 52s 83m 56s
View full comparison →

Results by reasoning effort

Select two or three configurations to compare scores and costs.

Swipe to see all results →

Model / effortOverallBackend & testingFrontend & interactionKnowledge & reasoningEvaluation timeReference costCompare
low88.988948561m 27s$3.70
medium88.691898572m 41s$3.76
high91.192968579m 52s$4.40
xhigh95.395969583m 56s$7.01
max94.795949593m 26s$8.35

How to read these results

The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.

Evaluation details · Published 9/30/2026