gpt-6-astra
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score95.3
gpt-6-astra · xhigh
- Backend & testing
- 95
- Frontend & interaction
- 96
- Knowledge & reasoning
- 95
What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
high xhigh
- Overall score
- +4.2 pts 91.1 → 95.3
- Reference cost
- +$2.61 $4.40 → $7.01
- Evaluation time
- +5.1% 79m 52s → 83m 56s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| low | 88.9 | 88 | 94 | 85 | 61m 27s | $3.70 | |
| medium | 88.6 | 91 | 89 | 85 | 72m 41s | $3.76 | |
| high | 91.1 | 92 | 96 | 85 | 79m 52s | $4.40 | |
| xhigh | 95.3 | 95 | 96 | 95 | 83m 56s | $7.01 | |
| max | 94.7 | 95 | 94 | 95 | 93m 26s | $8.35 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026