grok-4.6
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score79.5
grok-4.6 · xhigh
- Backend & testing
- 75
- Frontend & interaction
- 80
- Knowledge & reasoning
- 85
What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
high xhigh
- Overall score
- +5.9 pts 73.6 → 79.5
- Reference cost
- +$0.054 $2.46 → $2.51
- Evaluation time
- +42.7% 101m 5s → 144m 16s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| high | 73.6 | 67 | 76 | 80 | 101m 5s | $2.46 | |
| xhigh | 79.5 | 75 | 80 | 85 | 144m 16s | $2.51 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026