k2.8 preview
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score64.2
k2.8 preview · high
- Backend & testing
- 54
- Frontend & interaction
- 72
- Knowledge & reasoning
- 70
What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
high max
- Overall score
- — 64.2 → —
- Reference cost
- — $0.66 → —
- Evaluation time
- — 70m 58s → —
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| high | 64.2 | 54 | 72 | 70 | 70m 58s | $0.66 | |
| max | — | — | — | 70 | — | — |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026