deepseek-v4.1-flash
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
Highest overall score68.7
deepseek-v4.1-flash · max
- Backend & testing
- 57
- Frontend & interaction
- 88
- Knowledge & reasoning
- 65
What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
high max
- Overall score
- +20.5 pts 48.2 → 68.7
- Reference cost
- +$0.12 $0.29 → $0.41
- Evaluation time
- +19.7% 29m 47s → 35m 40s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| high | 48.2 | 50 | 44 | 50 | 29m 47s | $0.29 | |
| max | 68.7 | 57 | 88 | 65 | 35m 40s | $0.41 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Evaluation details · Published 9/30/2026