- Published
- Aug 26, 2026, 11:00 PM UTC
- Question pack
- coding-fast-v4.10
- Coverage
- 2 measured configurations
- Batch
- snapshot-2026-08-26T23-00-00Z-r3
Measured configurations
| Configuration | Rank | Score | Elapsed | Reference cost | Route |
|---|---|---|---|---|---|
| DeepSeek V4 Flash / High | #11 | 66/100 | 21m 51s | $0.12 | Custom endpoint |
| DeepSeek V4 Pro / High | #14 | 64/100 | 23m 34s | $0.32 | Custom endpoint |
Five-question profile
| Question | DeepSeek V4 Flash / High | DeepSeek V4 Pro / High |
|---|---|---|
| Black-box Regression Audit | 10/20 | 12/20 |
| Retry Planner Counterexamples | 14/20 | 13/20 |
| CI Adversarial Audit | 13/20 | 12/20 |
| Transaction Regression Design | 14/20 | 14/20 |
| Cache Regression Test Design | 15/20 | 13/20 |
Flash and Pro are not interchangeable rows
DeepSeek V4 Flash and Pro are distinct model identities. Each effort level is another configuration. Combining them into one DeepSeek score would hide which variant produced the result.
Use the table to find whether quality, elapsed time, and cost point to the same configuration. When they do not, apply a task-specific quality floor before optimizing the other two.