DeepSeek V4
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
deepseek-v4-pro · max
- Backend & testing
- 67
- Frontend & interaction
- 63
- Knowledge & reasoning
- 65
- Evaluation time
- 93m 15s
- Reference cost
- $1.15
Totals for all three evaluations at this configuration; cost estimated from API usage.
Overall rank 22 / 33 models · Best tested effort
Backend tested · 9/21/2026
Published summary scores use the same effort. Axes may be evaluated on different dates; a backend date refers only to backend testing.
View task coverage and scoring →What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
- Overall score
- +2.1 pts 63.1 → 65.2
- Reference cost
- — — → $1.15
- Evaluation time
- +52.7% 61m 4s → 93m 15s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| deepseek-v4-prohigh | 63.1 | 61 | 54 | 75 | 61m 4s | $0.98Partial cost | |
| deepseek-v4-flashhigh | 62.5 | 58 | 76 | 55 | 98m 25s | $0.42 | |
| deepseek-v4-promax | 65.2 | 67 | 63 | 65 | 93m 15s | $1.15 | |
| deepseek-v4-flashmax | 56.4 | 54 | 61 | 55 | 89m 12s | $0.60 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Data snapshot published 10/1/2026