grok-4.7
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
grok-4.7 · high
- Backend & testing
- 87
- Frontend & interaction
- 80
- Knowledge & reasoning
- 85
- Evaluation time
- 196m 18s
- Reference cost
- $5.25
Totals for all three evaluations at this configuration; cost estimated from API usage.
Overall rank 6 / 33 models · Best tested effort
Backend tested · 9/22/2026
Published summary scores use the same effort. Axes may be evaluated on different dates; a backend date refers only to backend testing.
View task coverage and scoring →What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
- Overall score
- −2.8 pts 84.3 → 81.5
- Reference cost
- +$0.51 $5.25 → $5.75
- Evaluation time
- +19.6% 196m 18s → 234m 43s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| high | 84.3 | 87 | 80 | 85 | 196m 18s | $5.25 | |
| xhigh | 81.5 | 86 | 77 | 80 | 234m 43s | $5.75 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Data snapshot published 10/1/2026