gpt-6-luna
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
gpt-6-luna · xhigh
- Backend & testing
- 62
- Frontend & interaction
- 80
- Knowledge & reasoning
- 50
- Evaluation time
- 63m 52s
- Reference cost
- $0.061
Totals for all three evaluations at this configuration; cost estimated from API usage.
Overall rank 24 / 33 models · Best tested effort
Backend tested · 9/23/2026
Published summary scores use the same effort. Axes may be evaluated on different dates; a backend date refers only to backend testing.
View task coverage and scoring →What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
- Overall score
- +16.2 pts 47.6 → 63.8
- Reference cost
- +$0.019 $0.042 → $0.061
- Evaluation time
- +33.0% 48m 0s → 63m 52s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| low | 31.4 | 23 | 49 | 25 | 39m 27s | — | |
| medium | 33.9 | 24 | 56 | 25 | 39m 46s | $0.032 | |
| high | 47.6 | 44 | 50 | 50 | 48m 0s | $0.042 | |
| xhigh | 63.8 | 62 | 80 | 50 | 63m 52s | $0.061 | |
| max | 56.5 | 61 | 77 | 30 | 72m 51s | — |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Data snapshot published 10/1/2026