AI model benchmarks & comparisons

Private datasets.
Compare model performance.

Check new model results, compare generations, and see how reasoning effort affects scores, evaluation time and cost.

35 models · 83 configurations

Top five · Overall

Out of 100
1
2
3
4
5

Generations & reasoning levels

Overall score

Backend 40% · Frontend 30% · Reasoning 30%, from the same tested configuration.

Explore results by task

Explore the evaluation →

Tested on our private datasets; scores reflect these tasks. Why we test this way →

Model performance and evaluation cost

Full rankings →

Upper left: less cost, higher score · Each point is one model and effort

Pareto frontierNew model resultsOther configurations66 / 83 configurations plotted
Overall score ↑020406080100$0.030$0.10$0.30$1.00$3.00$10.00$30.00Reference cost · All three evaluations · USD, log scale

Select a point to see its scores and cost

gpt-6-astra
OpenAImax
Overall95.5Reference cost$8.35

Total estimated API cost across three evaluations. Log scale; only complete, positive costs are plotted. Comparison method →