AI model benchmarks & comparisons
Private datasets.
Compare model performance.
Check new model results, compare generations, and see how reasoning effort affects scores, evaluation time and cost.
35 models · 83 configurations
Top five · Overall
Out of 1001
gpt-6-astramax95.5
2
gpt-6.1-solmax93.8
3
claude-opus-5-5max90.3
4
claude-fable-5-1max89.8
5
claude-sonnet-5-5xhigh86.6
Generations & reasoning levels
Overall scoreBackend 40% · Frontend 30% · Reasoning 30%, from the same tested configuration.
Explore results by task
Explore the evaluation →Backend & Testing Service contracts, error handling and regression testsFrontend & Interaction UI behavior, interaction state and responsive layoutsKnowledge & Reasoning Knowledge and reasoning across disciplines
Tested on our private datasets; scores reflect these tasks. Why we test this way →
Model performance and evaluation cost
Full rankings →Upper left: less cost, higher score · Each point is one model and effort
Pareto frontierNew model resultsOther configurations66 / 83 configurations plotted
Select a point to see its scores and cost
Overall95.5Reference cost$8.35
Total estimated API cost across three evaluations. Log scale; only complete, positive costs are plotted. Comparison method →