ModelDial · Methodology

Evaluation methods and result updates

The rankings report completed evaluations. Here is how individual scores, the overall score and result updates work.

Each score has a defined task scope

Backend & testing

Identify conflicting or missing rules and produce verifiable solutions, with checks for constraints, edge cases and regression coverage.

The backend set contains five questions, each worth 20 points, for a total of 100. Versioned graders check answers, constraints, error handling and regression evidence, and retain each question score.

Frontend & interaction

Check behavior, state management and interaction feedback against the scoring rules for the task.

The frontend set evaluates a web application and its interaction flows out of 100. Dimensions and weights depend on the scoring version. For example, r11 assigns 20 points each to layout, interaction and continuity, 15 each to accessibility/feedback and state predictions, and 10 to regression. Expand comparison details for the dimensions and maxima in each record.

Knowledge & reasoning

Score answer correctness and retain results by discipline. Scores describe performance on the evaluated tasks.

The current reasoning set contains 20 cross-disciplinary questions. Its score is correct answers ÷ 20 × 100, with question and correct-answer counts retained by domain. Valid incorrect answers count as wrong; failed requests are not published as zero-score answers.

Valid completed answers are graded and retained, including low scores. Failed requests, incomplete output and unfinished evaluations do not create a valid new score or overwrite an earlier valid result.

Explore benchmark coverage

One configuration for the overall score

Backend 40%Frontend 30%Reasoning 30%

The weighted sum is displayed to one decimal place. A missing axis is not filled with zero and does not produce a complete overall score.

Reasoning effort, time and cost

Reasoning effort is a model request setting, such as High or Max. The same label can mean different things across providers. By default, each model shows its highest-scoring effort for the selected category. Expand it to see other efforts, or adjust them in the comparison.

Reference costs use recorded usage and pricing; evaluation time is the time spent on these tasks. Incomplete cost coverage is labeled. These figures do not measure everyday chat latency or subscription spending.

How new results update the rankings

New models are evaluated promptly after release, and verified results update their configurations. Existing results remain available; routine retesting is not promised. Missing, failed or unfinished evaluations do not erase valid scores. Pages check for data updates.

A score difference describes these tasks, not a permanent lead or capability in every domain. Historical records retain their scoring identities.

Access results and sources through the API