Evaluation method

How ModelDial compares models

Each capability is evaluated separately. The overall score weights Backend & Testing 40%, Frontend & Interaction 30%, and Knowledge & Reasoning 30%.

Checking the latest verified result

The standard behind the current ranking

See how the scores are weighted, which results can be compared, and when the current results were published.

Latest publication
Capabilities
Backend & Testing, Frontend & Interaction, Knowledge & Reasoning
Overall weighting
Backend & Testing 40% · Frontend & Interaction 30% · Knowledge & Reasoning 30%
Comparison basis
Same capability, same standard, separate ranking

Comparison rules

Three simple rules

01

Keep capabilities separate

Backend, frontend, and reasoning rankings remain visible beneath the overall score so one number does not hide meaningful differences.

02

Compare like with like

Rank and trend changes appear only when the evaluation conditions are comparable.

03

Keep efficiency separate

The weighted overall score measures capability only. Time and reference cost stay separate, while stability is checked before a personal switch.

Decision policy

How to read the ranking before a switch

The overall leader is a useful starting point. Your choice also depends on the task, waiting time, and budget.

Start with the overall leader
Use the weighted ranking as the broad shortlist, then open the capability that matches your work.
Check comparability
Compare earlier results that used the same tasks and scoring rules. One lower score may just be normal variation.
Test locally if needed
Consider completion time and reference cost. If needed, test with your own account and connection before switching.

Publication

Results publish after verification

ModelDial runs evaluations regularly and checks the results before updating the rankings. Previous verified results remain visible until the update is ready.

  1. 01

    Evaluate

    Run the current configurations under consistent conditions.

  2. 02

    Verify

    Confirm that the result is complete, comparable, and free of known anomalies.

  3. 03

    Publish

    Update the current ranking and retain the earlier verified result.

Interpretation

How to use these results

Public Radar

A quick view of model performance in ModelDial's public evaluation. It does not use your provider quota.

Local evaluation

A check of whether your own account and route fit your work. Keys and history stay on your Mac.

Limits

One result is not a permanent conclusion, and reference cost is not your actual bill.