Evaluation method

How ModelDial compares models

Each capability is evaluated separately. The overall score weights Backend & Testing 50%, Frontend & Interaction 30%, and Knowledge & Reasoning 20%.

Checking the latest verified result

The standard behind the current ranking

The public page shows the weighting, comparison boundary, publication state, and source identity needed to interpret the current result.

Latest publication
Capabilities
Backend & Testing, Frontend & Interaction, Knowledge & Reasoning
Overall weighting
Backend & Testing 50% · Frontend & Interaction 30% · Knowledge & Reasoning 20%
Comparison basis
Same capability, same standard, separate ranking

Comparison rules

Three simple rules

01

Keep capabilities separate

Backend, frontend, and reasoning rankings remain visible beneath the overall score so one number does not hide meaningful differences.

02

Compare like with like

Rank and trend changes appear only when the evaluation conditions are comparable.

03

Keep efficiency separate

The weighted overall score measures capability only. Time and reference cost stay separate, while stability is checked before a personal switch.

Decision policy

How to read the ranking before a switch

The current overall leader is the first row of the 50/30/20 ranking, not an automatic instruction to change your configuration.

Start with the overall leader
Use the weighted ranking as the broad shortlist, then open the capability that matches your work.
Check comparability
Use protocol-matched history to judge change. One swing does not prove a model regression.
Verify your route
Compare time and reference cost separately, then verify your exact account and route before switching.

Publication

Results publish after verification

ModelDial verifies on a schedule and publishes only when new evidence passes validation. The prior reliable result remains visible until then.

  1. 01

    Evaluate

    Run the current configurations under consistent conditions.

  2. 02

    Verify

    Confirm that the result is complete, comparable, and free of known anomalies.

  3. 03

    Publish

    Update the current ranking and retain the earlier verified result.

Interpretation

How to use these results

Public Radar

A quick view of model performance in ModelDial's public evaluation. It does not use your provider quota.

Local evaluation

A check of whether your own account and route fit your work. Keys and history stay on your Mac.

Limits

One result is not a permanent conclusion, and reference cost is not your actual bill.