Evaluation method
How ModelDial compares models
Each capability is evaluated separately. The overall score weights Backend & Testing 50%, Frontend & Interaction 30%, and Knowledge & Reasoning 20%.
Checking the latest verified result
The standard behind the current ranking
The public page shows the weighting, comparison boundary, publication state, and source identity needed to interpret the current result.
- Latest publication
- Capabilities
- Backend & Testing, Frontend & Interaction, Knowledge & Reasoning
- Overall weighting
- Backend & Testing 50% · Frontend & Interaction 30% · Knowledge & Reasoning 20%
- Comparison basis
- Same capability, same standard, separate ranking
Comparison rules
Three simple rules
Keep capabilities separate
Backend, frontend, and reasoning rankings remain visible beneath the overall score so one number does not hide meaningful differences.
Compare like with like
Rank and trend changes appear only when the evaluation conditions are comparable.
Keep efficiency separate
The weighted overall score measures capability only. Time and reference cost stay separate, while stability is checked before a personal switch.
Decision policy
How to read the ranking before a switch
The current overall leader is the first row of the 50/30/20 ranking, not an automatic instruction to change your configuration.
- Start with the overall leader
- Use the weighted ranking as the broad shortlist, then open the capability that matches your work.
- Check comparability
- Use protocol-matched history to judge change. One swing does not prove a model regression.
- Verify your route
- Compare time and reference cost separately, then verify your exact account and route before switching.
Publication
Results publish after verification
ModelDial verifies on a schedule and publishes only when new evidence passes validation. The prior reliable result remains visible until then.
- 01
Evaluate
Run the current configurations under consistent conditions.
- 02
Verify
Confirm that the result is complete, comparable, and free of known anomalies.
- 03
Publish
Update the current ranking and retain the earlier verified result.
Interpretation
How to use these results
Public Radar
A quick view of model performance in ModelDial's public evaluation. It does not use your provider quota.
Local evaluation
A check of whether your own account and route fit your work. Keys and history stay on your Mac.
Limits
One result is not a permanent conclusion, and reference cost is not your actual bill.