Evaluation method
How ModelDial compares models
Each capability is evaluated separately. The overall score weights Backend & Testing 40%, Frontend & Interaction 30%, and Knowledge & Reasoning 30%.
Checking the latest verified result
The standard behind the current ranking
See how the scores are weighted, which results can be compared, and when the current results were published.
- Latest publication
- Capabilities
- Backend & Testing, Frontend & Interaction, Knowledge & Reasoning
- Overall weighting
- Backend & Testing 40% · Frontend & Interaction 30% · Knowledge & Reasoning 30%
- Comparison basis
- Same capability, same standard, separate ranking
Comparison rules
Three simple rules
Keep capabilities separate
Backend, frontend, and reasoning rankings remain visible beneath the overall score so one number does not hide meaningful differences.
Compare like with like
Rank and trend changes appear only when the evaluation conditions are comparable.
Keep efficiency separate
The weighted overall score measures capability only. Time and reference cost stay separate, while stability is checked before a personal switch.
Decision policy
How to read the ranking before a switch
The overall leader is a useful starting point. Your choice also depends on the task, waiting time, and budget.
- Start with the overall leader
- Use the weighted ranking as the broad shortlist, then open the capability that matches your work.
- Check comparability
- Compare earlier results that used the same tasks and scoring rules. One lower score may just be normal variation.
- Test locally if needed
- Consider completion time and reference cost. If needed, test with your own account and connection before switching.
Publication
Results publish after verification
ModelDial runs evaluations regularly and checks the results before updating the rankings. Previous verified results remain visible until the update is ready.
- 01
Evaluate
Run the current configurations under consistent conditions.
- 02
Verify
Confirm that the result is complete, comparable, and free of known anomalies.
- 03
Publish
Update the current ranking and retain the earlier verified result.
Interpretation
How to use these results
Public Radar
A quick view of model performance in ModelDial's public evaluation. It does not use your provider quota.
Local evaluation
A check of whether your own account and route fit your work. Keys and history stay on your Mac.
Limits
One result is not a permanent conclusion, and reference cost is not your actual bill.