Backend & Testing
40%Code, contracts, debugging, and test design
AI coding models, tested
Use test results to choose a model and reasoning level. If needed, test your own connection locally on your Mac.
A weighted result: 40% Backend & Testing, 30% Frontend & Interaction, and 30% Knowledge & Reasoning
Current overall leadergpt-6-astraXHigh·All three capabilities verified
How we measure
Overall score weights Backend & Testing 40%, Frontend & Interaction 30%, and Knowledge & Reasoning 30%. Only complete configurations enter the ranking.
Code, contracts, debugging, and test design
Interface implementation and visual behavior
Research depth, synthesis, and hard reasoning
Each Backend & Testing publication is checked against the previous comparable result, separating rank movement, score changes, and recommendation changes from ordinary reshuffling.
Read the evaluation methodLocal verification
Use the public rankings to compare three capabilities. If you want to check your own account or connection, run the Backend & Testing evaluation on your Mac. Keys and test history stay local.
Test results
Verified before publication
Must hold
Requests go directly through the route you chose.
Keys, configuration, scan history, and recommendations remain local.
Learn how scores work, when rankings update, and when a local test may help.
ModelDial tests models at different reasoning levels and shows their scores, completion times, and reference costs. Use the public rankings to shortlist models. If you want to check whether your own account or connection performs differently, you can run a local test.
Backend & Testing covers code implementation, debugging, and test design. Frontend & Interaction covers interface implementation, interaction details, and visual results. Knowledge & Reasoning covers information understanding, synthesis, and complex problem solving. Each capability is ranked separately.
Each capability has its own scoring rules, with a maximum of 100 points. Scores are not real-world success rates or directly comparable with unrelated benchmarks. The overall score weights Backend & Testing at 40%, Frontend & Interaction at 30%, and Knowledge & Reasoning at 30%. A model and effort level enter the overall ranking only when all three results are verified. Missing scores are never counted as zero.
Results appear after they have been checked. Until then, the page shows the previous verified score, or a pending label if there is no earlier result. The timestamp shows when a result was published, not when the test started.
No. Radar's reference cost is the approximate cost of a comparable reference run, not your actual bill. Your actual bill depends on your account, caching, and custom route.
Public Radar uses ModelDial-controlled runs across three capabilities and uses no quota from your provider account. A local scan applies the Backend & Testing method to your own account, route, or endpoint only when you start it. Keys, configuration, scan history, and recommendations remain on your Mac. Website analytics uses a first-party random browser cookie, expiring after 180 days, and honors GPC and DNT.