Model performance monitoring

How to detect AI model regressions without overreacting

Use repeated, comparable evidence to separate one noisy batch from a real decline in an AI coding model.

Short answer

Start with the selected Radar snapshot and keep the model, reasoning effort, route, and comparison standard fixed. Investigate only when comparable batches repeatedly decline, then verify locally before changing your default.

Establish the comparison frame

Record the selected snapshot and publication time before comparing results. Use the Radar filters to keep the model, reasoning effort, and comparison standard constant.

A changed route, endpoint, or comparison standard starts a new series. It is not another point in the old trend.

Read the Radar signals together

Use the selected-batch score with the 24-hour range, then open the row to inspect the question profile. Compare elapsed time and reference cost only after the quality signal remains comparable.

  • Look for repeated movement across protocol-matched batches, not one isolated low score.
  • Check whether the change is concentrated in one question, one reasoning level, one route, or the whole model family.
  • Treat stable quality with rising latency differently from a score decline accompanied by failures.

Verify locally before switching

Public Radar contains completed first-party results. It cannot diagnose failures, authentication changes, or routing conditions on your own account.

If the decline repeats, rerun the affected configuration locally, inspect failures and route identity, and compare a nearby effort level or alternate route. Change the default only after the replacement clears the same quality guardrail.