Model providers update weights, adjust safety layers, and revise defaults, often without a version bump you can pin. Drift Monitor watches the behaviour of your production system continuously and raises an alert when it changes — whatever the cause.
Highlights
- Behavioural baselines. Drift Monitor builds a statistical profile of your system’s outputs from production traces: task success signals, refusal rates, output length and structure, tool-use patterns, latency, and cost.
- Continuous probes. A fixed probe set drawn from your eval suites runs against production models on a schedule, separating changes in the model from changes in your traffic.
- Change-point alerts. When observed behaviour departs from baseline, you get an alert with the affected metrics, the estimated onset time, and example traces — not a dashboard you have to remember to check.
- Provider correlation. Detected shifts are annotated against known provider release windows, so you can tell a silent model revision from a change in what your users are asking.
- Segmented views. Track drift per route, per model, and per tenant, since an upgrade that helps one workload can degrade another.
Why it matters
A system you evaluated in March is not the system you are running in July, even if you changed nothing. Drift detection closes the gap between when behaviour shifts and when you find out — turning provider updates from something that happens to you into something you observe, measure, and respond to on your own schedule.