When a model is wrong but reports no errors: build your own data drift detection with Prometheus and Grafana
A model can keep returning HTTP 200 in a few milliseconds even after customer data has moved far from the data it learned on. This guide shows how to build a system that catches the change before the customer does.
In brief
- A model that fails because of drift raises no errors, so you have to measure the data distribution yourself and feed that number into your monitoring system.
- Evidently computes drift scores and serves them on their own endpoint, Prometheus periodically pulls the metrics over HTTP, and Grafana draws the dashboard and sends alerts when a threshold is crossed.
- Drift scores must use a Gauge, not a Counter. Alert thresholds must also be calibrated on real data, not copied from a tutorial.
A fraud-scoring model can return results in a few milliseconds, answer every request with HTTP 200 and write not a single error line to the log. It can still be wrong. When production data has moved significantly away from the training data, which is what data drift means, the model has no way of knowing and nothing to report.
Only a system built specifically to measure the data distribution will notice.
For an FDE, this is part of the daily job. You deploy a model onto a customer’s infrastructure, and a few weeks later someone asks: “Why has the model been so poor lately?” If you already have a dashboard showing which feature started to shift and on which day, you answer with data. If you do not, all you can do is guess.
This article builds exactly that system with three tools. Evidently collects and computes the metrics, Prometheus stores them, and Grafana displays them and raises alerts.
What you will build, and what you need
The end result is two Python processes. The model-serving service exposes metrics on prediction counts and latency through an HTTP endpoint. A drift job that runs on a schedule exposes drift scores through a second endpoint.
Prometheus visits both endpoints at regular intervals to pull the numbers, Grafana has a panel charting each feature’s drift score over time, and you receive an alert when a score crosses the threshold.
You need Python 3, a trained model (any scikit-learn classifier will do), and Prometheus and Grafana installed locally or running in containers. You also need a reference dataset, usually the data the model was trained on. Without it there is nothing to compare against.
One architectural point is worth grasping before writing any code. Prometheus collects time series using a pull model over HTTP: each process only has to expose its metrics on an endpoint, and Prometheus comes to fetch them. So you do not write any code to push metrics anywhere.
Step 1: pick the right metric type for each question
Prometheus provides client libraries for instrumenting application code, and a model-serving service is just another application. To monitor a model you need to answer three questions, and each fits one metric type.
| Question | Metric type | Why |
|---|---|---|
| How many predictions has the model returned? | Counter | Only increases monotonically, which suits counting |
| How long does inference take? | Histogram | Counts each observation into configurable buckets |
| How far is the data drifting? | Gauge | The value can go up or down freely |
The first two questions are answered inside the model-serving service. Below is a minimal sketch using the Python client library. Function names can differ between versions, so check against the docs for the version you have installed.
# serve.py — sketch; check the API against the client library docs
from prometheus_client import Counter, Histogram, start_http_server
PREDICTIONS = Counter("predictions_total", "Number of predictions returned")
LATENCY = Histogram("inference_seconds", "Inference time")
start_http_server(8000)
def predict(x):
with LATENCY.time():
y = model.predict(x)
PREDICTIONS.inc()
return y
Check: send a few requests to the predict function, then open localhost:8000/metrics in a browser. You should see predictions_total rise after each request, along with the bucket lines for inference_seconds.
Step 2: compute drift in batches, in a process with its own endpoint
Drift is a property of a distribution, and a single request has no distribution. In the architecture Evidently describes, Evidently reads the model’s logs, compares recent data with the reference set, and then exposes an endpoint for Prometheus to collect from.
The “own endpoint” detail matters more than it looks. A Gauge lives in the memory of the process that created it. A scheduled job running in a different process cannot write to the model-serving service’s Gauge. So the drift job has to create its own Gauge and expose its own metrics on a different port, here 8001.
You could also run the drift calculation as a background thread inside the service itself, but keeping it separate stops the heavy computation from slowing down inference.
# drift_job.py — pseudocode. compute_drift() stands in for Evidently's Data Drift Report;
# the Evidently API changes between versions, see the current docs.
from prometheus_client import Gauge, start_http_server
DRIFT = Gauge("feature_drift_score", "Drift score per feature", ["feature"])
start_http_server(8001)
while True:
current_window = load_recent_prediction_log() # read the service's log
for feature in FEATURES:
score = compute_drift(reference[feature], current_window[feature])
DRIFT.labels(feature=feature).set(score)
sleep_until_next_run() # e.g. once an hour
Picture a fraud model with 20 features. Each feature is one label value on the same Gauge, so you get 20 time series. Evidently notes that this Data Drift example applies in the same way to its other Reports, so later you can add data-quality metrics without changing the architecture.
Check: after the first run, localhost:8001/metrics should contain 20 lines of feature_drift_score{feature="..."}, each carrying a value.
Step 3: configure Prometheus to pull the metrics
In the configuration file, scrape_interval sets how often Prometheus pulls metrics. Because there are two processes, you declare two targets. The file below is trimmed for illustration.
# prometheus.yml (trimmed)
global:
scrape_interval: 15s
scrape_configs:
- job_name: "fraud-model"
static_configs:
- targets: ["localhost:8000"]
- job_name: "drift-monitor"
static_configs:
- targets: ["localhost:8001"]
A quick calculation. At a 15-second interval, each time series receives 240 samples an hour. But if the drift job runs only once an hour, 239 of those samples repeat the same value. That is not wrong, but you need to understand it when reading the chart: the drift line moves in steps that follow the job’s schedule, not the scrape interval.
Start Prometheus with this file (the Getting Started page has the right command for your version), open the web interface and type the simplest possible PromQL query, the metric name feature_drift_score.
Check: the targets page should show both fraud-model and drift-monitor as up. If a job is missing, the cause is usually a wrong port or a process that is not running.
Step 4: build the panel and alerts in Grafana
In Grafana, add Prometheus as a data source, create a time series panel and use the query feature_drift_score. Each feature becomes its own line. This is the chart you will open when the customer asks “is something wrong with the model?”
Next come alerts. Grafana lets you set alerts via email, Slack or SMS on custom thresholds. DataCamp’s tutorial uses a threshold of 0.026, meaning the system sends an alert when the drift score exceeds that level. To try it out, you can write a condition such as feature_drift_score > 0.026.
Check: take the test data, double the values of one feature and feed it into the current data window. That feature’s line should jump and the alert should arrive on the channel you configured.
The right threshold comes from the customer’s own data
The safe approach is to run the system in observe-only mode for a few weeks while the model is performing well, then set the threshold from those numbers. Suppose the drift job runs every hour for 4 weeks: each feature gets 4 × 7 × 24 = 672 baseline points.
Sort those 672 points and take the 99th percentile. Since 1% of 672 is 6.72, this value sits around the seventh-highest point. Say it is 0.04: that means in 99% of normal hours the drift score did not exceed 0.04.
You can set the threshold somewhat higher, say 0.05, so that only genuinely abnormal shifts trigger an alert. The numbers here are hypothetical, but the method works on real data.
Calculate it separately for each important feature, because some features naturally fluctuate more than others. In PromQL you can use quantile_over_time over the baseline period. Check the syntax in the Prometheus docs before using it.
Mistakes that make a monitoring system useless
The first is using a Counter for drift scores. A Counter only goes up, so when drift falls the metric cannot reflect it and your chart will mislead you. Any value that represents “current state”, such as a drift score or accuracy, must use a Gauge.
The second is copying a threshold straight from a tutorial. A threshold that is too low floods the customer’s Slack with alerts, and after a week nobody reads them. The percentile calculation above costs a few lines of code and saves you from this.
The third is labelling by something with a very large number of values, such as customer ID. As in step 2, each label value produces its own time series: 20 features means 20 readable lines, while a per-user label makes the number of time series grow with the number of users.
Keep labels to dimensions with a small, known set of values, and read the Prometheus docs carefully before adding a new label.
The last one few people notice: the reference set no longer matches the model in production, because the model has been retrained while the reference set was left unchanged.
Where this skill shows up in customer work
In customer work, the hardest part is usually not the code. You need to sit down with the business side and agree on three things: which features matter enough to warrant their own alert, who receives the alerts, and what happens next when an alert fires. Without a follow-up step, the dashboard is decoration.
So in your first session with a customer, ask whether they use Prometheus, Grafana or some other monitoring system. Plugging the model’s metrics into infrastructure they already have is usually far easier to get accepted than proposing a new stack.
When reading FDE or MLOps job descriptions, look for phrases such as “model monitoring”, “observability” or “drift detection”. On a CV, “set up Prometheus and Grafana” says little. “Detected drift across 20 features, alerting via Slack, with thresholds taken from the 99th percentile of baseline data” tells a recruiter you understand the problem.
Every deployed model will go wrong at some point. The difference is whether you find out from an alert as the data starts to shift, or wait until the customer calls to tell you.