This is part one of three. It covers everything you need to start using Alibi Detect for real work: asking, in code, whether the data reaching your model still looks like the data the model learned from. By the end you can install the library, run a drift check on a table of numbers, read the result without guessing, tune the one setting that controls false alarms, save a detector to disk and load it back, and explain to a manager what the whole thing does and does not prove. Mid-level and Senior take the same topics further; nothing here is thrown away.
Each section ends with a Try it task. Do them as you go. They take a few minutes each, and drift detection only makes sense once you have watched a detector stay quiet on normal data and then fire on data you deliberately damaged.
Why a model that worked yesterday can be wrong today
A trained model is a summary of the data it saw. It learned that, in that data, certain input patterns go with certain outputs. The moment you deploy it, the world keeps moving and the model does not. Customers change their behaviour, a sensor gets recalibrated, an upstream team renames a category, a marketing campaign brings in a different kind of user, a currency devalues. The model has no way to notice any of this. It receives numbers, produces numbers, and returns them with the same confidence as on its first day.
This is the quiet failure that makes machine learning different from ordinary software. A web service that breaks returns errors, and your monitoring lights up. A model that has gone stale returns perfectly valid-looking answers that are simply worse. Nothing crashes. The only symptom is that the business outcome slowly degrades, and by the time someone notices, weeks of bad predictions are already in the world.
There are two ways to find out. The first is to wait for the true answers (the labels) to arrive and measure accuracy. That is the gold standard, but labels are often late: you learn whether a loan defaulted after months, whether a customer churned after a quarter. The second way is to watch the inputs. If the data arriving today no longer resembles the data the model was trained on, you have a strong early warning that the model is operating outside the territory it understands, long before any label arrives.
Alibi Detect is a Python library for that second way. Its own documentation describes it as a source-available Python library focused on outlier, adversarial and drift detection. You give it a sample of data that represents "normal", and it gives you back a yes-or-no answer, plus supporting numbers, about whether new data is different enough to worry about.
Notice what the library does not do. It does not retrain your model, it does not fix anything, and it does not tell you why the data changed. It raises a hand and says "this looks different". What you do next is a human or pipeline decision, and a good part of the later guides is about making that decision well.
What people did before, and where this fits
Before dedicated libraries, teams monitored models in one of three ways, and you will still meet all of them.
The first was no monitoring at all, relying on users or analysts to complain. This works until it does not, and it fails silently.
The second was hand-rolled summary statistics. An engineer computes the mean and standard deviation of each feature every week, plots them, and eyeballs the chart. It catches a feature that doubles in value. It misses a change in the relationship between features, and it has no principled answer to the question "is this wobble normal?", because nothing tells you how much a mean should wobble from sampling noise alone.
The third was hand-rolled statistical tests, calling a two-sample test from SciPy for each column. This is the right idea, and it is exactly what the simplest Alibi Detect detector does. What the library adds is the surrounding machinery: tests that work on many features at once with proper correction for repeated testing, tests that work on images and text by first compressing them, versions that process one instance at a time instead of a batch, a consistent return format, and a way to save and reload a configured detector.
Alibi Detect also sits next to, not instead of, other tools in this catalogue. If you want ready-made drift reports and dashboards for tabular data, look at Evidently and NannyML. If you want to track experiments and models, MLflow is the usual choice. Alibi Detect is the lower-level, more programmable building block, and it is the one with the widest set of detector types, which is why people reach for it when the data is images, text, or a stream.
https://docs.seldon.ai/alibi-detect and find the algorithm overview page. Count how many drift methods the table lists. You are not expected to learn them all; the point is to see how wide the menu is.The mental model: four nouns
Everything in the library is built from four ideas. Learn these and the documentation becomes readable.
The detector. A detector is an ordinary Python object. You create it, and then you call its predict() method on data. There are three families. Drift detectors ask whether a whole batch of data differs from the reference. Outlier detectors ask whether a single instance looks unlike the training data. Adversarial detectors ask whether an input was deliberately crafted to fool a model. This guide is mostly about drift, because it is the most commonly useful and the easiest first step.
Reference data, written x_ref in the code. This is the sample the detector treats as normal. It is usually a slice of your training or validation data. The quality of this sample decides the quality of everything else, so it deserves more thought than beginners give it. A reference set that is too small gives a jumpy detector. A reference set that does not represent the real normal gives you alarms about a "problem" that was always there.
The test and its threshold. A drift detector runs a statistical test comparing the new batch with the reference. The test produces a p-value: roughly, the probability of seeing a difference this large if the two datasets really came from the same source. You choose a threshold, called p_val in the code, commonly 0.05. If the p-value falls below the threshold, the detector declares drift. We will unpack what that really means shortly, because the usual misreading causes most beginner confusion.
The result. Calling predict() returns a dictionary with two parts: data, which holds the verdict and the numbers, and meta, which describes the detector that produced it. The single most important field is is_drift, which is 1 when drift was found and 0 when it was not.
Two more words you will see early. Data drift, also called covariate drift, means the distribution of the inputs changed. This is what the main detectors test. Concept drift means the relationship between inputs and the correct answer changed even if the inputs look the same, for example a fraud pattern that evolves. Input-only tests cannot confirm concept drift on their own, and you usually need labels for that. Alibi Detect has detectors based on model uncertainty that help without labels, but treat them as hints rather than proof.
Finally, detectors come in two styles. Offline detectors take a whole batch and compare it with the reference each time you call them. Online detectors take one instance at a time, keep a sliding window, and track a running timestep. Online detectors are for streams and they introduce two extra settings, ert and window_size, which we will meet briefly later. Start offline.
is_drift. Then check yourself against this section.Read this before you build anything: the licence
Most tools in this catalogue are open source in the usual sense. Alibi Detect changed. Versions up to 0.11.4 were released under the Apache 2.0 licence. Starting with version 0.11.5, released in January 2024, the project moved to the Business Source License 1.1. The current release is 0.13.0, from December 2025, and it is under the Business Source License too. PyPI lists it as "Other/Proprietary License", and the project describes itself as source-available, not open source.
In plain terms, the licence lets you read the code, copy it, modify it and redistribute it, and it lets you use it for non-production purposes. It also contains an additional grant that allows production use by non-profit educational institutions, as long as that use does not involve commercialising the library or putting it into models, products or services sold, licensed or marketed to third parties. Production use by anyone else needs a commercial licence from Seldon, the company behind the project, whose pricing page is linked in the sources.
What does this mean for you as a student? Learning, coursework, personal experiments and portfolio projects are non-production use, and you are fine. The care is needed later. If you join an employer or take a client and you put Alibi Detect into a live system, somebody must check that the licence covers that use. This is not a hypothetical worry: the MLServer project, which is also maintained by Seldon, has an open issue about removing Alibi Detect from its base container image because of exactly this licence mismatch.
The other practical consequence is about alternatives. Versions up to 0.11.4 are Apache 2.0, but they date from the Python 3.7 to 3.11 era and carry dependency and security costs, so pinning an old version is a poor default. For a real product, comparing with other drift tools in this catalogue is a sensible step. Everything you learn here about drift (reference data, tests, thresholds, false alarms) transfers to those tools, so your effort is not wasted either way.
Installing and checking your setup
Alibi Detect needs Python 3.9 or newer, and Python 3.12 support arrived in release 0.13.0. As with every Python project, work in a virtual environment so that the library and its dependencies do not collide with anything else on your machine.
python -m venv .venv
source .venv/bin/activate # on Windows: .venv\Scripts\activate
pip install alibi-detect
That default install gives you the basic detectors and, importantly, no deep-learning framework. That is enough for everything in this guide. The simplest drift detector, the Kolmogorov-Smirnov test, runs on plain NumPy and SciPy.
The library also offers extras for heavier detectors. You will need these later for image or text data, or for detectors that train a small neural network.
pip install "alibi-detect[tensorflow]" # TensorFlow backend
pip install "alibi-detect[torch]" # PyTorch backend
pip install "alibi-detect[all]" # everything
Two details trip people up. First, quote the brackets on macOS and in zsh: without quotes the shell tries to expand the square brackets as a pattern and the command fails. Second, the pip extra is called torch, but when you later pass a backend to a detector the string is 'pytorch'. The names differ, and mixing them up is a classic first-week error.
Platform notes. On Linux everything works with pip or conda. On macOS pip works, but do not use the keops extra, which is not officially supported there. On Windows the docs advise installing and testing PyTorch before installing Alibi Detect if you want GPU use, and the keops extra is Linux only. If you are a beginner on any platform, the default install is the right start.
One more trap: the TensorFlow version the library allows is capped below 2.19 as of release 0.13.0. If you install a very new TensorFlow first and then add Alibi Detect, pip may fight you. If you only need the default install, you can ignore this entirely, which is one more reason to start there.
Now check that it works:
import alibi_detect
print(alibi_detect.__version__)
from alibi_detect.cd import KSDrift
You should see a version string such as 0.13.0 and no error on the import. The alibi_detect.cd package holds the drift detectors. The outlier detectors live in alibi_detect.od.
Your first drift check
Time to build something. We will fake a tiny, honest scenario. Imagine a model that takes five numeric features per customer. Your reference data is a thousand past customers. We will generate it randomly so you can run this anywhere, then show the detector a "today" batch that is normal, and another that is shifted.
import numpy as np
from alibi_detect.cd import KSDrift
rng = np.random.default_rng(42)
# Reference data: 1000 instances, 5 features, drawn from a standard normal
x_ref = rng.standard_normal((1000, 5)).astype(np.float32)
# Build the detector. Drift detectors take the reference data at construction.
cd = KSDrift(x_ref, p_val=0.05)
# A batch that comes from the same source as the reference
x_same = rng.standard_normal((200, 5)).astype(np.float32)
print("same source:", cd.predict(x_same)["data"]["is_drift"])
# A batch where every feature has been shifted by 0.5
x_shifted = (rng.standard_normal((200, 5)) + 0.5).astype(np.float32)
print("shifted: ", cd.predict(x_shifted)["data"]["is_drift"])
Run it. The shifted batch should print 1, meaning drift. The same-source batch should print 0 on most runs, and the reason it is not guaranteed is explained in the next two sections, so do not be alarmed if you ever see a 1 there.
Walk through what each part does, because every later example is a variation on this one.
The reference array has shape (1000, 5): one row per instance, one column per feature. This rows are instances, columns are features layout is how every detector in the library expects tabular data. The call to astype(np.float32) converts numbers to 32-bit floats. For the simple KS detector this is optional, but it is a good habit, because the deep-learning backends you will meet later are fussy about it.
KSDrift(x_ref, p_val=0.05) constructs the detector. Notice that there is no fit step. Unlike many scikit-learn objects, a drift detector receives its reference data at construction and is ready immediately. KS stands for Kolmogorov-Smirnov, a classic test that compares the shape of two distributions one feature at a time. It is cheap, it needs no neural network, and it is the right default for a first look at tabular data.
cd.predict(x) runs the test on the new batch and returns the result dictionary. We read ["data"]["is_drift"] to get the verdict.
x_shifted. At what size of shift does the detector stop reliably noticing? That boundary is the detector's sensitivity, and finding it by experiment is the best way to build intuition.Reading the result without fooling yourself
The verdict is one field in a larger dictionary. Print the whole thing once to see what is available:
preds = cd.predict(x_shifted)
print(preds["meta"])
print(preds["data"])
The meta part describes the detector: its name, its detector_type (which is 'drift' for these, 'outlier' or 'adversarial' for the others), its data_type, whether it is online, and the library version. This is useful when you log results from several detectors and need to know afterwards which one said what.
The data part holds the verdict and numbers. For KSDrift, you get is_drift, the threshold, and, if you ask, the p-value and the test statistic for each feature. You ask by passing flags to predict:
preds = cd.predict(x_shifted, drift_type="feature", return_p_val=True, return_distance=True)
print(preds["data"]["p_val"]) # one p-value per feature
print(preds["data"]["distance"]) # one K-S statistic per feature
Here drift_type="feature" tells the detector to report per feature rather than a single combined answer. Look at the two arrays. The p-value array says, for each column, how surprising the difference is: smaller means more surprising. The distance array holds the K-S statistic, which is the largest gap between the two distributions' cumulative curves for that column: bigger means a larger difference. The p-value tells you whether to believe a difference exists; the distance tells you how big it is. Both matter, because with a huge sample even a trivial difference produces a tiny p-value, and the distance is what tells you whether it is a trivial difference.
Now the misreading. A p-value of 0.03 does not mean "there is a 97 percent chance the data has drifted". It means that if there were no real difference, you would see a gap this large only about 3 percent of the time. That is a statement about the test under the assumption of no drift. It is a useful alarm trigger, but it is not a probability of drift, and it says nothing about whether the drift is large enough to hurt your model.
And the corollary that matters most for operations: with a threshold of 0.05, a perfectly stable system will still fire about one time in twenty per test, simply by chance. If you run a check every hour on stable data, you should expect roughly one false alarm a day, and with a batch of features each tested separately, more than that. This is not a bug. It is what the number means. The next section shows how the library helps you cope.
is_drift == 1 as an incident. Drift detection tells you the data looks different, not that the model is worse. Check the size of the difference, check whether it persists across several batches, and where possible check the model's real performance before waking anyone up.is_drift is 1. With p_val=0.05 and five features and the default correction, what fraction do you see? Compare it with 5 percent and think about why they might differ.Many features at once: p_val and correction
Our detector tested five columns. When you test several features separately and ask "did any of them drift", chance works against you. If each column has a 5 percent chance of a false alarm, the chance that at least one of five fires is much larger than 5 percent. With a hundred columns, a false alarm on any given batch becomes nearly certain.
Statisticians call this the multiple comparisons problem, and the standard remedy is to make the per-column threshold stricter in a principled way. KSDrift offers this through the correction argument, which takes 'bonferroni' or 'fdr':
cd = KSDrift(x_ref, p_val=0.05, correction="bonferroni")
Bonferroni is the conservative option. It divides the threshold by the number of features, so that the chance of any false alarm across all features stays near your chosen p_val. Its cost is that it can miss a real but modest drift in one column. FDR, short for false discovery rate, is the more permissive option. Instead of controlling the chance of any false alarm, it controls the expected fraction of flagged features that are false alarms, which suits cases where you are happy to follow up a few spurious flags in exchange for catching more real ones.
Which should you choose? A reasonable beginner rule is to start with Bonferroni when an alarm pages a human and false alarms are expensive, and with FDR when you are exploring and will look at the per-feature numbers anyway. Either is better than ignoring the issue.
The other control is alternative, which defaults to 'two-sided', meaning the test flags a change in either direction. You can pass 'less' or 'greater' to look for one direction only, but until you have a reason, leave the default.
You also decide how to read the result. With drift_type="batch" the detector combines the per-feature outcomes, after correction, into a single yes-or-no answer. With drift_type="feature" you see each column's own result. A practical pattern is to alert on the batch answer and log the feature answer, so that when an alert fires, the first thing you read tells you which columns moved.
correction="bonferroni". Compare the false alarm rates. Then shift a single feature by 1.0 and see which setting still catches it.Choosing and caring for your reference data
The detector is only as good as its idea of normal, so spend real thought here. Several decisions are hiding in the phrase "a sample of training data".
Size. Too few instances and the reference distribution is poorly estimated, so the detector fires on noise or misses real change. A few hundred to a few thousand instances is a sensible range for the simple tests. More is not always better: for the heavier kernel-based detectors, cost grows quickly with the reference size, as the later guides explain.
Representativeness. The reference should look like the data the model is expected to see in healthy operation, not just the training set. If your training data was rebalanced or filtered, the reference should reflect what production normally looks like, otherwise you will get alarms from day one and learn to ignore them.
Freshness and the refresh trap. Reality moves, so some people refresh the reference with recent data. The library supports this through update_x_ref, which takes forms such as {'last': 1000} or {'reservoir_sampling': 1000} to keep a rolling or sampled reference. This is powerful and dangerous. If the reference always follows today's data, a slow drift is absorbed and never alerts, because each day looks like the previous day. Only update the reference when you deliberately want the detector to accept the new normal. As a beginner, keep a fixed reference and revisit it on purpose.
Privacy. The reference is a copy of part of your training data. If that data contains personal information, so does the file you save. In the Gulf and in Egypt, employers in banking, telecoms and health care often have strict rules about where such data may be stored and who may touch it. Treat the reference file with the same care as the training set: restricted access, and no casual copies to a laptop or a public notebook.
Order and structure. Rows should be independent instances. If your data is a time series in which neighbouring rows are strongly related, a basic per-feature test will overstate its confidence. That is an advanced topic, but be aware that the simple tests assume the rows are not copies of each other.
Outliers and adversarial inputs: the other two families
Drift detection answers a question about a whole batch. Sometimes the question is about one instance: is this single request so strange that I should not trust the model's answer? That is outlier detection, and it lives in alibi_detect.od. The catalogue includes Isolation Forest, Mahalanobis distance, auto-encoders (AE and VAE), the AEGMM and VAEGMM variants, likelihood ratios, a Prophet-based detector for time series, a spectral residual detector, and a sequence-to-sequence detector.
The idea differs from drift in one important way. Many outlier detectors are fitted on normal data first, and then they need a threshold on an instance score. You choose the threshold so that a chosen share of normal data would be flagged, for example the 99th percentile of scores on clean data. The library provides an infer_threshold method for this on the detectors that need it. The result of predict is again a dictionary, with an is_outlier flag per instance rather than a single is_drift flag.
Because the exact constructor arguments vary a lot between these detectors, and they change the examples considerably, this guide does not reproduce them. Use the outlier methods page in the Sources, pick the detector that matches your data type, and copy its example. For tabular data, Isolation Forest and Mahalanobis distance are the lightest starting points and need no deep-learning backend.
Adversarial detection is the third family, for inputs crafted on purpose to fool a model, usually images with tiny, human-invisible changes. The library offers an Adversarial Auto-Encoder detector and a Model Distillation detector. This is a specialised area; as a beginner, it is enough to know it exists and where it lives.
Why not start with outliers? Because drift gives a more actionable first signal for a deployed model: it is cheap, it needs no training for the simple tests, and it matches the operational question "should I worry about my model today?". Outlier detection becomes valuable when single bad requests matter, such as fraud, safety, or rejecting garbage input before it reaches an expensive model.
Images, text and big tables: preprocessing in one paragraph
The simple KS test works on numeric columns. Real data is often richer. An image has thousands of pixels; a sentence is not a number at all. Testing raw pixels directly works poorly, because pixel space is huge and noisy and almost any two image batches look "different" there.
The standard remedy is a preprocessing function, passed as preprocess_fn. Its job is to reduce each instance to a compact, meaningful representation before the statistical test runs. Typical choices are the output of an encoder network, or the probabilities your own classifier assigns, since a shift in what your model sees is exactly what you care about. The detector then compares the compressed representations instead of the raw inputs.
This is where the optional backends come in: building and running an encoder needs TensorFlow or PyTorch. It is also where the library is most powerful, because the same drift test now works on images and text. You do not need it for today's tabular work, and the mid-level guide takes this up properly. The takeaway for now is the shape of the idea: reduce first, then test.
Offline versus online, briefly
Everything so far was offline: a batch arrives, you compare it with the reference, you get a verdict. This fits a nightly job that checks yesterday's traffic, which is the most common setup in practice.
Some systems need a verdict as data streams in, one instance at a time. Online detectors exist for this, with names such as MMDDriftOnline, LSDDDriftOnline, CVMDriftOnline and FETDriftOnline. They keep a sliding window and a counter of how many instances they have seen. Two settings control them. The expected run-time, ert, is the average number of instances you are willing to wait before a false alarm on stable data. A larger value means fewer false alarms but slower detection. The window size, window_size, sets how much recent data each test looks at: small windows react fast to severe drift, larger windows are more sensitive to subtle drift.
Online detectors have a one-time setup cost, because the constructor simulates many runs to calibrate its threshold for your chosen ert and window_size. That can take noticeable time on a large reference set, so you do it once at start-up, not on every request. Online detectors are also the ones where state matters: the detector remembers where it is, and after an alert you reset it with reset_state().
If you are new, do not start here. Master the offline pattern first. The mid-level and senior guides cover online detectors, their state handling and their operational traps.
Saving a detector and loading it back
A detector you build in a notebook is useless to a scheduled job on another machine unless you can store it. Alibi Detect provides two functions for this:
from alibi_detect.saving import save_detector, load_detector
save_detector(cd, "./my_detector/")
cd = load_detector("./my_detector/")
The first call writes a folder; the second reads it and returns a ready detector. Try the round trip and call predict on the loaded detector to prove it behaves the same.
There are two storage formats, and understanding them avoids a lot of confusion. The config format is the default for drift detectors. It writes a human-readable file called config.toml plus supporting files such as the reference data as a .npy array. Because it is readable text, you can open it, edit a setting, and keep it in version control, which makes changes reviewable. A minimal config looks like this:
name = "KSDrift"
x_ref = "x_ref.npy"
p_val = 0.05
The name line is the detector class and every other line is a constructor argument. The example above is the shape of the file; open the one your own save produces to see exactly what your version writes.
The legacy format uses a Python serialisation library called dill. It is the only format available for outlier and adversarial detectors, and you can request it for drift detectors with save_detector(cd, path, legacy=True). It is not supported for some detectors, including the online ones. One further note from the 0.13.0 release: saving in legacy format with recent TensorFlow needs the environment variable TF_USE_LEGACY_KERAS=1 so that TensorFlow uses Keras 2.
Two safety points every beginner should absorb now. First, match versions. The docs warn that cross-version loading is not guaranteed: loading a detector saved by a different Alibi Detect version issues a warning, and the docs strongly recommend matching the library, Python and dependency versions when you save and load. Pin them in your requirements file. Second, only load detector files from sources you trust. Dill-based files can execute arbitrary code when loaded, exactly like a pickle file, so an untrusted detector folder is as dangerous as an untrusted program.
alibi_detect.__version__ next to every saved detector.KSDrift detector, open the folder, read the config.toml, then load it in a fresh Python process and confirm that it gives the same verdict on the shifted batch.When things go wrong: configuration and common errors
Most beginner problems fall into a short list. Reading them in advance saves an evening.
The backend is missing. If you build a detector that needs a deep-learning framework, for example by passing a backend argument, and the matching extra is not installed, the construction fails, typically with an import-related error. The cure is to install the extra, alibi-detect[tensorflow] or alibi-detect[torch]. Remember the naming quirk: the extra is torch but the backend string is 'pytorch'.
pip cannot resolve a TensorFlow version. Alibi Detect caps TensorFlow below 2.19 in release 0.13.0. If the environment already holds a newer TensorFlow, the resolver complains or pulls an older one. Use a clean environment, or choose the PyTorch backend.
Binary incompatibility with NumPy. If you see an error that talks about binary incompatibility or a changed dtype size, two compiled packages were built against different NumPy generations. Align the versions of NumPy and the libraries that depend on it, or rebuild the environment from scratch.
A hang when importing both TensorFlow and PyTorch. PyTorch 2.0.0 to 2.0.1 could hang in this situation; the docs say it is fixed in 2.1.0 and later.
Shape and type problems in predict. The new data must have the same feature layout as the reference: same number of columns, same order. Offline detectors expect a batch, so passing a single row gives a meaningless test. Cast the arrays to float32 if a deep backend complains.
A loaded detector warns about versions. This is the expected behaviour described above. Match versions, or re-save the detector with the current library.
Alerts that never come, or never stop. These are not errors. A detector that never fires may have a reference that drifted along with the data, or a threshold that is too strict. A detector that fires constantly may have a reference that does not represent normal, or too many features and no correction. Reread the sections on reference data and correction before blaming the library.
When you meet an error not on this list, read the last line of the traceback first: it names the real problem. Then search the repository's issue tracker, and check the release notes in the changelog for something that changed in your version. A habit worth building early: always report your versions of Python, Alibi Detect, NumPy and any deep-learning library when you ask for help, because most problems in this area are version problems.
predict with an array that has the wrong number of columns. Read the error, find the line that names the cause, and write down what it tells you.A short playbook for when an alert fires
Sooner or later your detector will say is_drift == 1. What you do in the next ten minutes decides whether the monitor earns trust or gets muted. Here is a calm sequence a beginner can follow, in order, and each step is cheaper than the one after it.
First, check that the check itself is sound. Did the batch contain enough rows? A batch of twelve instances gives a test very little to work with. Did the columns arrive in the same order as the reference? Did a file fail to load and get replaced by zeros or empty values? A surprising share of alerts are caused by a broken pipeline step rather than by the world changing, and finding that is good news, because the fix is ordinary engineering.
Second, look at which features moved. Run the per-feature report with drift_type="feature" and sort by distance. If one column carries nearly all the difference, ask what feeds that column. A unit change (kilometres to metres), a new category code, a join that started returning nulls, or a timezone shift each leave a clear fingerprint in a single column. If many columns moved a little together, the cause is more likely a change in who your users are, or a change in how the data is collected for everyone.
Third, look at the size, not just the p-value. Compare the K-S distances with those from earlier healthy batches. A distance that is only slightly above the usual range with a tiny p-value usually means the batch was large and the test was very powerful, not that anything serious happened. A distance several times the usual range deserves attention.
Fourth, check persistence. Run the detector on the next batch, and the one after. A single alert among many quiet days is the one-in-twenty chance you read about earlier. Three alerts in a row, all pointing at the same column, is a pattern.
Fifth, connect it to the model. If labels have arrived for any of the recent data, measure the model on them. If you have no labels yet, look at the distribution of the model's own predictions: a model that suddenly predicts one class far more often than before is telling you something. Drift in the inputs raises a question about the model, and only performance evidence answers it.
Finally, decide, and write it down. The options are to do nothing and keep watching, to fix the upstream problem, to retrain on recent data, or to update the reference because the new data is the new normal. Record the choice and the reason. Teams that write these decisions down learn which kinds of alert matter, and over a few months they tune their thresholds with evidence instead of instinct.
Where Alibi Detect fits in a real system
A detector on a laptop is a learning tool. In a real system it is one part of a loop. The common shape is a scheduled batch job: a scheduler such as Airflow starts a script every night. The script loads the saved detector with load_detector, reads the last day of production inputs from wherever they are logged, calls predict, and writes the verdict, the p-values and the distances to a store. A dashboard in Grafana or an alert rule in Prometheus turns those numbers into something a person sees.
Alibi Detect also has an official deployment path with Seldon Core and Knative eventing, in which model requests are logged as events and fanned out to detector services. The documentation for that path describes Seldon Core version 1, so check how it fits current Seldon and KServe setups before copying it. For a beginner, the batch job is the right model: it is simple, it is easy to test, and it uses only what you already learned.
Packaging the job in a container makes it reproducible; the Docker guide covers that. And if you track model versions in MLflow, store the detector folder and the reference file with the model version they belong to. A detector separated from the model it protects is a common source of confusion later.
Two habits make the loop trustworthy. Log the metadata (the detector name and library version from meta) alongside every result, so you can tie any alert to the code that raised it. And alert on sustained drift, not a single batch, for the reason covered earlier: a lone is_drift == 1 is often chance.
Putting it all together
Here is one small end-to-end project that uses everything above. It builds a reference, saves a detector, then simulates a week of daily batches in which drift begins on day five, and prints a daily report. Run it and watch the detector stay quiet and then fire.
import numpy as np
from alibi_detect.cd import KSDrift
from alibi_detect.saving import save_detector, load_detector
rng = np.random.default_rng(7)
# 1. Reference data: 2000 healthy instances, 8 features
x_ref = rng.standard_normal((2000, 8)).astype(np.float32)
# 2. Build, save, and reload (as a scheduled job would)
detector = KSDrift(x_ref, p_val=0.05, correction="bonferroni")
save_detector(detector, "./kst_detector/")
detector = load_detector("./kst_detector/")
# 3. Simulate seven days of production batches
for day in range(1, 8):
batch = rng.standard_normal((500, 8)).astype(np.float32)
if day >= 5:
batch[:, 2] += 0.6 # feature 2 slowly goes wrong from day 5
preds = detector.predict(batch, drift_type="feature",
return_p_val=True, return_distance=True)
p_vals = preds["data"]["p_val"]
dist = preds["data"]["distance"]
worst = int(np.argmax(dist))
print(f"day {day}: drift={preds['data']['is_drift']} "
f"worst feature={worst} distance={dist[worst]:.3f} p={p_vals[worst]:.4f}")
Read the output the way an operator would. Days one to four should show no drift, and from day five the worst feature should be feature 2 with a large distance and a tiny p-value. Notice what you did without effort: you built the detector once, saved it, reloaded it as a separate job would, and produced a report that names the column responsible. That is the core of a real drift monitor.
Now extend it, as an exercise rather than a copy-paste. Make the shift on day five small (0.1) and see whether you still catch it with 500 instances per batch. Add correction="fdr" and compare. Write each day's line to a CSV file instead of printing it. Wrap the detection in a function that returns a dictionary of the fields you want to log, including the library version from preds["meta"]. Each change teaches a different lesson about sensitivity, correction or logging.
What you can now do, and what comes next
You can now explain what drift is and why a model fails silently without it. You know the three detector families and the four nouns: detector, reference data, test with a threshold, and result. You can install the library in a clean environment and check it. You have run a Kolmogorov-Smirnov drift check, read the verdict, p-values and distances, and you know that a p-value is not a probability of drift. You can control false alarms with p_val and correction, you understand the risk of refreshing the reference, you can save and reload a detector, and you can name the common errors and their causes. You also know the licence, which matters more than any technical detail when the work moves from learning to shipping.
What the next level adds: heavier detectors such as MMD and learned-kernel drift, preprocessing with encoders so that images and text can be tested, the details of the config files and the registry, the online detectors in depth, and testing and performance. After that, the senior guide covers designing monitoring for a team, failure modes at scale, security of serialised detectors, and when to choose something else. If your interest is the wider monitoring picture, read Evidently and NannyML next, and compare how they answer the same questions.
The most valuable thing you can do now is to use a detector on data you care about. Take a public dataset, split it by time or by category so that the halves differ, and see what the detector says. Intuition for drift comes from watching real data misbehave.
Sources
- Alibi Detect documentation
- Getting started and installation
- Algorithm overview
- Drift detection overview
- Kolmogorov-Smirnov drift detector
- Online MMD drift detector
- Outlier detection methods
- Adversarial detection
- Saving and loading detectors
- Config files
- Deployment with Seldon Core and Knative
- GitHub repository
- Changelog
- Licence
- PyPI package page