Skip to content
Back to student guides
NannyMLMLOpsMonitoring3 levels95 sectionsCovers NannyML 0.13

The Complete NannyML Guide

Estimate model performance without labels and detect drift with NannyML. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
17sections
13examples

This is part one of three. It covers everything you need to do real work with NannyML, not a teaser. By the end you can install the library, load a reference set and a production set, estimate how well a classification model is doing before any labels arrive, find which input columns have drifted, compare your estimates with the real numbers once labels show up, and read every alert the library raises without guessing what it means. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and the ideas only stick once you have watched your own chunks turn into a chart and seen an alert fire.

This guide is written against NannyML 0.13.1, the latest release on PyPI at the time of writing, which supports Python 3.9 to 3.12. A section near the end is honest about the project's maintenance status, because you should know it before you build on the library.

The problem: a model that fails without telling anyone

A model in production is a strange kind of software. A web service that breaks returns errors, and somebody notices within minutes. A model that breaks keeps returning answers. The answers are well-formed, they arrive on time, the HTTP status is 200, and every dashboard about latency and uptime stays green. Meanwhile the answers have quietly become worse, because the world the model was trained on is no longer the world it is being asked about.

Here is a concrete case. A bank in Cairo trains a model to predict whether a loan applicant will repay. It works well on last year's applicants. Then a new salary-transfer partnership brings in a different kind of customer, interest rates move, and the typical income in the application form shifts. The model still produces a probability of default for every applicant. Nobody sees an error. Three months later the risk team notices that defaults are higher than expected, and by then the bad loans are on the books.

The slow part of that story is the labels. You only know whether a borrower repaid after months have passed. For fraud, you often learn the truth after a chargeback dispute weeks later. For a churn model, you learn after the customer has already left. This delay between a prediction and the truth is the central difficulty of monitoring, and it is the reason a plain "compute accuracy every day" job does not work for most real models: the data for it does not exist yet.

NannyML exists for exactly this gap. It is an open-source Python library whose headline idea is to estimate the performance of a model without access to the true labels, and to tell you when the estimate drops. It also detects changes in the input data (called drift) and, importantly, helps you separate drift that matters from drift that does not.

Try it. Think of one model you know (your own, a course project, or a public one) and write down how long you would wait to learn whether a single prediction was right. If the answer is more than a day, you have just described the situation NannyML is built for.

What came before NannyML, and why it was not enough

Before libraries like this, teams monitored models in two common ways, and both have a flaw you can feel.

The first is monitoring the outputs and inputs with simple statistics: track the average prediction, track the mean of each input column, and alert if they move. This is cheap and it is a reasonable start. The flaw is that it produces a flood of alerts. Data always moves a little. A feature's average changes on Fridays, at month end, during Ramadan, in summer. Most of those movements do not hurt the model, so the team learns to ignore the alerts, and then the one that mattered is ignored too.

The second is waiting for labels and computing realized performance. This is the gold standard because it measures the thing you actually care about. The flaw is the delay described above. By the time the number is available, the damage is done.

Drift detection tools improved the first approach by applying proper statistical tests to distributions, and several of them are good. But a statistical test answers the question "did this distribution change?" and not the question "did my model get worse?". Those are different questions. A feature can shift dramatically and the model can be unaffected, if the model hardly uses that feature or the shift stays inside the range the model already knows. The reverse also happens: a small shift in an important feature can wreck accuracy.

NannyML's contribution is to put performance first. You start from an estimate of how the model is doing, and you use drift detection as the tool for explaining a drop in that estimate, not as the alarm itself. In the docs' own framing, the workflow is to estimate performance, check whether there is a problem, and then look for the cause in the data.

If you want to compare with a tool that focuses on rich drift reports and dashboards, read the Evidently guide at /student-guides/evidently, and for a library centred on statistical detectors, /student-guides/alibi-detect. Seeing two approaches side by side is one of the best ways to understand what each is for.

The mental model: four nouns

NannyML has a lot of classes, but a handful of ideas carry nearly everything. Learn these four nouns and the rest of the API will look familiar.

REFERENCEdata where the model was good
→
FITlearn baselines and thresholds
→
ANALYSISproduction data to monitor
→
CHUNKSresults per slice, with alerts

Reference data is a period where you trust the model. It is usually your held-out test set, or an early stretch of production for which you already have labels. It must contain the true outcomes, because the library uses it to learn what "normal" looks like: how much the metrics wobble from one slice to the next, and how the data usually looks. Think of it as the calibration measurement of an instrument.

Analysis data is the data you are monitoring: recent production inputs and the model's predictions. For performance estimation it does not need labels. That is the whole point. If labels exist for some analysis data, you can use them later to check the estimates, but they are not required to produce the estimates.

Chunks are the unit of everything. NannyML does not produce one number for your whole analysis set. It cuts the data into groups, either by number of rows, by number of groups, or by calendar period such as a week or a quarter, and computes each metric once per chunk. The result is a time series. A chunk has a key, a start and end, and a label saying whether it came from the reference or the analysis period. Chunking is the single idea most beginners underestimate, so a later section is devoted to it.

Thresholds and alerts turn numbers into decisions. During fitting on the reference data, NannyML measures how much each metric normally varies across reference chunks. By default the thresholds sit three standard deviations around the reference mean. In the analysis period, a chunk whose value falls outside those bounds is marked with an alert. So an alert does not mean "this is bad in an absolute sense". It means "this chunk is outside the range the reference data taught us to expect".

Two more names will appear immediately. Calculators and estimators are the objects you create: CBPE estimates performance for classification, UnivariateDriftCalculator measures drift column by column, and PerformanceCalculator computes realized performance. All of them follow the same rhythm you may know from scikit-learn: create the object, call fit() on reference data, then call estimate() or calculate() on analysis data. The result is a results object that you can filter, turn into a DataFrame, or plot.

Try it. Without running any code, sketch a timeline for a model you know. Mark where your reference period would end and your analysis period would begin, and decide whether you would chunk by week or by a fixed number of rows. You will compare your choice with the guidance in the chunking section.

Installing NannyML and checking the setup

NannyML is a normal Python package. The one decision that matters is the Python version: 0.13.1 declares support for Python 3.9 through 3.12. Python 3.13 is not listed, and on a newer interpreter pip may refuse to find a compatible release. If you have a choice, use 3.11 or 3.12.

Always create a virtual environment. A monitoring library pulls in pandas, scikit-learn, LightGBM, Plotly and others, and you do not want to disturb the rest of your machine.

BASH
python3.12 -m venv .venv
source .venv/bin/activate
pip install nannyml

On Windows, the recommended route for a course like this is WSL2, which gives you a Linux shell and avoids most binary-wheel surprises. If you prefer conda, the documentation also shows conda install -c conda-forge nannyml.

Verify the install by importing the package and printing its version:

BASH
python -c "import nannyml; print(nannyml.__version__)"

You should see 0.13.1. If you see something older, an earlier environment is shadowing the new one, so check which python you are running.

NannyML depends on LightGBM, which needs a native OpenMP runtime. On a Mac that means Homebrew's libomp; on a slim Linux container it means the libgomp1 package. If the import fails with a message about a missing dynamic library, install the matching runtime and try again:

BASH
brew install libomp                      # macOS
sudo apt-get install -y libgomp1         # Debian or Ubuntu

There is also an optional extra. Since version 0.12, the database dependencies are not installed by default. You only need pip install "nannyml[db]" if you plan to write results into a SQL database, which is not a beginner task.

If you work behind a company network that inspects TLS traffic, pip install can fail with CERTIFICATE_VERIFY_FAILED. Do not turn verification off. Ask your IT team for the proxy's root certificate and point pip at it with the PIP_CERT environment variable.

Trap. Installing with a bare pip that belongs to a different Python than the one you run. If import nannyml fails right after a successful install, use python -m pip install nannyml, which guarantees both come from the same interpreter.
Try it. Create the virtual environment, install NannyML and print the version. If the import fails, read the last line of the error: it names the missing library, which is almost always the OpenMP runtime.

Your first project: estimating performance without labels

Now the main event. NannyML ships with sample datasets so you can practice without preparing your own. We will use the one from the official quickstart: a model that predicts whether a person in Massachusetts is employed, built from US census records. The data is split into a reference set from 2015 and an analysis set covering 2016 to 2018. A third table holds the true outcomes for the analysis rows, which we will keep aside to play the role of "labels that arrive later".

PYTHON
import nannyml as nml

reference_df, analysis_df, analysis_targets_df = nml.load_us_census_ma_employment_data()

print(reference_df.shape, analysis_df.shape)
print(reference_df.columns.tolist())
print(reference_df.head())

Spend a minute on that output. Among the columns you will see the input features, a timestamp, the model's predicted_probability (how sure it is that the person is employed), its prediction (the final yes or no), and employed, the true outcome. The reference table includes the true outcome. The analysis table does not, and that is deliberate.

The estimator for a binary classifier is called CBPE, short for Confidence-Based Performance Estimation. The idea is simple to say. If the model's predicted probabilities are well calibrated, then a prediction of 0.9 is right about 90 percent of the time. If you know the probability attached to every prediction, you can work out the expected number of true positives, false positives and so on, without seeing a single label, and from those compute a metric such as ROC AUC. The technique rests on two assumptions you must remember: the model's probabilities are calibrated, and the relationship between inputs and outcomes has not changed (no concept drift). We return to these in the limits section.

PYTHON
estimator = nml.CBPE(
    problem_type='classification_binary',
    y_pred_proba='predicted_probability',
    y_pred='prediction',
    y_true='employed',
    timestamp_column_name='timestamp',
    metrics=['roc_auc'],
    chunk_size=5000,
)

estimator.fit(reference_df)
results = estimator.estimate(analysis_df)

Read the constructor argument by argument, because each one is a promise about your data.

problem_type tells NannyML which family of metrics applies. The values are classification_binary, classification_multiclass and regression. y_pred_proba is the name of the column holding predicted probabilities, y_pred the column with the predicted class, and y_true the column with the real outcome (needed in the reference data only). timestamp_column_name points at the date column, which gives your results real dates on the chart axis. metrics is the list of metrics to estimate, here only ROC AUC, and chunk_size is the number of rows per chunk.

Then the rhythm: fit() learns from the reference data, and estimate() applies the estimator to the analysis data. Calling estimate() before fit() raises an error telling you the estimator was not fitted.

Finally, look at the output as a table and a picture:

PYTHON
print(results.filter(period='analysis').to_df())
results.plot().show()

The plot opens in your browser, or inline in a notebook. You will see a line of estimated ROC AUC over chunks, a shaded confidence band around it, and horizontal threshold lines. When the line leaves the thresholds, the chunk is highlighted. That highlight is your alert.

Shortcut. If the plot does not show in a plain script, save it instead with results.plot().write_html("estimated_roc_auc.html") and open the file. The plots are Plotly figures, so they are interactive.
Try it. Run the code above in a notebook or script. Then change chunk_size from 5000 to 2000 and rerun. Notice how the line becomes more jagged and the shaded band wider. You have just seen sampling error, which the next section explains.

Reading the results: values, bands, thresholds and alerts

A results object is more than a chart. Each row of the DataFrame describes one chunk, and knowing the columns is how you stop treating the library as a black box.

For an estimated metric, a row carries the value (the estimate), the sampling error, the confidence boundaries (upper and lower), the thresholds (upper and lower) and a boolean alert. Chunk metadata sits next to them: the chunk key, start and end index, start and end date, and the period, which is reference or analysis. Call .to_df() and print .columns to see exact names in your installed version, since the DataFrame groups them by metric.

These columns answer different questions, and mixing them up is a classic beginner error.

Sampling error is the statistical noise in a metric computed from a limited number of rows. If you compute a ROC AUC from 200 rows, it could easily land a few points above or below the true value just by luck. NannyML estimates this noise for each chunk and draws a confidence band of plus or minus three sampling errors around the value. A small chunk gives a wide band, a big chunk a narrow one. The practical lesson is that a wobble inside the band is not evidence of anything.

Thresholds are different. They are not about noise in one chunk; they are about the normal range of the metric learned from the reference data. By default they are set at three standard deviations around the reference mean of the metric across reference chunks.

An alert is raised when the estimated value is outside the thresholds. Notice what this means: alerts are relative to what the reference taught the library. A reference set that is too short, or chunks that are too small, produce thresholds that are too narrow or too noisy, and you will see alerts on healthy data. A reference that is unrepresentative gives the opposite problem.

You can pull out only the alerting chunks with ordinary pandas. With multilevel=False the DataFrame is flat and easier to filter:

PYTHON
df = results.filter(period='analysis').to_df(multilevel=False)
print(df.columns.tolist())
alert_cols = [c for c in df.columns if c.endswith('alert')]
print(df[df[alert_cols[0]]][['chunk_key']])

Beyond the main plot, results.filter(period='analysis') limits the view to production chunks, and results.filter(metrics=['roc_auc']) limits it to one metric. These filters return a new results object, so you can chain them with .plot() or .to_df().

Another useful habit is looking at the reference period on the plot too. You want to see the reference chunks sitting comfortably inside the thresholds. If the reference chunks themselves already alert, the setup is wrong before you even look at production: the reference is too small, the chunks too small, or the reference data was not as clean as you thought.

Try it. Print the column names of your results DataFrame, then write one sentence for each saying whether it describes noise, normal range, or a decision. Then plot the full results (without the period filter) and check that the reference chunks are all inside the thresholds.

Chunking: the setting that changes every result

Because every number in NannyML is per chunk, how you chunk decides what you can and cannot see. The documentation describes three ways to define chunks, and you pick one per calculator.

By size with chunk_size=5000: every chunk has that many rows. Use this when your traffic is steady and you want comparable statistical power per chunk. If the last group has fewer rows, it is kept as a smaller chunk by default, which is worth knowing because a tiny final chunk has a wide band.

By period with chunk_period, for example chunk_period='W' for weeks or 'M' for months. The documented aliases include 'D', 'W', 'M', 'Q' and 'A' among others. This needs a timestamp column, because the library has to know which period each row belongs to. Use this when your business thinks in calendar terms: a weekly report, a monthly review. The cost is that chunks have different sizes, so a quiet week gets a wider uncertainty band than a busy one.

By count with chunk_number, for example chunk_number=9: the data is divided into that many chunks of equal size. Use it when you care about having a fixed number of points on a chart.

If you pass none of them, NannyML falls back to a count-based split into ten chunks. That default is fine for a first look and a poor choice for monitoring, so set your own.

How do you choose a size? Think about two opposing forces. Larger chunks give narrower confidence bands and smoother lines, so alerts are more trustworthy, but they average over a longer time, which delays detection and can hide a short incident. Smaller chunks react faster but are noisier, and the sampling error can be wider than the effect you are looking for. A reasonable beginner rule: choose the smallest chunk for which the confidence band is narrow enough that a drop you care about would clearly leave it. For a fraud model that needs a daily answer but sees 800 transactions a day, chunking by day may give a band too wide to be useful, and a three-day or weekly chunk is the honest choice.

You also need enough chunks in the reference period. Thresholds come from the spread of the metric across reference chunks, and a spread estimated from two or three chunks is unreliable. If your reference set is small, that is a reason to use smaller chunks there, not a reason to shrink the band by hand.

PYTHON
estimator = nml.CBPE(
    problem_type='classification_binary',
    y_pred_proba='predicted_probability',
    y_pred='prediction',
    y_true='employed',
    timestamp_column_name='timestamp',
    metrics=['roc_auc'],
    chunk_period='M',
)

Passing both chunk_size and chunk_number is an error, since they describe incompatible plans. Pick one chunking style per calculator.

Trap. Using chunk_period without a timestamp_column_name. The library has no way to assign rows to weeks or months, so it raises an error about the missing timestamp. Add the column name or switch to chunk_size.
Try it. Rerun your estimator three times with chunk_size=1000, chunk_size=5000 and chunk_period='M'. For each, count how many chunks you get, and write down how wide the confidence band looks. Decide which you would trust for a weekly review.

Detecting data drift, one column at a time

An alert on estimated performance tells you that something is wrong. The next question is why, and that is where drift detection comes in. Data drift means the distribution of an input column in production differs from the distribution in the reference data. The tool for beginners is the univariate drift calculator: "univariate" means it examines one column at a time.

Start by choosing which columns to examine. The sample data has many features, and for a first pass it is fine to look at a few. Selecting them from the frame's own columns avoids typing a name that does not exist:

PYTHON
non_features = ['timestamp', 'predicted_probability', 'prediction', 'employed']
feature_columns = [c for c in reference_df.columns if c not in non_features]
print(feature_columns)

Then build the calculator. It follows the same create, fit, calculate pattern:

PYTHON
drift_calculator = nml.UnivariateDriftCalculator(
    column_names=feature_columns[:4],
    timestamp_column_name='timestamp',
    continuous_methods=['kolmogorov_smirnov', 'jensen_shannon'],
    categorical_methods=['chi2', 'jensen_shannon'],
    chunk_size=5000,
)
drift_calculator.fit(reference_df)
drift = drift_calculator.calculate(analysis_df)

Two arguments deserve an explanation. NannyML needs to know whether each column is continuous (numbers on a scale, like age) or categorical (labels, like a marital status code). It infers this from the data type, and a numeric column that really holds categories can be declared with treat_as_categorical. The continuous_methods and categorical_methods lists choose the statistical measures to apply to each kind.

The measures answer "how different are these two distributions?" in different ways. For continuous columns the documentation lists kolmogorov_smirnov, jensen_shannon, wasserstein and hellinger. For categorical columns it lists chi2, jensen_shannon, l_infinity and hellinger. The Kolmogorov-Smirnov and chi-squared options are classic statistical tests that produce a statistic and a p-value style signal, and are sensitive to small changes in large samples. Jensen-Shannon distance, Wasserstein distance and Hellinger distance are distance measures, which give a magnitude of change instead of a yes or no. If you only remember one rule: with very large chunks, tests flag changes too small to matter, and a distance measure lets you judge size. Using Jensen-Shannon alongside a test is a sensible default.

Now look at the output:

PYTHON
print(drift.filter(period='analysis', methods=['jensen_shannon']).to_df(multilevel=False))
drift.filter(methods=['jensen_shannon']).plot(kind='drift').show()
drift.filter(methods=['jensen_shannon']).plot(kind='distribution').show()

The drift plot is a line chart of the measure per chunk with thresholds and alert markers, one panel per column. The distribution plot shows what the column actually looks like in each chunk: a joyplot for continuous columns and stacked bars for categorical ones. Always look at both. The line tells you when the measure crossed the line, and the distribution plot shows you what changed, such as an age group growing, a category appearing more often, or the middle of the range shifting. Without the second plot you know that a column drifted and not what that means for your business.

This is the part where discipline matters. Drift in a column is a suspect, not a verdict. Many columns will alert, and few of them are why the model got worse. The correct sequence is: look at the estimated performance first, and when it drops, use the drift results to find candidate causes. NannyML even offers rankers that order drifting columns by how many alerts they raised or by how strongly they correlate with the performance change, but those are mid-level tools. For now, the habit is enough: performance drops first, drift explains.

Try it. Run the drift calculator on four features. Find one that alerts and look at its distribution plot. Write one sentence describing in plain language how its distribution changed between the reference and the analysis period.

Realized performance: checking the estimate when labels arrive

An estimate that you never compare with reality is a belief, not a measurement. Eventually labels arrive, and NannyML lets you compute the real metric with PerformanceCalculator. It uses the same pattern, but it needs the true outcomes for the analysis rows.

In the sample data the targets live in a separate table. Join them to the analysis predictions using the index, which is how the two frames line up:

PYTHON
analysis_with_targets = analysis_df.merge(
    analysis_targets_df, left_index=True, right_index=True
)

calculator = nml.PerformanceCalculator(
    problem_type='classification_binary',
    y_pred_proba='predicted_probability',
    y_pred='prediction',
    y_true='employed',
    metrics=['roc_auc'],
    timestamp_column_name='timestamp',
    chunk_size=5000,
)
calculator.fit(reference_df)
realized = calculator.calculate(analysis_with_targets)

print(realized.filter(period='analysis').to_df())
realized.plot().show()

Use the same chunk_size as the estimator. If the chunks differ, the two curves describe different slices of time, and you cannot compare them point by point. Then plot the realized line next to the estimated one, or print the two DataFrames and look at the value columns for matching chunks.

What do you expect to see? When the CBPE assumptions hold, the estimate should track the realized line closely, and both should fall in the same chunks. That agreement is the evidence that the method works on your data. When they diverge, the divergence is itself information. Typical causes are a model whose probabilities are not calibrated, or a change in the relationship between inputs and outcomes that estimation cannot see because no labels are involved.

This comparison also builds trust with the people who read your reports. A team lead is far more likely to act on "the estimate said performance fell in March, and the labels that arrived in June confirmed it" than on a drift chart. In practice, you will run estimation continuously and realized calculation whenever a batch of labels arrives, and keep both.

Shortcut. Treat realized performance as your regression test for the monitor itself. Every few weeks, compare estimated and realized values for the chunks that now have labels, and write down the gap. A gap that grows is the signal that your calibration assumption is slipping.
Try it. Compute realized ROC AUC on the sample data with the same chunk size as your estimate. Plot both. In which chunks do they agree most closely, and is there any chunk where the estimate was visibly too optimistic?

Other metrics, other problem types, and the limits you must remember

So far we used ROC AUC. For binary classification, CBPE supports a family of metrics. The documentation lists, among others, roc_auc, accuracy, f1, precision, recall, specificity, average_precision, confusion_matrix and business_value. You pass several at once in a list and get one result per metric, so a single estimator can feed a whole dashboard. Choose metrics that match how the model is used: for a fraud model with rare positives, precision and recall tell you far more than accuracy, which stays high even for a useless model.

For problems that are not binary classification, NannyML changes the estimator. Regression models use DLE, Direct Loss Estimation, which trains an internal LightGBM model to predict your model's error, then aggregates the predicted errors into metrics like mean absolute error. Multiclass classification uses problem_type='classification_multiclass', where the probability columns are given as a mapping from class to column. Both appear in the Mid-level guide. The mental model you learned here, reference, fit, analysis, chunks and thresholds, carries over unchanged.

Now the limits. They matter more than any feature, because a monitoring tool that is trusted beyond its limits is more dangerous than none.

Estimation assumes calibrated probabilities. If your model says 0.9 when the true rate is 0.7, the estimate inherits that error. Check calibration on your reference set before trusting CBPE, and consider calibrating the model's scores first.

Estimation assumes no concept drift. Concept drift means the rule linking inputs to outcomes changes, for example, because customer behaviour changes in response to a new policy. Estimation looks only at inputs and predictions, so it cannot see this, and the estimate may stay flat while real performance collapses. This is why realized performance, once labels arrive, remains essential.

Covariate shift can push the model outside what it knows. If production inputs land in regions that the reference never covered, the calibration learned there no longer applies. Drift detection helps you notice this.

Estimates are early warnings, not proof. Treat an alert as a prompt to investigate. Confirm with labels, with a sample you review by hand, or with a business metric.

Trap. Reading a flat estimated line as proof that the model is healthy. A flat estimate under concept drift is exactly the case where it is wrong. Keep collecting labels and compute realized performance on whatever you can get.
Try it. Add 'f1' and 'accuracy' to the metrics list of your estimator, rerun, and plot. Which of the three metrics reacts earliest, and which one would you want on the dashboard for a model with very rare positives?

Configuration you will touch, and the errors you will meet

Most of what you configure lives in the constructor arguments you have already seen. A short map helps you remember where to look when something is off.

If the results look wrong, the first suspects are the column names. y_pred_proba, y_pred, y_true and timestamp_column_name must match the DataFrame exactly, including case. A typo gives you an error about a missing column. Second is the reference data: is it long enough, and does it really represent the good period? Third is the chunking discussed above. Only after those three should you doubt the method.

The thresholds can also be changed. By default they come from the reference data, but you can supply your own threshold objects from nannyml.thresholds, such as a constant lower bound for a metric, which suits cases where the business has a fixed minimum, for example "ROC AUC below 0.7 is unacceptable". That belongs to the Mid-level guide, but it is useful to know that the automatic thresholds are a default and not a law.

Here are the errors you are most likely to meet, how to read them, and what to do.

  • Missing native library on import. A message such as Library not loaded: libomp.dylib on macOS or libgomp.so.1: cannot open shared object file on Linux means LightGBM's OpenMP runtime is missing. Install libomp with Homebrew or libgomp1 with apt.
  • No matching distribution on a new Python. If pip reports no matching distribution for nannyml on Python 3.13, the interpreter is outside the supported range. Create a 3.12 environment.
  • Not fitted. An error saying the calculator was not fitted means you called calculate() or estimate() before fit(). Always fit on the reference data first.
  • Invalid arguments. NannyML raises an invalid-arguments error for conflicting or missing settings, for example giving both chunk_size and chunk_number, or omitting y_pred_proba for CBPE. Read the message: it names the argument.
  • Missing columns. A column name that is not in the DataFrame produces an error that lists the missing names. Print df.columns and compare.
  • Period chunking without timestamps. Using chunk_period needs the timestamp column, as described earlier.
  • Saving a figure as an image. Calling write_image on a Plotly figure needs pip install kaleido. Use write_html if you do not want another dependency.
  • Pydantic import errors. If something like cannot import name 'BaseSettings' from 'pydantic' appears, another package in the environment pinned an incompatible Pydantic. A fresh virtual environment is the cleanest fix.
  • SSL errors from pip. CERTIFICATE_VERIFY_FAILED behind a TLS-inspecting proxy needs the proxy's CA certificate, never disabled verification.

There are also errors that are not exceptions but wrong-looking results. An alert on nearly every reference chunk means the reference is too small or the chunks too small. No alerts ever means the chunks may be so large that they smooth away all variation. An estimate far from the realized value means calibration or concept drift.

Try it. Break your own script on purpose. Misspell y_pred_proba's column name, then call estimate() before fit(), then give both chunk_size and chunk_number. Read each error and say which fix applies.

Where NannyML stands today: a note on maintenance

A responsible guide tells you what you are building on. In June 2025, the data quality company Soda announced that it had acquired NannyML. Soda stated that the open-source library would remain open and maintained, and it mentioned a future 1.0 release. The latest release on PyPI, 0.13.1, is from July 2025. At the time of writing, a community issue on the project's GitHub repository asks about the pace of activity, and no newer release has appeared.

What does that mean for you as a beginner? It is not a reason to avoid the library. The version you install works, the ideas are sound, and the code is Apache 2.0 licensed. It is a reason to be careful in three ways. First, pin the version (nannyml==0.13.1) in your requirements file so that a surprise upgrade cannot change behaviour. Second, pin its neighbours, such as pandas, NumPy and scikit-learn, with a lock file, because the library's constraints on them were last updated in 2025 and newer releases of them may not fit. Third, stay on Python 3.9 to 3.12. Check the repository's releases page before starting a long project, in case that situation has changed.

The same maintenance question is a good habit for every tool in this catalogue. Look at the last release date, the open issues, and who owns the project before you commit a team to it. Learning what NannyML does well, especially label-free performance estimation, remains valuable even if you later move to another tool, because the ideas of reference, chunks, thresholds and sampling error exist in all of them. For experiment tracking that pairs well with monitoring results, see /student-guides/mlflow.

Try it. Open the project's releases page on GitHub and find the date of the newest release. Then open the issues list and read one recent thread. Write two sentences on whether you would be comfortable running this in production for a year.

Preparing your own data for NannyML

The sample datasets are convenient, but the day will come when you point NannyML at your own model. Most of the trouble at that point has nothing to do with the library. It comes from the shape of the data, so it is worth knowing what NannyML expects before you start.

You need two tables, each as a pandas DataFrame. The reference table has one row per prediction from a period you trust. It holds the input features, the model's predicted probability, the predicted class, the true outcome and a timestamp. The analysis table has the same columns for the period you want to watch, usually without the true outcome. Both tables must use identical column names and compatible types. If a feature is an integer in the reference and a string in production, you will get either an error or a nonsense comparison, so fix the types in your own loading code and not by hand in a notebook.

Where do the predictions come from? The cleanest answer is that your serving system logs every prediction together with its inputs and a timestamp, to a table or a set of files. If your model does not log, start doing it today. A model that does not record what it saw and what it said cannot be monitored by anyone, with any tool. The predicted probability matters most for NannyML: store the probability, not only the final class. A model that outputs only a yes or no label gives CBPE nothing to work with.

Think about the reference period with the same care as the model itself. A common and sound choice is the held-out test set the model was evaluated on, because you know how the model performed there and it has labels. Another is the first weeks of production, once their labels have arrived. A poor choice is a period that includes a known incident, a holiday that will not repeat, or a data pipeline bug, because the thresholds will learn that abnormal behaviour as normal. Also make sure that the reference contains enough rows to form several chunks. A useful rule of thumb is at least ten chunks at the chunk size you plan to use, and more is better, since the thresholds come from the spread across those chunks.

Time matters too. Keep the timestamps in a real datetime type and in one time zone. A model serving customers in Riyadh, Cairo and Dubai across several time zones can end up with rows from the same day falling into different calendar chunks if timestamps are a mixture of local times and UTC. Convert to UTC when you log, and convert for display only.

Finally, think about privacy before the data leaves its home. The tables hold copies of production inputs, which may contain personal data. Many employers in the Gulf and Egypt require that such data stays in a given region or inside a given network. Run the monitoring job where the data lives, and keep the outputs, which can reveal distributions of sensitive columns, under the same access rules as the data itself.

A quick checklist before running anything: print reference_df.dtypes and analysis_df.dtypes and confirm they match, confirm the timestamp column is a datetime, confirm the probability column holds values between zero and one, and confirm the reference has labels while the analysis does not need them. Five minutes here saves the confusing errors that appear later as strange charts.

Try it. Take a small dataset you own, even a classifier from a course, and build a reference and an analysis DataFrame from it by splitting on a date. Print the dtypes of both and fix any mismatch. Then run CBPE on them with a chunk size that gives at least ten reference chunks.

Writing up a finding that someone will act on

A chart is not a result until a person understands it and does something. NannyML gives you numbers and plots, and your value as an engineer is turning them into a short message that a product owner or a risk manager can use. A good habit is to write every finding with the same five parts.

First, what happened, in one sentence and in business words: "Estimated ROC AUC fell below its lower threshold in the chunks starting in March." Avoid describing the library; describe the model.

Second, how sure you are. State the chunk size, mention whether the confidence band is wide, and say whether the effect is a single chunk or several in a row. One alerting chunk out of forty is often noise within three standard deviations, since by construction a small share of healthy chunks will cross the line. Several consecutive chunks drifting downward is a pattern.

Third, what changed in the data. Use the drift results and the distribution plots: "The share of applicants in the youngest age group doubled, and the income column shifted upward." This is the explanation part, and you give it as a candidate cause and not as a proven cause.

Fourth, how you checked. If labels have arrived for any part of the period, say what the realized metric shows. If not, say when labels are expected, and what other check you used, such as a manual review of a sample of predictions.

Fifth, what you recommend. Retrain on newer data, investigate an upstream data change, add a threshold in a downstream rule, or keep watching. Be ready to say what evidence would change your recommendation.

Writing in this shape protects you in two ways. It keeps you honest about uncertainty, which matters because the estimate rests on assumptions you cannot fully verify. And it makes your work easy to audit later: when the labels arrive and settle the question, you can check your own message against the truth and learn how reliable your monitor is on this model. Teams that do this for a few months tend to know exactly how much to trust each alert, and that knowledge is worth more than any tuning of thresholds.

Try it. Using the outputs of your own run, write a five-part finding of under 150 words for a non-technical reader. Ask a friend to read it and tell you what they would do next. If they cannot answer, rewrite it.

Putting it all together: a small end-to-end project

Let us finish with one script that does everything in this guide, saves its outputs, and is safe to rerun. You will estimate performance, compute realized performance, run drift detection on a few columns, and write the findings to files you can open later.

monitor.py
import nannyml as nml

# 1. Load the sample data: reference (with labels), analysis (without), and the late-arriving targets.
reference_df, analysis_df, analysis_targets_df = nml.load_us_census_ma_employment_data()

CHUNK = 5000
non_features = ['timestamp', 'predicted_probability', 'prediction', 'employed']
features = [c for c in reference_df.columns if c not in non_features][:4]

# 2. Estimate performance without labels.
estimator = nml.CBPE(
    problem_type='classification_binary',
    y_pred_proba='predicted_probability',
    y_pred='prediction',
    y_true='employed',
    timestamp_column_name='timestamp',
    metrics=['roc_auc', 'f1'],
    chunk_size=CHUNK,
)
estimator.fit(reference_df)
estimated = estimator.estimate(analysis_df)

# 3. Realized performance, using the labels that "arrived later".
full = analysis_df.merge(analysis_targets_df, left_index=True, right_index=True)
calculator = nml.PerformanceCalculator(
    problem_type='classification_binary',
    y_pred_proba='predicted_probability',
    y_pred='prediction',
    y_true='employed',
    metrics=['roc_auc', 'f1'],
    timestamp_column_name='timestamp',
    chunk_size=CHUNK,
)
calculator.fit(reference_df)
realized = calculator.calculate(full)

# 4. Drift on a few columns, to explain any drop.
drift_calc = nml.UnivariateDriftCalculator(
    column_names=features,
    timestamp_column_name='timestamp',
    continuous_methods=['jensen_shannon'],
    categorical_methods=['jensen_shannon'],
    chunk_size=CHUNK,
)
drift_calc.fit(reference_df)
drift = drift_calc.calculate(analysis_df)

# 5. Save tables and interactive charts.
estimated.filter(period='analysis').to_df(multilevel=False).to_csv('estimated.csv', index=False)
realized.filter(period='analysis').to_df(multilevel=False).to_csv('realized.csv', index=False)
drift.filter(period='analysis').to_df(multilevel=False).to_csv('drift.csv', index=False)

estimated.plot().write_html('estimated.html')
realized.plot().write_html('realized.html')
drift.plot(kind='drift').write_html('drift.html')
drift.plot(kind='distribution').write_html('distribution.html')

print('Done: open estimated.html, realized.html, drift.html and distribution.html')

Run it with python monitor.py and then work through the outputs in the order a real investigation would follow. Open estimated.html first: does the estimate leave its thresholds, and in which chunk? Open realized.html and ask whether the real numbers agree. If performance dropped, open drift.html and distribution.html and look for columns whose alerts start in the same chunk. Write a short note with the structure that a stakeholder would want: what dropped, since when, which inputs changed, and what you would check next.

Notice how the structure mirrors the practice rather than the library. The estimate raises the question, the realized numbers test it, the drift results suggest a cause, and a human decides what to do. NannyML does the statistics; the judgement stays with you.

To make it yours, replace the sample loader with your own pandas DataFrames. The reference frame needs the features, the predicted probabilities, the predicted classes, the true labels and a timestamp. The analysis frame needs the same columns except the labels. If the model you monitor is hosted in a cloud region such as the Gulf or Egypt for data-residency reasons, run this job inside the same environment as the data. The inputs to a monitoring job are copies of your production inputs, so they carry the same privacy obligations.

Try it. Run monitor.py, then change CHUNK to 2500 and run it again into a different folder. Compare the two sets of charts and write down which conclusions stayed the same and which changed.

What you can now do, and what comes next

You now understand the problem NannyML was built to solve: models degrade silently, and labels arrive late. You can install the library on a supported Python and diagnose the usual native-library problems. You know the four nouns, reference, analysis, chunks and thresholds, and you can explain why an alert is relative to the reference rather than an absolute verdict. You can estimate classification performance with CBPE, read the value, confidence band, thresholds and alert columns, choose a sensible chunking style and size, run univariate drift detection and read its line and distribution plots, compute realized performance, and compare it with your estimate. You also know the assumptions that can make the estimate wrong, and the maintenance situation of the library, so you can pin versions and decide with open eyes.

The Mid-level guide takes you into the machinery. It covers custom chunkers and thresholds, regression with DLE, multiclass models, multivariate drift with PCA reconstruction and domain classifiers, data quality checks, and the rankers that link drift to performance impact. The Senior guide covers running monitoring as a scheduled production job, persisting fitted calculators, storing results in a database, scaling limits and governance.

Natural neighbours in this catalogue: if you want to log the monitor's results next to your experiments, read /student-guides/mlflow. If you need to schedule the monitoring script to run every day, /student-guides/airflow and /student-guides/prefect are the usual choices. If you want to package the job so it behaves the same everywhere, /student-guides/docker is the place to start. And to validate the data before it ever reaches the model, see /student-guides/great-expectations.

Sources