تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
MLflowTracking & registry3 levels111 sectionsدليل بالإنجليزية

The Complete MLflow Guide

Taught at three levels, written three ways. Pick Beginner, Mid-level, or Senior, then read the full explanation, a fast interview review, or the practical tips and traps for that level.

15sections
32examples

This is part one of three. It covers everything you need to do real work with MLflow, not a teaser. By the end you can track an experiment without restructuring your code, log parameters, metrics, and artifacts on purpose, compare thirty runs in a table and a chart, package a model so someone else can load it without knowing your framework, register and promote it through stages, and serve it behind an HTTP endpoint. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and these ideas only stick once you have watched a metric appear in the MLflow UI while your script is still running.

What MLflow is, and the problem it solves

You trained a model last Tuesday. It scored 0.94 AUC. Today you cannot reproduce it. The notebook has moved on, the learning rate in the cell is not the one that ran, the feature list changed, and the only surviving record of 0.94 is a number in a Slack message.

Carelessness was never the problem. A system was missing: nothing was writing anything down. Every fact about that run lived in your terminal scrollback and your memory.

MLflow writes it down. You add a handful of lines, often just one, and it records the parameters, the metrics over time, the artifacts, the source file, the git commit, and the environment needed to load the model again, into a store you and your team can query.

YOUR CODE+ autolog
TRACKINGruns, params, metrics
MODELSpackaged + registered
SERVINGload anywhere

MLflow's central idea is a framework-agnostic contract. A logged model is a directory with a declared flavour, a signature describing inputs and outputs, and a dependency list rather than a bare pickle from someone's laptop. Anything that understands that contract can load it: a batch job, a REST server, a Spark UDF, a Docker image. That single property is what turns a training script into something deployable.

That gives you four components, and they map exactly onto four problems:

📋

Tracking

Runs with parameters, metrics, tags, and artifacts. The answer to "what did we try and what happened?"

📦

Models

A packaging format with flavours, a signature, and an environment. The answer to "how do I load this?"

🗂️

Model Registry

Named models with versions, aliases, and stage transitions. The answer to "which one is in production?"

🧪

Projects

A declared entry point and environment so a run is reproducible by command. The answer to "how do I rerun it?"

The vocabulary is small, and getting it straight now prevents most later confusion:

Term What it is
Run One execution: parameters, metrics, tags, artifacts, and a status. The atomic unit
Experiment A named container of runs. The unit you compare within
Tracking server The service that stores run metadata and serves the UI
Backend store Where run metadata lives: a database, or local files
Artifact store Where files live: S3, GCS, Azure, or a local directory
Artifact Any file attached to a run: a plot, a CSV, a model directory
Flavour How a logged model can be loaded: sklearn, pytorch, pyfunc, and others
Signature The declared input and output schema of a model
Registered model A named entry in the registry with numbered versions

You need little to follow along: Python, a script that trains something small, and a folder. A tiny dataset and a five-line model are the best place to start, because a broken run costs you nothing.

Try it
  1. Find your most recent training script. Write down, from memory, the exact learning rate and feature set it last ran with.
  2. Check the script and see whether you were right.
  3. Find the model file it produced and try to name the git commit that trained it.
  4. Count the files on your disk matching model_final*, *_v2*, or a date in the name.
at least one of those is unanswerable, and usually all three. The third one, tying a model file to a commit, is the question that turns into an afternoon of archaeology during a code review, and it is the one MLflow answers with a run id.

The four components, and the two stores behind them

MLflow is often described as one tool. It is four that share a store, and you can adopt them in order.

1TrackingLog params, metrics, and artifacts. Useful on day one, with no infrastructure.
2ModelsLog the trained object in a loadable format with a signature and environment.
3RegistryPromote a specific model version by alias, so consumers stop hard-coding paths.
4Serving / ProjectsLoad that version in a batch job, a REST server, or a container.

Two stores sit behind everything, and confusing them causes a whole category of errors:

your process
mlflow clientlog_param · log_metric · log_model
Autolog patchessklearn, torch, xgboost: params and metrics for free
tracking server: two separate stores
Backend storeRuns, params, metrics, tags, and the registry. Needs a database for aliases
Artifact storeFiles: models, plots, reports. S3, GCS, Azure, or a directory
Web UIReads both. Compare, filter, chart
Registry versionImmutable, numbered, linked to its run
Alias → @championMoveable pointer. Promotion is a tag move
ConsumersBatch job · REST server · container · Spark UDF

The detail that surprises people: by default the client reads and writes the artifact store directly, and the tracking server only hands out a URI. That is why metrics can appear in the UI while artifacts 404, and it means every user and CI job needs storage credentials unless you enable proxied access.

Store Holds Configured by
Backend store Runs, params, metrics, tags, experiment and registry metadata --backend-store-uri
Artifact store Files: models, plots, CSVs, anything you log as an artifact --default-artifact-root
BASH
# Local files only — no server, no database. Creates ./mlruns
python train.py

# A local server with a SQLite backend — needed for the Model Registry
mlflow server --backend-store-uri sqlite:///mlflow.db \
              --default-artifact-root ./mlartifacts \
              --host 127.0.0.1 --port 5000

# A team setup: Postgres for metadata, S3 for files
mlflow server --backend-store-uri postgresql://user:pass@host/mlflow \
              --default-artifact-root s3://my-bucket/mlflow \
              --host 0.0.0.0 --port 5000
The Model Registry requires a database backend With the default file store, tracking works and the registry does not. You get a clear error the first time you call register_model. SQLite is enough for one person; Postgres or MySQL is what a team needs. Decide this early, because migrating a file store into a database later is a chore nobody enjoys.

Your code points at a tracking server with one line, or one environment variable:

PYTHON
import mlflow
mlflow.set_tracking_uri("http://127.0.0.1:5000")   # or "sqlite:///mlflow.db", or a path
mlflow.set_experiment("churn-baseline")
BASH
export MLFLOW_TRACKING_URI=http://127.0.0.1:5000
Try it
  1. Run pip install mlflow, then mlflow server --backend-store-uri sqlite:///mlflow.db --default-artifact-root ./mlartifacts --port 5000.
  2. Open http://127.0.0.1:5000 and note that it is empty: no runs, no experiments except Default.
  3. Sketch the two stores from memory and label which one a metric goes to and which one a plot goes to.
  4. Now run a script with no set_tracking_uri at all and find the mlruns/ directory it created.
step four is the one worth doing: MLflow works with zero configuration by writing to a local folder, which is why people accidentally end up with runs scattered across five mlruns directories. Knowing that the default exists is how you avoid it.

Your first tracked run, line by line

Start from a script with no MLflow in it:

train.py (before)
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score

n_estimators, max_depth = 200, 5

X, y = load_breast_cancer(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

model = RandomForestClassifier(n_estimators=n_estimators, max_depth=max_depth, random_state=42)
model.fit(X_train, y_train)
print("auc", roc_auc_score(y_test, model.predict_proba(X_test)[:, 1]))

Here it is tracked:

train.py (after)
import mlflow
import mlflow.sklearn
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score

mlflow.set_tracking_uri("http://127.0.0.1:5000")
mlflow.set_experiment("churn-baseline")

params = {"n_estimators": 200, "max_depth": 5, "random_state": 42}

with mlflow.start_run(run_name="rf baseline") as run:
    mlflow.log_params(params)

    X, y = load_breast_cancer(return_X_y=True, as_frame=True)
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

    model = RandomForestClassifier(**params).fit(X_train, y_train)
    auc = roc_auc_score(y_test, model.predict_proba(X_test)[:, 1])

    mlflow.log_metric("auc", auc)
    mlflow.sklearn.log_model(model, name="model", input_example=X_train.head(3))

    print("run id:", run.info.run_id, "auc:", auc)

Run it, then open the UI. Four things happened:

  1. A run was createdWith a run id, an experiment, a status, a start and end time, and the user and source file that produced it.
  2. Parameters were recordedImmutable strings. A parameter is set once and describes the configuration of the run.
  3. A metric was recordedA float with a step and a timestamp, so a metric is a series rather than a single number.
  4. A model was loggedNot a bare pickle, but a directory with an MLmodel file declaring its flavours, its signature, and its dependencies.

The with block matters more than it looks. It sets the run's status to FINISHED on success and FAILED on an exception, and it guarantees the run is closed. Without it, an interrupted script leaves a run stuck in RUNNING forever.

Use the context manager, always with mlflow.start_run(): is the correct default. A bare mlflow.start_run() followed by mlflow.end_run() works until something raises in between, and then you have a run that claims to be running days later. In a notebook, an un-ended run also silently captures your next cell's logs.
Try it
  1. Run the tracked script and open the run in the UI. Visit each tab: Overview, Parameters, Metrics, Artifacts.
  2. Under Artifacts, open model/MLmodel and read it. Note the flavours and the signature.
  3. Raise an exception halfway through the with block and confirm the run's status becomes FAILED.
  4. Now do the same without the context manager and see the run stuck in RUNNING.
the MLmodel file is the thing to read. It is the entire "why MLflow" argument in twenty lines: a declared loader, a declared schema, and a declared environment, which is what makes the artifact portable rather than personal.

Params, metrics, tags, artifacts: choosing correctly

Four things you log, with four different meanings. Choosing correctly is most of what separates a readable experiment table from an unusable one.

PYTHON
# Parameters: configuration. Immutable, stored as strings.
mlflow.log_param("learning_rate", 3e-4)
mlflow.log_params({"batch_size": 64, "optimizer": "adamw", "features": "v3"})

# Metrics: numbers that can move. Float, with an optional step.
mlflow.log_metric("train_loss", 0.412, step=epoch)
mlflow.log_metric("val_loss", 0.503, step=epoch)
mlflow.log_metrics({"precision": 0.91, "recall": 0.88}, step=epoch)

# Tags: mutable, searchable labels for organisation.
mlflow.set_tag("stage", "baseline")
mlflow.set_tags({"team": "risk", "dataset": "v3", "owner": "amina"})

# Artifacts: files.
mlflow.log_artifact("outputs/confusion_matrix.png")
mlflow.log_artifact("outputs/report.html", artifact_path="reports")
mlflow.log_artifacts("outputs/plots", artifact_path="plots")   # a whole directory

# Convenience helpers that avoid touching disk yourself
mlflow.log_figure(fig, "plots/roc.png")
mlflow.log_table(predictions_df, "predictions.json")
mlflow.log_text(json.dumps(config, indent=2), "config.json")
mlflow.log_dict(config, "config.yaml")
Kind Mutable Type Use for
Param No: logging twice with a different value errors String Configuration: hyperparameters, feature-set version, seed
Metric Yes: appends a new step Float Anything measured: loss, AUC, latency, training minutes
Tag Yes: overwrites String Organisation and search: team, purpose, ticket, dataset
Artifact Append-only File Plots, reports, predictions, the model itself

Two rules resolve nearly every "which one?" question. If it is a number you might plot, it is a metric. If it is something you might change your mind about after the run, it is a tag.

Params are immutable, and that bites in loops Calling log_param("lr", x) twice with different values raises. It is deliberate, a run has one configuration, but it surprises people who log parameters inside a training loop or a retry. Log parameters once, at the top of the run, and use metrics for anything that varies during it.
Try it
  1. Log the same parameter twice with different values and read the error.
  2. Log a metric with step=0..9 in a loop and look at the Metrics tab, and note that you get a chart, not a number.
  3. Set a tag, then set the same tag to a different value, and confirm it overwrites.
  4. Log a matplotlib figure with log_figure and confirm it appears under Artifacts without you writing a file.
the immutability difference in steps one and three is the whole model: params describe the run, tags describe how you think about the run, and only one of those is allowed to change afterwards.

Autologging: one line, and its boundaries

Most of the code in the previous section is unnecessary for common frameworks. One line replaces it.

PYTHON
import mlflow

mlflow.autolog()            # everything MLflow can detect

# or per-framework, which is more predictable
mlflow.sklearn.autolog()
mlflow.pytorch.autolog()
mlflow.tensorflow.autolog()
mlflow.xgboost.autolog()
mlflow.lightgbm.autolog()
mlflow.transformers.autolog()

With sklearn.autolog() active, calling .fit() logs the estimator's parameters, training metrics, the fitted model with a signature, and, for search objects like GridSearchCV, child runs for every candidate configuration:

PYTHON
import mlflow
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.model_selection import GridSearchCV

mlflow.set_experiment("gb-sweep")
mlflow.sklearn.autolog(max_tuning_runs=10)

grid = GridSearchCV(
    GradientBoostingClassifier(random_state=42),
    {"max_depth": [2, 3, 4], "learning_rate": [0.05, 0.1]},
    scoring="roc_auc", cv=3,
)

with mlflow.start_run(run_name="gb grid search"):
    grid.fit(X_train, y_train)          # one parent run, plus child runs per candidate
Autologged Typically includes
Parameters Every constructor argument of the estimator or the training call
Metrics Training scores, and per-epoch metrics for deep learning frameworks
Model Logged with a signature inferred from the training data
Artifacts Feature importances, metric plots, and framework-specific summaries
Child runs One per candidate in a hyperparameter search, up to max_tuning_runs
Datasets For some flavours, a dataset reference describing the training input

What autologging does not do: log your own custom evaluation metrics, log artifacts you generated yourself, or know which of your variables were meaningful. So the realistic pattern is autolog plus a few deliberate calls:

PYTHON
mlflow.sklearn.autolog(log_models=True, log_datasets=True, silent=True)

with mlflow.start_run(run_name="rf + business metric"):
    model.fit(X_train, y_train)                     # autologged
    mlflow.log_metric("expected_loss_usd", cost)     # yours
    mlflow.set_tag("dataset", "v3")                  # yours
    mlflow.log_artifact("outputs/segment_report.html")
Call autolog before you create the estimator Autologging patches the framework, so it has to be active before the training call, and for some frameworks before the object is constructed. Putting mlflow.autolog() at the top of the entry point, next to the imports, removes an entire class of "why did nothing get logged".
Try it
  1. Delete every logging call from your script, add mlflow.sklearn.autolog(), and compare what gets recorded.
  2. Run a GridSearchCV under autolog and look at the parent run's child runs in the UI.
  3. Add one custom metric of your own alongside autolog and confirm both appear.
  4. Move the autolog() call to after the estimator is created and see what changes.
the child-run structure in step two is the payoff: a six-candidate search becomes six comparable runs in one table, for free. Step four teaches you the ordering rule the hard way, which is the way that sticks.

The MLflow Model format: flavours and signatures

This is the component that makes MLflow more than a logbook, and everything downstream depends on understanding it properly.

log_model does not save a pickle. It creates a directory:

mlruns///artifacts/model/
model/
├── MLmodel                 # the manifest — flavours, signature, dependencies
├── model.pkl               # the serialised estimator
├── conda.yaml              # a conda environment that can load it
├── python_env.yaml         # a pip/venv equivalent
└── requirements.txt        # the pinned pip requirements
MLmodel
artifact_path: model
flavors:
  python_function:
    loader_module: mlflow.sklearn
    model_path: model.pkl
    python_version: 3.11.9
  sklearn:
    pickled_model: model.pkl
    sklearn_version: 1.5.1
    serialization_format: cloudpickle
mlflow_version: 2.16.0
model_uuid: 8f2c4b19e0a7d3f1c6b8a2e4d7091f3b
run_id: a1b2c3d4e5f64718b9c0d1e2f3a4b5c6
signature:
  inputs: '[{"name": "mean radius", "type": "double"}, {"name": "mean texture", "type": "double"}]'
  outputs: '[{"type": "long"}]'

Two ideas in that file carry the whole design:

Flavours are loaders. A model can declare several. The sklearn flavour lets you load it back as a scikit-learn object with all its methods. The python_function (pyfunc) flavour lets anything load it as a generic object with a predict method and no knowledge of scikit-learn at all. That second one is why the same artifact works in a REST server, a Spark job, and a Docker image.

PYTHON
import mlflow

# Native flavour: the real object, with its own API
sk_model = mlflow.sklearn.load_model("runs:/a1b2c3d4/model")
sk_model.feature_importances_

# Generic flavour: just predict(). Framework-agnostic.
pyfunc = mlflow.pyfunc.load_model("runs:/a1b2c3d4/model")
pyfunc.predict(X_test)

A signature is a contract. It declares the input columns and types and the output type, so a wrong-shaped or wrong-typed request is rejected at the boundary rather than producing nonsense predictions.

PYTHON
from mlflow.models import infer_signature

signature = infer_signature(X_train, model.predict(X_train))
mlflow.sklearn.log_model(
    model,
    name="model",
    signature=signature,
    input_example=X_train.head(3),      # also enables the UI's sample request
)
Piece Why it matters
flavors How the model can be loaded, and by whom
signature Input/output schema: validated at serving time
input_example A concrete sample, used for docs and for signature inference
requirements.txt The pinned environment, so serving does not guess
run_id The link back to the run, and therefore to the params and metrics

The URI schemes you will use constantly:

URI Points at
runs:/<run_id>/model A model logged by a specific run
models:/<name>/<version> A specific registry version
models:/<name>@champion Whatever version currently holds that alias
s3://bucket/path/model A model directory in object storage
file:///abs/path/model A local directory
Log a signature, or serving will accept anything Without a signature, an endpoint takes columns in the wrong order or of the wrong type and returns predictions that are wrong. infer_signature costs one line and converts a class of silent production bugs into a loud 400 at the boundary. Always pass it, or pass input_example so MLflow can infer it.
Try it
  1. Log a model, then open its MLmodel file and identify every flavour it declares.
  2. Load it twice (once with mlflow.sklearn.load_model, once with mlflow.pyfunc.load_model) and note what each object can do.
  3. Log one model with a signature and one without, then feed both a DataFrame with the columns in the wrong order.
  4. Open the requirements.txt inside the model directory and compare it against your environment.
step three is the demonstration: the unsigned model returns confident nonsense, the signed one refuses. That difference is the argument for signatures, and no amount of reading it is as convincing as seeing both outputs.

Comparing runs: the table, the chart, the query

Tracking pays off here, and it is largely a UI skill. Run your script three or four times with different parameters before reading on.

In the experiment table:

Action How
Show a param or metric as a column The columns selector; metric columns show the latest value
Sort by a metric Click the column header
Filter The search box, using the run-filter syntax
Compare Tick two or more runs → Compare
Chart across runs The Chart view: parallel coordinates, scatter, or bar

Learn the search syntax, because the Python API uses the same language:

TEXT
metrics.auc > 0.9
params.max_depth = '5'
tags.dataset = 'v3' and metrics.auc > 0.88
attributes.status = 'FINISHED'
metrics.rmse < 1.5 and params.model_type != 'linear'

In code, which is how you build a report or a promotion gate:

PYTHON
import mlflow

runs = mlflow.search_runs(
    experiment_names=["churn-baseline"],
    filter_string="metrics.auc > 0.9 and tags.dataset = 'v3'",
    order_by=["metrics.auc DESC"],
    max_results=20,
)
print(runs[["run_id", "params.max_depth", "metrics.auc", "tags.mlflow.runName"]])

best = runs.iloc[0]
print("best run:", best.run_id, best["metrics.auc"])

search_runs returns a pandas DataFrame with flattened params.*, metrics.*, and tags.* columns, which makes "find the best run and do something with it" a three-line script rather than a project.

What comparison gives you

  • A diff of parameters across selected runs
  • Overlaid metric curves by step
  • Parallel coordinates to see which parameter drives the metric
  • A queryable table you can drive from Python

What it cannot do for you

  • Compare numbers you only print()ed
  • Compare parameters you never logged
  • Distinguish runs you never named or tagged
  • Explain a difference caused by unlogged data changes
Name and tag every run, in code run_name="rf depth=5 feats=v3" plus a couple of tags makes a table of two hundred runs navigable. Set them in the script rather than by hand, because the run you forget to label is always the one you need to find later.
Try it
  1. Run one script four times with different depths, each with a descriptive run_name.
  2. Add the depth and your metric as columns and sort by the metric.
  3. Select all four and open the Compare view, then the Chart view's parallel coordinates.
  4. Reproduce the ranking in Python with mlflow.search_runs and a filter_string.
parallel coordinates in step three is the underused one: with four runs it looks decorative, and with forty it tells you immediately which parameter matters. Step four matters because that same query is what a CI promotion check runs.

The Model Registry: versions and aliases

Tracking answers "what did we try". The registry answers "which one is live", and it is the boundary between experimentation and production.

Register a model in one of three ways:

PYTHON
# 1. At log time — the most common
mlflow.sklearn.log_model(model, name="model", registered_model_name="churn-classifier")

# 2. From an existing run afterwards
mlflow.register_model("runs:/a1b2c3d4e5f64718/model", "churn-classifier")

# 3. Through the client, with explicit control
from mlflow import MlflowClient
client = MlflowClient()
client.create_registered_model("churn-classifier")
client.create_model_version(
    name="churn-classifier",
    source="runs:/a1b2c3d4e5f64718/model",
    run_id="a1b2c3d4e5f64718",
)

Each registration creates a numbered version (version 1, 2, 3) which is immutable in content and permanently linked back to the run that produced it.

Then you point at a version by alias rather than by number, which is the part that makes deployment sane:

PYTHON
from mlflow import MlflowClient
client = MlflowClient()

client.set_registered_model_alias("churn-classifier", "champion", version=7)
client.set_registered_model_alias("churn-classifier", "challenger", version=8)

# Consumers never mention a version number
model = mlflow.pyfunc.load_model("models:/churn-classifier@champion")
1LoggedA run produces a model artifact with a signature and an environment.
2RegisteredIt becomes version N of a named registered model, linked to its run.
3EvaluatedA separate job scores it and records the result, often as a tag on the version.
4Aliasedchampion moves to that version. Every consumer follows the alias.
5SupersededA newer version takes the alias; the old one stays queryable by number forever.
Concept Meaning
Registered model The name: one per problem, e.g. churn-classifier
Version An immutable numbered entry, linked to the producing run
Alias A moveable pointer: champion, challenger, shadow
Tags Metadata on the model or a version, searchable
Description Free text: what it is, who owns it, known limitations
Stages are deprecated: use aliases Older MLflow used fixed stages: None, Staging, Production, Archived, moved with transition_model_version_stage. You will see them in older code and tutorials, and they still appear in some UIs. Aliases replaced them because they are arbitrary, multiple, and not tied to a fixed vocabulary. Write new code against aliases; recognise stages when you read old code.
Try it
  1. Register two models from two different runs and confirm you get versions 1 and 2.
  2. Set the champion alias to version 1, then load models:/name@champion and print a prediction.
  3. Move the alias to version 2 and rerun the exact same loading code. Nothing in your code changed.
  4. From a registered version in the UI, click through to the run that produced it and find its parameters.
step three is the entire point of a registry: a deployment changed without a code change, a redeploy, or a file path. Step four is the audit path, weights to run to parameters to commit, and it should take you seconds.

Loading and serving: batch, REST, container

A registered model is only useful once something consumes it. Three ways, in increasing order of infrastructure.

In a batch job, which is the most common and the easiest:

batch_score.py
import mlflow
import pandas as pd

model = mlflow.pyfunc.load_model("models:/churn-classifier@champion")
frame = pd.read_parquet("data/today.parquet")
frame["score"] = model.predict(frame[model.metadata.get_input_schema().input_names()])
frame.to_parquet("out/scored.parquet")

Behind a local REST endpoint, with one command:

BASH
mlflow models serve -m "models:/churn-classifier@champion" --host 127.0.0.1 --port 5001 --env-manager local
BASH
curl -X POST http://127.0.0.1:5001/invocations \
  -H 'Content-Type: application/json' \
  -d '{"dataframe_split": {"columns": ["mean radius", "mean texture"], "data": [[17.99, 10.38]]}}'
Payload format Shape
dataframe_split {"columns": [...], "data": [[...]]}: the usual choice for tabular
dataframe_records [{"col": value, ...}]: row dicts
instances [[...], [...]]: tensor-style input
inputs A named-tensor dict, for deep learning models

As a container, which is what you hand to a platform team:

BASH
mlflow models build-docker -m "models:/churn-classifier@champion" -n churn-scorer
docker run -p 5001:8080 churn-scorer

The --env-manager flag decides how the serving environment is built, and it is the setting people trip over:

Value Behaviour When
virtualenv Builds a fresh env from python_env.yaml. The default Correct, slower start
conda Builds from conda.yaml If your stack is conda-based
local Uses the current environment as-is Fast iteration, and only safe when versions match
A local MLflow endpoint has no authentication mlflow models serve starts an unauthenticated HTTP server that will run predictions for anyone who can reach it. Bind it to 127.0.0.1 for local work, and never expose it directly. Put it behind an ingress or gateway that terminates TLS and enforces auth. Mid and Senior levels cover the production shape properly.
Mismatched dependencies are the top serving failure A model pickled with scikit-learn 1.5 and loaded under 1.2 may fail loudly or, worse, load and behave differently. That is what the logged requirements.txt is for, so let the default virtualenv manager use it in anything that matters, and reserve --env-manager local for quick local checks.
Try it
  1. Serve your registered champion locally and call it with curl using dataframe_split.
  2. Send a request with a column missing and read the error the signature produces.
  3. Serve with --env-manager virtualenv and watch it build the environment from the logged requirements.
  4. Write the batch scoring script and run it against the same alias.
step two is why signatures exist: a clear 400 instead of a wrong prediction that arrives with no warning. Step three shows you where model portability comes from: the environment travels with the artifact, not with your laptop.

Evaluating with mlflow.evaluate

Logging your own metrics is fine. For standard tasks there is a built-in evaluator that produces a consistent, comparable set, plus plots and an explainability summary, from one call.

PYTHON
import mlflow
from mlflow.models import infer_signature

with mlflow.start_run(run_name="rf + eval"):
    model.fit(X_train, y_train)
    signature = infer_signature(X_train, model.predict(X_train))
    info = mlflow.sklearn.log_model(model, name="model", signature=signature)

    eval_data = X_test.copy()
    eval_data["label"] = y_test

    result = mlflow.evaluate(
        model=info.model_uri,
        data=eval_data,
        targets="label",
        model_type="classifier",
        evaluators=["default"],
    )
    print(result.metrics["roc_auc"], result.metrics["f1_score"])
model_type Produces
classifier Accuracy, precision, recall, F1, ROC AUC, log loss, plus ROC/PR/confusion plots
regressor MAE, MSE, RMSE, R², plus residual and prediction-error plots
question-answering / text Text metrics such as exact match, toxicity, and token counts

You can also add your own metric alongside the defaults, which is how a business number gets into the same comparable table:

PYTHON
from mlflow.models import make_metric

def _expected_cost(eval_df, _builtin_metrics):
    fp = ((eval_df["prediction"] == 1) & (eval_df["target"] == 0)).sum()
    fn = ((eval_df["prediction"] == 0) & (eval_df["target"] == 1)).sum()
    return 50 * fp + 500 * fn

expected_cost = make_metric(eval_fn=_expected_cost, greater_is_better=False, name="expected_cost")

result = mlflow.evaluate(
    model=info.model_uri, data=eval_data, targets="label",
    model_type="classifier", extra_metrics=[expected_cost],
)
One evaluator means every run is comparable Hand-rolled metrics drift: one run computes AUC on probabilities, another on labels, a third excludes a class. mlflow.evaluate computes the same set the same way for every run, which is what makes a leaderboard trustworthy rather than tidy.
Try it
  1. Run mlflow.evaluate on a classifier and look at the metrics and plots it added to the run.
  2. Compare two runs evaluated this way and confirm every metric lines up.
  3. Add a custom cost metric with make_metric and sort your experiment table by it.
  4. Run it on a regressor and note that the metric and plot set changes automatically.
the custom cost metric in step three is the one that changes conversations: sorting by expected loss in dollars rather than by AUC often reorders your leaderboard, and that reordering is the actual result.

Reproducibility: what a run records, and what it does not

A tracked run is not automatically a reproducible one. MLflow captures some of what you need and leaves the rest to you.

Recorded automatically Tag or field
Source file or notebook mlflow.source.name
Git commit, if run inside a repository mlflow.source.git.commit
Whether the tree was dirty mlflow.source.git.repoURL, plus a dirty warning
User mlflow.user
Run type mlflow.source.type: LOCAL, PROJECT, JOB, NOTEBOOK
Model dependencies requirements.txt inside the model directory
Not recorded unless you do it How
The data you trained on mlflow.log_input with a dataset, or a tag naming a version
Random seeds Log them as parameters and set them
Uncommitted code MLflow records the commit, not a diff: commit before real runs
System packages, CUDA Use a container; the pip list is not the whole environment
Feature definitions Log the feature-set version as a param or tag
logging the dataset, so the run names its input
import mlflow.data
import pandas as pd

frame = pd.read_parquet("data/train_v3.parquet")
dataset = mlflow.data.from_pandas(frame, source="s3://bucket/data/train_v3.parquet", name="churn-train", targets="label")

with mlflow.start_run():
    mlflow.log_input(dataset, context="training")
    mlflow.log_params({"seed": 42, "feature_set": "v3"})
    ...
MLflow records your commit, not your uncommitted changes Unlike some tools, MLflow does not store a diff of your working tree. A run tagged with commit abc123 that was executed with unsaved edits is not reproducible from that commit, and nothing in the UI shouts about it. Commit before any run whose result you might defend later, and treat a dirty-tree run as a scratch experiment.
Try it
  1. Run a script inside a git repository and find mlflow.source.git.commit in the run's tags.
  2. Make an uncommitted edit, rerun, and confirm the commit tag is unchanged, because the edit is invisible.
  3. Add mlflow.log_input with a dataset and find it in the run's Datasets section.
  4. Log your seed as a parameter, then try to reproduce a run exactly from what MLflow recorded. Note what was missing.
step two is the important discovery, and it is the single biggest difference between MLflow and tools that store a diff. Step four usually reveals two gaps: the data version and the environment beyond pip.

MLflow Projects: a run as one command

A Project is a directory with a declared entry point and environment, so a run becomes a command anyone can execute, including MLflow itself, straight from a git URL.

MLproject
name: churn

python_env: python_env.yaml

entry_points:
  main:
    parameters:
      max_depth: {type: int, default: 5}
      n_estimators: {type: int, default: 200}
      data_version: {type: string, default: "v3"}
    command: >
      python train.py --max-depth {max_depth}
                      --n-estimators {n_estimators}
                      --data-version {data_version}

  evaluate:
    parameters:
      run_id: {type: string}
    command: "python evaluate.py --run-id {run_id}"
python_env.yaml
python: "3.11"
build_dependencies:
  - pip==24.2
dependencies:
  - -r requirements.txt
BASH
# Run locally, with overrides
mlflow run . -P max_depth=8 -P n_estimators=400

# Run a specific entry point
mlflow run . -e evaluate -P run_id=a1b2c3d4

# Run straight from git, at a pinned commit
mlflow run https://github.com/org/repo.git --version 3f9a1c2 -P max_depth=8

# Skip environment creation when you know the current env is right
mlflow run . --env-manager local
Field Purpose
name Shows up on the run
python_env / conda_env / docker_env How the environment is built
entry_points Named commands; main is the default
parameters Typed, with defaults: overridable with -P
command The shell command, with {param} substitution

A run launched this way gets mlflow.source.type = PROJECT and records the entry point and its parameters, so the run itself tells you the command that produced it.

A Project is the cheapest reproducibility win available Three small files turn "clone the repo, install something, hope, then run a script with the right flags" into one command with typed parameters and a pinned commit. Even if you never use mlflow run yourself, having the entry point declared is documentation that cannot go stale.
Try it
  1. Add an MLproject and python_env.yaml to your project and run it with mlflow run. -P max_depth=8.
  2. Watch it build an isolated environment, then check the run's source type in the UI.
  3. Push the repo and run it from the git URL at a pinned commit.
  4. Add a second entry point and invoke it with -e.
step three is the one that feels like a different tool: a training run reproduced on a machine that never had your code, from a URL and a commit. That is the property a Project exists to give you.

Reading a failed run, in order

Debugging MLflow has an order, and following it beats guessing.

  1. Check the run's status and whether it exists at allA run stuck in RUNNING means the process died without closing it. No run at all means the tracking URI was wrong and you wrote to a local mlruns/ somewhere.
  2. Confirm the tracking URI and experimentmlflow.get_tracking_uri() and mlflow.get_experiment_by_name(...). Runs landing in the wrong place is the most common "MLflow is broken".
  3. Read the run's tagsSource file, git commit, and user. A surprising commit explains a lot.
  4. Read the params, not your codeEspecially with autolog, the recorded parameters are what ran.
  5. Open the model's MLmodel and requirements.txtServing and loading failures are nearly always environment mismatches, and this is where the truth is.
  6. Reload the model in a clean environmentmlflow models predict -m ... --env-manager virtualenv reproduces the serving path without standing up a server.

Six failures cover most of what you will hit at this level:

Symptom Cause Fix
No runs in the UI Wrote to a local mlruns/, not the server Set MLFLOW_TRACKING_URI or set_tracking_uri
register_model fails File-based backend store Use SQLite/Postgres for the backend
Artifacts missing but metrics present Artifact root not reachable from the client Check --default-artifact-root and credentials
Run stuck in RUNNING No context manager, process died Use with mlflow.start_run()
Param already logged error log_param called twice with different values Log params once, use metrics for varying values
Serving fails on load Dependency version mismatch Let virtualenv build from the logged requirements
useful in a debugging session
import mlflow
from mlflow import MlflowClient

print(mlflow.get_tracking_uri())
client = MlflowClient()

run = client.get_run("a1b2c3d4e5f64718b9c0d1e2f3a4b5c6")
print(run.info.status, run.info.artifact_uri)
print(run.data.params)
print(run.data.metrics)
print({k: v for k, v in run.data.tags.items() if k.startswith("mlflow.")})
print([f.path for f in client.list_artifacts(run.info.run_id)])
client.set_terminated(run.info.run_id, status="FAILED")   # close a stuck run
Try it: cause each failure on purpose
  1. Unset your tracking URI, run a script, and find the local mlruns/ it created instead.
  2. Point at a file-based store and call register_model. Read the error.
  3. Kill a script mid-run and find the stuck RUNNING run, then close it with set_terminated.
  4. For each one, diagnose it from the client API rather than from memory of what you broke.
the first one is by far the most common real-world confusion: two people "using MLflow" while writing to two different local folders. Recognising it from a missing experiment saves hours.

Putting it all together

Everything above in one project. Nothing here is new. Read it as a whole and you should be able to justify every line.

project layout
.
├── MLproject                  # declared entry points and parameters
├── python_env.yaml            # the environment MLflow builds
├── requirements.txt           # pinned
├── data/
│   └── train_v3.parquet
├── src/
│   ├── train.py               # the tracked run
│   ├── evaluate.py            # mlflow.evaluate against a registered version
│   ├── promote.py             # moves the champion alias
│   └── batch_score.py         # loads models:/name@champion
└── README.md                  # the three commands to reproduce
src/train.py
import argparse
import mlflow
import mlflow.sklearn
import pandas as pd
from mlflow.models import infer_signature
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split

ap = argparse.ArgumentParser()
ap.add_argument("--max-depth", type=int, default=5)
ap.add_argument("--n-estimators", type=int, default=200)
ap.add_argument("--data-version", default="v3")
ap.add_argument("--seed", type=int, default=42)
args = ap.parse_args()

mlflow.set_experiment("churn")                       # 1. named experiment
mlflow.sklearn.autolog(log_models=False, silent=True) # 2. autolog params/metrics only

frame = pd.read_parquet(f"data/train_{args.data_version}.parquet")
dataset = mlflow.data.from_pandas(                   # 3. the run names its data
    frame, source=f"data/train_{args.data_version}.parquet",
    name=f"churn-train-{args.data_version}", targets="label",
)

X = frame.drop(columns=["label"])
y = frame["label"]
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=args.seed, stratify=y,
)

run_name = f"rf d={args.max_depth} n={args.n_estimators} data={args.data_version}"
with mlflow.start_run(run_name=run_name) as run:      # 4. context manager, named
    mlflow.log_input(dataset, context="training")
    mlflow.log_params({"seed": args.seed, "data_version": args.data_version})
    mlflow.set_tags({"team": "risk", "purpose": "baseline"})   # 5. tags in code

    model = RandomForestClassifier(
        max_depth=args.max_depth, n_estimators=args.n_estimators,
        random_state=args.seed,
    ).fit(X_train, y_train)

    signature = infer_signature(X_train, model.predict(X_train))   # 6. a contract
    info = mlflow.sklearn.log_model(                               # 7. registered
        model, name="model", signature=signature,
        input_example=X_train.head(3),
        registered_model_name="churn-classifier",
    )

    eval_data = X_test.copy()
    eval_data["label"] = y_test
    result = mlflow.evaluate(                                      # 8. one evaluator
        model=info.model_uri, data=eval_data, targets="label",
        model_type="classifier",
    )
    print(run.info.run_id, result.metrics["roc_auc"])
src/promote.py
import mlflow
from mlflow import MlflowClient

NAME, MARGIN = "churn-classifier", 0.002
client = MlflowClient()

def auc_of(version):
    return client.get_run(version.run_id).data.metrics.get("roc_auc", 0.0)

versions = client.search_model_versions(f"name = '{NAME}'")
best = max(versions, key=auc_of)

try:
    current = client.get_model_version_by_alias(NAME, "champion")
except mlflow.exceptions.MlflowException:
    current = None

if current is None or auc_of(best) > auc_of(current) + MARGIN:
    client.set_registered_model_alias(NAME, "champion", best.version)  # 9. alias move
    client.set_model_version_tag(NAME, best.version, "promoted_auc", f"{auc_of(best):.4f}")
    print("promoted version", best.version)
else:
    print("no promotion: improvement below margin")
BASH
# One-time setup
pip install mlflow
mlflow server --backend-store-uri sqlite:///mlflow.db \
              --default-artifact-root ./mlartifacts --port 5000
export MLFLOW_TRACKING_URI=http://127.0.0.1:5000

# Everyday loop
mlflow run . -P max_depth=8 -P n_estimators=400     # 10. reproducible by command
python src/promote.py                                # threshold, then move the alias
python src/batch_score.py                            # consumes models:/…@champion

Ten decisions in there are the whole lesson of this page:

Decision Section
A named experiment, not Default The four components
Autolog for framework params, explicit calls for yours Autologging
The dataset logged with log_input Reproducibility
with mlflow.start_run() and a descriptive run_name Your first tracked run
Tags set in code, never by hand Comparing runs
A signature inferred and logged The MLflow Model format
Registered at log time with registered_model_name The Model Registry
mlflow.evaluate so every run is comparable Evaluating models
Promotion by moving an alias, with a margin The Model Registry
An MLproject entry point, so a run is one command MLflow Projects
Try it: the one that matters
  1. Take this structure into a project you work on, adapting the training body to your own code.
  2. Get one run green with autolog, a custom metric, a signature, a registered model, and an evaluation.
  3. Run it twice with different parameters and pick a winner from the experiment table.
  4. Promote the winner by moving the champion alias, then serve models:/name@champion locally and call it with curl.
  5. Move the alias to the other version and call the endpoint again without changing a line of code.
step five is the acceptance test: a deployment that changed by moving a pointer. If anything in your loop required editing a path or a version number, that is the spot to fix. It is exactly where production drift comes from.

What you can now do, and what comes next

You can track runs with parameters, metrics, tags, and artifacts; autolog whole frameworks in one line; log models in a portable format with a signature and an environment; compare dozens of runs in the UI and in Python; register versions and move aliases so deployment is a pointer change; evaluate models consistently with built-in and custom metrics; serve a model locally, in a batch job, or as a container; declare a Project so a run is one reproducible command; and debug a failed run in a fixed order. That is a working practitioner's toolkit, enough to own the experiment-tracking and model-delivery story on a real project.

Can you…
Name MLflow's four components? Tracking, Models, Registry, Projects
Give the two stores and what each holds? Backend for metadata, artifact store for files
Say why the registry needs a database? The file store cannot back it
Explain param vs metric vs tag? Immutable config, movable number, mutable label
Say what log_model writes? A directory with MLmodel, flavours, signature, requirements
Explain a flavour, and what pyfunc buys? A loader; pyfunc is the framework-agnostic one
Say why a signature matters? Wrong-shaped input fails loudly instead of silently
Name the model URI schemes? runs:/, models:/name/1, models:/name@alias
Say why aliases replaced stages? Arbitrary, multiple, not a fixed vocabulary
Explain what MLflow does not record? Uncommitted diffs, data unless you log it, system packages
Say what mlflow.evaluate gives you? One consistent metric and plot set per task type
Name the first check when runs are missing? The tracking URI. You probably wrote to local mlruns/

Mid-level takes every one of those topics further: nested runs and sweep structure, pyfunc custom models and wrappers, model dependencies and environment management in depth, dataset tracking and lineage, the registry as a promotion workflow with webhooks and CI, mlflow.evaluate with custom metrics and validation thresholds, MLflow Tracing and LLM evaluation, deployment targets and the deployments API, autologging internals, the client API and bulk operations, and the CI patterns that make a training run a pull-request check.

Senior then covers what you own when MLflow is your team's platform: the self-hosted deployment and its stores, authentication and multi-tenancy, artifact-store access control and credential handling, backend database scaling and what breaks first, backup and upgrade procedure, cost and retention policy for artifacts, lineage and audit for regulated work, promotion gates and approval trails, incident playbooks, and where MLflow stops and a feature store, an orchestrator, or a dedicated serving platform begins.