This is part one of three. It covers everything you need to do real work with Great Expectations, not a teaser. By the end you can install the library, connect it to a dataframe, a folder of files or a database, write tests for your data, run them as a single repeatable step, read the results, publish a readable report, and make a pipeline stop when the data is bad. Mid-level and Senior take the same topics further; nothing here is thrown away.
Each section ends with a Try it task. Do them as you go. They take a few minutes each, and these ideas only stick once you have watched your own data fail a test and worked out why.
One warning before we start. Most blog posts, videos and chat answers about this tool describe a version that no longer exists. Great Expectations changed its entire API in 2024, and its hosted product was shut down in 2026. This guide teaches the 1.x library, called GX Core, and is checked against release 1.23.2. If a tutorial tells you to type great_expectations init or to call context.sources, it is describing the old tool, and that advice will not work here.
What Great Expectations is, and the problem it solves
Great Expectations, usually shortened to GX, is a Python library for testing data. You write down what you believe is true about a table or a file ("every order has an ID, IDs are never repeated, prices are positive, the country column only contains countries we ship to"), and GX checks the data against those beliefs and tells you precisely which ones failed and how badly.
Each of those beliefs is called an Expectation. They are assertions about data, in the same spirit as the assertions you write in a unit test for code, but aimed at the rows and columns flowing through your pipeline instead of at a function's return value.
Why would you need this? Think about how data problems actually reach you. A column that was always filled starts arriving empty because an upstream team renamed a field. A currency amount that used to be in riyals is suddenly in fils, a hundred times larger. A nightly export delivers half the usual number of rows because a job timed out. A date column quietly starts containing the string "N/A". None of these crash your code. Your pipeline happily reads the data, trains a model on it, or publishes a dashboard from it. The first sign of trouble is usually a person asking why the numbers look strange, days later, after decisions have been made.
Code fails loudly and data fails silently. That sentence is the whole reason this tool exists. A typo in your Python raises an exception on the first run. A typo in somebody else's data produces no exception at all; it produces wrong answers that look plausible.
Before tools like this, teams handled the problem in one of three ways. Some wrote ad hoc assert statements or if df["price"].min() < 0 checks scattered through their scripts, which works until the scripts multiply and nobody remembers what is being checked. Some wrote SQL queries that count bad rows and pasted them into a scheduled job, which works until the person who wrote them leaves. And many simply did nothing and relied on users to notice. GX gives those checks a shared vocabulary, a standard place to live, a standard format for results, and a way to produce documentation that non-programmers can read.
For machine learning work the stakes are higher than for reporting. A model trained on data with a silently broken column does not raise an error; it learns the breakage. A feature that is suddenly always null can drag accuracy down for weeks before anyone connects the drop to the data. Checking data at the boundaries, where it enters your pipeline and where it enters training, is one of the cheapest reliability improvements available.
Three facts about how GX works will save you confusion later.
It is a library, not a service. There is no server to install, no daemon running in the background and no web application to log into. You pip install it, import it into Python, and call it from a notebook, a script, an Airflow task or a CI job. Everything runs inside your own Python process.
It tests data where the data lives. If your data is a pandas dataframe in memory, GX computes its checks with pandas. If it is a table in PostgreSQL or Snowflake, GX translates the checks into SQL and lets the database do the counting, so you do not pull a hundred million rows into your laptop. If it is a Spark dataframe, the checks run on the Spark cluster.
It is free and open source. GX Core is released under the Apache 2.0 licence. There used to be a paid hosted product called GX Cloud. It was shut down in June 2026 (release 1.18.0), and the library now refuses to start in cloud mode, raising an error that says so. Anything you read about a GX Cloud account, an access token, or a hosted dashboard is obsolete.
great_expectations command-line tool any more
Older tutorials begin with great_expectations init, great_expectations suite new and great_expectations checkpoint run. That command-line interface was removed in version 1.0. Today you do everything from Python code. The only thing the package exposes as a command is python -m great_expectations skills, which installs helper files for AI coding assistants and has nothing to do with testing your data. If a guide tells you to run great_expectations init, close it.
- Think of one dataset you work with. Write down three things that must always be true about it, in plain English, without any code.
- For each one, write down what would go wrong downstream if it stopped being true and nobody noticed for a week.
- Keep this list. You will turn it into Expectations in a few sections.
The mental model: five nouns and a loop
GX has a fairly long vocabulary, but the day-to-day picture needs only a handful of words. Learn them in the order the data flows through them.
A Data Source tells GX how to reach a place where data lives. It might be "the pandas library in this Python process", "a folder on disk", "this PostgreSQL database" or "this S3 bucket". It holds the connection details and nothing else.
A Data Asset is a specific collection of records inside a Data Source: one table, one SQL query, one family of CSV files, or one dataframe. If the Data Source is the warehouse, the asset is one shelf in it.
A Batch Definition says how to slice the asset into pieces you will test. You might test the whole table every time, or one specific file, or only one day's rows, or one month's. The definition is the rule for slicing.
A Batch is what you get when you apply that rule: the actual group of records ready to be tested. You usually ask the Batch Definition for one, sometimes giving it parameters such as {"year": 2024, "month": 2} to say which slice you want.
Now the other half of the picture: what you assert, and how you run the assertions.
An Expectation is a single verifiable statement about data, such as "the column order_id has no null values". In code it is a Python class you create with arguments.
An Expectation Suite is a named collection of Expectations that together describe one dataset. "orders_suite" might hold twelve Expectations about the orders table.
A Validation Definition links one Suite to one Batch Definition. It answers the question "test this data against these rules". Running it produces a result.
A Checkpoint bundles one or more Validation Definitions and adds Actions, which are things to do after the tests run: send a Slack message, rebuild the documentation site, post to a webhook. The Checkpoint is what you run in production. The docs call it the primary means for validating data in a deployment.
Holding all of these together is the Data Context, the object you create first and use to reach everything else. It stores the configuration of your project and gives you methods to add and fetch Data Sources, Suites, Validation Definitions and Checkpoints.
That is a lot of nouns, so here is the same story told as a loop, which is how you will actually work:
- Create a context.
- Tell it where the data is (source, asset, batch definition).
- Write Expectations and collect them into a suite.
- Try the suite on a batch and read the results.
- Wrap the working suite into a Validation Definition and a Checkpoint.
- Run the Checkpoint whenever new data arrives, and react to the outcome.
You do not have to use every noun on day one. For exploration you can skip straight to testing one Expectation on one batch. The extra objects earn their keep when you want a repeatable, automated check.
Two kinds of context
The Data Context comes in two forms that matter to a beginner.
An Ephemeral Data Context lives in memory. When your Python process ends, everything you defined is gone. It is ideal for notebooks, tests and pipeline jobs that rebuild their setup from code each run.
A File Data Context saves its configuration into a folder named gx/ in your project. Suites, Validation Definitions, Checkpoints and Data Source settings persist between runs, and you can commit most of the folder to git. It is ideal when you want to build up a project over time.
There used to be a third, the Cloud Data Context. It is dead, as explained above.
- Without looking back, draw the two chains of nouns on paper: data on the left, rules on the right.
- Mark which object holds connection details, which holds the rules, and which holds the notifications.
- Check your drawing against the diagrams above.
Installing and checking your setup
GX needs Python 3.10, 3.11, 3.12 or 3.13. Python 3.14 is not supported at the time of writing, and 3.9 was dropped in late 2025. The official documentation lists macOS and Linux as supported platforms, and says that although Windows is not officially supported, many people run it there successfully. If you are on Windows and something behaves oddly, the Windows Subsystem for Linux (WSL2) is a sensible fallback, though that is our suggestion rather than an official statement.
First check your Python version:
python --version
If it prints something between 3.10 and 3.13, you are fine. On some systems the command is python3. Then create a virtual environment, which is a private folder of installed packages for one project, so that GX's dependencies do not collide with anything else on your machine. On macOS and Linux:
python -m venv my_venv
source my_venv/bin/activate
python -m pip install --upgrade pip
pip install great_expectations
On Windows PowerShell the same idea looks slightly different:
py -3.12 -m venv my_venv
.\my_venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install great_expectations
After activation your prompt usually shows (my_venv) at the start. Everything you install now goes into that folder and is deleted if you delete the folder.
The core install gives you pandas support and SQLite through Python's standard library. To talk to other systems you install extras, optional groups of dependencies named in square brackets. You quote the name so your shell does not try to interpret the brackets:
pip install 'great_expectations[postgresql]'
pip install 'great_expectations[snowflake]'
pip install 'great_expectations[spark]'
pip install 'great_expectations[s3]'
Use double quotes on Windows. The list of extras is long; it includes bigquery, databricks, sql-server, oracle, redshift, trino, gcp, azure and others. Notice that the old mssql extra was replaced by sql-server, and that there is no cloud extra any more.
Installing with the extra you need is not cosmetic. In September 2026 a new SQLAlchemy release (2.1) broke older GX versions on Python 3.11 and above. The fix in GX 1.23.2 is partly a version cap applied through the extras, so installing great_expectations[snowflake] gets you the protection while a bare install plus a separate Snowflake package might not. If you cannot upgrade, pinning sqlalchemy<2.1 works around it.
Now verify the install from Python:
import great_expectations as gx
print(gx.__version__)
context = gx.get_context()
print(type(context).__name__)
Expected output looks like this:
1.23.2
EphemeralDataContext
The version may be newer when you read this; releases arrive every couple of weeks. The type will be EphemeralDataContext if no gx/ folder was found, or FileDataContext if one was. That single line of output is worth understanding, because it is the source of the most common beginner confusion, which we return to in the errors section.
From a shell you can also run pip show great_expectations to see the installed version and where it lives.
great_expectations==1.23.2 (with your extras, such as great_expectations[postgresql]==1.23.2) in requirements.txt, so the code that works today still works when a colleague installs it next month. Upgrade deliberately, read the changelog, and re-run your suites when you do.
- Create a fresh folder and virtual environment, then install GX.
- Run the three-line verification script and note the version and context type.
- Run
pip show great_expectationsand confirm the same version appears.
Your first validation in five minutes
Before building any project structure, let us see GX catch a real problem. Create a small CSV file called orders.csv with a deliberate mess in it:
order_id,customer,country,amount
1001,Amal,EG,120.50
1002,Omar,SA,89.00
1003,Layla,AE,45.25
1003,Youssef,EG,300.00
1005,,SA,-15.00
1006,Nour,XX,60.00
Look at it for a moment. Order 1003 appears twice. Order 1005 has no customer and a negative amount. Order 1006 has a country code, XX, that does not exist. A human notices these by reading six rows; nobody notices them in six million. Now let GX do the noticing.
import great_expectations as gx
context = gx.get_context()
batch = context.data_sources.pandas_default.read_csv("orders.csv")
print(batch.head())
pandas_default is a built-in Data Source that GX creates for you so that you can explore without any setup. Calling read_csv on it reads the file into pandas and hands you back a Batch directly. The head() call shows the first rows, just as in pandas.
Now create an Expectation and test it:
not_null = gx.expectations.ExpectColumnValuesToNotBeNull(column="customer")
result = batch.validate(not_null)
print(result)
The printed result is a structured object. The important parts look like this:
{
"success": false,
"expectation_config": {
"type": "expect_column_values_to_not_be_null",
"kwargs": { "column": "customer", "batch_id": "pandas_default-#ephemeral_pandas_asset" }
},
"result": {
"element_count": 6,
"unexpected_count": 1,
"unexpected_percent": 16.666666666666664,
"partial_unexpected_list": [null]
}
}
Read it top to bottom. success: false is the verdict. expectation_config repeats what you asked: the snake_case type string is the same Expectation as the class name ExpectColumnValuesToNotBeNull, written the way it appears in saved files. element_count is how many rows were examined, unexpected_count is how many broke the rule, and unexpected_percent turns that into a percentage. The exact batch identifier will differ on your machine.
Try a range check on the amount column:
positive = gx.expectations.ExpectColumnValuesToBeBetween(
column="amount", min_value=0, max_value=10000
)
result = batch.validate(positive)
print(result.success)
print(result.result["partial_unexpected_list"])
That prints False and [-15.0], which is the offending value. You can tweak an Expectation and re-test it instantly. Set the minimum to -100 and run it again and it passes, because the rule is now looser:
positive.min_value = -100
print(batch.validate(positive).success)
This exploratory loop (create, validate, read, tweak) is the most useful habit in GX. You are not committing anything; you are asking questions of the data and learning what "normal" looks like. Only when an Expectation captures something you truly believe do you keep it.
A few more will demonstrate the range of what is available. Each is a class under gx.expectations:
unique_ids = gx.expectations.ExpectColumnValuesToBeUnique(column="order_id")
valid_country = gx.expectations.ExpectColumnValuesToBeInSet(
column="country", value_set=["EG", "SA", "AE", "KW", "QA"]
)
row_count = gx.expectations.ExpectTableRowCountToBeBetween(min_value=1, max_value=1000000)
for exp in (unique_ids, valid_country, row_count):
r = batch.validate(exp)
print(exp.__class__.__name__, r.success)
You should see that the uniqueness and country checks fail (order 1003 repeats and XX is not in the set) while the row count passes. In three minutes you found all three problems in the file.
ExpectColumnValuesToBeUnique. In results and saved JSON you will see expect_column_values_to_be_unique. Old tutorials use the snake_case form as a method call on a validator object, such as validator.expect_column_values_to_be_unique("id"). That style belongs to the 0.x API. In 1.x you instantiate the class and pass it to batch.validate(...).
- Save the CSV above as
orders.csvand run the snippets in a notebook or script. - Fix the file so that every Expectation above passes, then re-run.
- Add an Expectation of your own, for example that
customervalues are not in the set["", "unknown"].
Building a reusable suite from a dataframe
Testing one Expectation at a time is fine for exploring, but you want to keep the good ones together and run them as a unit. That is what an Expectation Suite is for. We will also stop using the built-in shortcut and set up the objects properly, because that is what you will do in every real project.
In a real pipeline your data usually arrives as a pandas dataframe that an earlier step produced. Here is the full setup, written to work in an in-memory context so that nothing is left on disk:
import pandas as pd
import great_expectations as gx
df = pd.read_csv("orders.csv")
context = gx.get_context(mode="ephemeral")
data_source = context.data_sources.add_pandas(name="my_pandas")
asset = data_source.add_dataframe_asset(name="orders_df")
batch_definition = asset.add_batch_definition_whole_dataframe("whole_orders")
batch = batch_definition.get_batch(batch_parameters={"dataframe": df})
Walk through it. add_pandas creates a Data Source for pandas. add_dataframe_asset adds an asset meaning "a dataframe". add_batch_definition_whole_dataframe says "the whole thing, no slicing". And get_batch needs to know which dataframe, since a dataframe asset has no fixed content, so you pass it at run time in batch_parameters. The same asset and definition can be reused tomorrow with tomorrow's dataframe.
Notice the argument mode="ephemeral". Passing the mode explicitly removes guesswork. With no argument, gx.get_context() looks for an existing project and falls back to in-memory if it finds none; with mode="ephemeral" you always get memory, even if a gx/ folder exists nearby.
Now build the suite:
suite = context.suites.add(gx.ExpectationSuite(name="orders_suite"))
suite.add_expectation(gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id"))
suite.add_expectation(gx.expectations.ExpectColumnValuesToBeUnique(column="order_id"))
suite.add_expectation(gx.expectations.ExpectColumnValuesToBeBetween(
column="amount", min_value=0, max_value=10000))
suite.add_expectation(gx.expectations.ExpectColumnValuesToBeInSet(
column="country", value_set=["EG", "SA", "AE", "KW", "QA"]))
suite.add_expectation(gx.expectations.ExpectTableRowCountToBeBetween(min_value=1))
context.suites.add(...) registers the suite with the context and returns it. Each add_expectation call attaches an Expectation and, in a context that persists, saves it. Now run the whole suite against the batch:
result = batch.validate(suite)
print(result.success)
print(result.statistics)
The first line prints False (the file has problems) and the second prints a small summary with counts such as evaluated_expectations, successful_expectations, unsuccessful_expectations and success_percent. One failing Expectation makes the whole suite fail.
You can look at each Expectation's individual result:
for r in result.results:
print(r.expectation_config.type, r.success)
That lists one line per Expectation, which is the quickest way to see which rules broke.
Changing an Expectation you already saved
Beginners often try to fix a rule by editing the object and then wonder why nothing changed. To change an Expectation that is already in a suite, edit its attributes and call save() on it:
for exp in suite.expectations:
if exp.__class__.__name__ == "ExpectColumnValuesToBeBetween":
exp.max_value = 50000
exp.save()
To fetch a suite you saved earlier, use context.suites.get(name="orders_suite"). If you edit the suite itself, for example by deleting an Expectation, call suite.save() afterwards. GX will remind you with an error saying the suite "has changed since it has last been saved" if you forget.
RuntimeError: Cannot add Expectation because it already belongs to an ExpectationSuite. Create a new object, or use copy.copy(expectation) and set its id to None, before adding it elsewhere.
- Build the suite above from your own dataframe.
- Run it, then list each Expectation with its success value.
- Loosen exactly one rule with
exp.max_value = ...andexp.save(), and re-run to see the statistics change.
Choosing and tuning Expectations
GX ships around sixty core Expectation classes, and you can browse every one, with examples, in the Expectation Gallery on the GX website. You will not memorise them. What you should develop is a sense of the families, so that when you have a data worry you know which family to reach for.
Shape of the table. ExpectTableRowCountToBeBetween and ExpectTableRowCountToEqual check volume, which catches the half-empty export. ExpectTableColumnsToMatchSet and ExpectTableColumnsToMatchOrderedList check that the expected columns exist, which catches a renamed or dropped field. ExpectColumnToExist is the gentlest version of that check.
Missing and unique values. ExpectColumnValuesToNotBeNull and ExpectColumnValuesToBeNull are the basics. ExpectColumnValuesToBeUnique protects keys. ExpectCompoundColumnsToBeUnique protects composite keys such as the pair (customer_id, order_date).
Allowed values and ranges. ExpectColumnValuesToBeInSet and ExpectColumnValuesToNotBeInSet handle categories. ExpectColumnValuesToBeBetween handles numeric and date ranges.
Text patterns and lengths. ExpectColumnValuesToMatchRegex checks formats such as phone numbers or IDs. ExpectColumnValueLengthsToBeBetween catches truncated or runaway text. ExpectColumnValuesToMatchStrftimeFormat checks that a string column holds dates in a particular format.
Types. ExpectColumnValuesToBeOfType and ExpectColumnValuesToBeInTypeList check that a column holds the type you expect, which catches numbers arriving as strings.
Statistics of a whole column. ExpectColumnMeanToBeBetween, ExpectColumnMedianToBeBetween, ExpectColumnMinToBeBetween, ExpectColumnMaxToBeBetween and ExpectColumnStdevToBeBetween check aggregates. These are the ones that catch the silent unit change: if amounts were averaging 120 and suddenly average 12,000, a range check per row may still pass but the mean check fails.
Relationships between columns. ExpectColumnPairValuesAToBeGreaterThanB catches an end date before its start date. ExpectMulticolumnSumToEqual checks that parts add up to a total.
There are two kinds of Expectation worth distinguishing, because it affects how you read results. Some judge each row: "every value in amount is between 0 and 10000". These produce a count of unexpected rows. Others judge the column or table as a whole: "the mean of amount is between 50 and 500". These produce a single observed value, which the result shows under observed_value.
The mostly argument
Real data is rarely perfect, and a rule that demands 100 percent purity fails every day and gets ignored. Row-level Expectations accept a mostly argument, a fraction between 0 and 1 that says how much of the data must satisfy the rule:
gx.expectations.ExpectColumnValuesToNotBeNull(column="customer", mostly=0.95)
This passes if at least 95 percent of customers are filled in. Use mostly when the business accepts a small, known amount of mess, and keep it tight enough that a real regression still trips it. The number is a statement about your data, so write it down where a colleague can question it.
Severity and descriptions
Two optional arguments make Expectations kinder to the humans reading results. severity can be "critical" (the default), "warning" or "info", and later you can tell actions to notify only on a given severity. description is a sentence that appears in the documentation site:
gx.expectations.ExpectColumnValuesToBeBetween(
column="amount",
min_value=0,
max_value=10000,
severity="warning",
description="Order amounts are positive and below 10,000 in local currency.",
)
A description turns a technical check into something a product manager can read and challenge.
Brittle
min_value=44.99, max_value=312.40copied from last week's minimum and maximum- The first unusually large order breaks it
Meaningful
min_value=0, max_value=10000from a business rule: no negative orders, nothing above the approval limit- A failure means something
- Take the three facts you wrote at the start of this guide and pick the Expectation class that matches each.
- Look each class up in the Expectation Gallery and note its arguments.
- Add a
mostlyvalue to one of them and explain in a sentence why that number is right.
Reading validation results
A result object carries a lot of information. Knowing which parts to read keeps you from drowning in it.
At the top level a suite result has success (true only if every Expectation passed), statistics (the counts), results (one entry per Expectation) and meta (run information). Each entry in results has its own success, the expectation_config you asked for, and a result dictionary whose contents depend on the result format.
The result format controls how much detail you get back. There are four levels:
| Level | What you get |
|---|---|
BOOLEAN_ONLY |
Only success for each Expectation. The cheapest. |
BASIC |
Adds counts, percentages, the observed_value, and a short list of unexpected values. |
SUMMARY |
Adds frequency counts of unexpected values and their row positions. This is the default for Checkpoints. |
COMPLETE |
Adds the full list of unexpected values, all their row indexes, and a query to fetch them. The most detailed and the largest. |
You choose the level when you validate:
result = batch.validate(suite, result_format="COMPLETE")
In a Checkpoint you pass result_format to the constructor, and in a Validation Definition run you pass it as result_format={"result_format": "SUMMARY"}. The choice is a trade-off. Detailed results help you debug, but they can be very large on a big table, and they contain actual data values, which matters if the data is sensitive. A sensible habit is SUMMARY by default and COMPLETE only when you are investigating.
Finding the bad rows
The failing rows are what you really want. By default results list only the offending values, which is not enough when a value like -15.0 could appear in a hundred rows. Ask GX to report which rows by naming a column that identifies them:
result = batch.validate(
suite,
result_format={
"result_format": "COMPLETE",
"unexpected_index_column_names": ["order_id"],
},
)
Now each failing Expectation includes an unexpected_index_list, a list of dictionaries holding the order_id and the bad value for each failing row. A few other options are worth knowing: partial_unexpected_count controls how many example values are shown (set it to 0 to show none), and include_unexpected_rows returns up to 200 complete failing rows for the row-level Expectation types.
You can print a readable summary of a Checkpoint result with result.describe(), which returns a JSON-style string of what passed and failed. There is also describe_dict() if you would rather work with a Python dictionary.
element_count first
When a result looks wrong, check element_count before anything else. If it is 0 or far from what you expected, the problem is not your rules: the batch is empty, the wrong file was loaded, or a filter removed everything. Many mysterious passes are just Expectations evaluating zero rows.
- Run your suite with each of the four result formats and compare the size of the output.
- Use
unexpected_index_column_namesto get the order IDs of the failing rows. - Fix just those rows in your dataframe and confirm the suite passes.
Connecting to files and databases
Dataframes are convenient, but in production your data more often sits in files or in a database. The setup has the same shape (source, asset, batch definition) with different method names.
A folder of files
For CSV, Parquet, JSON and similar files on disk, use the pandas filesystem Data Source pointed at a base folder:
data_source = context.data_sources.add_pandas_filesystem(
name="files", base_directory="./data"
)
csv_asset = data_source.add_csv_asset(name="taxi_csv")
There are sibling methods for other formats, such as add_parquet_asset, add_json_asset and add_excel_asset. Now define how to slice it. The simplest is one specific file:
january = csv_asset.add_batch_definition_path(
name="january", path="yellow_tripdata_sample_2019-01.csv"
)
batch = january.get_batch()
When the files follow a naming pattern with a date in it, define a monthly batch definition with a regular expression that has named groups called year and month:
monthly = csv_asset.add_batch_definition_monthly(
name="monthly",
regex=r"yellow_tripdata_sample_(?P<year>\d{4})-(?P<month>\d{2})\.csv",
)
batch = monthly.get_batch(batch_parameters={"year": 2019, "month": 1})
A regular expression is a compact pattern language for matching text; \d{4} means "four digits". The named groups (?P<year>...) tell GX which part of the file name is the year. The pattern is matched relative to the base directory. Batch parameters should be integers; older versions accepted digit strings like "2019", and 1.21.0 deprecated that. If you call get_batch() with no parameters on a partitioned definition, you receive the latest matching batch.
A database table
For SQL databases, use the Data Source matching your database. PostgreSQL is a good example:
data_source = context.data_sources.add_postgres(
name="warehouse",
connection_string="${POSTGRES_CONNECTION_STRING}",
)
asset = data_source.add_table_asset(name="trips", table_name="public.trips")
whole = asset.add_batch_definition_whole_table("FULL_TABLE")
Two details matter here. First, the connection string uses the ${POSTGRES_CONNECTION_STRING} placeholder instead of the literal password; the configuration section explains why. Second, the table name is schema-qualified (public.trips). Older code used a separate schema_name argument, which was deprecated in 1.14.0 and is scheduled for removal in 2.0, so always put the schema inside table_name.
To test only part of a table, define a date-partitioned batch definition using a date column:
daily = asset.add_batch_definition_daily(name="DAILY", column="pickup_datetime")
batch = daily.get_batch(batch_parameters={"year": 2020, "month": 1, "day": 14})
This makes GX add a filter on pickup_datetime so that only one day's rows are tested, and the database does the work. Validating yesterday's increment instead of the whole table keeps queries fast and your database bill small. There are also add_batch_definition_monthly and add_batch_definition_yearly.
Other SQL factories exist on context.data_sources, including add_sqlite, add_snowflake, add_bigquery, add_databricks_sql, add_redshift, add_sql_server and a generic add_sql that accepts any SQLAlchemy connection URL. You can also test the result of a query instead of a table with add_query_asset, provided the SQL starts with SELECT (a query beginning with WITH is rejected; wrap it in a subquery or create a view).
GX issues queries as the database user you give it. Use an account with read-only access (SELECT on the tables you validate) so that a mistake in a test can never change data.
context.data_sources.get("warehouse"), or use the add_or_update_* variant (for example add_or_update_postgres) so the script is safe to re-run. The same applies to suites, Validation Definitions and Checkpoints, which have add_or_update too.
- Put two or three CSV files in a
datafolder and add a filesystem Data Source for them. - Create one path-based and one monthly batch definition, and fetch a batch from each.
- If you have a local PostgreSQL or a SQLite file, add a table asset with a whole-table batch definition and validate one Expectation against it.
Validation Definitions and Checkpoints
So far you have called batch.validate(suite) by hand. That is a good way to explore but not a good way to run something every night. Two objects turn it into a repeatable operation.
A Validation Definition pairs one Batch Definition with one Suite and gives the pair a name:
validation_definition = context.validation_definitions.add(
gx.ValidationDefinition(
name="orders_whole_vd",
data=batch_definition,
suite=suite,
)
)
You run it by supplying batch parameters, if the definition needs any:
vd_result = validation_definition.run(
batch_parameters={"dataframe": df},
result_format={"result_format": "SUMMARY"},
)
print(vd_result.success)
A Validation Definition is all you need when you just want the verdict for one dataset. A Checkpoint adds two things: it can group several Validation Definitions, and it can trigger Actions afterwards.
checkpoint = context.checkpoints.add(
gx.Checkpoint(
name="orders_checkpoint",
validation_definitions=[validation_definition],
result_format={"result_format": "SUMMARY"},
)
)
checkpoint_result = checkpoint.run(batch_parameters={"dataframe": df})
print(checkpoint_result.success)
The batch parameters you pass to checkpoint.run apply to every Validation Definition inside it. So group definitions together only when they accept the same kind of parameters. A monthly definition expects year and month; handing it dataframe produces an error about the keys not being valid.
Using the result to control a pipeline
The reason to have a Checkpoint is to make a decision. The simplest decision is to stop:
if not checkpoint_result.success:
print(checkpoint_result.describe())
raise SystemExit(1)
Exiting with a non-zero status code makes schedulers such as Airflow, cron wrappers or CI systems mark the step as failed, which stops downstream steps from consuming bad data. That is the moment where data testing stops being a report and becomes a safety mechanism.
add_or_update in setup scripts
A script that calls context.suites.add(...) works once and then fails with "already exists". In scripts that rebuild the project on every run, use context.suites.add_or_update(...), and the matching methods on validation_definitions and checkpoints, so that the script can be repeated safely.
- Wrap your suite and batch definition in a Validation Definition and run it.
- Wrap that in a Checkpoint and run it with your dataframe.
- Add the
if not checkpoint_result.successblock, run it against the broken CSV, and check your shell's exit code withecho $?.
Data Docs: a report people can read
Validation results are data, and most people would rather read a web page. Data Docs are static websites GX generates from your Expectations and validation results. They show each suite as a readable list of rules (using your description text), and each validation run as a pass or fail table with the offending values. You can open them in a browser without any server, and you can share them by publishing the folder to a web host.
Data Docs need somewhere to live, which means a File Data Context. Create one:
import great_expectations as gx
context = gx.get_context(mode="file", project_root_dir="./my_project")
If the folder ./my_project/gx does not exist, GX creates it. Otherwise it loads what is there. Mode "file" without a root directory looks for the nearest existing project, or creates gx/ in the current directory.
The folder layout
After you create a File Data Context, your project contains:
my_project/
gx/
great_expectations.yml
expectations/
validation_definitions/
checkpoints/
plugins/
uncommitted/
config_variables.yml
validations/
data_docs/local_site/
The pieces matter. great_expectations.yml is the main configuration file: it lists the stores and also records your Data Sources, assets and batch definitions. The expectations, validation_definitions and checkpoints folders hold your suites and other definitions as JSON files. You commit all of these to git, and a colleague who clones the repository gets your whole setup.
The uncommitted folder is for things that should not go into git: the secrets file config_variables.yml, the stored validation results, and the generated Data Docs. GX writes a .gitignore that excludes it. Validation results and Data Docs can contain real data values, so treat them as sensitive.
Building and opening the docs
The UpdateDataDocsAction is the normal way to rebuild the site after each run. You can also trigger a build by hand:
context.build_data_docs()
context.open_data_docs()
build_data_docs regenerates the site under gx/uncommitted/data_docs/local_site/, and open_data_docs launches it in your default browser. The home page lists suites and recent validation runs. Click a run to see each Expectation with a green tick or red cross, the observed value, and the unexpected values.
A limitation worth knowing: since version 1.13.0 (February 2026) the only place GX can write stores and Data Docs is a local or network file system. The old S3, GCS, Azure and database store backends were removed. If you need your docs on a shared website, you copy or sync the folder yourself, for example to a protected bucket or an internal web server. The official Data Docs guidance also states that only the base_directory key of the site configuration is supported.
gx.get_context() from a different folder, or an environment variable named GX_HOME points at a folder without a great_expectations.yml, GX silently falls back to an in-memory context and your saved work appears to be gone. Nothing was deleted; you are just not looking at the project. Use gx.get_context(mode="file", project_root_dir="./my_project") so that the location is explicit, and print type(context).__name__ when in doubt.
- Create a File Data Context with
project_root_dirand list the folders it makes. - Rebuild your suite and Checkpoint inside it, run the Checkpoint, then call
context.build_data_docs()andcontext.open_data_docs(). - Open the
orders_suitepage and find yourdescriptiontext.
Actions: doing something with the result
An Action runs after a Checkpoint's validations finish and reacts to the result. You list Actions when creating the Checkpoint. They execute in order, once per Checkpoint run.
Version 1 of GX has no default Actions. In the old API, storing results and updating documentation happened automatically. Now if you want the Data Docs to refresh after a run, you must add UpdateDataDocsAction yourself. Forgetting it is a classic source of "why is my report out of date".
The built-in Actions include UpdateDataDocsAction, SlackNotificationAction, MicrosoftTeamsNotificationAction, EmailAction, PagerdutyAlertAction, OpsgenieAlertAction, SNSNotificationAction and APINotificationAction, which posts to a URL. The official compatibility reference lists Email, Microsoft Teams, Slack and custom Actions as supported; treat the rest as available but less formally covered.
A Checkpoint with documentation and a Slack alert looks like this:
from great_expectations.checkpoint import (
SlackNotificationAction,
UpdateDataDocsAction,
)
checkpoint = context.checkpoints.add_or_update(
gx.Checkpoint(
name="orders_checkpoint",
validation_definitions=[validation_definition],
actions=[
UpdateDataDocsAction(name="update_docs"),
SlackNotificationAction(
name="slack_on_fail",
slack_token="${SLACK_BOT_TOKEN}",
slack_channel="${SLACK_CHANNEL}",
notify_on="failure",
show_failed_expectations=True,
),
],
result_format={"result_format": "COMPLETE"},
)
)
The notify_on argument decides when a notification is sent. The accepted values are all, success, failure, critical, warning and info. The failure value fires on any failure regardless of severity. critical fires on a critical failure, or when an Expectation could not execute at all. warning fires when the worst failure is a warning, and info when the worst is informational. The highest severity present wins, so one critical failure in a run of mostly warnings still counts as critical.
Tokens never appear as literals in the code. The ${SLACK_BOT_TOKEN} placeholder is replaced at run time from an environment variable or the secrets file, which is the topic of the next section.
A note of realism: alerting well is harder than alerting at all. If every run posts a message, people mute the channel within a week. Start with notify_on="failure", include the failed Expectations so the message is actionable, and route the alert to the people who own the data source, not to a general channel.
- Add
UpdateDataDocsAction(name="update_docs")to your Checkpoint, run it twice with different data, and confirm the docs show both runs. - Read the Slack, Teams or Email action in the official docs and decide which your team would actually use.
- Explain in one sentence when you would choose
notify_on="critical"over"failure".
Configuration, credentials and project settings
Your first configuration problem will be credentials. A connection string for a real database contains a password, and you must never put that password in a file that goes into git.
GX solves this with substitution variables. Anywhere a credential field accepts it, you write ${NAME} instead of the literal value. When GX loads the configuration it replaces ${NAME} with the value from one of two places: an environment variable called NAME, or an entry called NAME in gx/uncommitted/config_variables.yml. If both exist the environment variable wins. You can also pass a runtime_environment dictionary to get_context, which overrides both.
POSTGRES_CONNECTION_STRING: "postgresql+psycopg2://report_user:s3cret@db.example.com:5432/analytics"
SLACK_BOT_TOKEN: "xoxb-your-token"
SLACK_CHANNEL: "data-quality-alerts"
Because that file lives under uncommitted/, git ignores it. In containers and CI, prefer environment variables:
export POSTGRES_CONNECTION_STRING="postgresql+psycopg2://report_user:s3cret@db.example.com:5432/analytics"
python run_checkpoint.py
When a password contains a literal dollar sign, escape it as \$, or GX will try to substitute it and fail. GX also restricts substitution to fields that are designed for it, such as connection_string and tokens. If you try it elsewhere you get an error saying that only certain fields may use config substitution. For larger teams, GX can also read secrets from AWS Secrets Manager, Google Secret Manager and Azure Key Vault, which the Mid-level guide covers.
- Move a connection string out of your code into an environment variable and use
${...}in the Data Source definition. - Open
gx/great_expectations.ymland find where your Data Source was saved. - Check that
gx/uncommittedappears in the generated.gitignoreand thatconfig_variables.ymldoes not show ingit status.
Common errors and how to read them
GX errors are usually specific, and the message usually contains the fix. Here are the ones beginners meet most, with the real text to match against.
GreatExpectationsError: GX Cloud has been shut down, so this no longer functions and will be removed in great_expectations 2.0. You asked for cloud mode, directly or by accident. The accident is usually leftover environment variables whose names begin with GX_CLOUD_, which make get_context() try cloud mode. Unset them, or call gx.get_context(mode="file") or mode="ephemeral".
MissingConfigVariableError: Unable to find a match for a config substitution variable. A ${NAME} placeholder has no value. Export the environment variable or add it to config_variables.yml. If the value itself contains a dollar sign, escape it as \$.
DataContextError: ExpectationSuite with name orders_suite was not found. The suite does not exist in the context you are using. Nine times out of ten you are in an ephemeral context that forgot it, or you loaded a different project. Run context.suites.all() to see what exists, and use mode="file" with project_root_dir so you always load the same project. The same message pattern exists for Validation Definitions and Checkpoints.
DataContextError: Cannot add ExpectationSuite with name orders_suite because it already exists. You re-ran setup code. Use add_or_update, or fetch with get. Data Sources produce a similar message: "Can not write the fluent datasource ... because a datasource of that name already exists".
ExpectationSuite 'orders_suite' has changed since it has last been saved. Please update with <SUITE_OBJECT>.save(), then try your action again. You edited a suite in memory and then used it. Call suite.save().
ExpectationSuite ... must be added to the DataContext before it can be updated. You built a suite object but never registered it. Call context.suites.add(suite) first.
TestConnectionError: Path: ... does not exist. The base_directory of a filesystem Data Source is wrong. Relative paths resolve from your current working directory, which differs between a notebook and a scheduled job. Use an absolute path in anything that runs unattended.
query must start with 'SELECT' followed by a whitespace. A query asset begins with WITH, a comment or some other keyword. Begin with SELECT, or move the logic into a view.
Batch parameters should only contain keys from the following set: ... which is not valid. You passed parameters the Batch Definition does not accept: day to a monthly definition, or dataframe to a table. Match the keys to the definition, and remember that all definitions in a Checkpoint share one set of parameters.
ModuleNotFoundError: No module named 'psycopg2' (or snowflake, pyodbc, pyspark, boto3). You are missing an extra. Install great_expectations[postgresql] or whichever extra matches. SQL Server additionally needs the Microsoft ODBC driver installed on the operating system, which pip cannot do.
A different class of problem produces no error at all. An Expectation passes on an empty batch. Always check element_count. Saved objects disappear between runs. You are in an ephemeral context. A regular expression Expectation behaves differently after upgrading to 1.23.2 on Snowflake. That release made Snowflake use substring matching like the other back ends, where it previously required the whole value to match. Anchor patterns with ^ and $ when you mean a full match.
type(context).__name__)? Which folder is it reading (project_root_dir, GX_HOME, the current directory)? And what version is installed (gx.__version__)? Most beginner problems are answered by one of these.
- Deliberately trigger three errors: add the same suite twice, read a missing suite, and pass a wrong key in
batch_parameters. - For each, find the line of the message that names the fix.
- Fix each one with the method the message suggests.
Putting it all together
Let us build one small project from start to finish, the kind you could put in a repository and run every night. It validates a daily orders CSV, rebuilds the documentation, and fails the pipeline if the data is bad.
The folder looks like this:
orders_quality/
data/orders_2024-06-01.csv
run_quality.py
requirements.txt
The requirements file pins the version:
great_expectations==1.23.2
pandas
The script uses a File Data Context so that the docs and results persist, and add_or_update everywhere so that it can be re-run:
import sys
import pandas as pd
import great_expectations as gx
from great_expectations.checkpoint import UpdateDataDocsAction
df = pd.read_csv("data/orders_2024-06-01.csv")
context = gx.get_context(mode="file", project_root_dir=".")
data_source = context.data_sources.add_or_update_pandas(name="orders_source")
asset = data_source.add_dataframe_asset(name="orders_df")
batch_definition = asset.add_batch_definition_whole_dataframe("whole_orders")
suite = gx.ExpectationSuite(name="orders_suite")
suite.add_expectation(gx.expectations.ExpectColumnToExist(column="order_id"))
suite.add_expectation(gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id"))
suite.add_expectation(gx.expectations.ExpectColumnValuesToBeUnique(column="order_id"))
suite.add_expectation(gx.expectations.ExpectColumnValuesToNotBeNull(
column="customer", mostly=0.98,
description="At least 98 percent of orders name a customer."))
suite.add_expectation(gx.expectations.ExpectColumnValuesToBeBetween(
column="amount", min_value=0, max_value=10000,
description="Order amounts are positive and below the approval limit."))
suite.add_expectation(gx.expectations.ExpectColumnValuesToBeInSet(
column="country", value_set=["EG", "SA", "AE", "KW", "QA"],
severity="warning"))
suite.add_expectation(gx.expectations.ExpectTableRowCountToBeBetween(min_value=100))
suite = context.suites.add_or_update(suite)
validation_definition = context.validation_definitions.add_or_update(
gx.ValidationDefinition(name="orders_vd", data=batch_definition, suite=suite)
)
checkpoint = context.checkpoints.add_or_update(
gx.Checkpoint(
name="orders_checkpoint",
validation_definitions=[validation_definition],
actions=[UpdateDataDocsAction(name="update_docs")],
result_format={
"result_format": "SUMMARY",
"unexpected_index_column_names": ["order_id"],
},
)
)
result = checkpoint.run(batch_parameters={"dataframe": df})
print(result.describe())
if not result.success:
print("Data quality check failed; see gx/uncommitted/data_docs/local_site/index.html")
sys.exit(1)
Read it as a story. The script loads the day's data, opens the project, and declares its Data Source, asset and batch definition using calls that are safe to repeat. It declares the rules, each with a business-facing description where it helps, and uses mostly for the one rule where small gaps are tolerated and severity="warning" for the country list, which we want to hear about but not block on. It registers the suite, links it to the batch definition in a Validation Definition, and wraps that in a Checkpoint whose only Action rebuilds the documentation. Finally it runs the Checkpoint, prints a description and exits with status 1 if anything failed.
Note one subtlety in this script: a warning-severity failure still makes result.success false, so it blocks here. That is deliberate for a beginner project; blocking and then choosing what to relax is safer than silently letting bad data through. When you later want warnings not to block, you would inspect severities in the result instead of the single flag.
Run it from a scheduler. A cron line such as 0 6 * * * cd /srv/orders_quality && ./venv/bin/python run_quality.py would run it at 06:00 daily. In Airflow the same script becomes a task, and the non-zero exit marks it failed; the Airflow guide shows how tasks fit into a larger workflow, and GX also publishes an official Airflow provider. If you run pipelines in containers, the Docker guide covers packaging this script with its dependencies.
Commit the gx/ folder, excluding what .gitignore already excludes, so teammates get the same suites and Checkpoints. Open gx/uncommitted/data_docs/local_site/index.html after a failing run and look at how the report names the failing rows by order_id.
Where does this sit in a larger pipeline? The official guidance suggests three natural places: at ingestion, to quarantine bad raw records before they spread; after transformation, to decide whether downstream jobs should run; and at delivery, to tell genuine business anomalies apart from data problems. For machine learning, the same idea applies at the point where training data is assembled. If you track experiments, put the data check before the run starts, so a failed check prevents a wasted training job; the MLflow guide shows how to record what data a run used, and Evidently covers a different question, whether live data has drifted from the training data. GX asks "is this data valid?"; drift tools ask "is this data different?". They complement each other.
- Build the project above with your own small dataset and run it twice, once clean and once with a deliberate defect.
- Confirm the exit code is 0 for the clean run and 1 for the bad one.
- Open the Data Docs, find the failing rows by order ID, and write one sentence describing what a colleague should do about them.
What you can now do, and what comes next
If you worked through the Try it tasks, you can now do things that many working data teams still do not. You can explain what GX is for and why the old tutorials are wrong. You can install it with the right extras and confirm the version. You can explore a dataset with throwaway Expectations, collect the good ones into a suite, and change them safely. You can connect to a dataframe, a folder of files, or a database table and slice the data into batches. You can turn a suite into a Validation Definition and a Checkpoint, choose a result format, find the exact rows that failed, and make a pipeline stop when the data is bad. You can keep credentials out of git, build a readable Data Docs site, and read the common errors instead of guessing.
Before you move on, two habits are worth forming. First, treat Expectations as code that deserves review: a rule is a claim about the business, so someone who understands the business should agree with it. Second, treat failures as information. A check that nobody ever sees fail is either testing something trivial or not being run. A check that fails daily and is ignored is worse than no check. Aim for a small set of meaningful rules that people trust.
Natural next steps in this catalogue are the Delta Lake guide for storage that already enforces some constraints, the Spark guide for validating large datasets, and the Dagster and Prefect guides for orchestrators that can run a Checkpoint as one step. If you work on Databricks, see the Databricks guide.
Sources
- Great Expectations documentation home
- GX Core introduction
- GX overview
- Try GX
- Changelog and deprecation timeline
- Compatibility reference
- Glossary
- Install Python
- Install GX
- Install additional dependencies
- Create a Data Context
- Connect to data
- Connect to SQL data
- Connect to filesystem data
- Connect to dataframes
- Create an Expectation
- Test an Expectation
- Organize Expectation Suites
- Create a Validation Definition
- Run a Validation Definition
- Create a Checkpoint with Actions
- Run a Checkpoint
- Choose a result format
- Configure credentials
- Configure Data Docs
- GX in your data pipeline
- Expectation Gallery
- great-expectations on PyPI
- Source repository and releases