Skip to content
Back to student guides
lakeFSMLOpsData & features3 levels93 sectionsCovers lakeFS 1.87

The Complete lakeFS Guide

Version data lakes with Git-like branches, commits and merges using lakeFS. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
16sections
44examples

This is part one of three. It takes you from never having seen lakeFS to using it confidently on a real dataset. By the end you can start a lakeFS server on your laptop, create a repository, put files into it, branch it, change data safely on the branch, compare the change with the original, merge it or throw it away, and tag the result so you can find it again in six months. You will also understand which parts of lakeFS are free, which are not, and what changed in the release this guide was checked against, lakeFS 1.87.

Each section ends with a Try it task. Do them as you go. Version control for data is an idea that feels abstract until you have watched a branch appear instantly next to a dataset that holds a few gigabytes, so please run the commands rather than only reading them.

What lakeFS is and what you can do with it

lakeFS puts a Git-like layer of version control on top of an object store. An object store is the kind of storage that holds files as "objects" inside "buckets": Amazon S3, Google Cloud Storage, Azure Blob Storage, MinIO, and other storage that speaks the S3 protocol. Your data stays in your bucket, in the format you already use. lakeFS does not convert it, does not hold a second copy of it, and does not care whether the files are Parquet, CSV, images, audio, model checkpoints, or Delta and Iceberg table files. What lakeFS adds is a set of pointers and metadata that say which version of which file belongs to which branch and which commit.

The best way to feel why this matters is to picture a data team on a normal Tuesday. A nightly job rewrites a folder of Parquet files that a dashboard and two machine learning models read from. One night the job has a bug, and the folder now holds half-written files and wrong numbers. Everything downstream is broken, nobody is sure what the folder looked like yesterday, and the only "backup" is a copy somebody made three weeks ago. With lakeFS the nightly job writes to a branch, a check runs against the branch, and only if the check passes is the branch merged into main. Readers who point at main never see the broken state. If something slips through anyway, you revert to the previous commit, and the folder is back exactly as it was.

OBJECT STOREyour bucket, your files
→
lakeFSbranches, commits, tags
→
YOUR TOOLSSpark, Python, S3 clients

In practice people use it for four things. First, isolated experiments: give every engineer or every pipeline run a private branch of the whole dataset, at almost no cost. Second, safe publishing: write new data on a branch, validate it, then merge it in one atomic step so readers see either the old version or the new one and never a half-finished mix. Third, reproducibility: a commit is an immutable snapshot, so a model trained on lakefs://my-repo/3f9a.../data/ can be retrained on exactly the same data a year later. Fourth, fast recovery: undoing a bad change is a metadata operation that takes moments, however large the data is.

Everything in this guide is checked against lakeFS v1.87.0, released on 22 September 2026. You will see the word "Community" a lot. It names the free edition you download and run yourself. The commercial editions are called Team, Enterprise and Cloud, and they add things such as multiple users, permissions and single sign-on. Section four explains the difference, because one change in 1.87 affects almost every beginner tutorial you will find on the internet.

Try it
  1. Think of one dataset you have overwritten by mistake, or been afraid to overwrite. Write down what you would have needed in order to undo the change.
  2. Write one sentence describing what "a safe way to test a change on production data" would look like for that dataset. Keep it; you will rewrite it at the end of this guide.

The problem lakeFS solves, and what came before

Code has had good version control for decades. You branch, you commit, you review a diff, you merge, and if you broke something you revert. Data mostly has not had any of that. A folder in a bucket is a single mutable thing: when you write to it, the old contents are gone, everybody who reads it sees the change at once, and there is no record of what it looked like before.

People have worked around this in several ways, and it helps to know them because they explain what lakeFS is competing with. The simplest is copying: before a risky change, copy the folder to a new prefix such as data_backup_2026_09_30/. This works until the data is large. Copying several terabytes is slow and expensive, nobody remembers which copy is which, and the copies are never cleaned up. The second workaround is naming conventions: writing each run into data/run=2026-09-30/ and keeping a pointer file that says which run is current. This is better, but the convention lives in people's heads, there is no atomic way to switch a whole dataset from one run to the next, and different teams invent different conventions.

A third approach is the versioning feature of the bucket itself, such as S3 object versioning. It keeps old versions of single objects, which helps if you delete one file by accident. But it has no concept of a consistent set of files at one moment. If a dataset is a thousand Parquet files, bucket versioning cannot tell you "the thousand files as they were last Monday at 9 a.m.", and it offers no way to try a change in isolation and then publish it as one unit.

A fourth approach is to use a table format such as Delta Lake or Apache Iceberg, which keeps a transaction log and supports time travel. These are excellent, and lakeFS works alongside them, but they version one table at a time. Real pipelines often touch many tables, raw files, images, and model artifacts together. lakeFS versions the whole repository as one unit, whatever the files are, so a single commit can capture the raw input, the cleaned tables, and the trained model that came from them.

What makes lakeFS practical is that it does not copy data to do any of this. Creating a branch is a metadata-only, zero-copy operation that takes constant time, so a branch of a repository holding a hundred million objects is as cheap as a branch of one holding ten. The sizing guide in the lakeFS documentation puts branch operations under about 30 milliseconds at the 90th percentile. That is what turns "give every pipeline run its own branch" from a nice idea into something you can actually do.

Try it
  1. For a dataset you know, list which of the four workarounds (copying, naming conventions, bucket versioning, a table format) your team uses today.
  2. For each one, write down the thing it cannot do. Compare your list with the four benefits named in the previous section.

The mental model: repositories, branches, commits and tags

If you know Git, most of lakeFS will feel familiar, and the differences are worth noticing. If you do not know Git, the nouns below are all you need. There are only a handful.

An object is a file managed by lakeFS. The bytes live in the underlying object store, and lakeFS keeps a pointer and some metadata. A repository is a named collection of objects, together with all their branches, commits and tags. Repository names must start with a lowercase letter or a number, may contain only lowercase letters, numbers and hyphens, and must be between 3 and 63 characters long. So my-repo and sales-2026 are fine and My_Repo is not.

Every repository is given a storage namespace when you create it. This is the location in your object store where all of the repository's data lives, for example s3://my-bucket/my-repo. Committed data and not-yet-committed data are both written there, and lakeFS keeps its own metadata files under a _lakefs/ folder inside it. Two repositories may not share a namespace. The underlying location of a particular object is called its physical address, and lakeFS maps the logical paths you use to these physical addresses. Several logical paths, across different branches and commits, can point at one physical object, which is exactly why branching costs nothing: a new branch points at the same physical objects as the branch it came from. Objects that lakeFS has written are never modified in place.

A branch is a pointer to a commit, plus a set of uncommitted changes, also called staged changes. When you upload a file to a branch it becomes an uncommitted change on that branch and is visible to anyone reading that branch, but it is not yet part of history. A commit is an immutable snapshot of the entire repository at one moment, with a committer, a timestamp, a message, and optional key-value metadata. Each commit has a commit ID, which is a long hexadecimal digest. You can refer to a commit by a unique prefix of it. The default branch of a new repository is called main.

A tag is an immutable, human-friendly name for one commit, such as v1.0. Branches move as you commit; tags stay put. Use a tag when you want to say "this is the exact data the paper used".

A ref is the general name for any of those things: a branch, a tag, or a commit ID. lakeFS also understands Git-style suffixes on a ref. main~1 means the parent of the current head of main, and main~2 means the grandparent. You will use these when comparing "now" against "a bit ago".

Every object has a lakeFS URI with a fixed shape:

TEXT
lakefs://<repository>/<ref>/<path>

So lakefs://my-repo/main/data/x.csv is the file data/x.csv on the main branch of my-repo, and lakefs://my-repo/v1.0/data/x.csv is the same path as it was at tag v1.0. If you have seen older tutorials that write lakefs://repo@branch/path, that @ syntax was replaced by a slash long ago, and using it gives a "not found" error.

Finally, a merge combines the changes from one ref (the source) into a branch (the destination). lakeFS does a three-way merge against the common ancestor. The important beginner fact is that conflicts are detected per whole object: lakeFS never tries to merge the inside of two different versions of a CSV. If the same file was changed differently on both sides, or changed on one side and deleted on the other, that is a conflict and you have to choose.

Repository my-repo with storage namespace s3://my-bucket/my-repo
Branches main, feature: pointers to commits plus uncommitted changes
Commits and tags immutable snapshots with IDs; tags give them names
Objects your files, stored once, in your bucket
Try it
  1. Without looking back, write the lakeFS URI for the file reports/q3.csv on a branch called cleanup in a repository called finance.
  2. Explain in one sentence why creating a branch does not copy any data.
  3. Decide which name you would use, a branch or a tag, for "the dataset we shipped to the customer on 1 October", and why.

Editions and the 1.87 licence change

Before you install anything, you should know what you are installing, because 1.87 changed the ground rules and most tutorials were written before it.

lakeFS Community is the edition you can download and run yourself, and it is what this guide uses. It is free to use under the Business Source License 1.1 from version 1.87 onward. In plain words, the licence lets you use the unmodified software in production, but only for your own organisation's internal business purposes. You may not host it as a service for third parties, whether they pay or not. Changing behaviour through the documented configuration, flags, plug-ins and interfaces does not count as modifying it. The licence also says that each version converts to Apache 2.0 four years after that version was first published. If you are a student or a data engineer at a company using lakeFS for its own data, you are well inside those terms. If you are planning to build a product that offers lakeFS to customers, read the LICENSE file in the repository and ask someone who can give legal advice. Some pages of the official FAQ still say "Apache 2.0"; the licence file and the editions page are the current word.

The second change that matters to a beginner is about users. In 1.87, the Community edition runs with a single administrator user who has a single set of credentials: one access key ID and one secret access key. The earlier support for many users, groups, policies and a pluggable IAM service was removed. The web pages that managed users and groups are gone, and commands such as lakectl auth groups return a "not implemented" error against a Community server. Multiple users, fine-grained permissions, single sign-on and audit features belong to the commercial editions (Team, Enterprise and Cloud).

For a beginner, this has a practical meaning. Everyone and everything that talks to your Community server uses the same key pair, so treat that key pair like a password to the whole system, and do not paste it into notebooks you will share. If you are on a team that needs different people to have different permissions on different repositories, that is the moment to look at the commercial editions rather than trying to work around it.

Old tutorials may not match 1.87 Guides written before September 2026 often show creating extra users, groups and policies from the UI or with lakectl auth. On a Community 1.87 server that does not work, and it is not your mistake. Also do not pin version 1.74.1 or 1.48.0 in anything you write: 1.74.1 changed lakectl's endpoint default and was quickly followed by 1.74.2, and 1.48.0 shipped with squash-merge accidentally on by default.
Try it
  1. Open https://docs.lakefs.io/editions/ and read what each edition includes. Note one feature only the commercial editions have.
  2. Decide whether your intended use (learning, a personal project, internal company data) fits the Community terms as described above.

Installing lakeFS and checking the setup

There are several ways to get lakeFS, and which one you choose depends on what you already have. For learning, I recommend the route the current documentation recommends first: the Python quickstart. For everything else, Docker is the most portable, Homebrew is convenient on macOS and Linux, and plain binaries work everywhere.

Two programs come in every installation. lakefs is the server, the thing that keeps running and answers requests. lakectl is the command-line client that you type commands into, and it talks to a running server over HTTP. Beginners often mix these up, so keep the distinction in mind: lakefs starts lakeFS, and lakectl uses it.

Option one: the Python quickstart

If you have Python 3.10 or newer, this is the shortest path on Linux, macOS and Windows:

BASH
pip install lakefs
python -m lakefs.quickstart

The second command downloads a lakefs server and a lakectl client that match the installed SDK version, places them in your virtual environment's bin/ folder (or Scripts\ on Windows, or ~/.lakefs/bin if you are not in a virtual environment), and starts the server in quickstart mode. When it is ready you see a line like lakeFS running in quickstart mode. Login at http://127.0.0.1:8000/. The quickstart uses fixed credentials that are the same for everyone, which is why it must never be used for real data:

TEXT
Access Key ID:     AKIAIOSFOLQUICKSTART
Secret Access Key: wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY

Quickstart mode stores its metadata in a local embedded database and its data on your local disk, so it needs no cloud account, no bucket and no PostgreSQL. That is exactly what you want for learning.

Option two: Docker

If you already know Docker (see the Docker guide if you do not), you can run the server with one command. This is the equivalent of the quickstart:

BASH
docker run --name lakefs --pull always --rm -p 8000:8000 treeverse/lakefs:1.87.0 run --quickstart

The -p 8000:8000 publishes the server's port, and everything after the image name is passed to the lakefs program, so run --quickstart is the same as running lakefs run --quickstart on the host. If you want the data to survive a restart, run it with a local database, a local blockstore (the place objects are stored), a mounted folder, and a secret used to encrypt stored credentials:

BASH
docker run --name lakefs -p 8000:8000 \
  -v /path/on/host:/home/lakefs/lakefs \
  -e LAKEFS_DATABASE_TYPE=local \
  -e LAKEFS_BLOCKSTORE_TYPE=local \
  -e LAKEFS_AUTH_ENCRYPT_SECRET_KEY="change-me-to-a-long-random-string" \
  treeverse/lakefs:1.87.0 run

Notice the naming rule for environment variables: every configuration key becomes LAKEFS_ followed by the key in upper case, with dots replaced by underscores. So auth.encrypt.secret_key becomes LAKEFS_AUTH_ENCRYPT_SECRET_KEY. Inside a container, localhost means the container itself, so a client running on your laptop reaches the server at http://localhost:8000, but a second container would not.

Option three: Homebrew or a downloaded binary

On macOS and Linux, Homebrew installs both programs at once:

BASH
brew tap treeverse/lakefs
brew install lakefs
lakefs run --local-settings

--local-settings uses a local database and local storage, like the quickstart, but without the fixed credentials: the first time you open the web page, a setup wizard asks you to create your administrator. Alternatively, download lakeFS_1.87.0_<OS>_<arch>.tar.gz (or the .zip for Windows) from the GitHub releases page, check it against checksums.txt, unpack it, and put lakefs and lakectl on your PATH. Windows binaries exist, but the sizing guide says they are not rigorously tested and that you should not run lakeFS in production on Windows. For learning on Windows, use the Python quickstart or Docker with WSL2.

Checking that it works

Whichever way you started the server, open a second terminal and run these checks:

BASH
lakefs --version
lakectl --version
curl -i http://localhost:8000/_health

Both version commands should print 1.87.0, and the health endpoint should return HTTP 200. That _health URL is what load balancers use later in production. You can also open http://localhost:8000/ in a browser, where the web interface lives. The server banner in its terminal reads lakeFS 1.87.0 - Up and running (^C to shutdown)....

Quickstart mode is for learning only Quickstart mode refuses to run with anything but local settings, and fails with FATAL: quickstart mode can only run with local settings if you try to combine it with a real database or blockstore. The shared, public credentials shown above are known to everybody who has read the documentation. Never put data you care about into a quickstart server.
Try it
  1. Start lakeFS using one of the three options and open http://localhost:8000/.
  2. Run lakefs --version, lakectl --version and the curl health check, and confirm the output matches the description above.
  3. Stop the server with Ctrl+C, start it again, and notice whether the data survived. Quickstart with --rm in Docker does not keep it, and a local mounted folder does.

Connecting lakectl to your server

The server is running, but lakectl does not yet know where it is or who you are. It reads this from a small YAML file in your home directory. The simplest way to create that file is the interactive command:

BASH
lakectl config

It asks for three things: the access key ID, the secret access key, and the server endpoint URL. For the quickstart, type the two fixed credentials shown earlier and http://localhost:8000 as the endpoint. The command writes ~/.lakectl.yaml. You can have lakectl read a different file with --config (or -c), or with the environment variable LAKECTL_CONFIG_FILE. You can also skip the file entirely and use environment variables, which is handy in scripts and containers:

BASH
export LAKECTL_CREDENTIALS_ACCESS_KEY_ID="AKIAIOSFOLQUICKSTART"
export LAKECTL_CREDENTIALS_SECRET_ACCESS_KEY="wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY"
export LAKECTL_SERVER_ENDPOINT_URL="http://localhost:8000"

The configuration file itself looks like this:

~/.lakectl.yaml
credentials:
  access_key_id: AKIAIOSFOLQUICKSTART
  secret_access_key: wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY
server:
  endpoint_url: http://localhost:8000
Always set the endpoint explicitly Older lakectl versions assumed http://localhost:8000 if you did not say otherwise. Since 1.74.1 there is no default at all, even though the generated command reference still suggests one. If you forget it, requests fail with a message like unsupported protocol scheme "". The fix is always the same: run lakectl config, or set LAKECTL_SERVER_ENDPOINT_URL.

Now test the connection. These two commands ask the server who you are and then run a broader self-check of configuration, credentials and connectivity:

BASH
lakectl identity
lakectl doctor
lakectl repo list

On a fresh server repo list prints an empty table, and that is success: it means authentication worked and there is simply nothing there yet. If you see an authorization error, the key pair is wrong. If you see a connection error, the server is not running or the endpoint is wrong.

If you started the server with --local-settings or from a Docker image with a local database, rather than the quickstart, you create your own administrator the first time you open the web page. The setup wizard asks for a user name and some contact preferences, and it shows you the access key and secret once. Save them somewhere safe immediately, and notice that the wizard also offers a ready-made .lakectl.yaml. The same thing can be done without a browser using lakefs setup --user-name admin, which prints the new credentials.

Try it
  1. Run lakectl config and enter your values. Open ~/.lakectl.yaml and check that it matches the example.
  2. Run lakectl identity and lakectl doctor, then deliberately break the endpoint by one character and read the error so you recognise it next time.

Your first repository, step by step

The quickest way to get a repository with data in it is the sample repository that the web interface creates for you. Open http://localhost:8000/, log in with your credentials, and click Create Sample Repository. It creates a repository with example data and branches, which is perfect for clicking around. But the skill you need is doing it yourself from the command line, so we will build a small repository from scratch. It has a clear shape, and you will see every step.

Step one: create the repository

A repository needs a name and a storage namespace. For the local quickstart server, the namespace uses the local:// scheme. On a real deployment it would be an s3://, gs:// or Azure https:// address of a bucket you own. The scheme must match the blockstore the server was started with.

BASH
lakectl repo create lakefs://my-repo local://my-repo

If you are running against a real S3-backed server, the equivalent is:

BASH
lakectl repo create lakefs://my-repo s3://my-bucket/my-repo

The command prints a confirmation and the default branch name, main. Now check that it exists:

BASH
lakectl repo list

Step two: put a file in

Create a tiny dataset on your laptop and upload it to main. The -s flag says where the source is on your disk.

BASH
mkdir -p ~/lakefs-demo && cd ~/lakefs-demo
printf 'id,name,score\n1,Amal,71\n2,Bassem,64\n3,Carla,88\n' > scores.csv
lakectl fs upload lakefs://my-repo/main/data/scores.csv -s ./scores.csv

The object is now on main, but it is an uncommitted change. Ask lakeFS what has changed on the branch:

BASH
lakectl diff lakefs://my-repo/main

The output lists one added object, data/scores.csv. This is the same idea as git status. Nothing is in history yet. Take the snapshot with a commit, and always give it a message:

BASH
lakectl commit lakefs://my-repo/main -m "Add first scores file"

The command prints the new commit ID and message. An empty message is refused unless you add --allow-empty-message, which is a good default to keep.

Step three: look around

Three small commands let you inspect what is in the repository. Listing shows what objects exist, cat prints an object's contents, and stat shows its metadata such as size and checksum:

BASH
lakectl fs ls lakefs://my-repo/main/data/
lakectl fs cat lakefs://my-repo/main/data/scores.csv
lakectl fs stat lakefs://my-repo/main/data/scores.csv
lakectl log lakefs://my-repo/main

lakectl log shows the history, newest first. You should see one commit, yours. You have just made a repository, added data, and recorded a version.

Try it
  1. Create a repository called my-repo with the correct namespace for your setup, and upload scores.csv to main.
  2. Run lakectl diff lakefs://my-repo/main before committing and again after. Explain why the second run is empty.
  3. Find your commit ID with lakectl log and open the commit in the web interface under the repository's Commits view.

Branching: the feature that makes lakeFS worth using

Now comes the part that justifies the whole tool. Suppose you want to fix a data problem, say the score for Bassem was entered wrongly, and you want to try the fix without any risk to main. You create a branch:

BASH
lakectl branch create lakefs://my-repo/fix-scores -s lakefs://my-repo/main

The -s flag names the source: the ref the new branch starts from. The command returns instantly, however big the repository is, because it copies nothing. Confirm the branches:

BASH
lakectl branch list lakefs://my-repo
lakectl branch show lakefs://my-repo/fix-scores

Both branches now point at the same commit. Edit the file locally, then upload it to the branch, not to main:

BASH
printf 'id,name,score\n1,Amal,71\n2,Bassem,74\n3,Carla,88\n' > scores.csv
lakectl fs upload lakefs://my-repo/fix-scores/data/scores.csv -s ./scores.csv

Read the file from both places and notice they differ. This is the magic moment:

BASH
lakectl fs cat lakefs://my-repo/main/data/scores.csv
lakectl fs cat lakefs://my-repo/fix-scores/data/scores.csv

main still says 64 and fix-scores says 74. Anyone reading main, a dashboard or a training job, is completely unaffected. Now look at exactly what changed on the branch, commit it, and compare the branch with main:

BASH
lakectl diff lakefs://my-repo/fix-scores
lakectl commit lakefs://my-repo/fix-scores -m "Correct Bassem's score"
lakectl diff lakefs://my-repo/main lakefs://my-repo/fix-scores

The first diff shows the uncommitted change. The second, with two refs, compares two branches. By default lakeFS shows the changes the second ref introduced compared with the merge base. Adding --two-way compares the two heads directly, and --prefix data/ narrows it to a folder.

Merging, or throwing it away

If you are happy with the change, merge the branch into main. The source comes first and the destination second:

BASH
lakectl merge lakefs://my-repo/fix-scores lakefs://my-repo/main -m "Merge the score fix"

The merge is atomic: readers of main see either the old data or the new data, never a mixture. It creates a merge commit on main. You can check the history again with lakectl log lakefs://my-repo/main.

If you are not happy, you simply do not merge. Delete the branch and nothing else in the repository is affected:

BASH
lakectl branch delete lakefs://my-repo/fix-scores

The command asks you to confirm, and -y skips the question in scripts. After the merge you would delete the branch as well, because it has done its job. Deleting a branch deletes the pointer, not the commits that were already merged.

One branch per task The habit that pays off immediately is to do every change on a short-lived branch and to merge when a check passes. Because branches cost nothing, there is no reason to touch main directly. The section on protecting branches shows how to make the server enforce it.
Try it
  1. Create a branch called fix-scores, change the file on it, and prove with two fs cat commands that main is unchanged.
  2. Commit on the branch, look at lakectl diff between the two branches, and then merge.
  3. Create a second branch, change something, and delete the branch without merging. Confirm with lakectl log that main has no trace of it.

Everyday commands, grouped by what you want to do

By now you have met most of the everyday commands. This section groups them by intent so that you can find the right one quickly, and adds the ones you have not seen. All of them take lakeFS URIs, and all of them print help with --help.

Working with files

Uploading a single file uses -s. To upload a whole folder add -r (recursive), and to read from standard input use -s -:

BASH
lakectl fs upload -r lakefs://my-repo/feature/images/ -s ./images
echo "hello" | lakectl fs upload lakefs://my-repo/feature/notes/hello.txt -s -
lakectl fs ls -r lakefs://my-repo/feature/images/
lakectl fs download -r lakefs://my-repo/main/data/ ./out
lakectl fs rm lakefs://my-repo/feature/notes/hello.txt
lakectl fs rm -r lakefs://my-repo/feature/images/

Uploads and downloads go straight between your machine and the object store using pre-signed URLs, which are temporary addresses that lakeFS creates for you. The data does not flow through the lakeFS server, which is why it scales well. Deleting (fs rm) only removes an object from that branch's current state. Older commits still contain it, and that is a feature.

Seeing what changed

BASH
lakectl diff lakefs://my-repo/feature
lakectl diff lakefs://my-repo/main lakefs://my-repo/feature
lakectl log lakefs://my-repo/main --amount 5
lakectl log lakefs://my-repo/main --first-parent --no-merges
lakectl show commit lakefs://my-repo/<commit-id>

log accepts filters. --amount limits how many commits to print, --since takes an RFC 3339 timestamp such as 2026-09-01T00:00:00Z, and --prefixes data/,reports/ shows only commits that touched those folders. lakectl show commit prints the details of one commit, including the metadata you attached.

Naming a version

Commits accept key-value metadata that travels with the snapshot, which is useful for recording what produced it:

BASH
lakectl commit lakefs://my-repo/main -m "Nightly load" --meta job=nightly --meta source=crm
lakectl tag create lakefs://my-repo/v1.0 lakefs://my-repo/main
lakectl tag list lakefs://my-repo
lakectl fs cat lakefs://my-repo/v1.0/data/scores.csv

The last command reads the file through the tag, so you get the data exactly as it was when you tagged it, no matter how main has moved since. To replace an existing tag you need -f, and to delete one you use lakectl tag delete lakefs://my-repo/v1.0, with --yes to skip confirmation. Because tags are meant to be stable, treat moving one as a big decision.

Undoing things

There are two very different kinds of "undo", and confusing them is a classic beginner trap. Reset throws away uncommitted changes. Revert creates a new commit that undoes an earlier commit.

BASH
lakectl branch reset lakefs://my-repo/feature
lakectl branch reset lakefs://my-repo/feature --prefix data/
lakectl branch reset lakefs://my-repo/feature --object data/scores.csv
lakectl branch revert lakefs://my-repo/main <commit-id>

The first command discards every uncommitted change on the branch, the next two limit it to a folder or a single object, and all of them ask for confirmation unless you add -y. The revert command leaves history intact and adds a commit that cancels the target commit's changes. If the commit you want to revert is a merge commit, you must say which parent to keep, using -m 1; without it you get must specify 1-based parent number for reverting merge commit. Both undo operations need a clean branch, and if the branch has uncommitted changes you get an error about an uncommitted changes (dirty branch). Commit or reset first.

Reset is not reversible Uncommitted changes are not in any commit, so once you reset them there is nothing to bring back. If there is any doubt, commit first. A commit you later regret is easy to revert; an uncommitted change you reset is gone.
Try it
  1. On a branch, upload a whole folder with -r and then delete one file from it. Run lakectl diff and read the result.
  2. Commit with two --meta pairs, then use lakectl show commit to find them.
  3. Make a mistake on purpose, commit it on a branch, and undo it with branch revert. Then make another uncommitted mistake and undo it with branch reset. Say in one sentence how the two differ.

Using lakeFS from Python and from S3 tools

Most data work is done in code rather than in a terminal, and lakeFS offers two ways in that do not involve lakectl at all. The first is the Python SDK, and the second is the S3 gateway, which lets almost any tool that can talk to S3 talk to lakeFS unchanged.

The high-level Python SDK

Install the package, called lakefs (the version checked here is 0.16.0, which needs Python 3.10 or newer). It automatically reads the same ~/.lakectl.yaml or LAKECTL_* environment variables you already set up, so in many cases you need no extra configuration.

BASH
pip install lakefs
sdk_demo.py
import lakefs

repo = lakefs.repository("my-repo")
main = repo.branch("main")

# a private branch, created instantly
exp = repo.branch("add-dave").create(source_reference="main")

# write an object on the branch
exp.object("data/extra.csv").upload(data=b"id,name,score\n4,Dave,55\n")

# read it back
with exp.object("data/extra.csv").reader(mode="r") as f:
    print(f.read())

# list uncommitted changes, commit, merge
for change in exp.uncommitted():
    print(change)

exp.commit(message="Add Dave", metadata={"source": "manual"})
exp.merge_into(main)

repo.tag("v1.1").create(source_ref="main")

The shape is the same as in lakectl: you get a repository, pick a branch, create objects, commit, merge, tag. If you need to point at a server explicitly, for example in a script that runs somewhere without a config file, build a client:

PYTHON
import lakefs
from lakefs.client import Client

clt = Client(host="http://localhost:8000",
             username="AKIAIOSFOLQUICKSTART",
             password="wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY")
repo = lakefs.repository("my-repo", client=clt)

Errors come from lakefs.exceptions. The ones you will meet first are NotFoundException (the repository, branch or object does not exist), ConflictException (a merge conflict or a name that already exists), and ForbiddenException (your credentials are not allowed). Wrapping the merge in a try block that catches ConflictException is a good habit.

Transactions

One feature worth learning early is the transaction. It creates a hidden temporary branch, lets you make several changes, commits them, and merges them back atomically, cleaning up afterwards. If anything inside raises an error, everything is rolled back and main is untouched:

PYTHON
with repo.branch("main").transact(commit_message="Load daily files") as tx:
    tx.object("data/day1.csv").upload(data=b"a,b\n1,2\n")
    tx.object("data/day2.csv").upload(data=b"a,b\n3,4\n")

That single block is the safe-publishing pattern from the start of this guide in a dozen lines: readers of main see both files or neither.

The S3 gateway

The lakeFS server also exposes an S3-compatible endpoint on the same port. In this view the S3 "bucket" is the repository, and the start of the key is the ref (the branch or tag), so the path is s3://<repo>/<branch>/<path>. That means any tool that can be pointed at a custom S3 endpoint works with lakeFS: the AWS CLI, boto3, many query engines and Spark. For example, with the AWS CLI, using a profile that holds your lakeFS keys:

BASH
aws --profile lakefs --endpoint-url http://localhost:8000 s3 ls s3://my-repo/main/

And in boto3:

PYTHON
import boto3

s3 = boto3.client(
    "s3",
    endpoint_url="http://localhost:8000",
    aws_access_key_id="AKIAIOSFOLQUICKSTART",
    aws_secret_access_key="wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY",
    region_name="us-east-1",
)
print(s3.list_objects_v2(Bucket="my-repo", Prefix="main/data/"))

Two things to remember. Use path-style addressing (the bucket in the path, not in the host name), and use the region us-east-1, which is the gateway's default. If an S3 client reports SignatureDoesNotMatch, suspect the region, the addressing style, or a computer clock that has drifted, since the gateway rejects requests whose timestamps are too far from its own.

The gateway is the path of least resistance for existing tools. The native clients (lakectl and the Python SDK) are more efficient at scale, because they send only metadata through lakeFS and move the bytes directly to and from the object store. For a beginner the difference is invisible, but it is worth knowing that there are two data paths.

Try it
  1. Run the Python example above, changing the branch name, and confirm with lakectl log that the commit and merge show up.
  2. Write a transaction that raises an exception half-way through (for example raise RuntimeError("stop") after the first upload) and check that no file appeared on main.
  3. List the contents of s3://my-repo/main/ through the gateway with the AWS CLI or boto3.

Protecting branches and checking data with hooks

Once several people or pipelines write to the same repository, you want rules. lakeFS gives you two simple, beginner-friendly ones. The first is branch protection, which stops anybody from writing directly to an important branch. The second is hooks, which run automatic checks when something happens, such as a commit or a merge.

Branch protection

A protection rule is a pattern. Any branch that matches it rejects direct writes, deletes, commits and resets, so the only way to change it is to merge another branch into it:

BASH
lakectl branch-protect add lakefs://my-repo 'main'
lakectl branch-protect list lakefs://my-repo

After adding the rule, try to upload to main directly. You get cannot write to protected branch. That is the rule working as intended. The right response is the workflow you already know: branch, change, commit, merge. If you ever need to remove the rule, lakectl branch-protect delete lakefs://my-repo 'main' does it.

Hooks: automatic checks

A hook is a small piece of automation that lakeFS runs when an event happens. You describe hooks in a YAML action file stored in the repository under the folder _lakefs_actions/. Each action file names the events it reacts to, and the ordered list of hooks to run. A hook can be a webhook (lakeFS calls a URL you run), a Lua script (run inside lakeFS), or an Airflow DAG trigger. The events that matter first are pre-commit and pre-merge. A pre- hook that fails aborts the operation, and that is the mechanism behind "validate before you publish". A post- hook that fails does not undo anything.

Here is a pre-merge action that calls a webhook on your own validation service before anything is merged into main:

_lakefs_actions/validate.yaml
name: validate before merge
on:
  pre-merge:
    branches:
      - main
hooks:
  - id: quality_check
    type: webhook
    description: Ask the validation service to approve this merge
    properties:
      url: "http://validator.internal:8080/check"
      timeout: 1m

Because the action file lives in the repository, it is versioned like data. You can check a file's syntax before committing it:

BASH
lakectl actions validate ./_lakefs_actions/validate.yaml

Action files are otherwise only read when an event fires, so a typo in one might surprise you later, which is why validating first is a good habit. After a run you can see what happened:

BASH
lakectl actions runs list lakefs://my-repo
lakectl actions runs describe lakefs://my-repo <run-id>

If a merge is rejected, the error tells you which hook failed, and the hook's log is stored in the repository under _lakefs/actions/log/. Hooks that need a secret, such as an API token, can read it from a server environment variable with the {{ ENV.NAME }} syntax in the webhook's headers or query parameters. The server only exposes variables that start with the configured prefix, which defaults to LAKEFSACTION_. This is a topic for the next level, but it is good to know the pattern exists.

Together, branch protection and a pre-merge hook give you a real data quality gate: nobody writes to main, and nothing gets merged unless a check passed.

Try it
  1. Add protection to main and try to upload a file to it. Read the error, then do the change properly on a branch and merge it.
  2. Write the validate.yaml file above with an unreachable URL, commit it on a branch, and see what a merge into main reports. This shows you what a failing gate looks like.

Configuration: what you set and where

So far you have run a server that configured itself. The moment you use a real bucket and a real database, you will write a configuration. You do not need every key. A small handful does the job, and lakeFS ships sensible defaults for the rest.

Configuration can come from a YAML file, from environment variables, or both. If you do not pass --config (or -c), the server looks for config.yaml in the current folder, then $HOME/lakefs/config.yaml, then /etc/lakefs/config.yaml, and finally $HOME/.lakefs.yaml. Any key can also be an environment variable: LAKEFS_ plus the key in upper case with dots turned into underscores. A minimal production-style file has just three blocks, the database that holds mutable metadata, the secret key, and the blockstore:

config.yaml
database:
  type: postgres
  postgres:
    connection_string: "postgres://lakefs:secret@db.example.com:5432/lakefs?sslmode=require"
auth:
  encrypt:
    secret_key: "a-long-random-string-you-keep-forever"
blockstore:
  type: s3
  s3:
    region: me-south-1

Run it with lakefs --config config.yaml run. A few notes on what each block means.

The database is called the KV store, and it holds the things that change often: where each branch points, which files are staged but not committed, and the credentials. The options are PostgreSQL, DynamoDB, Azure CosmosDB, or local, an embedded database for a single machine. PostgreSQL is the usual choice, and version 11 or newer is required. The object store holds the data and the immutable commit metadata. These two together are all the state lakeFS has, which is what allows you to run several identical lakeFS servers behind a load balancer.

The secret key under auth.encrypt.secret_key is required, and it is used to encrypt stored credentials. If you lose it or change it, the stored credentials stop working. Keep it in a secret manager, pass it as LAKEFS_AUTH_ENCRYPT_SECRET_KEY, and never commit it to Git. The development default is used only by the local and quickstart modes.

The blockstore is the place your data lives. Its type is one of s3, gs (Google Cloud Storage), azure, or local for learning. For S3 the key setting is the region; if you are working in the Middle East, for example in AWS region me-south-1 (Bahrain) or me-central-1 (UAE), set it here to keep your data in the region your employer's data-residency rules require. lakeFS uses the AWS credential chain, so an instance role or a Kubernetes service account is better than putting keys in the file. For S3-compatible stores such as MinIO, you add endpoint and usually force_path_style: true.

There is one rule worth memorising because it causes a confusing error. A repository remembers the type of blockstore it was created on. If you create repositories on a local blockstore and later restart the server with blockstore.type: s3, the server reports Mismatched adapter detected and refuses to serve those repositories. Switch back rather than forcing it.

You can also deploy lakeFS on Kubernetes with the official Helm chart:

BASH
helm repo add lakefs https://charts.lakefs.io
helm install my-lakefs lakefs/lakefs -f conf-values.yaml

The chart's defaults are for development only: a local database and local storage. Real deployments supply the PostgreSQL connection string and the encryption key through the chart's secret values. Helm itself is covered in the Helm guide, and the cloud side of the setup is where a tool such as Terraform is usually used.

Do not run quickstart or a local blockstore for real data The local database and local storage are for one machine and for learning. They are not designed to be shared, backed up or scaled. Before you move anything important to lakeFS, use PostgreSQL (or DynamoDB) and a real bucket, and keep the encryption key somewhere you can recover it.
Try it
  1. Write a config.yaml that uses database.type: local and blockstore.type: local with a secret key, and start the server with lakefs --config config.yaml run.
  2. Set the secret key through LAKEFS_AUTH_ENCRYPT_SECRET_KEY instead and confirm the server starts without it in the file.
  3. Read the configuration reference page and find the default value of listen_address.

Common errors and how to read them

lakeFS errors are usually clear once you know the pattern: the HTTP status says the category, and the message says the specific cause. Here are the ones beginners meet most, with what each really means.

repository not found, branch not found, commit not found, tag not found (404). Nine times out of ten this is a typo, or the old URI syntax with an @. Check that the URI has the form lakefs://repo/branch/path, and list what actually exists with lakectl repo list and lakectl branch list.

failed to create repository: storage namespace already in use (400). Another repository, or the leftover _lakefs data of a repository you deleted, occupies that location. lakeFS deliberately does not delete your data when you delete a repository. Pick a new prefix.

failed to create repository: failed to access storage (400). lakeFS could not write to and read from the namespace. The cause is almost always permissions, the wrong region, or a typo in the bucket name. The identity lakeFS runs as needs permission to read, write and list on that bucket prefix, and the server log has a reason field with the details. Fix the bucket policy or IAM role, then retry.

An invalid_namespace message ending with must match: <blockstore type>. The scheme of your namespace does not match the server's blockstore, for example a gs:// address on an S3 server. Use s3://, gs://, an Azure https://...blob.core.windows.net/... address, or local:// to match.

cannot write to protected branch. A protection rule applies. Do the work on another branch and merge.

conflict found (409). Two branches changed or deleted the same object differently. Redo the change from the current main, or choose a merge strategy deliberately.

commit with no message without specifying the "--allow-empty-message" flag. You ran lakectl commit without -m. Add a message.

uncommitted changes (dirty branch). The operation needs a clean branch. Commit your work or reset it.

unsupported protocol scheme "" from lakectl. No endpoint is configured. Run lakectl config.

HTTP 410 Gone when reading an object. The object was physically removed by garbage collection while an old commit still pointed to it. That is expected behaviour after a cleanup, not a bug. It is covered at the next level.

A server that refuses to start after an upgrade, with lakeFS has no administrator of its own. This is new in 1.87. If you upgraded an installation that already holds repositories but does not have a built-in administrator of its own (for example because its users lived in an external authorization service), the server stops and tells you the fix. Stop the server and run lakefs superuser --user-name admin. Add --access-key-id and --secret-access-key with your existing values if you want to keep the keys your clients already use, or omit them to get new ones. Then start the server again. For a fresh installation you will not see this.

Try it
  1. Cause three errors on purpose: read a branch that does not exist, create a repository with an upper-case name, and commit with no message. Match each message to this list.
  2. Try to create a second repository on the same namespace as your first, and read the message.

Putting it all together: a safe nightly update

Here is one small project that uses everything above. The story: a table of exam scores in data/scores.csv is published on main and read by a dashboard. Each night you need to add new rows, check them, and publish them, and you want a guarantee that a bad file never reaches the dashboard. You will also keep a tagged copy of the monthly version.

First, set up the repository with protection and a first version. Assume the server is running and lakectl is configured:

BASH
lakectl repo create lakefs://exam-results local://exam-results
printf 'id,name,score\n1,Amal,71\n2,Bassem,64\n3,Carla,88\n' > scores.csv
lakectl fs upload lakefs://exam-results/main/data/scores.csv -s ./scores.csv
lakectl commit lakefs://exam-results/main -m "Initial scores" --meta job=setup
lakectl tag create lakefs://exam-results/v2026-09 lakefs://exam-results/main
lakectl branch-protect add lakefs://exam-results 'main'

The tag v2026-09 freezes the September state. Protection now means nobody can write to main directly. Next, the nightly job is a short Python script. It works on a branch named after the date, validates the new data, and merges only if the validation passes:

nightly.py
import csv
import io
import datetime
import lakefs
from lakefs.exceptions import ConflictException

repo = lakefs.repository("exam-results")
main = repo.branch("main")
name = "nightly-" + datetime.date.today().isoformat()
branch = repo.branch(name).create(source_reference="main")

# 1. read the current file, add tonight's rows
with main.object("data/scores.csv").reader(mode="r") as f:
    rows = list(csv.DictReader(f))
rows.append({"id": "4", "name": "Dave", "score": "55"})

# 2. validate before publishing
for row in rows:
    if not (0 <= int(row["score"]) <= 100):
        branch.delete()
        raise SystemExit("bad score found; branch discarded, main untouched")

# 3. write the new version to the branch
out = io.StringIO()
writer = csv.DictWriter(out, fieldnames=["id", "name", "score"])
writer.writeheader()
writer.writerows(rows)
branch.object("data/scores.csv").upload(data=out.getvalue())

# 4. commit with metadata, then merge atomically
branch.commit(message="Nightly load " + name, metadata={"job": "nightly"})
try:
    branch.merge_into(main)
except ConflictException:
    raise SystemExit("conflict: someone else changed scores.csv; investigate")
branch.delete()
print("published", name)

Run it with python nightly.py. Then look at the results with the commands you know:

BASH
lakectl log lakefs://exam-results/main --amount 3
lakectl fs cat lakefs://exam-results/main/data/scores.csv
lakectl fs cat lakefs://exam-results/v2026-09/data/scores.csv

The first shows the new merge commit. The second shows four rows on main. The third reads the tag and shows the original three rows, proving that the September version is intact and readable. If you change the script so the new score is 150, the validation fails, the branch is deleted, and main never changes: the dashboard never sees a bad file. That is the entire idea of lakeFS in one script. If a bad change ever does get merged, you can put things right with lakectl branch revert lakefs://exam-results/main <commit-id>, but because the merge was a merge commit you would add -m 1.

Try it
  1. Run the whole project from scratch on your own server, including the protection rule and the tag.
  2. Change the script's new row to have a score of 150 and confirm that main is unchanged afterwards.
  3. Return to the sentence you wrote in the first Try it. Rewrite it using what you now know, naming the branch, the check and the merge.

What you can now do, and what comes next

You can now explain why data needs version control, and you can name lakeFS's core nouns: repository, storage namespace, branch, commit, tag, ref, and merge. You can install and start a server in several ways and verify it. You can configure lakectl, create a repository, upload, inspect, commit, branch, diff, merge, tag, reset and revert, and you know the difference between the last two. You can do the same from Python, including transactions, and through the S3 gateway with standard tools. You can protect a branch, sketch a validation hook, read a configuration file, and read the common error messages. You also know what changed in 1.87: the licence is the Business Source License, and Community has a single administrator.

Mid-level takes these topics further. It covers importing existing data into lakeFS without copying it, how commits and merges work under the hood (ranges and the metarange), cherry-picking, richer hooks with Lua and Airflow, working on a local copy with lakectl local, Spark and Delta Lake integration, garbage collection and retention rules, and using lakeFS with pipelines in CI.

Senior covers running lakeFS for a team: the architecture and its failure modes, sizing and the key-value store, security and the trust model, upgrades and migrations, backup and disaster recovery, cost control, and when the commercial editions or another tool are the better choice.

Where to go next in this catalogue: lakeFS is often used alongside DVC, which takes a different approach to versioning data; with MLflow to tie a model run to a data commit; and with Delta Lake for table-level versioning inside a lakeFS repository.

Sources