This is part one of three. It takes you from never having seen lakeFS to using it confidently on a real dataset. By the end you can start a lakeFS server on your laptop, create a repository, put files into it, branch it, change data safely on the branch, compare the change with the original, merge it or throw it away, and tag the result so you can find it again in six months. You will also understand which parts of lakeFS are free, which are not, and what changed in the release this guide was checked against, lakeFS 1.87.
Each section ends with a Try it task. Do them as you go. Version control for data is an idea that feels abstract until you have watched a branch appear instantly next to a dataset that holds a few gigabytes, so please run the commands rather than only reading them.
What lakeFS is and what you can do with it
lakeFS puts a Git-like layer of version control on top of an object store. An object store is the kind of storage that holds files as "objects" inside "buckets": Amazon S3, Google Cloud Storage, Azure Blob Storage, MinIO, and other storage that speaks the S3 protocol. Your data stays in your bucket, in the format you already use. lakeFS does not convert it, does not hold a second copy of it, and does not care whether the files are Parquet, CSV, images, audio, model checkpoints, or Delta and Iceberg table files. What lakeFS adds is a set of pointers and metadata that say which version of which file belongs to which branch and which commit.
The best way to feel why this matters is to picture a data team on a normal Tuesday. A nightly job rewrites a folder of Parquet files that a dashboard and two machine learning models read from. One night the job has a bug, and the folder now holds half-written files and wrong numbers. Everything downstream is broken, nobody is sure what the folder looked like yesterday, and the only "backup" is a copy somebody made three weeks ago. With lakeFS the nightly job writes to a branch, a check runs against the branch, and only if the check passes is the branch merged into main. Readers who point at main never see the broken state. If something slips through anyway, you revert to the previous commit, and the folder is back exactly as it was.
In practice people use it for four things. First, isolated experiments: give every engineer or every pipeline run a private branch of the whole dataset, at almost no cost. Second, safe publishing: write new data on a branch, validate it, then merge it in one atomic step so readers see either the old version or the new one and never a half-finished mix. Third, reproducibility: a commit is an immutable snapshot, so a model trained on lakefs://my-repo/3f9a.../data/ can be retrained on exactly the same data a year later. Fourth, fast recovery: undoing a bad change is a metadata operation that takes moments, however large the data is.
Everything in this guide is checked against lakeFS v1.87.0, released on 22 September 2026. You will see the word "Community" a lot. It names the free edition you download and run yourself. The commercial editions are called Team, Enterprise and Cloud, and they add things such as multiple users, permissions and single sign-on. Section four explains the difference, because one change in 1.87 affects almost every beginner tutorial you will find on the internet.
- Think of one dataset you have overwritten by mistake, or been afraid to overwrite. Write down what you would have needed in order to undo the change.
- Write one sentence describing what "a safe way to test a change on production data" would look like for that dataset. Keep it; you will rewrite it at the end of this guide.
The problem lakeFS solves, and what came before
Code has had good version control for decades. You branch, you commit, you review a diff, you merge, and if you broke something you revert. Data mostly has not had any of that. A folder in a bucket is a single mutable thing: when you write to it, the old contents are gone, everybody who reads it sees the change at once, and there is no record of what it looked like before.
People have worked around this in several ways, and it helps to know them because they explain what lakeFS is competing with. The simplest is copying: before a risky change, copy the folder to a new prefix such as data_backup_2026_09_30/. This works until the data is large. Copying several terabytes is slow and expensive, nobody remembers which copy is which, and the copies are never cleaned up. The second workaround is naming conventions: writing each run into data/run=2026-09-30/ and keeping a pointer file that says which run is current. This is better, but the convention lives in people's heads, there is no atomic way to switch a whole dataset from one run to the next, and different teams invent different conventions.
A third approach is the versioning feature of the bucket itself, such as S3 object versioning. It keeps old versions of single objects, which helps if you delete one file by accident. But it has no concept of a consistent set of files at one moment. If a dataset is a thousand Parquet files, bucket versioning cannot tell you "the thousand files as they were last Monday at 9 a.m.", and it offers no way to try a change in isolation and then publish it as one unit.
A fourth approach is to use a table format such as Delta Lake or Apache Iceberg, which keeps a transaction log and supports time travel. These are excellent, and lakeFS works alongside them, but they version one table at a time. Real pipelines often touch many tables, raw files, images, and model artifacts together. lakeFS versions the whole repository as one unit, whatever the files are, so a single commit can capture the raw input, the cleaned tables, and the trained model that came from them.
What makes lakeFS practical is that it does not copy data to do any of this. Creating a branch is a metadata-only, zero-copy operation that takes constant time, so a branch of a repository holding a hundred million objects is as cheap as a branch of one holding ten. The sizing guide in the lakeFS documentation puts branch operations under about 30 milliseconds at the 90th percentile. That is what turns "give every pipeline run its own branch" from a nice idea into something you can actually do.
- For a dataset you know, list which of the four workarounds (copying, naming conventions, bucket versioning, a table format) your team uses today.
- For each one, write down the thing it cannot do. Compare your list with the four benefits named in the previous section.
The mental model: repositories, branches, commits and tags
If you know Git, most of lakeFS will feel familiar, and the differences are worth noticing. If you do not know Git, the nouns below are all you need. There are only a handful.
An object is a file managed by lakeFS. The bytes live in the underlying object store, and lakeFS keeps a pointer and some metadata. A repository is a named collection of objects, together with all their branches, commits and tags. Repository names must start with a lowercase letter or a number, may contain only lowercase letters, numbers and hyphens, and must be between 3 and 63 characters long. So my-repo and sales-2026 are fine and My_Repo is not.
Every repository is given a storage namespace when you create it. This is the location in your object store where all of the repository's data lives, for example s3://my-bucket/my-repo. Committed data and not-yet-committed data are both written there, and lakeFS keeps its own metadata files under a _lakefs/ folder inside it. Two repositories may not share a namespace. The underlying location of a particular object is called its physical address, and lakeFS maps the logical paths you use to these physical addresses. Several logical paths, across different branches and commits, can point at one physical object, which is exactly why branching costs nothing: a new branch points at the same physical objects as the branch it came from. Objects that lakeFS has written are never modified in place.
A branch is a pointer to a commit, plus a set of uncommitted changes, also called staged changes. When you upload a file to a branch it becomes an uncommitted change on that branch and is visible to anyone reading that branch, but it is not yet part of history. A commit is an immutable snapshot of the entire repository at one moment, with a committer, a timestamp, a message, and optional key-value metadata. Each commit has a commit ID, which is a long hexadecimal digest. You can refer to a commit by a unique prefix of it. The default branch of a new repository is called main.
A tag is an immutable, human-friendly name for one commit, such as v1.0. Branches move as you commit; tags stay put. Use a tag when you want to say "this is the exact data the paper used".
A ref is the general name for any of those things: a branch, a tag, or a commit ID. lakeFS also understands Git-style suffixes on a ref. main~1 means the parent of the current head of main, and main~2 means the grandparent. You will use these when comparing "now" against "a bit ago".
Every object has a lakeFS URI with a fixed shape:
lakefs://<repository>/<ref>/<path>
So lakefs://my-repo/main/data/x.csv is the file data/x.csv on the main branch of my-repo, and lakefs://my-repo/v1.0/data/x.csv is the same path as it was at tag v1.0. If you have seen older tutorials that write lakefs://repo@branch/path, that @ syntax was replaced by a slash long ago, and using it gives a "not found" error.
Finally, a merge combines the changes from one ref (the source) into a branch (the destination). lakeFS does a three-way merge against the common ancestor. The important beginner fact is that conflicts are detected per whole object: lakeFS never tries to merge the inside of two different versions of a CSV. If the same file was changed differently on both sides, or changed on one side and deleted on the other, that is a conflict and you have to choose.
my-repo with storage namespace s3://my-bucket/my-repomain, feature: pointers to commits plus uncommitted changes- Without looking back, write the lakeFS URI for the file
reports/q3.csvon a branch calledcleanupin a repository calledfinance. - Explain in one sentence why creating a branch does not copy any data.
- Decide which name you would use, a branch or a tag, for "the dataset we shipped to the customer on 1 October", and why.
Editions and the 1.87 licence change
Before you install anything, you should know what you are installing, because 1.87 changed the ground rules and most tutorials were written before it.
lakeFS Community is the edition you can download and run yourself, and it is what this guide uses. It is free to use under the Business Source License 1.1 from version 1.87 onward. In plain words, the licence lets you use the unmodified software in production, but only for your own organisation's internal business purposes. You may not host it as a service for third parties, whether they pay or not. Changing behaviour through the documented configuration, flags, plug-ins and interfaces does not count as modifying it. The licence also says that each version converts to Apache 2.0 four years after that version was first published. If you are a student or a data engineer at a company using lakeFS for its own data, you are well inside those terms. If you are planning to build a product that offers lakeFS to customers, read the LICENSE file in the repository and ask someone who can give legal advice. Some pages of the official FAQ still say "Apache 2.0"; the licence file and the editions page are the current word.
The second change that matters to a beginner is about users. In 1.87, the Community edition runs with a single administrator user who has a single set of credentials: one access key ID and one secret access key. The earlier support for many users, groups, policies and a pluggable IAM service was removed. The web pages that managed users and groups are gone, and commands such as lakectl auth groups return a "not implemented" error against a Community server. Multiple users, fine-grained permissions, single sign-on and audit features belong to the commercial editions (Team, Enterprise and Cloud).
For a beginner, this has a practical meaning. Everyone and everything that talks to your Community server uses the same key pair, so treat that key pair like a password to the whole system, and do not paste it into notebooks you will share. If you are on a team that needs different people to have different permissions on different repositories, that is the moment to look at the commercial editions rather than trying to work around it.
lakectl auth. On a Community 1.87 server that does not work, and it is not your mistake. Also do not pin version 1.74.1 or 1.48.0 in anything you write: 1.74.1 changed lakectl's endpoint default and was quickly followed by 1.74.2, and 1.48.0 shipped with squash-merge accidentally on by default.
- Open
https://docs.lakefs.io/editions/and read what each edition includes. Note one feature only the commercial editions have. - Decide whether your intended use (learning, a personal project, internal company data) fits the Community terms as described above.
Installing lakeFS and checking the setup
There are several ways to get lakeFS, and which one you choose depends on what you already have. For learning, I recommend the route the current documentation recommends first: the Python quickstart. For everything else, Docker is the most portable, Homebrew is convenient on macOS and Linux, and plain binaries work everywhere.
Two programs come in every installation. lakefs is the server, the thing that keeps running and answers requests. lakectl is the command-line client that you type commands into, and it talks to a running server over HTTP. Beginners often mix these up, so keep the distinction in mind: lakefs starts lakeFS, and lakectl uses it.
Option one: the Python quickstart
If you have Python 3.10 or newer, this is the shortest path on Linux, macOS and Windows:
pip install lakefs
python -m lakefs.quickstart
The second command downloads a lakefs server and a lakectl client that match the installed SDK version, places them in your virtual environment's bin/ folder (or Scripts\ on Windows, or ~/.lakefs/bin if you are not in a virtual environment), and starts the server in quickstart mode. When it is ready you see a line like lakeFS running in quickstart mode. Login at http://127.0.0.1:8000/. The quickstart uses fixed credentials that are the same for everyone, which is why it must never be used for real data:
Access Key ID: AKIAIOSFOLQUICKSTART
Secret Access Key: wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY
Quickstart mode stores its metadata in a local embedded database and its data on your local disk, so it needs no cloud account, no bucket and no PostgreSQL. That is exactly what you want for learning.
Option two: Docker
If you already know Docker (see the Docker guide if you do not), you can run the server with one command. This is the equivalent of the quickstart:
docker run --name lakefs --pull always --rm -p 8000:8000 treeverse/lakefs:1.87.0 run --quickstart
The -p 8000:8000 publishes the server's port, and everything after the image name is passed to the lakefs program, so run --quickstart is the same as running lakefs run --quickstart on the host. If you want the data to survive a restart, run it with a local database, a local blockstore (the place objects are stored), a mounted folder, and a secret used to encrypt stored credentials:
docker run --name lakefs -p 8000:8000 \
-v /path/on/host:/home/lakefs/lakefs \
-e LAKEFS_DATABASE_TYPE=local \
-e LAKEFS_BLOCKSTORE_TYPE=local \
-e LAKEFS_AUTH_ENCRYPT_SECRET_KEY="change-me-to-a-long-random-string" \
treeverse/lakefs:1.87.0 run
Notice the naming rule for environment variables: every configuration key becomes LAKEFS_ followed by the key in upper case, with dots replaced by underscores. So auth.encrypt.secret_key becomes LAKEFS_AUTH_ENCRYPT_SECRET_KEY. Inside a container, localhost means the container itself, so a client running on your laptop reaches the server at http://localhost:8000, but a second container would not.
Option three: Homebrew or a downloaded binary
On macOS and Linux, Homebrew installs both programs at once:
brew tap treeverse/lakefs
brew install lakefs
lakefs run --local-settings
--local-settings uses a local database and local storage, like the quickstart, but without the fixed credentials: the first time you open the web page, a setup wizard asks you to create your administrator. Alternatively, download lakeFS_1.87.0_<OS>_<arch>.tar.gz (or the .zip for Windows) from the GitHub releases page, check it against checksums.txt, unpack it, and put lakefs and lakectl on your PATH. Windows binaries exist, but the sizing guide says they are not rigorously tested and that you should not run lakeFS in production on Windows. For learning on Windows, use the Python quickstart or Docker with WSL2.
Checking that it works
Whichever way you started the server, open a second terminal and run these checks:
lakefs --version
lakectl --version
curl -i http://localhost:8000/_health
Both version commands should print 1.87.0, and the health endpoint should return HTTP 200. That _health URL is what load balancers use later in production. You can also open http://localhost:8000/ in a browser, where the web interface lives. The server banner in its terminal reads lakeFS 1.87.0 - Up and running (^C to shutdown)....
FATAL: quickstart mode can only run with local settings if you try to combine it with a real database or blockstore. The shared, public credentials shown above are known to everybody who has read the documentation. Never put data you care about into a quickstart server.
- Start lakeFS using one of the three options and open
http://localhost:8000/. - Run
lakefs --version,lakectl --versionand thecurlhealth check, and confirm the output matches the description above. - Stop the server with Ctrl+C, start it again, and notice whether the data survived. Quickstart with
--rmin Docker does not keep it, and a local mounted folder does.
Connecting lakectl to your server
The server is running, but lakectl does not yet know where it is or who you are. It reads this from a small YAML file in your home directory. The simplest way to create that file is the interactive command:
lakectl config
It asks for three things: the access key ID, the secret access key, and the server endpoint URL. For the quickstart, type the two fixed credentials shown earlier and http://localhost:8000 as the endpoint. The command writes ~/.lakectl.yaml. You can have lakectl read a different file with --config (or -c), or with the environment variable LAKECTL_CONFIG_FILE. You can also skip the file entirely and use environment variables, which is handy in scripts and containers:
export LAKECTL_CREDENTIALS_ACCESS_KEY_ID="AKIAIOSFOLQUICKSTART"
export LAKECTL_CREDENTIALS_SECRET_ACCESS_KEY="wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY"
export LAKECTL_SERVER_ENDPOINT_URL="http://localhost:8000"
The configuration file itself looks like this:
credentials:
access_key_id: AKIAIOSFOLQUICKSTART
secret_access_key: wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY
server:
endpoint_url: http://localhost:8000
http://localhost:8000 if you did not say otherwise. Since 1.74.1 there is no default at all, even though the generated command reference still suggests one. If you forget it, requests fail with a message like unsupported protocol scheme "". The fix is always the same: run lakectl config, or set LAKECTL_SERVER_ENDPOINT_URL.
Now test the connection. These two commands ask the server who you are and then run a broader self-check of configuration, credentials and connectivity:
lakectl identity
lakectl doctor
lakectl repo list
On a fresh server repo list prints an empty table, and that is success: it means authentication worked and there is simply nothing there yet. If you see an authorization error, the key pair is wrong. If you see a connection error, the server is not running or the endpoint is wrong.
If you started the server with --local-settings or from a Docker image with a local database, rather than the quickstart, you create your own administrator the first time you open the web page. The setup wizard asks for a user name and some contact preferences, and it shows you the access key and secret once. Save them somewhere safe immediately, and notice that the wizard also offers a ready-made .lakectl.yaml. The same thing can be done without a browser using lakefs setup --user-name admin, which prints the new credentials.
- Run
lakectl configand enter your values. Open~/.lakectl.yamland check that it matches the example. - Run
lakectl identityandlakectl doctor, then deliberately break the endpoint by one character and read the error so you recognise it next time.
Your first repository, step by step
The quickest way to get a repository with data in it is the sample repository that the web interface creates for you. Open http://localhost:8000/, log in with your credentials, and click Create Sample Repository. It creates a repository with example data and branches, which is perfect for clicking around. But the skill you need is doing it yourself from the command line, so we will build a small repository from scratch. It has a clear shape, and you will see every step.
Step one: create the repository
A repository needs a name and a storage namespace. For the local quickstart server, the namespace uses the local:// scheme. On a real deployment it would be an s3://, gs:// or Azure https:// address of a bucket you own. The scheme must match the blockstore the server was started with.
lakectl repo create lakefs://my-repo local://my-repo
If you are running against a real S3-backed server, the equivalent is:
lakectl repo create lakefs://my-repo s3://my-bucket/my-repo
The command prints a confirmation and the default branch name, main. Now check that it exists:
lakectl repo list
Step two: put a file in
Create a tiny dataset on your laptop and upload it to main. The -s flag says where the source is on your disk.
mkdir -p ~/lakefs-demo && cd ~/lakefs-demo
printf 'id,name,score\n1,Amal,71\n2,Bassem,64\n3,Carla,88\n' > scores.csv
lakectl fs upload lakefs://my-repo/main/data/scores.csv -s ./scores.csv
The object is now on main, but it is an uncommitted change. Ask lakeFS what has changed on the branch:
lakectl diff lakefs://my-repo/main
The output lists one added object, data/scores.csv. This is the same idea as git status. Nothing is in history yet. Take the snapshot with a commit, and always give it a message:
lakectl commit lakefs://my-repo/main -m "Add first scores file"
The command prints the new commit ID and message. An empty message is refused unless you add --allow-empty-message, which is a good default to keep.
Step three: look around
Three small commands let you inspect what is in the repository. Listing shows what objects exist, cat prints an object's contents, and stat shows its metadata such as size and checksum:
lakectl fs ls lakefs://my-repo/main/data/
lakectl fs cat lakefs://my-repo/main/data/scores.csv
lakectl fs stat lakefs://my-repo/main/data/scores.csv
lakectl log lakefs://my-repo/main
lakectl log shows the history, newest first. You should see one commit, yours. You have just made a repository, added data, and recorded a version.
- Create a repository called
my-repowith the correct namespace for your setup, and uploadscores.csvtomain. - Run
lakectl diff lakefs://my-repo/mainbefore committing and again after. Explain why the second run is empty. - Find your commit ID with
lakectl logand open the commit in the web interface under the repository's Commits view.
Branching: the feature that makes lakeFS worth using
Now comes the part that justifies the whole tool. Suppose you want to fix a data problem, say the score for Bassem was entered wrongly, and you want to try the fix without any risk to main. You create a branch:
lakectl branch create lakefs://my-repo/fix-scores -s lakefs://my-repo/main
The -s flag names the source: the ref the new branch starts from. The command returns instantly, however big the repository is, because it copies nothing. Confirm the branches:
lakectl branch list lakefs://my-repo
lakectl branch show lakefs://my-repo/fix-scores
Both branches now point at the same commit. Edit the file locally, then upload it to the branch, not to main:
printf 'id,name,score\n1,Amal,71\n2,Bassem,74\n3,Carla,88\n' > scores.csv
lakectl fs upload lakefs://my-repo/fix-scores/data/scores.csv -s ./scores.csv
Read the file from both places and notice they differ. This is the magic moment:
lakectl fs cat lakefs://my-repo/main/data/scores.csv
lakectl fs cat lakefs://my-repo/fix-scores/data/scores.csv
main still says 64 and fix-scores says 74. Anyone reading main, a dashboard or a training job, is completely unaffected. Now look at exactly what changed on the branch, commit it, and compare the branch with main:
lakectl diff lakefs://my-repo/fix-scores
lakectl commit lakefs://my-repo/fix-scores -m "Correct Bassem's score"
lakectl diff lakefs://my-repo/main lakefs://my-repo/fix-scores
The first diff shows the uncommitted change. The second, with two refs, compares two branches. By default lakeFS shows the changes the second ref introduced compared with the merge base. Adding --two-way compares the two heads directly, and --prefix data/ narrows it to a folder.
Merging, or throwing it away
If you are happy with the change, merge the branch into main. The source comes first and the destination second:
lakectl merge lakefs://my-repo/fix-scores lakefs://my-repo/main -m "Merge the score fix"
The merge is atomic: readers of main see either the old data or the new data, never a mixture. It creates a merge commit on main. You can check the history again with lakectl log lakefs://my-repo/main.
If you are not happy, you simply do not merge. Delete the branch and nothing else in the repository is affected:
lakectl branch delete lakefs://my-repo/fix-scores
The command asks you to confirm, and -y skips the question in scripts. After the merge you would delete the branch as well, because it has done its job. Deleting a branch deletes the pointer, not the commits that were already merged.
main directly. The section on protecting branches shows how to make the server enforce it.
- Create a branch called
fix-scores, change the file on it, and prove with twofs catcommands thatmainis unchanged. - Commit on the branch, look at
lakectl diffbetween the two branches, and then merge. - Create a second branch, change something, and delete the branch without merging. Confirm with
lakectl logthatmainhas no trace of it.
Everyday commands, grouped by what you want to do
By now you have met most of the everyday commands. This section groups them by intent so that you can find the right one quickly, and adds the ones you have not seen. All of them take lakeFS URIs, and all of them print help with --help.
Working with files
Uploading a single file uses -s. To upload a whole folder add -r (recursive), and to read from standard input use -s -:
lakectl fs upload -r lakefs://my-repo/feature/images/ -s ./images
echo "hello" | lakectl fs upload lakefs://my-repo/feature/notes/hello.txt -s -
lakectl fs ls -r lakefs://my-repo/feature/images/
lakectl fs download -r lakefs://my-repo/main/data/ ./out
lakectl fs rm lakefs://my-repo/feature/notes/hello.txt
lakectl fs rm -r lakefs://my-repo/feature/images/
Uploads and downloads go straight between your machine and the object store using pre-signed URLs, which are temporary addresses that lakeFS creates for you. The data does not flow through the lakeFS server, which is why it scales well. Deleting (fs rm) only removes an object from that branch's current state. Older commits still contain it, and that is a feature.
Seeing what changed
lakectl diff lakefs://my-repo/feature
lakectl diff lakefs://my-repo/main lakefs://my-repo/feature
lakectl log lakefs://my-repo/main --amount 5
lakectl log lakefs://my-repo/main --first-parent --no-merges
lakectl show commit lakefs://my-repo/<commit-id>
log accepts filters. --amount limits how many commits to print, --since takes an RFC 3339 timestamp such as 2026-09-01T00:00:00Z, and --prefixes data/,reports/ shows only commits that touched those folders. lakectl show commit prints the details of one commit, including the metadata you attached.
Naming a version
Commits accept key-value metadata that travels with the snapshot, which is useful for recording what produced it:
lakectl commit lakefs://my-repo/main -m "Nightly load" --meta job=nightly --meta source=crm
lakectl tag create lakefs://my-repo/v1.0 lakefs://my-repo/main
lakectl tag list lakefs://my-repo
lakectl fs cat lakefs://my-repo/v1.0/data/scores.csv
The last command reads the file through the tag, so you get the data exactly as it was when you tagged it, no matter how main has moved since. To replace an existing tag you need -f, and to delete one you use lakectl tag delete lakefs://my-repo/v1.0, with --yes to skip confirmation. Because tags are meant to be stable, treat moving one as a big decision.
Undoing things
There are two very different kinds of "undo", and confusing them is a classic beginner trap. Reset throws away uncommitted changes. Revert creates a new commit that undoes an earlier commit.
lakectl branch reset lakefs://my-repo/feature
lakectl branch reset lakefs://my-repo/feature --prefix data/
lakectl branch reset lakefs://my-repo/feature --object data/scores.csv
lakectl branch revert lakefs://my-repo/main <commit-id>
The first command discards every uncommitted change on the branch, the next two limit it to a folder or a single object, and all of them ask for confirmation unless you add -y. The revert command leaves history intact and adds a commit that cancels the target commit's changes. If the commit you want to revert is a merge commit, you must say which parent to keep, using -m 1; without it you get must specify 1-based parent number for reverting merge commit. Both undo operations need a clean branch, and if the branch has uncommitted changes you get an error about an uncommitted changes (dirty branch). Commit or reset first.
- On a branch, upload a whole folder with
-rand then delete one file from it. Runlakectl diffand read the result. - Commit with two
--metapairs, then uselakectl show committo find them. - Make a mistake on purpose, commit it on a branch, and undo it with
branch revert. Then make another uncommitted mistake and undo it withbranch reset. Say in one sentence how the two differ.
Using lakeFS from Python and from S3 tools
Most data work is done in code rather than in a terminal, and lakeFS offers two ways in that do not involve lakectl at all. The first is the Python SDK, and the second is the S3 gateway, which lets almost any tool that can talk to S3 talk to lakeFS unchanged.
The high-level Python SDK
Install the package, called lakefs (the version checked here is 0.16.0, which needs Python 3.10 or newer). It automatically reads the same ~/.lakectl.yaml or LAKECTL_* environment variables you already set up, so in many cases you need no extra configuration.
pip install lakefs
import lakefs
repo = lakefs.repository("my-repo")
main = repo.branch("main")
# a private branch, created instantly
exp = repo.branch("add-dave").create(source_reference="main")
# write an object on the branch
exp.object("data/extra.csv").upload(data=b"id,name,score\n4,Dave,55\n")
# read it back
with exp.object("data/extra.csv").reader(mode="r") as f:
print(f.read())
# list uncommitted changes, commit, merge
for change in exp.uncommitted():
print(change)
exp.commit(message="Add Dave", metadata={"source": "manual"})
exp.merge_into(main)
repo.tag("v1.1").create(source_ref="main")
The shape is the same as in lakectl: you get a repository, pick a branch, create objects, commit, merge, tag. If you need to point at a server explicitly, for example in a script that runs somewhere without a config file, build a client:
import lakefs
from lakefs.client import Client
clt = Client(host="http://localhost:8000",
username="AKIAIOSFOLQUICKSTART",
password="wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY")
repo = lakefs.repository("my-repo", client=clt)
Errors come from lakefs.exceptions. The ones you will meet first are NotFoundException (the repository, branch or object does not exist), ConflictException (a merge conflict or a name that already exists), and ForbiddenException (your credentials are not allowed). Wrapping the merge in a try block that catches ConflictException is a good habit.
Transactions
One feature worth learning early is the transaction. It creates a hidden temporary branch, lets you make several changes, commits them, and merges them back atomically, cleaning up afterwards. If anything inside raises an error, everything is rolled back and main is untouched:
with repo.branch("main").transact(commit_message="Load daily files") as tx:
tx.object("data/day1.csv").upload(data=b"a,b\n1,2\n")
tx.object("data/day2.csv").upload(data=b"a,b\n3,4\n")
That single block is the safe-publishing pattern from the start of this guide in a dozen lines: readers of main see both files or neither.
The S3 gateway
The lakeFS server also exposes an S3-compatible endpoint on the same port. In this view the S3 "bucket" is the repository, and the start of the key is the ref (the branch or tag), so the path is s3://<repo>/<branch>/<path>. That means any tool that can be pointed at a custom S3 endpoint works with lakeFS: the AWS CLI, boto3, many query engines and Spark. For example, with the AWS CLI, using a profile that holds your lakeFS keys:
aws --profile lakefs --endpoint-url http://localhost:8000 s3 ls s3://my-repo/main/
And in boto3:
import boto3
s3 = boto3.client(
"s3",
endpoint_url="http://localhost:8000",
aws_access_key_id="AKIAIOSFOLQUICKSTART",
aws_secret_access_key="wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY",
region_name="us-east-1",
)
print(s3.list_objects_v2(Bucket="my-repo", Prefix="main/data/"))
Two things to remember. Use path-style addressing (the bucket in the path, not in the host name), and use the region us-east-1, which is the gateway's default. If an S3 client reports SignatureDoesNotMatch, suspect the region, the addressing style, or a computer clock that has drifted, since the gateway rejects requests whose timestamps are too far from its own.
The gateway is the path of least resistance for existing tools. The native clients (lakectl and the Python SDK) are more efficient at scale, because they send only metadata through lakeFS and move the bytes directly to and from the object store. For a beginner the difference is invisible, but it is worth knowing that there are two data paths.
- Run the Python example above, changing the branch name, and confirm with
lakectl logthat the commit and merge show up. - Write a transaction that raises an exception half-way through (for example
raise RuntimeError("stop")after the first upload) and check that no file appeared onmain. - List the contents of
s3://my-repo/main/through the gateway with the AWS CLI or boto3.
Protecting branches and checking data with hooks
Once several people or pipelines write to the same repository, you want rules. lakeFS gives you two simple, beginner-friendly ones. The first is branch protection, which stops anybody from writing directly to an important branch. The second is hooks, which run automatic checks when something happens, such as a commit or a merge.
Branch protection
A protection rule is a pattern. Any branch that matches it rejects direct writes, deletes, commits and resets, so the only way to change it is to merge another branch into it:
lakectl branch-protect add lakefs://my-repo 'main'
lakectl branch-protect list lakefs://my-repo
After adding the rule, try to upload to main directly. You get cannot write to protected branch. That is the rule working as intended. The right response is the workflow you already know: branch, change, commit, merge. If you ever need to remove the rule, lakectl branch-protect delete lakefs://my-repo 'main' does it.
Hooks: automatic checks
A hook is a small piece of automation that lakeFS runs when an event happens. You describe hooks in a YAML action file stored in the repository under the folder _lakefs_actions/. Each action file names the events it reacts to, and the ordered list of hooks to run. A hook can be a webhook (lakeFS calls a URL you run), a Lua script (run inside lakeFS), or an Airflow DAG trigger. The events that matter first are pre-commit and pre-merge. A pre- hook that fails aborts the operation, and that is the mechanism behind "validate before you publish". A post- hook that fails does not undo anything.
Here is a pre-merge action that calls a webhook on your own validation service before anything is merged into main:
name: validate before merge
on:
pre-merge:
branches:
- main
hooks:
- id: quality_check
type: webhook
description: Ask the validation service to approve this merge
properties:
url: "http://validator.internal:8080/check"
timeout: 1m
Because the action file lives in the repository, it is versioned like data. You can check a file's syntax before committing it:
lakectl actions validate ./_lakefs_actions/validate.yaml
Action files are otherwise only read when an event fires, so a typo in one might surprise you later, which is why validating first is a good habit. After a run you can see what happened:
lakectl actions runs list lakefs://my-repo
lakectl actions runs describe lakefs://my-repo <run-id>
If a merge is rejected, the error tells you which hook failed, and the hook's log is stored in the repository under _lakefs/actions/log/. Hooks that need a secret, such as an API token, can read it from a server environment variable with the {{ ENV.NAME }} syntax in the webhook's headers or query parameters. The server only exposes variables that start with the configured prefix, which defaults to LAKEFSACTION_. This is a topic for the next level, but it is good to know the pattern exists.
Together, branch protection and a pre-merge hook give you a real data quality gate: nobody writes to main, and nothing gets merged unless a check passed.
- Add protection to
mainand try to upload a file to it. Read the error, then do the change properly on a branch and merge it. - Write the
validate.yamlfile above with an unreachable URL, commit it on a branch, and see what a merge intomainreports. This shows you what a failing gate looks like.
Configuration: what you set and where
So far you have run a server that configured itself. The moment you use a real bucket and a real database, you will write a configuration. You do not need every key. A small handful does the job, and lakeFS ships sensible defaults for the rest.
Configuration can come from a YAML file, from environment variables, or both. If you do not pass --config (or -c), the server looks for config.yaml in the current folder, then $HOME/lakefs/config.yaml, then /etc/lakefs/config.yaml, and finally $HOME/.lakefs.yaml. Any key can also be an environment variable: LAKEFS_ plus the key in upper case with dots turned into underscores. A minimal production-style file has just three blocks, the database that holds mutable metadata, the secret key, and the blockstore:
database:
type: postgres
postgres:
connection_string: "postgres://lakefs:secret@db.example.com:5432/lakefs?sslmode=require"
auth:
encrypt:
secret_key: "a-long-random-string-you-keep-forever"
blockstore:
type: s3
s3:
region: me-south-1
Run it with lakefs --config config.yaml run. A few notes on what each block means.
The database is called the KV store, and it holds the things that change often: where each branch points, which files are staged but not committed, and the credentials. The options are PostgreSQL, DynamoDB, Azure CosmosDB, or local, an embedded database for a single machine. PostgreSQL is the usual choice, and version 11 or newer is required. The object store holds the data and the immutable commit metadata. These two together are all the state lakeFS has, which is what allows you to run several identical lakeFS servers behind a load balancer.
The secret key under auth.encrypt.secret_key is required, and it is used to encrypt stored credentials. If you lose it or change it, the stored credentials stop working. Keep it in a secret manager, pass it as LAKEFS_AUTH_ENCRYPT_SECRET_KEY, and never commit it to Git. The development default is used only by the local and quickstart modes.
The blockstore is the place your data lives. Its type is one of s3, gs (Google Cloud Storage), azure, or local for learning. For S3 the key setting is the region; if you are working in the Middle East, for example in AWS region me-south-1 (Bahrain) or me-central-1 (UAE), set it here to keep your data in the region your employer's data-residency rules require. lakeFS uses the AWS credential chain, so an instance role or a Kubernetes service account is better than putting keys in the file. For S3-compatible stores such as MinIO, you add endpoint and usually force_path_style: true.
There is one rule worth memorising because it causes a confusing error. A repository remembers the type of blockstore it was created on. If you create repositories on a local blockstore and later restart the server with blockstore.type: s3, the server reports Mismatched adapter detected and refuses to serve those repositories. Switch back rather than forcing it.
You can also deploy lakeFS on Kubernetes with the official Helm chart:
helm repo add lakefs https://charts.lakefs.io
helm install my-lakefs lakefs/lakefs -f conf-values.yaml
The chart's defaults are for development only: a local database and local storage. Real deployments supply the PostgreSQL connection string and the encryption key through the chart's secret values. Helm itself is covered in the Helm guide, and the cloud side of the setup is where a tool such as Terraform is usually used.
- Write a
config.yamlthat usesdatabase.type: localandblockstore.type: localwith a secret key, and start the server withlakefs --config config.yaml run. - Set the secret key through
LAKEFS_AUTH_ENCRYPT_SECRET_KEYinstead and confirm the server starts without it in the file. - Read the configuration reference page and find the default value of
listen_address.
Common errors and how to read them
lakeFS errors are usually clear once you know the pattern: the HTTP status says the category, and the message says the specific cause. Here are the ones beginners meet most, with what each really means.
repository not found, branch not found, commit not found, tag not found (404). Nine times out of ten this is a typo, or the old URI syntax with an @. Check that the URI has the form lakefs://repo/branch/path, and list what actually exists with lakectl repo list and lakectl branch list.
failed to create repository: storage namespace already in use (400). Another repository, or the leftover _lakefs data of a repository you deleted, occupies that location. lakeFS deliberately does not delete your data when you delete a repository. Pick a new prefix.
failed to create repository: failed to access storage (400). lakeFS could not write to and read from the namespace. The cause is almost always permissions, the wrong region, or a typo in the bucket name. The identity lakeFS runs as needs permission to read, write and list on that bucket prefix, and the server log has a reason field with the details. Fix the bucket policy or IAM role, then retry.
An invalid_namespace message ending with must match: <blockstore type>. The scheme of your namespace does not match the server's blockstore, for example a gs:// address on an S3 server. Use s3://, gs://, an Azure https://...blob.core.windows.net/... address, or local:// to match.
cannot write to protected branch. A protection rule applies. Do the work on another branch and merge.
conflict found (409). Two branches changed or deleted the same object differently. Redo the change from the current main, or choose a merge strategy deliberately.
commit with no message without specifying the "--allow-empty-message" flag. You ran lakectl commit without -m. Add a message.
uncommitted changes (dirty branch). The operation needs a clean branch. Commit your work or reset it.
unsupported protocol scheme "" from lakectl. No endpoint is configured. Run lakectl config.
HTTP 410 Gone when reading an object. The object was physically removed by garbage collection while an old commit still pointed to it. That is expected behaviour after a cleanup, not a bug. It is covered at the next level.
A server that refuses to start after an upgrade, with lakeFS has no administrator of its own. This is new in 1.87. If you upgraded an installation that already holds repositories but does not have a built-in administrator of its own (for example because its users lived in an external authorization service), the server stops and tells you the fix. Stop the server and run lakefs superuser --user-name admin. Add --access-key-id and --secret-access-key with your existing values if you want to keep the keys your clients already use, or omit them to get new ones. Then start the server again. For a fresh installation you will not see this.
- Cause three errors on purpose: read a branch that does not exist, create a repository with an upper-case name, and commit with no message. Match each message to this list.
- Try to create a second repository on the same namespace as your first, and read the message.
Putting it all together: a safe nightly update
Here is one small project that uses everything above. The story: a table of exam scores in data/scores.csv is published on main and read by a dashboard. Each night you need to add new rows, check them, and publish them, and you want a guarantee that a bad file never reaches the dashboard. You will also keep a tagged copy of the monthly version.
First, set up the repository with protection and a first version. Assume the server is running and lakectl is configured:
lakectl repo create lakefs://exam-results local://exam-results
printf 'id,name,score\n1,Amal,71\n2,Bassem,64\n3,Carla,88\n' > scores.csv
lakectl fs upload lakefs://exam-results/main/data/scores.csv -s ./scores.csv
lakectl commit lakefs://exam-results/main -m "Initial scores" --meta job=setup
lakectl tag create lakefs://exam-results/v2026-09 lakefs://exam-results/main
lakectl branch-protect add lakefs://exam-results 'main'
The tag v2026-09 freezes the September state. Protection now means nobody can write to main directly. Next, the nightly job is a short Python script. It works on a branch named after the date, validates the new data, and merges only if the validation passes:
import csv
import io
import datetime
import lakefs
from lakefs.exceptions import ConflictException
repo = lakefs.repository("exam-results")
main = repo.branch("main")
name = "nightly-" + datetime.date.today().isoformat()
branch = repo.branch(name).create(source_reference="main")
# 1. read the current file, add tonight's rows
with main.object("data/scores.csv").reader(mode="r") as f:
rows = list(csv.DictReader(f))
rows.append({"id": "4", "name": "Dave", "score": "55"})
# 2. validate before publishing
for row in rows:
if not (0 <= int(row["score"]) <= 100):
branch.delete()
raise SystemExit("bad score found; branch discarded, main untouched")
# 3. write the new version to the branch
out = io.StringIO()
writer = csv.DictWriter(out, fieldnames=["id", "name", "score"])
writer.writeheader()
writer.writerows(rows)
branch.object("data/scores.csv").upload(data=out.getvalue())
# 4. commit with metadata, then merge atomically
branch.commit(message="Nightly load " + name, metadata={"job": "nightly"})
try:
branch.merge_into(main)
except ConflictException:
raise SystemExit("conflict: someone else changed scores.csv; investigate")
branch.delete()
print("published", name)
Run it with python nightly.py. Then look at the results with the commands you know:
lakectl log lakefs://exam-results/main --amount 3
lakectl fs cat lakefs://exam-results/main/data/scores.csv
lakectl fs cat lakefs://exam-results/v2026-09/data/scores.csv
The first shows the new merge commit. The second shows four rows on main. The third reads the tag and shows the original three rows, proving that the September version is intact and readable. If you change the script so the new score is 150, the validation fails, the branch is deleted, and main never changes: the dashboard never sees a bad file. That is the entire idea of lakeFS in one script. If a bad change ever does get merged, you can put things right with lakectl branch revert lakefs://exam-results/main <commit-id>, but because the merge was a merge commit you would add -m 1.
- Run the whole project from scratch on your own server, including the protection rule and the tag.
- Change the script's new row to have a score of 150 and confirm that
mainis unchanged afterwards. - Return to the sentence you wrote in the first Try it. Rewrite it using what you now know, naming the branch, the check and the merge.
What you can now do, and what comes next
You can now explain why data needs version control, and you can name lakeFS's core nouns: repository, storage namespace, branch, commit, tag, ref, and merge. You can install and start a server in several ways and verify it. You can configure lakectl, create a repository, upload, inspect, commit, branch, diff, merge, tag, reset and revert, and you know the difference between the last two. You can do the same from Python, including transactions, and through the S3 gateway with standard tools. You can protect a branch, sketch a validation hook, read a configuration file, and read the common error messages. You also know what changed in 1.87: the licence is the Business Source License, and Community has a single administrator.
Mid-level takes these topics further. It covers importing existing data into lakeFS without copying it, how commits and merges work under the hood (ranges and the metarange), cherry-picking, richer hooks with Lua and Airflow, working on a local copy with lakectl local, Spark and Delta Lake integration, garbage collection and retention rules, and using lakeFS with pipelines in CI.
Senior covers running lakeFS for a team: the architecture and its failure modes, sizing and the key-value store, security and the trust model, upgrades and migrations, backup and disaster recovery, cost control, and when the commercial editions or another tool are the better choice.
Where to go next in this catalogue: lakeFS is often used alongside DVC, which takes a different approach to versioning data; with MLflow to tie a model run to a data commit; and with Delta Lake for table-level versioning inside a lakeFS repository.
Sources
- lakeFS documentation home: https://docs.lakefs.io/
- Editions: https://docs.lakefs.io/editions/
- Quickstart: https://docs.lakefs.io/quickstart/launch/
- Architecture concepts: https://docs.lakefs.io/concepts/architecture/
- Glossary: https://docs.lakefs.io/resources/glossary/
- Installation: https://docs.lakefs.io/admin/install/
- Sizing guide: https://docs.lakefs.io/admin/sizing-guide/
- Protecting branches: https://docs.lakefs.io/guides/protect-branches/
- Hooks: https://docs.lakefs.io/guides/hooks/
- lakectl command reference: https://docs.lakefs.io/reference/cli/oss/
- Configuration reference: https://docs.lakefs.io/reference/configuration/
- S3 gateway reference: https://docs.lakefs.io/reference/s3/
- Python SDK: https://docs.lakefs.io/reference/python/getting-started/
- Release notes for v1.87.0: https://github.com/treeverse/lakeFS/releases/tag/v1.87.0
- Licence file (Business Source License 1.1): https://github.com/treeverse/lakeFS/blob/v1.87.0/LICENSE
- Helm chart: https://github.com/treeverse/charts/tree/master/charts/lakefs