تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
DatabricksMLOpsCloud ML platforms3 مستويات99 قسمًايغطّي Databricks Runtime 18 LTS (Spark 4.1)دليل بالإنجليزية

The Complete Databricks Guide

Run data engineering, ML and GenAI on the Databricks lakehouse platform. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
17sections
23examples

This is part one of three. It covers everything you need to do real work on Databricks, starting from a blank browser tab. By the end you can open a workspace, run code in a notebook, create governed tables with the three-level naming that Unity Catalog uses, load a file, travel back in time on a table, schedule a job, track a first machine learning run, and talk to the platform from your terminal with the command line interface. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. Databricks is a platform you learn by clicking and running, and the ideas stick only after you have seen your own table appear in the catalog and your own query come back with rows.

The guide is written against Databricks Runtime 18 LTS, which runs Apache Spark 4.1, and the Databricks command line interface in its 1.x line. The platform itself is updated continuously, so button labels in the web interface move around from month to month. Where a label matters I describe what the control does as well as what it is called, so you can still find it after a redesign.

What Databricks is, and what you can do by the end

Databricks is a managed platform for working with data at scale. You log in through a browser, you get a place to write code, a way to run that code on computers somebody else operates, a catalog where your tables live, and a scheduler that reruns your work every night. On top of that sit the machine learning and generative AI features: experiment tracking, model registry, model serving, vector search and more. You do not install a server, you do not patch a database, and you do not size a cluster before you can run your first line.

The platform is built on a handful of open technologies. Apache Spark does the distributed processing. Delta Lake is the table format that gives files on cheap cloud storage the behaviour of database tables, including transactions and history. MLflow tracks experiments and models. Unity Catalog governs who can see what. Databricks did not invent most of these pieces; it integrates them and runs them for you.

YOUR CODEnotebook, SQL, job
→
COMPUTEserverless or cluster
→
UNITY CATALOGnames and permissions
→
DELTA TABLESfiles in cloud storage

That diagram is the whole platform at the level a beginner needs. Code runs on compute, compute reads and writes tables, and the catalog decides which names exist and who may touch them. Every later section fills in one box.

What people actually use it for falls into a few groups.

🏗️

Data engineering

Ingest files and streams, clean them in stages, and publish tables other people can trust.

📊

Analytics

Run SQL against the same tables from dashboards, notebooks and BI tools without copying data around.

🤖

Machine learning

Train, track, register and serve models with experiment history kept next to the data.

✨

GenAI applications

Build retrieval and agent applications with vector search, model serving and a gateway to model providers.

This guide stays on the foundations that every one of those groups shares: notebooks, compute, Unity Catalog, Delta tables, files, jobs, a first MLflow run and the command line. A reader who has these can pick up any of the four groups afterwards.

You need very little to start: a web browser, an email address, and a willingness to type small snippets of Python and SQL. Python knowledge at the level of variables, functions and loops is plenty. SQL at the level of SELECT, WHERE and GROUP BY is plenty too. If you have never written either, spend a weekend on the basics first, because Databricks assumes you can already read a query.

Try it
  1. Think of a dataset you have worked with, even a spreadsheet.
  2. Write down who is allowed to see it, how often it changes, and where the single trusted copy lives.
  3. Note every place a second, drifting copy exists.
a list with at least one stale copy and at least one unanswered question about permissions. Those two problems, scattered copies and unclear access, are what the platform is organised around.

The problem it solves and what came before

To see why Databricks looks the way it does, look at what teams did before it. For a long time there were two separate worlds. The first was the data warehouse: a database engine built for analytical queries over clean, structured tables. Warehouses were fast and well governed, but loading them was slow and expensive, they handled raw files, images and text poorly, and machine learning code could not run inside them. The second was the data lake: a pile of files in cheap object storage such as Amazon S3 or Azure storage. Lakes held anything and cost little, but they had no transactions, no schema enforcement, no easy permissions and no history. Two jobs writing to the same folder could leave it half updated, and nobody could tell which version of a file a model had been trained on.

Most organisations ran both. Raw data landed in the lake, a pipeline copied a cleaned subset into the warehouse, and data scientists pulled a third copy onto a laptop or a separate machine learning cluster. Every copy was a place for numbers to disagree and for access rules to be forgotten.

The lakehouse idea is to keep the data in the lake, in open file formats, and add the missing warehouse features on top. Delta Lake adds a transaction log to Parquet files, which gives you atomic writes, schema checks and a history of every change. Unity Catalog adds names, permissions and lineage. SQL engines and Spark both read the same tables. One copy of the data serves reporting, engineering and machine learning.

Databricks is not a single program you install There is no pip install databricks that gives you the platform. Databricks is a service you sign up for, running on Amazon Web Services, Microsoft Azure or Google Cloud. What you install on your own machine are optional tools: the command line interface, a Python software development kit, and a connector that lets your editor talk to remote compute. We cover the command line interface later in this guide.

The practical consequence is that "which version of Databricks" is a slightly odd question. The thing with a version number is the Databricks Runtime, the bundle of Spark, Delta Lake and libraries that your compute runs. At the time of writing the newest long-term-support release is Runtime 18 LTS, based on Spark 4.1, with support to June 2029. Runtime 17.3 LTS is still supported too. Serverless compute is different: it is called versionless, because Databricks upgrades the runtime for everyone at once.

Try it
  1. Open the supported runtimes page.
  2. Find the newest entry marked LTS and write down its Spark version and its end-of-support date.
  3. Find one older runtime that has already reached end of support.
a table with release dates and support dates. LTS releases get three years of support, which is why teams pin to them.

The mental model: account, workspace, compute and catalog

Databricks has a lot of product names, but the beginner's mental model needs only a few nouns. Learn these five and every page of documentation becomes easier to place.

The account is the top-level entity that your organisation signs up for. It holds the people, groups, billing, and the list of workspaces. Account administrators work in a separate web page called the account console. Most beginners never need it, but you will hear the word.

A workspace is the environment you log into. It has its own web address that looks like https://<something>.cloud.databricks.com on Amazon Web Services, and inside it live notebooks, jobs, compute and queries. An account usually has several workspaces, for example one for development and one for production, so that experiments cannot break a nightly pipeline.

The control plane and the compute plane explain where things physically run. The control plane is operated by Databricks. It serves the web interface, stores your notebooks, and runs the job scheduler. The compute plane is where your data is actually processed. With classic compute the machines run inside your organisation's own cloud account. With serverless compute they run inside Databricks' account and you never see them. Your data stays in cloud object storage either way.

Compute is the set of machines that run your code. A notebook has no power of its own. It is a document, and it needs to be attached to compute before any cell can run. Section six covers the options.

Unity Catalog is the governance layer. It decides what tables, files and models exist, what they are called, and who may use them. Every object has a three-level name: catalog, then schema, then the object itself. You will type names like main.default.my_table for the rest of your Databricks career.

Control plane

  • Run by Databricks
  • The web interface, APIs and job scheduler
  • Stores notebook source and settings
  • You do not manage it

Compute plane

  • Where Spark actually runs your code
  • Classic: inside your own cloud account
  • Serverless: inside Databricks' account
  • Reads and writes your cloud storage

If you are working in the Middle East or North Africa, one more word matters early: region. A workspace is created in one cloud region, and the metastore that holds Unity Catalog metadata is tied to a region too. Where your data physically lives can be a legal question for Gulf and Egyptian employers. Cloud providers offer regions in the area, but which Databricks features are available in each region varies, so check the current cloud and region documentation for your provider before you promise a client that data never leaves the country.

Try it
  1. Draw five boxes on paper: account, workspace, compute, catalog, table.
  2. Draw arrows for "contains" and "reads from".
  3. Label which box the web address in your browser belongs to.
a workspace inside an account, compute reading tables, and the catalog naming those tables. If you drew the address on the workspace box, you have it.

Getting a workspace

The platform needs no installation, but you do need somewhere to log in. For learning, Databricks offers a Free Edition that you can sign up for without a company account. The sign-up page is the authority on its current limits, which change, so read what it says about compute allowances and supported features before you plan a project around it. Another route is a trial through a cloud marketplace, and a third is the workspace your employer or course provides.

The steps are the same everywhere in spirit.

  1. Go to the Databricks sign-up page and choose the free option for learning.
  2. Sign in with an email address or an existing Google or Microsoft account and follow the prompts.
  3. Wait for the workspace to be created. When it is ready you land on a home page inside your own workspace.
  4. Copy the address from your browser bar and save it somewhere. This is your workspace URL, and the command line interface will ask for it later.

Unity Catalog is turned on automatically for workspaces created after 8 November 2023, so a new workspace starts with a catalog called main and, inside it, a schema called default. You may also see a catalog called samples containing small ready-made datasets that Databricks provides for tutorials. We will use one of them shortly.

Use a throwaway mindset for the free tier A free learning workspace is for learning. Do not load real customer data into it, and do not build the only copy of something you care about there. Treat anything you create as something you might need to recreate, and keep your code in Git as soon as you have code worth keeping.

Once you are in, spend a few minutes just looking at the left sidebar. It lists the main areas: a workspace area for your files and notebooks, a catalog browser, compute, jobs and workflows, and a place for SQL. The exact names and order change as the product evolves. If something is not where this guide says, use the search box at the top of the page, which finds notebooks, tables and settings by name.

Try it
  1. Sign up for a workspace and save its URL.
  2. Open the catalog browser from the sidebar.
  3. Expand main and samples, and look for a schema containing taxi trips.
a tree of catalogs, schemas and tables. You have just met the three-level hierarchy before reading its definition.

Your first notebook

A notebook is the main place you write code on Databricks. It is a document made of cells. Each cell holds a few lines of code or some text, and you run cells one at a time. The output appears directly below the cell, whether it is a number, a table of rows or a chart. Notebooks suit data work because you rarely know the answer in advance. You look at the data, try a transformation, look again, and keep the conversation with the data on the page.

A notebook has a default language, usually Python, but each cell can switch language with a magic command on its first line. %sql makes the cell SQL, %python makes it Python, %md makes it formatted text. This lets one notebook mix a SQL query with a Python function, which is common in practice.

Create one now. In the workspace area choose to create a new notebook, give it a name such as first-steps, and attach it to compute using the selector at the top. On a new workspace the selector offers serverless compute, which needs no setup. Then type this into the first cell and run it with the run button or Shift+Enter.

PYTHON
print("hello from Databricks")
spark.version

The first line prints text. The second asks the Spark session, a ready-made object named spark that every notebook has, which version of Spark it is connected to. You never create spark yourself in a notebook. Databricks creates it and hands it to you, which is one of the small conveniences that makes the first hour pleasant.

Now read real data. The samples catalog contains a taxi trips table you can query without uploading anything.

PYTHON
df = spark.read.table("samples.nyctaxi.trips")
df.display()

spark.read.table loads a table by its three-level name into a DataFrame, which is Spark's name for a table of rows and named columns held in a form Spark can process in parallel. display() is a Databricks addition that renders a DataFrame as an interactive table with a button for turning it into a chart. Plain Spark has show(), which prints text. You will see both in examples online, and display() only works inside Databricks.

A detail that surprises beginners is that Spark is lazy. Reading a table or filtering a DataFrame does not immediately do the work. Spark records what you asked for as a plan, and only starts computing when something demands a result, such as display(), show(), a write, or a count. This lets Spark look at your whole plan, reorder it, and avoid reading data it does not need. The practical effect is that an error in an early step may only appear when a later step finally runs.

Switch to SQL in a new cell to see the same data from another angle.

SQL
%sql
SELECT pickup_zip, COUNT(*) AS trips, ROUND(AVG(fare_amount), 2) AS avg_fare
FROM samples.nyctaxi.trips
GROUP BY pickup_zip
ORDER BY trips DESC
LIMIT 10

The first line, %sql, is the magic command. The query counts trips per pickup postal code and shows the ten busiest. SQL and the DataFrame API compile to the same plan underneath, so choose whichever reads better for the question in front of you. Analysts usually start in SQL, engineers often prefer Python, and a good notebook uses each where it is clearer.

Run order is not page order Cells run in the order you click them, not the order they appear. If you run cell five, edit cell two and never rerun it, the page shows a result that no longer matches the code above it. Before you share a notebook or trust a result, restart the session and run all cells from the top. If it still works, it is honest.
Try it
  1. Create a notebook and attach it to compute.
  2. Run the DataFrame cell and the SQL cell above.
  3. Click the plus icon on the table output to add a bar chart of trips per pickup zone.
a table of ten rows and then a chart built from it. Notice that both cells took about as long to run as the first row took to appear, because Spark only did the work when you asked.

Compute: what actually runs your code

A notebook is only text until it is attached to compute. Beginners often ask why Databricks has so many kinds, so it helps to see them as answers to different questions.

Serverless compute is the default for most new work. You attach a notebook or a job to it and it starts in moments, because Databricks keeps capacity ready and manages everything for you. There is no cluster to size, no runtime to choose, and no machine left running after you stop. Databricks upgrades the underlying runtime for everyone at once, which is why it is called versionless. There are limits: older low-level Spark interfaces such as RDDs and the Spark context are not available, and some Spark settings are not allowed. For a beginner's work these rarely matter.

All-purpose clusters are classic compute that you create and control. You choose the runtime, the machine type and the number of workers, and you can leave the cluster running for interactive work with many people. It is flexible and it is also the easiest way to waste money, because it bills while it runs, including while you are at lunch. Always set auto-termination so the cluster shuts itself down after a period of inactivity.

Job clusters are classic clusters that exist only for the length of a scheduled job. They start when the job starts and disappear when it ends. For repeated work they are cheaper than leaving an all-purpose cluster up.

SQL warehouses are compute specialised for SQL, dashboards and business intelligence tools. They come in sizes and in serverless, pro and classic flavours, and they stop themselves after a quiet period. When an analyst opens a dashboard or connects Power BI, a SQL warehouse answers.

Serverless

  • Starts in moments
  • No sizing, no runtime to pick
  • Upgrades itself
  • Some low-level Spark APIs unavailable

Classic cluster

  • You choose runtime and machines
  • Starts in minutes
  • You must set auto-termination
  • Full control, including custom settings

When you do create a classic cluster, you pick a Databricks Runtime version. Choose the newest LTS release unless a library forces you elsewhere, and choose the ML variant, written for example as "18 LTS ML", if you will train models, because it bundles MLflow, scikit-learn, PyTorch and similar libraries so you do not install them yourself. You will also see a checkbox for Photon, Databricks' native vectorised query engine. It speeds up SQL and DataFrame work, and it is billed at a higher rate per hour, so it pays off for heavy queries and not for tiny experiments.

An idle cluster still costs money A running all-purpose cluster bills whether or not you are typing. Set auto-termination to a short period when you create it, and stop clusters you no longer need. On a company account, a forgotten cluster over a weekend is a classic first-week mistake, and the cost shows up on someone else's report.
Try it
  1. Open the compute page in your workspace.
  2. Look at the options for creating compute and note which runtime versions are offered.
  3. Find the auto-termination setting and the Photon option, without creating anything yet.
a runtime list headed by an LTS version, and a termination timer. Knowing where those two controls live is worth more than any default.

Unity Catalog and the three-level name

Every table, volume, function and model in modern Databricks lives in Unity Catalog. It is a single place that stores the names of your data objects and the rules about who may use them. Understanding its shape early saves a great deal of confusion, because nearly every error message in the first week is about a name or a permission.

The hierarchy has three levels you use every day, under a container that you rarely touch.

METASTOREone per region
→
CATALOGoften one per environment
→
SCHEMAa folder of related objects
→
OBJECTtable, view, volume, model

The metastore is the top-level container for metadata in a region. Your account has one per region, and workspaces attach to it. A catalog is the first level you name, and teams commonly use one per environment or per business domain. A schema, sometimes called a database, groups related tables. The object at the bottom might be a table, a view, a volume for files, a function, or a registered model.

You always address an object as catalog.schema.object. A table called events in schema bronze in catalog course is course.bronze.events. Writing the full name every time is the safest habit, because it works from any notebook without depending on settings. You may also set a default with USE CATALOG course and USE SCHEMA bronze, after which events alone resolves, but defaults make notebooks fragile when someone else runs them with different settings.

Tables come in two ownership styles. A managed table is fully controlled by Unity Catalog, which chooses where its files live and removes them when the table is dropped. An external table points at storage you own, and Unity Catalog governs access to it but does not own the files. Begin with managed tables. They are simpler, and they let the platform apply maintenance for you.

Permissions are granted to users, groups and service principals, a service principal being a non-human identity for automation. Access follows the hierarchy, and one rule trips up almost everyone. To read a table you need SELECT on the table, plus USE SCHEMA on its schema, plus USE CATALOG on its catalog. Having the table permission alone is not enough, because you must be allowed to walk down the tree to reach it. As an owner of your own catalog you hold these automatically, so you will first meet the rule when you share your work with someone else.

SQL
GRANT USE CATALOG ON CATALOG course TO `data-team`;
GRANT USE SCHEMA, SELECT ON SCHEMA course.bronze TO `data-team`;
SHOW GRANTS ON SCHEMA course.bronze;

That example gives a group called data-team the ability to find and read everything in the bronze schema. It is shown here so that you recognise the shape. Granting to groups and not to individuals is the standard practice, because people join and leave and the group stays.

Try it
  1. Open the catalog browser and find samples.nyctaxi.trips.
  2. Read its column list and the sample data tab.
  3. Write its full three-level name on paper and underline each of the three parts.
a table you can query without having created it, and a clear sense of which part of the name is the catalog, which is the schema and which is the table.

Your first project: from raw trips to a table you own

Reading a sample table is useful, but the real skill is creating your own governed table. This section builds a small project from start to finish. Everything runs in a notebook on serverless compute, and everything is a few lines.

The goal is a table of the busiest taxi pickup zones that you own, stored in your own catalog, with history you can inspect. First create a catalog and a schema. If your workspace does not allow you to create a catalog, use the main catalog and your own schema instead, and replace course with main in every example that follows.

SQL
CREATE CATALOG IF NOT EXISTS course;
CREATE SCHEMA IF NOT EXISTS course.bronze;

IF NOT EXISTS makes the statement safe to rerun, which matters because notebooks get rerun constantly. A command that fails on its second execution is a command you will dread.

Next load a cleaned slice of the sample trips into a new managed table. The CREATE TABLE ... AS SELECT pattern creates a table from a query's result in one step.

SQL
CREATE OR REPLACE TABLE course.bronze.trips_clean AS
SELECT pickup_zip, dropoff_zip, trip_distance, fare_amount
FROM samples.nyctaxi.trips
WHERE fare_amount > 0 AND trip_distance > 0;

Run it, then check the result with a query.

SQL
SELECT COUNT(*) AS rows_loaded FROM course.bronze.trips_clean;

The same thing in Python uses the DataFrame writer. This is the form you will use when the transformation is easier to express in Python than in SQL.

PYTHON
from pyspark.sql import functions as F

trips = spark.read.table("course.bronze.trips_clean")

busiest = (
    trips.groupBy("pickup_zip")
         .agg(F.count("*").alias("trips"),
              F.round(F.avg("fare_amount"), 2).alias("avg_fare"))
         .orderBy(F.desc("trips"))
)

busiest.write.mode("overwrite").saveAsTable("course.bronze.busiest_zones")
spark.read.table("course.bronze.busiest_zones").display()

Read it slowly. groupBy and agg compute a count and an average per pickup zone, orderBy sorts the result, write.mode("overwrite") replaces the table if it already exists, and saveAsTable registers the result in Unity Catalog under the name you gave it. The final line reads it back, which proves that it exists in the catalog and not just in memory.

Notice what you did not do. You did not choose a file format, because Delta is the default for tables in Databricks. You did not decide where the files go, because the table is managed. You did not set up permissions, because you own what you created. The lakehouse made several decisions on your behalf, and the defaults are sensible.

Name tables for the stage they are in A common layout names schemas after stages of cleaning: bronze for raw copies of the source, silver for cleaned and joined data, and gold for tables shaped for a report or a model. It is called the medallion pattern, and it works because anyone reading a name can tell how far a table is from the raw data.
Try it
  1. Create the catalog and schema, then the trips_clean table.
  2. Run the Python cell to produce busiest_zones.
  3. Find both tables in the catalog browser and open the Details tab of one.
two tables listed under course.bronze, with type MANAGED and format Delta shown on the details page.

Delta tables: history, time travel and safe mistakes

The tables you just created are Delta tables. A Delta table is a folder of Parquet data files in cloud storage plus a transaction log that records every change as an ordered list of commits. The log is what turns a pile of files into something with database behaviour. A write either appears completely or not at all, a reader never sees half of a change, and the table remembers its past.

That memory is a gift for beginners, because it makes mistakes recoverable. Overwrite your table by accident and the old version is still there. Look at the history first.

SQL
DESCRIBE HISTORY course.bronze.busiest_zones;

The output has one row per commit, with a version number starting at zero, a timestamp, the user, and the operation, such as CREATE OR REPLACE TABLE AS SELECT or WRITE. Each time you ran your Python cell with overwrite, a new version was added. Nothing was lost.

To read an older version, ask for it by number.

SQL
SELECT * FROM course.bronze.busiest_zones VERSION AS OF 0;

This is time travel. You can also ask for a table as it was at a timestamp. And if a bad write needs undoing, you can restore the table to an earlier state rather than rebuild it.

SQL
RESTORE TABLE course.bronze.busiest_zones TO VERSION AS OF 0;

Restoring is itself a new commit in the log, so it does not erase history. That is worth noticing, because it means that even the undo is undoable.

Time travel is not unlimited. Old data files are cleaned up by a command called VACUUM, which removes files no longer referenced by recent versions, with a default retention of seven days. After vacuum has run, versions older than the retention window can no longer be read. Do not treat Delta history as a backup system. It is a short safety net for human error, not a replacement for proper retention of data you must keep.

VACUUM shortens your safety net History lets you travel back only as far as the data files still exist. Running VACUUM with a very short retention to save storage means you cannot time travel past it, and it can break readers that are still using older files. Leave the default retention alone until you understand why you would change it.

There are two other maintenance commands you will see often. OPTIMIZE compacts many small files into fewer larger ones so queries read less overhead. Small files accumulate quickly when a stream or frequent small writes feed a table. On managed Unity Catalog tables, Databricks can run this housekeeping for you automatically through a feature called predictive optimization, so as a beginner you mostly need to recognise the names and not schedule them yourself.

Try it
  1. Run the Python overwrite cell twice, then run DESCRIBE HISTORY on the table.
  2. Query VERSION AS OF 0 and compare it with the current version.
  3. Delete half the rows with a DELETE statement, then RESTORE the previous version and count the rows again.
a history table with a growing list of versions, and a row count that returns to its earlier value after the restore.

Files, volumes and secrets

Tables are not the only thing you work with. Sooner or later someone hands you a CSV, a JSON export or a folder of images. Unity Catalog stores files in volumes, which are governed folders that sit inside a schema, so files get the same names and permissions as tables. A volume path looks like /Volumes/<catalog>/<schema>/<volume>/<file>.

Create one and upload a file.

SQL
CREATE VOLUME IF NOT EXISTS course.bronze.raw;

Then use the catalog browser to open the volume and upload a small CSV through the interface. Once it is there, read it from a notebook.

PYTHON
df = (spark.read
        .option("header", True)
        .option("inferSchema", True)
        .csv("/Volumes/course/bronze/raw/my_file.csv"))
df.display()

header says the first line holds column names and inferSchema asks Spark to guess column types by scanning the data. Inference is convenient for exploration and risky for production, because a column that happens to look numeric in today's file might hold text tomorrow. Later you will declare schemas explicitly.

To look around, list the contents of a volume with dbutils, the utilities object that every notebook has.

PYTHON
dbutils.fs.ls("/Volumes/course/bronze/raw")

You will find old tutorials that use dbfs:/ paths, /FileStore, or mounts created with dbutils.fs.mount. These belong to an older way of working on the Databricks File System root. For new work, prefer Unity Catalog volumes and tables, because they carry permissions and lineage and the older locations do not.

The dbutils object does more than list files. Its secrets utility is how code reads a password or an API key without writing it into the notebook. A secret lives in a secret scope, and you fetch it by scope and key.

PYTHON
api_key = dbutils.secrets.get(scope="my-scope", key="api-key")

If you print a secret in a notebook, Databricks shows [REDACTED] instead of the value, a safeguard against leaking it into shared output. That safeguard is a courtesy and not a security boundary, so never rely on it to protect a value you have already stored in plain text somewhere else. Secrets are created with the command line interface or the API by someone with the right permission, and you should never paste a real token into a cell, even in a throwaway notebook, because notebooks get shared and exported.

Try it
  1. Create a volume and upload any small CSV through the catalog browser.
  2. Read it with spark.read and display it.
  3. Write it out as a managed Delta table with saveAsTable.
a CSV file turned into a governed table with history. That is the smallest useful ingestion pipeline there is.

The command line interface

Everything so far happened in a browser. The Databricks command line interface, usually just "the CLI", lets you do some of the same things from a terminal, which is where automation begins. You will use it to check that you are logged into the right workspace, to list resources, to trigger jobs, and later to deploy whole projects.

Install it for your system. On macOS with Homebrew:

BASH
brew tap databricks/tap
brew install databricks

The current documentation also shows a brew trust databricks/tap step on recent Homebrew versions, so run it between the two lines if your Homebrew asks for it. On Linux the project provides an install script:

BASH
curl -fsSL https://raw.githubusercontent.com/databricks/setup-cli/main/install.sh | sh

On Windows use WinGet or Chocolatey:

BASH
winget install Databricks.DatabricksCLI
choco install databricks-cli

Alternatively download a release archive from the project's GitHub releases page and put the binary on your path. Then check the installation.

BASH
databricks -v

You want a version in the 1.x line. The CLI reached 1.0 and is now updated often, so upgrade with the same tool you installed it with, for example brew upgrade databricks. If a team shares a project, everyone should use compatible CLI versions. A recent change to the way deployment state is stored means that very old clients cannot work with projects deployed by new ones.

Next, log in. The recommended way for a person is OAuth in the browser, called user-to-machine authentication.

BASH
databricks auth login --host https://<your-workspace-url>

The command opens your browser, you approve, and the CLI stores a token so you are not asked again. Since version 1.0 the token goes into your operating system's secure store, such as the macOS Keychain or Windows Credential Manager. On a headless Linux machine with no such store you can set the environment variable DATABRICKS_AUTH_STORAGE=plaintext to fall back to a plain file, which is acceptable for a throwaway machine and a poor habit anywhere else.

The CLI saves what it learns in a profile inside a file named .databrickscfg in your home directory. A profile is a named set of connection details, so you can keep one for development and one for production and switch between them with -p.

BASH
databricks auth profiles
databricks current-user me
databricks clusters list
databricks workspace list /Users/<you>
databricks fs ls /Volumes/course/bronze/raw/
databricks jobs list
databricks clusters list -p <profile>

current-user me is the best first command because it answers "who am I, and on which workspace?" in one line. If it returns your email, authentication works. The other commands list clusters, workspace folders, volume contents and jobs. Output is plain text by default and can be switched to JSON when you want to process it with other tools.

Check which workspace you are talking to The CLI acts on whichever profile it resolves. Environment variables such as DATABRICKS_HOST take priority over a profile in the config file, so a stale variable in your shell can send commands to the wrong workspace. When a result looks wrong, run databricks current-user me first.
Try it
  1. Install the CLI and run databricks -v.
  2. Log in with databricks auth login --host and your workspace URL.
  3. Run databricks current-user me and databricks clusters list.
your own email in the first output, and a possibly empty list in the second, since serverless needs no cluster. An empty list is a success, not an error.

Jobs: running your work on a schedule

A notebook you run by hand is an experiment. A notebook that runs every morning without you is a pipeline. Databricks calls the scheduling feature Jobs, part of the product family now branded Lakeflow. A job is a set of tasks with a trigger. A task can run a notebook, a Python file, a SQL query, a dbt project or a declarative pipeline. Tasks can depend on each other, so "ingest" runs before "clean", which runs before "publish".

Create one from the web interface. Open the jobs area in the sidebar and choose to create a job. Give the first task a name, choose the type notebook, point it at your notebook, and select serverless compute if it is offered. Run it once with the run-now control and watch the run page. It shows the status of each task, the output of the notebook, and the time taken.

To run it on a schedule, add a trigger. The usual one is a schedule written as a time and frequency, such as every day at six in the morning in a named time zone. Other triggers start a job when new files arrive in a location, or when another job finishes. You can also attach notifications so that a failure sends an email, which is the single most important setting on any job nobody watches.

From the terminal you can inspect and trigger the same job.

BASH
databricks jobs list
databricks jobs run-now <job-id>
databricks jobs get-run <run-id>

run-now starts a run and prints its identifier, and get-run reports on it. These commands become the building blocks for scripts and continuous integration pipelines. If you later keep jobs as code in a project deployed with the CLI, which the Declarative Automation Bundles feature supports, you will define the job in a file and deploy it with databricks bundle deploy. Older articles call that feature Databricks Asset Bundles. It is the same tool under a new name, and the commands did not change.

One principle matters from the first job: make the notebook safe to run twice. Schedulers retry failed runs, and humans click run again. A notebook that appends the same rows every time it runs will double your data. Prefer writes that replace a defined slice, or use merge statements that update matching rows, so that rerunning gives the same result as running once. This property is called idempotence, and it is the difference between a pipeline that recovers from failures and one that needs a person to clean up after it.

Try it
  1. Turn your busiest-zones code into a notebook and create a one-task job that runs it.
  2. Run it from the interface and open the run page.
  3. Run databricks jobs list in your terminal and find the job.
a successful run and a matching entry in the terminal listing. Because your write uses overwrite, running it twice leaves the same table.

MLflow: your first tracked run

Machine learning adds one new problem to data work: you will train the same model many times with different settings, and weeks later you will need to know which run produced which result. MLflow solves this by recording every attempt. Databricks hosts a managed MLflow, so there is no server to set up. The ML runtimes include it, and on serverless compute you can install the library with %pip install mlflow at the top of a notebook.

MLflow has a few nouns. An experiment is a named collection of runs, usually one per problem. A run is one execution of your training code. It records parameters, the settings you chose, metrics, the numbers you measured, and artifacts, the files produced, including the model itself. Here is the smallest useful example, using a tiny built-in dataset.

PYTHON
import mlflow
from sklearn.datasets import load_diabetes
from sklearn.linear_model import Ridge
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error

X, y = load_diabetes(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)

with mlflow.start_run():
    alpha = 0.5
    model = Ridge(alpha=alpha).fit(X_train, y_train)
    rmse = mean_squared_error(y_test, model.predict(X_test)) ** 0.5
    mlflow.log_param("alpha", alpha)
    mlflow.log_metric("rmse", rmse)

Run the cell, then open the experiments area from the sidebar, or follow the link that appears in the cell output. Your run is listed with its parameter and metric. Change alpha, run again, and the two runs sit side by side so you can compare them in a chart. That comparison, repeated across dozens of attempts, is what makes experiment tracking valuable.

Models you want to keep go into the model registry, and on current Databricks that registry lives in Unity Catalog, so a model gets a three-level name like any table, for example course.models.churn. Versions of a model are labelled with aliases such as champion, which your serving or batch code loads by name. The older Workspace Model Registry with its fixed stages named Staging and Production is deprecated, so if a tutorial calls transition_model_version_stage, it describes the legacy way. The Mid-level guide covers registering and promoting models, and the MLflow guide goes deeper into the library itself.

Try it
  1. Run the example three times with different alpha values.
  2. Open the experiment and select all three runs.
  3. Compare them and note which alpha gives the lowest RMSE.
three rows in an experiment table and one clearly best setting. You now have a record that survives closing the notebook.

Common errors and how to read them

Databricks errors look intimidating because they combine Spark, Delta, Unity Catalog and cloud messages in one stack trace. The trick is to read from the top for the error class in square brackets, such as [TABLE_OR_VIEW_NOT_FOUND], and then to scan for the first line that mentions something you wrote. The long Java frames below it are almost never the cause. The exact wording varies by runtime version, so use the patterns below as a guide to what to look for and check the error reference in the documentation for the precise text.

Table or view not found. The message says that a table or view cannot be found. The usual causes are a missing catalog or schema prefix, a typo, or a missing permission, because Unity Catalog may report "not found" for objects you are not allowed to see. Check the three-level name against the catalog browser and then check your grants.

Insufficient privileges. The message names a missing privilege such as USE SCHEMA on a schema. Remember that SELECT on a table is not sufficient on its own. The reader also needs USE CATALOG and USE SCHEMA on the parents. Ask the owner for the grant to the group you belong to.

Cannot configure default credentials. The CLI or the Python SDK prints that it cannot configure default credentials and points you at the authentication documentation. It means no profile, host or token could be found. Run databricks auth login --host <url> or set DATABRICKS_HOST and a credential, then retry.

Unauthorized, 401 or invalid access token. Your token expired or was revoked. Log in again with databricks auth login.

Cluster terminated, cloud provider launch failure. A classic cluster could not start, typically because your cloud account lacks quota or capacity for the machine type. Choose a different instance type or ask your cloud administrator to raise the limit. Serverless avoids this class of problem for you.

Schema mismatch when writing to a Delta table. You tried to write columns that differ from the table. Decide whether the new shape is intended. If it is, enable mergeSchema for that write, and if it is not, fix the DataFrame.

Read the first line, then the first line of yours Copy the bracketed error class into a search of the Databricks error-message reference. Each class has a page that explains the condition and shows an example. That is faster than guessing from a stack trace and teaches you the vocabulary the platform uses.
Try it
  1. Query a table that does not exist, such as course.bronze.nope.
  2. Note the error class and the sentence that follows it.
  3. Fix it by correcting the name and rerun.
an error that begins with a bracketed class and names the object it could not find. Learning to spot that shape is the biggest time saver in the first month.

Putting it all together

Here is one small project that uses everything above. You will ingest a file, build a table, schedule it, and track a model. Budget an hour and use a fresh notebook.

  1. Prepare the home. Create a catalog course with schemas bronze and models, and a volume course.bronze.raw.
  2. Ingest. Upload a small CSV of your choice into the volume, read it with spark.read, and save it as course.bronze.my_data with saveAsTable.
  3. Inspect. Run DESCRIBE HISTORY on the table, then overwrite it once and read the earlier version with VERSION AS OF 0.
  4. Summarise. Build a second table, course.bronze.my_summary, with a groupBy aggregation that answers one question about your data.
  5. Schedule. Put steps two to four in a notebook, create a one-task job on serverless compute, add an email notification on failure, and run it once.
  6. Track. In another notebook, train a small scikit-learn model on your summary or on the diabetes dataset, and log a parameter and a metric with MLflow.
  7. Drive it from the terminal. Log in with the CLI, run databricks current-user me, list your jobs, and trigger the job with databricks jobs run-now.

The finished result is modest, and it contains every moving part of a real project: governed storage, a repeatable transformation, a schedule, a tracked experiment and a scripted entry point. A production system is the same shape with more tables, more tasks, stronger permissions and a deployment process.

Try it
  1. Complete the seven steps in order.
  2. Write a three-line description of the project that a colleague could read before opening it.
  3. Put the notebooks in a Git folder and commit them.
a project that survives you closing your laptop: the data is in tables, the logic is in a job, the experiments are in MLflow and the code is in Git.

What you can now do, and what comes next

You can now explain what the lakehouse is and why it replaced the lake-plus-warehouse pair. You can describe an account, a workspace, the two planes and the catalog, and say which of them a given problem belongs to. You can run code in a notebook, choose between serverless compute, a cluster and a SQL warehouse, and avoid the most expensive beginner mistake of leaving compute running. You can create catalogs, schemas, tables and volumes with correct three-level names, read the history of a Delta table, travel back in time and restore it. You can read the common errors and fix the usual causes. You can schedule a notebook as a job, track a first MLflow run, and use the command line interface to authenticate and drive the platform.

The Mid-level guide picks up where this one stops. It covers Unity Catalog permissions in depth, incremental ingestion with Auto Loader, merge statements and table maintenance, registering models in Unity Catalog with aliases, the Python SDK, Databricks Connect for developing from your own editor, and how to run jobs with service principals instead of your personal identity. The Senior guide then covers deployment with Declarative Automation Bundles, environment separation, cost governance, security, and the points where Databricks is not the right tool.

Neighbouring guides help here: Delta Lake and Spark explain the engines underneath, and Terraform and GitHub Actions pair with bundles.

The best next step is to keep the project from the last section alive and extend it a little each week: add a second source, add a permission, add a test. Platforms are learned by owning something small that has to keep working.

Sources