Skip to content
Back to student guides
Apache SparkMLOpsDistributed compute3 levels111 sectionsCovers Apache Spark 4.2

The Complete Apache Spark Guide

Process big data and train models at scale with Apache Spark and PySpark. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
17sections
28examples

This is part one of three. It covers everything you need to do real work with Apache Spark on a single machine, not a teaser. By the end you can install Spark, start a session, load a CSV file, clean it, summarise it with both the DataFrame API and SQL, write the result as Parquet, read the web interface to see what your code actually did, and recognise the handful of errors that stop almost every newcomer. Mid-level and Senior take the same topics further; nothing here is thrown away.

The guide targets Apache Spark 4.2, which the official downloads page lists as the current stable release at the time of writing. Patch numbers move quickly, so check https://spark.apache.org/downloads.html on the day you install. Everything below uses Python (PySpark), because that is what most data and machine learning teams reach for first.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and Spark's ideas stick only once you have watched your own job run, finish, and show up in the web interface.

What Spark is, and the problem it solves

Apache Spark is an engine for processing data that is too big, or too slow to process, on one core of one machine. You describe a computation (read these files, keep the rows where this is true, group by that column, average this number) and Spark splits the data into pieces, spreads the pieces across many processor cores or many machines, runs your computation on every piece at once, and stitches the answer back together.

YOUR CODEPySpark, SQL
→
PLANwhat to compute
→
TASKSone per data piece
→
RESULTfiles or a table

The diagram is the whole idea, and the rest of this guide fills in each box. Notice what is missing: you never write a loop over the pieces, you never open threads, you never decide which machine handles which rows. You write what you want, and Spark works out how.

Why would you need that? Picture a retailer in Cairo or Riyadh with three years of order history: a few hundred million rows across thousands of files. A script that loops through them in plain Python takes hours and may run out of memory halfway. The same question in Spark, running on a laptop with eight cores, already finishes several times faster, because eight pieces are processed at once. Point the same code at a cluster of fifty machines and nothing in your program changes; it just finishes sooner. That property, the same code on a laptop and on a cluster, is the reason Spark became the default tool for large-scale data work.

What people use it for:

🧹

Cleaning and joining big data

Turn raw event logs, exports and tables into tidy datasets that analysts and models can use.

📊

Large aggregations

Daily, weekly and per-customer summaries over billions of rows, written with SQL or DataFrames.

🤖

Feature engineering for ML

Build the input columns for a model at full scale, and run batch scoring over the whole customer base.

🌊

Streaming

Process data as it arrives, using the same DataFrame code you wrote for files (covered in the Mid-level guide).

Spark is also a family of parts that share one engine: Spark SQL and DataFrames for tables, Structured Streaming for data that keeps arriving, MLlib for machine learning, and GraphX for graphs. This guide stays with the first one, because it is the foundation the others stand on.

Try it
  1. Think of one data task you or your team does that takes too long, or does not fit in memory, on a single machine.
  2. Write one sentence describing it in the form: read X, keep Y, group by Z, produce W.
a sentence made of verbs like read, filter, group and join. Those are exactly the operations Spark is built to spread across many cores, and you will write that sentence as code by the end of this guide.

What came before Spark

To appreciate the design, you need the story in one paragraph. In the mid-2000s, companies with huge datasets used a model called MapReduce, popularised by Hadoop. You wrote two kinds of functions, a map that processes each record and a reduce that combines results, and the system ran them across a cluster of cheap machines. It worked and it scaled, but every step wrote its intermediate results to disk, and real analyses needed many steps chained together. A ten-step analysis meant ten rounds of writing to and reading from disk, and writing the steps in the MapReduce style was tedious.

Spark kept the central insight, splitting data across machines and moving the computation to the data, and changed two things. First, it keeps intermediate data in memory where it can, so chained steps do not pay the disk round trip each time. Second, it offers a high-level API: instead of hand-writing map and reduce functions you call familiar operations such as filter, join and groupBy, or write SQL, and an optimiser decides how to execute them efficiently.

Two terms you will meet constantly belong here. Hadoop is an ecosystem; its file system, HDFS, and its resource manager, YARN, are still widely used, and Spark can run on YARN and read from HDFS. A cluster is just a group of machines that cooperate on one job. You do not need a cluster to learn Spark. Everything in this guide runs on your laptop, in what Spark calls local mode, where one program plays every role.

One practical note on history, because tutorials online span many years. Older material uses an API called the RDD (resilient distributed dataset), a low-level distributed collection. You will still see it in old blog posts and Stack Overflow answers. For nearly all work today you use DataFrames instead, which are higher level, easier to write and usually faster, because Spark can optimise them. If a tutorial starts with sc.parallelize, it is teaching the older style, and you can skip ahead to its DataFrame section.

A second warning about old tutorials: Spark 4 only supports Scala 2.13, Java 17 or newer, and it removed support for Mesos as a cluster manager. If an article tells you to download a _2.12 jar or run Spark on Mesos, it was written for Spark 3 or earlier.

Check the date on every tutorial Spark 4.0 changed real defaults (most importantly, ANSI SQL mode is now on, covered later in this guide). A tutorial written for Spark 2 or 3 can show output that Spark 4 will not reproduce. Look at the version first.
Try it
  1. Search for any "Spark tutorial" that interests you.
  2. Find whether it uses RDDs (sc.parallelize, .map(lambda ...) on raw collections) or DataFrames (spark.read, .select), and what Spark version it names.
you can now classify a tutorial as current or dated within a minute, which saves hours of confusion.

The mental model: driver, executors, partitions

Spark has a small vocabulary, and nearly every confusing moment later traces back to one of these nouns being fuzzy. Learn them now.

A running Spark program is an application. Every application has exactly one driver: the process that runs your main code, builds the plan for your computation, and hands out work. The driver holds your SparkSession, the object you type spark. on. Work itself happens in executors: processes, usually on other machines, that run tasks and can hold data in memory for reuse. The software that hands out machines and memory to the application is the cluster manager. Spark 4 supports Standalone (Spark's own simple one), Hadoop YARN and Kubernetes; and in local mode, no external manager at all.

DRIVERplans and schedules
→
CLUSTER MANAGERlocal, YARN, K8s
→
EXECUTORSrun tasks, cache data

The data side has its own nouns. A DataFrame is a distributed table: named columns with types, rows spread over many machines. You treat it like one table, and Spark handles the spreading. The pieces it is spread into are partitions; a partition is the chunk of data that a single task processes. If a DataFrame has 8 partitions and you have 8 free cores, 8 tasks run simultaneously. This is the single most important relationship in Spark: parallelism equals partitions processed at the same time.

Now the part that surprises everyone coming from pandas. Spark operations come in two kinds.

  • Transformations describe a new DataFrame from an old one: select, filter, withColumn, groupBy, join. They are lazy. Calling one does not touch any data; it just adds a step to the plan.
  • Actions ask for an actual result: show, count, collect, write. An action makes Spark take the whole accumulated plan, optimise it, and execute it.

Think of ordering at a restaurant. Transformations are you telling the waiter what you want, one dish at a time; nothing is cooked. The action is "bring the bill": at that point the kitchen reads the whole order and does the work in the sensible order. The benefit of laziness is that Spark sees your whole recipe before cooking, so it can skip unneeded columns, push filters as early as possible, and combine steps.

Executing an action produces a job. Spark splits a job into stages, and splits each stage into tasks, one per partition. Stage boundaries happen at a shuffle: any step that needs rows with the same key to end up together, such as a groupBy or a join, forces data to be redistributed across executors. Shuffles are the main cost in Spark, because they move data over the network and write it to disk. You do not need to tune them yet, but you should be able to say "this line causes a shuffle" when you see a groupBy.

Summarising the nouns in one table:

Term Meaning
Application One SparkSession's lifetime, one program run
Driver The process running your code and planning the work
Executor A process that runs tasks and caches data
Cluster manager Allocates machines and memory (Standalone, YARN, Kubernetes)
DataFrame A distributed table with named, typed columns
Partition The slice of data one task works on
Transformation A lazy step that describes a new DataFrame
Action A step that triggers execution and returns or writes a result
Job, stage, task Action, shuffle-separated section of it, one-partition unit of work
Shuffle Redistributing data across executors (joins, group by)
Local mode hides the cluster, not the concepts On your laptop the driver and executors run inside one Java process, so there is no network and no second machine. The vocabulary, the laziness and the shuffles are all still there, which is why practising locally transfers to a real cluster.
Try it
  1. Without running any code, classify each as a transformation or an action: filter, count, select, show, write, join.
  2. Then mark which of them would force a shuffle (hint: two of the transformations need rows with the same key together, only one of them is in this list).
filter, select and join are transformations; count, show and write are actions; join is the shuffling one. If that took you more than a minute, re-read the two bullet points above.

Installing Spark and checking it works

The easiest route is pip, because the PySpark package bundles Spark itself. You need two things on the machine first: a Java Development Kit and Python. Spark 4.2 documents support for Java 17, 21 or 25, and Python 3.10 or newer. Install a JDK from your package manager or a vendor you trust, and make sure the JAVA_HOME environment variable points to it.

Check both before anything else:

BASH
java -version
python3 --version
echo $JAVA_HOME

If java -version fails, or JAVA_HOME is empty, fix that first; it is the cause of most failed first starts. Now create an isolated Python environment and install PySpark:

BASH
python3 -m venv .venv
source .venv/bin/activate          # Windows PowerShell: .venv\Scripts\Activate.ps1
pip install pyspark

A virtual environment (the .venv folder) is a private copy of Python's packages for this project, so installing PySpark does not disturb anything else on your machine. Always activate it in each new terminal before working.

Then run a one-line smoke test that starts Spark, prints its version and prints five numbers:

BASH
python -c "from pyspark.sql import SparkSession; s=SparkSession.builder.master('local[*]').getOrCreate(); print(s.version); s.range(5).show()"

You should see log lines, then a version string and a small table:

TEXT
+---+
| id|
+---+
|  0|
|  1|
|  2|
|  3|
|  4|
+---+

If you see that table, Spark works. spark.range(5) generates a one-column DataFrame named id; it is the quickest way to get test data with no files at all.

There are other ways to install, and each suits a different person. With conda, conda install -c conda-forge pyspark inside a new environment does the same job. If you want the full Spark distribution with its command-line tools, download the tarball from the downloads page, unpack it, set SPARK_HOME to that folder and add its bin directory to your PATH:

BASH
tar xzvf spark-4.2.0-bin-hadoop3.tgz
cd spark-4.2.0-bin-hadoop3
export SPARK_HOME=$(pwd)
export PATH=$SPARK_HOME/bin:$PATH
pyspark --master "local[2]"

That last command opens an interactive Python shell with a ready-made spark session. The same distribution gives you spark-submit, which is how scripts are launched on clusters, and you can confirm any installation with:

BASH
spark-submit --version

If you prefer containers, the project publishes the apache/spark image on Docker Hub; see the Docker guide at /student-guides/docker for the basics before trying it.

Windows: prefer WSL2 Spark runs on Windows, but some local file operations depend on Hadoop helper binaries (winutils.exe) that are awkward to set up. The least painful path for a beginner is WSL2 with Ubuntu, then following the Linux steps above. If you do stay on native Windows, expect file-writing problems until Hadoop's Windows helpers are installed and HADOOP_HOME is set.
Try it
  1. Run the three check commands, then create the virtual environment and install PySpark.
  2. Run the smoke test.
  3. Change range(5) to range(20) and rerun. Notice how long the first start takes compared with the second.
the five-row table, then twenty rows. Startup takes a few seconds because Spark launches a Java process; that cost is paid once per session, not once per query.

Your first session and first DataFrame

Everything in Spark starts from a SparkSession, the entry point object. Create a file called first.py:

first.py
from pyspark.sql import SparkSession

spark = (SparkSession.builder
         .appName("first-steps")
         .master("local[*]")
         .getOrCreate())

print("Spark version:", spark.version)

Read the builder chain from top to bottom. appName is a label that appears in the web interface, which helps when several applications run. master("local[*]") says "run here, on this machine, using as many threads as there are cores"; local[2] would use two. getOrCreate() returns the existing session if one is running, or creates one. Using it everywhere keeps you from accidentally starting two.

Now make a tiny DataFrame by hand, so you can experiment without any files:

first.py
people = spark.createDataFrame(
    [("Aya", "Egypt", 29), ("Omar", "Saudi Arabia", 41), ("Layla", "Egypt", 35), ("Khaled", "UAE", 17)],
    schema=["name", "country", "age"],
)

people.printSchema()
people.show()

Run it with python first.py. The output has two parts. printSchema() prints the column names and types, like name: string and age: long, along with whether each may contain nulls. show() prints the first rows as a table. The schema (the list of column names and types) matters a great deal in Spark, because the optimiser and the file formats rely on it.

Two methods you will use hundreds of times for orientation: people.columns returns the column names as a Python list, and people.count() returns the number of rows. Note which is an action. count() runs a job; columns only reads the schema, which Spark already knows, so it is instant.

Use show() to look, never collect() show() prints a handful of rows. collect() and toPandas() pull every row into the driver's memory, which is fine for ten rows and fatal for ten million. When you only want a peek, use show(5) or limit(5).

When the script ends, the session ends too. In a notebook or an interactive shell, you can stop it explicitly with spark.stop(), which releases memory and the port the web interface used.

Try it
  1. Add two more people to the list, one with a country you invent.
  2. Print people.columns and people.count().
  3. Change one age to a string such as "thirty" and run again to see what Spark does with mixed types.
the counts update as you add rows. Mixing a string into the age column makes creation fail with a type error rather than silently storing text, which is Spark protecting the schema.

Reading real data: a first project, step by step

Toy data is for learning syntax. Now build something real: a small analysis of a CSV file. Create a file orders.csv in your project folder with these contents (add more rows if you like; the shape is what matters):

orders.csv
order_id,customer,country,category,amount,order_date
1001,Aya,Egypt,books,120.50,2026-01-04
1002,Omar,Saudi Arabia,electronics,899.00,2026-01-04
1003,Layla,Egypt,books,35.25,2026-01-05
1004,Khaled,UAE,electronics,1250.00,2026-01-06
1005,Aya,Egypt,electronics,299.99,2026-01-07
1006,Noor,Jordan,home,64.10,2026-01-07
1007,Omar,Saudi Arabia,home,210.00,2026-01-08
1008,Layla,Egypt,books,18.75,2026-01-09

Step one is reading it. A CSV file has no embedded types, so you must tell Spark two things: that the first line is a header, and how to work out column types.

analysis.py
from pyspark.sql import SparkSession, functions as F

spark = SparkSession.builder.appName("orders").master("local[*]").getOrCreate()

orders = spark.read.csv("orders.csv", header=True, inferSchema=True)
orders.printSchema()
orders.show(5)

header=True uses the first line as column names; without it you get _c0, _c1 and a header row mixed into your data. inferSchema=True makes Spark scan the file to guess types, so amount becomes a decimal-capable number rather than text. The cost is an extra pass over the data, which is trivial here and noticeable on huge files. For anything beyond learning, it is better to declare the schema yourself, which is faster and safer:

analysis.py
from pyspark.sql.types import StructType, StructField, StringType, IntegerType, DoubleType, DateType

schema = StructType([
    StructField("order_id", IntegerType()),
    StructField("customer", StringType()),
    StructField("country", StringType()),
    StructField("category", StringType()),
    StructField("amount", DoubleType()),
    StructField("order_date", DateType()),
])

orders = spark.read.csv("orders.csv", header=True, schema=schema)

Declaring a schema also protects you: if next month's file has "N/A" in the amount column, you find out loudly instead of discovering that a column was quietly inferred as text. Other common readers follow the same pattern: spark.read.json(...), spark.read.parquet(...), and spark.read.format("jdbc") with options for reading from a database.

Reading is lazy too A wrong file path does not always fail at the line that names it. Some problems only appear when the first action runs. If an error mentions a path, look at the read statement even if the traceback points at a later show() or count().
Try it
  1. Create orders.csv and run the read with inferSchema=True. Print the schema.
  2. Remove header=True and run again.
  3. Restore it and switch to the explicit schema.
with the header option, real column names; without it, columns named _c0 to _c5 and your header row as data. Seeing that wrong version once means you will recognise it forever.

Transforming data: select, filter, withColumn, groupBy

With the DataFrame loaded, you now describe the analysis. The functions live in pyspark.sql.functions, conventionally imported as F, and a column is referenced with F.col("name").

Selecting picks columns, and filtering keeps rows:

analysis.py
big_orders = (orders
    .select("order_id", "customer", "country", "amount")
    .filter(F.col("amount") > 100))
big_orders.show()

Adding or changing a column uses withColumn, which takes the new column's name and an expression. Here we add a rounded amount with a flag column:

analysis.py
enriched = (orders
    .withColumn("amount_rounded", F.round(F.col("amount"), 0))
    .withColumn("is_large", F.col("amount") >= 500))
enriched.show(3)

DataFrames are immutable: withColumn does not change orders, it returns a new DataFrame. If you write orders.withColumn(...) on a line by itself and never assign the result, nothing is kept. This catches nearly everyone once.

Grouping and aggregating is where Spark earns its keep. groupBy names the key, and agg lists the summary calculations, each given a readable name with alias:

analysis.py
by_country = (orders
    .groupBy("country")
    .agg(F.count("*").alias("orders"),
         F.sum("amount").alias("revenue"),
         F.avg("amount").alias("avg_order"))
    .orderBy(F.desc("revenue")))
by_country.show()

The expected output is a table with one row per country, with Saudi Arabia and the UAE at the top for revenue in the sample data. Every groupBy is a shuffle, so this is the first line where Spark has to move rows between partitions so that all of one country's rows meet. On eight rows you cannot feel it; on eight hundred million rows it is the line you would watch.

Two more everyday operations. Joining combines two DataFrames on a key, for example attaching a table of country regions to the orders:

analysis.py
regions = spark.createDataFrame(
    [("Egypt", "North Africa"), ("Saudi Arabia", "Gulf"), ("UAE", "Gulf"), ("Jordan", "Levant")],
    schema=["country", "region"],
)
with_region = orders.join(regions, on="country", how="left")
with_region.select("order_id", "country", "region").show()

The how argument is the join type; left keeps every order even when no region matches, filling the missing region with null. And handling missing data has its own methods: df.dropna() removes rows containing nulls, and df.fillna({"amount": 0}) replaces nulls in named columns with a value. Choose deliberately; silently dropping rows is the usual way analyses go wrong.

Chaining all of this reads like a sentence, and nothing has run yet. Until you call show, count or write, Spark has only a plan.

Wrap long chains in parentheses Python needs the outer parentheses to let you break a method chain across lines without backslashes. Put one operation per line, and the code reads as a pipeline you can comment out one step at a time while debugging.
Try it
  1. Build a DataFrame of orders in the books category only, with columns customer and amount.
  2. Compute total spend per customer and show the top three.
  3. Write orders.withColumn("x", F.lit(1)) on its own line, then print orders.columns.
Aya and Layla appear in the books list; orders.columns does not contain x, which proves the DataFrame is immutable.

Spark SQL: the same engine, in SQL

If you know SQL, you can use it directly. Spark SQL and the DataFrame API are two front doors to the same engine, and mixing them is normal. First give the DataFrame a name that SQL can see, called a temporary view:

analysis.py
orders.createOrReplaceTempView("orders")

spark.sql("""
    SELECT country,
           count(*)      AS orders,
           round(sum(amount), 2) AS revenue
    FROM orders
    WHERE order_date >= '2026-01-05'
    GROUP BY country
    ORDER BY revenue DESC
""").show()

spark.sql(...) returns a DataFrame, so you can keep chaining DataFrame operations on its result, or the other way round. A temporary view is only a name for a plan inside this session; it copies no data and disappears when the session ends.

When should you choose which? Use SQL when the logic is a clear query, particularly joins and window-style analyses that you or analysts already write in SQL. Use the DataFrame API when you are assembling steps programmatically, looping over column names, or want your editor to catch typos in function names. Because the optimiser sees both the same way, performance is not the deciding factor.

You can also look at what Spark intends to do. Call explain() on any DataFrame to print the plan:

analysis.py
by_country.explain()

The output is a tree read from the bottom up: a scan of the file, a partial aggregation on each partition, an exchange (that word means shuffle), then the final aggregation. You do not need to understand every line today. What matters is that you can find the word Exchange and know a shuffle happens there. With explain("formatted") you get a more readable layout. Note that in Spark 4.2 the plan may also mention adaptive execution; that is a runtime optimiser that revisits the plan using real data sizes, and it is switched on by default, so there is nothing for you to do.

DataFrame API

  • Build queries from code, in loops and functions
  • Editor completion and linting help
  • Easy to unit test small steps

SQL

  • Familiar to analysts and anyone who knows databases
  • Best for one clear, self-contained query
  • Errors in column names appear only when run
Try it
  1. Register orders as a temp view and write a SQL query for the average amount per category.
  2. Call .explain() on its result and find the line containing Exchange.
three categories with averages, and a plan that includes an exchange, which is the shuffle caused by the GROUP BY.

Writing results, and the Parquet format

Analysis is only useful if the result lands somewhere. The write side mirrors the read side:

analysis.py
by_country.write.mode("overwrite").parquet("out/revenue_by_country")

Three things to understand here. First, write is an action: it is the line that actually runs your whole pipeline. Second, mode("overwrite") tells Spark what to do when the target exists; the options are overwrite, append, ignore and the default error/errorifexists, which refuses to write if the path exists. Third, the path is a directory, not a file.

Open out/revenue_by_country and you will find files with names like part-00000-....snappy.parquet, plus an empty file named _SUCCESS. Each part- file is written by one task for one partition, which is parallelism leaking into your file system. That is why a big job produces many files, and why every Spark reader accepts a directory as input.

Parquet is the format you should default to. It is a columnar format: values of one column are stored together, with compression and the schema embedded. Compared with CSV, a Parquet file is smaller, keeps its types, and lets Spark read only the columns a query needs. Read it back with spark.read.parquet("out/revenue_by_country") and printSchema() shows the original types with no inference. Spark can also write CSV and JSON with .csv(...) and .json(...), which is useful for handing data to people or tools that need it, but for data you will read again, choose Parquet.

When a file-per-partition layout is a problem, you can control how many files appear with coalesce, which reduces the partition count without a full shuffle:

analysis.py
by_country.coalesce(1).write.mode("overwrite").csv("out/revenue_csv", header=True)

That produces a single part file, which is convenient for a small summary and harmful for a large dataset, since one task then does all the writing. For big outputs, leave the partitioning alone, or organise files by a column with partitionBy:

analysis.py
orders.write.mode("overwrite").partitionBy("country").parquet("out/orders_by_country")

This creates sub-folders such as country=Egypt/, so later queries that filter on country read only the matching folders. Choose partition columns with a modest number of distinct values; partitioning by something unique like order_id creates millions of tiny files, a classic performance trap.

For data residency: if your employer in the Gulf or Egypt requires data to stay in-country, the same write call works against cloud object storage in a regional bucket, so "where the files land" is decided by the path you write to and the bucket's region, not by Spark.

Overwrite means overwrite mode("overwrite") replaces everything at the target path. Double-check the path before you run it, and never point it at a folder that holds anything you cannot regenerate.
Try it
  1. Write by_country to Parquet and list the output directory.
  2. Read it back and run printSchema().
  3. Write the orders partitioned by category and look at the folder names.
part files plus _SUCCESS, types preserved on read-back, and folders named category=books and so on.

Watching your job: the web interface

When a Spark session is alive, it serves a web interface at http://localhost:4040. If that port is busy, because another session is running, Spark takes 4041, then 4042, and so on. This page is how you stop guessing about what Spark did.

To use it, you need a session that stays alive long enough to look at. Add this to the end of a script, run it, and visit the address:

analysis.py
by_country.show()
input("Press Enter to stop Spark... ")
spark.stop()

The page has tabs. Jobs lists one entry per action you ran, so the show() above appears as a job, and a write appears as another. Click into a job and you reach Stages, which are the shuffle-separated sections, each with a count of tasks and the time they took. SQL / DataFrame shows the query plan as a diagram with row counts, which is the friendliest place to see where data shrinks or grows. Executors lists the processes (in local mode, just the driver), with memory use. Storage shows cached data, and Environment lists every configuration value in effect.

What to look for as a beginner: number of tasks in a stage (does it match your partition count?), the longest task compared with the median (a big gap hints that some partition has far more data than the others), and any red failed tasks. Do not try to read everything. Learn to answer "which stage was slow?", because that tells you which line of code to blame.

Spark records its own plan too, which is why explain() and the SQL tab agree. When the session ends, the web interface disappears. Keeping a record of finished applications requires event logging and a History Server, which belong in the Mid-level guide.

Name your application The appName you set appears at the top of the interface. When you run three experiments side by side, a name like orders-cleanup saves you from opening the wrong tab.
Try it
  1. Add the input(...) pause to your script and open http://localhost:4040.
  2. Find the job for your show() and open its stages.
  3. Open the SQL tab and find the aggregation in the diagram.
a job with two stages (one before the shuffle, one after), and a plan diagram whose rows-output counts shrink after the aggregation.

Lazy evaluation and ANSI mode: the two behaviours that surprise people

Two behaviours deserve their own section, because they produce the most "but that worked in pandas" confusion.

Laziness, revisited with an example. Run this and notice what is slow:

analysis.py
step = orders.filter(F.col("amount") > 100)      # returns instantly
step2 = step.withColumn("tax", F.col("amount") * 0.14)   # returns instantly
print(step2.count())                              # the work happens here

The first two lines return immediately because they only extend the plan. The third line executes everything. A consequence is that errors can appear far from their cause: a typo in the column name on line one raises an error when Spark analyses the plan, which for DataFrames is usually at line one, but a data problem such as a malformed value only appears when an action runs. Another consequence is repetition: if you call two actions on the same DataFrame, Spark recomputes the plan from the source for both. For an expensive DataFrame that you will reuse, you can mark it for caching with step2.cache() and release it with step2.unpersist(). Caching is an optimisation for later; for now, know that it exists and that recomputation is the default.

ANSI mode. Since Spark 4.0, spark.sql.ansi.enabled is true by default. ANSI mode means Spark follows the SQL standard strictly: invalid operations raise errors instead of quietly returning null. Dividing by zero, casting the text abc to a number, and arithmetic overflow all fail loudly. Try it:

analysis.py
spark.sql("SELECT 1/0").show()

On Spark 4 this raises an error whose class is DIVIDE_BY_ZERO, and the message itself tells you what to do: use try_divide to get null instead, or turn the setting off. Similarly, SELECT CAST('abc' AS INT) raises a CAST_INVALID_INPUT error. If you have used Spark 3, where these returned null and carried on, this is the biggest behaviour change you will notice. Loud failures are usually a good thing for data quality, because bad rows no longer slip through as nulls.

You have two sane ways to deal with it. Use the tolerant functions when you expect bad values:

analysis.py
spark.sql("SELECT try_divide(10, 0) AS safe_division, try_cast('abc' AS INT) AS safe_cast").show()

Both return null rather than failing. Or, for old code you cannot change yet, set the legacy behaviour for the session:

analysis.py
spark.conf.set("spark.sql.ansi.enabled", "false")

Prefer the first approach in new code. Turning ANSI off hides exactly the problems it was introduced to expose, and a fix that makes the error vanish is not the same as a fix that makes the data correct.

Old tutorials assume ANSI is off If a Spark 2 or 3 tutorial shows a null where Spark 4 shows an error, this setting is almost always the reason. Check spark.version and the setting before assuming your code is wrong.
Try it
  1. Run spark.sql("SELECT 1/0").show() and read the error message from top to bottom.
  2. Run the try_divide version.
  3. Time a pipeline: put import time and timestamps around the transformation lines and around the count().
an error class name plus a hint pointing at try_divide; then a null; and timings showing almost all the time is spent in the action.

Configuration you will actually touch

Spark has hundreds of settings, and you will use a handful. They are set in three places. In code, with builder.config("key", "value") before getOrCreate(), or at runtime with spark.conf.set for SQL settings. On the command line, with --conf key=value passed to spark-submit or pyspark. And in a file, conf/spark-defaults.conf, in a full Spark distribution. When the same key appears in several places, the one set in code wins over command-line flags, which win over the defaults file.

A short list worth knowing as a beginner:

Setting What it does Default
spark.master / .master(...) Where to run: local[*], local[2], a cluster URL none; must be set locally
spark.sql.shuffle.partitions Partitions after a shuffle (join, group by) 200
spark.driver.memory Memory for the driver process set at launch
spark.sql.ansi.enabled Strict SQL error behaviour true in Spark 4
spark.sql.adaptive.enabled Runtime plan re-optimisation true

The shuffle partitions default of 200 is a good example of why defaults need thought. On a laptop with a tiny dataset, 200 tiny tasks after each shuffle is wasteful, and you may notice small jobs take longer than they should. For learning, lowering it is reasonable:

analysis.py
spark = (SparkSession.builder
         .appName("orders")
         .master("local[*]")
         .config("spark.sql.shuffle.partitions", "8")
         .getOrCreate())

Adaptive query execution, on by default, already merges small post-shuffle partitions in many cases, so the effect is smaller than it once was, but eight is a sensible number on a laptop. Memory deserves a note too. The driver in local mode is the whole application, so if you are processing a file larger than your RAM you may need to raise it, and for the driver this must be set before the JVM starts: pass --driver-memory 4g on the command line, or --conf spark.driver.memory=4g. Setting it with builder.config after Python has already launched Spark's JVM in an interactive shell does not take effect.

You can inspect any live setting with spark.conf.get("spark.sql.shuffle.partitions"), and the web interface's Environment tab lists them all.

Do not copy configuration from the internet blindly A tuning line that helped one job on a 200-node cluster may slow yours down. Change one setting at a time, and prove it helped by checking the web interface before and after.
Try it
  1. Print spark.conf.get("spark.sql.shuffle.partitions").
  2. Add the config line above, rerun, and print it again.
  3. Open the Environment tab and find the value there.
200 before, 8 after, and the same 8 visible in the interface.

Common errors and how to read them

A Spark error is usually long, because it carries a Python traceback on top of a Java stack trace. Read from the bottom up, and look for the part that starts with a bracketed error class or a line after Caused by:. The real message is nearly always short and near the end.

What you see What it means Fix
JAVA_HOME is not set, or Java gateway process exited before sending its port number PySpark could not start the Java process: Java is missing or JAVA_HOME points to the wrong place Install JDK 17, 21 or 25 and set JAVA_HOME; rerun the java -version check
A master URL must be set in your configuration A script launched without a master Add .master("local[*]") or pass --master to spark-submit
Python in worker has different version ... than that in driver Spark's workers found a different Python from your driver Set PYSPARK_PYTHON and PYSPARK_DRIVER_PYTHON to the same interpreter, ideally your virtual environment's
[UNRESOLVED_COLUMN.WITH_SUGGESTION] You named a column that does not exist Read the "Did you mean" list; check spelling and case with df.columns
[DIVIDE_BY_ZERO] or [CAST_INVALID_INPUT] ANSI mode rejected bad data Use try_divide or try_cast, or clean the data first
Task N in stage S failed 4 times A wrapper around a real problem inside a task Scroll to Caused by: for the actual cause
java.lang.OutOfMemoryError: Java heap space The driver or an executor ran out of memory, often after collect() or toPandas() Stop pulling big data to the driver; raise --driver-memory; write results to files instead
Address already in use / Service 'SparkUI' could not bind on port 4040. Attempting port 4041. Another session holds the web port Harmless warning; Spark picks the next port
ModuleNotFoundError: No module named 'pyarrow' An optional dependency is missing pip install "pyspark[sql]" in the active environment

Walk through one example to practise. You run orders.select("amout") with a typo. The error begins with a Python AnalysisException, then the bracketed class UNRESOLVED_COLUMN.WITH_SUGGESTION, a sentence saying that a column with name amout cannot be resolved, and a "Did you mean one of the following?" list containing amount. The suggestion is the answer, and you found it in the first lines because Spark's current errors are designed to be read, not just searched.

Another worth rehearsing: the file path error. If you write spark.read.csv("order.csv") and the file is orders.csv, the error says the path does not exist. Relative paths resolve from the folder you launched Python in, not from where your script lives, so a script that works in one terminal folder can fail in another. When in doubt, print os.getcwd() or use an absolute path.

And one that looks scary but is not: lots of WARN and INFO lines at startup. These are log messages, not failures. If they drown your output, lower the noise with spark.sparkContext.setLogLevel("WARN") after the session starts; the call is available in a classic session and is simply a convenience here.

Do not paste the entire stack trace into a chat first Find the bracketed error class and the Caused by line, and search for those. Pasting two hundred lines of Java frames buries the one line that matters.
Try it
  1. Deliberately misspell a column in a select and read the error.
  2. Point spark.read.csv at a file that does not exist.
  3. Run the script from a different folder than the one it lives in.
a suggestion list naming the right column; a path-not-found message; and a lesson that relative paths depend on where you launch from.

Spark next to pandas: when each is the right tool

Most beginners arrive from pandas, so it helps to say plainly how the two relate. pandas runs in one process on one machine and holds the whole table in memory; it is superb for data that fits comfortably in RAM, with a huge ecosystem and immediate feedback. Spark exists for the case where one machine is not enough, or where a job is slow enough that spreading it over many cores is worth the setup.

The practical rule for a student is simple. If your data is a few hundred megabytes and you are exploring, pandas is faster to start and easier to debug, because there is no session to launch and no plan to wait for. If your data is several gigabytes or more, arrives as many files, or has to be processed on a schedule by a team, Spark is the better fit. Spark also has a startup cost: even a tiny query takes a second or two longer than pandas because of the Java process and the planning step. Do not conclude from a ten-row test that Spark is slow; the overhead is fixed, and it is repaid only at scale.

There is a bridge between the two worlds. The Pandas API on Spark, imported as pyspark.pandas, lets you write pandas-style code that runs as Spark jobs; it needs the pyspark[pandas_on_spark] extra. And any small Spark result can be turned into a pandas DataFrame with toPandas(), which is the right way to plot a summary after Spark has shrunk a billion rows down to a hundred. The key word is small: aggregate first, convert second.

PYTHON
summary_small = by_country.limit(100).toPandas()   # safe: at most 100 rows reach the driver
print(summary_small.head())

Notice the order. The heavy lifting (filtering, joining, grouping) happens in Spark, close to the data, and only the tiny final table crosses into pandas. Reversing it, by converting first and aggregating after, throws away everything Spark is good at, and it is the usual cause of the out-of-memory errors listed in the previous section.

Try it
  1. Load orders.csv with pandas using pandas.read_csv and compute revenue per country.
  2. Compare the answer with your Spark by_country result.
identical numbers. Having both versions side by side makes the translation between the two APIs concrete, and gives you a trusted check for every Spark pipeline you write later.

Putting it all together

Here is one small end-to-end project that uses every piece. It reads the orders, cleans them, computes revenue per country and category, joins region names, and writes Parquet. Put it in a file called pipeline.py and run it with python pipeline.py.

pipeline.py
from pyspark.sql import SparkSession, functions as F
from pyspark.sql.types import (StructType, StructField, StringType,
                               IntegerType, DoubleType, DateType)

spark = (SparkSession.builder
         .appName("orders-pipeline")
         .master("local[*]")
         .config("spark.sql.shuffle.partitions", "8")
         .getOrCreate())

schema = StructType([
    StructField("order_id", IntegerType()),
    StructField("customer", StringType()),
    StructField("country", StringType()),
    StructField("category", StringType()),
    StructField("amount", DoubleType()),
    StructField("order_date", DateType()),
])

orders = spark.read.csv("orders.csv", header=True, schema=schema)

clean = (orders
    .dropna(subset=["order_id", "amount"])
    .filter(F.col("amount") > 0)
    .dropDuplicates(["order_id"]))

regions = spark.createDataFrame(
    [("Egypt", "North Africa"), ("Saudi Arabia", "Gulf"),
     ("UAE", "Gulf"), ("Jordan", "Levant")],
    schema=["country", "region"])

summary = (clean
    .join(regions, on="country", how="left")
    .groupBy("region", "country", "category")
    .agg(F.count("*").alias("orders"),
         F.round(F.sum("amount"), 2).alias("revenue"))
    .orderBy(F.desc("revenue")))

summary.show(truncate=False)
summary.write.mode("overwrite").parquet("out/summary")

print("rows written:", spark.read.parquet("out/summary").count())
spark.stop()

Walk through it as Spark does. The read and every transformation until summary.show build a plan without touching data. show is the first action: it triggers a job, and the plan contains one shuffle for the join (possibly a broadcast, which Spark chooses automatically for a small table like regions) and one for the group by. The write is a second job that recomputes the same plan, because we did not cache summary. The final line reads the Parquet back as a sanity check, a third job. You can open the web interface, if you add a pause before stop(), to see all three.

Notice the habits embedded in the file. The schema is explicit, so bad input fails loudly. Cleaning steps are separate and named, so you can inspect clean on its own. The aggregate columns use alias, so downstream readers get meaningful names. And the output goes to Parquet, not CSV, so types survive.

This is also the shape of production Spark jobs: read, clean, join, aggregate, write, with an orchestrator scheduling the script. When you reach that stage, tools like Airflow (/student-guides/airflow) or Dagster (/student-guides/dagster) run the script on a schedule, and the same code can be submitted to a cluster with spark-submit.

Try it
  1. Run pipeline.py and check the out/summary folder.
  2. Add a duplicate row and a row with a negative amount to orders.csv and confirm the cleaning removes them.
  3. Extend the pipeline with a second output: the top customer per country, written to out/top_customers.
a summary table sorted by revenue, Parquet files on disk, and unchanged row counts after you add bad rows, because the cleaning step discards them.

What you can now do, and what comes next

You can now install Spark and verify it, explain the driver, executors, partitions, transformations and actions in your own words, read CSV, JSON and Parquet data, apply an explicit schema, filter, select, add columns, group, aggregate and join, express the same logic in SQL, write results to Parquet, read the web interface to find the slow stage, and decode the common error messages. You also know the two traps that bite beginners hardest: pulling everything to the driver with collect(), and forgetting that Spark 4 has ANSI mode on.

The things we deliberately left for later are the ones that matter once your data and your team grow. Mid-level covers how Spark executes a query in detail, shuffles and partition sizing, broadcast joins and skew, caching rules, spark-submit against real clusters, Structured Streaming, MLlib, Spark Connect and testing. Senior covers running Spark on Kubernetes as a platform: security, cost, upgrades, multi-tenancy and incident playbooks.

Natural neighbours in this catalogue: store and version data with Delta Lake (/student-guides/delta-lake), track the models you train from Spark-built features with MLflow (/student-guides/mlflow), and package your jobs with Docker (/student-guides/docker) before running them on Kubernetes (/student-guides/kubernetes).

Your next practical step: take a dataset you genuinely care about, anything from a few hundred thousand to a few million rows, and rebuild one analysis you already do in pandas or a spreadsheet. Compare the results, then look at the web interface to see what Spark did. That comparison is the fastest way to build intuition.

Sources