This is part one of three. It covers everything you need to build real pipelines with GitLab CI/CD, not a teaser. By the end you can write a .gitlab-ci.yml from a blank file, run tests on every push, pass build output from one job to the next, speed up slow jobs with a cache, keep a password out of your repository, put a manual approval button in front of a deploy, and read a red pipeline well enough to fix it in minutes. Mid-level and Senior take the same topics further; nothing you learn here is thrown away.
Each section ends with a Try it task. Do them as you go. They take a few minutes each, and pipelines only start to make sense once you have watched your own job fail, read its log, and fix it.
One note on versions before we start. This guide was checked against GitLab 19.4, released on 17 September 2026, with GitLab Runner 19.4.1. GitLab ships a minor release every month and a major release every May, so a few things in older tutorials are now wrong. Wherever that matters, this guide says so and shows the current form.
What GitLab CI/CD is, and the problem it solves
CI/CD stands for continuous integration and continuous delivery (or deployment). The words sound heavier than the idea. Continuous integration means that every time someone pushes code, a machine automatically builds it and runs the tests, so that a broken change is noticed in minutes rather than at the end of the month. Continuous delivery means that the steps between "the tests passed" and "it is running for users" are also automated, so releasing is a button press rather than a ritual.
GitLab CI/CD is the part of GitLab that does this. You describe what should happen in a text file called .gitlab-ci.yml, at the root of your repository. GitLab reads that file on every push, creates a pipeline, and hands the individual steps to machines that execute them. The results (green ticks, red crosses, logs, test reports, downloadable files) appear in the same web page where you review the code.
To see why this exists, picture the world before it. A team of five shares a repository. Each developer runs the tests on their own laptop, more or less carefully, before pushing. Some forget. Some have a different Python version. Someone has a slightly stale database schema. A change that passes on one laptop fails on another, and a change that passes on all of them fails in production because production runs something else again. Releases are done by whoever remembers the steps, typing commands from a wiki page, and the wiki page is two versions out of date.
The first generation of fix was a dedicated build server, usually Jenkins, that watched the repository and ran a script. That worked, but the build server was a separate system with its own login, its own plug-ins, its own configuration living in a web interface, and its own person who understood it. When the server broke, nobody could release. When the configuration changed, nobody could say who changed it or why.
GitLab's approach is to make the pipeline part of the repository. The recipe lives in .gitlab-ci.yml, committed next to the code it builds. It is versioned, reviewed in merge requests, and diffed like any other file. When a pipeline breaks, git log on that file tells you who changed it and when. A new team member gets the whole automation setup by cloning the project. This is the same idea as keeping a Dockerfile next to your source, applied to the whole delivery process, and it is the single biggest reason teams adopt GitLab CI/CD.
What people use it for:
Automatic tests
Every push and every merge request runs the test suite on a clean machine, so "it passes here" stops being an opinion.
Building artifacts
Compile a package, build a container image, or generate a report, and keep the result attached to the pipeline that produced it.
Repeatable deploys
Deploy the exact commit that passed the tests, with a manual approval step where you want one.
Scheduled jobs
Nightly data refreshes, dependency checks, and model retraining runs without anyone remembering to start them.
For MLOps work specifically, the same machinery runs your data-validation checks, unit tests on feature code, model evaluation gates, and image builds for serving. A pipeline is simply an automated checklist that never gets tired and never skips a step.
You need very little to follow along. A free GitLab.com account, a web browser, and either a terminal with Git or the web editor that GitLab provides. If your employer runs a self-managed GitLab (common in regulated organisations in the Gulf and in Egypt, where data residency rules push teams to host their own), everything here applies to it as well, and the setup section explains the one difference.
- Write down every manual step someone performs between "I finished the code" and "it is live" on a project you know.
- Mark each step as check (run tests, lint, scan) or action (build, upload, deploy).
- Circle the steps that depend on one person remembering to do them.
The mental model: pipelines, stages, jobs and runners
GitLab CI/CD has four nouns. Everything else in this guide is a detail of one of them, so spend a few minutes on these before writing any YAML.
A job is the smallest unit. It is one named task with a list of shell commands to run, such as "install the dependencies and run the tests". A job has a name you choose (test, build-image, lint) and, at minimum, a script key listing the commands.
A stage is a named group of jobs that belong together and run in parallel. Stages run in order. By default the stages are build, then test, then deploy: all the jobs in build run at the same time, and only when they all succeed does GitLab start the jobs in test. If any job in a stage fails, the later stages do not start (you can change that, but it is the default, and it is usually what you want: there is no point deploying code whose tests failed).
A pipeline is the whole run: every stage and every job created for one event, such as a push to a branch or a merge request being opened. A pipeline belongs to one specific commit. You see it in the project under Build > Pipelines as a row with a status, and clicking it shows the stages as columns and the jobs as boxes inside them.
A runner is the machine (or container, or Kubernetes pod) that actually executes a job's commands. GitLab itself never runs your scripts. It decides what should run and when, then a program called GitLab Runner asks GitLab "is there a job for me?", receives the commands, runs them, streams the log back, and reports success or failure. This split matters, and it explains many beginner surprises: the job runs on a runner's machine, not on GitLab's web server and not on your laptop.
Two facts about how these pieces communicate will save you confusion later. First, the runner polls GitLab: it makes outbound HTTPS requests asking for work. GitLab never connects into the runner. That is why a runner can sit on a laptop, behind a corporate firewall, or inside a private network and still work, because it needs only outbound access. Second, GitLab reads your .gitlab-ci.yml and freezes the configuration when the pipeline is created. If you click "retry" on a failed job after fixing the YAML, the retry still uses the old configuration, because it belongs to the old pipeline. To test a config change you need a new pipeline, which a new commit gives you.
Jobs also start from a clean slate. Each job runs in a fresh environment, typically a fresh container, with your repository checked out at the pipeline's commit. Nothing installed or created by one job is visible to the next unless you explicitly pass it along with artifacts (files kept for later jobs, covered in their own section) or speed it up with a cache. Most "it worked in the previous job" confusion comes from forgetting this.
Here is the vocabulary in one table, to come back to:
| Term | What it is | Where you see it |
|---|---|---|
| Pipeline | One run of your whole recipe for one commit | Build > Pipelines |
| Stage | An ordered group of jobs that run in parallel | The columns in the pipeline graph |
| Job | A named task with a script |
A box in the graph, with its own log |
| Runner | The machine that executes a job | Settings > CI/CD > Runners |
| Executor | How a runner runs jobs (Docker, shell, Kubernetes) | Chosen when the runner is set up |
| Artifact | Files a job saves for later jobs or for download | The job page, "Browse" and "Download" |
| Cache | A best-effort store that speeds up repeated installs | Invisible, except in the log |
needs keyword lets a job start as soon as the specific jobs it depends on finish, regardless of stage. Mid-level covers it. For this guide, think in stages.
- Sketch a pipeline on paper for a project of yours: three stages, and two or three jobs in each.
- Under each job, write the single command you would run by hand to do it.
- For each stage, answer: if one job here fails, should the next stage still run?
.gitlab-ci.yml. Stages become the stages list, jobs become top-level keys, and the commands become their script.
Getting set up: an account, a project, and a runner
You do not install GitLab to learn GitLab CI/CD. The fastest route is a project on GitLab.com, the hosted service, where machines to run your jobs are already provided.
Create an account at gitlab.com, then create a new blank project (the New project button, then Create blank project). Give it a name such as ci-playground, leave it private, and tick the option to initialise the repository with a README so that it has a first commit.
Now the machines. GitLab.com offers GitLab-hosted runners: a fleet of virtual machines that GitLab operates and that every project can use without any setup. Each job gets a brand-new virtual machine, which is deleted when the job ends. A few facts worth knowing on day one:
- Jobs that do not ask for a specific kind of runner go to a small Linux machine with two virtual CPUs and 8 GB of memory (the runner tag is
saas-linux-small-amd64). Larger sizes exist on paid plans. - If your
.gitlab-ci.ymldoes not name an image, the job runs in a default container image (ruby:3.1). You will almost always name your own image, which is one of the first things the first pipeline does. - A single job can run for at most three hours on hosted runners.
- Hosted runners spend compute minutes, a monthly allowance that depends on your plan. A pipeline that runs for four minutes across two parallel jobs uses the minutes of both jobs, with a multiplier for larger machines.
- On the Free tier, the first time you run a pipeline on GitLab.com you may be asked for identity verification, a one-time check that exists to stop abuse of the free compute. Without it, jobs fail with
Identity verification is required in order to run CI jobs. Complete the verification the page asks for and re-run the pipeline.
If your organisation uses a self-managed GitLab (installed on its own servers), there is no hosted fleet. Someone has to provide runners, and that may be you. Runner installation is short and worth seeing once, even if you never need it, because it makes the word "runner" concrete. The modern order of operations is:
- In the project, open Settings > CI/CD > Runners and choose Create project runner. Give it a description and tags (tags are covered later). GitLab shows a runner authentication token that starts with
glrt-. Copy it; it is shown once. - Install the
gitlab-runnerprogram on a machine that is not the GitLab server itself. - Register it, which links that machine to the runner you created in step 1.
- Start the service.
On a Debian or Ubuntu machine that looks like this:
curl -L "https://packages.gitlab.com/install/repositories/runner/gitlab-runner/script.deb.sh" -o script.deb.sh
less script.deb.sh # read what you are about to run as root
sudo bash script.deb.sh
sudo apt install gitlab-runner
sudo gitlab-runner register --url https://gitlab.com --token glrt-YOUR_TOKEN_HERE
The register command asks which executor to use. The executor is how the runner runs a job. For a first runner, choose docker, which runs each job in a fresh container, and give a default image such as alpine:latest. The answers are written to a file called config.toml (at /etc/gitlab-runner/config.toml when run as root). Then confirm that everything is alive:
gitlab-runner --version
sudo gitlab-runner status
sudo gitlab-runner verify
verify should print Verifying runner... is alive, and in the GitLab web page the runner shows a green Online status. If you install on a Mac, Windows machine, inside Docker, or through Helm on Kubernetes, the commands differ slightly, and the official install pages (linked in Sources) cover each of them.
--registration-token. That workflow is deprecated and is scheduled for removal in GitLab 20.0 (May 2027). The current flow is the one above: create the runner in the web interface first, then register with the glrt- authentication token using --token. If you ever see Check registration token or 410 Gone - runner registration disallowed, that is the old flow hitting the new rules.
Two more rules. Keep the runner's version close to your GitLab's version: the docs recommend matching the major and minor numbers, and this guide's runner is 19.4.1 to match GitLab 19.4. And never install the runner on the GitLab server itself, because a job can run arbitrary commands, and you do not want that happening on the machine that stores your code.
- Create a blank private project with a README on GitLab.com.
- Open Settings > CI/CD > Runners and find the list of available runners. Note whether hosted (instance) runners are enabled for your project.
- If you use a self-managed GitLab with no runner, create a project runner and register one on a spare machine, then run
gitlab-runner verify.
Your first pipeline
The smallest possible pipeline is three lines. Create a file named .gitlab-ci.yml at the root of your project (the leading dot matters, and the name must match exactly) with this content:
hello:
script:
- echo "Hello from $CI_COMMIT_SHORT_SHA on $CI_COMMIT_BRANCH"
You can create it in the web editor by choosing New file from the repository page, or on your laptop with git add, git commit and git push. The moment the commit lands on the project, GitLab creates a pipeline. Open Build > Pipelines and you will see a row with a status badge moving from pending (waiting for a runner to pick it up) to running to passed. Click the badge, click the hello job, and read its log.
Every job log has the same shape, and learning to read it is the most useful early skill. From the top you will see lines about the runner starting, about the image being pulled and the repository being cloned (the job's setup), then a green $ echo "Hello from ..." line for each command of your script, followed by the command's output, and finally Job succeeded or ERROR: Job failed: exit code 1.
Let us read the file, since three lines hide a lot:
hello:is the job name. Any top-level key that is not a reserved keyword is a job. You choose the name.script:is the only required key of a job. It is a list, and each item is one shell command, run in order.$CI_COMMIT_SHORT_SHAand$CI_COMMIT_BRANCHare predefined variables. GitLab fills in dozens of these for every job (the commit, the branch, the project name, the pipeline id) without you declaring anything. They are how a pipeline knows what it is building.
Notice what you did not write: no stage, no image, no runner selection. GitLab used the default stage (test), the default image on hosted runners, and any available runner. Defaults get you started, but real pipelines state these things explicitly, so let us do that. Replace the file with a real pipeline for a small Python project. First the project files; create these at the root of the repository:
def greet(name: str) -> str:
"""Return a friendly greeting."""
if not name:
raise ValueError("name must not be empty")
return f"Hello, {name}!"
import pytest
from app import greet
def test_greet_returns_greeting():
assert greet("Aya") == "Hello, Aya!"
def test_greet_rejects_empty_name():
with pytest.raises(ValueError):
greet("")
pytest>=8,<9
And now the pipeline that tests it:
default:
image: python:3.12
stages:
- test
unit-tests:
stage: test
script:
- python --version
- pip install -r requirements.txt
- pytest -v
Commit and push. The default: block sets values that apply to every job; here it makes every job run inside the official python:3.12 container image, so you get a known Python regardless of what the runner's own machine has. stages: declares the stage names in order. The job joins the test stage, installs the dependencies, and runs pytest. In the log you will see pytest's own output, 2 passed, and a green job.
Now break it on purpose, because you will meet red pipelines constantly and it is better to meet your first one calmly. Edit app.py so that greet returns f"Hi, {name}!", commit, and push. The new pipeline fails: the job turns red, and the log ends with pytest's assertion diff and ERROR: Job failed: exit code 1. The reason is simple and worth memorising: a job fails when any command in its script exits with a non-zero status. Pytest exits with 1 when a test fails, the runner sees it, stops the job, and marks the pipeline failed. Nothing magical is happening: the same rule applies to every shell command, which is why a script that ends with a failing grep or a curl that returns 404 can also fail a job.
.gitlab-ci.yml with live validation, a visual graph of the stages and jobs, and a "Full configuration" tab showing the merged result. A typo in a keyword appears as an error before you commit. As of GitLab 19.4 the editor also validates keywords that previously only the backend checked.
- Add the four files above to a fresh project and push. Watch the pipeline go green.
- Change the expected string in
test_app.pyso a test fails, push again, and open the red job's log. - Find the line in the log that names the failing assertion, and the final line that reports the exit code. Fix the test and push once more.
Stages, parallel jobs, and what happens when one fails
A single job is useful; several jobs in stages are where pipelines earn their keep. Extend the project so it checks code style and builds something, not only tests:
default:
image: python:3.12
stages:
- build
- test
- deploy
build-package:
stage: build
script:
- mkdir -p dist
- cp app.py dist/
- echo "$CI_COMMIT_SHA" > dist/VERSION
lint:
stage: test
script:
- pip install ruff
- ruff check .
unit-tests:
stage: test
script:
- pip install -r requirements.txt
- pytest -v
The stages list defines the order: build, then test, then deploy. Every job names its stage with stage:. A job with no stage: lands in test, which is why the very first pipeline worked without one. Notice that nothing in this file uses the deploy stage yet. A stage with no jobs is simply skipped.
Open the pipeline graph after pushing. You will see three columns. build-package is alone in the first. lint and unit-tests sit together in the second, and they run at the same time, on separate runners or separate slots of a runner, because two jobs in one stage are independent by design. That parallelism is free speed: if linting takes one minute and tests take three, the stage takes three, not four. The third column is empty, so it does not appear.
The waiting rule is the part to internalise. The test stage starts only after every job in build has succeeded. If build-package fails, lint and unit-tests never run, and they show as grey "skipped" or not at all. This is what you want for a chain where later steps need the earlier ones to have worked. It is not what you want for a check that is merely informational, and there are two keywords for that.
allow_failure: true lets a job fail without failing the pipeline. The job shows an orange warning icon instead of a red cross, and later stages still run. Use it for checks you are introducing gradually, such as a new linter that currently complains about old code:
lint:
stage: test
allow_failure: true
script:
- pip install ruff
- ruff check .
when: controls whether a job runs based on what happened before it. Its values are on_success (the default: run only if everything earlier passed), on_failure (run only if something earlier failed, useful for a notification job), always (run regardless), manual (wait for a human to press a play button), delayed (run after a wait you specify with start_in), and never. A common use is cleanup that must happen even after a failure:
notify-failure:
stage: .post
when: on_failure
script:
- echo "The pipeline for $CI_COMMIT_SHORT_SHA failed. Check $CI_PIPELINE_URL"
That job uses a special stage, .post, which always runs last. Its sibling .pre always runs first. You do not declare them in stages; GitLab adds them. They are handy for one-off setup and teardown that does not belong to any of your own stages.
Two other job-level keys you will meet early. timeout: sets how long a job may run before GitLab kills it (for example timeout: 15 minutes); the project default is 60 minutes, and a job that hangs forever otherwise burns your compute allowance. retry: re-runs a job automatically when it fails, for example retry: 2, which is useful for flaky infrastructure problems (a package mirror timing out) and dangerous as a cover for flaky tests, because a retried failure that later passes hides a real bug.
allow_failure: true the pipeline passes even though the job failed. Teams add it "temporarily" and forget. When a job stays orange for weeks, either fix what it reports or delete the job. A check everyone has learned to ignore is worse than no check.
- Add the
build-package,lintandunit-testsjobs above to your project and push. - Watch the graph: confirm the two test-stage jobs run side by side.
- Change
build-packageso its script ends withexit 1, push, and observe that the test stage never starts. Then remove theexit 1.
Scripts, images and services: what a job really runs
A job is three things: an environment, some preparation, and commands. This section covers how you control each one.
The environment comes from the image: key. On a Docker-executor runner (the usual case, including hosted runners), every job runs inside a fresh container started from the image you name. image: python:3.12 means a Linux container with Python 3.12 installed, pulled from Docker Hub. This is what makes a pipeline reproducible: the job does not depend on what happens to be installed on the runner. Pick the image that has the tools your script needs. A Node project uses node:22, a Go project golang:1.23, a job that only runs shell commands can use a tiny alpine:3.20. If your script says command not found for a tool you expected, nine times in ten the image is wrong.
You can set image: per job or once under default:. A per-job value overrides the default, which is how a Python pipeline can still have one job that uses a Node image for a front-end check:
default:
image: python:3.12
frontend-lint:
image: node:22
script:
- npm ci
- npm run lint
image, services, cache, before_script and after_script at the very top of the file, outside any job. That global form is deprecated. It still works, but the current way is to put them under default:, as in the examples here. If you see the top-level form in an existing project, it is safe to move it under default:.
The preparation is what before_script is for. Commands listed there run before the script of each job, in the same shell, so anything they set up (an activated virtual environment, an exported variable) is visible to the script. A typical use is installing dependencies once for every job:
default:
image: python:3.12
before_script:
- python --version
- pip install -r requirements.txt
unit-tests:
script:
- pytest -v
after_script is its counterpart for cleanup. It runs after the main script even when that script failed, but in a separate shell, so it does not see variables you exported in script. Use it for things like printing diagnostics or stopping a helper process, not for passing values.
One sharp edge is how a job decides success. The runner runs your script lines in a shell that stops at the first command returning non-zero. That is the behaviour you want, and it is also why a script that looks fine can fail: a grep that finds nothing returns 1, and a test -f somefile for a missing file returns 1. If a command is allowed to fail, handle it explicitly in the shell (grep pattern file || true) rather than hoping.
For anything longer than a line or two, resist the urge to write a mini-program in YAML. A multi-line shell block is written with |:
report:
script:
- |
echo "Branch: $CI_COMMIT_BRANCH"
if [ -f dist/VERSION ]; then
echo "Version: $(cat dist/VERSION)"
else
echo "No version file"
fi
Once a script grows past a few lines, move it to a file in the repository (scripts/report.sh), make it executable, and call it from the job. Files can be linted and run on your laptop; YAML cannot.
Services are the third piece. A services: entry starts an extra container next to your job, reachable by hostname, which is how a test job gets a database without installing one:
integration-tests:
image: python:3.12
services:
- name: postgres:16
alias: db
variables:
POSTGRES_PASSWORD: example-only-not-a-real-secret
script:
- pip install -r requirements.txt
- pytest tests/integration -v
Inside the job, the database is at host db, port 5432. The service container is started before the script and thrown away after. A service that does not come up is one of the classic confusing failures, so if a job cannot reach its database, read the top of the log for the service's own startup messages.
python:3.12 is better than python:latest, because latest moves under you and a pipeline that passed yesterday can fail today with no code change. For stricter reproducibility pin a full patch version or an image digest. The price is updating it deliberately, which is the price of reproducibility.
- Move the
pip installline out of theunit-testsjob and into abefore_scriptunderdefault:. Confirm the pipeline still passes. - Add a job with
image: alpine:3.20whose script runspython --version. Read the failure and explain it. - Fix it by choosing a better image or by installing Python in the script, and note which fix you prefer.
python: not found error in the alpine job, which shows that the image, not the runner, defines what tools a job can use.
Variables: passing values into a pipeline
Hard-coding values into scripts is the fastest way to write a pipeline nobody can reuse. CI/CD variables are environment variables that GitLab gives to each job, and they come from several places. Knowing which place to use for what is most of the skill.
Predefined variables come from GitLab and describe the context. You have met a few. These are the ones worth memorising:
| Variable | What it holds |
|---|---|
CI_COMMIT_SHA |
The full commit hash |
CI_COMMIT_SHORT_SHA |
The first eight characters, good for image tags |
CI_COMMIT_BRANCH |
The branch name (empty in tag and some merge request pipelines) |
CI_COMMIT_TAG |
The tag name, only in tag pipelines |
CI_COMMIT_REF_NAME |
The branch or tag name |
CI_COMMIT_REF_SLUG |
The ref name made safe for URLs and container tags |
CI_DEFAULT_BRANCH |
The project's default branch, usually main |
CI_PIPELINE_SOURCE |
Why the pipeline exists: push, merge_request_event, schedule, web, api, trigger |
CI_PROJECT_DIR |
The directory where the repository is checked out |
CI_REGISTRY_IMAGE |
The path of the project's built-in container registry |
CI_MERGE_REQUEST_IID |
The merge request number, in merge request pipelines |
YAML variables you declare yourself in the file, either for all jobs or for one:
variables:
APP_ENV: "staging"
PIP_CACHE_DIR: "$CI_PROJECT_DIR/.cache/pip"
deploy-preview:
stage: deploy
variables:
APP_ENV: "preview" # overrides the global value for this job only
script:
- echo "Deploying to $APP_ENV"
A variable defined on a job overrides the same name defined globally. Use YAML variables for configuration that is not secret and that belongs with the code: a Python version, a cache directory, a feature toggle.
Project and group variables, set in the web interface under Settings > CI/CD > Variables, are for values that should not live in the repository: API keys, passwords, tokens. Add a variable there, give it a key such as DEPLOY_TOKEN, and every job in the project sees it as $DEPLOY_TOKEN. The Add variable dialog has options that matter:
- Masked hides the value in job logs, replacing it with
[MASKED]. To be maskable, the value must be a single line with no spaces and at least eight characters. As of GitLab 18.3, new variables in the interface default to masked visibility. - Masked and hidden goes further and hides the value in the settings page too, after you save it. Use it for secrets you never need to read back.
- Protected makes the variable available only to pipelines running on protected branches and tags (for example
main). A variable holding a production credential should be protected, so a pipeline on someone's experimental branch cannot read it. - Type: File writes the value to a temporary file and gives you the file's path in the variable. This is right for certificates and key files, which do not work well as environment variable text.
- Expand variable reference controls whether
$OTHER_VARinside the value is substituted. Since GitLab 18.6 it is off by default for variables set in the interface, which means a value containing a literal dollar sign is now kept as written. Turn it on only when you want substitution.
Two behaviours to know. Variables are read by the shell inside the job, so they are referred to as $NAME (or ${NAME} when followed by other characters, like ${NAME}_suffix). And when the same name exists in several places, there is a precedence order. You do not need the full list yet, only the practical version: a value set when you run a pipeline by hand beats a project variable, which beats a variable in the YAML file, which beats a predefined one. When a variable "has the wrong value", search every place it could be defined before you debug anything else.
env, base64-encodes the value, or echoes it one character at a time defeats it. Anyone who can change .gitlab-ci.yml on a branch that can read the variable can make the pipeline reveal it. That is why protected variables, and keeping real production secrets out of pipelines that run on untrusted branches, matter more than masking does.
- In Settings > CI/CD > Variables, add a variable named
GREETING_NAMEwith the valueMENA-community, masked. - Add a job whose script is
echo "Hello $GREETING_NAME", push, and read the log. - Now change the script to
echo "$GREETING_NAME" | sed 's/./& /g'and read the log again.
Hello [MASKED] from the first job, and your value spelled out letter by letter in the second, which is exactly why masking is a convenience and not a security boundary.
Rules: deciding when a job runs
So far every job runs on every push to every branch. Real projects need finer control: run the deploy only on the main branch, skip the heavy tests when only the README changed, run the review checks only on merge requests. The keyword for this is rules.
A rules: list is read top to bottom. For each job GitLab finds the first rule that matches, and that rule decides what happens. If no rule matches, the job is not added to the pipeline at all. A rule can have an if: condition, a when: outcome, and a few other keys.
deploy-production:
stage: deploy
script:
- echo "Deploying $CI_COMMIT_SHORT_SHA"
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
when: manual
- when: never
Read it aloud: if the branch is the default branch, add this job and require a manual click; otherwise never add it. The final - when: never makes the fallback explicit. The job appears only on main, and even there it waits for a human.
The expressions in if: have a particular syntax that trips beginners constantly, so learn the rules now:
- Variables are written unquoted, with a dollar sign, and compared with a quoted string:
$CI_COMMIT_BRANCH == "main". - The wrong forms are
"$CI_COMMIT_BRANCH" == "main",${CI_COMMIT_BRANCH} == "main"and$CI_COMMIT_BRANCH == main. All three are invalid or do not behave as you expect. =~and!~match a regular expression written between slashes:$CI_COMMIT_BRANCH =~ /^feature\//.&&,||and parentheses combine conditions:$CI_COMMIT_BRANCH == "main" && $CI_PIPELINE_SOURCE == "push".- A bare variable like
if: $CI_COMMIT_TAGis true when that variable is set and not empty.
Two more rule keys cover the common cases. changes: runs a job only when certain files changed, which saves compute:
docs-check:
script:
- echo "Checking docs"
rules:
- changes:
- docs/**/*
- README.md
And exists: runs a job only when a file is present. Both can be combined with if: in the same rule, and all keys in one rule must match together.
Why pipelines run twice, and how to stop it. This is the most common puzzle in a real project. When you push commits to a branch that has an open merge request, GitLab can create two pipelines: a branch pipeline (for the push) and a merge request pipeline (for the merge request). If your jobs use rules with no if: on the pipeline source, both pipelines run the same jobs, which doubles your compute and confuses everyone, and GitLab warns Job may allow multiple pipelines to run for a single action.
The standard fix is a workflow:rules block at the top of the file. It decides whether a pipeline is created at all, before any job rules apply:
workflow:
rules:
- if: $CI_PIPELINE_SOURCE == "merge_request_event"
- if: $CI_COMMIT_BRANCH && $CI_OPEN_MERGE_REQUESTS
when: never
- if: $CI_COMMIT_BRANCH
Read it as three rules. A merge request event creates a pipeline. A branch push where a merge request is already open creates nothing (because the merge request pipeline already covers it). Any other branch push creates a pipeline. The result: exactly one pipeline per action. Copy this block into every project you write; it is boring and correct.
If workflow:rules matches nothing, the merge request shows Pipeline filtered out by workflow rules. That is not a bug but your own rules excluding the run, and the fix is to adjust the block.
rules existed, pipelines used only: and except:. They are deprecated. They still work, so you will see them in older projects, but they cannot express what rules can, and mixing the two in one job is an error. Write rules in new code and convert old jobs when you touch them.
- Add a
deploy-productionjob with the rules above, push to a feature branch, and confirm the job does not appear in the graph. - Merge to
main(or push there) and confirm the job appears with a play button instead of running. - Add the
workflow:rulesblock, open a merge request from a branch, push a second commit, and confirm only one pipeline runs for it.
Artifacts: passing files between jobs, and keeping results
Remember that every job starts in a clean environment. If build-package creates a dist/ folder, the deploy job, running later in a different container and possibly on a different machine, will not see it. Artifacts are the mechanism for moving files forward. You list paths in a job's artifacts: key, the runner uploads them to GitLab when the job finishes, and jobs in later stages download them automatically before they start.
build-package:
stage: build
script:
- mkdir -p dist
- cp app.py dist/
- echo "$CI_COMMIT_SHA" > dist/VERSION
artifacts:
paths:
- dist/
expire_in: 1 week
deploy-production:
stage: deploy
script:
- cat dist/VERSION
- echo "Uploading dist/ to the server"
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
when: manual
The deploy-production job never built anything, yet dist/VERSION is there, because the artifacts of earlier-stage jobs are downloaded into the working directory first. You can see them in the web interface too: open the build-package job page and use Browse or Download on the right-hand side. That is how a pipeline delivers a compiled binary, a model file, or a report to a human.
The keys to know:
pathslists files and directories to keep, relative to the project directory. Wildcards work (reports/*.html).expire_insays how long GitLab keeps them:30 minutes,1 week,3 months. Without it, the instance default applies (30 days out of the box). Artifacts take up storage, so set a short expiry on anything bulky.whendecides when to upload. By default artifacts are saved only if the job succeeds.when: alwayssaves them after failures too, which is what you want for test reports, because the failed run is the one you need to read.namenames the download archive, andexcluderemoves files matching a pattern from the set.reportsis a special kind of artifact that GitLab understands and displays instead of just storing.
The last one is worth seeing, because it turns a wall of log text into a proper result. If your tests write a JUnit XML file, GitLab shows the individual test results in the pipeline page and in merge requests:
unit-tests:
stage: test
script:
- pip install -r requirements.txt
- pytest -v --junitxml=report.xml
artifacts:
when: always
reports:
junit: report.xml
expire_in: 1 week
After this job runs, the pipeline page gains a Tests tab listing each test with its status and time, and a failing test in a merge request is named right there without anyone opening the log. The when: always matters: without it, a failing run (the one with the interesting report) would upload nothing.
Artifacts have limits. The default maximum size of one artifact file is 100 MB on most instances, so they are not the place for multi-gigabyte datasets; store large files in object storage (S3, GCS, or a regional equivalent) and pass only a reference. Anything a user can download from an artifact is visible to everyone with access to the job, so do not put secrets in artifacts.
- Give
build-packagetheartifactsblock above and push. - Open the finished build job, click Browse, and open
dist/VERSION. Compare its content with the commit's hash. - Add the JUnit report to
unit-testsand look for the Tests tab on the pipeline page.
dist/ attached to the build job, and a Tests tab listing your two tests individually.
Cache: making repeated work faster
Every job installs its dependencies from scratch in a clean container. For a project with a hundred Python packages or a large node_modules, that can be most of the job's duration, and most of your compute bill. A cache lets jobs reuse the downloaded packages from earlier runs.
variables:
PIP_CACHE_DIR: "$CI_PROJECT_DIR/.cache/pip"
unit-tests:
stage: test
cache:
key:
files:
- requirements.txt
paths:
- .cache/pip
script:
- pip install -r requirements.txt
- pytest -v
Two things are happening. First, the variable PIP_CACHE_DIR tells pip to keep its downloaded packages inside the project directory, because the cache can only save paths that are inside the project. Second, the cache: block says which directory to save (paths) and under what name (key). The runner downloads the cache before the script and uploads it after. The second pipeline's pip install finds the packages already local, and the job gets faster.
The key is how GitLab decides which cache belongs to which job. A fixed string such as key: my-cache shares one cache everywhere, and a stale cache can linger forever. The pattern used here, key:files: with requirements.txt, builds the key from a hash of that file. When your dependencies change, the file changes, the key changes, and a fresh cache is built. When they do not, the same cache is reused. That is almost always the right behaviour, and it keeps the cache honest without anyone clearing it.
Some options you will eventually want:
policy:pull-pushis the default (download at the start, upload at the end).pullonly downloads, which suits jobs that only read the cache, andpushonly uploads.fallback_keys: other keys to try if the exact key has no cache yet, so a changed dependency list can start from a nearly-right cache.when:on_successby default;alwaysalso saves after a failure.
The most important thing to understand about a cache is that it is not guaranteed. The cache may be empty because this is the first run, because it expired, because the job ran on a different runner that does not share the cache storage, or because someone cleared it. Your job must work with an empty cache, only more slowly. If a job passes only when the cache is warm, you have found a bug, not a feature. A job can use up to four caches at once.
| Cache | Artifacts | |
|---|---|---|
| Purpose | Speed up repeated downloads | Hand results to later jobs or people |
| Guaranteed to exist | No, best effort | Yes, until they expire |
| Typical contents | Package downloads, build tool caches | Built packages, reports, model files |
| Who can download it | Only the runner, automatically | Anyone with access, from the job page |
| If it is missing | Job is slower | Later jobs fail |
- Add the cache block and
PIP_CACHE_DIRtounit-testsand push. In the log, find the lines about the cache: the first run reports that it cannot find one. - Push a trivial change (edit a comment). Find the lines that show the cache being restored, and compare the install time.
- Add a package to
requirements.txtand watch the key change and a new cache be built.
Runners and tags: choosing where a job runs
Each job needs a runner, and a project often has several: a hosted fleet, a team runner with special hardware, a cheap one for linting. Tags are how you match jobs to runners. A runner is given tags when it is created (docker, gpu, linux-arm64), and a job lists the tags it needs with tags:. A job runs only on a runner that has all of the job's tags.
train-smoke-test:
stage: test
tags:
- gpu
script:
- nvidia-smi
- python train.py --epochs 1 --dry-run
This job waits for a runner tagged gpu. A job with no tags: can run on any runner that is set to accept untagged jobs. On GitLab.com, untagged jobs go to the small Linux hosted runner, and tags such as saas-linux-medium-amd64 select bigger ones on plans that include them.
Runners also have a scope, which is who may use them. Instance runners serve every project on the GitLab server, group runners serve all projects in a group, and project runners serve only the projects you attach them to. Hosted runners on GitLab.com are instance runners. A runner can also be marked protected, in which case it runs only pipelines for protected branches and tags, and locked, in which case it stays with its current project.
The failure you will see first is this one, and it is shown as a clock-like "pending" job that never starts:
This job is stuck because you don't have any active runners that can run this job.
It means no online runner satisfies the job. The causes, in order of likelihood:
- A tag mismatch. The job says
tags: [gpu]and no runner hasgpu(or the tag is spelled differently). - The runner refuses untagged jobs and your job has no tags.
- The runner is offline or paused. Check Settings > CI/CD > Runners for a green Online badge.
- The runner is protected-only and your branch is not protected, or locked to another project.
Fix it from the Runners page: compare the job's tags with the runner's, and enable Run untagged jobs on the runner if that is what you intend. If nothing matches for an hour, GitLab fails the job with the reason stuck_pending_no_matching_runners.
The executor is the other runner property worth knowing by name. The actively maintained ones are Docker (each job in a container, the common choice), Kubernetes (each job in a pod) and the autoscaling executors that create cloud machines on demand. The older Shell, SSH, VirtualBox, Parallels and Custom executors are in maintenance mode, receiving security fixes only, and the Docker Machine executor is deprecated with removal in GitLab 20.0. As a beginner you only need to know that the executor decides what image: means: on a shell executor, the image: key does nothing, because there is no container.
.gitlab-ci.yml can run anything. Use container-based runners for code you do not fully trust.
- Add a job with
tags: [nonexistent-tag]and push. Watch it sit in pending with the stuck message. - Open Settings > CI/CD > Runners and write down the tags of each runner available to you.
- Change the job to use one of those tags, or remove
tags:, and push again.
Environments and deploying safely
A pipeline that only tests is half the job. The other half gets the tested code to where it runs. This guide does not teach a particular hosting platform (that depends on your employer), but the GitLab side of deployment is the same everywhere, and worth learning with a stand-in script.
An environment is a named target such as staging or production. Declaring one on a job tells GitLab that running this job is a deployment, and GitLab records every deployment: which commit, which pipeline, when, and by whom. The project then has an Environments page (under the Operate section of the sidebar) that shows what is currently deployed where, with the history, which turns "what version is in production?" from a question into a click.
deploy-staging:
stage: deploy
environment:
name: staging
url: https://staging.example.com
script:
- echo "Deploying $CI_COMMIT_SHORT_SHA to staging"
- ./scripts/deploy.sh staging
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
deploy-production:
stage: deploy
environment:
name: production
url: https://www.example.com
script:
- echo "Deploying $CI_COMMIT_SHORT_SHA to production"
- ./scripts/deploy.sh production
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
when: manual
Staging deploys automatically when code reaches main. Production waits for a person to press the play button (when: manual). This two-step shape is how most teams work: automatic delivery to a safe place, human approval for the real one. A manual job shows in the graph with a play icon and the pipeline is marked as passed-with-manual-pending rather than blocked, so it never holds up the rest of your work.
Two extra protections are worth knowing about even now. First, a protected environment restricts who is allowed to run deployments to it, which is a paid-tier feature; if someone outside the allowed list clicks play, they see You are not authorized to run this manual job. Second, the resource_group keyword stops two deployments to the same place from running at once, which prevents two quick merges from racing each other:
deploy-production:
resource_group: production
environment:
name: production
script:
- ./scripts/deploy.sh production
Jobs that share a resource group run one at a time; the second waits for the first to finish.
If your deploy needs a credential, this is where your earlier variables lesson pays off. Store the token as a masked, protected project variable, reference it as $DEPLOY_TOKEN inside scripts/deploy.sh, and the value never appears in the repository. Because it is protected, only pipelines on protected branches such as main can read it.
- Add the two deploy jobs above, with an
echoin place ofdeploy.sh, and push tomain. - Watch staging deploy by itself, and find the play button on the production job.
- Click play, then open the project's Environments page and read the deployment history for each environment.
Reusing configuration without copy-paste
Once you have a dozen jobs, the same lines appear everywhere. YAML and GitLab give you three tools to stop that. Learn them in this order.
Hidden jobs and extends. A job whose name starts with a dot (.python-job) is never run; it is a template. A real job copies it with extends::
.python-job:
image: python:3.12
cache:
key:
files:
- requirements.txt
paths:
- .cache/pip
before_script:
- pip install -r requirements.txt
unit-tests:
extends: .python-job
stage: test
script:
- pytest -v
type-check:
extends: .python-job
stage: test
script:
- pip install mypy
- mypy app.py
Both jobs inherit the image, cache and install step, and each adds its own script. Keys defined on the job win over the template's. The merge is shallow for most keys: if the job defines its own cache, it replaces the template's whole cache block rather than merging into it, which surprises people, so check the Full configuration tab in the pipeline editor to see the result.
include. As the file grows, or as several projects need the same setup, you can split the configuration into files and pull them in:
include:
- local: /ci/test-jobs.yml
- template: Jobs/SAST.gitlab-ci.yml
local pulls a file from the same repository (the path starts with / and is relative to the repository root), and template uses one of GitLab's own ready-made configurations, here a security scan. There are other forms: project (a file from another GitLab project at a chosen ref), remote (a URL) and component. A project can include at most 150 files. The errors to recognise are Local file ... does not exist!, which means a wrong path, and Project ... not found or access denied, which means a wrong project path or missing access.
Components and the CI/CD Catalog. The modern, versioned way to share pipeline pieces is a CI/CD component: a reusable chunk of configuration, published in a catalog, that you include by name and version and can configure with inputs. Teams use them to give every project the same lint, test, and build steps without copying YAML. They belong to Mid-level; for now it is enough to know they exist, so that when a colleague writes include: - component: ..., you know what you are looking at.
$VAR in rules:if and in scripts, $[[ inputs.name ]] in components, and, in very recent releases, ${{ job.inputs.name }} for job-level inputs. They are evaluated at different times and are not interchangeable. As a beginner, use $VAR everywhere and treat the other two as things you will meet later.
- Create a hidden
.python-jobwith the image and install step, and makeunit-testsand a newtype-checkjob extend it. - In the pipeline editor, open the Full configuration tab and confirm both jobs received the image and the install step.
- Move the two jobs into
ci/test-jobs.ymlandincludethat file from the main config.
Validating and debugging a pipeline
Sooner or later a pipeline will refuse to start, or a job will fail in a way the log does not explain. Debugging pipelines has a method, and using it beats staring.
Step one: is the configuration valid? Before any job runs, GitLab validates the YAML. If it is wrong, the pipeline is not created, and you see a red banner such as This GitLab CI configuration is invalid: jobs:test-job:artifacts config contains unknown keys: path. The message names the job and the exact problem (here path should have been paths). Use the pipeline editor for instant feedback, or check from a terminal with the official CLI, glab:
glab auth login
glab ci lint
glab ci lint --dry-run --ref main
glab ci lint sends your local .gitlab-ci.yml to GitLab for validation and prints any errors. The --dry-run form goes further and simulates pipeline creation, catching rules that would leave you with no jobs. Another error you will see is jobs config should contain at least one visible job, meaning the file has only hidden jobs or your rules excluded everything.
Step two: did the pipeline get created, and with which jobs? If the configuration is valid but the job you expected is missing, the cause is almost always rules or workflow:rules. Open the Full configuration tab, then re-read your conditions with the actual values of CI_PIPELINE_SOURCE and CI_COMMIT_BRANCH for this kind of pipeline in mind. A job restricted to branches will not appear in a tag pipeline or a merge request pipeline, and that surprises nearly everyone once.
Step three: did the job start? A job stuck in pending is a runner-matching problem, covered above. A job that starts and fails instantly is usually the image or the clone step, so read the first twenty lines of the log.
Step four: read the failing command. Scroll to the first red line, not the last. The last line is always ERROR: Job failed: exit code 1, which tells you only that something returned non-zero. The command above it, with its own output, is the evidence. GitLab collapses sections of the log; expand them. Timestamps on each line help you spot a step that took suspiciously long.
Step five: reproduce it. Much of the time you can run the same commands locally in the same image. For the Python example:
docker run --rm -it -v "$PWD":/work -w /work python:3.12 bash
pip install -r requirements.txt
pytest -v
If it fails here too, you have a code problem that the pipeline was right to catch. If it passes locally but fails in CI, compare environment variables, file permissions, and which files exist in a clean checkout (did you forget to commit a file?).
The command-line companions. With glab you can drive pipelines without the browser:
| Command | What it does |
|---|---|
glab ci status |
Show the state of the latest pipeline for your branch |
glab ci status --live |
Watch it update until it finishes |
glab ci view |
A terminal view of stages and jobs |
glab ci trace |
Stream a job's log to your terminal |
glab ci run -b main |
Start a new pipeline on a branch |
glab ci retry |
Retry failed jobs |
glab ci cancel |
Cancel a running pipeline |
CI_DEBUG_TRACE to "true" makes the runner print every command and every variable value, including secrets, into the job log. It is the last resort for a puzzling failure, and anyone who can read the log can read your secrets. Delete the log afterwards, and never leave it on.
- Introduce a typo: rename
paths:topath:underartifacts, and runglab ci lint(or open the pipeline editor). - Read the error and locate the job and key it names. Fix it.
- Run the failing job's commands in the
python:3.12container locally, as shown above, and compare with the job log.
Common errors and how to read them
Most failures repeat. Here are the ones beginners meet, with what the message means and the fix. Keep this table open during your first month.
| Message | What it means | Fix |
|---|---|---|
This job is stuck because you don't have any active runners that can run this job. |
No online runner has all the job's tags, or none accepts untagged jobs | Check tags, enable "Run untagged jobs", confirm the runner is online |
This GitLab CI configuration is invalid: ... unknown keys: ... |
A typo or bad indentation in the YAML | Use the pipeline editor or glab ci lint |
jobs config should contain at least one visible job |
Only hidden jobs, or every job was excluded by rules | Add a visible job or fix the rules |
Pipeline filtered out by workflow rules. |
Your workflow:rules matched nothing or matched when: never |
Adjust the block |
Job may allow multiple pipelines to run for a single action |
Job rules cause duplicate branch and merge request pipelines | Use the standard workflow:rules block |
ERROR: Job failed: exit code 1 |
A command in the script returned non-zero | Read the lines above it |
ERROR: Job failed: execution took longer than <timeout> seconds |
The job timeout was reached | Find the hang, or raise timeout: deliberately |
Job's log exceeded limit of 4194304 bytes. |
The job printed more than the runner's log limit | Reduce output noise |
Identity verification is required in order to run CI jobs |
A GitLab.com Free account using hosted runners | Complete verification, or use your own runner |
Local file ... does not exist! |
A wrong include:local path |
Start the path with / from the repository root |
Project ... not found or access denied |
A wrong include:project path or missing access |
Fix the path and permissions |
x509: certificate signed by unknown authority |
A self-hosted GitLab or proxy with a private certificate authority | Add the authority's certificate to the runner |
WARNING: Failed to pull image ... denied: requested access to the resource is denied |
The image is private, or the job token is not allowed to pull it | Authenticate to the registry or allowlist the project |
Cannot connect to the Docker daemon at tcp://docker:2375 |
A Docker-in-Docker job without the dind service or without a privileged runner | Covered at Mid-level; see the note below |
Two of these deserve a paragraph. The timeout message reports the limit that was hit, so compare it with your timeout: value, the project default of 60 minutes, and any runner maximum before deciding whether the job hung or simply needs longer. And the Docker error appears the moment you try to run docker build inside a job. Building container images in a pipeline needs a Docker daemon the job can reach, usually through a docker:dind service on a privileged runner, and the GitLab.com hosted runners support it. That setup is a Mid-level topic, but knowing the error means you will recognise it when you get there.
Finally, a small rule for reading any error: find the first error in the log, not the last, and ask whether it is about the configuration (before any job), the runner (the job never starts), the environment (the image or the tools), or the script (your commands). Four categories cover nearly everything.
- Pick three rows from the table and reproduce each deliberately: a missing include file, an impossible tag, and a script that calls a command that does not exist.
- For each, note which of the four categories (configuration, runner, environment, script) it belongs to.
- Fix each and confirm the pipeline turns green.
Keeping secrets out of your repository
Every beginner pushes a password to a repository once. In a pipeline the stakes are higher, because the credential is now usable by anything that can run a job. A few habits cover most of the risk.
Never commit secrets. Not in .gitlab-ci.yml, not in a script, not in a .env file, not in a test fixture. Git history is permanent: deleting a file in the next commit leaves the secret in the previous one. If a real secret reaches the repository, treat it as leaked and rotate it, meaning generate a new one and revoke the old, rather than trying to scrub history.
Put them in variables. Project or group CI/CD variables, masked and protected, are the beginner-level solution, as you saw above. Use File-type variables for certificates and key files.
Use the job token for GitLab itself. Inside a job, GitLab provides CI_JOB_TOKEN, a short-lived credential tied to that one job that can do a limited set of things (pull from the project's container registry, call some API endpoints, clone allowed projects). The built-in registry variables use it: CI_REGISTRY_USER and CI_REGISTRY_PASSWORD log you in to the project's registry. Prefer them to creating a personal access token that outlives the job. Since GitLab 19.0 the job token uses the JWT format by default, and one consequence is that wrapping it in base64 with line breaks corrupts it (use base64 -w0).
Do not print secrets. echo $DEPLOY_TOKEN hides the value only when it is masked, and even then a transformed value escapes. Write scripts that pass the variable to the tool that needs it and print nothing.
Mind who can run what. Protected variables reach only pipelines on protected branches. A merge request from a fork runs in the fork's context without your secrets unless a maintainer deliberately runs it in the parent project, and that is a safety feature: read a stranger's changes to .gitlab-ci.yml before approving a pipeline.
The best modern approach, which Mid and Senior levels cover, is to avoid stored secrets altogether: a job asks GitLab for a short-lived OpenID Connect token (id_tokens:) and exchanges it with your cloud provider or vault for temporary credentials. Nothing long-lived is stored, so nothing long-lived can leak.
Do
- Store credentials as masked, protected variables
- Use File type for certificates and keys
- Use the built-in registry variables and job token
- Rotate any secret that touches the repository
- Review pipeline changes from outside contributors
Do not
- Write a password in
.gitlab-ci.ymlor a script - Rely on masking as your only protection
- Give protected secrets to every branch
- Leave
CI_DEBUG_TRACEon - Use a personal token where the job token would do
- Make a variable named
DEPLOY_TOKENthat is masked and protected, with a dummy value of at least eight characters. - Run a job on a non-protected branch that prints whether it is set:
test -n "$DEPLOY_TOKEN" && echo set || echo unset. - Run the same job on
main.
unset on the feature branch and set on main, which shows protection doing real work without ever printing the secret.
Putting it all together
Let us assemble everything into one small, complete project: a Python package with tests, a build that produces an artifact, a test report, a cache, a staging deploy on main, and a manual production deploy. Start from the four files you already have (app.py, test_app.py, requirements.txt and, optionally, a scripts/deploy.sh that just echoes). Here is the full pipeline:
default:
image: python:3.12
interruptible: true
stages:
- build
- test
- deploy
variables:
PIP_CACHE_DIR: "$CI_PROJECT_DIR/.cache/pip"
workflow:
rules:
- if: $CI_PIPELINE_SOURCE == "merge_request_event"
- if: $CI_COMMIT_BRANCH && $CI_OPEN_MERGE_REQUESTS
when: never
- if: $CI_COMMIT_BRANCH
.python-cache:
cache:
key:
files:
- requirements.txt
paths:
- .cache/pip
build-package:
stage: build
script:
- mkdir -p dist
- cp app.py dist/
- echo "$CI_COMMIT_SHA" > dist/VERSION
artifacts:
paths:
- dist/
expire_in: 1 week
lint:
extends: .python-cache
stage: test
allow_failure: true
script:
- pip install ruff
- ruff check .
unit-tests:
extends: .python-cache
stage: test
script:
- pip install -r requirements.txt
- pytest -v --junitxml=report.xml
artifacts:
when: always
reports:
junit: report.xml
expire_in: 1 week
deploy-staging:
stage: deploy
environment:
name: staging
script:
- echo "Deploying version $(cat dist/VERSION) to staging"
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
deploy-production:
stage: deploy
resource_group: production
environment:
name: production
script:
- echo "Deploying version $(cat dist/VERSION) to production"
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
when: manual
Read it once from the top; every line is something you have now learned. The default: block sets the image and marks jobs interruptible, so that a newer push can cancel an older, now-pointless pipeline and save minutes. workflow:rules guarantees one pipeline per action. The hidden .python-cache job carries the cache, and two jobs extend it. build-package produces a dist/ artifact, which the deploy jobs read without building anything. The unit tests publish a JUnit report. Staging deploys on main; production waits for a click and is serialised with a resource group.
Now walk it end to end, as a teammate would:
- Create a branch, change
greet, and push. A branch pipeline runs: build, then lint and tests in parallel. No deploy jobs appear, because the branch is notmain. - Open a merge request. The workflow rules switch to a merge request pipeline, with the tests shown in the merge request's widget. Push a fix and confirm only one pipeline runs.
- Merge. A pipeline on
mainruns everything, anddeploy-stagingruns automatically. - Open the pipeline and click play on
deploy-production. Open the Environments page and read the history. - Break a test deliberately in a new branch and watch the merge request show the failing test by name.
Commit it, push it, and then change one thing at a time to see how each keyword behaves. Remove the cache and compare times. Delete the workflow block and watch the duplicate pipelines appear. Add allow_failure: false and see how the lint job then blocks. That experimentation is the real lesson, and it costs almost nothing.
If you want to take the image-building route next, the same project becomes a container pipeline once you add a Dockerfile; the Docker guide covers writing one, and the Mid-level part of this guide shows how to build and push it from a job. For deploying onto a cluster, continue with Kubernetes, and for the infrastructure underneath, Terraform.
- Build the full project above in a fresh repository, following the five-step walk-through.
- Make three deliberate mistakes (a failing test, a missing tag, a wrong include path) and diagnose each using the four-category method.
- Write one sentence for each job in the file explaining why it exists.
What you can now do, and what comes next
You can read and write a .gitlab-ci.yml from nothing. You know the four nouns (pipeline, stage, job, runner) and how they relate. You can run tests on every push, choose images and services, pass files between jobs with artifacts, speed up installs with a cache keyed on the lockfile, and keep a secret in a masked, protected variable instead of the repository. You can control when jobs run with rules and workflow:rules, put a manual approval in front of a deployment, record deployments in environments, reuse configuration with hidden jobs and extends, and diagnose a red pipeline by category. You also know which older idioms to avoid: registration tokens, only/except, and top-level image and cache.
The Mid-level guide picks up exactly here. It opens the pipeline and shows how GitLab decides what to run, then teaches the DAG with needs, parallel and matrices, downstream and dynamic pipelines, components and inputs, variable precedence in full, merge request pipelines and merge trains, Docker-in-Docker builds, OpenID Connect secrets, and running your own runners well. The Senior guide treats GitLab CI/CD as a platform you operate for a team: runner fleets and autoscaling, the security and trust model, multi-tenancy, cost, upgrades (with 20.0 in May 2027 as the next breaking-change window), and governance.
Before moving on, keep three habits. Validate before you push, using the pipeline editor or glab ci lint. Read the first error in a log, not the last. And treat a pipeline as code: small commits, reviewed changes, and no secret ever near the repository. Those three habits are worth more than any single keyword.
The Interview and Tips & Tricks companions for this level turn what you just learned into answers you can say out loud and mistakes you can avoid.
Sources
The content of this guide was checked against the official documentation for GitLab 19.4 and GitLab Runner 19.4.1.
- GitLab CI/CD documentation
- CI/CD YAML syntax reference
- Deprecated keywords
- Workflow keyword and rules
- Validate CI/CD configuration
- Job rules
- Job artifacts
- Troubleshooting jobs
- Debugging CI/CD pipelines
- CI/CD variables
- Predefined CI/CD variables
- Caching in GitLab CI/CD
- The pipeline editor
- Hosted runners on GitLab.com
- Linux hosted runners
- Runner scopes
- Creating a runner (new workflow)
- CI/CD job token
- Building Docker images with GitLab CI/CD
- GitLab Runner documentation
- Install GitLab Runner
- Registering runners
- GitLab Runner commands
- Runner executors
- Deprecations and removals
- GitLab 19.4 release notes
- glab CLI, CI commands