This is part one of three, and it covers everything you need to do real work with Terraform. By the end you can install Terraform, write a configuration from scratch, read a plan the way a reviewer does, understand what the state file is and why everyone is nervous about it, use variables, outputs and loops, and create, change and destroy infrastructure with confidence. You will also know how to read the error messages that stop most beginners. Mid-level and Senior take the same ideas into teams, pipelines and production. Nothing on this page is thrown away later.
The main examples run entirely on your laptop, with no cloud account and no bill. They use two small providers that manage local files and random names, so you can make every mistake on purpose and lose nothing. One section then shows the same workflow against a real cloud, so you can see that nothing changes except the provider.
Each section ends with a Try it task. Do them as you go. Terraform only clicks once you have watched your own plan say "1 to add", applied it, broken it and watched the next plan notice.
What Terraform is, and the problem it solves
Terraform is an infrastructure as code tool. You describe the infrastructure you want (servers, networks, storage buckets, databases, DNS records, Kubernetes clusters, even a file on disk) in plain text files. Terraform works out what has to change to make the real world match that description, shows you the list of changes, and then makes them by calling the relevant APIs for you.
That sounds abstract, so start with what came before it.
The first way people build cloud infrastructure is by clicking. You open the AWS, Azure or Google Cloud console, create a storage bucket, tick a few boxes, create a virtual machine, attach it to a network, and write down what you did (or don't). This is sometimes called ClickOps, and it works well for one person on day one. It works badly for everyone after that. Six months later, nobody remembers why the bucket has public access turned on. The staging environment was built by a different person on a different afternoon, so it differs from production in twenty small ways, and one of them causes an outage. When someone asks "can we build this again in another region?", the honest answer is "roughly, over a week, with some surprises".
The second way is scripting. You write a shell script or a Python script that calls the cloud's command-line tool: create the bucket, create the machine, attach the network. Scripts are repeatable, which is an improvement, but they are imperative: they describe steps, not an end result. Run the script twice and the second run fails because the bucket already exists, or worse, creates a second machine. To make the script safe to re-run you add "if it doesn't exist, create it" checks everywhere. Then someone wants to change the machine size, and you need "if it exists but is the wrong size, resize it" logic too. Before long the script is mostly checking and very little doing, and it still cannot tell you what it is about to change before it changes it.
Terraform takes the third approach. You write down the desired end state, and Terraform does the comparison for you:
Three consequences of that design explain most of what follows in this guide, so notice them now.
The infrastructure is described in files you commit. Your .tf files live in Git next to your application code. Every change to the infrastructure is a diff that someone can review before it happens, and git log tells you who opened port 22 to the internet and why. Teams that adopt Terraform usually say this is the real win: infrastructure changes become ordinary code review.
You see the change before it happens. Terraform's plan command prints exactly what it will create, modify and destroy. You read it, and only then apply it. No script gives you that, and it is the habit that separates careful engineers from people who make headlines.
Terraform keeps a record of what it manages. To know that the bucket in your file corresponds to a particular bucket in your account, Terraform writes a state file. State is the most important and most misunderstood part of Terraform, and it gets its own section below.
Terraform works with far more than one cloud. It talks to each platform through a plugin called a provider, and there are providers for AWS, Azure, Google Cloud, Oracle Cloud, Kubernetes, GitHub, Cloudflare, Datadog, and thousands of other services on the public Terraform Registry. The workflow you learn on this page is identical for all of them.
What people in machine learning and data teams use it for:
Storage for data and models
Buckets for raw data, features and model artifacts, with the same encryption, versioning and access rules in every environment.
Compute that appears and disappears
A GPU machine or a Kubernetes node pool for a training run, created from a file and destroyed when the run ends, so it stops costing money.
Identical environments
Dev, staging and production built from the same code with different inputs, so "it worked in staging" means something.
Reviewed access
IAM roles, service accounts and network rules as code, so every permission change goes through a pull request.
A word on the name and the licence, because you will hear both in interviews. Terraform is made by HashiCorp, which IBM acquired in 2025. Since version 1.6 in 2023, Terraform has been released under the Business Source License rather than an open-source licence. In response, the community created OpenTofu, an open-source fork now run by the Linux Foundation. OpenTofu started from the same code and uses the same language, but it is a separate project with its own releases and features. This guide is about Terraform. Most of the beginner material transfers directly to OpenTofu, but don't assume newer features match.
HashiCorp also sells a hosted service that runs Terraform for teams. It used to be called Terraform Cloud and is now called HCP Terraform, and there is a self-hosted version called Terraform Enterprise. You don't need either to learn Terraform. Everything on this page uses the free command-line tool, usually called the Terraform CLI.
- Pick one piece of cloud infrastructure you or your team created by clicking in a console (a bucket, a virtual machine, a database).
- Write down every setting you would need to recreate it exactly: name, region, size, access rules, encryption, tags.
- Ask yourself where that information lives today, and who would notice if someone changed one of those settings tomorrow.
Declarative versus imperative
Every interview about Terraform starts here, and it is also the idea that explains Terraform's behaviour when it surprises you, so it is worth a section on its own.
An imperative tool takes instructions: create this, then change that, then delete the other. You are responsible for knowing the current situation and choosing the right steps. A declarative tool takes a description of the result: there should be one bucket called churn-model-artifacts, with versioning on. The tool is responsible for finding out the current situation and choosing the steps.
Declarative (Terraform)
- You write what should exist
- Running it twice changes nothing the second time
- Deleting a block from the file deletes the thing
- Shows the change before making it
- The file is the documentation
Imperative (a shell script)
- You write the steps to get there
- Running it twice often fails or duplicates things
- Deleting a line does nothing to what already exists
- Makes the change as it reads each line
- You need separate notes on what exists
The property in the second row has a name worth knowing: idempotence. An operation is idempotent if doing it once and doing it ten times gives the same result. terraform apply is idempotent. When the real infrastructure already matches your files, a second apply reports "No changes" and touches nothing. This is what makes it safe to run Terraform often, including automatically in a pipeline.
The third row catches people out. In a declarative tool, removing something from the description is an instruction to remove it from the world. If you delete a resource block from your file and apply, Terraform destroys that resource. This is exactly the right behaviour, because the file says it should not exist. But it means that "tidying up" a Terraform file is a real change, and the plan is where you catch it.
Declarative does not mean Terraform can do anything you describe. It can only manage what its providers know how to create, read, update and delete, and some settings of some resources cannot be changed in place, so Terraform has to destroy the old object and create a new one. The plan tells you when that will happen, and you will learn to spot it below.
- Write a three-line shell script that creates a folder called
demowithmkdir demo, then run it twice. - Read the error the second run gives you.
- Rewrite it so it is safe to run twice, and count how many extra lines it took.
The core model: provider, resource, state, plan
Four nouns carry almost all of Terraform. Learn them precisely, because every error message and every interview question uses them.
| Noun | What it is | Where it lives | Example |
|---|---|---|---|
| Provider | A plugin that knows how to talk to one API | Downloaded into .terraform/ by terraform init |
hashicorp/aws, hashicorp/local |
| Resource | One thing you want to exist, described in a block | Your .tf files |
resource "aws_s3_bucket" "artifacts" |
| State | Terraform's record of what it created and the real IDs | terraform.tfstate, or a remote backend |
"aws_s3_bucket.artifacts is the bucket named churn-…" |
| Plan | The list of changes needed to make reality match the files | Printed to your terminal, or saved to a file | "1 to add, 0 to change, 0 to destroy" |
A provider is a separate program, downloaded from a registry, that translates Terraform's generic "create this resource" into the specific API calls a platform understands. Terraform itself knows nothing about S3 or virtual machines. The AWS provider does. Every provider has an address in the form namespace/type, such as hashicorp/aws. The full form includes the registry's hostname, registry.terraform.io/hashicorp/aws, and you will see that longer form in error messages.
A resource is one infrastructure object that Terraform manages. You write it as a block with two labels: the resource type, which comes from the provider, and a local name that you choose. The combination is the resource's address, which is how you refer to it everywhere else: aws_s3_bucket.artifacts, local_file.readme. The name is only used inside Terraform. It is not the name of the bucket in AWS.
State is a JSON file in which Terraform records every resource it manages, the real-world ID behind it, and the values of its attributes as of the last run. Without state, Terraform would have no way to know that aws_s3_bucket.artifacts in your file is the bucket it created last Tuesday rather than some other bucket with a similar name.
A plan is the result of comparing three things: your configuration (the .tf files, which say what should exist), the state (which says what Terraform created last time), and the real infrastructure (which the providers read fresh at the start of every plan, in a step called refresh). The differences become a list of actions: create, update in place, replace, or destroy.
Why it matters: every surprising plan is explained by one of the three inputs. Either the file changed, the state is not what you think, or someone changed the real thing by hand.
A few more words you will meet in the first hour, defined now so they don't trip you later:
- Configuration: all the
.tffiles in one directory, read together. Terraform merges them, and the order of files and blocks does not matter. - Module: any directory of
.tffiles. The directory you run Terraform in is the root module. Mid-level covers calling other modules from it. - Data source: a read-only lookup of something that already exists, written as a
datablock. It reads and never changes anything. - Apply: the command that carries out a plan.
- Drift: a difference between the state and reality, usually because someone changed something outside Terraform.
- HCL: the HashiCorp Configuration Language, the syntax
.tffiles are written in.
- Without looking back, write one sentence each for provider, resource, state and plan.
- Then write the three inputs a plan compares.
- Check your answers against the table and the diagram.
Installing Terraform and checking your setup
Terraform is a single binary with no runtime dependencies, so installation is mostly about getting that one file onto your PATH from a trustworthy source. HashiCorp publishes official packages for every common operating system.
macOS. Use HashiCorp's own Homebrew tap. The name matters:
brew tap hashicorp/tap
brew install hashicorp/tap/terraform
brew install terraform gives you an old version
The terraform formula in Homebrew's main collection stopped at 1.5.7, the last release before the licence change, because Homebrew's core collection only carries open-source licences. If terraform version prints 1.5.7, you installed the wrong one. Run brew uninstall terraform, then use the hashicorp/tap commands above.
Ubuntu and Debian. Add HashiCorp's signing key and package repository, then install with apt:
wget -O - https://apt.releases.hashicorp.com/gpg | sudo gpg --dearmor -o /usr/share/keyrings/hashicorp-archive-keyring.gpg
echo "deb [arch=$(dpkg --print-architecture) signed-by=/usr/share/keyrings/hashicorp-archive-keyring.gpg] https://apt.releases.hashicorp.com $(grep -oP '(?<=UBUNTU_CODENAME=).*' /etc/os-release || lsb_release -cs) main" | sudo tee /etc/apt/sources.list.d/hashicorp.list
sudo apt update && sudo apt install terraform
The first line downloads HashiCorp's public key and stores it where apt can use it to check signatures. The second line tells apt where the packages live and to trust only packages signed with that key. The third installs. It looks like a lot of ceremony, but it is what lets sudo apt upgrade keep Terraform current later.
Red Hat, CentOS, Fedora and Amazon Linux have equivalent repositories. On RHEL:
sudo yum install -y yum-utils
sudo yum-config-manager --add-repo https://rpm.releases.hashicorp.com/RHEL/hashicorp.repo
sudo yum -y install terraform
Windows. The official route is to download the zip file for your architecture from releases.hashicorp.com/terraform, extract terraform.exe into a folder such as C:\tools\terraform, and add that folder to your PATH under System Properties, Environment Variables. Community package managers also work (winget install Hashicorp.Terraform or choco install terraform), but HashiCorp does not maintain those packages, so check the version they give you.
Any system, pinned exactly. For a build server where you want one precise version, download the zip plus the SHA256SUMS file from https://releases.hashicorp.com/terraform/<version>/, check the checksum, and unzip the binary onto your PATH. Teams that need different versions per repository often use a version manager such as tenv or tfenv. These are community tools, not HashiCorp products.
Now check the result. Three commands tell you everything:
terraform version
terraform -help
terraform -install-autocomplete
terraform version should print something like this:
Terraform v1.16.4
on darwin_arm64
The first line is the Terraform release. The second is your operating system and processor architecture, here macOS on Apple Silicon. You'll see linux_amd64 on most Linux laptops and servers. If a newer release exists, Terraform adds a line telling you your version is out of date. That check contacts HashiCorp's servers, and you can turn it off by setting the environment variable CHECKPOINT_DISABLE=1.
terraform -help lists every subcommand, with the main ones first. terraform plan -help prints the options for one command, and you should reach for it before searching the web, because it always matches the version you have. The third command adds tab completion for Terraform subcommands to bash or zsh. Open a new terminal afterwards for it to take effect.
About versions. Terraform releases a new minor version (1.14, 1.15, 1.16) every three to four months, with small patch releases (1.16.1, 1.16.2) in between. All 1.x releases follow HashiCorp's compatibility promise: a configuration that works on 1.10 keeps working on 1.16. The reverse isn't guaranteed, though. A state file written by a newer Terraform may not be readable by an older one. So everyone who works on the same project should use the same minor version, and you will pin it in your configuration shortly.
- Install Terraform using the official method for your system.
- Run
terraform versionand confirm it is 1.16 or newer, and note your platform string. - Run
terraform plan -helpand find the options-out,-varand-destroyin the list.
darwin_arm64 or linux_amd64. If you see 1.5.7 on a Mac, you installed Homebrew's frozen formula. Swap it for the HashiCorp tap before going further.
Your first project, step by step
Your first project creates two things: a random project name, and a text file that mentions it. That is deliberately small. It uses two official HashiCorp providers, random and local, which need no account and cost nothing, and it still exercises the whole workflow you will use against a real cloud.
Make an empty folder and open it in your editor:
mkdir hello-terraform && cd hello-terraform
Create one file called main.tf:
terraform {
required_version = "~> 1.16"
required_providers {
local = {
source = "hashicorp/local"
version = "~> 2.9"
}
random = {
source = "hashicorp/random"
version = "~> 3.9"
}
}
}
resource "random_pet" "project" {
length = 2
}
resource "local_file" "readme" {
filename = "${path.module}/out/README.txt"
content = "Project ${random_pet.project.id} is managed by Terraform.\n"
}
output "project_name" {
value = random_pet.project.id
}
Read it top to bottom before running anything.
The terraform block holds settings for Terraform itself. required_version = "~> 1.16" says this configuration needs Terraform 1.16 or any later 1.x release, and Terraform refuses to run it on anything older. required_providers lists the plugins this configuration needs, where to download each one (source) and which versions are acceptable (version). The ~> operator is called the pessimistic constraint: ~> 2.9 means "2.9 or newer, but below 3.0", which accepts bug-fix and feature releases while refusing the next major version, where breaking changes live.
The first resource block asks the random provider for a random_pet, a readable random name such as lucky-rooster, made of two words. The second asks the local provider for a file. Its filename uses path.module, a built-in value meaning "the directory this configuration lives in", so the file lands in an out folder next to your code. Its content refers to the pet name with random_pet.project.id. That reference is doing two jobs. It inserts the value, and it tells Terraform that the file depends on the pet, so the pet must be created first.
The output block prints a value after Terraform runs, and makes it available to scripts. Here it prints the generated name.
Step 1: terraform init
terraform init
init prepares the working directory. It reads required_providers, downloads each provider, and records exactly which versions it chose. The output tells you precisely what happened:
Initializing the backend...
Initializing provider plugins...
- Finding hashicorp/local versions matching "~> 2.9"...
- Finding hashicorp/random versions matching "~> 3.9"...
- Installing hashicorp/local v2.9.1...
- Installed hashicorp/local v2.9.1 (signed by HashiCorp)
- Installing hashicorp/random v3.9.1...
- Installed hashicorp/random v3.9.1 (signed by HashiCorp)
Terraform has created a lock file .terraform.lock.hcl to record the provider
selections it made above. Include this file in your version control repository
so that Terraform can guarantee to make the same selections by default when
you run "terraform init" in the future.
Terraform has been successfully initialized!
"Initializing the backend" refers to where state will be stored. You have not configured anything, so it is the default, a local file. The "Finding … Installing … Installed" lines show each provider constraint being resolved to a concrete version and downloaded. "Signed by HashiCorp" means the download's signature was checked. Your version numbers will be newer if new releases have come out since this was written, which is exactly what the ~> constraints allow.
init created two things in your folder. The .terraform/ directory is a local cache holding the downloaded provider binaries. It can always be recreated, so it never goes into Git. The .terraform.lock.hcl file is the dependency lock file. It records the exact provider versions chosen and checksums of their packages, so that a teammate, or your CI pipeline, running init next week gets the same versions and not whatever is newest that day. It always goes into Git.
You run init once when you start, and again whenever you add a provider, change a version constraint, or change where state is stored. Running it again when nothing changed is harmless.
Step 2: terraform fmt and terraform validate
terraform fmt
terraform validate
fmt rewrites your files into Terraform's standard layout: two-space indentation and aligned equals signs. It prints the names of files it changed, and nothing if they were already tidy. Every Terraform codebase uses it, so diffs show real changes rather than whitespace.
validate checks that the configuration is internally consistent: correct syntax, known argument names, references that point to things that exist, values of the right type. It doesn't contact any cloud API, so it is fast and needs no credentials. You should see:
Success! The configuration is valid.
Step 3: terraform plan
terraform plan
This is the command you will run most often. It compares your configuration with the state (empty so far) and with reality, and prints what it would do:
Terraform used the selected providers to generate the following execution
plan. Resource actions are indicated with the following symbols:
+ create
Terraform will perform the following actions:
# local_file.readme will be created
+ resource "local_file" "readme" {
+ content = (known after apply)
+ content_base64sha256 = (known after apply)
+ directory_permission = "0777"
+ file_permission = "0777"
+ filename = "./out/README.txt"
+ id = (known after apply)
# (other checksum attributes omitted here)
}
# random_pet.project will be created
+ resource "random_pet" "project" {
+ id = (known after apply)
+ length = 2
+ separator = "-"
}
Plan: 2 to add, 0 to change, 0 to destroy.
Changes to Outputs:
+ project_name = (known after apply)
Every + means "will be created". The last line is the summary to read first on every plan you ever run: two to add, nothing changed, nothing destroyed. Some values say (known after apply). The pet's name does not exist until the random provider generates it during apply, so Terraform cannot show it yet, and the file's content depends on the name, so that is unknown too. The plan also shows default values you never wrote, such as separator = "-", because the provider fills them in.
A plan changes nothing. You can run it a hundred times.
Step 4: terraform apply
terraform apply
apply computes the plan again, prints it, and then stops and asks:
Do you want to perform these actions?
Terraform will perform the actions described above.
Only 'yes' will be accepted to approve.
Enter a value: yes
Only the full word yes is accepted. Anything else cancels. After you confirm, Terraform creates the resources in dependency order and reports each one:
random_pet.project: Creating...
random_pet.project: Creation complete after 0s [id=lucky-rooster]
local_file.readme: Creating...
local_file.readme: Creation complete after 0s [id=e3ca33a3f60921936c06ea699a33c806463ec814]
Apply complete! Resources: 2 added, 0 changed, 0 destroyed.
Outputs:
project_name = "lucky-rooster"
Notice the order: the pet first, then the file, because the file refers to the pet. The value in square brackets is each resource's ID, the identifier the provider uses to find the real object again. For a cloud resource this would be something like a bucket name or an instance ID.
Look at what you have now:
cat out/README.txt
ls
The file exists with a generated name inside it. There is also a new file in your folder, terraform.tfstate. That is the state, and the section after next is about it.
Step 5: run it again
terraform plan
No changes. Your infrastructure matches the configuration.
Terraform has compared your real infrastructure against your configuration
and found no differences, so no changes are needed.
This is idempotence in practice. The pet name does not change, and the file is not rewritten. Terraform read the state, checked the real file, compared both with your configuration, and found nothing to do. A second apply would say the same.
Step 6: clean up
terraform destroy
destroy plans the removal of everything this configuration manages, shows you the plan with - symbols, and asks for yes. It removes resources in the reverse of creation order, so the file goes before the name it depends on. It is the same as terraform apply -destroy. Run terraform apply again afterwards and you get a new pet name, because the old one was destroyed with everything else.
- Build the project exactly as above and run
init,fmt,validate,planandapply. - Run
terraform plana second time and read the "No changes" message. - Run
terraform output project_name, thenterraform output -raw project_name, and compare the two. - Run
terraform destroy, thenapplyagain, and note the new name.
-raw (the form scripts want). After destroy and apply you have a different name, because destroying really does remove the resource.
Reading a plan like a reviewer
A plan is the one moment when you can stop a mistake before it reaches the real world, so reading it well is the single most valuable Terraform skill. At work, the plan is also what your teammates review in a pull request. People who skim plans are the people who delete production databases.
Every resource in a plan starts with a symbol and a comment line saying what will happen to it:
| Symbol | Comment line says | Meaning | How worried to be |
|---|---|---|---|
+ |
will be created |
A new object will be made | Normal |
~ |
will be updated in-place |
Some settings change, the object stays | Read which settings |
-/+ |
must be replaced |
Destroyed and a new one created | Stop and understand why |
+/- |
must be replaced |
New one created first, then the old one destroyed | Same, slightly safer order |
- |
will be destroyed |
The object will be deleted | Stop and confirm it is intended |
<= |
will be read during apply |
A data source that can only be read later | Normal |
The replacement symbols deserve the most attention. Some settings of some resources cannot be changed on an existing object. A bucket cannot be renamed, and a virtual machine cannot move to another availability zone, for example. When you change one of those settings, the provider tells Terraform the only way to get there is to destroy the object and create a new one. For a random name or a text file that is harmless. For a database it means an empty database.
Here is what a replacement looks like. Change the file content in the first project by adding "Owner: data team." to the string, then run terraform plan:
Terraform will perform the following actions:
# local_file.readme must be replaced
-/+ resource "local_file" "readme" {
~ content = <<-EOT # forces replacement
- Project lucky-rooster is managed by Terraform.
+ Project lucky-rooster is managed by Terraform. Owner: data team.
EOT
~ content_md5 = "d7ce13a0cc8c9a04f01ef8bf1ae31421" -> (known after apply)
~ id = "e3ca33a3f60921936c06ea699a33c806463ec814" -> (known after apply)
# (3 unchanged attributes hidden)
}
Plan: 1 to add, 0 to change, 1 to destroy.
Read it in this order. First, the summary line: one to add and one to destroy means a replacement, not a change, even though you only edited text. Second, the resource comment: "must be replaced". Third, find the attribute marked # forces replacement. That is the reason. The local provider treats a file's content as unchangeable, so a new content means a new file. Fourth, the individual attribute changes: ~ marks a changed value, shown as old -> new, and a multi-line string shows removed lines with - and added lines with +. Unchanged attributes are hidden and counted, which keeps big plans readable.
The plan also ends with a note you will see every time:
Note: You didn't use the -out option to save this plan, so Terraform can't
guarantee to take exactly these actions if you run "terraform apply" now.
This is a real warning, not boilerplate. terraform apply on its own computes a fresh plan. If something changed between your plan and your apply (a teammate applied, or someone edited a setting in the console), the apply plan can differ from the one you reviewed. The fix is to save the plan to a file and apply that exact file:
terraform plan -out=tfplan
terraform apply tfplan
Applying a saved plan does not ask for confirmation, because the review happened when you read it. If the state changed after you saved it, Terraform refuses with Saved plan is stale, and you plan again. Treat the tfplan file as sensitive and keep it out of Git: it is binary, and it can contain secret values in readable form. Pipelines at work almost always use this two-step form.
Two more plan options are worth knowing on day one. terraform plan -destroy shows what a destroy would remove without doing it. terraform plan -detailed-exitcode changes the exit code so scripts can tell the outcomes apart: 0 means no changes, 1 means an error, and 2 means there are changes. That is how scheduled jobs detect drift.
- In your first project, change the text in
contentand runterraform plan. - Find the summary line, the "must be replaced" line and the
# forces replacementmarker. - Now change
length = 2tolength = 3on the pet and plan again. Count how many resources are replaced, and explain why the file is replaced too. - Save a plan with
-out=tfplan, apply it, and notice that nothing asks you to typeyes.
State: Terraform's memory
Open terraform.tfstate in your editor. It is JSON, and the interesting part looks like this (shortened):
{
"version": 4,
"terraform_version": "1.16.4",
"serial": 3,
"lineage": "8f1c0a52-…",
"outputs": {
"project_name": { "value": "lucky-rooster", "type": "string" }
},
"resources": [
{
"mode": "managed",
"type": "random_pet",
"name": "project",
"provider": "provider[\"registry.terraform.io/hashicorp/random\"]",
"instances": [
{ "attributes": { "id": "lucky-rooster", "length": 2, "separator": "-" } }
]
}
]
}
Each entry maps a resource address in your configuration (type random_pet, name project) to the real object (ID lucky-rooster) and records its attributes as of the last apply. serial goes up with every write. lineage is a unique ID given to this state when it was first created, which stops Terraform from mixing up two unrelated state files.
State exists because the configuration alone cannot answer the question "which real object is this block about?". Your file says there should be a bucket. Your AWS account has forty buckets. State is the link between the two, and it also lets Terraform notice that a block has been deleted from your files, since the resource is still in state but no longer in the configuration, which means it should be destroyed.
You should almost never read the JSON directly. Terraform has commands for it:
terraform state list # every resource address in state
terraform state show random_pet.project # one resource's recorded attributes
terraform show # the whole state, human-readable
$ terraform state list
local_file.readme
random_pet.project
$ terraform state show random_pet.project
# random_pet.project:
resource "random_pet" "project" {
id = "lucky-rooster"
length = 2
separator = "-"
}
state list is the fastest way to answer "what does this configuration manage?". Mid-level covers the commands that change state (moving and removing entries), which you should not need yet.
Drift: when reality changes behind Terraform's back
At the start of every plan Terraform refreshes: each provider reads the real objects and compares them with what state recorded. If someone changed something by hand, that difference is called drift, and the plan will propose to put things back the way your configuration says. Try it:
echo "edited by hand" > out/README.txt
terraform plan
random_pet.project: Refreshing state... [id=lucky-rooster]
local_file.readme: Refreshing state... [id=e3ca33a3f60921936c06ea699a33c806463ec814]
# local_file.readme will be created
+ resource "local_file" "readme" {
+ content = <<-EOT
Project lucky-rooster is managed by Terraform.
EOT
...
}
Plan: 1 to add, 0 to change, 0 to destroy.
The "Refreshing state" lines are the provider reading reality. The local provider noticed that the file no longer has the content it wrote, so it treats the old file as gone, and Terraform plans to write it again. A cloud provider would more often show a ~ update that reverts the changed setting. Either way, the rule is the same: your configuration wins. If the manual change was actually correct, the fix is to put it in the .tf file, not to keep changing things by hand.
The rules for state
Beginners get hurt by state in four predictable ways, so learn the rules before you need them.
State can contain secrets in plain text. Any attribute a provider returns is written to state, and that includes database passwords, private keys and access tokens. Marking a value sensitive hides it from the terminal but not from the state file. Treat the state file like a password file.
Never commit state to Git. It holds secrets, and two people with two copies of state in two branches will each believe different things about what exists. Add *.tfstate and *.tfstate.* to .gitignore on day one.
Never edit state by hand. If the JSON becomes invalid, or an ID points at the wrong object, Terraform will make confident, wrong decisions. When state really needs changing, there are commands for it, covered at Mid-level.
Local state does not work for a team. The file on your laptop is only on your laptop. For shared work, state is stored in a backend: a remote store such as an S3 bucket, an Azure storage account or a Google Cloud Storage bucket, or HCP Terraform. A good backend also provides locking, which stops two people applying at the same time and corrupting the state. You will set one up at Mid-level. For now, know that terraform.tfstate in your folder is the local backend, the default, and that it is fine for learning and for one-person experiments only.
terraform.tfstate.backup is
Every time the local backend writes new state, it keeps the previous version in terraform.tfstate.backup. It is a safety net for disasters, not a history. It only goes back one step, and it contains the same secrets as the main file.
- Run
terraform state listandterraform state show local_file.readme. - Edit
out/README.txtby hand, then runterraform planand find the "Refreshing state" lines. - Run
terraform applyand check that the file is back to the configured content. - Delete the
outfolder completely, then plan again.
HCL in one sitting
Terraform files are written in HCL, the HashiCorp Configuration Language. It is small. You can learn enough of it to read almost any configuration in one sitting, and this section covers that much.
Blocks are the containers. A block has a type, zero or more labels in quotes, and a body in braces:
resource "local_file" "readme" { # type "resource", two labels
filename = "out/README.txt" # an argument
}
The block types you will use as a beginner are terraform, provider, resource, data, variable, locals and output. Some blocks contain nested blocks, like the required_providers inside terraform or a validation inside a variable.
Arguments assign a value to a name: filename = "out/README.txt". Which argument names are allowed depends on the block. For resources, the provider's documentation on the Terraform Registry lists every argument, which ones are required, and which attributes the resource exports after creation (such as the id you used). Get into the habit of opening the resource's documentation page every time you use a resource type for the first time.
Values have types:
| Type | Example | Notes |
|---|---|---|
string |
"eu-central-1" |
Always double quotes |
number |
3, 0.5 |
No quotes |
bool |
true, false |
No quotes |
list(...) |
["raw", "models"] |
Ordered, can repeat |
set(...) |
toset(["raw", "models"]) |
Unordered, no repeats |
map(...) |
{ team = "data", env = "dev" } |
Keys to values of one type |
object({...}) |
{ name = "x", size = 2 } |
Named fields, each with its own type |
References read values from elsewhere in the configuration. You have already seen random_pet.project.id, which reads the id attribute of a resource. Variables are var.NAME, locals are local.NAME, and data sources are data.TYPE.NAME.ATTRIBUTE. There are also a few built-ins, such as path.module for the configuration's directory.
Interpolation puts a value inside a string with ${ }: "Project ${random_pet.project.id}". When the whole value is a single reference, write it without quotes: value = random_pet.project.id, not value = "${random_pet.project.id}". Both work, but reviewers and linters expect the first, and older examples that wrap every reference in "${ }" date from before Terraform 0.12.
Expressions and functions calculate values. The conditional expression is condition ? value_if_true : value_if_false, as in var.environment == "prod" ? 3 : 1. Terraform has a library of built-in functions, called like upper("dev"), join("-", ["a", "b"]), lower(var.name), jsonencode({...}), file("path") and merge(map1, map2). You cannot write your own functions in HCL. For experimenting, terraform console opens an interactive prompt where you can type any expression and see its value:
$ terraform console
> upper("churn-model")
"CHURN-MODEL"
> join("-", ["churn", "dev", "01"])
"churn-dev-01"
> "dev" == "prod" ? 3 : 1
1
> exit
Multi-line strings use the "heredoc" form, which you will use for file contents and scripts. With <<-EOT, the leading indentation is removed, so you can indent the text to match your code:
content = <<-EOT
# ${var.project}
Managed by Terraform.
EOT
Comments use # (preferred) or // for a single line and /* ... */ for several lines.
Files and layout. Terraform reads every .tf file in the directory and treats them as one configuration, so the split between files is only for humans. The convention most teams follow, and the one used in the capstone below:
| File | Holds |
|---|---|
versions.tf (or terraform.tf) |
The terraform block: required Terraform and provider versions |
providers.tf |
provider blocks and their settings |
main.tf |
Resources and data sources |
variables.tf |
variable blocks |
outputs.tf |
output blocks |
terraform.tfvars |
Values for the variables |
Terraform ignores files in subdirectories. It only reads the directory you run it in, so modules/network/main.tf is not part of your configuration until you call it as a module. Terraform also ignores files with other extensions, which is a classic trap: a file called main.tf.txt or variables.tf.bak is silently skipped.
- In your first project, run
terraform console. - Evaluate
random_pet.project.id,upper(random_pet.project.id),length(["raw","features","models"])andjsonencode({ env = "dev" }). - Move the
outputblock into a new file calledoutputs.tf, runterraform plan, and confirm nothing changes.
Variables, locals and outputs
Your first project hard-coded everything. That is fine for one experiment, but the point of infrastructure as code is to build the same thing several times with small differences: a dev copy and a prod copy, a bucket in one region and another in a second region. Three block types handle this, and they map neatly onto a function in any programming language. Input variables are the parameters, locals are the intermediate values you compute inside the function, and outputs are the return values.
Input variables
A variable block declares an input. By convention they live in variables.tf:
variable "project" {
type = string
description = "Short name of the ML project, used in every resource name."
}
variable "environment" {
type = string
description = "Which copy of the project this is."
default = "dev"
validation {
condition = contains(["dev", "staging", "prod"], var.environment)
error_message = "environment must be dev, staging or prod."
}
}
variable "data_folders" {
type = list(string)
description = "Top-level folders every project gets."
default = ["raw", "features", "models"]
}
Every argument inside is optional, but write type and description every time. The type makes Terraform reject a wrong value early, with a clear message, instead of failing halfway through an apply. The description is the documentation a teammate reads when they wonder what to pass. A variable with a default is optional for the caller. A variable without one is required, and Terraform will not plan until it has a value.
The validation block adds your own rule on top of the type. condition is any expression that must be true, and error_message is what the user sees when it is not. Validation runs during plan, before anything changes, so a typo like environment = "prdo" stops at the first command instead of creating a misnamed bucket.
You read a variable anywhere in the configuration as var.NAME: var.project, var.environment.
Giving variables values
There are five ways to set a variable, and when the same variable is set in more than one place, the later one in this list wins:
- An environment variable named
TF_VAR_plus the variable name, for exampleexport TF_VAR_environment=staging. - A file called
terraform.tfvarsin the working directory, loaded automatically. - A file called
terraform.tfvars.json, also loaded automatically. - Any files ending in
.auto.tfvarsor.auto.tfvars.json, loaded automatically in alphabetical order. - The
-varand-var-fileoptions on the command line, in the order you type them, the last one winning.
A .tfvars file contains only assignments, no blocks:
project = "churn"
environment = "dev"
Most teams keep one file per environment and choose it explicitly on the command line, which leaves no doubt about which copy you are touching:
terraform plan -var-file=prod.tfvars
terraform plan -var 'environment=staging'
If a required variable still has no value, Terraform asks for it interactively:
var.project
Short name of the ML project, used in every resource name.
Enter a value:
That prompt is convenient on a laptop and a problem in automation, where nobody is there to type. Pipelines pass -input=false, and then a missing value becomes an error instead of a prompt: Error: No value for required variable.
sensitive = true hides a value from your screen, not from the state
Marking a variable or output sensitive makes Terraform print (sensitive value) in plans and outputs instead of the value. That stops it leaking into terminal logs and CI output. The value is still written to the state file in plain text, and a saved plan file contains it too. Real secret handling (secret managers and values that are never stored) is Mid-level and Senior material. For now: keep secrets out of .tf and .tfvars files that go into Git, and treat state as secret.
Locals
A locals block names a value you compute once and use in several places. Locals cannot be set from outside, which is exactly what you want for derived values:
locals {
name_prefix = "${var.project}-${var.environment}"
common_tags = {
project = var.project
environment = var.environment
managed_by = "terraform"
}
}
You read them as local.name_prefix (note: the block is locals, plural, and the reference is local., singular, which trips everyone once). The rule of thumb is simple. If a value comes from the person running Terraform, it is a variable. If it is built from other values, it is a local. A configuration where the same "${var.project}-${var.environment}" string appears in ten places has a local waiting to be written.
Outputs
An output block publishes a value after apply. Outputs are printed at the end of terraform apply, can be read again later with terraform output, and are how scripts and other tools get information out of Terraform, such as a bucket name for a training job or an endpoint for a smoke test.
output "name_prefix" {
description = "Prefix used for every resource in this environment."
value = local.name_prefix
}
terraform output # all outputs, human-readable
terraform output name_prefix # one output, quoted
terraform output -raw name_prefix # one output, bare, for scripts
terraform output -json # everything, for tools like jq
Outputs also matter for a reason you will meet at Mid-level: when one configuration is called as a module by another, its outputs are the only values the caller can see.
- Add the
projectandenvironmentvariables to your first project, and change the file name to"${path.module}/out/${var.project}-${var.environment}.txt". - Run
terraform planwith no values and answer the prompt. Then createterraform.tfvarsand plan again. - Run
terraform plan -var 'environment=prdo'and read the validation error. - Run
TF_VAR_environment=staging terraform planwithenvironmentalso set interraform.tfvars, and work out which value won.
terraform.tfvars value beating the environment variable, because files come later in the precedence list. A -var on the command line would beat both.
Making several of something: count and for_each
Real configurations rarely need exactly one of each thing. An ML project needs a folder or bucket for raw data, one for features and one for models, and copying the same resource block three times is how bugs get in: somebody fixes the first copy and forgets the other two. Terraform has two meta-arguments (arguments that work on any resource, whatever the provider) for making several instances from one block.
count takes a whole number and makes that many instances, numbered from zero:
resource "random_pet" "worker" {
count = 3
length = 2
}
The instances are addressed as random_pet.worker[0], random_pet.worker[1] and random_pet.worker[2]. Inside the block, count.index gives the current number. count is ideal for "N identical things", and it has a second common use as an on/off switch: count = var.create_bucket ? 1 : 0.
for_each takes a map or a set of strings and makes one instance per element, addressed by key rather than number:
resource "local_file" "folder_readme" {
for_each = toset(var.data_folders)
filename = "${path.module}/out/${each.key}/README.md"
content = "# ${each.key}\n\nPart of ${local.name_prefix}.\n"
}
Inside the block, each.key is the current element (and for a map, each.value is its value). The instances are local_file.folder_readme["raw"], local_file.folder_readme["features"] and local_file.folder_readme["models"]. toset() is there because for_each refuses a plain list. A list can contain duplicates and has an order, and for_each needs unique, unordered keys.
Which one to use matters more than it looks, and the reason is what happens when the list changes:
for_each over names
- Instances are keyed
["raw"],["models"] - Remove "features" from the middle: one instance destroyed
- Everything else keeps its address
- The right default for named things
count over a list
- Instances are numbered
[0],[1],[2] - Remove the middle item: index 1 now means "models"
- Terraform changes or replaces everything after the gap
- Fine only for identical, interchangeable things
With count, the identity of each instance is its position. Take "features" out of a three-item list and "models" moves from position 2 to position 1, so Terraform sees position 1 change from "features" to "models" and position 2 disappear. With files that is a rewrite. With buckets full of data it is a replacement and a deletion. With for_each, removing "features" removes exactly ["features"] and nothing else moves.
In a shell, quote addresses that contain brackets and quotes, or the shell will mangle them: terraform state show 'local_file.folder_readme["raw"]'.
- Add the
data_foldersvariable and thefor_eachresource above, then apply and runterraform state list. - Remove
"features"from the default list and plan. Count the actions. - Put it back. Now rewrite the resource with
count = length(var.data_folders)andvar.data_folders[count.index], apply, remove"features"again, and plan.
for_each, one file destroyed and nothing else touched. With count, the "models" file is replaced at index 1 and index 2 is destroyed: more changes for the same edit. That difference is harmless here and expensive in a real account.
Data sources, references and the dependency graph
So far every value came from your own configuration. Real infrastructure has to fit into what already exists: a network another team built, the ID of the latest machine image, the account you are logged in to. A data source is a read-only lookup for exactly that. It is written as a data block, it never creates or changes anything, and its result is read during plan (or during apply, if it depends on something not yet created).
The local provider has a data source that reads a file, which keeps the example offline:
data "local_file" "team" {
filename = "${path.module}/team.txt"
}
resource "local_file" "owners" {
filename = "${path.module}/out/OWNERS"
content = "Owned by: ${trimspace(data.local_file.team.content)}\n"
}
A data source is referenced with the data. prefix: data.TYPE.NAME.ATTRIBUTE. In a cloud configuration the same pattern looks like data "aws_caller_identity" "me" {}, which tells you which AWS account your credentials belong to, or a lookup of an existing network by its tags. The provider's documentation lists data sources separately from resources, under "Data Sources".
How Terraform decides the order
You never tell Terraform what order to do things in. It works it out from references. Whenever one block mentions another (random_pet.project.id inside the file's content, data.local_file.team.content inside another file), Terraform draws an edge between them in a dependency graph. It then creates things in an order that respects every edge, destroys them in the reverse order, and runs anything without an edge between them in parallel, ten operations at a time by default.
This is why the order of blocks in your files does not matter, and why you should never try to control ordering by rearranging them. References are the ordering mechanism.
Very occasionally, one thing must wait for another without any value flowing between them. The classic cloud example is a server that needs an access policy attached before it starts, where nothing in the server's arguments mentions the policy. For that case there is depends_on:
resource "local_file" "done_marker" {
filename = "${path.module}/out/.done"
content = "ready\n"
depends_on = [local_file.folder_readme]
}
Use it sparingly. A reference is always better when one exists, because it documents why the dependency is there. An unnecessary depends_on slows plans down and can make a data source wait until apply to be read, which turns known values in your plan into (known after apply).
If you want to see the graph, terraform graph prints it in the DOT format used by Graphviz, and since Terraform 1.16 terraform graph -format=mermaid prints a Mermaid diagram you can paste into a Markdown file on GitHub.
- Create
team.txtcontaining your team name, add the data source and theownersfile, and apply. - Change the text in
team.txt(not the.tffiles) and plan. - Delete
team.txtand plan again, and read the error. - Run
terraform graph -format=mermaidand find the edge from the data source to the file.
The same workflow against a real cloud
Everything so far ran on your laptop. This section shows that moving to a real cloud changes the provider and the credentials, and nothing else. The example creates a private, versioned storage bucket for model artifacts on AWS. It costs almost nothing while empty, but it does need an AWS account, and you should destroy it at the end. If you don't have an account, read it anyway: the shape is identical on Azure, Google Cloud and Oracle Cloud.
Credentials first
A provider needs to authenticate to its API, and credentials never go in .tf files. The AWS provider looks for them in the same places the AWS command-line tool does: environment variables, the shared profile files in ~/.aws/, or a single sign-on session. The simplest setup on a laptop is a named profile:
aws configure sso # or: aws configure, for access keys
export AWS_PROFILE=my-sandbox
aws sts get-caller-identity
If the last command prints your account ID, Terraform will find the same credentials. If it fails, Terraform will too, with this message:
Error: configuring Terraform AWS Provider: no valid credential sources for Terraform AWS Provider found.
That error always means the provider found no credentials in its environment, never that your configuration is wrong. Fix the shell, not the code.
The configuration
terraform {
required_version = "~> 1.16"
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 6.0"
}
random = {
source = "hashicorp/random"
version = "~> 3.9"
}
}
}
provider "aws" {
region = var.region
default_tags {
tags = {
project = "churn"
managed_by = "terraform"
}
}
}
variable "region" {
type = string
description = "AWS region for the artifact bucket."
default = "me-central-1"
}
resource "random_pet" "suffix" {
length = 2
}
resource "aws_s3_bucket" "artifacts" {
bucket = "churn-artifacts-${random_pet.suffix.id}"
}
resource "aws_s3_bucket_versioning" "artifacts" {
bucket = aws_s3_bucket.artifacts.id
versioning_configuration {
status = "Enabled"
}
}
resource "aws_s3_bucket_public_access_block" "artifacts" {
bucket = aws_s3_bucket.artifacts.id
block_public_acls = true
block_public_policy = true
ignore_public_acls = true
restrict_public_buckets = true
}
output "bucket_name" {
value = aws_s3_bucket.artifacts.bucket
}
Three things are new. The provider "aws" block configures the provider: which region to work in, and default_tags, which the AWS provider adds to every resource that supports tags, so nobody forgets the ownership tags. Second, bucket settings such as versioning and public-access blocking are separate resources in version 4 and later of the AWS provider, each pointing at the bucket by its ID. That is a provider design choice, and it is why reading the provider documentation beats guessing. Third, the random suffix exists because S3 bucket names are global across every AWS account in the world, and churn-artifacts is certainly taken.
The workflow is the one you already know: terraform init (which now downloads the AWS provider, a much bigger binary), terraform plan, read the plan, terraform apply, and finally terraform destroy.
me-central-1 is AWS's UAE region and me-south-1 is Bahrain. Both are opt-in regions: an account administrator has to enable them before any API call works, and until then you will see authorization errors that look like credential problems. Many employers in the Gulf and in Egypt have data-residency rules for customer or health data, so the region is often a compliance decision rather than a latency one, and making it a variable (as above) keeps that decision visible in review. The other clouds have regions in the area too, for example Azure's UAE North and Google Cloud's Doha and Dammam regions, and the provider block is where you choose them.
What changes and what doesn't
Compared with the local project, you changed the provider and added credentials. The commands, the plan symbols, the state file, variables, for_each and the dependency graph are all the same. The state file now holds the bucket's name and ARN, and for databases it would hold connection details and sometimes passwords, which is why the state rules earlier in this guide matter more from here on.
Destroy is also more serious. terraform destroy on this configuration fails if the bucket still contains objects, because S3 refuses to delete a non-empty bucket. That refusal is a feature. The AWS provider has a force_destroy argument that empties the bucket first, and you should not set it on anything holding data you care about.
- If you have an AWS sandbox account, confirm
aws sts get-caller-identityworks, then init, plan and apply the configuration above (use a region your account has enabled). - Find the bucket in the S3 console and check its tags: they came from
default_tags. - Turn versioning off by hand in the console, then run
terraform planand read the drift. - Run
terraform destroyand confirm the bucket is gone.
~ update back to Enabled. Nothing about the workflow changed except where the resources live.
The everyday commands, grouped by what you are doing
You have now used most of the commands a beginner needs. Here they are in one place, grouped by the question you are trying to answer. Every one accepts -help.
| You want to… | Command | Notes |
|---|---|---|
| Prepare a new or changed folder | terraform init |
After cloning, adding a provider, or changing versions or backend |
| Accept newer provider versions | terraform init -upgrade |
Updates the lock file within your constraints |
| Tidy the formatting | terraform fmt |
-recursive for subfolders, -check in CI |
| Catch mistakes without any API calls | terraform validate |
Needs init first |
| See what would change | terraform plan |
-var, -var-file, -out=tfplan |
| Make the changes | terraform apply |
Or terraform apply tfplan for a saved plan |
| Rebuild one thing on purpose | terraform apply -replace=ADDR |
The modern replacement for taint |
| Update state to match reality, change nothing | terraform apply -refresh-only |
Replaces the old terraform refresh |
| Remove everything | terraform destroy |
Same as terraform apply -destroy |
| Preview a destroy | terraform plan -destroy |
Changes nothing |
| Read outputs | terraform output |
-raw NAME or -json for scripts |
| List what is managed | terraform state list |
Addresses only |
| Inspect one resource | terraform state show ADDR |
Recorded attributes |
| Inspect everything | terraform show |
State, or a saved plan with terraform show tfplan |
| Try an expression | terraform console |
Uses real values from state |
| Check versions | terraform version |
Also lists provider versions |
| Run from another folder | terraform -chdir=envs/dev plan |
-chdir goes before the subcommand |
Two of those rows replace older commands you will still see in blog posts. terraform taint marked a resource to be rebuilt on the next apply; today you pass -replace to plan or apply, which shows you the replacement in the plan before it happens. terraform refresh updated state from reality without a review step; today terraform apply -refresh-only shows you what it would record and asks first. Both old commands still exist, but the new forms are safer.
A safe daily rhythm is short enough to memorise: edit, fmt, validate, plan, read, apply. The cheapest commands come first because they fail fastest.
- EditChange the
.tffiles, never the real infrastructure. - fmt and validateSeconds, offline. They catch syntax errors and typos before any API call.
- planRead the summary line first, then every
-and-/+. - applyIdeally the saved plan you just read, so what runs is what you reviewed.
- CommitThe
.tffiles and the lock file, never the state.
- In your project, run
terraform plan -detailed-exitcode, thenecho $?. - Make a change, and repeat.
- Run
terraform apply -replace=random_pet.projectand read how the plan marks the resource.
-replace plan marks the pet with -/+ and a note that it is being replaced "as requested", and the file follows because it refers to the pet.
Configuration: versions, the lock file and your repository
A Terraform project that works on your laptop has to work identically for a teammate and a pipeline. Three things make that happen, and all three should be in place before your first commit.
Pin versions with constraints. required_version pins the Terraform CLI and required_providers pins each provider. Use the pessimistic operator on both: ~> 1.16 for Terraform, and ~> 6.0 or similar for a provider. The operators are =, !=, >, >=, <, <= and ~>. The rule for ~> is that only the rightmost number you wrote may increase: ~> 1.16 allows 1.17 and 1.20 but not 2.0, while ~> 1.16.0 allows 1.16.9 but not 1.17. Leaving a provider unpinned means the next init -upgrade may bring a new major version with breaking changes.
Commit the lock file. Constraints say what is acceptable. .terraform.lock.hcl says what was actually chosen, with checksums of the provider packages. With the lock file in Git, terraform init on another machine installs exactly the same versions and refuses a package whose checksum does not match. You change it on purpose with terraform init -upgrade, and review the diff like any other code.
Ignore the right files. Add this .gitignore before the first commit:
# Local provider cache, recreated by terraform init
.terraform/
# State and its backups: may contain secrets
*.tfstate
*.tfstate.*
# Saved plans: binary, may contain secrets
tfplan
*.tfplan
# Crash logs
crash.log
crash.*.log
# Personal overrides
override.tf
override.tf.json
*_override.tf
*_override.tf.json
Notice what is not ignored: .terraform.lock.hcl, and your .tfvars files. Whether to commit .tfvars depends on what is in them. Environment settings such as region and instance sizes belong in Git, because they are part of the reviewed configuration. Anything secret does not belong in a .tfvars file at all.
Environment variables Terraform reads are worth knowing because pipelines use them:
| Variable | Effect |
|---|---|
TF_VAR_name |
Sets input variable name |
TF_LOG |
Turns on logging: TRACE, DEBUG, INFO, WARN, ERROR or JSON |
TF_LOG_PATH |
Writes the log to a file instead of the terminal |
TF_INPUT |
Set to 0 or false to disable interactive prompts, like -input=false |
TF_IN_AUTOMATION |
Makes output assume no human is reading, for CI |
TF_PLUGIN_CACHE_DIR |
Shares downloaded providers between projects instead of re-downloading |
CHECKPOINT_DISABLE |
Turns off the "new version available" check |
.terraform/, and the AWS provider alone is several hundred megabytes. Setting TF_PLUGIN_CACHE_DIR to a folder such as $HOME/.terraform.d/plugin-cache (create it first) lets every project on your machine reuse one download.
- Open
.terraform.lock.hcland find the version and theh1:hash for therandomprovider. - Change the constraint to
version = "~> 2.0", runterraform init, and read the error. - Put it back, run
git init, add the.gitignoreabove, and rungit status.
init, which is the lock file doing its job. git status lists your .tf files and the lock file, and not the state or .terraform/.
Common errors, and how to read them
Terraform's error messages are long, and that is good news: they almost always say what is wrong, where, and often what to do. Read them in a fixed order. The first line after Error: is the category. The on main.tf line 12 line is the location, with the offending line quoted and underlined. The paragraph after that is the explanation. Beginners scroll to the bottom looking for the answer, when the answer is in the first three lines.
Error: Reference to undeclared input variable
on main.tf line 12, in resource "local_file" "readme":
12: filename = "${path.module}/out/${var.projct}.txt"
An input variable with the name "projct" has not been declared. Did you mean "project"?
These are the ones you will meet in your first weeks:
| Message | What it means | Fix |
|---|---|---|
Reference to undeclared resource or input variable |
A typo, or a block that doesn't exist in this directory | Correct the name or declare it. Read the "Did you mean" hint |
Unsupported argument … An argument named "x" is not expected here. |
A typo, or an argument that doesn't exist in this provider version | Check the resource's page in the provider docs for your version |
Missing required argument |
A required argument is absent | Add it. The docs mark required arguments |
No value for required variable |
A variable with no default got no value, and prompts are off | Pass -var, a .tfvars file or TF_VAR_ |
Inconsistent dependency lock file |
The lock file doesn't match the providers the configuration asks for | terraform init, or init -upgrade if you changed a constraint |
Failed to query available provider packages |
Terraform couldn't reach the registry, or the source address is wrong | Check the network or proxy, then the source spelling |
Unsupported Terraform Core version |
required_version excludes the CLI you are running |
Install a matching version |
Invalid for_each argument … must be a map, or set of strings |
You passed a list | Wrap it in toset() |
Invalid for_each argument … cannot be determined until apply |
The keys depend on values that only exist after apply | Build keys from variables or fixed strings, not resource attributes |
Cycle: A, B |
A refers to B and B refers to A | Break one of the references |
Error acquiring the state lock |
Another run is using the state, or a crashed run left a lock | Wait for it. Mid-level covers force-unlock for stale locks |
Saved plan is stale |
State changed after plan -out |
Plan again |
no valid credential sources for Terraform AWS Provider found |
No cloud credentials in this shell | Log in, or set AWS_PROFILE |
When the message itself isn't enough, turn on logging for one run and read what the provider actually asked the API:
TF_LOG=DEBUG TF_LOG_PATH=./tf.log terraform plan
The log is long and noisy. Search it for error and for the resource address. Delete it afterwards, because debug logs can contain secrets.
apply creates three resources and fails on the fourth, the three stay created and are recorded in state. Terraform does not undo them. Fix the cause and run apply again, and the next plan picks up where it stopped. Resources that failed partway through creation are marked tainted in state, and the next plan replaces them.
- Misspell a variable reference, run
terraform validate, and find the file, the line and the "Did you mean" hint. - Change
for_each = toset(var.data_folders)tofor_each = var.data_foldersand plan. - Make two locals refer to each other and read the cycle error.
Putting it all together
This project uses everything above to lay out a local workspace for an ML project: a folder per data stage with a README, a JSON config file a training script could read, and outputs a script can consume. It runs offline, and every line maps to a section of this guide.
terraform {
required_version = "~> 1.16"
required_providers {
local = {
source = "hashicorp/local"
version = "~> 2.9"
}
random = {
source = "hashicorp/random"
version = "~> 3.9"
}
}
}
variable "project" {
type = string
description = "Short name of the ML project."
validation {
condition = can(regex("^[a-z][a-z0-9-]{2,20}$", var.project))
error_message = "project must be 3-21 lowercase letters, digits or hyphens."
}
}
variable "environment" {
type = string
description = "dev, staging or prod."
default = "dev"
validation {
condition = contains(["dev", "staging", "prod"], var.environment)
error_message = "environment must be dev, staging or prod."
}
}
variable "data_folders" {
type = set(string)
description = "Data stages that each get their own folder."
default = ["raw", "features", "models"]
}
resource "random_pet" "run" {
length = 2
}
locals {
name_prefix = "${var.project}-${var.environment}"
root = "${path.module}/workspace/${local.name_prefix}"
}
resource "local_file" "stage_readme" {
for_each = var.data_folders
filename = "${local.root}/${each.key}/README.md"
content = <<-EOT
# ${each.key}
Stage folder for ${local.name_prefix}.
Managed by Terraform. Do not edit by hand.
EOT
}
resource "local_file" "config" {
filename = "${local.root}/config.json"
content = jsonencode({
project = var.project
environment = var.environment
run_name = random_pet.run.id
paths = { for stage in var.data_folders : stage => "${local.root}/${stage}" }
})
}
output "workspace_root" {
description = "Where the workspace was created."
value = local.root
}
output "config_path" {
description = "Path to the generated config, for training scripts."
value = local_file.config.filename
}
project = "churn"
environment = "dev"
terraform init
terraform fmt -check
terraform validate
terraform plan -var-file=dev.tfvars -out=tfplan
terraform apply tfplan
cat "$(terraform output -raw config_path)"
terraform destroy -var-file=dev.tfvars
| Line | Why it is there |
|---|---|
required_version and ~> on providers |
Everyone runs compatible versions, and the lock file records exact ones |
type, description and validation on variables |
Wrong input fails at plan with a readable message |
data_folders typed as set(string) |
for_each can use it directly, no toset() needed |
locals for the prefix and root |
One place to change naming |
for_each over names, not count |
Removing a stage removes only that stage |
jsonencode and a for expression |
The config is valid JSON built from the same inputs |
The reference to random_pet.run.id |
Orders the pet before the config, with no depends_on |
-var-file=dev.tfvars |
The environment is explicit on the command line |
plan -out then apply tfplan |
What runs is exactly what you read |
output -raw in the shell |
Scripts get values without parsing |
- Build this project from scratch, typing it rather than pasting it, and run the full command sequence.
- Add a
prod.tfvars, apply it too, and notice that the second apply wants to replace the dev workspace. Work out why from the plan (one configuration, one state). - Add a stage called
"evaluation", plan, and confirm only new files appear. - Commit it to Git with the
.gitignorefrom this guide, and check that no state file was committed.
What you can now do, and what comes next
You can explain what Terraform is and why declarative infrastructure beats clicking and scripting, install it correctly, write a configuration with providers, resources, variables, locals and outputs, run the init, plan and apply cycle, and read a plan closely enough to catch a replacement or a destroy before it happens. You know what state is, why it holds secrets, and why it never goes in Git. You can make many resources with for_each, read existing things with data sources, let references order your resources, pin versions with a committed lock file, and go from an error message to its fix.
| Can you… | |
|---|---|
| Explain declarative versus imperative? | Describe the end state; Terraform computes the steps |
| Name the three inputs to a plan? | Configuration, state, real infrastructure |
Say what init does? |
Downloads providers, sets up the backend, writes the lock file |
| Spot a replacement? | -/+ and # forces replacement |
| Say why state is sensitive? | It stores attribute values in plain text, secrets included |
| Say what to commit? | .tf files and .terraform.lock.hcl, never state or .terraform/ |
Choose between count and for_each? |
for_each for named things |
| Force a rebuild the modern way? | terraform apply -replace=ADDR |
| Apply exactly what you reviewed? | plan -out=tfplan, then apply tfplan |
| Explain what happens after a failed apply? | No rollback; fix and apply again |
Mid-level builds directly on this page: remote backends with locking so a team can share state, modules for reusing configuration, separate environments with separate state, importing resources that were created by hand, refactoring with moved blocks, terraform test, and running plan and apply in CI with pull-request review. Senior covers owning Terraform for an organisation: state architecture and blast radius, secrets that never touch state, policy as code, supply-chain controls, upgrades, and when to reach for something else.
Terraform usually creates the platform that other tools run on. The natural next guides are Kubernetes, to run workloads on the cluster Terraform builds, and then Helm and Argo CD for delivering applications onto it. If you haven't containerised anything yet, the Docker guide comes first.
Sources
- Terraform documentation
- Install Terraform
- Terraform releases
- Terraform v1.16 changelog
- Terraform v1.x compatibility promises
- CLI commands overview
- terraform plan command
- Resource block reference
- The for_each meta-argument
- The lifecycle meta-argument
- Manage sensitive data
- Dependency lock file
- S3 backend
- hashicorp/local provider
- hashicorp/random provider
- hashicorp/aws provider