Skip to content
Back to student guides
AxolotlLLMsFine-tuning & training3 levels91 sectionsCovers Axolotl 0.20

The Complete Axolotl Guide

Fine-tune LLMs from a single YAML config with Axolotl: SFT, LoRA, multi-GPU and more. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
17sections
24examples

This is part one of three. It covers everything you need to fine-tune a language model with Axolotl for the first time, not a teaser. By the end you can install the tool, write a configuration file, check that your data is being read the way you think it is, train a small model, talk to the result, and merge it into something you can hand to a colleague. Mid-level and Senior take the same ground further, into multi-GPU training, reinforcement learning and running Axolotl as a platform. Nothing here is thrown away.

This guide is written against Axolotl 0.20.0, released on 30 September 2026. The project moves quickly and renames configuration keys every few weeks, so a tutorial found on a blog from last year is probably wrong in at least one place. Where a key changed recently, this guide shows the current form and says what the old one was.

Each section ends with a Try it task. Do them in order. Fine-tuning is a craft you learn by watching a loss number fall (or refuse to), and no amount of reading replaces one real run.

What Axolotl is, and the problem it solves

A large language model such as Llama, Qwen or Gemma arrives from its publisher as a set of weights: billions of numbers that encode what the model learned from reading a vast amount of text. Out of the box it is a general-purpose assistant. Often that is not what you want. You may need it to answer in a particular tone, follow a strict output format, know the vocabulary of a specific industry, or handle Arabic dialect text better than the publisher's release does. Fine-tuning means continuing the model's training on your own examples so that its behaviour shifts towards them.

Doing that by hand is more work than it first appears. You must load the model, load and clean the data, turn each example into the exact token sequence the model expects, decide which tokens count towards the learning signal, pick a method that fits in your GPU's memory, choose an optimizer and a learning rate schedule, run the loop, log the results, save checkpoints, and later load everything back for testing. Each of those steps has a library behind it: transformers for models and the trainer, peft for efficient adapters, trl for preference and reinforcement methods, accelerate for running on several GPUs, datasets for data. Wiring them together in a Python script is possible, and many people do it, but every project ends up with its own slightly different script, and every script has its own bugs.

Axolotl is an open-source framework that wraps those libraries behind one YAML file. You describe the whole run in a configuration (which model, which data, which method, which hyperparameters, where to save) and Axolotl does the wiring. The same file then drives every other step: preparing the data, training, trying the model, merging the adapter, evaluating, quantizing and exporting.

YAML CONFIGone file describes the run
→
PREPROCESStokenize and cache data
→
TRAINload model, run the loop
→
INFERENCE / MERGEtest, then fold in the adapter

That matters for a reader in the Gulf or Egypt. Many employers cannot send customer text to a third-party API, either because of data-residency rules or because the contract forbids it. A config-driven tool that runs entirely on hardware you control, including a rented GPU in a region you choose, is a practical answer to that constraint. Axolotl does collect anonymous usage telemetry by default, and we cover how to turn it off in the configuration section; in a regulated environment, turn it off from the first run.

Two facts about how the tool is built explain much of what follows:

Try it
  1. Write down one behaviour you would want a model to have that a general assistant does not (a tone, a format, a domain).
  2. Write three example inputs and the exact outputs you would want for them.
  3. Ask yourself whether a better prompt alone could produce those outputs.
a clear sense of what you would be teaching. If a prompt can already do it reliably, fine-tuning may not be worth the effort; if it cannot, the three examples you wrote are the seed of your training data.

The mental model: six nouns

Axolotl has a lot of keys in its configuration, hundreds in the reference page, but a beginner can work with six ideas. Learn these and the rest of the file becomes readable.

Base model. The pretrained model you start from, named by the key base_model. It is either a repository id on the Hugging Face Hub, such as NousResearch/Llama-3.2-1B, or a path to a folder on disk. It is the only key Axolotl strictly requires.

Adapter. A small set of extra trainable weights attached to a frozen base model. Instead of changing billions of numbers, you train a few million that sit alongside them. The most common kind is LoRA (low-rank adaptation). The variant QLoRA does the same thing but loads the frozen base in a compressed 4-bit form to save memory. If your config has no adapter: line, Axolotl does a full fine-tune, where every parameter is updated.

Dataset and its type. Your examples live in files or on the Hub, but a model cannot read a spreadsheet row. The dataset type tells Axolotl how to turn a row into tokens, and which of those tokens the model should learn to produce. Types include alpaca for instruction and response pairs, chat_template for multi-turn conversations, and completion for plain text.

Prepared dataset. Turning text into tokens takes time, so Axolotl does it once, in a step called preprocessing, and caches the result on disk. Training then reads the cache.

Sequence length and packing. sequence_len is the maximum number of tokens in one training row. Sample packing glues several short examples into one row so that the GPU is not spending its time on padding.

Output directory. Where checkpoints and the final adapter land, set with output_dir.

config.ymlthe one file you edit
Hugging Face Hubbase model, tokenizer, datasets
your dataJSONL, CSV, Parquet or a Hub id
what the axolotl command does
validatePydantic checks every key
prepare datatokenize, mask, pack, cache
traintransformers + peft + accelerate
output_dircheckpoints, adapter, tokenizer, config copy

Why it matters: every box on the middle row reads the same YAML, so one mistake in the file shows up early, at validation, rather than an hour into training.

One more idea deserves a place in your head from the start: label masking. During training the model is shown a sequence of tokens and asked to predict each next token. The loss, the number that says how wrong it was, is computed over tokens whose label is not -100. Tokens labelled -100 are ignored. For instruction data you normally want the model to learn the answer, not to memorise the question, so the prompt tokens are masked. Axolotl's default, train_on_inputs: false, does exactly that. Most surprising results from a first fine-tune trace back to masking being different from what the author assumed, which is why the preprocessing step has a debug mode that lets you see it.

Try it
  1. Without looking back, write one sentence each for: base model, adapter, dataset type, prepared dataset, sequence length, output directory.
  2. Then describe in your own words what masking prompt tokens achieves.
six short sentences that you can say aloud. If masking is fuzzy, reread the last paragraph before moving on.

What you need before you start

Fine-tuning needs a GPU, and the first practical question is which one. Axolotl supports NVIDIA GPUs, and an AMD path exists, though its documentation is dated and you should treat it as community-grade. For the example in this guide you want an NVIDIA card from the Ampere generation or newer (RTX 30-series, A10, A100, L4, H100 and later). Ampere or newer matters because bf16, the 16-bit number format most recipes use, and Flash Attention both need it. Older cards such as the T4 or V100 can work with fp16 instead, but expect more friction.

How much memory? The project's own guidance for choosing a method gives rough figures for short sequences and a micro-batch of one or two:

Model size QLoRA LoRA in bf16 Full fine-tune in bf16
7 to 8 billion parameters 10 to 14 GB 16 to 24 GB 60 to 80 GB
70 to 72 billion parameters 40 to 48 GB 2 x 80 GB 4 to 8 x 80 GB

Read that table as a reason to start small. This guide uses a one-billion-parameter model, which fits comfortably on a modest card, so you can finish every step without renting anything expensive. The method (LoRA instead of full fine-tuning) matters more than any other single decision for memory.

You also need:

  • Linux, or Windows through WSL2. Axolotl does not support Windows natively; the documentation recommends WSL2 or Docker.
  • Python 3.12 or newer. This is a hard minimum since version 0.20. Tutorials that say Python 3.10 are out of date.
  • PyTorch 2.13 or newer. The release pins a range from 2.13.0 to 2.14.0.
  • A Hugging Face account, for downloading models and, if you choose, uploading results. Some models are gated, meaning you must accept their licence on the model page before the download works.
  • Disk space. Models and datasets are cached, and checkpoints are written during training. Leave tens of gigabytes free, more for larger models.

A Mac is for authoring, not training. Apple Silicon is partly supported: full training, LoRA and sample packing work, but 4-bit and 8-bit loading (so no QLoRA), mixed precision, Flash Attention and DeepSpeed do not. Use a Mac to write configs and check data on a tiny model, then train on a Linux GPU.

If you do not own a GPU, the documentation names providers such as RunPod, Vast.ai and Modal that work with Axolotl's Docker images, and guides for SkyPilot and Hugging Face Jobs. Rented machines bill while they run, so smoke-test on a few steps first.

Heads up Blackwell GPUs (B100, B200, B300 and RTX 50-series) need CUDA 13.0. The documentation notes that CUDA 12.8 cannot compile for the B300. If you rent a very new card and the install fails during compilation, check the CUDA version first.
Try it
  1. On the machine you plan to use, run nvidia-smi and note the GPU name and its memory.
  2. Run python3 --version.
  3. Using the table above, decide whether a 1B model with LoRA fits (it will) and what the largest model you could LoRA-tune is.
the GPU name, its memory and your Python version written down. If Python is older than 3.12, the next section's virtual environment will fix that.

Installing Axolotl

The recommended installer is uv, a fast Python package and environment manager. It replaces the old combination of pip, venv and friends for this purpose, and the project's own instructions and Docker images are built around it. On Linux the sequence is:

BASH
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env

export UV_TORCH_BACKEND=cu130
uv venv --python 3.12
source .venv/bin/activate
uv pip install --no-build-isolation 'axolotl[deepspeed]'

Take it line by line. The first two install uv and put it on your path. UV_TORCH_BACKEND tells uv which CUDA build of PyTorch to fetch; cu130 means CUDA 13.0, and the documentation also lists cu128 for CUDA 12.8. Pick the one that matches your driver. uv venv --python 3.12 creates an isolated environment so Axolotl's pinned libraries cannot clash with anything else on your machine, and source .venv/bin/activate switches into it. The final line installs Axolotl itself, together with the optional deepspeed extra.

The --no-build-isolation flag is required. Some of Axolotl's dependencies compile code against the PyTorch that is already installed, so the build must be allowed to see it. Leaving the flag out is the most common cause of a failed install.

The square brackets in axolotl[deepspeed] name an extra, an optional bundle of dependencies. Quote them (as above) because the zsh shell on macOS and many Linux setups treats square brackets as a pattern. Other extras you will meet later include flash-attn, vllm and ringmaster. Beginners can ignore them for now.

If you prefer plain pip, install PyTorch first, then:

BASH
pip3 install -U packaging setuptools wheel ninja
pip3 install --no-build-isolation 'axolotl[deepspeed]'

If you want the newest code instead of the last release, clone the repository and install it in editable mode:

BASH
git clone https://github.com/axolotl-ai-cloud/axolotl.git
cd axolotl
uv pip install --no-build-isolation -e '.[deepspeed]'

The Docker route avoids almost all of this. The project publishes images on Docker Hub that already contain a matching Python, CUDA, PyTorch and Axolotl:

BASH
docker run --gpus '"all"' --rm -it --ipc=host axolotlai/axolotl:main-latest

The --gpus '"all"' flag exposes your GPUs to the container (the odd double quoting is deliberate), and --ipc=host gives it the shared memory that data loaders need. The main-latest tag follows the development branch, which is fine for learning but wrong for anything you must reproduce; for that, use a release tag such as 0.20.0-py3.12-cu130-2.13.0. Since version 0.18 all published images are uv-based. If you are new to containers, the Docker guide explains what that command is doing. Rental providers usually want the axolotlai/axolotl-cloud image, which adds Jupyter and SSH access.

Tip Pin the version you learn on. uv pip install --no-build-isolation 'axolotl[deepspeed]==0.20.0' keeps a tutorial's config and your installed tool in agreement, which matters because keys are renamed between releases.
Try it
  1. Install Axolotl using the uv commands above, or start the Docker image.
  2. If the install fails, read the last twenty lines of the error. Look for a mention of CUDA, a missing compiler or the build-isolation flag.
an environment where axolotl is on your path. Verification comes next.

Checking that the setup works

Never start a three-hour run on an unverified install. Four quick commands tell you almost everything:

BASH
axolotl --version
axolotl --help
python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
nvidia-smi

The first should print axolotl, version 0.20.0. The second lists the commands the tool offers. The third prints your PyTorch version and then True if PyTorch can see a GPU. If it prints False, training will not use your GPU: the cause is nearly always a PyTorch build that does not match the CUDA driver, or running inside a container started without --gpus.

Next, log in to Hugging Face so the tool can download models:

BASH
hf auth login

Paste an access token from your account settings when prompted. The older huggingface-cli login command does the same job. You can also set the environment variable HF_TOKEN. Axolotl checks the token when it starts, and if something is wrong you will see a message beginning "Error verifying HuggingFace token. Remember to log in using hf auth login". If you are working with no internet on purpose, setting HF_HUB_OFFLINE=1 skips that check.

One behaviour surprises people: when you run any axolotl command, it first prints an ASCII-art banner, and it loads a .env file from the current folder if one exists. That means a token or key in .env is picked up automatically. It also means you must never commit that file.

Telemetry. Axolotl sends anonymous usage data (operating system, library versions, hardware, and sanitized error types) by default, and on first start it pauses for ten seconds to tell you so. To switch it off, set an environment variable before running:

BASH
export AXOLOTL_DO_NOT_TRACK=1

DO_NOT_TRACK=1 works as well. Setting AXOLOTL_DO_NOT_TRACK=0 explicitly turns it on and silences the notice. If your employer works in a regulated sector, ask which you should do, and do it deliberately rather than by accident.

Finally, fetch the example configs that ship with the project. They are the best teaching material you have:

BASH
axolotl fetch examples

This downloads a folder called examples/ into the current directory, with ready-made YAML files for many models and methods. Use examples from the same release as your install. An example from the development branch may use a key your installed version does not know.

Try it
  1. Run the four verification commands and confirm the version, the command list and True for CUDA.
  2. Run hf auth login.
  3. Run axolotl fetch examples and list the folder with ls examples. Open examples/llama-3/lora-1b.yml.
a working install, a logged-in Hub account and a config file you are about to read line by line.

Your first run: a LoRA fine-tune of a small model

The official quickstart trains a LoRA adapter on NousResearch/Llama-3.2-1B, a one-billion-parameter model that is not gated, using a public instruction dataset. We will follow it, then take it apart. Here is the config, saved as lora-1b.yml:

lora-1b.yml
base_model: NousResearch/Llama-3.2-1B

adapter: lora
lora_r: 16
lora_alpha: 32
lora_dropout: 0.05
lora_target_modules:
  - gate_proj
  - down_proj
  - up_proj
  - q_proj
  - v_proj
  - k_proj
  - o_proj

datasets:
  - path: teknium/GPT4-LLM-Cleaned
    type: alpaca
dataset_prepared_path: last_run_prepared
val_set_size: 0.1
output_dir: ./outputs/lora-out

sequence_len: 2048
sample_packing: true

micro_batch_size: 2
gradient_accumulation_steps: 2
num_epochs: 1
optimizer: adamw_8bit
lr_scheduler: cosine
learning_rate: 0.0002
warmup_ratio: 0.1

bf16: auto
gradient_checkpointing: true
attn_implementation: flash_attention_2

evals_per_epoch: 4
saves_per_epoch: 1
logging_steps: 1

special_tokens:
  pad_token: "<|end_of_text|>"

The quickstart page also shows a load_in_8bit: true line in its snippet, while the example file in the repository omits it. Both work; leaving it out is simpler and is what we do here.

The first run: do the cheap smoke test first, then the real run.

BASH
axolotl train lora-1b.yml --max-steps 5
axolotl train lora-1b.yml

If the five steps finish, your install, GPU, token and config all work. You can also take the file straight from the repository without saving it first, because the CLI accepts a local path or an HTTPS link to raw YAML. Pin such a link to a release tag, not to the moving main branch.

While it trains, watch the terminal. You will see a progress bar and, because logging_steps: 1, a line of numbers each step. The two you care about first are loss (how wrong the model currently is on the training data; it should trend downwards) and grad_norm (the size of the update signal; steady and modest is good). We read these properly in a later section. Also open a second terminal and run watch -n 2 nvidia-smi to see GPU memory in use. Watching it once teaches you more about the relationship between batch size and memory than any table.

A note on the first thing that may go wrong. If the log says the tokenizer has no padding token, the special_tokens block is what fixes it, and the quickstart config already includes it. A tokenizer is the component that turns text into token numbers, and training in batches requires one extra token to pad short rows; many base models do not define one, so you nominate an existing token for the job.

Try it
  1. Save the config as lora-1b.yml and run the five-step smoke test.
  2. In another terminal, run watch -n 2 nvidia-smi during the run and note peak memory.
  3. Run the full job. Note the final training loss and where the output was saved.
a folder at ./outputs/lora-out containing checkpoints and your adapter. The loss at the end should be lower than at the beginning, though how much depends on the data.

Reading the config, line by line

A config you cannot read is a config you cannot debug. Here is what every line in that file does, grouped by purpose.

The model. base_model names the starting point. Axolotl downloads it through the Hugging Face libraries on first use and caches it, so the second run starts faster.

The adapter. adapter: lora switches from full fine-tuning to LoRA. lora_r is the rank: the size of the small matrices that are trained. Higher rank means more capacity and more memory; 8 to 32 is the common range, and 16 is a sensible beginner default. lora_alpha scales how strongly the adapter's output is mixed in; a common convention is to set it to twice the rank, as the quickstart does. lora_dropout randomly zeroes a fraction of adapter activations during training to reduce overfitting, and 0.05 is a gentle value. lora_target_modules lists which layers of the model get an adapter; the seven names here cover the attention and feed-forward projections of a Llama-style model. The shortcut lora_target_linear: true targets every linear layer without you listing names.

The data. datasets is a list, so you can mix several sources. Each has a path (a Hub id or a local file) and a type. dataset_prepared_path is where the tokenized cache goes. val_set_size: 0.1 holds out ten percent of the data, never trained on, so the evaluation steps can measure how the model does on examples it has not seen.

Sequence handling. sequence_len: 2048 caps each row at 2,048 tokens. The default is only 512, which silently shortens many real examples, so always set it deliberately. sample_packing: true packs short examples together. With Flash Attention it does so safely, so packed examples cannot attend to each other.

Batching. micro_batch_size is how many examples each GPU processes in one forward pass; gradient_accumulation_steps is how many such passes are added up before one weight update. The effective batch size is the product of the two and the number of GPUs: here 2 x 2 x 1 = 4. If you run out of memory, lower the micro-batch and raise the accumulation to keep the effective size the same. Do not set batch_size yourself; it is derived.

The schedule. num_epochs: 1 means one pass over the data. If you also set max_steps, that wins. optimizer: adamw_8bit stores the optimizer's state in 8 bits to save memory. lr_scheduler: cosine shapes the learning rate along a cosine curve. learning_rate: 0.0002 is 2e-4, normal for LoRA; the project's stability guide gives 1e-4 to 3e-4 for LoRA and 1e-5 to 5e-5 for full fine-tuning. warmup_ratio: 0.1 ramps the rate up over the first tenth of training so early steps do not jolt the model.

Precision and speed. bf16: auto uses bfloat16 where the hardware supports it. gradient_checkpointing: true recomputes some values during the backward pass instead of storing them, trading roughly thirty percent more time for a large memory saving. attn_implementation: flash_attention_2 selects the attention algorithm. This is the current form. Older tutorials write flash_attention: true, which is still accepted but marked deprecated since version 0.17.

Evaluation and saving. evals_per_epoch: 4 evaluates four times per epoch; saves_per_epoch: 1 writes one checkpoint per epoch. These cannot be combined with the step-based versions (eval_steps, save_steps): Axolotl rejects a config that uses both of a pair.

Special tokens. The pad_token entry sets the padding token as discussed.

Note Axolotl validates the file with Pydantic when it loads, so a bad key or a contradictory pair of settings fails immediately with a readable message, before any model downloads. Run axolotl config-schema to see the full schema, or open the config reference page linked in Sources.
Try it
  1. Change lora_r from 16 to 8 and run five steps. Compare peak GPU memory with before.
  2. Change micro_batch_size to 1 and gradient_accumulation_steps to 4. The effective batch stays at 4; see what happens to memory.
  3. Put back the original values.
a feel for which knobs trade memory for capacity, and the proof that micro-batch and accumulation are two ways to reach the same effective batch.

Your own data: formats and dataset types

The quickstart dataset is someone else's. Soon you will want to train on your own examples, and the single most important skill is matching your data's shape to the right type. Axolotl accepts local files in JSON Lines, CSV, Parquet and Arrow formats (it infers the format from the extension), folders of files, Hub repositories, and paths on cloud storage such as s3:// and gs://. JSON Lines, one JSON object per line, is the easiest to produce and inspect.

Instruction data (alpaca). Each row has an instruction, an optional input and an output. A file my_data.jsonl looks like this:

JSON
{"instruction": "Translate to French.", "input": "Good morning", "output": "Bonjour"}
{"instruction": "Name a prime number between 10 and 20.", "input": "", "output": "13"}

and the config entry is:

YAML
datasets:
  - path: my_data.jsonl
    type: alpaca

Conversation data (chat_template). For multi-turn chat, each row holds a list of messages in the OpenAI style:

JSON
{"messages": [{"role": "user", "content": "What is MLOps?"}, {"role": "assistant", "content": "MLOps is the practice of running machine learning reliably in production."}]}

with this config:

YAML
chat_template: tokenizer_default
datasets:
  - path: my_chats.jsonl
    type: chat_template
    field_messages: messages
    message_property_mappings:
      role: role
      content: content
    roles_to_train: ["assistant"]
    train_on_eos: turn

A chat template is a small Jinja program, supplied by the model's publisher, that renders a list of messages into the exact text (with special marker tokens) that the model was trained to expect. chat_template: tokenizer_default means "use the one the tokenizer carries". Use it when your base model is a chat model that has one. If the base model has no template, you will get an error saying the tokenizer's chat_template is null, and the fix is to name a built-in one such as chatml or llama3, or to supply your own with chat_template_jinja.

roles_to_train: ["assistant"] is label masking for chat: only assistant turns count towards the loss, and system and user turns are context. train_on_eos: turn additionally trains the end-of-turn token for each assistant message, so the model learns when to stop talking.

If your data came from a ShareGPT-style export, with keys called conversations, from and value, you map them rather than renaming your file:

YAML
datasets:
  - path: sharegpt_export.jsonl
    type: chat_template
    field_messages: conversations
    message_property_mappings:
      role: from
      content: value

Do not use type: sharegpt. It was removed, and today it stops the run with the message "type: sharegpt.* is deprecated. Please use type: chat_template instead." Many older blog posts still show it.

Plain text (completion). If you just want to continue pretraining on raw text, a completion type reads a text field. This is for domain adaptation, not for teaching a model to follow instructions.

Heads up A dataset mapping mistake often shows up as jinja2.exceptions.UndefinedError: 'dict object' has no attribute 'content'. It means the template looked for a key your rows do not have. Check field_messages and message_property_mappings against the real key names in your file.
Try it
  1. Create a file of ten examples of the behaviour you chose in the first section, in the format that fits (alpaca for single turns, chat_template for chats).
  2. Validate it with python -c "import json; [json.loads(l) for l in open('my_data.jsonl')]", which fails if any line is not valid JSON.
a clean file that parses. Ten examples is too few to teach a model anything, but it is plenty to test the pipeline.

Preprocessing, and looking at what the model will actually see

This is the step beginners skip and later wish they had not. axolotl preprocess runs only the data stage: it tokenizes your dataset, applies masking and packing, and writes the result to the cache.

BASH
axolotl preprocess lora-1b.yml --debug

With --debug, it prints a few examples as a stream of tokens, each shown with its label and its token id. Tokens the model will be trained to predict show a real label; tokens that are masked show -100. Read a few examples and ask the plain question: is the part I want the model to learn the part that is not masked, and do the special marker tokens look right? You can ask for more examples with --debug-num-examples.

Why do this separately when train would do it anyway? Three reasons. It surfaces data problems in seconds, before you have loaded a model. It can run on a machine with no GPU at all (setting CUDA_VISIBLE_DEVICES="" keeps it off the GPU), which is cheaper when your GPU is rented by the hour. And it makes the cache explicit, so you know where it is.

The cache has a subtlety that causes real bugs. The tokenized data is stored under dataset_prepared_path, in a folder named by a hash of the relevant config. When you leave dataset_prepared_path empty, output goes to ./last_run_prepared/ and the old cache is ignored, so every run re-tokenizes. When you set it explicitly, as the quickstart does, the cache is reused on the next run. That is faster, but if you edit your data file or a custom prompt strategy without changing the config, you may silently train on the stale version. When in doubt, delete the folder:

BASH
rm -rf last_run_prepared
Tip Make preprocess --debug a reflex whenever you change the dataset type, the chat template, the special tokens or the roles to train. A one-minute look prevents the commonest class of silent failure, where training succeeds but teaches the wrong thing.
Try it
  1. Run axolotl preprocess lora-1b.yml --debug on the quickstart config.
  2. Find one example. Identify which tokens carry a label and which show -100.
  3. Set train_on_inputs: true in the config, run it again, and compare.
with the default, the instruction is masked and the response is trained; with train_on_inputs: true, almost everything carries a label. Put the setting back afterwards.

Training and reading the logs

When axolotl train is running you are looking at the output of the Hugging Face trainer. The three numbers worth understanding at first are loss, gradient norm and learning rate.

Loss. It is the average error per predicted token. For chat fine-tuning on a reasonable dataset the project's stability guide says to expect values between roughly 0.5 and 2.0. What matters more than the absolute number is the trend. A loss that drops quickly and then flattens is normal. A loss that stays flat from the start suggests the learning rate is too low, the data is not what you think, or nearly everything is masked. A loss that hits nan (not a number) means numerical trouble, often an over-high learning rate or a precision problem.

Evaluation loss. Because you set val_set_size, Axolotl periodically reports eval_loss on held-out examples. Compare it to the training loss. If training loss keeps falling while evaluation loss rises, the model is overfitting: memorising the training set instead of learning something general. The cures are more data, fewer epochs, a lower learning rate or a smaller rank.

Gradient norm. grad_norm measures how large the update is. The stability guide gives a healthy range of about 0.1 to 10, with spikes above 100 indicating instability. The config key max_grad_norm (1.0 is the usual value) clips updates so a rare bad batch cannot wreck the model.

Learning rate. With a warmup and a cosine scheduler you will see it climb, peak, then glide towards zero. If the log shows it stuck at a constant, check that lr_scheduler is what you intended.

Training writes into output_dir. You will find checkpoint folders (one per save), and, at the end, the adapter weights and tokenizer files. A copy of your config is placed there too, which is why tokens must never be pasted into the YAML: the file travels with the output, and it can end up on the Hub if you push.

For tracking runs over time, Axolotl logs to Weights and Biases, MLflow, Comet, Trackio and TensorBoard. Setting up one tracker early is worth it, because comparing a dozen runs by scrolling terminal history is miserable. The MLflow guide and the Weights and Biases guide cover those tools. In Axolotl you enable them with a few keys, for example wandb_project: my-first-finetune, with your key supplied through the WANDB_API_KEY environment variable rather than the file.

Try it
  1. Run a short training of about 50 steps with --max-steps 50 and logging every step.
  2. Copy the loss from the first and last steps into a note.
  3. Add use_tensorboard: true to the config and run again, then open the logs in TensorBoard if you have it.
a loss that moves downward over the short run, and a taste of what a tracker gives you that a terminal does not.

Trying the model, then merging the adapter

A trained adapter is not a standalone model. It is a small file that only makes sense on top of its base model. To test it, Axolotl loads both:

BASH
axolotl inference lora-1b.yml --lora-model-dir ./outputs/lora-out

You get an interactive prompt: type something, press enter, and the model answers. Since version 0.18 there is a multi-turn chat mode, which keeps context across messages and suits a model trained on conversations:

BASH
axolotl inference lora-1b.yml --lora-model-dir ./outputs/lora-out --chat

and a browser interface through --gradio. The inference command reads the same YAML, so it knows the base model and the prompt format. It can also read prompts from standard input, which lets you script a batch of test questions.

Prompt your model the way you will use it in practice, including some inputs that look nothing like your training data. A model that handles your ten training-like examples and falls apart on everything else has memorised, not learned. Keep a small fixed list of test prompts, so you can compare run to run.

Once you are satisfied, merge the adapter into the base weights to produce a normal, standalone model:

BASH
axolotl merge-lora lora-1b.yml --lora-model-dir ./outputs/lora-out

The merged model is written to ./outputs/lora-out/merged. After this, it is an ordinary Hugging Face model folder that can be loaded by the transformers library or served by an inference engine such as vLLM. If your GPU is busy, setting CUDA_VISIBLE_DEVICES="" makes the merge run on CPU. Always merge with this command and not with a custom script: Axolotl handles details such as embeddings that grew during training.

There is one error to recognise here. Putting merge_lora: true into a config that also sets load_in_4bit produces "Can't merge qlora if loaded in 4bit". The merge command handles quantized bases itself, so run axolotl merge-lora separately, optionally with --dequant to get a bf16 result.

To use the merged model on a laptop with tools such as Ollama or llama.cpp, version 0.20 added a GGUF exporter, axolotl export. It needs a built copy of llama.cpp and it refuses an adapter-only folder, so merge first. That is a mid-level topic; for now just know the route exists.

If you want to share your work, hub_model_id in the config pushes the output to your Hugging Face account.

Try it
  1. Run axolotl inference on your adapter and try five prompts, two of them unlike your training data.
  2. Run axolotl merge-lora and list the contents of the merged folder.
  3. Compare the folder sizes of the adapter and the merged model.
the adapter is tiny next to the merged model, because it holds only the trained delta, while the merged folder holds the full set of weights.

The everyday commands, grouped by purpose

Axolotl's command line follows one pattern: axolotl <command> <config.yml> [options]. The config can be a local path or an HTTPS URL. Options use dashes (--lora-model-dir). By contrast the legacy form python -m axolotl.cli.train uses underscores. Learn them by what you are trying to do.

Getting started.

Command What it does
axolotl fetch examples Downloads the example configs
axolotl fetch deepspeed_configs Downloads the DeepSpeed JSON files (used later)
axolotl config-schema Prints the JSON schema of every config key
axolotl agent-docs Bundled documentation, readable offline

Preparing data.

Command What it does
axolotl preprocess cfg.yml Tokenizes and caches the dataset
axolotl preprocess cfg.yml --debug Also prints tokens with their labels

Training and measuring.

Command What it does
axolotl train cfg.yml Runs training
axolotl train cfg.yml --resume-from-checkpoint DIR Continues from a checkpoint
axolotl evaluate cfg.yml Computes loss on the train and eval sets
axolotl lm-eval cfg.yml Runs the EleutherAI evaluation harness, with its plugin enabled

Using the result.

Command What it does
axolotl inference cfg.yml --lora-model-dir DIR Talk to the model
axolotl merge-lora cfg.yml Fold the adapter into the base weights
axolotl quantize cfg.yml Post-training quantization with torchao
axolotl export cfg.yml Export to GGUF (new in 0.20)

Two ideas are worth adopting immediately. First, any config key can be overridden on the command line as --key-name value, using dashes. That is how the --max-steps 5 smoke test worked without editing the file. Nested keys use dot notation, for example --trl.beta 0.1. Use this for one-off experiments and keep the file as the record of what you ran. Second, the launcher decides how worker processes start. The default is accelerate; with several GPUs you can pass --launcher torchrun, and arguments for the launcher go after a double dash:

BASH
axolotl train cfg.yml --launcher torchrun -- --nproc_per_node=2

You will not need that until you have more than one GPU, which is the subject of the next level. The older style, accelerate launch -m axolotl.cli.train cfg.yml, still works and appears in many tutorials.

When a run misbehaves, set AXOLOTL_LOG_LEVEL=DEBUG and keep a copy of the output:

BASH
AXOLOTL_LOG_LEVEL=DEBUG axolotl train cfg.yml 2>&1 | tee run.log

The tee command shows the output while also saving it, so you can search the log afterwards or paste it into a bug report.

Try it
  1. Run axolotl --help and match every command to one row of the tables above.
  2. Run axolotl config-schema --field adapter and read what it says about that key.
  3. Run a five-step training with a command-line override, for example --learning-rate 0.0001.
comfort with the command shape, and confirmation that the file does not need editing for a quick experiment.

Choosing a method: LoRA, QLoRA or full fine-tuning

The main decision in any run is how much of the model to train. Everything else follows from it.

LoRA

  • Trains small adapter matrices on a frozen base
  • Base loaded in 16-bit
  • Output is a small adapter file
  • The usual starting point
  • Learning rate around 1e-4 to 3e-4

Full fine-tune

  • Updates every weight
  • Needs several times the memory
  • Output is a complete model
  • Most capacity, most cost
  • Learning rate around 1e-5 to 5e-5

LoRA is the right default. It uses far less memory, trains faster, and for teaching a style, a format or a narrow domain it usually gets you most of the way there. QLoRA is LoRA with the frozen base compressed to 4-bit, which lets a model fit on a smaller card at some cost in speed and a small cost in quality. Choose it when plain LoRA does not fit. To switch, change the adapter and ask for 4-bit loading:

YAML
adapter: qlora
load_in_4bit: true

Both lines are needed. If you set adapter: qlora alone, validation fails with "Require cfg.load_in_4bit to be True for qlora". QLoRA also relies on the bitsandbytes library, which is not available on macOS, so it is a Linux GPU technique. Full fine-tuning is the config with no adapter: line at all. It can teach the model things an adapter cannot, but it needs many times the memory, produces a full-size output and is easier to ruin the model with, because every weight can move. Reach for it only when LoRA has demonstrably failed.

How do you decide? If LoRA fits on your GPU, use it; if not, try QLoRA; only if LoRA demonstrably cannot learn your task, consider a larger rank and then a full fine-tune. The library behind adapters is covered in the PEFT and TRL guide, and DeepSpeed is what Axolotl uses to spread large runs across GPUs.

Tip Axolotl switches on faster LoRA kernels for you when it can ("Auto-enabling LoRA kernel optimizations for faster training" appears in the log). Do not hunt for flags to speed up a basic LoRA run; most of the available speed-ups are already on.
Try it
  1. Copy your config to qlora-1b.yml, change the adapter to qlora and add load_in_4bit: true.
  2. Run five steps with each config and compare peak GPU memory and step time.
lower memory and somewhat slower steps for QLoRA. That trade is the entire reason the method exists.

Common errors and how to read them

Error messages in Axolotl usually come from one of three places: the config validator (immediate, readable), the data pipeline (during preprocessing) or PyTorch (during model load or training). Identify the stage first, then match the message.

Message What it means Fix
CUDA out of memory The GPU ran out of VRAM Lower micro-batch, enable gradient checkpointing, use QLoRA, shorten sequence_len
exitcode: -9 The OS killed the process for using too much RAM, often while tokenizing Lower dataset_num_proc, or use less data
Asking to pad but the tokenizer does not have a padding token Batches need a pad token and yours has none Add special_tokens: {pad_token: "<|end_of_text|>"} (often the end-of-sequence token)
`type: sharegpt.*` is deprecated. Please use `type: chat_template` instead. Old dataset type Switch to chat_template with a mapping
`chat_template` choice is `tokenizer_default` but tokenizer's `chat_template` is null The base model has no template Set chat_template: chatml or another built-in name
Require cfg.load_in_4bit to be True for qlora QLoRA without 4-bit loading Add load_in_4bit: true
save_steps and saves_per_epoch are mutually exclusive You set both of a pair Keep one (same for eval_steps and evals_per_epoch)
bf16 requested, but AMP is not supported on this GPU Pre-Ampere card Use fp16: true and bf16: false
Error verifying HuggingFace token Not logged in, or a gated model hf auth login, and accept the model licence on the Hub
max_steps must be set when using streaming datasets Streaming data has no known length Set max_steps
eval dataset split is too small for sample_packing Tiny validation set with packing on Set eval_sample_packing: false or enlarge val_set_size
Data processing error: CAS service error A hiccup in the Hub's storage backend export HF_HUB_DISABLE_XET=1

The padding error exists because a batch holds rows of different lengths and the filler token is the pad token. The bf16 error is a hardware fact: bfloat16 needs an Ampere-generation GPU or newer. For problems not in the table, read the first error in the log, decide which stage failed, and reproduce with the smallest config.

Try it
  1. Copy your config and deliberately break it: set adapter: qlora without load_in_4bit.
  2. Run it and read the message. Then add a mutually exclusive pair, for example save_steps: 10 next to saves_per_epoch: 1.
  3. Fix each one.
two quick failures from the validator, each with a message that names the problem. Validation errors are your friends: they arrive before any GPU time is spent.

Putting it all together: a small end-to-end project

Here is a complete workflow that uses everything above. The goal is to teach a small model a consistent house style for support replies, from a dataset you build yourself.

  1. Set up and pin.Create the virtual environment, install axolotl==0.20.0, run axolotl --version, log in with hf auth login, and set AXOLOTL_DO_NOT_TRACK=1 if your employer requires it.
  2. Build the data.Write at least a few hundred examples of the style you want in support.jsonl, in the OpenAI messages format, with a system message that stays identical across rows. Quality beats quantity; ten careful examples beat a hundred sloppy ones.
  3. Write the config.Copy the quickstart, set the dataset to type: chat_template with field_messages: messages, set chat_template: chatml if the base model has no template, and name the output ./outputs/support-lora.
  4. Look at the data.Run axolotl preprocess support.yml --debug and confirm only assistant turns are trained.
  5. Smoke test.Run axolotl train support.yml --max-steps 5, fix whatever appears, then launch the full run.
  6. Watch it.Follow loss and eval loss. Stop and rethink if eval loss rises while training loss falls.
  7. Test it.Run axolotl inference support.yml --lora-model-dir ./outputs/support-lora --chat with a fixed set of test prompts, including ones outside the style.
  8. Merge and keep.Run axolotl merge-lora support.yml --lora-model-dir ./outputs/support-lora, and store the config, the data and a note of the version next to the result.

Here is the config for step three, compact enough to read in one go:

support.yml
base_model: NousResearch/Llama-3.2-1B

adapter: lora
lora_r: 16
lora_alpha: 32
lora_dropout: 0.05
lora_target_linear: true

chat_template: chatml
datasets:
  - path: support.jsonl
    type: chat_template
    field_messages: messages
    message_property_mappings:
      role: role
      content: content
    roles_to_train: ["assistant"]
    train_on_eos: turn

dataset_prepared_path: last_run_prepared
val_set_size: 0.1
output_dir: ./outputs/support-lora

sequence_len: 2048
sample_packing: true
eval_sample_packing: false

micro_batch_size: 2
gradient_accumulation_steps: 4
num_epochs: 2
optimizer: adamw_8bit
lr_scheduler: cosine
learning_rate: 0.0002
warmup_ratio: 0.1

bf16: auto
gradient_checkpointing: true
attn_implementation: flash_attention_2

evals_per_epoch: 4
saves_per_epoch: 1
logging_steps: 1

special_tokens:
  pad_token: "<|end_of_text|>"

When the result disappoints, the cause is almost always one of three things: not enough good data, a prompt format at test time that differs from training, or masking that is not what you assumed. Check them in that order, and change one thing per experiment.

Try it
  1. Follow the eight steps with a dataset of your own, even a small one.
  2. Keep a log of each command you ran, the loss at the end, and one thing you would change next time.
a merged model that answers in your style, and a written record. The record is as valuable as the model, because it is what lets you improve on it.

What you can now do, and what comes next

You can now explain the problem Axolotl solves and the six nouns of its model. You can install it, verify the install, and avoid the early traps (the Python and PyTorch minimums, --no-build-isolation, the token login). You can read a LoRA config line by line, match your data to a dataset type, check masking before spending GPU time, run and monitor a training job, test the result and merge it, and recognise the common errors by their messages.

The beginner's rules of thumb: start small; smoke-test with a few steps; look at the tokens before you train; change one thing at a time; keep the config and the version with every result; never put secrets in the file. When a run runs out of memory, lower micro_batch_size and raise gradient_accumulation_steps, turn on gradient_checkpointing, switch to QLoRA, then shorten sequence_len. The dataset's size does not change peak memory.

The mid level goes deeper into the machinery: multi-GPU training with DeepSpeed and FSDP2, plugins and kernels, preference methods such as DPO, tracking and evaluation. Neighbouring guides worth reading: the Hugging Face guide, the PEFT and TRL guide and the vLLM guide. The senior level covers running Axolotl as a shared platform.

The quality of a fine-tuned model is decided mostly by the quality of the data and how carefully you checked it. Spend your next hour on the examples, not the knobs.

Sources