This is part one of three. It covers everything you need to fine-tune a language model with Axolotl for the first time, not a teaser. By the end you can install the tool, write a configuration file, check that your data is being read the way you think it is, train a small model, talk to the result, and merge it into something you can hand to a colleague. Mid-level and Senior take the same ground further, into multi-GPU training, reinforcement learning and running Axolotl as a platform. Nothing here is thrown away.
This guide is written against Axolotl 0.20.0, released on 30 September 2026. The project moves quickly and renames configuration keys every few weeks, so a tutorial found on a blog from last year is probably wrong in at least one place. Where a key changed recently, this guide shows the current form and says what the old one was.
Each section ends with a Try it task. Do them in order. Fine-tuning is a craft you learn by watching a loss number fall (or refuse to), and no amount of reading replaces one real run.
What Axolotl is, and the problem it solves
A large language model such as Llama, Qwen or Gemma arrives from its publisher as a set of weights: billions of numbers that encode what the model learned from reading a vast amount of text. Out of the box it is a general-purpose assistant. Often that is not what you want. You may need it to answer in a particular tone, follow a strict output format, know the vocabulary of a specific industry, or handle Arabic dialect text better than the publisher's release does. Fine-tuning means continuing the model's training on your own examples so that its behaviour shifts towards them.
Doing that by hand is more work than it first appears. You must load the model, load and clean the data, turn each example into the exact token sequence the model expects, decide which tokens count towards the learning signal, pick a method that fits in your GPU's memory, choose an optimizer and a learning rate schedule, run the loop, log the results, save checkpoints, and later load everything back for testing. Each of those steps has a library behind it: transformers for models and the trainer, peft for efficient adapters, trl for preference and reinforcement methods, accelerate for running on several GPUs, datasets for data. Wiring them together in a Python script is possible, and many people do it, but every project ends up with its own slightly different script, and every script has its own bugs.
Axolotl is an open-source framework that wraps those libraries behind one YAML file. You describe the whole run in a configuration (which model, which data, which method, which hyperparameters, where to save) and Axolotl does the wiring. The same file then drives every other step: preparing the data, training, trying the model, merging the adapter, evaluating, quantizing and exporting.
That matters for a reader in the Gulf or Egypt. Many employers cannot send customer text to a third-party API, either because of data-residency rules or because the contract forbids it. A config-driven tool that runs entirely on hardware you control, including a rented GPU in a region you choose, is a practical answer to that constraint. Axolotl does collect anonymous usage telemetry by default, and we cover how to turn it off in the configuration section; in a regulated environment, turn it off from the first run.
Two facts about how the tool is built explain much of what follows:
- Write down one behaviour you would want a model to have that a general assistant does not (a tone, a format, a domain).
- Write three example inputs and the exact outputs you would want for them.
- Ask yourself whether a better prompt alone could produce those outputs.
The mental model: six nouns
Axolotl has a lot of keys in its configuration, hundreds in the reference page, but a beginner can work with six ideas. Learn these and the rest of the file becomes readable.
Base model. The pretrained model you start from, named by the key base_model. It is either a repository id on the Hugging Face Hub, such as NousResearch/Llama-3.2-1B, or a path to a folder on disk. It is the only key Axolotl strictly requires.
Adapter. A small set of extra trainable weights attached to a frozen base model. Instead of changing billions of numbers, you train a few million that sit alongside them. The most common kind is LoRA (low-rank adaptation). The variant QLoRA does the same thing but loads the frozen base in a compressed 4-bit form to save memory. If your config has no adapter: line, Axolotl does a full fine-tune, where every parameter is updated.
Dataset and its type. Your examples live in files or on the Hub, but a model cannot read a spreadsheet row. The dataset type tells Axolotl how to turn a row into tokens, and which of those tokens the model should learn to produce. Types include alpaca for instruction and response pairs, chat_template for multi-turn conversations, and completion for plain text.
Prepared dataset. Turning text into tokens takes time, so Axolotl does it once, in a step called preprocessing, and caches the result on disk. Training then reads the cache.
Sequence length and packing. sequence_len is the maximum number of tokens in one training row. Sample packing glues several short examples into one row so that the GPU is not spending its time on padding.
Output directory. Where checkpoints and the final adapter land, set with output_dir.
Why it matters: every box on the middle row reads the same YAML, so one mistake in the file shows up early, at validation, rather than an hour into training.
One more idea deserves a place in your head from the start: label masking. During training the model is shown a sequence of tokens and asked to predict each next token. The loss, the number that says how wrong it was, is computed over tokens whose label is not -100. Tokens labelled -100 are ignored. For instruction data you normally want the model to learn the answer, not to memorise the question, so the prompt tokens are masked. Axolotl's default, train_on_inputs: false, does exactly that. Most surprising results from a first fine-tune trace back to masking being different from what the author assumed, which is why the preprocessing step has a debug mode that lets you see it.
- Without looking back, write one sentence each for: base model, adapter, dataset type, prepared dataset, sequence length, output directory.
- Then describe in your own words what masking prompt tokens achieves.
What you need before you start
Fine-tuning needs a GPU, and the first practical question is which one. Axolotl supports NVIDIA GPUs, and an AMD path exists, though its documentation is dated and you should treat it as community-grade. For the example in this guide you want an NVIDIA card from the Ampere generation or newer (RTX 30-series, A10, A100, L4, H100 and later). Ampere or newer matters because bf16, the 16-bit number format most recipes use, and Flash Attention both need it. Older cards such as the T4 or V100 can work with fp16 instead, but expect more friction.
How much memory? The project's own guidance for choosing a method gives rough figures for short sequences and a micro-batch of one or two:
| Model size | QLoRA | LoRA in bf16 | Full fine-tune in bf16 |
|---|---|---|---|
| 7 to 8 billion parameters | 10 to 14 GB | 16 to 24 GB | 60 to 80 GB |
| 70 to 72 billion parameters | 40 to 48 GB | 2 x 80 GB | 4 to 8 x 80 GB |
Read that table as a reason to start small. This guide uses a one-billion-parameter model, which fits comfortably on a modest card, so you can finish every step without renting anything expensive. The method (LoRA instead of full fine-tuning) matters more than any other single decision for memory.
You also need:
- Linux, or Windows through WSL2. Axolotl does not support Windows natively; the documentation recommends WSL2 or Docker.
- Python 3.12 or newer. This is a hard minimum since version 0.20. Tutorials that say Python 3.10 are out of date.
- PyTorch 2.13 or newer. The release pins a range from 2.13.0 to 2.14.0.
- A Hugging Face account, for downloading models and, if you choose, uploading results. Some models are gated, meaning you must accept their licence on the model page before the download works.
- Disk space. Models and datasets are cached, and checkpoints are written during training. Leave tens of gigabytes free, more for larger models.
A Mac is for authoring, not training. Apple Silicon is partly supported: full training, LoRA and sample packing work, but 4-bit and 8-bit loading (so no QLoRA), mixed precision, Flash Attention and DeepSpeed do not. Use a Mac to write configs and check data on a tiny model, then train on a Linux GPU.
If you do not own a GPU, the documentation names providers such as RunPod, Vast.ai and Modal that work with Axolotl's Docker images, and guides for SkyPilot and Hugging Face Jobs. Rented machines bill while they run, so smoke-test on a few steps first.
- On the machine you plan to use, run
nvidia-smiand note the GPU name and its memory. - Run
python3 --version. - Using the table above, decide whether a 1B model with LoRA fits (it will) and what the largest model you could LoRA-tune is.
Installing Axolotl
The recommended installer is uv, a fast Python package and environment manager. It replaces the old combination of pip, venv and friends for this purpose, and the project's own instructions and Docker images are built around it. On Linux the sequence is:
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
export UV_TORCH_BACKEND=cu130
uv venv --python 3.12
source .venv/bin/activate
uv pip install --no-build-isolation 'axolotl[deepspeed]'
Take it line by line. The first two install uv and put it on your path. UV_TORCH_BACKEND tells uv which CUDA build of PyTorch to fetch; cu130 means CUDA 13.0, and the documentation also lists cu128 for CUDA 12.8. Pick the one that matches your driver. uv venv --python 3.12 creates an isolated environment so Axolotl's pinned libraries cannot clash with anything else on your machine, and source .venv/bin/activate switches into it. The final line installs Axolotl itself, together with the optional deepspeed extra.
The --no-build-isolation flag is required. Some of Axolotl's dependencies compile code against the PyTorch that is already installed, so the build must be allowed to see it. Leaving the flag out is the most common cause of a failed install.
The square brackets in axolotl[deepspeed] name an extra, an optional bundle of dependencies. Quote them (as above) because the zsh shell on macOS and many Linux setups treats square brackets as a pattern. Other extras you will meet later include flash-attn, vllm and ringmaster. Beginners can ignore them for now.
If you prefer plain pip, install PyTorch first, then:
pip3 install -U packaging setuptools wheel ninja
pip3 install --no-build-isolation 'axolotl[deepspeed]'
If you want the newest code instead of the last release, clone the repository and install it in editable mode:
git clone https://github.com/axolotl-ai-cloud/axolotl.git
cd axolotl
uv pip install --no-build-isolation -e '.[deepspeed]'
The Docker route avoids almost all of this. The project publishes images on Docker Hub that already contain a matching Python, CUDA, PyTorch and Axolotl:
docker run --gpus '"all"' --rm -it --ipc=host axolotlai/axolotl:main-latest
The --gpus '"all"' flag exposes your GPUs to the container (the odd double quoting is deliberate), and --ipc=host gives it the shared memory that data loaders need. The main-latest tag follows the development branch, which is fine for learning but wrong for anything you must reproduce; for that, use a release tag such as 0.20.0-py3.12-cu130-2.13.0. Since version 0.18 all published images are uv-based. If you are new to containers, the Docker guide explains what that command is doing. Rental providers usually want the axolotlai/axolotl-cloud image, which adds Jupyter and SSH access.
uv pip install --no-build-isolation 'axolotl[deepspeed]==0.20.0' keeps a tutorial's config and your installed tool in agreement, which matters because keys are renamed between releases.
- Install Axolotl using the uv commands above, or start the Docker image.
- If the install fails, read the last twenty lines of the error. Look for a mention of CUDA, a missing compiler or the build-isolation flag.
axolotl is on your path. Verification comes next.
Checking that the setup works
Never start a three-hour run on an unverified install. Four quick commands tell you almost everything:
axolotl --version
axolotl --help
python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
nvidia-smi
The first should print axolotl, version 0.20.0. The second lists the commands the tool offers. The third prints your PyTorch version and then True if PyTorch can see a GPU. If it prints False, training will not use your GPU: the cause is nearly always a PyTorch build that does not match the CUDA driver, or running inside a container started without --gpus.
Next, log in to Hugging Face so the tool can download models:
hf auth login
Paste an access token from your account settings when prompted. The older huggingface-cli login command does the same job. You can also set the environment variable HF_TOKEN. Axolotl checks the token when it starts, and if something is wrong you will see a message beginning "Error verifying HuggingFace token. Remember to log in using hf auth login". If you are working with no internet on purpose, setting HF_HUB_OFFLINE=1 skips that check.
One behaviour surprises people: when you run any axolotl command, it first prints an ASCII-art banner, and it loads a .env file from the current folder if one exists. That means a token or key in .env is picked up automatically. It also means you must never commit that file.
Telemetry. Axolotl sends anonymous usage data (operating system, library versions, hardware, and sanitized error types) by default, and on first start it pauses for ten seconds to tell you so. To switch it off, set an environment variable before running:
export AXOLOTL_DO_NOT_TRACK=1
DO_NOT_TRACK=1 works as well. Setting AXOLOTL_DO_NOT_TRACK=0 explicitly turns it on and silences the notice. If your employer works in a regulated sector, ask which you should do, and do it deliberately rather than by accident.
Finally, fetch the example configs that ship with the project. They are the best teaching material you have:
axolotl fetch examples
This downloads a folder called examples/ into the current directory, with ready-made YAML files for many models and methods. Use examples from the same release as your install. An example from the development branch may use a key your installed version does not know.
- Run the four verification commands and confirm the version, the command list and
Truefor CUDA. - Run
hf auth login. - Run
axolotl fetch examplesand list the folder withls examples. Openexamples/llama-3/lora-1b.yml.
Your first run: a LoRA fine-tune of a small model
The official quickstart trains a LoRA adapter on NousResearch/Llama-3.2-1B, a one-billion-parameter model that is not gated, using a public instruction dataset. We will follow it, then take it apart. Here is the config, saved as lora-1b.yml:
base_model: NousResearch/Llama-3.2-1B
adapter: lora
lora_r: 16
lora_alpha: 32
lora_dropout: 0.05
lora_target_modules:
- gate_proj
- down_proj
- up_proj
- q_proj
- v_proj
- k_proj
- o_proj
datasets:
- path: teknium/GPT4-LLM-Cleaned
type: alpaca
dataset_prepared_path: last_run_prepared
val_set_size: 0.1
output_dir: ./outputs/lora-out
sequence_len: 2048
sample_packing: true
micro_batch_size: 2
gradient_accumulation_steps: 2
num_epochs: 1
optimizer: adamw_8bit
lr_scheduler: cosine
learning_rate: 0.0002
warmup_ratio: 0.1
bf16: auto
gradient_checkpointing: true
attn_implementation: flash_attention_2
evals_per_epoch: 4
saves_per_epoch: 1
logging_steps: 1
special_tokens:
pad_token: "<|end_of_text|>"
The quickstart page also shows a load_in_8bit: true line in its snippet, while the example file in the repository omits it. Both work; leaving it out is simpler and is what we do here.
The first run: do the cheap smoke test first, then the real run.
axolotl train lora-1b.yml --max-steps 5
axolotl train lora-1b.yml
If the five steps finish, your install, GPU, token and config all work. You can also take the file straight from the repository without saving it first, because the CLI accepts a local path or an HTTPS link to raw YAML. Pin such a link to a release tag, not to the moving main branch.
While it trains, watch the terminal. You will see a progress bar and, because logging_steps: 1, a line of numbers each step. The two you care about first are loss (how wrong the model currently is on the training data; it should trend downwards) and grad_norm (the size of the update signal; steady and modest is good). We read these properly in a later section. Also open a second terminal and run watch -n 2 nvidia-smi to see GPU memory in use. Watching it once teaches you more about the relationship between batch size and memory than any table.
A note on the first thing that may go wrong. If the log says the tokenizer has no padding token, the special_tokens block is what fixes it, and the quickstart config already includes it. A tokenizer is the component that turns text into token numbers, and training in batches requires one extra token to pad short rows; many base models do not define one, so you nominate an existing token for the job.
- Save the config as
lora-1b.ymland run the five-step smoke test. - In another terminal, run
watch -n 2 nvidia-smiduring the run and note peak memory. - Run the full job. Note the final training loss and where the output was saved.
./outputs/lora-out containing checkpoints and your adapter. The loss at the end should be lower than at the beginning, though how much depends on the data.
Reading the config, line by line
A config you cannot read is a config you cannot debug. Here is what every line in that file does, grouped by purpose.
The model. base_model names the starting point. Axolotl downloads it through the Hugging Face libraries on first use and caches it, so the second run starts faster.
The adapter. adapter: lora switches from full fine-tuning to LoRA. lora_r is the rank: the size of the small matrices that are trained. Higher rank means more capacity and more memory; 8 to 32 is the common range, and 16 is a sensible beginner default. lora_alpha scales how strongly the adapter's output is mixed in; a common convention is to set it to twice the rank, as the quickstart does. lora_dropout randomly zeroes a fraction of adapter activations during training to reduce overfitting, and 0.05 is a gentle value. lora_target_modules lists which layers of the model get an adapter; the seven names here cover the attention and feed-forward projections of a Llama-style model. The shortcut lora_target_linear: true targets every linear layer without you listing names.
The data. datasets is a list, so you can mix several sources. Each has a path (a Hub id or a local file) and a type. dataset_prepared_path is where the tokenized cache goes. val_set_size: 0.1 holds out ten percent of the data, never trained on, so the evaluation steps can measure how the model does on examples it has not seen.
Sequence handling. sequence_len: 2048 caps each row at 2,048 tokens. The default is only 512, which silently shortens many real examples, so always set it deliberately. sample_packing: true packs short examples together. With Flash Attention it does so safely, so packed examples cannot attend to each other.
Batching. micro_batch_size is how many examples each GPU processes in one forward pass; gradient_accumulation_steps is how many such passes are added up before one weight update. The effective batch size is the product of the two and the number of GPUs: here 2 x 2 x 1 = 4. If you run out of memory, lower the micro-batch and raise the accumulation to keep the effective size the same. Do not set batch_size yourself; it is derived.
The schedule. num_epochs: 1 means one pass over the data. If you also set max_steps, that wins. optimizer: adamw_8bit stores the optimizer's state in 8 bits to save memory. lr_scheduler: cosine shapes the learning rate along a cosine curve. learning_rate: 0.0002 is 2e-4, normal for LoRA; the project's stability guide gives 1e-4 to 3e-4 for LoRA and 1e-5 to 5e-5 for full fine-tuning. warmup_ratio: 0.1 ramps the rate up over the first tenth of training so early steps do not jolt the model.
Precision and speed. bf16: auto uses bfloat16 where the hardware supports it. gradient_checkpointing: true recomputes some values during the backward pass instead of storing them, trading roughly thirty percent more time for a large memory saving. attn_implementation: flash_attention_2 selects the attention algorithm. This is the current form. Older tutorials write flash_attention: true, which is still accepted but marked deprecated since version 0.17.
Evaluation and saving. evals_per_epoch: 4 evaluates four times per epoch; saves_per_epoch: 1 writes one checkpoint per epoch. These cannot be combined with the step-based versions (eval_steps, save_steps): Axolotl rejects a config that uses both of a pair.
Special tokens. The pad_token entry sets the padding token as discussed.
axolotl config-schema to see the full schema, or open the config reference page linked in Sources.
- Change
lora_rfrom 16 to 8 and run five steps. Compare peak GPU memory with before. - Change
micro_batch_sizeto 1 andgradient_accumulation_stepsto 4. The effective batch stays at 4; see what happens to memory. - Put back the original values.
Your own data: formats and dataset types
The quickstart dataset is someone else's. Soon you will want to train on your own examples, and the single most important skill is matching your data's shape to the right type. Axolotl accepts local files in JSON Lines, CSV, Parquet and Arrow formats (it infers the format from the extension), folders of files, Hub repositories, and paths on cloud storage such as s3:// and gs://. JSON Lines, one JSON object per line, is the easiest to produce and inspect.
Instruction data (alpaca). Each row has an instruction, an optional input and an output. A file my_data.jsonl looks like this:
{"instruction": "Translate to French.", "input": "Good morning", "output": "Bonjour"}
{"instruction": "Name a prime number between 10 and 20.", "input": "", "output": "13"}
and the config entry is:
datasets:
- path: my_data.jsonl
type: alpaca
Conversation data (chat_template). For multi-turn chat, each row holds a list of messages in the OpenAI style:
{"messages": [{"role": "user", "content": "What is MLOps?"}, {"role": "assistant", "content": "MLOps is the practice of running machine learning reliably in production."}]}
with this config:
chat_template: tokenizer_default
datasets:
- path: my_chats.jsonl
type: chat_template
field_messages: messages
message_property_mappings:
role: role
content: content
roles_to_train: ["assistant"]
train_on_eos: turn
A chat template is a small Jinja program, supplied by the model's publisher, that renders a list of messages into the exact text (with special marker tokens) that the model was trained to expect. chat_template: tokenizer_default means "use the one the tokenizer carries". Use it when your base model is a chat model that has one. If the base model has no template, you will get an error saying the tokenizer's chat_template is null, and the fix is to name a built-in one such as chatml or llama3, or to supply your own with chat_template_jinja.
roles_to_train: ["assistant"] is label masking for chat: only assistant turns count towards the loss, and system and user turns are context. train_on_eos: turn additionally trains the end-of-turn token for each assistant message, so the model learns when to stop talking.
If your data came from a ShareGPT-style export, with keys called conversations, from and value, you map them rather than renaming your file:
datasets:
- path: sharegpt_export.jsonl
type: chat_template
field_messages: conversations
message_property_mappings:
role: from
content: value
Do not use type: sharegpt. It was removed, and today it stops the run with the message "type: sharegpt.* is deprecated. Please use type: chat_template instead." Many older blog posts still show it.
Plain text (completion). If you just want to continue pretraining on raw text, a completion type reads a text field. This is for domain adaptation, not for teaching a model to follow instructions.
jinja2.exceptions.UndefinedError: 'dict object' has no attribute 'content'. It means the template looked for a key your rows do not have. Check field_messages and message_property_mappings against the real key names in your file.
- Create a file of ten examples of the behaviour you chose in the first section, in the format that fits (
alpacafor single turns,chat_templatefor chats). - Validate it with
python -c "import json; [json.loads(l) for l in open('my_data.jsonl')]", which fails if any line is not valid JSON.
Preprocessing, and looking at what the model will actually see
This is the step beginners skip and later wish they had not. axolotl preprocess runs only the data stage: it tokenizes your dataset, applies masking and packing, and writes the result to the cache.
axolotl preprocess lora-1b.yml --debug
With --debug, it prints a few examples as a stream of tokens, each shown with its label and its token id. Tokens the model will be trained to predict show a real label; tokens that are masked show -100. Read a few examples and ask the plain question: is the part I want the model to learn the part that is not masked, and do the special marker tokens look right? You can ask for more examples with --debug-num-examples.
Why do this separately when train would do it anyway? Three reasons. It surfaces data problems in seconds, before you have loaded a model. It can run on a machine with no GPU at all (setting CUDA_VISIBLE_DEVICES="" keeps it off the GPU), which is cheaper when your GPU is rented by the hour. And it makes the cache explicit, so you know where it is.
The cache has a subtlety that causes real bugs. The tokenized data is stored under dataset_prepared_path, in a folder named by a hash of the relevant config. When you leave dataset_prepared_path empty, output goes to ./last_run_prepared/ and the old cache is ignored, so every run re-tokenizes. When you set it explicitly, as the quickstart does, the cache is reused on the next run. That is faster, but if you edit your data file or a custom prompt strategy without changing the config, you may silently train on the stale version. When in doubt, delete the folder:
rm -rf last_run_prepared
preprocess --debug a reflex whenever you change the dataset type, the chat template, the special tokens or the roles to train. A one-minute look prevents the commonest class of silent failure, where training succeeds but teaches the wrong thing.
- Run
axolotl preprocess lora-1b.yml --debugon the quickstart config. - Find one example. Identify which tokens carry a label and which show
-100. - Set
train_on_inputs: truein the config, run it again, and compare.
train_on_inputs: true, almost everything carries a label. Put the setting back afterwards.
Training and reading the logs
When axolotl train is running you are looking at the output of the Hugging Face trainer. The three numbers worth understanding at first are loss, gradient norm and learning rate.
Loss. It is the average error per predicted token. For chat fine-tuning on a reasonable dataset the project's stability guide says to expect values between roughly 0.5 and 2.0. What matters more than the absolute number is the trend. A loss that drops quickly and then flattens is normal. A loss that stays flat from the start suggests the learning rate is too low, the data is not what you think, or nearly everything is masked. A loss that hits nan (not a number) means numerical trouble, often an over-high learning rate or a precision problem.
Evaluation loss. Because you set val_set_size, Axolotl periodically reports eval_loss on held-out examples. Compare it to the training loss. If training loss keeps falling while evaluation loss rises, the model is overfitting: memorising the training set instead of learning something general. The cures are more data, fewer epochs, a lower learning rate or a smaller rank.
Gradient norm. grad_norm measures how large the update is. The stability guide gives a healthy range of about 0.1 to 10, with spikes above 100 indicating instability. The config key max_grad_norm (1.0 is the usual value) clips updates so a rare bad batch cannot wreck the model.
Learning rate. With a warmup and a cosine scheduler you will see it climb, peak, then glide towards zero. If the log shows it stuck at a constant, check that lr_scheduler is what you intended.
Training writes into output_dir. You will find checkpoint folders (one per save), and, at the end, the adapter weights and tokenizer files. A copy of your config is placed there too, which is why tokens must never be pasted into the YAML: the file travels with the output, and it can end up on the Hub if you push.
For tracking runs over time, Axolotl logs to Weights and Biases, MLflow, Comet, Trackio and TensorBoard. Setting up one tracker early is worth it, because comparing a dozen runs by scrolling terminal history is miserable. The MLflow guide and the Weights and Biases guide cover those tools. In Axolotl you enable them with a few keys, for example wandb_project: my-first-finetune, with your key supplied through the WANDB_API_KEY environment variable rather than the file.
- Run a short training of about 50 steps with
--max-steps 50and logging every step. - Copy the loss from the first and last steps into a note.
- Add
use_tensorboard: trueto the config and run again, then open the logs in TensorBoard if you have it.
Trying the model, then merging the adapter
A trained adapter is not a standalone model. It is a small file that only makes sense on top of its base model. To test it, Axolotl loads both:
axolotl inference lora-1b.yml --lora-model-dir ./outputs/lora-out
You get an interactive prompt: type something, press enter, and the model answers. Since version 0.18 there is a multi-turn chat mode, which keeps context across messages and suits a model trained on conversations:
axolotl inference lora-1b.yml --lora-model-dir ./outputs/lora-out --chat
and a browser interface through --gradio. The inference command reads the same YAML, so it knows the base model and the prompt format. It can also read prompts from standard input, which lets you script a batch of test questions.
Prompt your model the way you will use it in practice, including some inputs that look nothing like your training data. A model that handles your ten training-like examples and falls apart on everything else has memorised, not learned. Keep a small fixed list of test prompts, so you can compare run to run.
Once you are satisfied, merge the adapter into the base weights to produce a normal, standalone model:
axolotl merge-lora lora-1b.yml --lora-model-dir ./outputs/lora-out
The merged model is written to ./outputs/lora-out/merged. After this, it is an ordinary Hugging Face model folder that can be loaded by the transformers library or served by an inference engine such as vLLM. If your GPU is busy, setting CUDA_VISIBLE_DEVICES="" makes the merge run on CPU. Always merge with this command and not with a custom script: Axolotl handles details such as embeddings that grew during training.
There is one error to recognise here. Putting merge_lora: true into a config that also sets load_in_4bit produces "Can't merge qlora if loaded in 4bit". The merge command handles quantized bases itself, so run axolotl merge-lora separately, optionally with --dequant to get a bf16 result.
To use the merged model on a laptop with tools such as Ollama or llama.cpp, version 0.20 added a GGUF exporter, axolotl export. It needs a built copy of llama.cpp and it refuses an adapter-only folder, so merge first. That is a mid-level topic; for now just know the route exists.
If you want to share your work, hub_model_id in the config pushes the output to your Hugging Face account.
- Run
axolotl inferenceon your adapter and try five prompts, two of them unlike your training data. - Run
axolotl merge-loraand list the contents of themergedfolder. - Compare the folder sizes of the adapter and the merged model.
The everyday commands, grouped by purpose
Axolotl's command line follows one pattern: axolotl <command> <config.yml> [options]. The config can be a local path or an HTTPS URL. Options use dashes (--lora-model-dir). By contrast the legacy form python -m axolotl.cli.train uses underscores. Learn them by what you are trying to do.
Getting started.
| Command | What it does |
|---|---|
axolotl fetch examples |
Downloads the example configs |
axolotl fetch deepspeed_configs |
Downloads the DeepSpeed JSON files (used later) |
axolotl config-schema |
Prints the JSON schema of every config key |
axolotl agent-docs |
Bundled documentation, readable offline |
Preparing data.
| Command | What it does |
|---|---|
axolotl preprocess cfg.yml |
Tokenizes and caches the dataset |
axolotl preprocess cfg.yml --debug |
Also prints tokens with their labels |
Training and measuring.
| Command | What it does |
|---|---|
axolotl train cfg.yml |
Runs training |
axolotl train cfg.yml --resume-from-checkpoint DIR |
Continues from a checkpoint |
axolotl evaluate cfg.yml |
Computes loss on the train and eval sets |
axolotl lm-eval cfg.yml |
Runs the EleutherAI evaluation harness, with its plugin enabled |
Using the result.
| Command | What it does |
|---|---|
axolotl inference cfg.yml --lora-model-dir DIR |
Talk to the model |
axolotl merge-lora cfg.yml |
Fold the adapter into the base weights |
axolotl quantize cfg.yml |
Post-training quantization with torchao |
axolotl export cfg.yml |
Export to GGUF (new in 0.20) |
Two ideas are worth adopting immediately. First, any config key can be overridden on the command line as --key-name value, using dashes. That is how the --max-steps 5 smoke test worked without editing the file. Nested keys use dot notation, for example --trl.beta 0.1. Use this for one-off experiments and keep the file as the record of what you ran. Second, the launcher decides how worker processes start. The default is accelerate; with several GPUs you can pass --launcher torchrun, and arguments for the launcher go after a double dash:
axolotl train cfg.yml --launcher torchrun -- --nproc_per_node=2
You will not need that until you have more than one GPU, which is the subject of the next level. The older style, accelerate launch -m axolotl.cli.train cfg.yml, still works and appears in many tutorials.
When a run misbehaves, set AXOLOTL_LOG_LEVEL=DEBUG and keep a copy of the output:
AXOLOTL_LOG_LEVEL=DEBUG axolotl train cfg.yml 2>&1 | tee run.log
The tee command shows the output while also saving it, so you can search the log afterwards or paste it into a bug report.
- Run
axolotl --helpand match every command to one row of the tables above. - Run
axolotl config-schema --field adapterand read what it says about that key. - Run a five-step training with a command-line override, for example
--learning-rate 0.0001.
Choosing a method: LoRA, QLoRA or full fine-tuning
The main decision in any run is how much of the model to train. Everything else follows from it.
LoRA
- Trains small adapter matrices on a frozen base
- Base loaded in 16-bit
- Output is a small adapter file
- The usual starting point
- Learning rate around 1e-4 to 3e-4
Full fine-tune
- Updates every weight
- Needs several times the memory
- Output is a complete model
- Most capacity, most cost
- Learning rate around 1e-5 to 5e-5
LoRA is the right default. It uses far less memory, trains faster, and for teaching a style, a format or a narrow domain it usually gets you most of the way there. QLoRA is LoRA with the frozen base compressed to 4-bit, which lets a model fit on a smaller card at some cost in speed and a small cost in quality. Choose it when plain LoRA does not fit. To switch, change the adapter and ask for 4-bit loading:
adapter: qlora
load_in_4bit: true
Both lines are needed. If you set adapter: qlora alone, validation fails with "Require cfg.load_in_4bit to be True for qlora". QLoRA also relies on the bitsandbytes library, which is not available on macOS, so it is a Linux GPU technique. Full fine-tuning is the config with no adapter: line at all. It can teach the model things an adapter cannot, but it needs many times the memory, produces a full-size output and is easier to ruin the model with, because every weight can move. Reach for it only when LoRA has demonstrably failed.
How do you decide? If LoRA fits on your GPU, use it; if not, try QLoRA; only if LoRA demonstrably cannot learn your task, consider a larger rank and then a full fine-tune. The library behind adapters is covered in the PEFT and TRL guide, and DeepSpeed is what Axolotl uses to spread large runs across GPUs.
- Copy your config to
qlora-1b.yml, change the adapter toqloraand addload_in_4bit: true. - Run five steps with each config and compare peak GPU memory and step time.
Common errors and how to read them
Error messages in Axolotl usually come from one of three places: the config validator (immediate, readable), the data pipeline (during preprocessing) or PyTorch (during model load or training). Identify the stage first, then match the message.
| Message | What it means | Fix |
|---|---|---|
CUDA out of memory |
The GPU ran out of VRAM | Lower micro-batch, enable gradient checkpointing, use QLoRA, shorten sequence_len |
exitcode: -9 |
The OS killed the process for using too much RAM, often while tokenizing | Lower dataset_num_proc, or use less data |
Asking to pad but the tokenizer does not have a padding token |
Batches need a pad token and yours has none | Add special_tokens: {pad_token: "<|end_of_text|>"} (often the end-of-sequence token) |
`type: sharegpt.*` is deprecated. Please use `type: chat_template` instead. |
Old dataset type | Switch to chat_template with a mapping |
`chat_template` choice is `tokenizer_default` but tokenizer's `chat_template` is null |
The base model has no template | Set chat_template: chatml or another built-in name |
Require cfg.load_in_4bit to be True for qlora |
QLoRA without 4-bit loading | Add load_in_4bit: true |
save_steps and saves_per_epoch are mutually exclusive |
You set both of a pair | Keep one (same for eval_steps and evals_per_epoch) |
bf16 requested, but AMP is not supported on this GPU |
Pre-Ampere card | Use fp16: true and bf16: false |
Error verifying HuggingFace token |
Not logged in, or a gated model | hf auth login, and accept the model licence on the Hub |
max_steps must be set when using streaming datasets |
Streaming data has no known length | Set max_steps |
eval dataset split is too small for sample_packing |
Tiny validation set with packing on | Set eval_sample_packing: false or enlarge val_set_size |
Data processing error: CAS service error |
A hiccup in the Hub's storage backend | export HF_HUB_DISABLE_XET=1 |
The padding error exists because a batch holds rows of different lengths and the filler token is the pad token. The bf16 error is a hardware fact: bfloat16 needs an Ampere-generation GPU or newer. For problems not in the table, read the first error in the log, decide which stage failed, and reproduce with the smallest config.
- Copy your config and deliberately break it: set
adapter: qlorawithoutload_in_4bit. - Run it and read the message. Then add a mutually exclusive pair, for example
save_steps: 10next tosaves_per_epoch: 1. - Fix each one.
Putting it all together: a small end-to-end project
Here is a complete workflow that uses everything above. The goal is to teach a small model a consistent house style for support replies, from a dataset you build yourself.
- Set up and pin.Create the virtual environment, install
axolotl==0.20.0, runaxolotl --version, log in withhf auth login, and setAXOLOTL_DO_NOT_TRACK=1if your employer requires it. - Build the data.Write at least a few hundred examples of the style you want in
support.jsonl, in the OpenAI messages format, with a system message that stays identical across rows. Quality beats quantity; ten careful examples beat a hundred sloppy ones. - Write the config.Copy the quickstart, set the dataset to
type: chat_templatewithfield_messages: messages, setchat_template: chatmlif the base model has no template, and name the output./outputs/support-lora. - Look at the data.Run
axolotl preprocess support.yml --debugand confirm only assistant turns are trained. - Smoke test.Run
axolotl train support.yml --max-steps 5, fix whatever appears, then launch the full run. - Watch it.Follow loss and eval loss. Stop and rethink if eval loss rises while training loss falls.
- Test it.Run
axolotl inference support.yml --lora-model-dir ./outputs/support-lora --chatwith a fixed set of test prompts, including ones outside the style. - Merge and keep.Run
axolotl merge-lora support.yml --lora-model-dir ./outputs/support-lora, and store the config, the data and a note of the version next to the result.
Here is the config for step three, compact enough to read in one go:
base_model: NousResearch/Llama-3.2-1B
adapter: lora
lora_r: 16
lora_alpha: 32
lora_dropout: 0.05
lora_target_linear: true
chat_template: chatml
datasets:
- path: support.jsonl
type: chat_template
field_messages: messages
message_property_mappings:
role: role
content: content
roles_to_train: ["assistant"]
train_on_eos: turn
dataset_prepared_path: last_run_prepared
val_set_size: 0.1
output_dir: ./outputs/support-lora
sequence_len: 2048
sample_packing: true
eval_sample_packing: false
micro_batch_size: 2
gradient_accumulation_steps: 4
num_epochs: 2
optimizer: adamw_8bit
lr_scheduler: cosine
learning_rate: 0.0002
warmup_ratio: 0.1
bf16: auto
gradient_checkpointing: true
attn_implementation: flash_attention_2
evals_per_epoch: 4
saves_per_epoch: 1
logging_steps: 1
special_tokens:
pad_token: "<|end_of_text|>"
When the result disappoints, the cause is almost always one of three things: not enough good data, a prompt format at test time that differs from training, or masking that is not what you assumed. Check them in that order, and change one thing per experiment.
- Follow the eight steps with a dataset of your own, even a small one.
- Keep a log of each command you ran, the loss at the end, and one thing you would change next time.
What you can now do, and what comes next
You can now explain the problem Axolotl solves and the six nouns of its model. You can install it, verify the install, and avoid the early traps (the Python and PyTorch minimums, --no-build-isolation, the token login). You can read a LoRA config line by line, match your data to a dataset type, check masking before spending GPU time, run and monitor a training job, test the result and merge it, and recognise the common errors by their messages.
The beginner's rules of thumb: start small; smoke-test with a few steps; look at the tokens before you train; change one thing at a time; keep the config and the version with every result; never put secrets in the file. When a run runs out of memory, lower micro_batch_size and raise gradient_accumulation_steps, turn on gradient_checkpointing, switch to QLoRA, then shorten sequence_len. The dataset's size does not change peak memory.
The mid level goes deeper into the machinery: multi-GPU training with DeepSpeed and FSDP2, plugins and kernels, preference methods such as DPO, tracking and evaluation. Neighbouring guides worth reading: the Hugging Face guide, the PEFT and TRL guide and the vLLM guide. The senior level covers running Axolotl as a shared platform.
The quality of a fine-tuned model is decided mostly by the quality of the data and how carefully you checked it. Spend your next hour on the examples, not the knobs.
Sources
- Axolotl documentation home
- Installation
- Getting started
- Command line reference
- Config reference
- Dataset formats
- Conversation dataset format
- Instruction tuning format
- Dataset loading
- Dataset preprocessing
- Sample packing
- Batch size vs gradient accumulation
- Choosing a method
- LoRA and QLoRA
- Inference
- Docker images
- Running on a Mac
- Telemetry
- FAQ
- Debugging
- Training stability
- Axolotl releases on GitHub
- Axolotl v0.20.0 release
- Axolotl on PyPI
- Axolotl on Docker Hub