تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
PEFT & TRLLLMsFine-tuning & training3 مستويات100 قسمًايغطّي TRL 1.14 / PEFT 0.21دليل بالإنجليزية

The Complete PEFT & TRL Guide

Fine-tune LLMs efficiently with LoRA/QLoRA (PEFT) and align them with SFT, DPO and RLHF (TRL). Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
17sections
23examples

This is part one of three. It assumes you have never used PEFT or TRL, and it takes you from an empty folder to a language model that you fine-tuned yourself, saved as a small adapter, reloaded, and merged into a standalone model. By the end you will be able to read a training log, recognise the handful of errors that stop most first runs, and know which of the many trainers in TRL to reach for next. Mid-level and Senior go further on the same topics; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. Fine-tuning is a skill you learn by watching a loss curve fall, and a loss curve that falls on your own data teaches more than any amount of reading.

A word on versions before we start. This guide is written against TRL 1.14 and PEFT 0.21, both of which were current at the end of September 2026. Both libraries changed a great deal between 2024 and 2026. Blog posts, forum answers and the memory of many language models still describe APIs that no longer exist, such as PPOTrainer, DataCollatorForCompletionOnlyLM, and SFTTrainer(tokenizer=..., dataset_text_field=..., max_seq_length=...). If you paste an old snippet and it fails, that is almost certainly why. Every example below uses the current form.

What PEFT and TRL are, and what you can do with them

A language model such as Qwen3 or Llama arrives from its creators as a file of billions of numbers called weights. Those numbers were learned from a vast amount of text, and they encode what the model knows and how it writes. Fine-tuning means continuing the training of that model on your own, much smaller dataset so that it behaves the way you want: answering in a house style, following a particular output format, speaking Arabic dialect more naturally, or handling the vocabulary of your industry.

Two libraries from Hugging Face do most of the work in the open-source world, and they are separate packages that cooperate.

TRL stands for Transformers Reinforcement Learning, though most of what beginners do with it is not reinforcement learning at all. It is Hugging Face's library for post-training, a word that covers everything you do to a model after the expensive pretraining is finished: supervised fine-tuning, preference tuning, and reinforcement learning. TRL gives you ready-made trainer classes. You hand a trainer a model and a dataset, call train(), and it runs the loop for you.

PEFT stands for Parameter-Efficient Fine-Tuning. Instead of updating every one of the billions of weights, PEFT freezes the original model and trains a tiny set of extra parameters, called an adapter, that sit alongside it. The best-known technique is LoRA. The adapter is typically a few megabytes to a few hundred megabytes, and it can be trained on a single modest GPU.

BASE MODELfrozen, billions of weights
→
PEFT ADAPTERsmall, trainable
→
TRL TRAINERruns the training loop
→
YOUR MODELadapter, or merged weights

Put together, the two libraries let you do things that used to need a cluster. With a small model and a free or cheap GPU you can teach a model a new format in under an hour. With a rented GPU you can fine-tune a seven-billion-parameter model. And because the result is an adapter rather than a full copy of the model, you can keep one base model and a dozen adapters for a dozen different customers or tasks.

What people use them for:

🗣️

Style and format

Teach a model to answer in a fixed structure, tone or language, such as Gulf-dialect customer replies.

🏷️

Narrow specialist tasks

Classification, extraction or routing where a small tuned model beats a large general one on cost.

⚖️

Preference alignment

Nudge outputs toward answers humans or a scoring function prefer, using DPO or GRPO.

🔒

Data that stays put

Train on your own hardware or in your own cloud region, so sensitive text never goes to a third-party API.

That last card matters in this region. Many employers in the Gulf and Egypt have data-residency rules that make sending customer text to an overseas API awkward. Fine-tuning an open model inside a cloud region you control, or on a machine in your own office, is a legitimate answer, and these two libraries are how it is usually done.

You need very little to follow along: Python 3.10 or newer, a terminal, and ideally an NVIDIA GPU with a few gigabytes of memory. A free notebook GPU is enough for everything in this guide, because we use a model with only 0.6 billion parameters. A Mac can also run the examples on its own accelerator, only slowly.

Try it
  1. Write down one behaviour you wish a language model had, in one sentence.
  2. Decide whether it is a matter of knowledge (facts the model lacks) or behaviour (how it answers).
  3. Keep that sentence. You will turn it into training data later.
a behaviour sentence, for example "always reply in two short sentences". Fine-tuning is far better at behaviour than at memorising facts.

The problem these libraries solve

To see why PEFT exists, you need a feel for the cost of training a model the obvious way.

Training adjusts every weight a little, many thousands of times. For each weight the training process must keep not only the weight itself but also its gradient (the direction to nudge it) and, with the common AdamW optimiser, two further running statistics per weight. A rough rule is that full fine-tuning needs several times the memory of the model itself. A seven-billion-parameter model that occupies around fourteen gigabytes just to sit in memory in half precision can need well over a hundred gigabytes to train fully. That means a rack of expensive accelerators, and every fine-tuned variant is another complete multi-gigabyte copy of the model that you have to store and ship.

That is the situation PEFT was designed for. The key observation behind LoRA is that when you fine-tune a large model for a specific purpose, the change to each weight matrix does not need to be complicated. It can be well approximated by the product of two thin matrices. If a weight matrix is a big square grid of numbers, LoRA leaves that grid alone and learns two skinny rectangles that multiply together to form a correction of the same size. The skinniness is controlled by a number called the rank, written r. With a small rank, the skinny matrices contain a tiny fraction of the numbers in the original, so there is far less to train and store.

Before PEFT, practitioners had a few poor options. They could fine-tune the whole model if they had the hardware. They could freeze most layers and train only the last few, which is cheap but often disappointing. Or they could skip fine-tuning and rely on careful prompting, which is cheap and flexible but cannot make a model reliably follow a format it does not naturally produce, and makes every request longer and more expensive.

TRL solves a different problem. Even with the memory question answered, writing a correct training loop for a language model is surprisingly fiddly. You must turn raw text into token IDs, decide which tokens count toward the loss, pad or pack examples into batches, schedule the learning rate, log metrics, save checkpoints, and handle several GPUs. Hugging Face's transformers.Trainer already handles much of this for generic models. TRL builds on top of it and adds the language-model-specific parts. If you have used Trainer before, TRL's trainers will feel familiar, because they are subclasses of it.

Try it
  1. Look up the parameter count of a model you like, for example 7 billion.
  2. Multiply by 2 to get its size in gigabytes at 16-bit precision.
  3. Multiply that by 4 or more to get a ballpark for full fine-tuning memory.
a number far above the memory of the GPU you can rent for an afternoon. That gap is exactly what LoRA closes.

The mental model: layers and nouns

Fine-tuning vocabulary comes at you fast, so it helps to see how the pieces stack before you meet any code. Think of five layers, from the bottom up.

PYTORCHtensors, gradients
→
TRANSFORMERSmodels, tokenizers
→
ACCELERATEdevices
→
PEFTadapters
→
TRLtrainers, CLI

You will mostly touch the top two layers, but the lower ones leak into error messages, so knowing they exist saves confusion. When a message mentions accelerate, it is the layer that decides which device your tensors live on. When it mentions transformers, it is about the model or the tokenizer.

Here are the nouns you need, defined once so that you can recognise them everywhere.

Base model. The pretrained model you start from, for example Qwen/Qwen3-0.6B. It is hosted on the Hugging Face Hub and downloaded the first time you use it. During LoRA training it stays frozen.

Tokenizer. The component that splits text into tokens, the integer pieces the model actually reads. Every model comes with its own tokenizer, and mixing a model with the wrong tokenizer produces nonsense. In TRL the tokenizer is passed as processing_class. Older code called it tokenizer; that name is gone from the trainers.

Dataset. A table of training examples, loaded with the Hugging Face datasets library. The columns it must have depend on the trainer, which is the subject of its own section below.

Trainer and config pairs. Every TRL method comes as two classes. The config holds the settings, and the trainer does the work. SFTConfig pairs with SFTTrainer, DPOConfig with DPOTrainer, GRPOConfig with GRPOTrainer, and so on. The config classes subclass transformers' TrainingArguments, so everything you may know about TrainingArguments, such as output_dir, num_train_epochs and learning_rate, still applies.

SFT. Supervised fine-tuning. You show the model examples of the text you want it to produce, and it learns to predict each next token. The loss is called cross-entropy or negative log-likelihood (NLL): roughly, how surprised the model is by the correct next token. Lower is better. SFT is where every beginner starts, and it is what this guide covers in depth.

LoRA. Low-Rank Adaptation. Next to each chosen weight matrix W, PEFT adds a pair of thin matrices A and B. The layer's output becomes the original output plus a scaled correction from B times A. B starts at zero, so at step zero the model behaves exactly like the base model, and training gradually grows the correction.

Adapter. The set of those added matrices. It is the only thing training changes, and the only thing you save.

Rank (r) and alpha (lora_alpha). The rank sets how thin the matrices are, and so how much capacity the adapter has. Alpha scales the correction; the scale applied is alpha divided by rank. A common starting point is rank 16 and alpha 32.

Target modules. Which layers of the model get adapters. You can name them, but the beginner-friendly choice is the string "all-linear", which means every linear layer except the final output layer.

Offline and online methods. TRL's trainers split into two families. Offline methods, which are SFT, DPO, KTO and reward modelling, learn from a fixed dataset that already exists. Online methods, GRPO and RLOO, make the model generate answers during training and then score them. Online methods are heavier and need more setup. This guide stays offline.

The one sentence to remember PEFT decides what gets trained (a small adapter), and TRL decides how it is trained (the loss, the data handling, the loop). You can use PEFT without TRL and TRL without PEFT, but the pairing is the everyday workhorse.
Try it
  1. Without looking back, write the five layers in order from PyTorch up to TRL.
  2. For each noun in bold above, write one sentence in your own words.
the five layers in the right order and a sentence each for base model, adapter, rank, and trainer. If any sentence is hard, reread that paragraph.

Installing and checking the setup

The install is short but has one trap, and the trap is PyTorch.

Install PyTorch first, using the selector on pytorch.org to pick the build that matches your hardware (a specific CUDA version for NVIDIA, ROCm for AMD, or the default for Apple Silicon and CPU). If you skip this step, installing TRL will pull in whatever PyTorch pip chooses, which on some systems is a CPU-only build, and you will wonder why training is a hundred times slower than expected.

Then create an isolated environment and install the libraries.

BASH
python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install "trl[peft]"

The [peft] part is an extra: TRL on its own does not install PEFT, so pip install trl followed by a first LoRA run would stop with an ImportError telling you to install trl[peft]. Asking for the extra up front avoids that. If you prefer uv, the same command works as uv pip install "trl[peft]". Python 3.10 or newer is required by both libraries. One page of the PEFT documentation still says 3.9, but the package metadata and the release notes say 3.10, and 3.10 is what you should assume.

There are other extras you will meet later, and it helps to know their names now. trl[quantization] adds bitsandbytes for 4-bit loading. trl[vllm] adds the fast generation engine used by online methods. trl[deepspeed] and trl[liger] add memory and speed helpers. You do not need any of them for this guide except bitsandbytes in the QLoRA section.

What works on which operating system

Piece Linux macOS (Apple Silicon) Windows
TRL and PEFT core, such as SFT Yes Yes, on CPU or the MPS accelerator, slowly Yes, with an NVIDIA GPU
bitsandbytes (QLoRA) Yes, NVIDIA GPU CPU wheel only, so no 4-bit GPU training Yes, NVIDIA GPU
vLLM Yes Not officially Not natively; use WSL2

On a Mac, stick to a tiny model and expect minutes where a GPU takes seconds. On Windows, native Python works for the basics, and WSL2 is the smoother path once you want anything more. A free cloud notebook with an NVIDIA GPU is a perfectly good place to do this whole guide.

Verify the install

TRL ships a command that prints everything relevant about your environment.

BASH
trl env
python -c "import trl, peft, transformers, torch; print(trl.__version__, peft.__version__, transformers.__version__, torch.__version__, torch.cuda.is_available())"
trl sft --help

trl env prints your platform, Python, PyTorch, accelerator, and the versions of transformers, accelerate, datasets, TRL, PEFT and others. The official documentation asks you to paste this output when you report an issue, so get used to running it. The second command prints five values on one line; the last one, True or False, tells you whether PyTorch can see an NVIDIA GPU. False on a machine with a GPU means you have the wrong PyTorch build. The third command proves that the trl command-line tool is installed.

Some models on the Hub are gated: you must accept a licence on the model page and then prove who you are. The tiny Qwen model we use is not gated, so you can skip this today, but when you move to a gated model, run hf auth login or set the HF_TOKEN environment variable.

Anonymous usage statistics Since TRL 1.5, each trainer you create sends one anonymous ping to Hugging Face, recording things like the TRL version, the trainer class, whether PEFT is used, and the GPU model. If your employer forbids that, set HF_HUB_DISABLE_TELEMETRY=1 before you run, or HF_HUB_OFFLINE=1 if you are working fully offline with a pre-filled cache. Nothing is sent when the CI environment variable is set.
Try it
  1. Create a virtual environment, install PyTorch for your hardware, then run pip install "trl[peft]".
  2. Run trl env and find the lines for TRL, PEFT and your accelerator.
  3. Run the one-line Python check and read the final True or False.
version numbers for both libraries, and True if you have an NVIDIA GPU. If you see False, fix your PyTorch install before going further.

Understanding your data: dataset formats

Most fine-tuning failures are data failures, so it pays to learn the shape of the data before running anything. Training data in TRL comes in two formats and several types.

The format is about how the text is written. In standard format each field is a plain string. In conversational format each field is a list of messages, and each message is a small dictionary with a role (such as system, user or assistant) and a content. Conversational format is how chat models are used in practice, and it is what most modern datasets use.

The type is about which fields exist, and each TRL trainer expects a particular type.

Type Columns Used by
Language modelling text (standard) or messages (conversational) SFT
Prompt-completion prompt and completion SFT
Preference prompt, chosen, rejected DPO, reward modelling
Unpaired preference prompt, completion, label (true or false) KTO
Prompt-only prompt GRPO, RLOO

A single row of a conversational language-modelling dataset looks like this.

PYTHON
example = {
    "messages": [
        {"role": "user", "content": "Summarise: the meeting moved to Thursday."},
        {"role": "assistant", "content": "The meeting is now on Thursday."},
    ]
}

A prompt-completion row separates the two halves, which gives you more control, because SFT then computes its loss only on the completion by default. The model is not rewarded for reproducing the question, only for producing the answer.

PYTHON
example = {
    "prompt": [{"role": "user", "content": "Summarise: the meeting moved to Thursday."}],
    "completion": [{"role": "assistant", "content": "The meeting is now on Thursday."}],
}

If you use plain messages data, the loss by default covers every token, including the user's turn. That is acceptable for a first run, and later you can switch on assistant_only_loss=True to restrict the loss to the assistant's turns. That option needs a chat template that marks the assistant tokens, and TRL patches the templates of well-known model families such as Qwen3 for you.

A chat template deserves a sentence of its own. It is a small Jinja program, stored with the tokenizer, that turns a list of messages into the exact string of special tokens that model was trained to expect. You never need to write it for a mainstream model; the trainer applies it for you when it sees conversational data. This is the main reason conversational format is convenient: the formatting details are handled once, correctly.

You can load a dataset from the Hub in one line and inspect it before spending a minute of GPU time.

PYTHON
from datasets import load_dataset

dataset = load_dataset("trl-lib/Capybara", split="train")
print(dataset)               # column names and number of rows
print(dataset[0])            # the first example

Read the output. trl-lib/Capybara is a conversational dataset with a messages column, which is why SFT accepts it with no preparation. Whenever you swap in your own data, run these two lines first and check that the column names match the table above. If they do not, rename the columns with dataset.rename_column("old", "new") before you hand the dataset to a trainer. If you only have plain strings, put them in a text column.

For your own data, the simplest route is a JSON Lines file with one example per line, which you can load like this.

PYTHON
dataset = load_dataset("json", data_files="my_data.jsonl", split="train")
Try it
  1. Load trl-lib/Capybara and print the first example.
  2. Find the role and content keys and count how many turns the conversation has.
  3. Write three examples of your own behaviour sentence in the messages format as a Python list.
a list of dictionaries with role and content keys. Three examples will not teach a model anything, but writing them proves you understand the shape.

Your first training run: SFT in a few lines

Now the payoff. The smallest useful TRL program is this, taken from the official quickstart pattern.

train_basic.py
from datasets import load_dataset
from trl import SFTTrainer

trainer = SFTTrainer(
    model="Qwen/Qwen3-0.6B",
    train_dataset=load_dataset("trl-lib/Capybara", split="train"),
)
trainer.train()

Four things are worth noticing. First, model is just a string. TRL downloads the model and its tokenizer from the Hub and builds them for you, so you do not need to import AutoModelForCausalLM for the simple case. Second, there is no args at all, so TRL uses defaults. Third, there is no peft_config, so this run is full fine-tuning: every weight is trained. For a 0.6-billion-parameter model that is just about possible on a good GPU, but it is not what we want to teach you, and the full Capybara dataset is large enough to take a long time. Fourth, the tokenizer is quietly loaded and wired in as the trainer's processing_class.

TRL's configs differ from plain TrainingArguments in a few defaults that are worth knowing, because they explain behaviour you might otherwise find surprising.

  • logging_steps is 10, not 500, so you see the loss early and often.
  • gradient_checkpointing is on. This trades about a fifth of your speed for a large memory saving by recomputing some values instead of storing them.
  • bf16 is on unless you set fp16. This is a 16-bit number format that is stable for training on modern GPUs.
  • The default learning rate for SFT is 2e-5.
  • max_length defaults to 1024 tokens. Examples longer than that are truncated.

Do not run the basic script on the whole dataset as your first experiment. Let us shrink it so that the whole cycle takes a few minutes, and let us also use LoRA, which is the next section. For now, the point is only to know the shortest path exists.

What happens when you call train() The trainer tokenizes the dataset, groups examples into batches, runs the model forward to get predictions for every next token, compares them to the real next tokens to get the loss, runs backward to compute gradients, and lets the optimiser nudge the trainable weights. It repeats this for every batch. Each repetition is a step. Every ten steps, by default, it prints the loss and a few other numbers.

To keep your first experiment short, you can limit the work with the standard max_steps setting, and slice the dataset. This is the structure of every script you will write from now on.

train_small.py
from datasets import load_dataset
from trl import SFTConfig, SFTTrainer

dataset = load_dataset("trl-lib/Capybara", split="train[:1000]")

config = SFTConfig(
    output_dir="qwen3-sft-full-test",
    max_steps=20,
    per_device_train_batch_size=2,
    report_to="none",
)

trainer = SFTTrainer(
    model="Qwen/Qwen3-0.6B",
    args=config,
    train_dataset=dataset,
)
trainer.train()

Run it with python train_small.py. The slice train[:1000] is datasets syntax for the first thousand rows. Setting report_to="none" turns off experiment tracking, which we will come back to. The first run downloads the model, which takes a minute depending on your connection; later runs reuse the local cache.

You will see a progress bar and, every ten steps, a dictionary of numbers. We read those numbers in a later section. For now, confirm that the loss is a finite number and that the run finishes without an error.

Try it
  1. Save the script as train_small.py and run it.
  2. Watch for the first log line at step 10 and note the loss value.
  3. Open the qwen3-sft-full-test folder and list what is inside.
a loss somewhere above 1 at step 10 and a folder containing a checkpoint. If you run out of memory, move on to the LoRA version; it exists precisely for that.

Adding LoRA with PEFT

Now the central technique. To switch the run above from full fine-tuning to LoRA, you add one argument: a LoraConfig passed as peft_config. TRL sees it, wraps the model with adapters, freezes everything else, and trains only the adapters.

train_lora.py
from datasets import load_dataset
from peft import LoraConfig
from trl import SFTConfig, SFTTrainer

dataset = load_dataset("trl-lib/Capybara", split="train[:2000]")

peft_config = LoraConfig(
    r=16,
    lora_alpha=32,
    lora_dropout=0.05,
    target_modules="all-linear",
    task_type="CAUSAL_LM",
)

config = SFTConfig(
    output_dir="qwen3-sft-lora",
    learning_rate=2e-4,
    per_device_train_batch_size=2,
    gradient_accumulation_steps=8,
    num_train_epochs=1,
    report_to="none",
)

trainer = SFTTrainer(
    model="Qwen/Qwen3-0.6B",
    args=config,
    train_dataset=dataset,
    peft_config=peft_config,
)

trainer.model.print_trainable_parameters()
trainer.train()
trainer.save_model("qwen3-sft-lora")

Let us walk through each choice, because each is a lever you will turn later.

r=16 is the rank. Higher means more capacity and more memory. For learning a style or a format, 8 to 32 is typical. lora_alpha=32 is the scale numerator; with alpha at twice the rank, the correction is multiplied by 2. A common habit is to keep alpha at one or two times the rank and to tune the learning rate instead. lora_dropout=0.05 randomly drops a small share of the adapter's activations during training, a mild guard against overfitting.

target_modules="all-linear" adds adapters to every linear layer of the model except the output layer. Older tutorials target only the attention projections, such as q_proj and v_proj. TRL's own guidance, summarised in its "LoRA Without Regret" page, is that applying LoRA to all weight matrices works better than attention-only even when the total parameter count is matched, so "all-linear" is both the easiest and the recommended choice.

task_type="CAUSAL_LM" tells PEFT this is a next-token-prediction model, so it saves the right things. The other task types, such as sequence classification, matter when you build a classifier.

There is one detail that trips people up. LoraConfig has its own defaults, which are r=8, lora_alpha=8 and lora_dropout=0.0. The TRL command-line tool uses different defaults, 16, 32 and 0.05. So a LoraConfig() with no arguments is not the same as what you get from the CLI with --use_peft. In your own scripts, always write the numbers down explicitly so that you never depend on a default.

The learning rate is the other important change. For full fine-tuning the SFT default is 2e-5. For LoRA the TRL documentation recommends roughly ten times higher, 2e-4. This is a very common surprise: the adapter starts from zero effect and needs bigger steps to move. If you copy a full fine-tuning learning rate into a LoRA run, the loss will crawl.

per_device_train_batch_size=2 with gradient_accumulation_steps=8 gives an effective batch of 16 examples per update on one GPU, while only ever holding two in memory at once. Gradient accumulation means the trainer runs eight small batches, adds up their gradients, and only then updates the weights. It is the standard trick for getting a big batch on a small GPU. The effective batch size is per-device batch size times the number of devices times the accumulation steps.

When the script starts, print_trainable_parameters() prints a line like the following. The exact numbers depend on the model.

TEXT
trainable params: 10,092,544 || all params: 606,142,464 || trainable%: 1.6652

That single line is the whole point of PEFT. You are about to train under two percent of the model's weights. If it prints trainable%: 100, no adapter was attached, so check that you passed peft_config. If it prints a tiny number such as 0.0, your target_modules matched nothing useful.

Do not also pass an already-wrapped model Pass either a peft_config or a model that is already a PEFT model, never both. If you do, TRL stops with an error telling you to merge and unload the existing adapter first. The simple rule for beginners: give the trainer a model name and a peft_config, and let it do the wrapping.
Try it
  1. Run train_lora.py and read the trainable params line.
  2. Change r=16 to r=4, rerun just the print (stop with Ctrl+C after it appears), and compare the trainable percentage.
  3. Change it back to 16 and let the training finish.
a trainable percentage that shrinks roughly in proportion to the rank. Training time barely changes, because most of the cost is in the frozen model. The saving is mostly memory and file size.

Reading the training log

While the training runs, a line is printed every ten steps. A representative line looks like this.

TEXT
{'loss': 1.4213, 'grad_norm': 0.6821, 'learning_rate': 0.000187, 'entropy': 1.31, 'num_tokens': 163840.0, 'mean_token_accuracy': 0.6702, 'epoch': 0.08}

Here is how to read the SFT numbers.

loss is the number to watch. It should start somewhere around 1 to 3 for a chat model on general data and trend downward. It will jump around from line to line, because each line reflects different examples, so judge the trend over many lines rather than any single value. A loss that never moves at all is a warning sign, and so is a loss that falls to nearly zero on a small dataset, which usually means the model is memorising it.

mean_token_accuracy is the fraction of next tokens the model predicts exactly right. It rises as the model learns your data's patterns. It is a friendlier number than loss for a quick sanity check.

grad_norm is the size of the gradient. You mostly care that it is a finite, reasonable number. If it becomes nan or enormous, training has gone unstable, and lowering the learning rate is the first remedy.

To see these over time instead of scrolling a terminal, plug in an experiment tracker. TRL's documentation recommends Trackio, an open-source tracker from Hugging Face; install it and set report_to="trackio". Weights & Biases, TensorBoard and MLflow work through the same report_to setting. Our guides on MLflow and Weights & Biases cover those tools in depth. For a first run, the terminal is enough, and report_to="none" keeps it simple.

One more habit worth forming now: hold some data back. A model can drive its training loss down simply by memorising the examples. If you give the trainer an eval_dataset and set eval_strategy="steps" with an eval_steps value, it will also report an eval_loss on data it never trained on. When training loss keeps falling while eval loss rises, you are overfitting, and you should stop earlier, use less rank, or add more varied data.

Try it
  1. Rerun the LoRA script with train[:2000] split into two parts using dataset.train_test_split(test_size=0.1).
  2. Pass the two halves as train_dataset and eval_dataset, and add eval_strategy="steps" and eval_steps=20 to the config.
  3. Find the eval_loss lines in the output.
an eval_loss every twenty steps that tracks the training loss at first. That pairing is how you spot overfitting early.

Saving, loading and merging adapters

When training finishes, trainer.save_model("qwen3-sft-lora") writes the result. Because you used PEFT, it saves only the adapter, not the base model. Look inside the folder.

BASH
ls -lh qwen3-sft-lora

You will find adapter_config.json, which records the LoRA settings and the name of the base model, and adapter_model.safetensors, which holds the trained numbers. The folder is a few tens of megabytes instead of gigabytes. That is the practical magic: you can email it, keep fifty of them, or push each to the Hub.

An adapter is useless without its base model, so to use it you load both. The right tool for loading a trained adapter is PeftModel.from_pretrained.

load_adapter.py
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-0.6B", dtype="auto")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-0.6B")
model = PeftModel.from_pretrained(base, "qwen3-sft-lora")
model.eval()

messages = [{"role": "user", "content": "Explain gradient accumulation in two sentences."}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
).to(model.device)

output = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Note dtype="auto". Older code wrote torch_dtype; the current name is dtype. Note also model.eval(), which switches off dropout so that the output is deterministic in character. A shortcut is AutoPeftModelForCausalLM.from_pretrained("qwen3-sft-lora"), which reads the base model name from adapter_config.json and loads both in one call.

The classic mistake: get_peft_model on a trained adapter get_peft_model(model, config) creates a new, randomly initialised adapter. It is for starting training. To use an adapter you already trained, always use PeftModel.from_pretrained. If you mix them up, nothing crashes, the model simply behaves like the base model, and you spend an afternoon wondering why fine-tuning "did nothing".

Sometimes you want a single standalone model with no adapter, perhaps to hand to a serving system that does not know about PEFT. For that you merge: PEFT folds the adapter's correction into the base weights and removes the adapter machinery.

merge.py
merged = model.merge_and_unload()
merged.save_pretrained("qwen3-merged")
tokenizer.save_pretrained("qwen3-merged")

The qwen3-merged folder is now an ordinary Hugging Face model that loads with AutoModelForCausalLM.from_pretrained("qwen3-merged") and that tools such as vLLM can serve. The trade-off is size: a merged model is a full copy, gigabytes instead of megabytes. Keep the small adapter as your source of truth and merge only when you need to.

If you trained with a quantized base model (the QLoRA technique later in this guide), reload the base in full or bf16 precision before merging. Merging into a 4-bit base introduces rounding error.

Try it
  1. Run load_adapter.py and read the answer.
  2. Comment out the PeftModel.from_pretrained line, use the base model alone, and compare answers.
  3. Run the merge, and compare the folder sizes of the adapter and the merged model with du -sh.
an adapter folder in megabytes and a merged folder in gigabytes. The answers from base and tuned will differ in style if your data was distinctive.

Using the command line instead of a script

You do not have to write Python for standard runs. The trl command wraps the trainers, and it is ideal for quick experiments and for runs you want to record as a one-line command in a README. The subcommands are trl sft, trl dpo, trl grpo, trl rloo, trl kto, trl reward, trl distillation, and trl env, which you have already met.

The same LoRA run as our script looks like this.

BASH
trl sft \
  --model_name_or_path Qwen/Qwen3-0.6B \
  --dataset_name trl-lib/Capybara \
  --output_dir qwen3-sft-cli \
  --use_peft \
  --lora_r 16 \
  --lora_alpha 32 \
  --lora_target_modules all-linear \
  --learning_rate 2e-4 \
  --per_device_train_batch_size 2 \
  --gradient_accumulation_steps 8 \
  --num_train_epochs 1

The flag names are the same as the Python argument names, which is easy to remember. For longer runs, put the settings in a YAML file whose keys are the flag names without dashes, then point the command at it.

sft_config.yaml
model_name_or_path: Qwen/Qwen3-0.6B
dataset_name: trl-lib/Capybara
output_dir: qwen3-sft-cli
use_peft: true
lora_r: 16
lora_alpha: 32
lora_target_modules: all-linear
learning_rate: 2.0e-4
per_device_train_batch_size: 2
gradient_accumulation_steps: 8
num_train_epochs: 1
BASH
trl sft --config sft_config.yaml
Try it
  1. Run trl sft --help and scroll to find the LoRA flags.
  2. Write sft_config.yaml as above with max_steps: 20 added, and run it.
a short run that finishes and writes an adapter to qwen3-sft-cli, with no Python script involved.

The settings that matter most

Beyond the LoRA choices, a handful of config settings account for most of what you will tune. Learn these before the hundred others.

Learning rate. The single most influential number. For LoRA SFT, start at 2e-4. If the loss spikes or turns into nan, halve it. If the loss barely moves after a hundred steps, double it.

Epochs or steps. num_train_epochs counts passes over the data. max_steps counts updates and, if set, wins over epochs. For small datasets, one to three epochs is typical. More epochs on a small dataset are the commonest route to overfitting.

Batch size and accumulation. per_device_train_batch_size is limited by memory, so make it as large as fits, then use gradient_accumulation_steps to reach an effective batch of about 16 to 32. TRL's guidance for LoRA is to keep the effective batch below about 32.

max_length. The longest sequence, in tokens, that the trainer will keep. Default 1024. Longer examples are cut off at the end, which silently removes the end of your answers if your data is long. Memory grows quickly with this number, so raise it deliberately, not by habit.

packing. Setting packing=True concatenates several short examples into one max_length sequence so that no compute is wasted on padding. It can speed up training a lot when your examples are short. It needs max_length to be set, and it is not supported for image data.

loss_type. Since TRL 1.7 the SFT default is chunked_nll, which has the same maths as plain NLL but computes the loss in chunks so that peak memory is lower, by around 30 percent on average. You normally leave it alone. It does have one interaction worth knowing: it does not work if you put an adapter on the lm_head layer, which target_modules="all-linear" avoids by excluding the output layer.

Here is a config that sets the important choices explicitly, which is a good habit: it documents your run and protects you when a default changes in a future release. Several TRL defaults did change in the past year, which is the practical reason to be explicit.

config_explicit.py
from trl import SFTConfig

config = SFTConfig(
    output_dir="qwen3-sft-lora",
    learning_rate=2e-4,
    num_train_epochs=1,
    per_device_train_batch_size=2,
    gradient_accumulation_steps=8,
    max_length=1024,
    packing=False,
    logging_steps=10,
    save_steps=200,
    seed=42,
    report_to="none",
)
Try it
  1. Run the LoRA script twice with max_steps=60: once at learning_rate=2e-5 and once at 2e-4.
  2. Compare the final loss and mean_token_accuracy of the two runs.
the higher rate reaching a lower loss in the same number of steps. That is the "LoRA wants a bigger learning rate" lesson in numbers.

QLoRA: fitting a bigger model on a small GPU

LoRA shrinks what you train, but the frozen base model still has to sit in memory. For a small model that is fine. For a seven-billion-parameter model it is not, and that is where QLoRA comes in. QLoRA first quantizes the frozen base model to 4-bit numbers, which cuts its memory to roughly a quarter, and then trains LoRA adapters on top of it in higher precision. Quantization means storing each weight with fewer bits, accepting a small loss of precision in exchange for a large saving of memory. The library that does the 4-bit work is bitsandbytes.

In current TRL the way to ask for this is the quantization_config argument on the trainer, available since TRL 1.8. Install the extra first.

BASH
pip install "trl[quantization]"

Then pass a BitsAndBytesConfig.

train_qlora.py
import torch
from datasets import load_dataset
from peft import LoraConfig
from transformers import BitsAndBytesConfig
from trl import SFTConfig, SFTTrainer

bnb = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
)

trainer = SFTTrainer(
    model="Qwen/Qwen3-0.6B",
    args=SFTConfig(output_dir="qwen3-qlora", learning_rate=2e-4, report_to="none"),
    train_dataset=load_dataset("trl-lib/Capybara", split="train[:2000]"),
    quantization_config=bnb,
    peft_config=LoraConfig(r=16, lora_alpha=32, target_modules="all-linear", task_type="CAUSAL_LM"),
)
trainer.train()

nf4 is a 4-bit number format designed for the bell-shaped distribution of neural network weights. Double quantization squeezes the bookkeeping numbers too. bnb_4bit_compute_dtype says that the arithmetic itself is done in bf16, even though the stored weights are 4-bit.

Two rules keep you out of trouble. Set quantization_config as a trainer argument only, not also inside model_init_kwargs, or the trainer raises an error. And if you use raw PEFT without TRL, you must call prepare_model_for_kbit_training(model) before get_peft_model; TRL does that preparation for you.

Try it
  1. On an NVIDIA GPU, run train_qlora.py with max_steps=20 added to the config.
  2. Watch memory with nvidia-smi in a second terminal and compare to the plain LoRA run.
lower GPU memory use than the plain LoRA run. With a model this small the saving is modest; with a 7B model it decides whether the job fits at all.

Common errors and how to read them

Every beginner meets the same handful of failures. Read the last line of a Python traceback first, since the exception type and message sit there, then match it here.

torch.OutOfMemoryError: CUDA out of memory. The model, the sequence length or the batch is too big for the GPU. Work down this list: set per_device_train_batch_size=1 and raise gradient_accumulation_steps to compensate; lower max_length; use LoRA, then QLoRA; pick a smaller model. Check that nothing else is using the GPU with nvidia-smi, since a forgotten notebook can hold all the memory.

ImportError: You passed peft_config but the peft library is not installed. Install it with pip install trl[peft]. You installed trl without the extra. Run pip install "trl[peft]".

TypeError: peft_config must be a peft.PeftConfig instance (e.g. peft.LoraConfig), got dict. You passed a dictionary. Build a LoraConfig(...) object instead.

KeyError: The input examples must contain either 'messages' for conversational data or 'text' for standard data. Your dataset's columns are not what SFT expects. Print dataset.column_names and rename a column to text or messages, or use prompt and completion.

ValueError: You passed a PeftModel instance together with a peft_config to the trainer. Pick one: give a plain model name with peft_config, or give an already-wrapped model without it.

ValueError: loss_type='chunked_nll' is not supported when lm_head is wrapped by a PEFT adapter. Something targeted the output layer. Remove lm_head from target_modules, or set loss_type="nll".

ValueError: When packing is enabled, max_length can't be None. Give max_length a number when packing=True.

RuntimeError: You're using assistant_only_loss=True, but at least one example has no assistant tokens. The chat template has no markers for assistant tokens. Leave assistant_only_loss off, or use a model family whose template TRL patches.

ValueError: Attempting to unscale FP16 gradients. The model was loaded in 16-bit float and trained with fp16 mixed precision, leaving trainable weights in 16-bit. Prefer bf16 on modern GPUs, which is TRL's default.

ModuleNotFoundError: No module named 'trl.losses' or ImportError: cannot import name 'PPOTrainer' from 'trl'. You are running code written for an older TRL. Those pieces were removed in 2026. PPO was dropped in TRL 1.13, and for online reinforcement learning the library now points you to GRPO or RLOO. Do not fight this; port the code to a current trainer.

Then there are the silent problems, which print no error at all.

  • The loss barely moves. Some of your target_modules names did not match any layer. PEFT raises an error only if none match, and quietly skips partial mismatches. Prefer "all-linear", and check print_trainable_parameters().
  • The tuned model acts like the base model. You probably used get_peft_model where you needed PeftModel.from_pretrained, or forgot model.eval(), or pointed at the wrong folder.

Finally, a pragmatic tip: when you must ask for help, run trl env and paste its output together with the full traceback and the exact versions. The maintainers ask for that.

Try it
  1. Deliberately set per_device_train_batch_size=64 and max_length=4096 and run it.
  2. Read the last line of the traceback and note the memory numbers.
  3. Fix it using the steps above and confirm the run starts.
an out-of-memory error you can read without panic, followed by a run that works. Seeing a failure on purpose makes the next real one routine.

A first look beyond SFT: DPO and GRPO

SFT teaches a model to imitate. Sometimes you want to teach it to prefer. That is what the other TRL trainers are for, and you only need a taste of them now.

DPO, Direct Preference Optimization, learns from pairs. Each training row has a prompt, a chosen answer people liked and a rejected answer they did not. The model learns to make the chosen one more likely than the rejected one, without any separate scoring model. A frozen reference model acts as an anchor so that the model does not wander too far from where it started, and a number called beta, default 0.1, controls how tight that anchor is. When you use PEFT, TRL does not load a second copy of the model: the reference is simply the base model with the adapter switched off.

train_dpo.py
from datasets import load_dataset
from peft import LoraConfig
from trl import DPOConfig, DPOTrainer

trainer = DPOTrainer(
    model="Qwen/Qwen2.5-0.5B-Instruct",
    args=DPOConfig(output_dir="dpo-lora", learning_rate=5e-6, beta=0.1, report_to="none"),
    train_dataset=load_dataset("trl-lib/ultrafeedback_binarized", split="train[:1000]"),
    peft_config=LoraConfig(task_type="CAUSAL_LM"),
)
trainer.train()

Notice that the learning rate is far smaller than in SFT. DPO is sensitive, and TRL's recommendation for LoRA is about 5e-6.

GRPO, Group Relative Policy Optimization, is online. For each prompt the model writes several answers, a reward function that you write scores each one, and the model learns to produce more of the answers that scored above the group average. It is how models are taught to reason in ways that can be checked automatically, such as getting a maths answer right. A reward function is just Python: it receives the completions and returns one number per completion. GRPO needs more memory and, ideally, a generation engine such as vLLM, so leave it until the basics are second nature. The mid-level guide covers it properly.

Try it
  1. Load trl-lib/ultrafeedback_binarized and print one example.
  2. Identify the prompt, the chosen and the rejected answer, and decide in a sentence why the chosen one is better.
a row with three fields. Reading preference data by hand shows you what "better" is being taught, and how subjective it can be.

Putting it all together

Here is one small, complete project that uses every piece above. The goal: take your "behaviour sentence" from the first section, teach it to a small model, and prove that the change took.

First, write the data. Make about forty examples in a JSON Lines file called my_data.jsonl, one JSON object per line, each with a messages list. If your behaviour is "reply in two short sentences", write forty varied questions with two-sentence answers. Vary the topics, because a model trained on near-duplicates learns only the duplicates. Here are the first lines of such a file.

my_data.jsonl
{"messages": [{"role": "user", "content": "What is a GPU?"}, {"role": "assistant", "content": "A GPU is a chip built to do many simple calculations in parallel. That makes it ideal for graphics and for training neural networks."}]}
{"messages": [{"role": "user", "content": "What is Docker?"}, {"role": "assistant", "content": "Docker packages an application with everything it needs into a portable image. You can then run that image as a container on any machine."}]}

Second, the training script, which uses LoRA and holds back a few rows for evaluation.

final_project.py
from datasets import load_dataset
from peft import LoraConfig
from trl import SFTConfig, SFTTrainer

data = load_dataset("json", data_files="my_data.jsonl", split="train")
splits = data.train_test_split(test_size=0.1, seed=42)

trainer = SFTTrainer(
    model="Qwen/Qwen3-0.6B",
    args=SFTConfig(
        output_dir="two-sentence-adapter",
        learning_rate=2e-4,
        num_train_epochs=3,
        per_device_train_batch_size=2,
        gradient_accumulation_steps=2,
        eval_strategy="epoch",
        logging_steps=5,
        seed=42,
        report_to="none",
    ),
    train_dataset=splits["train"],
    eval_dataset=splits["test"],
    peft_config=LoraConfig(
        r=16, lora_alpha=32, lora_dropout=0.05,
        target_modules="all-linear", task_type="CAUSAL_LM",
    ),
)

trainer.model.print_trainable_parameters()
trainer.train()
trainer.save_model("two-sentence-adapter")

Third, compare the base model and the tuned model on a question that is not in your data. Reuse load_adapter.py with the new folder name, ask something new, then ask the same thing again with the adapter removed, by using the plain base object. If the tuned answer is two short sentences and the base answer rambles, your fine-tune worked.

Fifth, write down what you did. Record the library versions from trl env, the dataset file, the config values and the result of your before-and-after check in a short README beside the adapter. Six weeks from now, with different library versions installed, this note is the difference between a repeatable result and a mystery.

Try it
  1. Create your forty-row file, run the script, and compare base and tuned answers on three unseen questions.
  2. Increase to 120 rows, rerun, and see whether the tuned behaviour becomes more reliable.
a tuned model that follows your format on questions it never saw, and noticeably better consistency with more data. That is fine-tuning working end to end.

What you can now do, and what comes next

You can now install TRL and PEFT and verify the setup, explain in plain words what a base model, an adapter, LoRA, rank and a trainer are, and prepare a dataset in the shape each trainer expects. You can run supervised fine-tuning with LoRA or QLoRA from a script or from the command line, read the loss and token accuracy in the training log, and hold out data to spot overfitting. You can save an adapter, load it correctly, merge it into a standalone model, and recognise the errors that stop most first attempts. And you know where DPO and GRPO fit, so you will not mistake them for beginner topics or for things you must learn today.

The natural next steps follow from here. Mid-level covers preference tuning with DPO in earnest, GRPO with your own reward functions, memory tuning, multi-GPU training with accelerate, and experiment tracking. Senior covers running training and serving adapters as a platform for a team. Alongside this guide, the Hugging Face guide covers the Hub, datasets and transformers underneath, Unsloth offers a faster fine-tuning stack built around the same ideas, DeepSpeed is how very large models are spread over many GPUs, and vLLM is how you serve the model you just trained.

Sources