Skip to content
Back to student guides
UnslothLLMsFine-tuning & training3 levels100 sectionsCovers Unsloth 2026.9

The Complete Unsloth Guide

Fine-tune open LLMs faster and with less GPU memory using Unsloth. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
16sections
27examples

This is part one of three. It covers everything you need to fine-tune and run an open language model with Unsloth, assuming you have never used the tool and may never have trained a model at all. By the end you can install Unsloth in the way that suits your machine, chat with a model that runs entirely on your own hardware, fine-tune a small model on your own examples, test the result, and export it so that other tools can run it. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. Fine-tuning feels abstract until you have watched a loss curve fall on your own data, and that takes about ten minutes once the setup is done.

What Unsloth is, and what you can do by the end

Unsloth is an open-source toolkit that makes three jobs cheaper: fine-tuning an open model (teaching it your style, your format or your domain), reinforcement learning on a model (rewarding it for good answers), and running models locally. Its headline claim from the official documentation is that it trains models about twice as fast while using about 70% less video memory, with no loss in accuracy. Video memory, or VRAM, is the memory on your graphics card, and it is the scarce resource in this whole field. A card with 8 GB of VRAM cannot normally fine-tune an 8-billion-parameter model. With Unsloth's techniques it can come close, which is why students, hobbyists and small teams use it.

It is worth being precise about what Unsloth is not. It is not a model. It does not come with its own language model that you can talk to. It is also not a replacement for the Hugging Face ecosystem. It sits on top of Hugging Face transformers, trl and peft, patches parts of them with faster hand-written code, and hands you back objects that behave like the ones you already know. If you later read a tutorial about fine-tuning with plain Hugging Face libraries, almost everything transfers.

By the end of this guide you will be able to do the following.

  1. Choose between the three ways to install Unsloth and set one up.
  2. Check that your GPU is visible and that the install is healthy.
  3. Download a model, chat with it locally, and understand what a "quantization" label such as Q4_K_M means.
  4. Fine-tune a small model on a few dozen examples, first in the graphical Studio and then in about forty lines of Python.
  5. Read the training loss and know whether the run went well.
  6. Save your result as a small adapter, a merged model, or a GGUF file for tools such as Ollama.
  7. Recognise the dozen most common error messages and fix them.

One note on versions. Unsloth moves quickly: the Python package is versioned by date and a new release arrives most weeks. This guide was checked against the unsloth Python package 2026.9.12 and the Unsloth Studio and Desktop release v0.1.900-beta. If you are reading this months later, run the commands in the "Everyday commands" section to see your own version, and treat the official changelog as the source of truth whenever something here disagrees with your screen.

Try it
  1. Find out what graphics card your computer has. On Windows open Task Manager, then Performance, then GPU. On Linux run nvidia-smi if you have an NVIDIA card. On a Mac, open About This Mac.
  2. Write down the card name and its memory in gigabytes.
  3. Keep the number. Every decision later in this guide, from which model to pick to which method to use, starts from it.
a card name and a memory figure. If you have no dedicated GPU, do not stop: free Google Colab notebooks and Kaggle give you a T4 card, and the chat features in Studio work on a CPU.

The problem it solves, and what came before

To see why Unsloth exists you need a little background on why training a language model is so hungry for memory.

A language model is a very large collection of numbers, called parameters or weights. An "8B" model has eight billion of them. Stored in 16-bit precision, each takes two bytes, so the model alone needs roughly 16 GB just to sit in memory. That is before you do anything with it.

Training needs far more than the weights. During training the software also keeps gradients (one number per weight saying which way to nudge it), optimizer state (extra bookkeeping the training algorithm needs, often twice the size of the weights) and activations (the intermediate results of the forward pass, which grow with the length of your text). Full fine-tuning, where every weight is updated, therefore needs several times the model's own size in memory. For an 8B model that means well over 100 GB, which is data-centre hardware.

Two ideas from research brought this within reach of ordinary machines, and Unsloth builds on both.

LoRA (Low-Rank Adaptation) freezes the original weights and trains only a pair of thin extra matrices attached to each layer you choose. The extra matrices hold about 1% of the number of parameters, so gradients and optimizer state shrink to almost nothing. The result is a small file, called an adapter, that you apply on top of the original model.

Quantization stores the frozen weights in fewer bits. QLoRA combines the two: load the base model in 4-bit precision, then train a LoRA adapter on top. The docs say QLoRA uses over 75% less VRAM than full fine-tuning, and they recommend that you start there.

Before Unsloth, people ran QLoRA with the standard Hugging Face stack. It worked, but it was slower than it needed to be and wasteful with memory, because general-purpose libraries are written to support every model and every situation. Unsloth's authors rewrote the hot paths by hand: custom GPU kernels written in a language called Triton, a manually derived backward pass, and a memory-saving form of gradient checkpointing that moves activations out to system RAM while the GPU works. They also publish pre-quantized copies of popular models so that you do not have to quantize them yourself. The numbers they report, 2x speed and 70% less VRAM, come from these optimisations.

A newer part of the story is that Unsloth is no longer only a Python library. In 2026 it gained a browser-based interface called Studio and a native Desktop app, so you can chat with models, prepare data and launch training without writing code. Studio and Desktop are still labelled beta, while the Python library is the mature part and the one production pipelines use. This guide teaches both, because a beginner benefits from seeing a training run in a graphical interface first and then understanding the code that does the same thing.

Try it
  1. Think of one task you would like a small model to do consistently, such as replying to support emails in a fixed format or classifying messages into categories.
  2. Write one sentence describing the input and one describing the ideal output.
  3. Ask yourself whether the model needs to behave differently (fine-tune) or know something new (retrieve).
a clear one-line task. You will turn it into a dataset later in this guide.

The mental model: six nouns

Unsloth has a lot of vocabulary, but a beginner needs only six words. Learn them now and the rest of the documentation becomes readable.

BASE MODELfrozen weights
→
ADAPTERLoRA, trained
→
TRAINERruns the loop
→
EXPORTadapter, merged or GGUF

Base model. The pretrained model you start from, usually downloaded from the Hugging Face Hub, the public website where models and datasets are shared. Its identifier looks like unsloth/gemma-3-270m-it, meaning the organisation unsloth and a model called gemma-3-270m-it. The letters "it" mean instruction-tuned: the model was already taught to follow instructions and hold a conversation.

Adapter (LoRA). The small set of new weights you train. Two numbers describe its size. The rank, written r, sets its capacity; the docs recommend 16 or 32. The alpha, written lora_alpha, scales how strongly the adapter influences the model; the recommendation is to set it equal to the rank, or twice the rank. You also choose which layers receive adapters, called the target modules. The recommendation is to target every attention and feed-forward layer, which in Unsloth's notation is the seven names q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj and down_proj.

Quantization. Storing weights in fewer bits. Two flavours matter to you, and they are easy to confuse. Training-time quantization is the 4-bit loading used by QLoRA, controlled by load_in_4bit=True. Inference-time quantization is the GGUF file format used by llama.cpp, Ollama and Studio's chat, whose variants have names such as Q4_K_M. A GGUF model can be chatted with but cannot be trained. You train a normal model, then convert the result to GGUF for running.

Dataset. Your examples. For a chat model each example is a short conversation: a user message and the answer you want. Quality beats quantity here. A few hundred clean examples teach a consistent format better than tens of thousands of messy ones.

Chat template. Every model family expects its conversations written in a particular format, with special marker tokens separating who is speaking. The template is the rule for producing that format. Using a different template when you run the model than the one you trained with is the single most common reason a model "worked in training but produces gibberish afterwards", and it will appear again in the errors section.

Trainer. The object that runs the training loop. Unsloth does not ship its own: you use SFTTrainer from the TRL library, where SFT stands for supervised fine-tuning, learning from examples of the right answer. Unsloth patches it behind the scenes. Studio runs the same trainer for you.

You will also see three more terms in the documentation. Epoch means one pass over your whole dataset. Step means one update of the weights. Loss is the number that measures how wrong the model's predictions were on the training text; training tries to push it down.

Three loader classes appear in code. FastLanguageModel loads text models and is the one you will use here. FastVisionModel loads models that read images. FastModel is a general loader for other architectures, such as audio models. For this level you only need the first.

Finally, a rule that trips up everyone once: import unsloth before transformers, trl or peft. Unsloth works by patching those libraries when it is imported, so if they are imported first the patches cannot apply. A banner that begins "Unsloth: Will patch your computer to enable 2x faster free finetuning" confirms the patching happened.

Try it
  1. Without looking back, write one sentence each for: base model, adapter, dataset, chat template.
  2. Then look at the model name unsloth/Llama-3.1-8B-Instruct-unsloth-bnb-4bit and decide which of the six nouns it belongs to.
a base model, stored in 4-bit form. The "bnb-4bit" ending tells you it is already quantized for QLoRA.

Three ways to install it

The official install page says Unsloth can be installed in three distinct ways. They are not competitors; they serve different people, and you can have more than one on the same computer.

Install What it is Best for
Desktop A native app for macOS, Windows and Linux The easiest start; no terminal needed
Studio A web interface you launch from the terminal Chat, data prep and training in a browser, including on a remote machine
Core The Python package, pip install unsloth Notebooks, scripts and anything you will automate

Desktop was introduced in August 2026. Download the installer from unsloth.ai/download/mac, unsloth.ai/download/windows or unsloth.ai/download/linux. On macOS you drag the app to Applications; on Windows you run the .exe; on Linux you get a .deb package or an AppImage. It updates itself from Settings, then General, then Check for updates, except the Linux .deb, which you reinstall manually when a new version arrives.

Studio is installed by a script. On macOS, Linux and WSL (the Windows Subsystem for Linux) run:

BASH
curl -fsSL https://unsloth.ai/install.sh | sh

On Windows PowerShell:

POWERSHELL
irm https://unsloth.ai/install.ps1 | iex

Piping a downloaded script into a shell runs whatever the server sends, so read it first if you are cautious. You can open https://unsloth.ai/install.sh in a browser before running it.

You then launch Studio from a terminal:

BASH
unsloth studio -p 8888

Open http://127.0.0.1:8888 in your browser. On first launch you are sent to a /change-password page to set the admin password. Studio listens only on your own machine by default (127.0.0.1), which is the safe choice. Later you will see tutorials that add -H 0.0.0.0; that makes Studio reachable from other machines on your network over unencrypted HTTP, so leave it off until you understand the risk.

Requirements for Studio and Desktop: Windows 10 or 11 (no WSL needed), macOS 12 or later on Intel or Apple Silicon, or Linux such as Ubuntu 20.04 and later, and Python 3.11 to 3.13. For training on an NVIDIA GPU you need the driver installed and the CUDA toolkit version 12.4 or later (12.8 or later for the newest Blackwell cards). A machine with no GPU can still use the chat feature and Data Recipes.

Core is the Python package. The recommended setup uses uv, a fast Python package manager, and a virtual environment so that Unsloth's carefully pinned dependencies do not collide with other projects:

BASH
curl -LsSf https://astral.sh/uv/install.sh | sh
uv venv unsloth_env --python 3.13
source unsloth_env/bin/activate
uv pip install unsloth --torch-backend=auto

The --torch-backend=auto flag makes uv choose the PyTorch build that matches your GPU driver. On Windows, install uv with irm https://astral.sh/uv/install.ps1 | iex and activate with unsloth_env\Scripts\activate. Plain pip install unsloth also works if you prefer it.

Core needs an NVIDIA GPU with CUDA capability 7.0 or higher. That covers V100, T4, the RTX 20 series and everything newer. The docs note that Core support for Apple Silicon is still in the works; on a Mac, use Studio or Desktop, which support training and running models through Apple's MLX framework. AMD and Intel GPUs have their own installation guides.

If you would rather not install anything, the official Colab notebooks run Unsloth on a free T4 GPU in your browser, and there is also a Docker image, unsloth/unsloth, for NVIDIA cards. Docker is covered in the Docker guide; for now, know that the container bundles Studio, which listens on port 8000 inside it, and JupyterLab on 8888.

Do not update Studio with unsloth studio update The changelog says plainly not to use that command any more, because packaging will not fetch the latest updates. To update Studio, run the install one-liner again. This is a good example of why you should distrust a command from an old blog post: the tool you are reading about may have changed how it works since.
Try it
  1. Pick one install route that matches your machine: Desktop if you want to avoid the terminal, Studio if you are comfortable with one command, Core if you plan to write Python.
  2. Install it, following only the official steps above.
  3. Write down which route you picked and the exact command or installer you used.
an installed tool and a note you can repeat on another machine. Keep the note: it is your first reproducible setup.

Checking your hardware and setup

Before you train anything, spend five minutes verifying the foundations. Most "Unsloth is broken" reports are actually "my GPU is not visible to Python".

On an NVIDIA machine, start with the driver:

BASH
nvidia-smi

This prints a table with your card's name, its total and used memory, and the driver version. If the command is not found, the NVIDIA driver is not installed, and nothing else will work until it is.

Next, confirm that PyTorch, the library underneath everything, can see the card. Run this inside your Core environment:

BASH
python -c "import torch; print(torch.cuda.is_available())"

You want True. If you see False, PyTorch was installed without GPU support, or the driver and the CUDA runtime do not match. The Studio installer detects a CPU-only PyTorch and repairs it, but with Core you must reinstall, and the --torch-backend=auto flag exists to prevent exactly this.

For a Studio install, Unsloth ships its own checker:

BASH
unsloth studio verify-install

It exits with code 0 when the install is complete and 1 when something is missing. Add --json for machine-readable output. If you are missing build tools you will see messages such as nvcc not found, cmake not found or git not found. Studio uses cmake and a compiler to build the llama.cpp engine that runs GGUF models when no prebuilt binary is available, so on Ubuntu you install them with sudo apt install cmake git build-essential. A failed llama-server build is not fatal; it only means GGUF chat is unavailable until you fix it.

For Core, the quickest check is to import the library:

PYTHON
from unsloth import FastLanguageModel

A healthy import prints a banner that looks like this:

TEXT
🦥 Unsloth: Will patch your computer to enable 2x faster free finetuning.
==((====))==  Unsloth 2026.x.y: Fast Gemma3 patching. Transformers: ...
   \\   /|    NVIDIA GeForce RTX 3060. Num GPUs = 1. Max memory: 12.0 GB. Platform: Windows.
O^O/ \_/ \    Torch: ... CUDA: 8.6. CUDA Toolkit: 13.0. Triton: ...
\        /    Bfloat16 = TRUE. FA [Xformers = ... FA2 = False]

Read it as a status report. The second line names the Unsloth version. The third line names your GPU, how many Unsloth sees, and how much memory is available. The line starting O^O/ shows the PyTorch and CUDA versions in use. Bfloat16 = TRUE means your card supports a number format that makes training more stable; older cards such as the T4 will say FALSE and use float16 instead, which is fine.

Finally, pick a model size that fits. The documentation publishes absolute minimum VRAM for fine-tuning, and it is the most useful table in this guide:

Model size QLoRA (4-bit) LoRA (16-bit)
3B 3.5 GB 8 GB
7B 5 GB 19 GB
8B 6 GB 22 GB
14B 8.5 GB 33 GB
27B 22 GB 64 GB
70B 41 GB 164 GB

These are minimums. A longer context or a larger batch pushes the real figure higher, so leave a margin. A free Colab T4 has about 15 GB and therefore handles an 8B or even a 14B model with QLoRA, at modest context lengths.

Try it
  1. Run nvidia-smi and, if you installed Core, the PyTorch one-liner above.
  2. Using the table, decide the largest model you could QLoRA fine-tune on your card, keeping a 30% margin.
  3. If you installed Studio, run unsloth studio verify-install and confirm it exits cleanly.
a printed table of your GPU, True from PyTorch, and a model size chosen with evidence rather than hope.

Your first model: chatting in Studio

Before training, use Studio or Desktop as a local chat client. This teaches you what a "model" is in practice and confirms your install works end to end. The steps are the same in both.

Launch the app (Desktop from your applications list, Studio with unsloth studio) and start a new chat. Choose Select model, or open the model hub, and pick a small model. For a first run, anything described as a few billion parameters or fewer is good. Studio will offer GGUF variants, which are the inference-optimised files described earlier.

You will see names such as unsloth/gemma-3-270m-it-GGUF followed by a variant label. The labels follow a pattern worth decoding.

  • Q4_K_M is a 4-bit quantization of medium quality. It is the sensible default: small, fast, and close to the original in answer quality. Since release v0.1.812-beta, Studio's default GGUF export is Q4_K_M.
  • Q8_0 is 8-bit. Larger and slower, slightly more faithful.
  • F16 or BF16 keeps full 16-bit precision. Biggest file, best fidelity.
  • UD-Q4_K_XL and similar names starting with UD are Unsloth Dynamic quantizations. Unsloth keeps the most sensitive layers at higher precision and squeezes the rest harder, aiming for better accuracy at the same file size.

Studio marks variants that will not fit in your memory with an OOM label, short for out of memory. Avoid those. When in doubt, pick Q4_K_M.

When you send your first message, Studio downloads the model into the Hugging Face cache at ~/.cache/huggingface/hub and then starts the llama.cpp server that actually runs it. The first reply is slow because of the download and load; later replies are fast. Studio stores its own state, such as chat history and settings, under ~/.unsloth/studio.

There is also a terminal shortcut that does the same thing and prints an endpoint you can use from other programs:

BASH
unsloth run --model unsloth/gemma-3-270m-it-GGUF:Q4_K_M

The text after the colon selects the variant. The command loads the model, starts the server and interface, and prints the URL and an API key. Studio exposes an API compatible with OpenAI's, so any client written for that API can point at it. The endpoints include /v1/chat/completions and /v1/models on the same port, and the key you send in an Authorization: Bearer header starts with sk-unsloth-. You can create such keys in Settings, then API; a key is shown only once, so copy it when it appears.

Try one test of your own and keep it. Ask the model to do the task you chose earlier, using a plain prompt. Save its answer. This is your baseline: after fine-tuning you will ask the same question and compare. Without a baseline you cannot tell whether training helped.

Try it
  1. Load a small GGUF model in Studio or Desktop and send it your task prompt.
  2. Copy the answer into a text file called baseline.txt.
  3. Note which quantization variant you chose and how large the download was.
a working local chat and a saved baseline. Everything from here is measured against that file.

Your first fine-tune in Studio

Now the main event. Studio guides you through training in a handful of choices. We will walk through them in order and explain what each one controls, so that you are not just clicking defaults.

1. Pick the model type. Studio asks what kind of model you are training: Text, Vision, Audio or Embeddings. Choose Text.

2. Pick the base model. Choose a non-GGUF model, because GGUF files are inference-only and cannot be trained. A name such as unsloth/gemma-3-270m-it works. Prefer the unsloth/ uploads: they come with fixes for tokenizers and chat templates that the original uploads sometimes lack.

3. Pick the method. You choose between QLoRA, LoRA and Full. Choose QLoRA. It uses a 4-bit base model and the least memory, and the documentation recommends starting there. LoRA keeps a 16-bit base and needs much more VRAM, and Full updates every weight and needs the most.

4. Pick the dataset. You can pull one from the Hugging Face Hub by name, or upload your own file. Studio accepts PDF, DOCX, JSONL, JSON, CSV and Parquet. You also tell Studio the format of your data: auto, alpaca, chatml or sharegpt. Leave it on auto first. These names describe how the conversations are stored. In ChatML, for instance, each message is an object with a role (such as user or assistant) and content. We return to formats in the dataset section.

5. Set the hyperparameters. A hyperparameter is a setting you choose before training rather than something training learns. The defaults are sensible for a first run:

  • Max Steps: the default is 0, which means "use epochs instead". For a first test you can set this to a small number such as 60 to finish quickly.
  • Context Length: default 2048, adjustable from 512 to 32768. It is the longest text, counted in tokens, that the model sees at once. Longer context uses much more memory.
  • Learning rate: default 2e-4, which is 0.0002. This is the size of each weight update, and it is the standard starting point for LoRA and QLoRA.

6. Start training. Studio shows a status panel with live charts of loss, gradient norm, learning rate and GPU use. The page works from a phone, so you can watch a long run from the sofa.

7. Export. When training ends you can chat with the result inside Studio, then export. The export choices are Merged 16-bit, LoRA only or GGUF. We explain the difference in the saving section.

What should you see? On a tiny model and small dataset, the loss should start somewhere above 1.0 and drift downward over the first few dozen steps. Do not worry if the line is noisy; individual steps are measured on small batches. The documentation says a loss of roughly 0.5 to 1.0 is usually healthy for this kind of training. A loss that falls close to zero suggests the model is memorising your examples, a problem called overfitting, and it will perform badly on anything new.

Try it
  1. Pick a tiny non-GGUF model, choose Text and QLoRA, and select a small dataset from the Hub or upload a short JSONL file of your own.
  2. Set Max Steps to 60 and leave the learning rate at 2e-4.
  3. Start the run and watch the loss chart for a minute.
a loss curve that trends down. If the run stops with an out-of-memory message, lower the context length or choose a smaller model and try again.

Your first fine-tune in code

Studio is a good way to see the shape of a run. Real projects eventually move to code, because code can be version-controlled, repeated and run on a remote machine. Here is the complete flow with Unsloth Core, built up a step at a time. Run it in a notebook or a file, in the environment where you installed Core.

Step one: load the model.

PYTHON
from unsloth import FastLanguageModel   # import unsloth FIRST

max_seq_length = 2048
model, tokenizer = FastLanguageModel.from_pretrained(
    model_name = "unsloth/gemma-3-270m-it",
    max_seq_length = max_seq_length,
    load_in_4bit = True,        # QLoRA. False means 16-bit LoRA
    load_in_8bit = False,
    load_in_16bit = False,
    full_finetuning = False,    # True would mean full fine-tuning
)

from_pretrained downloads the model on first use and returns two objects: the model and its tokenizer, the component that converts text into the numeric tokens a model reads and back again. The max_seq_length sets the longest sequence you will train on. Only one of load_in_4bit, load_in_8bit, load_in_16bit and full_finetuning may be true at once; setting two is a mistake. For a gated model, one that requires you to accept a licence on the Hub, add token = "hf_..." with your Hugging Face access token.

Step two: attach the adapter.

PYTHON
model = FastLanguageModel.get_peft_model(
    model,
    r = 16,
    target_modules = ["q_proj", "k_proj", "v_proj", "o_proj",
                      "gate_proj", "up_proj", "down_proj"],
    lora_alpha = 16,
    lora_dropout = 0,
    bias = "none",
    use_gradient_checkpointing = "unsloth",
    random_state = 3407,
    max_seq_length = max_seq_length,
)

This is where the frozen model gets its trainable LoRA layers. The name peft stands for parameter-efficient fine-tuning, the Hugging Face library that implements LoRA; if you want to see that library on its own, the PEFT and TRL guide covers it. Taking the arguments in turn: r and lora_alpha are the rank and scale from earlier. lora_dropout = 0 and bias = "none" are the settings the docs describe as optimised, so leave them. use_gradient_checkpointing = "unsloth" selects Unsloth's own memory saver, which the docs credit with about 30% less VRAM. random_state is a seed that makes runs repeatable.

Step three: prepare a dataset. For a first project, a tiny hand-made dataset is clearer than downloading a large one. Suppose the task is to answer questions about a fictional company in a fixed, short style. We build conversations as lists of messages and convert them to the text the model expects:

PYTHON
from datasets import Dataset
from unsloth.chat_templates import get_chat_template

tokenizer = get_chat_template(tokenizer, chat_template = "gemma-3")

pairs = [
    ("What are your opening hours?", "We are open 9am to 5pm, Sunday to Thursday."),
    ("Do you ship to Egypt?", "Yes. Orders to Egypt arrive in 5 to 7 working days."),
    ("How do I reset my password?", "Choose 'Forgot password' on the sign-in page."),
    # add 30 to 100 more examples in the same style
]
conversations = [
    [{"role": "user", "content": q}, {"role": "assistant", "content": a}]
    for q, a in pairs
]
dataset = Dataset.from_dict({"conversations": conversations})

def to_text(examples):
    texts = [tokenizer.apply_chat_template(c, tokenize = False,
                                           add_generation_prompt = False)
             for c in examples["conversations"]]
    return {"text": texts}

dataset = dataset.map(to_text, batched = True)

get_chat_template installs the Gemma 3 template on the tokenizer so that training and inference agree. apply_chat_template then turns each conversation into one string with all the special marker tokens in place, stored in a column called text. You can print dataset[0]["text"] to see what the model will actually be trained on, and you should, because reading one real example catches more mistakes than any amount of theory. Three examples are far too few for a real run; the point is the shape. Aim for at least a few dozen clean, varied examples, and more for anything serious.

Step four: configure and run the trainer.

PYTHON
from trl import SFTTrainer, SFTConfig

trainer = SFTTrainer(
    model = model,
    tokenizer = tokenizer,
    train_dataset = dataset,
    args = SFTConfig(
        max_seq_length = max_seq_length,
        per_device_train_batch_size = 2,
        gradient_accumulation_steps = 4,
        warmup_steps = 10,
        max_steps = 60,
        learning_rate = 2e-4,
        logging_steps = 1,
        optim = "adamw_8bit",
        seed = 3407,
        output_dir = "outputs",
        dataset_num_proc = 1,
    ),
)
trainer.train()

Several of these settings need explaining. per_device_train_batch_size is how many examples the GPU processes at once. gradient_accumulation_steps makes the trainer collect gradients from that many batches before updating the weights, which simulates a larger batch without the memory cost. The effective batch size is the product: here 2 times 4, so 8 examples per update. If you run out of memory, lower the first number and raise the second; the product, and therefore the learning behaviour, stays the same. Unsloth fixed a long-standing gradient accumulation bug, so those two combinations now give matching loss curves.

warmup_steps ramps the learning rate up gently at the start. max_steps = 60 stops after sixty updates; the alternative is num_train_epochs = 1, which runs a full pass over the data and is the better choice once your dataset is real. optim = "adamw_8bit" is an optimizer that stores its state in 8 bits to save memory. logging_steps = 1 prints the loss every step.

When training runs, you see a table of step numbers and training loss. If your tutorial notes disagree with the tokenizer = argument, it is because the TRL library keeps evolving its argument names; the form above follows the current Unsloth documentation, and if you hit an argument error, check the changelog for the pairing of Unsloth and TRL versions you installed.

Step five: try the result.

PYTHON
from transformers import TextStreamer

FastLanguageModel.for_inference(model)   # switch to fast generation mode

messages = [{"role": "user", "content": "Do you ship to Egypt?"}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt = True,
    tokenize = True, return_tensors = "pt", return_dict = True,
).to("cuda")

_ = model.generate(**inputs, streamer = TextStreamer(tokenizer),
                   max_new_tokens = 128)

FastLanguageModel.for_inference flips the model into a mode the docs describe as twice as fast for generation. add_generation_prompt = True appends the marker that tells the model "now it is your turn", which is the one difference from training formatting. TextStreamer prints tokens as they arrive. Compare the answer to your baseline.txt.

Try it
  1. Type the five steps into a notebook, replacing the example pairs with at least 30 of your own.
  2. Print dataset[0]["text"] and read it carefully before training.
  3. Run training, then ask the model a question that is in your data and one that is not.
the trained model should answer the first in your style. The second tells you how well it generalises, which is often the more interesting result.

Choosing the numbers: a beginner's guide to hyperparameters

You now know the settings exist. This section explains how to choose among them without a research budget. The advice comes from the official hyperparameters guide.

Learning rate. Use 2e-4 for LoRA and QLoRA. If the loss swings wildly or jumps to a huge number, lower it; if it barely moves after many steps, raise it a little. The usual reinforcement-learning value is much smaller, about 5e-6, so do not copy a number from an RL tutorial into your SFT run.

Epochs. One to three passes over the data is the recommended range. More epochs on a small dataset mostly teach the model to memorise it. If your dataset is small, prefer collecting more varied examples over adding epochs.

Rank and alpha. Start with r = 16, lora_alpha = 16. If the model seems unable to learn your task, try r = 32. Bigger ranks use more memory and are more likely to overfit, so increase gradually.

Datasets and chat templates, the part that decides quality

If one thing separates good fine-tunes from bad ones, it is the data. Models learn precisely what you show them, including your typos and inconsistencies.

What a good example looks like. For a chat model, an example is a conversation with a clear instruction and the exact answer you want. Make the answers match what you would be proud to see in production: same format, same length, same tone. If half your answers are one word and half are paragraphs, the model will produce a random mixture. Remove duplicates and contradictions. A hundred carefully written examples routinely beat ten thousand scraped ones for style and format tasks.

Common storage formats. Studio's dataset step names three: alpaca has separate fields for an instruction, an optional input and an output; chatml stores a list of messages, each with a role and content; sharegpt stores a list of turns with from and value keys. All three express the same idea. Studio's auto setting guesses the format, and in code you can convert ShareGPT-style data to the role-and-content layout using the standardize_sharegpt helper from unsloth.chat_templates.

Chat templates, again. Every model family wraps turns in its own marker tokens. Gemma, Llama and Mistral all differ. Training with one template and running with another produces the classic failure: the model "forgets" what it learned, rambles, or never stops. Rules that avoid this:

  1. Use get_chat_template with the template that matches your model family, or the tokenizer's own template for models that already ship with one.
  2. Apply the same template at inference with apply_chat_template.
  3. If you export to another tool, check that it uses the same template. Unsloth's notebook exports to Ollama create the Modelfile automatically, including the template.
Read twenty examples with your own eyes Before every training run, print and read at least twenty random rows of the final formatted text. Look for broken formatting, truncated answers, leftover HTML and duplicates. Ten minutes here saves hours of GPU time.
Try it
  1. Write 30 examples for your chosen task in a spreadsheet, one column for the question and one for the answer.
  2. Export it as CSV and upload it in Studio's dataset step, or load it in code with the Hugging Face datasets library.
  3. Read ten rows after formatting.
formatted text where each row clearly shows who spoke and what the ideal answer is. If anything looks off, fix the data, not the training settings.

Saving, testing and exporting your model

A trained model that lives only in a notebook disappears when the session ends. Saving is not one operation but a choice among three, and each suits a different use.

LoRA adapter only. The smallest option, typically tens of megabytes. It holds only the trained weights and requires the original base model to run.

PYTHON
model.save_pretrained("lora_model")
tokenizer.save_pretrained("lora_model")

Use this while experimenting, and when you want to keep several fine-tunes of one base model.

Merged 16-bit. The adapter is folded into the base weights, producing a normal standalone model. This is the format that servers such as vLLM and libraries such as plain Transformers expect.

PYTHON
model.save_pretrained_merged("finetuned_model", tokenizer,
                             save_method = "merged_16bit")

The docs discourage merged_4bit, because merging into a 4-bit model loses accuracy; it needs an explicit merged_4bit_forced to be used at all. A merged 16-bit model is large, about two bytes per parameter, so an 8B model needs around 16 GB of disk. You can serve it with the vLLM guide tooling.

GGUF. A single file for llama.cpp-based runners, including Studio, Ollama and many desktop apps.

PYTHON
model.save_pretrained_gguf("gguf_model", tokenizer,
                           quantization_method = "q4_k_m")

Common quantization_method values include q4_k_m (the balanced default), q5_k_m, q8_0 and f16. To publish rather than save locally, use the matching push_to_hub_merged and push_to_hub_gguf methods, which take a repository name like "your-name/your-model" and a Hugging Face token. Never paste a token into a notebook you will share.

Saving is memory-hungry because merging happens on the GPU. If it crashes, pass maximum_memory_usage = 0.5 to lower the peak from its default of 0.75.

Using it with Ollama. The Unsloth notebooks export a GGUF and automatically create the Modelfile that Ollama needs, including the chat template. After that, ollama create followed by ollama run is all it takes. The Ollama guide explains that tool, and the llama.cpp guide explains the engine underneath. If your fine-tune behaves differently in Ollama than in your notebook, the template is the first suspect.

In Studio the same choices appear in the Export step: Merged 16-bit, LoRA only, or GGUF.

The command line has an equivalent. After a Studio or CLI training run, list your runs and export one:

BASH
unsloth list-checkpoints --outputs-dir ./outputs
unsloth export ./outputs/checkpoint-60 ./exported --format gguf --quantization q4_k_m

The --format option accepts merged-16bit (the default), merged-4bit, gguf and lora, and the --quantization option accepts q4_k_m (the default), q5_k_m, q8_0 and f16. Replace the checkpoint path with a real folder from the listing.

Testing the export. Whatever you produce, load it once in the tool where it will really run and ask your baseline question. A model that answers perfectly in the notebook and badly in Ollama has a template or end-of-sequence problem, not a training problem.

Try it
  1. Save your fine-tune three ways: adapter, merged 16-bit and GGUF (Q4_K_M).
  2. Compare the sizes of the three folders with du -sh or your file manager.
  3. Load the GGUF in Studio or Ollama and ask your baseline question.
an adapter of tens of megabytes, a merged model of many gigabytes and a GGUF in between, and the same answer from the GGUF as from the notebook.

Everyday commands, grouped by goal

By now you have used most of the interface. This section collects the commands in one place, organised by what you are trying to do. All of them come from the installed unsloth command or the Python API covered above.

I want to start and stop Studio.

BASH
unsloth studio -p 8888        # start on port 8888 (the default)
unsloth studio stop           # stop it
unsloth studio reset-password # set a new admin password

If you see Error: Unsloth Studio is already running on port 8888, an earlier instance is still up. Stop it, or choose another port with -p.

I want to chat with a model quickly.

BASH
unsloth run --model unsloth/gemma-3-270m-it-GGUF:Q4_K_M

unsloth run is a short alias for unsloth studio run. Put flags after the subcommand: unsloth studio run --secure works, but unsloth studio --secure run exits with an error. If you do not pass sampling options such as --temp, Unsloth applies the model's recommended settings, which is a convenient default.

I want to train without the graphical interface. unsloth train reads a YAML or JSON configuration file:

BASH
unsloth train --config my_run.yaml --dry-run

--dry-run checks the file without training. A minimal file looks like this:

my_run.yaml
model: unsloth/gemma-3-270m-it
data:
  local_dataset: ["./my_data.jsonl"]
  format_type: auto
training:
  training_type: lora
  max_seq_length: 2048
  load_in_4bit: true
  num_epochs: 1
  learning_rate: 2e-4
  output_dir: ./outputs
lora:
  lora_r: 16
  lora_alpha: 16

The schema rejects keys it does not know, so a typo becomes an immediate error instead of a silent misconfiguration. Note that the command-line schema uses names such as num_epochs and batch_size, which differ from the num_train_epochs and per_device_train_batch_size you see in Python and in Studio's saved files; the two are similar, not identical, so do not assume a file from one works in the other. Tokens are supplied through --hf-token or the HF_TOKEN environment variable rather than written in the file. The command-line default for rank is 64 with alpha 16, which differs from the notebook convention of 16 and 16, so set both explicitly as above.

I want to see what I have trained. Use unsloth list-checkpoints --outputs-dir ./outputs.

I want to export. Use unsloth export <checkpoint> <output_dir> as shown in the previous section.

I want to check versions.

BASH
pip show unsloth unsloth_zoo

I want to update. For Core, upgrade both packages together:

BASH
pip install --upgrade unsloth unsloth_zoo

For Studio, run the install one-liner again. For Desktop, use Check for updates in Settings.

BASH
unsloth start claude

This is an advanced convenience, so treat it as a preview of where the tool goes after you have the basics.

Try it
  1. Write the YAML file above for your own dataset.
  2. Run unsloth train --config my_run.yaml --dry-run and fix any error it reports.
  3. Then run it without --dry-run.
a validated config and a training run you could repeat next week from the same file.

Configuration, state and where things live

Knowing where Unsloth keeps things turns many mysteries into one-line fixes.

Models are cached. Downloaded models live in the Hugging Face cache, ~/.cache/huggingface/hub. You can move it by setting HF_HOME or HF_HUB_CACHE. Models are large, so on a small laptop disk this folder is usually what fills up. Deleting it frees space at the cost of re-downloading.

Studio keeps its state under ~/.unsloth/studio. That includes the authentication data, a small database, a cache and the built llama.cpp engine. The installer's command and the engine itself sit under ~/.local/bin/unsloth and ~/.unsloth/llama.cpp. To relocate the whole install, set the environment variable UNSLOTH_STUDIO_HOME to an absolute path before installing. Uninstalling is a script, and it never removes your model cache. Running rm -rf ~/.unsloth yourself, however, permanently deletes all chats, checkpoints and exports, so back up anything you care about first.

Installer options are environment variables. UNSLOTH_NO_TORCH=1 installs a GGUF-only Studio with no training libraries, UNSLOTH_PYTHON=3.12 pins the Python version, and UNSLOTH_CPU_THREADS=8 caps the CPU threads used at launch. On Unix, put them after the pipe:

BASH
curl -fsSL https://unsloth.ai/install.sh | UNSLOTH_NO_TORCH=1 sh

When it goes wrong: common errors and how to read them

Errors in this ecosystem look intimidating because they come from several libraries at once. Read them from the bottom up: the last lines name the actual failure, and the lines above only show how it was reached. The table below covers the messages beginners hit most, with causes taken from the official troubleshooting material.

What you see Likely cause What to do
Out-of-memory error during training Batch too large, context too long or model too big Batch size 1 to 3, raise gradient accumulation, shorten context, or pick a smaller model
Out-of-memory or crash when saving Peak memory while merging weights maximum_memory_usage = 0.5 in the save call
Good in your notebook, gibberish or endless output in Ollama or vLLM Chat template, end-of-sequence or beginning-of-sequence token differs Use the same template at inference as in training; check the exported Modelfile
All labels in your dataset are -100. Training losses will be all 0. train_on_responses_only markers do not match the template Use the exact marker strings for your model family
Error: Unsloth Studio is already running on port 8888 Port in use unsloth studio stop, or use -p with another port
Download stuck at 90 to 95 percent The fast parallel downloader stalls Set UNSLOTH_STABLE_DOWNLOADS=1 before importing Unsloth
401 Unauthorized from the API Missing, wrong or revoked key Create a new key in Settings, then API; old keys cannot be viewed again
Lost connection to the model server The llama.cpp server crashed or the tab was closed Reload the model from a new chat
Python version error during install Python outside 3.11 to 3.13 Install a supported Python, or set UNSLOTH_PYTHON
nvcc not found, cmake not found, git not found Missing build tools Install them; sudo apt install cmake git on Ubuntu
vLLM or GRPO fails on native Windows vLLM does not support native Windows Use WSL or Linux
Training looks slower than promised at first torch.compile warms up for a few minutes Judge speed after warm-up

A few of these deserve a longer explanation.

Out of memory is by far the most common. Think of memory as a budget with four spenders: the model weights, the optimizer, the activations and the batch. In rough order of effect, you reduce the budget by choosing a smaller model, switching from LoRA to QLoRA, shortening the context length and lowering the per-device batch size. Do those in that order, and do not conclude that your GPU is too small until you have tried all four.

**## Putting it all together

Here is one small end-to-end project that uses everything above. The goal is a tiny support assistant that answers questions about a made-up shop in a fixed style, then runs locally in Ollama or Studio.

  1. Choose the hardware route. Check nvidia-smi or open a free Colab T4 notebook. Note the memory.
  2. Install. Use the Core commands in a fresh uv environment, or open the official Colab notebook.
  3. Pick the model. unsloth/gemma-3-270m-it with QLoRA for the first pass; step up to a 3B or 8B model only if the small one cannot learn the task.
  4. Capture a baseline. Ask the untouched model five questions and save the answers to baseline.txt.
  5. Build the dataset. Write 50 question-and-answer pairs in your target style, with varied phrasing and no duplicates. Apply the chat template and read twenty formatted rows.
  6. Train. Use the starting values from the hyperparameter table, with max_steps = 60 first, then a full epoch if the loop works.
  7. Test. Ask the same five baseline questions plus five new ones. Note where the style improved and where the model invented facts.
  8. Save. Keep the adapter in version control (it is small) and export a Q4_K_M GGUF.
  9. Run it. Load the GGUF in Studio or Ollama and ask the same ten questions again. If the answers differ from the notebook, check the template first.
  10. Write it down. Record the Unsloth and unsloth_zoo versions, the model name, the dataset version and your settings in a README. Next month you will not remember them.
Try it
  1. Complete the ten steps with your own task.
  2. Share your README, adapter folder and five before-and-after answers with a friend or in the community.
a project another person could reproduce from your notes. That is the standard to aim for at every level.

What you can now do, and what comes next

You can now explain what Unsloth does and what it does not, choose among Desktop, Studio and Core, check your GPU and your install, chat with a model locally, decode a GGUF name, fine-tune a small model in Studio and in code, read the training loss with sensible suspicion, save the result in three formats, and diagnose the dozen most common problems.

You have not yet covered several things that matter at work. Holding out evaluation data and stopping early, resuming from checkpoints, training only on the assistant's replies, vision and audio models, serving a model to many users, and reinforcement learning with GRPO are all in the mid-level track. Multi-GPU training, security of a shared Studio server, upgrade policy for weekly releases and when Unsloth is the wrong tool belong to the senior track.

Natural next stops in this catalogue are the PEFT and TRL guide for the libraries underneath, the Hugging Face guide for the Hub, models and datasets you used throughout, the vLLM guide for serving a merged model, and the Ollama guide for running a GGUF locally. Keep the habit that will serve you longest: pin versions, keep a baseline, and read your data.

Sources