Fine-Tuning & Alignment
Fine-tuning takes a pretrained language model and keeps training it on your own data so that it adopts a new behaviour, format, style or skill; alignment is the family of post-training methods (RLHF, DPO, GRPO and friends) that shape a model towards what humans actually prefer. This page covers the whole lifecycle, from deciding whether to fine-tune at all, through data, LoRA/QLoRA, memory math and preference optimization, to evaluating, merging, distilling and serving the result.
- Try prompting first, add RAG when the model lacks knowledge, and fine-tune when it lacks behaviour (format, tone, skill, latency/cost targets).
- Supervised fine-tuning (SFT) is next-token prediction on curated prompt–response pairs, with the loss usually computed on the response tokens only.
- Full fine-tuning with AdamW costs about 16 bytes per parameter before activations (≈112 GB for 7B); LoRA trains a low-rank update
ΔW = (α/r)·BAon a frozen model, and QLoRA also stores the frozen model in 4-bit NF4, so a 7B model trains on one consumer GPU. - Alignment turns a helpful-looking model into a preferred one: classic RLHF is SFT → reward model → PPO with a KL penalty; DPO skips the reward model and RL loop; GRPO is the critic-free RL used for reasoning with verifiable rewards.
- Data quality beats quantity: a few thousand clean, deduplicated, correctly templated examples usually beat a large, noisy dump.
- The main risks are catastrophic forgetting, overfitting, safety regression, data leakage and licensing, so always compare against the base model on a regression suite.
- To deploy, either merge the adapter into the base weights (no latency overhead) or serve many LoRAs over one shared base, then quantize (GPTQ/AWQ/GGUF) for cheap or on-device inference.
The big picture: where fine-tuning fits
A modern chat model is built in stages. Pretraining teaches a network general language and world knowledge by predicting the next token over trillions of tokens of text. The result, a base model, is a powerful autocomplete engine: ask it a question and it may continue with more questions, because that is what the internet looks like. Post-training then turns the base model into an assistant: supervised fine-tuning on demonstrations teaches it the question-answer format, and preference optimization (RLHF, DPO and relatives) teaches it which of several plausible answers humans actually prefer. When you "fine-tune" as a practitioner, you are adding one more post-training step on top of an existing base or instruct model, aimed at your own use case.
Think of a medical graduate. Years of general schooling and medical school are pretraining: broad knowledge, expensive, done once. A residency in cardiology is fine-tuning: a much shorter, focused programme that turns the generalist into a specialist who follows the conventions of one department. Feedback from senior doctors on which diagnosis explanation was clearer is alignment: it does not add much new knowledge, it refines judgement and manner.
Mapping back: pretraining is the costly general education, SFT is the residency that teaches a specific way of working, and preference tuning is the senior-doctor feedback that sharpens which answers are preferred. Just like a resident, a fine-tuned model can also forget things from outside the specialty if it trains too narrowly.
THE LLM LIFECYCLE ┌──────────────┐ ┌──────────────┐ ┌──────────────────┐ ┌──────────────┐ │ Pretraining │──▶│ SFT / │──▶│ Preference │──▶│ Your domain │ │ next-token │ │ instruction │ │ optimization │ │ fine-tune │ │ on ~10T tok │ │ tuning │ │ RLHF / DPO / RL │ │ (LoRA/QLoRA) │ └──────────────┘ └──────────────┘ └──────────────────┘ └──────────────┘ base model instruct model chat/aligned model specialised model (knowledge) (follows format) (preferred answers) (your behaviour) months, $M+ hours-days days minutes-hours
What fine-tuning actually changes
Fine-tuning runs ordinary gradient descent on the model's weights (all of them, or a small added set) using a loss computed on your examples. After training, the change lives in the weights: no extra prompt is needed at inference time. This is the key difference from prompting and in-context learning, where the weights are frozen and the model adapts only through the text you give it, and from retrieval-augmented generation (RAG), where facts are fetched at query time and pasted into the context.
Behaviour and format
Always answer in a given JSON schema, adopt a brand voice, follow a house style, stop being verbose. Fine-tuning is excellent at this.
Skills
Classify tickets, extract fields, write SQL for your schema, call your tools correctly, reason in a domain-specific way. Fine-tuning is good at this with enough examples.
Knowledge
Memorizing facts, especially ones that change, is where fine-tuning is weakest: it is slow to update, hard to cite, and prone to hallucinating near-miss facts. Use RAG for this.
Cost and latency
A small fine-tuned model can replace a large general model with a long prompt, cutting per-call cost and latency by an order of magnitude.
Prompt vs RAG vs fine-tune: making the decision
Fine-tuning is the most expensive lever in terms of data, engineering and maintenance, so it should be pulled deliberately. The standard progression is: start with a strong prompt (zero-shot, then few-shot), add tools or structured output when you need determinism or actions, add RAG when the model is missing knowledge, and fine-tune when the desired behaviour is still unreliable or too expensive to achieve through prompts. The rule of thumb is short: new knowledge → RAG; new behaviour → fine-tune; both → both.
Imagine onboarding a capable new employee. Prompting is giving clear instructions in each email. RAG is handing them the company handbook to look things up as needed, which stays current as the handbook is edited. Fine-tuning is sending them on a training course so that the company's way of doing things becomes second nature, which is great for habits but useless for remembering a policy that changes next week.
Mapping back: instructions correspond to prompts (cheap, instant, but must be repeated every time), the handbook corresponds to a retrieval index (fresh and citable), and the training course corresponds to weight updates (durable habits, but frozen at training time).
| Need | Prompting | RAG | Fine-tuning |
|---|---|---|---|
| Fresh or frequently changing facts | Only if pasted in manually | Best choice: update the index | Poor: facts are frozen, retraining needed |
| Citations and traceability | No | Yes, cite retrieved chunks | No |
| Consistent format, tone or style | Works but drifts on long tails | Does not help | Best choice |
| New skill (domain classification, extraction, tool calling) | Few-shot may be enough | Limited help | Strong |
| Private data at scale (large corpus) | Context window limits | Designed for it | Continued pretraining possible but costly |
| Lower latency and cost per call | Long prompts cost tokens | Adds retrieval latency and tokens | Small tuned model can replace a big one |
| Time to first result | Minutes | Days | Days to weeks (data is the bottleneck) |
| Data needed | A few examples | Documents | Hundreds to tens of thousands of labelled examples |
| Maintenance | Edit prompt | Keep index fresh | Retrain on base-model upgrades, re-evaluate |
| Hallucination control | Moderate | Good (grounding) | Can increase if trained on facts it did not know |
A decision procedure
- Define success and build an eval set first. Without 50–200 labelled test cases and a metric, you cannot tell whether any lever helped.
- Baseline with prompting. Zero-shot, then few-shot, then a structured system prompt. Record scores, cost and latency. See Prompt Engineering.
- Diagnose the failures. Are they missing facts (retrieval problem) or wrong behaviour despite having the facts (format, reasoning, tone, refusal)?
- Missing facts → RAG. Build retrieval and re-evaluate. See Retrieval-Augmented Generation.
- Behaviour still wrong, or too costly → fine-tune. Start with LoRA/QLoRA SFT on a few thousand high-quality examples; consider preference tuning if the issue is "which answer is better" rather than "what format".
- Combine. Many production systems fine-tune for style, format and tool use, and use RAG for facts. A retrieval-aware fine-tune (training on examples that include retrieved context) teaches the model to use context faithfully.
Good reasons to fine-tune
- Strict output format that prompting breaks 2–5% of the time
- Distilling a big model's behaviour into a small, cheap one
- Domain language the base model handles poorly (clinical notes, legal clauses, internal codes)
- Tool/function calling for a private API
- Latency or on-device limits that force a small model
- Removing a very long system prompt from every call
Bad reasons to fine-tune
- "Teach it our product catalogue" that changes weekly
- No evaluation set yet
- Fewer than a hundred messy examples
- Prompting has not been seriously tried
- Hoping to fix factual hallucinations by training on facts
- A closed model that only allows limited API fine-tuning, when prompting already works
Types of fine-tuning
"Fine-tuning" is an umbrella term. The variants differ in what data is used, what loss is optimized and which parameters are updated. The first two axes decide what the model learns; the third (full vs parameter-efficient) decides cost.
Consider a chef. Reading hundreds of regional cookbooks cover to cover is continued pretraining. Practising exact recipes that customers order is instruction tuning. Specialising in only pastry is task-specific tuning. Having diners taste two versions of a dish and pick the better one is preference tuning. Whether the chef relearns every technique from scratch or just adds a few new habits is the full-vs-parameter-efficient choice.
Mapping back: raw domain text maps to continued pretraining, prompt–response pairs map to SFT, a narrow labelled task maps to task-specific tuning, and chosen-vs-rejected pairs map to preference optimization; the chef's retraining scope maps to updating all weights versus small adapters.
| Type | Data | Loss | Typical use |
|---|---|---|---|
| Continued pretraining (domain-adaptive pretraining) | Unlabelled domain text: papers, code, contracts, other languages | Next-token loss on all tokens | Teach vocabulary, style and knowledge of a domain or language; followed by SFT to restore chat ability |
| Supervised fine-tuning / instruction tuning | Instruction (+ optional input) → response pairs, or multi-turn chats | Next-token loss, usually on response tokens only | Turn a base model into an assistant, or teach a format/skill |
| Domain adaptation | Domain-specific instructions (medical Q&A, legal drafting) | SFT loss, sometimes after continued pretraining | Specialist assistant for a vertical |
| Task-specific fine-tuning | Labelled examples for one task (classification, NER, extraction, SQL) | Cross-entropy on labels, via a classification head or generative targets | High accuracy on one narrow task, often with a small model |
| Preference tuning | Prompt with chosen and rejected responses, or scalar rewards | RLHF (reward + PPO), DPO, ORPO, KTO, GRPO | Helpfulness, harmlessness, style preferences, reasoning |
| Distillation fine-tuning | Teacher model outputs (text or full probability distributions) | SFT on teacher text, or KL to teacher logits | Compress a large model's behaviour into a small one |
| Long-context fine-tuning | Long documents with adjusted positional scaling | Next-token loss | Extend context window (with RoPE scaling tricks) |
Full fine-tuning vs parameter-efficient fine-tuning
Full fine-tuning
- Every weight gets a gradient and optimizer state
- Highest capacity; best for large distribution shifts (new language, continued pretraining)
- Memory ≈ 16 bytes per parameter plus activations
- Produces a full new copy of the model per task
- Higher risk of catastrophic forgetting
- Needs small learning rates (around 1e-5)
Parameter-efficient fine-tuning (PEFT)
- Base weights frozen; train <1–2% extra parameters
- Matches full fine-tuning on most SFT tasks when rank and target modules are adequate
- Fits on one GPU; QLoRA fits 7B–8B in under 16 GB
- Adapter files are megabytes; many tasks share one base
- Forgets less, because the base is untouched
- Tolerates larger learning rates (around 2e-4)
Instruction tuning in detail
Instruction tuning trains on examples that look like the requests a user would make. A single record typically has three fields, instruction, an optional input (extra context) and output (the target answer). Diversity of instructions matters as much as volume: models tuned on many different task types generalize to unseen instructions, which is why instruction tuning produced the jump from "autocomplete" to "assistant". Instruction tuning is different from prompt engineering: it changes weights during training, while prompt engineering changes only the input at inference.
{"instruction": "Classify the sentiment of the review.",
"input": "The battery died after two days.",
"output": "negative"}
{"instruction": "Explain what an ACE inhibitor does in one sentence.",
"input": "",
"output": "An ACE inhibitor lowers blood pressure by blocking the enzyme that makes angiotensin II, a hormone that narrows blood vessels."}
Data preparation: the part that decides the outcome
In fine-tuning, the dataset is the specification. The model will imitate whatever it sees, including typos, inconsistent formats, wrong answers and a verbose tone, so most of the work in a successful project is data work. This section covers formats, chat templates, cleaning, deduplication, synthetic data, and how to split the data for honest evaluation.
An apprentice copies the master's work exactly. Show them fifty beautifully finished pieces and they learn craftsmanship; show them five thousand rushed pieces with inconsistent measurements and they learn to be sloppy at scale. They also cannot learn from a piece whose instructions are written in a notation they have never seen.
Mapping back: the finished pieces are your training examples (quality matters more than count), the inconsistent measurements are noisy labels and mixed formats, and the unfamiliar notation is a prompt that does not match the model's chat template.
Common dataset formats
| Format | Shape | Used for |
|---|---|---|
| Instruction (Alpaca-style) | instruction, input, output | Single-turn SFT |
| Conversational (messages) | messages: [{role, content}, ...] with system/user/assistant (and tool) roles | Multi-turn SFT, tool calling; the modern default |
| Prompt–completion | prompt, completion | Completion-only SFT, classification as text |
| Plain text | text | Continued pretraining |
| Preference pairs | prompt, chosen, rejected | Reward models, DPO, ORPO, IPO, SimPO |
| Unpaired feedback | prompt, completion, label (good/bad) | KTO |
| Prompts only (+ verifier) | prompt and a reference answer or test | Online RL such as PPO or GRPO |
{"messages": [
{"role": "system", "content": "You are a concise support assistant for a telecom company."},
{"role": "user", "content": "My bill is higher this month. Why?"},
{"role": "assistant", "content": "Bills usually rise for three reasons: a promotion ended, you used extra data, or a one-off fee was added. Check the 'Charges' tab for the exact line item."}
]}
Chat templates: format must match the model
Every instruct model was trained with a specific chat template: special tokens that mark where each role's turn begins and ends. For example, one popular family wraps turns as <|im_start|>user ... <|im_end|>, another uses [INST] ... [/INST], and another uses header tokens such as <|start_header_id|>assistant<|end_header_id|>. If you fine-tune an instruct model with a different layout, you waste capacity re-teaching the format and often break its existing chat skills. In Hugging Face Transformers the tokenizer carries its template, and tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) renders it.
<|im_start|>system
You are a helpful medical assistant. Answer accurately and safely.<|im_end|>
<|im_start|>user
What is the first-line treatment for mild hypertension?<|im_end|>
<|im_start|>assistant
Lifestyle changes come first: reduce salt, ...<|im_end|>
add_generation_prompt=Trueappends the opening of the assistant turn, so the model knows it should answer next. Use it when rendering prompts for training-with-masking and for inference.- End-of-turn token. The target must end with the template's end-of-turn or EOS token (for example
<|im_end|>). If you omit it, the model never learns when to stop and rambles untilmax_new_tokens. - No double special tokens. A rendered template often already contains BOS; tokenizing it again with
add_special_tokens=Truecan prepend a second BOS. - Base models have no template. You may choose one (a common choice is to adopt the template of the matching instruct model), but then add any new special tokens to the tokenizer, resize embeddings and train the embedding rows too.
Loss masking: train on answers, not prompts
In SFT we usually compute the loss only on the assistant tokens (completion-only or "train on responses"). Prompt tokens still pass through the model as context, but their labels are set to -100 (the ignore index in PyTorch cross-entropy), or multiplied by a zero mask. Otherwise the model spends capacity learning to reproduce user prompts and system text, which is not the behaviour you want, and long prompts dominate the gradient. For multi-turn data, mask every user and system turn and supervise every assistant turn.
tokens : [sys][sys][user][user][user][asst-hdr] A1 A2 A3 <|im_end|> [pad][pad]
labels : -100 -100 -100 -100 -100 -100 A1 A2 A3 <|im_end|> -100 -100
context only (no loss) loss here (incl. stop token) ignored
Quality over quantity
A consistent finding across instruction-tuning work is that a small set of carefully curated examples (around a thousand in some well-known experiments) can produce a strong assistant from a good base model, while larger noisy sets add little. The base model already has the knowledge; SFT mostly teaches format and style, which is sometimes called the superficial alignment hypothesis. Rough sizes that work in practice:
| Goal | Typical size |
|---|---|
| Output format or style | 200–2,000 examples |
| Simple classification / extraction | 500–5,000 examples |
| Domain assistant (SFT) | 5,000–50,000 examples |
| General instruction tuning from a base model | 10,000–1,000,000 examples |
| Preference tuning (DPO) | 2,000–100,000 pairs |
| Continued pretraining | Hundreds of millions to billions of tokens |
Cleaning checklist
- Normalize the schema. One format, consistent roles, placeholder tokens like a literal
<noinput>converted to an empty string, whitespace stripped. - Filter bad examples. Empty or truncated answers, wrong language, broken markup, refusals where you want answers (or the reverse), answers that contradict the instruction.
- Deduplicate. Exact duplicates via hashing; near-duplicates via MinHash/LSH or embedding similarity. Duplicates overweight some behaviours and leak between train and test.
- Remove sensitive data. Scrub personal information, secrets and credentials; models can memorize and regurgitate rare strings.
- Check lengths. Compute token-length percentiles and set
max_seq_lengthnear the 95th percentile; drop or split outliers instead of silently truncating answers. - Balance. Watch class balance and task mix; mix in some general-purpose data to reduce forgetting.
- Decontaminate. Remove any example that overlaps with your evaluation set or with public benchmarks you plan to report.
- Human spot-check. Read at least 100 random examples. Most dataset bugs are obvious to a human in ten minutes.
Synthetic data
Much modern fine-tuning data is generated by a stronger model: the teacher writes instructions (self-instruct style), answers them, rewrites human answers into a better format, or produces chosen/rejected pairs. Synthetic data is cheap and controllable but needs guardrails: seed it with diverse real prompts, filter with rules and an LLM judge, verify facts or run code/tests where possible, deduplicate aggressively (generators repeat themselves), and check the teacher's licence and terms of use, which often forbid training competing models. Beware model collapse: repeated training on model-generated text narrows diversity, so keep a share of real human data.
Train, validation and test splits
- Shuffle once with a fixed seed, then slice train/validation/test (for example 80/10/10, or at small scale a few hundred / a few dozen / a few dozen as in the worked example later).
- Split by group, not by row when examples are related (same customer, same document, same conversation), otherwise near-duplicates leak across splits.
- Validation drives early stopping and hyperparameter choices; test is touched once at the end.
- Keep a separate regression set of general tasks (math, coding, instruction following, safety prompts) to measure forgetting.
Packing and padding
Batches must be rectangular, so shorter sequences are padded. Padding wastes compute when lengths vary a lot. Packing concatenates several examples into one sequence of the maximum length, separated by EOS; with modern kernels, position IDs are reset per example and attention is blocked across boundaries so examples do not attend to each other. Packing can speed training several-fold on short-example datasets. Without packing, group similar lengths into the same batch to reduce padding.
input_ids != pad_token_id. This also masks the real end-of-sequence token, so the model never learns to stop. Build masks from sequence lengths instead, or use a dedicated pad token.How supervised fine-tuning works under the hood
SFT uses exactly the same objective as pretraining: the model predicts each next token and pays a cross-entropy penalty for putting low probability on the true one. The only differences are the data (curated examples instead of raw web text) and the mask (loss on answer tokens only). Understanding the shift-by-one and the masking removes most confusion about "how the model learns to answer".
Picture a language tutor who reads a model answer aloud one word at a time and, before each word, asks the student to guess it. The student hears the question in full but is only graded on guesses within the answer. Every wrong guess produces a correction proportional to how surprised the student was.
Mapping back: the question is the masked prompt, each guess is a next-token prediction, the grade is the cross-entropy loss on answer tokens, and the correction is the gradient update. Because the true previous word is always revealed before the next guess, this is teacher forcing.
The shift-by-one
A causal language model produces logits at every position; the logits at position t are a prediction for token t+1. So we compare logits[:, :-1] with input_ids[:, 1:] and shift the mask the same way. Hugging Face models do this internally when you pass labels; in a hand-written loop you do it yourself:
def batch_loss(model, input_ids, attn, comp_mask):
"""Mean next-token cross-entropy over answer tokens only."""
logits = model(input_ids=input_ids, attention_mask=attn).logits
shift_logits = logits[:, :-1, :].contiguous() # predictions for t+1
shift_labels = input_ids[:, 1:].contiguous() # the true next tokens
shift_mask = comp_mask[:, 1:].contiguous().float() # 1 on answer tokens
per_tok = F.cross_entropy(shift_logits.view(-1, shift_logits.size(-1)),
shift_labels.view(-1), reduction="none")
return (per_tok * shift_mask.view(-1)).sum() / shift_mask.sum().clamp(min=1.0)
Token-level vs sequence-level averaging
Averaging over all answer tokens in a batch weights long answers more heavily than short ones; averaging per sequence first weights every example equally. With gradient accumulation, the naive approach of averaging each micro-batch and then averaging micro-batches is subtly wrong when micro-batches contain different token counts; libraries now normalize by the total number of supervised tokens across the accumulated batch. For most SFT it is a minor effect, but it matters when lengths vary widely.
The training loop in one picture
for each epoch:
for each micro-batch:
forward ──▶ masked loss / grad_accum ──▶ backward (grads accumulate)
every grad_accum steps:
clip_grad_norm(1.0) ──▶ optimizer.step() ──▶ scheduler.step() ──▶ zero_grad()
evaluate on validation ──▶ keep best checkpoint ──▶ early-stop if val loss rises
Mixed precision and numerics
- bf16 has the same exponent range as fp32, so it rarely overflows; it is the default on GPUs that support it (Ampere and newer).
- fp16 has a narrow range and needs loss scaling; full fine-tuning in pure fp16 can overflow to
inf/NaN. Older GPUs such as the T4 lack fast bf16, so a common pattern there is fp32 master weights for full fine-tuning, and fp16 compute for LoRA/QLoRA (where only small adapters are trained). - Master weights in fp32 keep tiny updates from being rounded away; mixed-precision training keeps an fp32 copy for the optimizer while computing in 16-bit.
Memory math for fine-tuning
Most practical fine-tuning questions reduce to "will it fit on my GPU?" GPU memory during training holds four things: the weights, the gradients, the optimizer states, and the activations saved for the backward pass (plus some framework overhead and temporary buffers). The first three scale with the number of trainable parameters; activations scale with batch size, sequence length, hidden size and number of layers.
Think of renovating a large library. You need space for the books themselves (weights), a note on every book saying how to change it (gradients), and a running history for every book of past changes so edits are smooth (optimizer states). You also need scratch tables for the pages currently open (activations), and the number of tables grows with how many readers and how long each document is. Deciding to edit only a small index card per shelf instead of every book is LoRA; storing the books in a compressed format is QLoRA.
Mapping back: per-book notes and histories exist only for books you are allowed to edit, which is exactly why freezing the base model removes most of the gradient and optimizer memory, while compression shrinks the weight storage itself.
Bytes per parameter
| Component | Full FT, mixed precision + AdamW | Full FT, pure fp32 + AdamW | Notes |
|---|---|---|---|
| Weights | 2 (bf16) | 4 | Needed for forward/backward |
| Gradients | 2 (bf16) | 4 | Same size as trainable weights |
| fp32 master weights | 4 | (included above) | Only in mixed precision |
| Adam first moment m | 4 | 4 | 8-bit Adam: 1 byte |
| Adam second moment v | 4 | 4 | 8-bit Adam: 1 byte |
| Total | 16 bytes | 16 bytes | Before activations |
Worked example: a 7B model
Take a 7B-parameter decoder with 32 layers, hidden size 4096, 32 attention heads and an MLP width of 11008 (a common 7B shape). Parameter count is about 6.7B; use 7B for round numbers.
| Method | Weights | Grads + optimizer | Total before activations | Fits on |
|---|---|---|---|---|
| Inference only, bf16 | 14 GB | 0 | 14 GB (+ KV cache) | One 24 GB GPU |
| Full FT, mixed precision AdamW | 14 GB | 7B × 14 B = 98 GB | ≈112 GB | Multiple 80 GB GPUs with sharding (ZeRO-3/FSDP) |
| Full FT, 8-bit Adam + bf16 grads | 14 GB | 7B × (2 + 2 + 4 master) ≈ 56 GB | ≈70 GB | One 80 GB GPU, tight |
| LoRA r=16, all linear layers, bf16 base | 14 GB | 40M × 16 B ≈ 0.64 GB | ≈15 GB | One 24 GB GPU |
| QLoRA r=16, NF4 base with double quantization | ≈3.9 GB | ≈0.64 GB (less with paged 8-bit optimizer) | ≈4.5 GB | One 8–12 GB GPU (with checkpointing) |
Where the 40M LoRA parameters come from: with rank r, an adapter on a din × dout linear layer adds r·(din + dout) parameters. Per layer: four attention projections of 4096×4096 give 4 × 16 × 8192 = 524,288; three MLP projections between 4096 and 11008 (gate, up, down) give 3 × 16 × 15104 = 724,992. That is ≈1.25M per layer, ×32 layers ≈ 40M parameters, or about 0.6% of the model. In bf16 the adapter file is about 80 MB. This 7B shape is multi-head attention; grouped-query models have thinner k_proj and v_proj, so the attention term is smaller. Embeddings and the LM head are not in this 40M unless you list them in modules_to_save.
Where the 3.9 GB NF4 figure comes from: 4 bits per weight is 0.5 bytes, so 7B weights take 3.5 GB. Blockwise quantization stores one scale per 64 weights; in fp32 that is 32/64 = 0.5 extra bits per weight, and double quantization reduces it to about 0.127 bits. Embeddings and the output head are often kept in 16-bit, which adds a few hundred MB.
Activations: the part people forget
For each transformer layer the backward pass needs intermediate tensors from the forward pass. A widely used estimate for 16-bit activations per layer, without recomputation, is roughly s·b·h·(34 + 5·a·s/h) bytes, where s is sequence length, b micro-batch size, h hidden size and a number of heads. The second term is the attention score matrix, which memory-efficient attention kernels (FlashAttention and SDPA) avoid storing.
These are estimates: exact numbers depend on MLP width, dropout masks and the framework. The important scaling rules are that activations grow linearly with batch size and layers, linearly with sequence length once fused attention is used (quadratically without it), and shrink dramatically with gradient checkpointing at a cost of roughly 20–35% more compute (one extra forward pass).
Levers to reduce memory, in the order people usually pull them
- Smaller micro-batch, more gradient accumulation. Same effective batch, far fewer activations.
- Gradient checkpointing. Recompute activations during backward; large savings for about a third more compute.
- Fused/flash attention. Removes the quadratic attention-matrix term.
- Shorter
max_seq_lengthbased on the data's length percentiles. - LoRA instead of full fine-tuning: removes gradients and optimizer states for the frozen weights.
- QLoRA: 4-bit base weights, about a quarter of the bf16 footprint.
- 8-bit or paged optimizers to shrink optimizer states and survive spikes.
- Sharding and offload: ZeRO-2/3 or FSDP across GPUs, CPU/NVMe offload of optimizer states or parameters.
- Chunked loss / fused cross-entropy: the logits tensor (batch × seq × vocab) is huge for 150k-token vocabularies; computing the loss in chunks avoids materializing it all.
The PEFT family: training a tiny fraction of the model
Parameter-efficient fine-tuning (PEFT) freezes the pretrained weights and trains a small number of new or selected parameters. It works because the change needed to adapt a large model to a task appears to live in a low-dimensional space: research on "intrinsic dimension" showed that many tasks can be learned by optimizing only a few hundred or thousand directions in parameter space. PEFT methods differ in where the trainable parameters are placed.
You bought a well-built house and want it to suit your family. You could rebuild every wall (full fine-tuning). Or you could add a small extension in the hallway (adapters), leave notes on the front door that everyone reads on the way in (prompt and prefix tuning), install dimmer switches on existing lights (IA3 scaling), or fit a thin overlay on existing walls that shifts them slightly in a few directions (LoRA). The house stays intact and you can remove the changes later.
Mapping back: the house is the frozen pretrained network, each renovation style is a PEFT method that inserts trainable parameters at a different place, and "removing the changes" is detaching the adapter to recover the original model.
| Method | What is trained | Trainable params | Inference overhead | Notes |
|---|---|---|---|---|
| Adapters (bottleneck) | Small MLP (down-project, nonlinearity, up-project) inserted after attention and/or FFN, with a residual | 0.5–5% | Extra sequential layers add latency | The original PEFT method for transformers |
| Prefix tuning | Trainable key/value vectors prepended at every layer | ~0.1% | Uses context length | Often reparameterized with an MLP during training for stability |
| Prompt tuning (soft prompts) | A few dozen trainable embedding vectors prepended to the input only | <0.01% | Uses context length | Competitive only for very large models |
| P-tuning | Soft prompt tokens produced by a small encoder, possibly at several layers | Small | Uses context length | Variant of prompt tuning |
| BitFit | Only bias terms | ~0.1% | None | Surprisingly decent for encoders; weak for LLM SFT |
| IA3 | Learned vectors that rescale keys, values and FFN activations | ~0.01% | None once merged | Extremely small |
| LoRA | Low-rank matrices A and B added to chosen linear layers | 0.1–2% | None once merged | The de facto standard |
| QLoRA | LoRA on a 4-bit quantized frozen base | Same as LoRA | Dequantization cost if served in 4-bit | Largest memory savings |
| DoRA | LoRA on direction plus a trained magnitude vector | LoRA + tiny | None once merged | Often closes the gap to full fine-tuning at low rank |
| Selective / partial | Only the top k layers, or layer norms, or embeddings | Varies | None | Simple baseline; freezes lower layers |
Additive, sequential (adapters)
- New modules in the forward path
- Cannot be folded into existing weights
- Adds latency at batch size 1
Reparameterization (LoRA, DoRA, IA3)
- Learn a change to existing weights
- Can be merged: W' = W + ΔW
- Zero extra latency after merging
LoRA: low-rank adaptation in depth
LoRA (Low-Rank Adaptation) is based on a hypothesis: the update a pretrained weight matrix needs during fine-tuning has low intrinsic rank. So instead of learning a full dout × din update ΔW, LoRA learns two thin matrices whose product has rank at most r, typically 8 to 64, while the original weight stays frozen.
An audio engineer wants to change the sound of a full orchestra recording. Rather than editing every instrument track, she adds a small equalizer with a handful of sliders. Each slider moves a whole family of frequencies together, and a master volume knob controls how strongly the equalizer applies. With just a few sliders she can give the recording a noticeably different character.
Mapping back: the orchestra tracks are the frozen weight matrix W, the small set of sliders is the low-rank pair B and A (the rank r is the number of sliders), each slider's broad effect is a rank-one direction, and the master knob is the scaling factor α/r.
The math
- Parameter count per adapted matrix:
r·(din + dout)instead ofdin·dout. For a 4096×4096 projection with r = 16: 131,072 versus 16.8M, a 128× reduction. - Initialization: A is random (Kaiming/Gaussian), B is zero. So at step 0, BA = 0 and the adapted model is exactly the pretrained model; training starts from a known-good function and is stable.
- Why not both zero? If both were zero, the gradient with respect to each would be zero (each gradient depends on the other matrix), so nothing would ever learn.
- Scaling α/r: keeps the update size roughly constant when you change r, so you do not need to retune the learning rate for every rank. A common convention is α = r or α = 2r. rsLoRA argues the right scaling is α/√r, which keeps learning stable at high ranks.
- Merging: after training, compute
W' = W0 + (α/r)BAonce. The merged model has exactly the original architecture and zero added latency. - Dropout: LoRA dropout (for example 0.05) is applied to the input of the adapter path only, as light regularization.
x (d_in)
│
┌────────┴─────────┐
│ │
┌──┴───┐ ┌───┴───┐
│ W0 │ frozen │ A │ r × d_in (random init)
│d_out │ └───┬───┘
│× d_in│ │ r-dim bottleneck
└──┬───┘ ┌───┴───┐
│ │ B │ d_out × r (zero init)
│ └───┬───┘
│ │ × (α / r)
└──────── + ───────┘
│
h (d_out)
Why it works so well
Empirically, fine-tuning updates of large pretrained models have rapidly decaying singular values: most of the change is captured by a few directions. LoRA exploits this, and it also acts as a regularizer: a low-rank update cannot rewrite the whole matrix, which is part of why LoRA forgets less than full fine-tuning (a finding often summarized as "LoRA learns less and forgets less"). The flip side is that for large distribution shifts, such as teaching a new language or continued pretraining on code, full fine-tuning or high-rank LoRA learns more.
Choosing rank, alpha and target modules
| Knob | Typical values | Guidance |
|---|---|---|
| Rank r | 8, 16, 32, 64 | 8–16 is usually enough for format/style SFT; 32–128 for harder skills or larger shifts. Doubling r doubles adapter params; gains often flatten quickly. |
| Alpha α | r or 2r (16, 32) | Effective scale is α/r. Keep the ratio fixed when sweeping rank to isolate capacity from step size. |
| Target modules | All linear layers: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj | The original paper adapted only q and v; later work (including QLoRA) found adapting all linear layers matters more than raising rank. PEFT accepts target_modules="all-linear". |
| Dropout | 0.0–0.1 | 0.05 is common; use more on tiny datasets. |
| Bias | "none" | Training biases adds little. |
| Learning rate | 1e-4 to 3e-4 (2e-4 typical) | About 10–20× higher than full fine-tuning; the frozen base and zero-initialized B make large steps safe. |
modules_to_save | embed_tokens, lm_head | Only when you add new tokens (for example a chat template for a base model); those rows must be trained fully. |
LoRA in code with PEFT
from peft import LoraConfig, get_peft_model, TaskType
lora_cfg = LoraConfig(
task_type=TaskType.CAUSAL_LM,
r=16,
lora_alpha=32, # scale = alpha / r = 2
lora_dropout=0.05,
bias="none",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
)
model = get_peft_model(base_model, lora_cfg)
model.print_trainable_parameters()
# e.g. for a 0.5B model with all-linear r=16:
# trainable params: ~8.8M || all params: ~503M || trainable%: ~1.75
A from-scratch LoRA layer
import math, torch, torch.nn as nn
class LoRALinear(nn.Module):
def __init__(self, base: nn.Linear, r=16, alpha=32, dropout=0.05):
super().__init__()
self.base = base
for p in self.base.parameters():
p.requires_grad = False # freeze W0 (and bias)
self.A = nn.Parameter(torch.empty(r, base.in_features))
self.B = nn.Parameter(torch.zeros(base.out_features, r)) # zero init
nn.init.kaiming_uniform_(self.A, a=math.sqrt(5))
self.scale = alpha / r
self.drop = nn.Dropout(dropout)
def forward(self, x):
return self.base(x) + self.scale * (self.drop(x) @ self.A.T @ self.B.T)
@torch.no_grad()
def merge(self):
self.base.weight += self.scale * (self.B @ self.A) # W' = W0 + (alpha/r) BA
return self.base
Variants worth knowing
DoRA
Weight-Decomposed LoRA splits each weight into a magnitude vector and a direction: W' = m · (W0 + BA) / ‖W0 + BA‖c (column-wise norm). LoRA updates the direction, and m is trained separately. This mimics full fine-tuning's learning pattern more closely and often helps at low rank, at some extra training cost. Mergeable after training.
rsLoRA
Rank-stabilized LoRA uses scale α/√r instead of α/r so that higher ranks actually learn faster rather than being damped. Useful when sweeping large ranks.
LoRA+
Uses a larger learning rate for B than for A (for example 16×), which improves convergence speed because the two matrices play asymmetric roles.
AdaLoRA
Allocates rank adaptively across layers using an SVD-like parameterization and prunes unimportant singular values during training.
PiSSA / LoftQ
Smarter initialization: PiSSA initializes A and B from the top singular vectors of W; LoftQ initializes adapters to compensate for quantization error in QLoRA.
VeRA and tied variants
Share frozen random A and B across layers and train only small scaling vectors, shrinking adapters by another 10× or more.
QLoRA: fine-tuning on a 4-bit base
LoRA removed gradients and optimizer states for frozen weights, but the frozen weights themselves still occupy 2 bytes per parameter in bf16. QLoRA stores the frozen base model in 4-bit and trains ordinary 16-bit LoRA adapters on top. During each forward and backward pass, weights are dequantized block by block to the compute dtype, used in the matrix multiply and discarded. Gradients flow through the dequantized weights to the adapters, but the 4-bit weights never change. QLoRA introduced three ideas that make this work without losing quality.
A translator works from a huge reference encyclopedia but has a small desk. She keeps the encyclopedia as a compact pocket edition with abbreviated entries, expanding each entry to full text only while she reads it, and she writes her own corrections in a thin, full-quality notebook next to it. If the desk briefly overflows, she moves a few notebook pages to a shelf behind her and brings them back when needed.
Mapping back: the pocket edition is the NF4-quantized base, expanding an entry is on-the-fly dequantization to bf16/fp16, the full-quality notebook is the 16-bit LoRA adapters that are actually trained, and moving pages to the shelf is the paged optimizer spilling optimizer state to CPU memory during spikes.
1. 4-bit NormalFloat (NF4)
Pretrained weights are roughly normally distributed around zero. A uniform 4-bit grid (INT4) wastes levels in the sparse tails and has too few near zero, where most weights sit. NF4 places its 16 levels at the quantiles of a standard normal distribution, rescaled to [−1, 1], so each level is used about equally often: an information-theoretically near-optimal code for normally distributed data. It includes an exact zero.
Quantization is blockwise: weights are split into blocks of 64, each block is scaled by its absolute maximum (absmax) so values fall into [−1, 1], and each value is mapped to the nearest NF4 level. Blockwise scaling limits the damage an outlier can do to one block rather than a whole tensor.
2. Double quantization
Each block of 64 weights needs its own fp32 scale: 32 bits / 64 weights = 0.5 extra bits per parameter, which is 0.44 GB for a 7B model. Double quantization quantizes those scales too: the fp32 constants are grouped into blocks of 256 and quantized to 8-bit with one fp32 second-level constant per group. The overhead falls to about 8/64 + 32/(64·256) ≈ 0.127 bits per parameter, saving roughly 0.37 bits per parameter (about 0.3 GB on 7B, several GB on 65B).
3. Paged optimizers
Memory use during training is spiky, particularly with long sequences and gradient checkpointing. Paged optimizers allocate optimizer states in NVIDIA unified memory, which automatically pages them to CPU RAM when the GPU is about to run out and pages them back when needed, turning a crash into a small slowdown. In practice you select optim="paged_adamw_8bit" or "paged_adamw_32bit".
Compute dtype and the full recipe
- Storage dtype is NF4; compute dtype is bf16 (or fp16 on GPUs without bf16 support). Matmuls always happen in 16-bit.
prepare_model_for_kbit_trainingcasts layer norms (and usually the output head) to fp32 for stability, enables gradient checkpointing, and makes input embeddings require gradients so gradients can flow through a frozen, checkpointed network into the adapters.- Adapters on all linear layers were key to QLoRA matching 16-bit full fine-tuning in the original experiments; the headline result was fine-tuning a 65B model on a single 48 GB GPU.
- Speed: QLoRA is typically 20–40% slower per step than LoRA because of dequantization, a trade of time for memory.
import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
from peft import prepare_model_for_kbit_training, get_peft_model, LoraConfig
bnb_cfg = BitsAndBytesConfig(
load_in_4bit=True, # store frozen weights in 4-bit
bnb_4bit_quant_type="nf4", # NormalFloat4 levels (vs "fp4")
bnb_4bit_use_double_quant=True, # quantize the per-block scales too
bnb_4bit_compute_dtype=torch.bfloat16, # float16 on T4-class GPUs
)
base = AutoModelForCausalLM.from_pretrained(
"your-org/your-7b-model", quantization_config=bnb_cfg, device_map="auto")
base = prepare_model_for_kbit_training(base, use_gradient_checkpointing=True)
model = get_peft_model(base, LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,
target_modules="all-linear",
task_type="CAUSAL_LM"))
| LoRA | QLoRA | |
|---|---|---|
| Frozen base precision | bf16/fp16 (2 bytes) | NF4 (≈0.53 bytes with double quant) |
| Trainable parameters at the same rank | Identical | Identical |
| Peak memory, 7B | ≈16–20 GB | ≈6–10 GB |
| Step speed | Faster | Slower (dequantization) |
| Quality | Reference | Usually within noise of LoRA |
| Merging | Straightforward | Dequantize base to 16-bit first, then merge |
For the broader quantization picture (PTQ, GPTQ, AWQ, INT8/INT4 inference, quantization-aware training and on-device formats) see Edge AI. Quantization-aware training combined with LoRA is also used to recover accuracy for aggressively quantized on-device models.
Hyperparameters that matter
Fine-tuning has fewer knobs than pretraining, but the defaults differ by method, and a wrong learning rate is the most common reason a run either does nothing or destroys the model. Treat the table below as starting points, then tune on validation loss and, more importantly, on task metrics.
Tuning a fine-tuning run is like adjusting a car for a new track. The learning rate is how hard you press the accelerator; warmup is easing out of the pit lane before flooring it; the schedule is braking smoothly into the finish; batch size is how many laps you average before adjusting the setup; epochs are how many practice sessions you run before you start memorizing the track instead of learning to drive.
Mapping back: too much accelerator (learning rate) and the car spins off (divergence or forgetting), no warmup and the tires lose grip at the start (unstable early updates while Adam's statistics are still poor), and too many sessions produce a driver who only knows one track (overfitting).
| Hyperparameter | Full fine-tuning | LoRA / QLoRA | DPO and friends |
|---|---|---|---|
| Learning rate | 5e-6 to 2e-5 (smaller for bigger models) | 1e-4 to 3e-4 | 5e-7 to 5e-6 (full), ~5e-6 to 5e-5 (LoRA) |
| Epochs | 1–3 | 1–3 (up to 5 on tiny sets) | 1–3 |
| Effective batch size | 32–512 sequences | 16–128 | 32–128 |
| Warmup | 3–10% of steps | 3–10% | ~10% |
| Schedule | Cosine or linear decay | Cosine or linear decay (constant also works) | Linear or cosine |
| Weight decay | 0–0.1 | 0–0.01 | 0 |
| Gradient clipping | 1.0 | 0.3–1.0 | 1.0 |
| Optimizer | AdamW (fused), 8-bit AdamW | AdamW, paged 8-bit AdamW | AdamW |
| Other | Gradient checkpointing, ZeRO/FSDP | r 8–64, α = r or 2r, dropout 0.05 | β = 0.1 (range 0.01–0.5) |
The rules behind the numbers
- Learning rate is the most important knob. Full fine-tuning moves pretrained weights directly, so it needs tiny steps to avoid wrecking learned features. LoRA starts from a zero update on a frozen base and can take larger steps.
- Effective batch size = micro-batch × gradient-accumulation steps × number of GPUs. When you increase it substantially, raise the learning rate a bit (square-root scaling is a safe heuristic for Adam).
- Warmup protects the first steps, when Adam's second-moment estimates are poor and a few large updates can damage the model.
- Epochs: SFT data is often seen only 1–3 times. Many passes over a small set drive the training loss towards zero while validation loss rises, the classic overfitting sign, and the model starts repeating training answers verbatim.
- Sequence length: set from data percentiles; longer costs memory linearly (with fused attention) and wastes compute on padding if most examples are short.
- Steps, not epochs, are what matter for compute planning: steps per epoch = ⌈N / effective batch⌉.
- Seed and reproducibility: fix seeds for Python, NumPy and PyTorch; small-data fine-tuning can vary noticeably between seeds, so compare methods over several seeds if differences are small.
Reading the loss curves
| Pattern | Likely cause | Action |
|---|---|---|
| Train and validation both flat from the start | LR too low, adapters not attached, all params frozen, mask all zeros | Check trainable-parameter count and mask; raise LR |
| Loss spikes or becomes NaN | LR too high, fp16 overflow, bad example, no clipping | Lower LR, use bf16 or fp32, clip gradients, inspect batch |
| Train falls, validation falls then rises | Overfitting | Early stop at best validation checkpoint, fewer epochs, more data, more dropout |
| Initial loss already very low | Model already knows the data, or leakage | Check for duplicates between train and validation |
| Staircase drops each epoch | Memorization of repeated examples | Usually a warning sign; reduce epochs |
| Validation loss lower but outputs worse | Loss is a proxy; verbose or bland outputs | Evaluate with generation-based metrics and a judge |
The tooling landscape and distributed training
The open-source stack for fine-tuning is layered. At the bottom are PyTorch and GPU kernels; above that are model libraries; then training libraries that implement SFT and preference losses; then distributed-training backends; and finally configuration-driven wrappers. Knowing which layer each tool occupies makes them easy to compare.
Building a house involves raw materials, prefabricated components, specialist trades, heavy machinery for big jobs, and a general contractor who coordinates everything from a plan. You can hire a contractor with a standard plan or manage the trades yourself for full control.
Mapping back: PyTorch and kernels are the raw materials, Transformers and PEFT are the prefabricated components, TRL's trainers are the specialist trades, DeepSpeed and FSDP are the heavy machinery for multi-GPU jobs, and YAML-driven tools like Axolotl are the general contractor working from a plan.
| Tool | Layer | What it provides |
|---|---|---|
| Hugging Face Transformers | Models | Model classes, tokenizers with chat templates, Trainer, generation, quantization integrations |
| PEFT | Adapters | LoRA, DoRA, IA3, prefix/prompt tuning, adapter saving/loading, merging, multi-adapter management |
| TRL | Training recipes | SFTTrainer (packing, completion-only loss, chat formatting), RewardTrainer, DPOTrainer, ORPOTrainer, KTOTrainer, PPOTrainer, GRPOTrainer |
| bitsandbytes | Kernels | 8-bit and 4-bit (NF4/FP4) quantized linear layers, 8-bit and paged optimizers |
| Accelerate | Launch/distribution | One script that runs on CPU, one GPU, many GPUs or many nodes; mixed precision; wraps DeepSpeed and FSDP |
| DeepSpeed | Distribution | ZeRO stages 1–3, CPU/NVMe offload, pipeline parallelism |
| PyTorch FSDP | Distribution | Native fully sharded data parallelism (parameters, gradients and optimizer states sharded) |
| Unsloth | Optimized trainer | Hand-written kernels and memory tricks for faster, lower-memory LoRA/QLoRA, mainly single-GPU; notebooks for popular models; GGUF export |
| Axolotl | Config wrapper | YAML-configured SFT/DPO/QLoRA across many models and distributed backends |
| LLaMA-Factory, torchtune | Config wrappers / recipes | Similar config-driven recipes; torchtune is PyTorch-native |
| FlashAttention, Liger kernels | Kernels | Memory-efficient attention; fused RMSNorm/RoPE/cross-entropy to cut memory |
| Hosted fine-tuning APIs | Managed service | Upload JSONL, receive a tuned model endpoint; limited control over method and hyperparameters, no weight access |
Distributed strategies
DDP (data parallel)
Each GPU holds a full copy of model, gradients and optimizer; batches are split and gradients all-reduced. Simple and fast, but the model must fit on each GPU.
ZeRO-1
Shard optimizer states across GPUs. Memory per GPU for the 12-byte optimizer part divides by the number of GPUs.
ZeRO-2
Also shard gradients. Good default for LoRA or mid-size full fine-tuning.
ZeRO-3 / FSDP full shard
Also shard parameters; each layer is gathered just-in-time for compute. Lets you fully fine-tune models far larger than one GPU, at the cost of more communication.
Offload
Move optimizer states (and even parameters) to CPU RAM or NVMe. Big memory win, significant slowdown.
Tensor / pipeline parallel
Split individual layers (tensor) or groups of layers (pipeline) across GPUs. Needed for very large models; more common in pretraining frameworks.
With ZeRO-3 on 8 GPUs, the 112 GB of weights, gradients and optimizer states for full 7B fine-tuning becomes about 14 GB per GPU, leaving room for activations on 40–80 GB cards.
Minimal SFT with TRL
from datasets import load_dataset
from trl import SFTTrainer, SFTConfig
from peft import LoraConfig
ds = load_dataset("json", data_files={"train": "train.jsonl", "eval": "val.jsonl"})
cfg = SFTConfig(
output_dir="out/sft-lora",
num_train_epochs=2,
per_device_train_batch_size=4,
gradient_accumulation_steps=4, # effective batch 16 per GPU
learning_rate=2e-4,
lr_scheduler_type="cosine",
warmup_ratio=0.05,
max_grad_norm=1.0,
bf16=True,
gradient_checkpointing=True,
max_length=1024, # named max_seq_length in older versions
packing=False,
assistant_only_loss=True, # loss on assistant turns only (recent versions)
eval_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
logging_steps=10,
)
trainer = SFTTrainer(
model="your-org/your-8b-instruct",
args=cfg,
train_dataset=ds["train"], # records with a "messages" field
eval_dataset=ds["eval"],
peft_config=LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,
target_modules="all-linear", task_type="CAUSAL_LM"),
)
trainer.train()
trainer.save_model("out/sft-lora/adapter")
max_seq_length became max_length, and completion-only masking moved from a separate collator to config flags). Pin versions in requirements.txt and check the installed version's documentation rather than copying snippets blindly.Alignment with RLHF
After SFT, a model imitates demonstrations, but imitation has limits: there are many acceptable answers, demonstrators vary in quality, and it is much easier for people to compare two answers than to write a perfect one. Reinforcement Learning from Human Feedback (RLHF) uses those comparisons. It was the method behind the first widely used chat assistants, where annotators preferred a 1.3B aligned model over a 175B base model, showing that alignment can matter more than size for perceived helpfulness.
Training a new chef for a restaurant: first she follows the recipe book (SFT). Then a food critic tastes many pairs of dishes and says which one is better; over time an assistant learns to predict the critic's verdict (reward model). The chef then experiments, getting quick scores from the assistant, while the head chef insists she must not stray too far from the house style (KL penalty), because the assistant can be fooled by, say, piling on salt.
Mapping back: the recipe book is the SFT model, the critic is human labellers, the assistant is the learned reward model, experimenting with scores is PPO optimization, the house-style rule is the KL penalty to the reference model, and fooling the assistant with salt is reward hacking.
The three-stage pipeline
- Supervised fine-tuning. Train the base model on human-written demonstrations of good responses. This teaches format and produces the reference policy πref.
- Reward model training. Sample several responses per prompt from the SFT model; humans rank them. Train a reward model (usually the SFT model with a scalar head replacing the vocabulary head) to give preferred responses higher scores.
- RL optimization (PPO). Generate responses with the policy, score them with the reward model, and update the policy to increase reward while a KL penalty keeps it close to πref.
prompts ──▶ [SFT model] ──▶ y1, y2, ... ──▶ humans rank ──▶ (x, y_w, y_l) pairs
│
▼
[Reward model r_φ(x, y)]
│ score
prompts ──▶ [Policy π_θ] ──▶ response y ──────────────────────▶│
▲ ▼
│ PPO update ◀── reward = r_φ(x,y) − β·KL(π_θ ‖ π_ref)
└────────────────── (value model estimates advantages)
Reward model: Bradley–Terry loss
The RL objective
PPO optimizes this with a clipped surrogate objective. For each generated token it computes the probability ratio ρ = πθ(a|s) / πold(a|s) and an advantage estimate A (from a learned value model using generalized advantage estimation), then maximizes:
Classic PPO-based RLHF juggles four models in memory: the policy, the frozen reference, the reward model and the value (critic) model. That is expensive and notoriously finicky (sensitive to KL coefficient, reward normalization, generation length and many implementation details), which motivated the simpler methods in the next section.
RLAIF and Constitutional AI
Human labels are slow and costly (large programmes take weeks to months). RLAIF (RL from AI Feedback) replaces human preference labels with judgements from a capable LLM prompted with guidelines. Constitutional AI is a structured version: a written list of principles (the "constitution") is used first in a supervised phase where the model critiques and revises its own harmful responses, and then in an RL phase where an AI judge chooses between responses according to the principles. It scales harmlessness training with far less human labelling of harmful content, and makes the target values explicit and auditable.
Known failure modes of RLHF
- Reward hacking: the policy finds outputs the reward model overrates (overly long answers, flattering language, certain phrases). Mitigations: KL penalty, reward-model ensembles, length penalties, periodic reward-model retraining.
- Sycophancy: agreeing with the user because labellers tend to prefer agreeable answers.
- Verbosity bias: longer answers win comparisons, so the policy learns to pad.
- Mode collapse / reduced diversity: aligned models become less creative and less calibrated than base models.
- Alignment tax: small drops on some capability benchmarks after alignment; mixing pretraining-style loss into RL (as some pipelines do) reduces it.
Beyond PPO: DPO, ORPO, KTO, GRPO and friends
Since 2023 a wave of methods has simplified preference learning. They fall into two groups: offline, RL-free methods that learn directly from a fixed preference dataset with a classification-like loss (DPO, IPO, ORPO, KTO, SimPO), and online RL methods that still sample from the current policy but drop expensive components (GRPO, RLOO). Online methods with verifiable rewards (unit tests, exact answers) power modern reasoning models.
Learning to write better essays. PPO is having a tutor who predicts a teacher's grades and coaching you draft by draft. DPO is being handed a stack of essay pairs already marked "this one is better" and directly studying what distinguishes them, with no tutor. GRPO is writing eight drafts of the same essay, having an automatic checker score them, and learning from which of your own drafts beat your average.
Mapping back: the tutor is the reward and value models, the marked pairs are an offline preference dataset used by DPO, and comparing drafts against their own average is GRPO's group-relative advantage, which replaces the learned critic.
DPO: Direct Preference Optimization
DPO's insight is that the KL-regularized RLHF objective has a closed-form optimal policy: π*(y|x) ∝ πref(y|x)·exp(r(x,y)/β). Rearranging, the reward can be expressed through the policy itself: r(x,y) = β·log(π(y|x)/πref(y|x)) + const. Substituting this "implicit reward" into the Bradley–Terry loss eliminates the reward model and the RL loop.
- Needs: a preference dataset and two models (policy + frozen reference). With LoRA, the reference is just the base with adapters disabled, so only one copy is held in memory.
- Pros: simple, stable, cheap, no sampling during training.
- Cons: offline, so it cannot explore beyond the dataset; can overfit to preferences (driving both chosen and rejected likelihoods down); sensitive to how far the data is from the policy's own outputs; can increase verbosity.
- Metrics to watch: reward accuracy (fraction of pairs where implicit reward of chosen > rejected), reward margins, and the log-probabilities of chosen responses (if they fall steadily, the model may be degrading).
from trl import DPOTrainer, DPOConfig
# dataset rows: {"prompt": [...messages...], "chosen": [...], "rejected": [...]}
cfg = DPOConfig(output_dir="out/dpo", beta=0.1, learning_rate=5e-6,
per_device_train_batch_size=2, gradient_accumulation_steps=8,
num_train_epochs=1, warmup_ratio=0.1, bf16=True,
max_length=1024)
trainer = DPOTrainer(model=sft_model, ref_model=None, # None + PEFT: base acts as reference
args=cfg, train_dataset=pref_ds,
processing_class=tokenizer, peft_config=lora_cfg)
trainer.train()
The rest of the family
| Method | Data | Reference model? | Key idea |
|---|---|---|---|
| IPO | Pairs | Yes | Replaces the log-sigmoid with a squared loss towards a target margin, preventing DPO's tendency to push margins to infinity on deterministic preferences. |
| ORPO | Pairs | No | Combines SFT and preference learning in one stage: SFT loss on chosen plus an odds-ratio term that favours the odds of chosen over rejected. No separate SFT stage and no reference model. |
| KTO | Unpaired thumbs up/down | Yes | Inspired by prospect theory; learns from individual examples labelled desirable or undesirable, which is how production feedback is usually collected. |
| SimPO | Pairs | No | Uses length-normalized average log-probability as the implicit reward plus a target margin; reference-free and reduces length exploitation. |
| CPO, RRHF, SLiC | Pairs / rankings | Varies | Ranking or contrastive losses over candidate responses. |
| Online / iterative DPO | Fresh pairs from the current policy, labelled by a reward model or judge | Yes | Reduces the off-policy gap of standard DPO by regenerating preference data each round. |
| RLOO | Prompts + reward | Yes (KL) | REINFORCE with a leave-one-out baseline from k samples; no value model. |
| GRPO | Prompts + reward function | Yes (KL) | Samples a group of G responses per prompt and uses group-normalized rewards as advantages; no critic. |
GRPO and RL with verifiable rewards
Group Relative Policy Optimization keeps PPO's clipped objective and KL control but removes the value model. For each prompt, sample G responses (for example 8–16), score each with a reward function, and set each response's advantage to its reward standardized within the group. Responses better than their siblings are reinforced; worse ones are discouraged. This halves memory relative to PPO and works especially well when rewards are verifiable: a math answer matches the reference, code passes unit tests, output parses as valid JSON. Training on such rewards (often called RLVR) is how models learn long chain-of-thought reasoning, with behaviours like self-checking emerging from reward alone.
from trl import GRPOTrainer, GRPOConfig
def correctness_reward(completions, answer, **kwargs):
# 1.0 if the final line contains the reference answer, else 0.0
return [1.0 if a in c[-1]["content"].splitlines()[-1] else 0.0
for c, a in zip(completions, answer)]
def format_reward(completions, **kwargs):
return [0.2 if "<answer>" in c[-1]["content"] else 0.0 for c in completions]
cfg = GRPOConfig(output_dir="out/grpo", num_generations=8, learning_rate=1e-6,
max_completion_length=512, beta=0.04, bf16=True)
trainer = GRPOTrainer(model=sft_model, reward_funcs=[correctness_reward, format_reward],
args=cfg, train_dataset=math_prompts) # rows: prompt, answer
trainer.train()
GRPO versus PPO
This comparison is a favourite interview follow-up. Both maximize a clipped probability-ratio objective; they differ in how the advantage is estimated and what sits in GPU memory.
| PPO (classic RLHF) | GRPO | |
|---|---|---|
| Advantage | Learned value (critic) + GAE, typically per token | Group-normalized reward: (r − mean)/std over G siblings; same A for every token in a response |
| Models in memory | Policy, frozen reference, reward model, value model (four) | Policy and frozen reference (two); the scorer is often a cheap verifier, not a neural RM |
| Samples per prompt | Usually one (or a few) on-policy rollout | A group of G (commonly 8–16) |
| Clip and KL | Yes: clip(ρ, 1−ε, 1+ε) and a KL penalty to πref | Same clip; KL either as a loss term or folded into the reward before grouping |
| Best fit | Subjective quality where you already have (or will train) a reward model | Verifiable rewards: math answers, unit tests, schema checks (RLVR) |
| Typical failure | Reward hacking, critic instability, KL/length sensitivity | Zero-variance groups (all right or all wrong) give no gradient |
RLOO is the close cousin: it also drops the critic, but uses a leave-one-out mean of the other samples as the baseline instead of a group z-score. Use PPO when you need a learned reward for fuzzy preferences; use GRPO/RLOO when a checker can score each sample.
Choosing a method
| Situation | Reasonable choice |
|---|---|
| Have good demonstrations only | SFT |
| Have chosen/rejected pairs, want simplicity | SFT then DPO (or ORPO in one stage) |
| Only thumbs-up/down logs from production | KTO |
| Answers can be checked automatically (math, code, schema) | GRPO / RLOO with verifiable rewards |
| Subjective quality, large budget, strong RL team | Reward model + PPO, or online DPO with a reward model |
| Need harmlessness at scale without human red-team labels | Constitutional AI / RLAIF |
Evaluating a fine-tuned model
A lower loss is not the goal; better behaviour on the target task without breaking everything else is. Evaluation for fine-tuning therefore has two halves: did it learn the task? (held-out task metrics, human or judge ratings) and did it keep what it had? (general benchmarks, safety, regression tests). Always evaluate the base model with its best prompt on the same sets, so improvements are measured against a fair baseline. The broader methodology is covered in LLM Evaluation & Safety.
A hospital hiring a newly specialised surgeon checks two things: can she perform the new procedure well (a practical exam on unseen cases), and has she still got the general skills every doctor needs (routine competency checks)? A great specialty score is worthless if she has forgotten basic hygiene.
Mapping back: unseen cases are the held-out test set, the practical exam is task metrics and judge ratings, and the routine competency checks are general benchmarks, safety prompts and regression tests that detect catastrophic forgetting.
Layers of evaluation
| Layer | What it measures | Examples |
|---|---|---|
| Training signals | Optimization health | Train/validation loss, perplexity (eloss), gradient norm |
| Task metrics on held-out data | Did it learn the task? | Accuracy, F1, exact match, JSON validity rate, pass@k for code, ROUGE-L, BERTScore |
| LLM-as-judge | Open-ended quality | Pairwise win rate vs base model, rubric scores for correctness, helpfulness, safety |
| Human evaluation | Ground truth for subjective quality | Blind pairwise comparisons by domain experts on a sample |
| General benchmarks | Forgetting | MMLU-style knowledge, GSM8K-style math, HumanEval-style coding, IFEval-style instruction following, MT-Bench / Arena-style chat quality |
| Safety and red-teaming | Safety regression | Harmful-request refusal rate, jailbreak suites, over-refusal on benign prompts, toxicity, bias probes |
| Regression tests | Product behaviour | A fixed suite of real prompts with assertions (format, must-mention, must-not-say), run in CI for every new checkpoint |
| Online metrics | Real impact | A/B tests: resolution rate, thumbs-up rate, escalations, latency, cost |
Lexical vs semantic metrics
ROUGE-L
- Longest common subsequence of words between prediction and reference
- Rewards saying the same words in the same order
- Penalizes valid paraphrases
- Cheap, deterministic
BERTScore-F1
- Greedy matching of contextual token embeddings
- Rewards similar meaning with different wording
- Scores are compressed into a narrow high range; compare relatively
- Can miss factual errors that use similar words
Reporting both gives a fuller picture: if ROUGE rises but BERTScore does not, the model may just be copying phrasing; if BERTScore rises with flat ROUGE, it paraphrases correctly. When references are themselves machine-generated or noisy, weight relative differences between methods more than absolute values.
LLM-as-judge done right
- Use a pairwise format (model A vs model B) with a clear rubric; pairwise judgements are more reliable than absolute 1–10 scores.
- Swap positions and average to cancel position bias; control for verbosity bias (judges favour longer answers) and self-preference (judges favour their own family's style).
- Validate the judge against a few hundred human labels and report agreement.
- Use a judge that is stronger than, and different from, the model being evaluated.
Evaluation hygiene
- Same prompts, same decoding for every model (for example greedy, same
max_new_tokens), and decode only the newly generated tokens. - Use the chat template at evaluation exactly as in training.
- Check contamination: n-gram or embedding overlap between training data and test sets.
- Error analysis: read failures, bucket them (too short, too long, hallucinated fact, missing caveat, wrong format), and compute simple statistics such as the prediction/reference length ratio and counts of empty or truncated outputs.
- Report cost alongside quality: trainable parameters, peak memory, training time, inference latency.
- Several seeds when differences are small, with confidence intervals (bootstrap over test examples).
Risks and failure modes
Fine-tuning is powerful precisely because it changes the model, and every change can break something. The main risks are losing general capabilities, memorizing instead of generalizing, eroding safety behaviour, leaking data, and violating licences. Each has known mitigations.
A pianist who practises only one concerto for months may play it flawlessly yet stumble on pieces she once knew, may start playing that concerto's phrasing everywhere, and may even forget the etiquette of performing. If she learned the piece from a pirated score, she may also be barred from performing it publicly.
Mapping back: stumbling on old pieces is catastrophic forgetting, over-applied phrasing is overfitting, forgotten etiquette is safety regression, and the pirated score is a licensing problem with training data or base-model terms.
Catastrophic forgetting
Gradient updates for the new task overwrite weights that encoded older skills. Signs: general benchmark drops, loss of multilingual ability, broken chat format, worse reasoning. Mitigations: LoRA instead of full fine-tuning, lower LR, fewer epochs, mix 5–20% general instruction data (replay), keep the chat template, regularize towards the original weights (L2-SP, EWC), or merge the fine-tuned model back towards the base.
Overfitting and memorization
Small data plus many epochs yields verbatim regurgitation, repetitive phrasing and poor generalization. Mitigations: early stopping on validation, more diverse data, dropout, lower rank, deduplication.
Safety regression
Research has shown that fine-tuning even on a handful of harmful examples, and sometimes on entirely benign data, can weaken a model's refusal behaviour. Mitigations: include safety examples in the mix, run refusal and jailbreak suites before release, add inference-time guardrails, keep adapters small.
Hallucination from new facts
Training on facts the model did not know teaches it to answer confidently beyond its knowledge. Keep facts in RAG; use fine-tuning for behaviour; include "I don't know" examples.
Data leakage and privacy
Personal information, secrets or customer data in training sets can be extracted later through prompting (membership-inference and extraction attacks). Scrub data, restrict access to tuned weights, consider differential privacy for sensitive data.
Evaluation leakage
Test examples or benchmark items present in training data inflate scores. Decontaminate and split by group.
Licence and terms issues
Base-model licences range from permissive (Apache 2.0, MIT) to custom community licences with usage restrictions and attribution requirements. Datasets have their own licences. Many API providers' terms prohibit using outputs to train competing models. Record provenance for every data source.
Reward hacking and sycophancy
In preference tuning, the model exploits weaknesses of the reward or judge (length, flattery). Monitor length, use KL control, diverse judges, and human spot checks.
Template and tokenizer mismatch
Training with one chat format and serving with another, or adding tokens without resizing embeddings, silently degrades quality. Test the exact serving path end to end.
Data poisoning and backdoors
Malicious examples can implant trigger phrases that cause specific behaviour. Vet data sources, especially scraped or crowd-sourced data, and scan third-party adapters before use.
A mitigation recipe for forgetting
- Measure first. Build a general regression suite and score the base model.
- Prefer PEFT. LoRA with moderate rank; the base weights stay intact and the adapter can be scaled down (for example multiply the LoRA scale by 0.5) if it is too strong.
- Mix data. Add general instruction and chat examples (replay) alongside the domain data.
- Train gently. Lower LR, 1–2 epochs, early stopping.
- Consider model merging. Interpolating fine-tuned and base weights often recovers general ability with a small loss of task gains.
Merging adapters and merging models
"Merging" means two different things. Adapter merging folds a LoRA update into its base weights so the result is a normal model. Model merging combines several fine-tuned models (or adapters) that share a base into one model with the skills of each, using arithmetic on weights and no further training. Both are cheap and widely used.
Think of transparent overlays on a map. Adapter merging is permanently printing one overlay onto the map so you no longer need to carry it separately. Model merging is taking overlays drawn by different people on copies of the same base map (one marks hiking trails, another marks restaurants) and combining them onto one sheet, resolving places where they disagree.
Mapping back: the base map is the shared pretrained model, each overlay is a weight delta (a task vector or LoRA update), printing is merge_and_unload, and resolving disagreements is what TIES and DARE do with conflicting parameter updates.
Adapter merging
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("your-org/base-model",
torch_dtype=torch.bfloat16) # 16-bit, not 4-bit
model = PeftModel.from_pretrained(base, "out/sft-lora/adapter")
merged = model.merge_and_unload() # W' = W0 + (alpha/r) * B @ A, adapters removed
merged.save_pretrained("out/merged", safe_serialization=True)
AutoTokenizer.from_pretrained("out/sft-lora/adapter").save_pretrained("out/merged")
- Always merge into a 16-bit copy of the exact base used in training (same revision).
- Save the tokenizer and chat template with the merged model.
- PEFT can also combine several LoRAs into one adapter (
add_weighted_adapterwith methods such as linear, SVD, TIES or DARE), or keep many adapters loaded and switch withset_adapter.
Model merging methods
| Method | How it works | When to use |
|---|---|---|
| Linear / model soup | Weighted average of weights of models fine-tuned from the same base | Averaging runs with different seeds or hyperparameters; often improves robustness |
| SLERP | Spherical interpolation between two models' weights, preserving norm geometry | Blending exactly two models smoothly |
| Task arithmetic | Task vector τ = θft − θbase; add λτ to gain a skill, subtract to remove a behaviour | Combining or negating skills |
| TIES | Trim small deltas, elect a sign per parameter by majority mass, average only agreeing deltas | Merging several task models with interference |
| DARE | Randomly drop a large fraction (for example 90%) of delta parameters and rescale the rest by 1/(1−p) | Pre-processing before TIES or linear merges; reduces interference |
| Frankenmerges / passthrough | Stack layers from different models into a deeper model | Experimental; unpredictable |
Configuration-driven tools such as mergekit implement these methods. Merged models must be evaluated like any other: merges can silently lose instruction following or produce format mixtures.
Knowledge distillation
Distillation trains a small student model to mimic a large teacher. For LLMs it is the main way to obtain small models that behave far above their size: the student learns from the teacher's outputs (and ideally its full probability distribution), which carry much more information per example than one-hot labels.
A master chess coach does not just tell a student the best move; she explains that two other moves were almost as good and one looked tempting but loses. The student who hears "move A 60%, move B 30%, move C 10%" learns the coach's judgement far faster than one who only hears "play A".
Mapping back: the coach is the teacher model, the percentages are its soft probability distribution over next tokens ("dark knowledge"), and learning from the full distribution rather than just the top choice is logit distillation, versus sequence-level distillation that trains on the teacher's final text.
Forms of distillation
| Form | Signal | Needs teacher logits? | Notes |
|---|---|---|---|
| Sequence-level (black-box) | Teacher-generated text used as SFT targets | No | Works with API teachers; most synthetic-data fine-tuning is this |
| Logit / token-level (white-box) | KL divergence between teacher and student next-token distributions | Yes (same tokenizer, or alignment) | Richer signal; common in open model families |
| On-policy distillation (GKD) | Student generates, teacher scores each token | Yes | Fixes the train/inference mismatch of learning only on teacher text |
| Reasoning distillation | Teacher's chain-of-thought traces as targets | No | Transfers long reasoning into small models effectively |
| Feature / hidden-state | Match intermediate representations | Yes | Used for encoders (for example compact BERT variants) |
import torch.nn.functional as F
def kd_loss(student_logits, teacher_logits, labels, mask, T=2.0, lam=0.5):
# logits: [B, S, V]; labels: [B, S]; mask: [B, S] (1 on answer tokens)
ce = F.cross_entropy(student_logits.transpose(1, 2), labels, reduction="none")
kl = F.kl_div(F.log_softmax(student_logits / T, -1),
F.log_softmax(teacher_logits / T, -1),
log_target=True, reduction="none").sum(-1) * (T * T)
per_tok = (1 - lam) * ce + lam * kl
return (per_tok * mask).sum() / mask.sum().clamp(min=1)
Serving fine-tuned models
After training you have a base model plus an adapter (or a full checkpoint). Deployment choices trade off latency, memory, the number of variants you need, and target hardware. The main decisions are: merge or keep adapters separate, how to serve many variants, and how to quantize for cost or edge deployment.
A camera shop can sell each customer a camera with a lens permanently glued on, or sell one camera body with a bag of interchangeable lenses. Glued lenses are simplest to use; interchangeable lenses let one body serve a portrait photographer in the morning and a wildlife photographer in the afternoon, at the cost of a moment to swap.
Mapping back: the glued lens is a merged model (no overhead, one task per copy), the camera body is the shared base model in GPU memory, the lenses are LoRA adapters of a few megabytes each, and the swap time is the small per-request cost of applying a different adapter.
Merged weights
- Zero added latency; standard architecture
- Works with any inference engine and any quantization tool
- One full model copy per fine-tune (e.g. 14 GB each at 7B bf16)
- Best for one or a few variants at high traffic
Separate adapters (multi-LoRA)
- One base in memory, hundreds of adapters of tens of MB
- Small compute overhead for the extra low-rank matmuls
- Per-request adapter selection, per-tenant customization
- Best for many low-traffic variants
Multi-LoRA serving
Modern inference engines batch requests for different adapters together. The base matmul is shared across the batch, and specialized kernels (segmented or gathered batched matmuls) apply each request's own A and B. Adapters live in a CPU/GPU cache and are paged in on demand. Engines such as vLLM (with LoRA enabled at startup), dedicated multi-LoRA servers, and research systems like S-LoRA and Punica follow this design; managed inference products expose the same idea. Constraints to know: a maximum rank per server, a maximum number of adapters resident on GPU at once, and all adapters must share the same base.
# serve a base model with adapters that clients select by name (vLLM CLI)
vllm serve your-org/base-8b-instruct \
--enable-lora \
--lora-modules support=./adapters/support sql=./adapters/sql \
--max-lora-rank 64 --max-loras 8
# request the "sql" adapter through the OpenAI-compatible endpoint
curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" \
-d '{"model": "sql", "messages": [{"role": "user", "content": "Top 5 customers by revenue?"}]}'
Quantizing after fine-tuning
| Format / method | Typical bits | Where it runs | Notes |
|---|---|---|---|
| bf16/fp16 | 16 | GPU servers | Reference quality |
| FP8 | 8 | Recent data-centre GPUs | Near-lossless, about 2× memory saving |
| GPTQ | 4 (or 3/8) | GPU | Post-training quantization with second-order error correction using a calibration set |
| AWQ | 4 | GPU | Protects salient weight channels based on activation statistics |
| bitsandbytes NF4 | 4 | GPU | Easy, used in QLoRA; slower inference than GPTQ/AWQ kernels |
| GGUF (k-quants such as Q4_K_M, Q5_K_M, Q8_0) | 2–8 | CPU, Apple silicon, consumer GPUs, edge | Single-file format for llama.cpp-family runtimes and local model runners |
Use a calibration set that resembles your fine-tuning domain (not generic web text) when quantizing a specialised model, and re-run your regression suite: quantization errors hit rare, domain-specific behaviours first.
GGUF export for local and edge deployment
- Merge the adapter into a 16-bit base and save with the tokenizer.
- Convert to a 16-bit GGUF file with the llama.cpp conversion script.
- Quantize to the target type (Q4_K_M is a common balance of size and quality; Q8_0 for near-lossless).
- Run and evaluate with the same chat template; check that the template embedded in the GGUF metadata is correct.
# 1) merged model already saved in out/merged (see the merging section)
# 2) convert Hugging Face checkpoint to GGUF (16-bit)
python llama.cpp/convert_hf_to_gguf.py out/merged --outfile model-f16.gguf --outtype f16
# 3) quantize to 4-bit k-quant
./llama.cpp/build/bin/llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M
# 4) quick interactive test
./llama.cpp/build/bin/llama-cli -m model-Q4_K_M.gguf -cnv
# alternatively keep the adapter separate: convert it with convert_lora_to_gguf.py
# and load it at runtime next to the base GGUF (--lora adapter.gguf)
On-device stacks take this further: a shared base model ships once with the operating system or app, and small task-specific LoRA adapters (tens of MB) are hot-swapped per feature, while quantization-aware training plus LoRA recovers accuracy lost to aggressive 4-bit quantization. See Edge AI for runtimes, NPU constraints and memory budgets.
End-to-end walkthrough: full fine-tuning vs LoRA vs QLoRA
This worked example puts everything together on a realistic small-scale study: adapt a 0.5B-parameter instruct model to answer medical instructions, training the same model on the same data three ways (full instruction fine-tuning, LoRA and QLoRA, the last two at ranks 8 and 16) on a single 16 GB T4-class GPU, then compare quality against cost. It uses a hand-written PyTorch loop rather than a high-level trainer so every mechanism is visible. Swap in a 7B–8B model and more data and the same code is a production-shaped QLoRA pipeline.
It is a controlled cooking experiment: the same ingredients (data), the same oven (GPU) and the same recipe card (hyperparameters), with only one variable changed at a time, first how much of the dish you are allowed to modify, then whether the base stock is stored concentrated. Only then can you say which change caused a difference in taste or cost.
Mapping back: the ingredients are the fixed train/validation/test split, the variable being changed is the trainable-parameter scope (full, LoRA) and then the backbone precision (16/32-bit vs 4-bit), and "taste vs cost" is ROUGE-L/BERTScore against trainable %, peak memory and training time.
data (instruction, input, output)
│ clean ("<noinput>" → ""), shuffle(seed), split 300 / 30 / 30
▼
chat template ──▶ tokenize once with offsets ──▶ completion mask (answer = 1)
▼
┌──────────────┬──────────────────────┬─────────────────────────────┐
│ Full IFT │ LoRA r=8, r=16 │ QLoRA r=8, r=16 │
│ fp32, lr 1e-5│ fp32 base, lr 2e-4 │ NF4 base, fp16 compute, │
│ 100% trained │ adapters on 7 linears│ same adapters, lr 2e-4 │
└──────┬───────┴──────────┬───────────┴──────────────┬──────────────┘
└────────── shared loop: masked loss, accumulation, warmup+cosine,
clipping, per-epoch validation, keep best checkpoint
▼
greedy generation on test ──▶ ROUGE-L + BERTScore ──▶ error analysis
▼
cost table: % trainable, peak GPU memory, training time
Step 1: environment, seed and configuration
# pip install -q transformers datasets peft "bitsandbytes>=0.43" evaluate rouge_score bert_score
import gc, math, random, time
import numpy as np, pandas as pd, torch
import torch.nn.functional as F
from torch.utils.data import Dataset, DataLoader
from datasets import load_dataset
from transformers import (AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig,
get_cosine_schedule_with_warmup, set_seed)
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training, TaskType
SEED = 42
def seed_everything(seed=SEED):
random.seed(seed); np.random.seed(seed)
torch.manual_seed(seed); torch.cuda.manual_seed_all(seed); set_seed(seed)
seed_everything()
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
MODEL_ID = "Qwen/Qwen2.5-0.5B-Instruct" # any small instruct model works
N_TRAIN, N_VAL, N_TEST = 300, 30, 30 # tiny on purpose: fits one session
MAX_LEN = 256 # chosen from the token-length p95
BATCH_SIZE, GRAD_ACCUM, EPOCHS = 4, 1, 3 # effective batch = 4
WARMUP_RATIO, MAX_GRAD_NORM = 0.05, 1.0
LR_FULL, LR_LORA = 1e-5, 2e-4 # full FT needs a much smaller LR
LORA_DROPOUT = 0.05
TARGETS = ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]
Step 2: load, clean, split and inspect the data
The dataset has three fields: instruction, an optional input with extra context (stored as a literal <noinput> placeholder when absent) and output, a reference answer that was itself generated by an older LLM. That last point matters for evaluation: the references are a noisy proxy for truth, so relative comparisons between methods are more meaningful than absolute scores.
df = pd.DataFrame(load_dataset("your-org/medical-instructions")["train"])
df["input"] = df["input"].replace("<noinput>", "").fillna("")
df = df.sample(frac=1, random_state=SEED).reset_index(drop=True) # shuffle once
train_df = df.iloc[:N_TRAIN]
val_df = df.iloc[N_TRAIN:N_TRAIN + N_VAL]
test_df = df.iloc[N_TRAIN + N_VAL:N_TRAIN + N_VAL + N_TEST]
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_token
lens = train_df.apply(lambda r: len(tokenizer(r["instruction"] + " " + r["input"] + " " + r["output"],
add_special_tokens=False)["input_ids"]), axis=1)
print("rows with input:", (train_df["input"] != "").mean().round(2))
print(lens.describe(percentiles=[0.5, 0.95])) # choose MAX_LEN near the p95
MAX_LEN to the p95 of content length plus template overhead, or you will truncate the longest answers, which are exactly the ones that teach the model to finish properly.Step 3: chat-template prompts and an offset-based completion mask
SYSTEM_MSG = "You are a helpful medical assistant. Answer accurately and safely."
def build_prompt(instruction, input_text):
user = f"{instruction}\n{input_text}" if input_text else instruction
msgs = [{"role": "system", "content": SYSTEM_MSG},
{"role": "user", "content": user}]
# add_generation_prompt=True appends the assistant header so the model answers next
return tokenizer.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
class SFTDataset(Dataset):
def __init__(self, frame):
self.rows = [{"prompt": build_prompt(r["instruction"], r["input"]),
"answer": r["output"].strip()} for _, r in frame.iterrows()]
def __len__(self): return len(self.rows)
def __getitem__(self, i): return self.rows[i]
def tokenize_example(prompt, answer):
"""Tokenize prompt+answer ONCE; mark answer tokens using character offsets."""
full = prompt + answer + tokenizer.eos_token # end-of-turn so the model learns to stop
enc = tokenizer(full, return_offsets_mapping=True, add_special_tokens=False,
truncation=True, max_length=MAX_LEN)
start = len(prompt) # first answer character
comp_mask = [1 if s >= start else 0 for s, e in enc["offset_mapping"]]
return enc["input_ids"], comp_mask
def collate(batch):
ids, masks = zip(*(tokenize_example(b["prompt"], b["answer"]) for b in batch))
L = max(len(x) for x in ids)
input_ids = torch.full((len(ids), L), tokenizer.pad_token_id, dtype=torch.long)
attn = torch.zeros((len(ids), L), dtype=torch.long)
comp_mask = torch.zeros((len(ids), L), dtype=torch.long)
for i, (x, m) in enumerate(zip(ids, masks)):
input_ids[i, :len(x)] = torch.tensor(x)
attn[i, :len(x)] = 1 # from length, NOT from pad id
comp_mask[i, :len(m)] = torch.tensor(m)
return input_ids, attn, comp_mask
train_loader = DataLoader(SFTDataset(train_df), batch_size=BATCH_SIZE, shuffle=True, collate_fn=collate)
val_loader = DataLoader(SFTDataset(val_df), batch_size=BATCH_SIZE, shuffle=False, collate_fn=collate)
x, a, m = next(iter(train_loader))
print(x.shape, "supervised tokens:", int(m.sum()), "masked:", int((a - m).sum()))
Three design decisions are worth defending in an interview. The chat template reproduces the special tokens the instruct model was trained with. The completion mask restricts the loss to answer tokens. Tokenizing once with offsets avoids boundary drift: tokenizing the prompt separately and using its length can mis-align by a token when the tokenizer would merge characters across the boundary. Padding fills input_ids with the pad token but the masks with 0, so padded positions are neither attended to nor scored; and because the pad token equals EOS here, the attention mask is built from lengths so the real EOS is still attended and supervised.
Step 4: the shared training loop
def batch_loss(model, input_ids, attn, comp_mask):
logits = model(input_ids=input_ids, attention_mask=attn).logits
logits, labels = logits[:, :-1].float(), input_ids[:, 1:]
mask = comp_mask[:, 1:].float()
tok = F.cross_entropy(logits.reshape(-1, logits.size(-1)), labels.reshape(-1), reduction="none")
return (tok * mask.reshape(-1)).sum() / mask.sum().clamp(min=1.0)
@torch.no_grad()
def eval_loss(model, loader):
model.eval(); tot, n = 0.0, 0.0
for x, a, m in loader:
x, a, m = x.to(DEVICE), a.to(DEVICE), m.to(DEVICE)
k = m[:, 1:].sum().item()
tot += batch_loss(model, x, a, m).item() * k; n += k
model.train()
return tot / max(n, 1)
def gpu_mem_gb():
return torch.cuda.max_memory_allocated() / 1e9 if torch.cuda.is_available() else 0.0
def free_memory():
gc.collect(); torch.cuda.empty_cache(); torch.cuda.reset_peak_memory_stats()
def train_model(model, lr, tag, patience=1):
seed_everything(); model.to(DEVICE)
params = [p for p in model.parameters() if p.requires_grad]
opt = torch.optim.AdamW(params, lr=lr)
total = math.ceil(len(train_loader) / GRAD_ACCUM) * EPOCHS
sched = get_cosine_schedule_with_warmup(opt, int(WARMUP_RATIO * total), total)
hist, best_val, best_state, bad = {"train": [], "val": []}, float("inf"), None, 0
t0 = time.time(); model.train(); opt.zero_grad()
for ep in range(1, EPOCHS + 1):
run = 0.0
for i, (x, a, m) in enumerate(train_loader):
x, a, m = x.to(DEVICE), a.to(DEVICE), m.to(DEVICE)
loss = batch_loss(model, x, a, m) / GRAD_ACCUM
loss.backward()
if (i + 1) % GRAD_ACCUM == 0 or i + 1 == len(train_loader):
torch.nn.utils.clip_grad_norm_(params, MAX_GRAD_NORM)
opt.step(); sched.step(); opt.zero_grad()
run += loss.item() * GRAD_ACCUM
tr, va = run / len(train_loader), eval_loss(model, val_loader)
hist["train"].append(tr); hist["val"].append(va)
print(f"[{tag}] epoch {ep}: train {tr:.4f} val {va:.4f}")
if va < best_val:
best_val, bad = va, 0
best_state = {n: p.detach().cpu().clone() # trainable tensors only
for n, p in model.named_parameters() if p.requires_grad}
else:
bad += 1
if bad > patience: break # early stopping
model.load_state_dict(best_state, strict=False) # restore best checkpoint
return model, hist, time.time() - t0, gpu_mem_gb()
def count_trainable(model):
t = sum(p.numel() for p in model.parameters() if p.requires_grad)
n = sum(p.numel() for p in model.parameters())
return t, n, 100 * t / n
Every standard mechanism appears here: gradient accumulation (divide loss, step every GRAD_ACCUM micro-batches), linear warmup then cosine decay, gradient-norm clipping at 1.0, per-epoch validation, keeping and restoring the best checkpoint, and early stopping. Storing only trainable tensors keeps checkpoints tiny for LoRA; for full fine-tuning it is the whole model.
Step 5: run 1, full instruction fine-tuning
free_memory()
full = AutoModelForCausalLM.from_pretrained(MODEL_ID, torch_dtype=torch.float32)
print("trainable %:", round(count_trainable(full)[2], 2)) # 100.0
full, full_hist, full_time, full_mem = train_model(full, LR_FULL, "full-IFT")
The model is loaded in fp32 because pure fp16 full fine-tuning overflows on this model family and the T4 has no fast bf16; the tiny learning rate (1e-5) prevents the update from overwriting pretrained features. Memory: 0.49B parameters × 16 bytes ≈ 7.9 GB for weights, gradients and Adam states, plus activations, which makes full fine-tuning the most memory-hungry run even at this small scale.
Step 6: runs 2–3, LoRA at r = 8 and r = 16
def lora_cfg(r):
return LoraConfig(task_type=TaskType.CAUSAL_LM, r=r, lora_alpha=2 * r, # alpha/r fixed at 2
lora_dropout=LORA_DROPOUT, bias="none", target_modules=TARGETS)
lora_runs = {}
for r in (8, 16):
free_memory(); seed_everything()
base = AutoModelForCausalLM.from_pretrained(MODEL_ID, torch_dtype=torch.float32)
model = get_peft_model(base, lora_cfg(r))
pct = count_trainable(model)[2] # about 0.9% at r=8, 1.75% at r=16
model, hist, t, mem = train_model(model, LR_LORA, f"LoRA r={r}")
lora_runs[r] = dict(model=model, hist=hist, time=t, mem=mem, pct=pct)
Keeping α/r fixed isolates the effect of rank. The parameter percentage is higher than the "well under 1%" often quoted for 7B models because small models have narrow layers: for this architecture (hidden size 896, MLP width 4864, 24 layers, grouped-query attention with narrow key/value projections) rank 16 on all seven linear layers adds about 8.8M parameters to about 494M.
Step 7: runs 4–5, QLoRA at r = 8 and r = 16
bnb_cfg = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.float16) # T4: fp16 compute
qlora_runs = {}
for r in (8, 16):
free_memory(); seed_everything()
base = AutoModelForCausalLM.from_pretrained(MODEL_ID, quantization_config=bnb_cfg,
torch_dtype=torch.float16)
base = prepare_model_for_kbit_training(base) # fp32 norms, checkpointing, input grads
model = get_peft_model(base, lora_cfg(r)) # identical adapters to the LoRA runs
pct = count_trainable(model)[2]
model, hist, t, mem = train_model(model, LR_LORA, f"QLoRA r={r}")
qlora_runs[r] = dict(model=model, hist=hist, time=t, mem=mem, pct=pct)
The only change from the LoRA runs is the backbone's precision, which makes this a clean controlled comparison. Trainable parameters match LoRA at the same rank exactly. Two small-model subtleties: first, the embedding matrix (vocabulary of about 152k × 896 ≈ 136M parameters) is not quantized, so on a 0.5B model the 4-bit saving is proportionally smaller than on 7B; second, gradient checkpointing (enabled by prepare_model_for_kbit_training) cuts activation memory but adds recomputation time, so the QLoRA runs are typically the slowest per epoch while using the least memory.
Step 8: generation-based evaluation
import evaluate
rouge, bertscore = evaluate.load("rouge"), evaluate.load("bertscore")
@torch.no_grad()
def generate_answers(model, frame, max_new_tokens=200):
model.eval(); preds, refs = [], []
for _, row in frame.iterrows():
enc = tokenizer(build_prompt(row["instruction"], row["input"]), return_tensors="pt",
add_special_tokens=False).to(DEVICE)
out = model.generate(**enc, max_new_tokens=max_new_tokens, do_sample=False,
pad_token_id=tokenizer.pad_token_id)
preds.append(tokenizer.decode(out[0, enc["input_ids"].shape[1]:], # new tokens only
skip_special_tokens=True).strip())
refs.append(row["output"].strip())
model.train()
return preds, refs
def score(preds, refs):
rl = rouge.compute(predictions=preds, references=refs, use_stemmer=True)["rougeL"]
bs = np.mean(bertscore.compute(predictions=preds, references=refs, lang="en")["f1"])
return {"rougeL": round(float(rl), 4), "bertscore_f1": round(float(bs), 4)}
results = {}
runs = {"full_ift": dict(model=full, time=full_time, mem=full_mem, pct=100.0),
**{f"lora_r{r}": v for r, v in lora_runs.items()},
**{f"qlora_r{r}": v for r, v in qlora_runs.items()}}
for name, run in runs.items():
p, refs = generate_answers(run["model"], test_df)
results[name] = {**score(p, refs), "pct_trainable": round(run["pct"], 2),
"peak_mem_gb": round(run["mem"], 2), "train_time_s": round(run["time"], 1)}
print(pd.DataFrame(results).T)
Step 9: error analysis
best = max(results, key=lambda k: results[k]["rougeL"])
preds, refs = generate_answers(runs[best]["model"], test_df)
pl = np.array([len(tokenizer(p, add_special_tokens=False)["input_ids"]) for p in preds])
rl = np.array([len(tokenizer(r, add_special_tokens=False)["input_ids"]) for r in refs])
ratio = pl / np.maximum(rl, 1)
print("median length ratio:", np.median(ratio).round(2),
"| empty:", int((pl == 0).sum()),
"| much shorter (<0.5):", int((ratio < 0.5).sum()),
"| much longer (>2):", int((ratio > 2).sum()))
for p, r in list(zip(preds, refs))[:3]:
print("REF:", r[:300], "\nGEN:", p[:300], "\n" + "-" * 60)
What results to expect and how to read them
| Run | Trainable % | Relative peak memory | Relative time | Typical quality |
|---|---|---|---|---|
| Full IFT (fp32) | 100 | Highest | Medium | Similar to adapters; small data limits the gain from extra capacity |
| LoRA r=8 | ≈0.9 | Much lower | Fastest | Close to full FT |
| LoRA r=16 | ≈1.75 | Much lower | Fast | Often indistinguishable from r=8 on 300 examples |
| QLoRA r=8 | ≈0.9 | Lowest | Slowest | Within noise of LoRA |
| QLoRA r=16 | ≈1.75 | Lowest | Slowest | Within noise of LoRA |
With only 300 examples and 3 epochs, differences between methods are usually small and can flip with the random seed, which is itself the key lesson: on a free GPU, LoRA or QLoRA gets essentially the same answer quality as full fine-tuning for a fraction of memory. Other typical observations: validation loss bottoms out after one or two epochs for full fine-tuning; lower training loss does not reliably predict better generated answers; the dominant failure modes are answers that are shorter or less specific than the references, missing safety caveats, and truncation when max_new_tokens is small. Reasonable next experiments are more data, a longer MAX_LEN, a larger model with QLoRA, or a light repetition penalty at decoding.
Step 10: save, merge and export
best_run = runs["qlora_r16"]["model"]
best_run.save_pretrained("out/qlora-r16-adapter") # adapter only: a few tens of MB
tokenizer.save_pretrained("out/qlora-r16-adapter")
# merge into a 16-bit base (never into the 4-bit weights), then export or quantize
from peft import PeftModel
base16 = AutoModelForCausalLM.from_pretrained(MODEL_ID, torch_dtype=torch.float16)
merged = PeftModel.from_pretrained(base16, "out/qlora-r16-adapter").merge_and_unload()
merged.save_pretrained("out/merged-fp16"); tokenizer.save_pretrained("out/merged-fp16")
# next: convert to GGUF and quantize for local/edge use (see the serving section)
prepare_model_for_kbit_training does, why QLoRA has the same trainable-parameter count as LoRA, whether r=16 beat r=8, and why lower training loss did not mean better test answers. Have a one-sentence answer ready for each, backed by your own numbers.input_ids != pad_token_id when pad equals EOS (the stop token is hidden), and saving the "best" state but never loading it back before evaluation. Also never treat a 0.5B model fine-tuned on machine-generated medical answers as clinically reliable; it is a methods exercise.Quick revision
- Fine-tuning continues gradient training of a pretrained model on your data; inference and prompting never change weights.
- Rule of thumb: new knowledge calls for RAG, new behaviour calls for fine-tuning, and many systems use both.
- Always baseline with the best prompt you can write and build an evaluation set before fine-tuning.
- Continued pretraining uses raw domain text; SFT uses prompt–response pairs; preference tuning uses chosen/rejected pairs or rewards.
- SFT is next-token cross-entropy with teacher forcing, usually on response tokens only (prompt labels set to -100).
- Logits at position t predict token t+1, so shift logits and labels (and the mask) by one.
- Use the model's own chat template via
apply_chat_template; mismatched templates degrade quality. - Append the end-of-turn/EOS token to every target, or the model never learns to stop.
- Tokenize prompt + answer once and use offsets to locate the answer boundary.
- Quality beats quantity: a thousand clean, diverse examples can beat a hundred thousand noisy ones.
- Deduplicate, scrub personal data, decontaminate against test sets, and split by group to avoid leakage.
- Set maximum sequence length near the 95th percentile of tokenized length plus template overhead.
- Packing concatenates short examples to cut padding waste; block cross-example attention when possible.
- Full fine-tuning with mixed-precision AdamW needs about 16 bytes per parameter before activations: 112 GB for 7B.
- Activations scale with batch × sequence × hidden × layers; gradient checkpointing trades roughly a third more compute for large savings.
- Large vocabularies make the logits tensor a hidden memory hog; chunked or fused cross-entropy helps.
- PEFT freezes the base and trains a small set of parameters: adapters, prefix/prompt tuning, IA3, BitFit, LoRA.
- LoRA:
h = W0x + (α/r)BAx, with A random and B zero so training starts from the pretrained model. - LoRA adds r·(din + dout) parameters per adapted matrix; rank 16 on all linear layers of a 7B model is about 40M parameters.
- Targeting all linear layers usually matters more than increasing rank.
- Keep α/r fixed when sweeping rank; rsLoRA uses α/√r for stability at high rank.
- Merged LoRA adds zero inference latency; unmerged adapters add a small overhead but can be swapped.
- DoRA separates magnitude and direction, training LoRA on the direction, and often helps at low rank.
- QLoRA = NF4 4-bit frozen base + double quantization + paged optimizers + 16-bit LoRA adapters; compute is in bf16/fp16.
- NF4 places 16 levels at normal-distribution quantiles with blockwise (64) absmax scaling.
- Double quantization cuts scale-constant overhead from 0.5 to about 0.127 bits per parameter.
prepare_model_for_kbit_trainingupcasts norms, enables checkpointing and input gradients for k-bit training.- Typical LRs: full FT about 1e-5, LoRA about 2e-4, DPO about 5e-7 to 5e-6 (full) with β about 0.1.
- 1–3 epochs, warmup 3–10%, cosine or linear decay, clipping at 1.0, and keep the best validation checkpoint.
- RLHF = SFT, then a Bradley–Terry reward model from human comparisons, then PPO maximizing reward minus β·KL to the reference.
- PPO-based RLHF keeps four models in memory: policy, reference, reward and value. The clip applies to ρ then multiplies by A; the objective is maximized.
- DPO loss: −log σ(β [log(πθ(yw)/πref(yw)) − log(πθ(yl)/πref(yl))]); no reward model or sampling; smaller β allows more deviation from the reference.
- ORPO and SimPO are reference-free; KTO uses unpaired thumbs-up/down data; IPO regularizes DPO's margins.
- GRPO keeps PPO's clip but replaces the critic with group-standardized rewards over G samples; two models in memory instead of four; stalls when a group has zero reward variance.
- RLAIF and Constitutional AI replace human preference labels with AI judgements guided by written principles.
- Reward hacking, sycophancy and verbosity are classic preference-tuning failure modes; the KL penalty limits them.
- Evaluate on held-out task data, with a validated LLM judge, on general benchmarks, safety suites and a regression suite, all against the base model.
- Catastrophic forgetting is mitigated by PEFT, data replay, lower LR, fewer epochs, regularization and model merging.
- Even benign fine-tuning can weaken safety refusals, so re-run safety evaluations on every release.
- Merge adapters into a 16-bit base, never directly into 4-bit weights, then re-quantize if needed.
- Model merging (linear, SLERP, task arithmetic, TIES, DARE) only works for models sharing a base.
- Distillation trains a smaller student on teacher outputs or soft logits with temperature; check the teacher's terms of use.
- Serve one merged model for high-traffic single tasks, or multi-LoRA over a shared base for many variants.
- For local/edge deployment: merge, convert to GGUF, quantize (for example Q4_K_M), and re-evaluate with the right template.
Glossary
- Adapter
- A small trainable module inserted into a frozen network; in the narrow sense, a bottleneck MLP added after attention or feed-forward blocks.
- AdamW
- Adam optimizer with decoupled weight decay; stores two moment estimates per trainable parameter.
- Alignment
- Post-training that shapes a model's behaviour towards human intentions and values, such as helpfulness, honesty and harmlessness.
- Alpha (LoRA)
- Scaling hyperparameter; the LoRA update is multiplied by α/r.
- Base model
- A pretrained model before instruction tuning or alignment; it continues text rather than following instructions.
- bitsandbytes
- Library of 8-bit and 4-bit quantized layers and memory-efficient optimizers used by QLoRA.
- Bradley–Terry model
- Probability model for pairwise preferences: P(A preferred over B) = σ(rA − rB).
- Catastrophic forgetting
- Loss of previously learned abilities when a network is trained on new data.
- Chat template
- The exact special-token layout a chat model uses to mark system, user, assistant and tool turns.
- Completion-only loss
- Computing the training loss only on response tokens, masking prompt tokens.
- Constitutional AI
- Alignment approach in which a model critiques and revises outputs, and an AI judge ranks them, according to a written set of principles.
- Continued pretraining
- Further next-token training on large unlabelled domain text; also called domain-adaptive pretraining.
- DARE
- Model-merging pre-processing that randomly drops most delta parameters and rescales the remainder.
- DeepSpeed ZeRO
- Memory optimization that shards optimizer states (stage 1), gradients (stage 2) and parameters (stage 3) across GPUs.
- Distillation
- Training a smaller student model to imitate a larger teacher's outputs or probability distributions.
- DoRA
- Weight-decomposed LoRA: learns a magnitude vector and applies LoRA to the normalized direction of each weight.
- Double quantization
- Quantizing the per-block quantization constants themselves to save memory.
- DPO
- Direct Preference Optimization: aligns a model with preference pairs using a closed-form loss on policy-to-reference log-probability ratios.
- Early stopping
- Halting training when validation performance stops improving, keeping the best checkpoint.
- Effective batch size
- Micro-batch size × gradient-accumulation steps × number of data-parallel GPUs.
- FSDP
- Fully Sharded Data Parallel: PyTorch's native sharding of parameters, gradients and optimizer states.
- GGUF
- Single-file model format with metadata and many quantization types, used by llama.cpp-family runtimes.
- Gradient accumulation
- Summing gradients over several micro-batches before one optimizer step to simulate a larger batch.
- Gradient checkpointing
- Saving only some activations and recomputing the rest during backward to save memory.
- GRPO
- Group Relative Policy Optimization: PPO-style RL without a value model, using rewards standardized within a group of samples per prompt.
- IA3
- PEFT method that learns vectors rescaling keys, values and feed-forward activations.
- Instruction tuning
- Supervised fine-tuning on diverse instruction–response pairs so a model follows natural-language commands.
- KL penalty
- Term penalizing divergence of the trained policy from a reference model during preference optimization.
- KTO
- Preference method learning from unpaired examples labelled desirable or undesirable.
- LLM-as-judge
- Using a strong language model with a rubric to rate or compare outputs.
- LoRA
- Low-Rank Adaptation: trains a low-rank update BA added to frozen weight matrices.
- Loss masking
- Excluding tokens from the loss, typically by setting their labels to -100.
- Merge and unload
- Folding a LoRA update into the base weights and removing the adapter modules.
- Mixed precision
- Computing in 16-bit while keeping an fp32 master copy of weights for the optimizer.
- Model merging
- Combining weights of several fine-tuned models that share a base, without further training.
- Multi-LoRA serving
- Serving many LoRA adapters over one shared base model, batching requests for different adapters together.
- NF4
- 4-bit NormalFloat: a 16-level code at normal-distribution quantiles, used for storing QLoRA base weights.
- ORPO
- Odds Ratio Preference Optimization: single-stage SFT plus odds-ratio preference loss, with no reference model.
- Packing
- Concatenating multiple short examples into one training sequence to reduce padding.
- Paged optimizer
- Optimizer whose states live in unified memory and page to CPU RAM during GPU memory spikes.
- PEFT
- Parameter-Efficient Fine-Tuning: methods that train a small fraction of parameters; also the name of the Hugging Face library.
- Perplexity
- Exponential of the average per-token cross-entropy; lower means the model finds the text less surprising.
- PPO
- Proximal Policy Optimization: RL algorithm using a clipped probability-ratio objective and a value model.
- Prefix tuning
- PEFT method that learns key/value vectors prepended at every attention layer.
- Prompt tuning
- PEFT method that learns a few soft embedding vectors prepended to the input.
- QLoRA
- LoRA fine-tuning on a base model stored in 4-bit NF4, with double quantization and paged optimizers.
- Rank (LoRA)
- Inner dimension r of the LoRA matrices; bounds the rank of the update and sets adapter size.
- Reference model
- Frozen copy of the starting policy used for the KL penalty or DPO log-ratios.
- Replay
- Mixing general or earlier-task data into fine-tuning to reduce forgetting.
- Reward hacking
- A policy exploiting flaws in a reward signal to score highly without truly improving.
- Reward model
- Model trained on human (or AI) comparisons to output a scalar score for a prompt–response pair.
- RLAIF
- Reinforcement Learning from AI Feedback: preference labels come from an AI judge instead of humans.
- RLHF
- Reinforcement Learning from Human Feedback: SFT, reward model from human preferences, then RL against that reward.
- RLVR
- RL with verifiable rewards: rewards from automatic checks such as correct final answers or passing tests.
- rsLoRA
- Rank-stabilized LoRA using scale α/√r.
- SFT
- Supervised Fine-Tuning: training on input–target demonstrations with next-token cross-entropy.
- SFTTrainer
- TRL's trainer for supervised fine-tuning, with chat formatting, packing and PEFT integration.
- SimPO
- Reference-free preference method using length-normalized log-probability as the implicit reward plus a margin.
- Sycophancy
- Tendency of a model to agree with or flatter the user rather than be accurate.
- Synthetic data
- Training examples generated or rewritten by a model rather than written by people.
- Target modules
- The layers (for example q_proj, v_proj, gate_proj) that receive LoRA adapters.
- Task vector
- Difference between fine-tuned and base weights, used for task arithmetic.
- Teacher forcing
- Training by feeding the true previous tokens rather than the model's own predictions.
- TIES merging
- Merging method that trims small deltas, elects a sign per parameter and averages agreeing deltas.
- TRL
- Hugging Face library of trainers for SFT, reward modelling, DPO, PPO, GRPO and related methods.
- Warmup
- Gradually increasing the learning rate from near zero over the first steps of training.
Interview questions
Fundamentals
What is fine-tuning an LLM?
Continuing to train a pretrained model with gradient descent on a smaller, task- or domain-specific dataset so that its weights (all of them, or a small added set) change to produce the desired behaviour. The result persists in the weights, unlike prompting, which only changes the input.
How is fine-tuning different from pretraining?
Pretraining is self-supervised next-token prediction on trillions of tokens of raw text, costs months and millions of dollars, and produces general knowledge. Fine-tuning starts from those weights, uses thousands to millions of curated examples, runs for minutes to days, and targets specific behaviour. The loss can be identical (next-token cross-entropy); data, scale, masking and learning rate differ.
What is the difference between a base model and an instruct (chat) model?
A base model is the output of pretraining: it continues text and may answer a question with more questions. An instruct model has additionally gone through SFT on instruction–response data and usually preference tuning, so it follows instructions, uses a chat template and refuses harmful requests.
When should you fine-tune instead of prompting or using RAG?
Fine-tune when the problem is behaviour: consistent format, tone, a specialised skill such as classification or tool calling, or when you need a small cheap model to replace a big one with a long prompt. Use RAG when the problem is knowledge, especially fresh, large or citable facts. Try prompting first because it is cheapest, and often combine fine-tuning (style) with RAG (facts).
What is instruction tuning?
Supervised fine-tuning on many diverse instruction–response pairs (instruction, optional input, output) so that the model learns to follow natural-language commands in general, including unseen ones. It is what turns a base model into an assistant.
Is instruction tuning the same as prompt engineering?
No. Instruction tuning happens at training time and changes weights. Prompt engineering happens at inference time and only changes the input text; the model is unchanged once the prompt is gone.
What is SFT?
Supervised Fine-Tuning: training on input–target demonstrations using next-token cross-entropy with teacher forcing, typically computing the loss only on the target (response) tokens.
What is RLHF in one paragraph?
Reinforcement Learning from Human Feedback is a post-training pipeline: first SFT on demonstrations; then humans compare pairs of model responses and a reward model is trained to predict their preferences; finally the model is optimized with RL (classically PPO) to maximize the reward while a KL penalty keeps it close to the SFT model. It aligns the model with what people prefer rather than just what demonstrators wrote.
What is PEFT and why is it useful?
Parameter-Efficient Fine-Tuning freezes the pretrained weights and trains a small number of extra or selected parameters (often under 1–2%). It cuts GPU memory (no gradients or optimizer states for frozen weights), produces tiny adapter files, lets many tasks share one base model, and reduces catastrophic forgetting, while matching full fine-tuning on most SFT tasks.
What is LoRA?
Low-Rank Adaptation keeps a weight W0 frozen and learns an update ΔW = (α/r)·BA, where A is r × din and B is dout × r with small rank r (such as 8–64). Only A and B are trained; after training they can be merged into W0 with no latency cost.
What is QLoRA?
LoRA on top of a base model stored in 4-bit NormalFloat (NF4) with double quantization, plus paged optimizers to absorb memory spikes. The frozen weights are dequantized to bf16/fp16 on the fly for each matmul; only the 16-bit adapters are trained. It lets a 7B–8B model be fine-tuned on a single 12–16 GB GPU.
What does the LoRA rank r control?
The inner dimension of A and B, which bounds the rank of the learned update and sets adapter size (r·(din+dout) parameters per matrix). Higher r means more capacity and memory; for most SFT tasks 8–32 is enough.
What does lora_alpha do?
It scales the update by α/r. It acts like a learning-rate multiplier for the adapter path and lets you change r without retuning the learning rate much. Common settings are α = r or α = 2r.
What are target modules in LoRA?
The layers that receive adapters, named by module, for example q_proj, k_proj, v_proj, o_proj in attention and gate_proj, up_proj, down_proj in the MLP. Adapting all linear layers generally works better than adapting only query and value projections.
What is a chat template and why does it matter for fine-tuning?
It is the exact special-token format a chat model uses to delimit system, user and assistant turns. Fine-tuning an instruct model with a different format wastes capacity re-learning structure and can break its chat behaviour; serving with a different template than training causes rambling or garbled output. Use tokenizer.apply_chat_template.
What is catastrophic forgetting?
The loss of previously learned capabilities when a network is trained on new data, because the same weights encode both old and new skills and new gradients overwrite the old solutions. In LLMs it shows up as worse general reasoning, instruction following, other languages or safety after narrow fine-tuning.
How much data do you need to fine-tune?
It depends on the goal: a format or style change can work with a few hundred to two thousand high-quality examples; simple classification with 500–5,000; a domain assistant with 5k–50k; preference tuning with a few thousand pairs or more. Quality, diversity and consistency matter more than raw count.
What is overfitting in fine-tuning and how do you spot it?
The model memorizes training examples instead of learning the general behaviour. Training loss keeps falling while validation loss rises; generations copy training answers verbatim or become repetitive. Fix with early stopping, fewer epochs, more or more diverse data, dropout and lower rank.
Why is the learning rate for LoRA higher than for full fine-tuning?
LoRA's base is frozen and B starts at zero, so the adapter begins as a no-op and its update is confined to a low-rank subspace; larger steps (about 2e-4) are safe. Full fine-tuning changes every pretrained weight directly, so it needs small steps (about 1e-5) to avoid destroying learned features or diverging.
What is an epoch, a step and gradient accumulation?
An epoch is one pass over the training set. A step is one optimizer update. Gradient accumulation sums gradients over several micro-batches before one step, so the effective batch size is micro-batch × accumulation steps × number of GPUs, without the activation memory of a large batch.
What is learning-rate warmup and why use it?
Increasing the learning rate linearly from near zero over the first few percent of steps. Early on, Adam's moment estimates are unreliable and a few large updates can damage the pretrained model; warmup keeps early updates small.
Can you fine-tune closed-source models?
Not in the full sense: you cannot access or update their weights. Some providers offer managed fine-tuning APIs for selected models, where you upload data and receive a private endpoint, with limited control over method and hyperparameters and no weight download. Open-weight models give full control.
What does merging a LoRA adapter mean?
Computing W' = W0 + (α/r)BA for every adapted layer and saving the result as an ordinary model, so no adapter modules remain and there is no inference overhead. In PEFT it is merge_and_unload().
What is DPO?
Direct Preference Optimization trains a model directly on chosen/rejected pairs with a logistic loss on the difference of policy-to-reference log-probability ratios, scaled by β. It achieves the RLHF objective without training a reward model or running RL, so it is simpler, cheaper and more stable.
What is a reward model?
A model (usually an LLM with a scalar output head) trained on preference comparisons to give higher scores to responses humans prefer. It serves as a learned proxy for human judgement during RL.
What is knowledge distillation?
Training a small student model to imitate a larger teacher, either on the teacher's generated text (sequence-level) or by matching its softened next-token distributions (logit-level, with a KL loss and temperature). It yields small models that perform far above their size.
Does fine-tuning add knowledge reliably?
Only weakly. Models absorb new facts slowly through fine-tuning, cannot cite them, and training on facts they do not know can increase hallucinations. Fine-tuning is best for behaviour and skills; retrieval is better for facts, especially changing ones.
Which libraries are commonly used to fine-tune open models?
Hugging Face Transformers (models, tokenizers, Trainer), PEFT (LoRA and other adapters), TRL (SFT, DPO, PPO, GRPO trainers), bitsandbytes (4/8-bit quantization and optimizers), Accelerate with DeepSpeed or FSDP for multi-GPU, and higher-level tools such as Unsloth (optimized single-GPU LoRA/QLoRA) and Axolotl or similar YAML-driven wrappers.
What is the difference between full fine-tuning and LoRA in terms of outputs saved?
Full fine-tuning saves a complete new model (for example 14 GB for 7B in bf16). LoRA saves only the adapter matrices, usually tens to a few hundred megabytes, which must be loaded on top of the exact same base model.
What is continued pretraining and when would you use it?
Further next-token training on large amounts of unlabelled domain text (papers, code, legal documents, a new language). Use it when the domain vocabulary and style are far from the base model's training data; follow it with SFT to restore instruction following.
Going deeper
Why do we compute the SFT loss only on response tokens?
We want to learn p(response | prompt), not to model prompts. Including prompt tokens wastes capacity, lets long prompts dominate the gradient, and can teach the model to produce user-like or system-prompt text. Prompt tokens still serve as context; only their labels are masked (-100). On very short-answer datasets some teams keep a small prompt-loss weight, but completion-only is the default.
Explain the shift-by-one in causal LM loss.
The logits at position t are the model's prediction for token t+1. So the loss compares logits[:, :-1] against input_ids[:, 1:], and the mask must be shifted the same way. Hugging Face models do this internally when you pass labels; custom loops must do it explicitly or they train the model to copy the current token.
Why tokenize prompt and answer together rather than separately?
Subword tokenizers can merge characters across the boundary differently when strings are joined, so separately tokenized pieces do not equal the tokenization of the full text the model sees at inference. Tokenize the concatenation once and use character offsets (return_offsets_mapping) to find which tokens belong to the answer.
Walk through the memory needed to fully fine-tune a 7B model.
With mixed-precision AdamW: 2 bytes bf16 weights + 2 bytes gradients + 4 bytes fp32 master weights + 8 bytes Adam moments = 16 bytes per parameter, so about 112 GB, plus activations (several GB to tens of GB depending on batch, sequence length, fused attention and checkpointing). It needs multiple 80 GB GPUs with ZeRO-3/FSDP, or aggressive tricks (8-bit Adam, offload) on one.
How much memory does QLoRA need for a 7B model?
About 3.5 GB for 4-bit weights plus roughly 0.1 GB of quantization constants with double quantization and a few hundred MB for unquantized embeddings/head; about 40M LoRA parameters (rank 16, all linear) cost around 0.6 GB with Adam states; activations with gradient checkpointing add 1–4 GB depending on sequence length. Total is roughly 6–10 GB, so a 12–16 GB GPU is comfortable.
Calculate LoRA parameters for a 4096 × 4096 projection at rank 8 and 16.
r·(din + dout) = 8 × 8192 = 65,536 at rank 8, and 131,072 at rank 16, compared with 16,777,216 for the full matrix (256× and 128× fewer).
Why is B initialized to zero and A randomly in LoRA?
With B = 0 the product BA is zero, so the adapted model starts exactly equal to the pretrained model, making early training stable. A must be random (non-zero); if both were zero, each matrix's gradient would be zero because it depends on the other, and nothing would learn.
Does LoRA reduce training compute as much as memory?
No. The forward pass still runs through the full model, and the backward pass still propagates activation gradients through every frozen layer; only the weight-gradient computation and optimizer update for frozen weights are skipped. Memory drops dramatically; step time drops modestly (and QLoRA is often slower than LoRA due to dequantization).
Explain NF4 quantization.
NormalFloat4 uses 16 levels placed at quantiles of a standard normal distribution normalized to [−1, 1], so that normally distributed weights use each level about equally (near-optimal for that distribution), with an exact zero. Weights are quantized in blocks of 64, each scaled by its absmax, which limits outlier damage.
What is double quantization and how much does it save?
Blockwise quantization stores one fp32 scale per 64 weights (0.5 bits per parameter). Double quantization quantizes those scales to 8-bit in groups of 256 with one fp32 second-level scale per group, reducing overhead to about 0.127 bits per parameter, saving roughly 0.37 bits per parameter (about 0.3 GB on a 7B model).
What are paged optimizers?
Optimizers whose state tensors are allocated in NVIDIA unified memory, so that when GPU memory spikes (for example on an unusually long batch) pages are automatically evicted to CPU RAM and brought back later. They turn out-of-memory crashes into small slowdowns.
What does prepare_model_for_kbit_training do?
For a quantized model it casts layer norms (and typically the output head) to fp32 for numerical stability, enables gradient checkpointing, and makes input embeddings require gradients (via a hook) so gradients can flow through the frozen, checkpointed network to the LoRA adapters.
Why is compute dtype fp16 rather than bf16 on some GPUs?
Older GPUs such as the T4 (Turing) lack native fast bf16 support, so compute uses fp16; on Ampere and newer, bf16 is preferred because its wider exponent range avoids overflow. Full fine-tuning in pure fp16 can overflow, which is why fp32 (or bf16 mixed precision) is used for it.
Why do QLoRA and LoRA have the same number of trainable parameters at the same rank?
The adapters are identical; QLoRA only changes the storage precision of the frozen backbone. Trainable parameter count depends on rank and target modules, not on how frozen weights are stored.
Compare adapters, prefix tuning, prompt tuning and LoRA.
Bottleneck adapters insert small MLPs into each block (sequential, not mergeable, adds latency). Prefix tuning learns key/value vectors at every layer; prompt tuning learns soft input embeddings only (both consume context length and prompt tuning needs very large models). LoRA learns low-rank weight updates that merge into existing weights, adding no latency, and works across sizes, which is why it dominates.
How do you pick the LoRA rank?
Start with 16 on all linear layers with α = 2r. Sweep 8/16/32/64 with α/r fixed and compare validation task metrics. Increase rank for large shifts (new language, complex skills, lots of data); keep it low for small datasets to reduce overfitting. Gains usually plateau quickly.
What are the three stages of classic RLHF and what models are involved?
SFT (produces the reference policy), reward-model training on pairwise human preferences (Bradley–Terry loss), and PPO optimization of the policy against the reward with a KL penalty. PPO needs the policy, a frozen reference, the reward model and a value model: four models in memory.
Write the reward model loss.
L = −E[log σ(r(x, yw) − r(x, yl))], the Bradley–Terry negative log-likelihood that the chosen response yw beats the rejected yl. Only reward differences matter, so rewards are shift-invariant and are often normalized.
What is the purpose of the KL penalty in RLHF?
It keeps the policy close to the SFT reference, which prevents reward hacking (exploiting reward-model errors in regions it never saw), preserves fluency and general capabilities, and stabilizes training. β controls the trade-off between reward and staying close.
How does DPO differ from PPO-based RLHF?
DPO is offline and RL-free: it uses a fixed preference dataset and a classification-style loss on log-probability ratios, needing only the policy and a frozen reference. PPO is online: it samples from the current policy, scores with a reward model, estimates advantages with a value model, and can explore beyond the dataset but is expensive and unstable.
Write the DPO loss.
LDPO = −E [ log σ( β ( log(πθ(yw|x)/πref(yw|x)) − log(πθ(yl|x)/πref(yl|x)) ) ) ]. Log-probabilities are summed over response tokens. Smaller β allows more deviation from the reference. Equivalent form: −log σ(β (log πθ(yw) − log πθ(yl) − log πref(yw) + log πref(yl))).
What does β mean in DPO?
It is the strength of the implicit KL constraint to the reference. Small β lets the policy move further from the reference (stronger preference fitting, more risk of degradation); large β keeps it close. Typical values are 0.05–0.5, with 0.1 a common default.
What are ORPO and KTO and when would you use them?
ORPO combines SFT and preference learning into one stage with an odds-ratio term and needs no reference model: useful when you want one cheap training run from a base or lightly tuned model. KTO learns from unpaired examples labelled good or bad, suitable for production thumbs-up/down logs where you rarely have paired comparisons.
What is GRPO?
Group Relative Policy Optimization samples a group of responses for each prompt, scores them, and uses each response's reward standardized within its group as the advantage, in a PPO-style clipped objective with a KL penalty. The advantage is sequence-level (one A per response) and is applied to every token; the probability ratio is still per token. It removes the value model, so typical memory is policy + reference rather than PPO's four models, and is well suited to verifiable rewards such as correct math answers or passing unit tests.
Compare GRPO and PPO.
Both maximize min(ρA, clip(ρ, 1−ε, 1+ε)A) with a KL penalty to a reference. PPO estimates A with a learned value model and GAE and, in RLHF, also keeps a reward model: four networks in memory, usually one rollout per prompt. GRPO drops the critic and sets A to the group z-score of G samples (8–16) of the same prompt; the scorer is often a deterministic verifier. Choose PPO when preferences are subjective and you already train a reward model; choose GRPO when answers can be checked automatically. GRPO's distinctive failure is a zero-variance group (all correct or all wrong), which contributes no gradient.
What is RLAIF / Constitutional AI?
RLAIF uses an AI model's judgements instead of human labels to create preference data. Constitutional AI structures this with a written set of principles: the model critiques and revises its own outputs against them (supervised phase), then an AI judge chooses between responses by the principles to train a preference model for RL. It scales harmlessness training and makes the values explicit.
How do you evaluate a fine-tuned model?
On a held-out test set with task metrics (accuracy, F1, exact match, format validity, ROUGE/BERTScore for free text), a validated LLM judge with pairwise comparisons against the base model, general benchmarks and safety suites to detect regression, a product regression suite in CI, and ultimately online A/B tests. Always compare against the best-prompted base model.
Why report both ROUGE-L and BERTScore?
ROUGE-L measures lexical overlap (longest common subsequence), rewarding the same wording; BERTScore measures embedding similarity, rewarding the same meaning with different wording. Together they distinguish copying from correct paraphrasing. Neither checks factual correctness reliably, so add judges or human review for high-stakes domains.
What is packing and what can go wrong with it?
Packing concatenates several examples into one full-length sequence to avoid padding waste. If attention is not blocked at example boundaries and position IDs are not reset, examples can attend to unrelated previous examples (cross-contamination), which slightly hurts quality; modern implementations handle this with per-example position IDs and variable-length attention kernels.
How do DeepSpeed ZeRO stages differ?
ZeRO-1 shards optimizer states across data-parallel GPUs; ZeRO-2 also shards gradients; ZeRO-3 also shards parameters, gathering each layer just in time. Higher stages save more memory per GPU at the cost of more communication. Offload variants move states to CPU or NVMe. FSDP is PyTorch's native equivalent of ZeRO-3.
What is gradient checkpointing and what does it cost?
Instead of storing every intermediate activation for the backward pass, store only checkpoints (such as each layer's input) and recompute the rest during backward. Activation memory drops from proportional to all intermediates to roughly one layer's worth plus the checkpoints, at the cost of roughly one extra forward pass (about 20–35% more compute).
How would you create synthetic fine-tuning data safely?
Seed a strong teacher with diverse real prompts, generate instructions and responses, filter with rules and an LLM judge, verify facts or run code/tests where possible, deduplicate heavily, balance topics, keep a share of human data to avoid collapse, and confirm that the teacher's licence and terms allow training on its outputs.
Advanced
Derive the DPO loss from the RLHF objective.
The KL-regularized objective maxπ E[r(x,y)] − β·KL(π ‖ πref) has the closed-form optimum π*(y|x) = πref(y|x)·exp(r(x,y)/β) / Z(x). Solving for the reward gives r(x,y) = β·log(π*(y|x)/πref(y|x)) + β·log Z(x). Substituting into the Bradley–Terry likelihood of the preference yw ≻ yl, the Z(x) terms cancel because they depend only on x, leaving L = −log σ(β[log(πθ(yw)/πref(yw)) − log(πθ(yl)/πref(yl))]). The policy itself parameterizes the reward: "your language model is secretly a reward model".
What does the DPO gradient do, intuitively?
It increases the log-probability of the chosen response and decreases that of the rejected one, weighted by σ(implicit reward of rejected − chosen), i.e. by how wrongly the current model ranks the pair. Pairs already ranked correctly by a wide margin contribute little; misranked pairs contribute most. This weighting is what prevents the degenerate behaviour of naive "unlikelihood" training.
What are known failure modes of DPO?
Both chosen and rejected likelihoods can fall (the model just widens the margin by making everything less likely), leading to degraded or out-of-distribution outputs; overfitting to deterministic preferences (motivating IPO); length exploitation, since longer chosen answers accumulate larger log-ratios (motivating length normalization in SimPO or length-controlled evaluation); and sensitivity to off-policy data generated by a different model. Remedies: SFT on chosen first, add an SFT/NLL term on chosen, tune β, use on-policy or iterative DPO.
Explain the PPO clipped objective and why clipping matters.
PPO maximizes E[min(ρA, clip(ρ, 1−ε, 1+ε)A)], where ρ is the new/old probability ratio for an action (token) and A the advantage. Clip the ratio, then multiply by A; do not write clip(ρA). When A > 0 the clip stops you rewarding ρ > 1+ε; when A < 0 it stops you rewarding ρ < 1−ε. Either way you lose the incentive to move the policy further in the improving direction, which is the trust region. Without clipping, a few high-|A| samples can cause large destructive updates. The full PPO loss also has a value-function term and often an entropy bonus.
How are advantages computed in PPO-RLHF versus GRPO?
In PPO-RLHF, the reward model score is given at the end of the sequence, a per-token KL penalty is added at each token, and a learned value model estimates expected return so generalized advantage estimation (GAE) produces per-token advantages. In GRPO, there is no value model: G responses are sampled per prompt and each response's advantage is its reward minus the group mean, divided by the group standard deviation, applied to all its tokens. RLOO similarly uses the leave-one-out mean of the other samples as a baseline.
Why can GRPO training stall, and how do you fix it?
If all responses in a group get the same reward (all correct or all wrong), standardized advantages are zero and the prompt contributes no gradient. Too-easy or too-hard prompt sets therefore stall learning. Fixes: curriculum or difficulty filtering to keep prompts where the model succeeds sometimes, larger group size, partial-credit rewards, and dropping zero-variance groups from the batch.
What is reward hacking and how do you detect and mitigate it?
The policy finds outputs the reward model overrates without being better: excessive length, flattering tone, repeated phrases, or for verifiable rewards, formatting tricks that fool the checker. Detect by tracking reward vs held-out human or judge preference (reward rising while true quality flattens), response length, KL from reference, and reading samples. Mitigate with KL penalties, reward-model ensembles and uncertainty, length penalties or normalization, robust verifiers, periodic reward-model retraining on new policy samples, and early stopping.
Why does fine-tuning on new factual knowledge increase hallucination?
Examples containing facts the model does not already know are learned slowly and, when learned, teach the model to produce confident answers not grounded in its internal knowledge. The model generalizes that behaviour to other questions it cannot answer. Fine-tuning mostly teaches how to use existing knowledge; new facts are better supplied via retrieval, and "I don't know" examples help calibration.
What is the superficial alignment hypothesis?
The idea that almost all knowledge and capability comes from pretraining, and alignment/SFT mainly teaches the format and style for interacting with users, which is why a small, very high-quality SFT set can produce a strong assistant. It implies that data quality and diversity matter far more than volume in SFT, and that SFT cannot compensate for a weak base model.
Compare LoRA and full fine-tuning in what they learn and forget.
Studies find that full fine-tuning learns larger-rank weight changes and achieves more on large distribution shifts (for example continued pretraining on code or math), while LoRA learns less but forgets less of the base model's general abilities, acting as a regularizer. For typical instruction tuning with modest data, LoRA on all linear layers with adequate rank performs on par with full fine-tuning.
Explain DoRA and why it can outperform LoRA.
DoRA decomposes each weight into a magnitude vector m (column norms) and a direction V/‖V‖. It trains m directly and applies LoRA to the direction: W' = m·(W0 + BA)/‖W0 + BA‖c. Analysis showed full fine-tuning tends to change magnitude and direction in a decoupled way while LoRA couples them; DoRA restores that flexibility, improving quality especially at low ranks, with extra compute for the norm and still mergeable after training.
Why does the α/r scaling slow learning at high rank, and what does rsLoRA change?
With α/r, the update's magnitude shrinks as r grows, and analysis of learning dynamics shows that the gradients' effect on the output scales poorly, so high-rank adapters effectively learn with a smaller step and fail to use their extra capacity. rsLoRA uses α/√r, which keeps the output update stable as rank increases, making higher ranks actually beneficial.
How do you estimate activation memory for a transformer, and what reduces it most?
A standard estimate for 16-bit activations per layer is s·b·h·(34 + 5·a·s/h) bytes (s sequence, b batch, h hidden, a heads). The 5·a·s2·b attention-score term vanishes with FlashAttention-style kernels; full gradient checkpointing reduces storage to about 2·s·b·h bytes per layer plus one layer's recompute. Micro-batch reduction scales linearly; sequence parallelism and activation offload help at extreme lengths.
How does multi-LoRA serving batch requests for different adapters efficiently?
The base model's matmul is computed once for the whole batch. For the low-rank path, requests are grouped by adapter and a segmented/gathered batched matmul kernel applies each request's own A and B (as in Punica's SGMV or S-LoRA). Adapters are stored in a unified memory pool with paging between CPU and GPU, alongside the KV cache. Limits include maximum rank, the number of GPU-resident adapters, and some overhead versus a merged model.
How do TIES and DARE improve model merging?
Naive averaging of task vectors suffers interference: redundant small changes add noise and opposite-sign changes cancel. TIES trims each task vector to its largest-magnitude entries, elects a sign per parameter by total magnitude, and averages only values agreeing with that sign. DARE randomly drops most delta entries (for example 90%) and rescales the rest by 1/(1−p), exploiting delta redundancy; it is often combined with TIES.
Why use temperature and the T2 factor in logit distillation? Forward or reverse KL?
Temperature T > 1 softens both distributions so the student learns the teacher's relative preferences among non-top tokens ("dark knowledge"). Gradients of the softened KL scale as 1/T2, so multiplying by T2 keeps their magnitude comparable to the hard-label term. Forward KL(teacher ‖ student) is mode-covering: the student spreads mass over all teacher modes and can produce odd samples; reverse KL is mode-seeking, focusing on high-probability teacher regions, often better for generation, and used in on-policy distillation.
Why can merging a LoRA into a 4-bit model harm quality, and what is the right procedure?
The adapter was trained against the dequantized 4-bit weights. Adding its small update to 4-bit values and re-rounding to the same 4-bit grid loses most of the update. Merging into the original 16-bit weights instead introduces a mismatch: the adapter compensated for quantization errors that are no longer there, but in practice that mismatch is small. Standard practice: load the base in 16-bit, merge, evaluate, then apply a fresh post-training quantization (GPTQ, AWQ, GGUF) with a domain calibration set and evaluate again. Methods like LoftQ initialize adapters to reduce this gap.
How does the gradient flow through a 4-bit frozen layer in QLoRA?
For y = Ŵx + s·BAx, backward needs ∂L/∂x = ŴT(∂L/∂y) + s·ATBT(∂L/∂y), which uses dequantized Ŵ in 16-bit, plus gradients for A and B. No gradient is computed or stored for Ŵ itself. The 4-bit weights are dequantized per block both in forward and in backward (and again during checkpoint recomputation), which is where the extra time goes.
How would you add new special tokens (for example tool-call tags) when fine-tuning?
Add them to the tokenizer, call model.resize_token_embeddings(len(tokenizer)), initialize new rows sensibly (for example the mean of existing embeddings rather than random), and make the embedding and output-head matrices trainable (with LoRA, list them in modules_to_save). Save the tokenizer with the adapter. Forgetting to train the new rows leaves the model unable to emit the new tokens.
What is the alignment tax and how can it be reduced?
The drop in some capabilities or benchmark scores after alignment (for example calibration, creativity or certain NLP tasks). Reductions: mixing pretraining-style next-token loss into RL updates, stronger KL constraints, model averaging between SFT and RL checkpoints, better preference data, and targeted capability data during alignment.
How does sequence-length normalization interact with preference losses?
DPO sums log-probabilities over tokens, so longer responses produce larger log-ratio magnitudes, giving systematic advantages to length and encouraging verbosity when chosen responses tend to be longer. SimPO and some DPO variants average per token and add a target margin; evaluations use length-controlled win rates. Data-side fixes include balancing lengths between chosen and rejected.
Explain on-policy versus off-policy preference data.
On-policy data consists of responses sampled from the current policy; off-policy data came from other models or earlier checkpoints. PPO and GRPO are on-policy by design. Offline DPO usually trains off-policy, and its effectiveness drops when the data is far from what the policy would generate, because it adjusts likelihoods of sequences the model would never produce. Iterative/online DPO regenerates and relabels pairs each round to stay on-policy.
What role does the reference model play in DPO when training with LoRA?
The reference provides log πref(y|x) for chosen and rejected responses. With a LoRA policy, the reference is simply the same model with adapters disabled, so no second copy is needed; alternatively, reference log-probabilities can be precomputed once over the dataset. The reference must be the model the policy started from (usually the SFT model), otherwise the implicit reward is miscalibrated.
How do you fine-tune for long context?
Adjust rotary position scaling (position interpolation, NTK-aware or YaRN scaling) to extend the positional range, then continue training on long documents so the model adapts; use FlashAttention, sequence packing with correct attention boundaries, gradient checkpointing and sequence/context parallelism to fit memory. Evaluate with long-context retrieval and reasoning tests, and check that short-context performance did not regress. Some LoRA variants additionally train embeddings and norms for long-context adaptation.
How can fine-tuning compromise safety even with benign data, and what would you do about it?
Safety behaviour is a learned, relatively shallow layer of the model; any gradient update that pushes it towards "always comply helpfully" can erode refusals, and research has shown degradation even on benign instruction data, with very few adversarial examples enough to largely remove guardrails. Mitigations: include safety demonstrations and refusals in the training mix, use PEFT with modest learning rates, evaluate harmful-request refusal and over-refusal before and after, add inference-time moderation, and restrict who can fine-tune hosted models.
Scenario & debugging
Your fine-tuned model is great at the new task but forgot general skills. What do you do?
This is catastrophic forgetting. First quantify it with a general regression suite against the base. Then: switch from full fine-tuning to LoRA (or lower the rank), reduce learning rate and epochs, mix 5–20% general instruction data into training (replay), make sure the original chat template is used, consider interpolating fine-tuned and base weights (or scaling down the LoRA contribution), and retrain with early stopping on a validation set that includes general tasks. If the product permits, serve the task adapter only for the task's traffic and the base model for everything else.
You hit OOM fine-tuning an 8B model on a 24 GB GPU. How do you fix it?
Do the math: bf16 weights alone are 16 GB, so full fine-tuning is impossible and even bf16 LoRA is tight. Use QLoRA (about 5 GB of weights), enable gradient checkpointing, micro-batch 1–2 with gradient accumulation for the effective batch, use FlashAttention/SDPA, cap max_seq_length at the data's p95, use a paged 8-bit optimizer, and consider a chunked/fused cross-entropy since large vocabularies make logits huge. Check that nothing else occupies the GPU (a leftover model from a previous run) and free memory between runs.
After fine-tuning, the model never stops generating and repeats itself. Why?
Almost always the training targets lacked an end-of-turn/EOS token, or the EOS token was masked out of the loss or attention (common when pad = EOS and masks are built from input_ids != pad_token_id), or the serving stop tokens differ from the template's end-of-turn. Fix the data to append the correct token, build masks from lengths, configure stop tokens and the generation config, and retrain.
The training loss does not decrease at all. What do you check?
Print trainable parameters (adapters may not be attached or everything is frozen); check that the completion mask is not all zeros (for example the prompt consumed the whole max_length due to truncation); confirm labels are shifted correctly; confirm the optimizer received the trainable parameters and the scheduler is not keeping LR at zero; check the learning rate is not tiny; and decode a batch to verify the text and mask are what you expect.
Loss becomes NaN a few hundred steps into training. What happened?
Likely fp16 overflow (especially full fine-tuning in fp16), a learning rate too high without warmup, missing gradient clipping, a zero-token mask producing division by zero, or a corrupted example. Switch to bf16 or fp32 master weights, add warmup and clip gradients at 1.0, clamp the loss denominator, lower LR, and bisect to the offending batch.
Training loss is near zero but test answers are worse than the base model. Explain.
Overfitting/memorization: too many epochs on a small dataset, possibly with duplicates. The model regurgitates training answers and loses generality. Use validation-based early stopping and restore the best checkpoint, reduce epochs, deduplicate, add data diversity, lower rank or LR, and evaluate with generation metrics rather than loss.
Validation loss improved but human raters prefer the base model. Why could that be?
Loss measures likelihood of references, not quality. If references are short, bland or machine-generated, the model learns to be bland; it may drop caveats, become terse or copy reference style. Evaluation prompts may differ from real traffic. Revisit data quality, use pairwise human or judge evaluations during model selection, and consider preference tuning to optimize what raters actually prefer.
The model works in your notebook but produces garbage when served. What do you check?
Template and tokenizer mismatch (serving applies a different chat template or none), missing special tokens, a different base revision than the adapter was trained on, adapter not actually loaded, merging into a quantized model, different stop tokens, or a quantized export that lost too much precision. Run the exact serving path on the evaluation set and diff outputs against the notebook.
You have 500 labelled examples for a classification task. Fine-tune or prompt?
Start with few-shot prompting and measure on a held-out slice. If accuracy or consistency is insufficient, or per-call cost matters, fine-tune a small model with LoRA (or a small encoder with a classification head) using 400 for training and 100 for validation/test with stratification. 500 clean examples are usually enough for a narrow classification task, and a fine-tuned small model is often both cheaper and more accurate.
A support bot must follow a fixed JSON format and know policies that change monthly. Design the solution.
RAG for the policies (update the index monthly, cite sources) and structured output/constrained decoding or a fine-tune for the format and tone. Fine-tune with LoRA on transcripts where the context includes retrieved passages, so the model learns to ground answers in them and emit the schema. Evaluate format validity, grounding and correctness; retrain only when behaviour changes, not when policies change.
DPO training shows rising reward accuracy but outputs become longer and worse. What is going on?
Likely length exploitation and likelihood collapse: chosen responses are longer, so the model learns that length increases the margin, while log-probabilities of both chosen and rejected fall. Check chosen log-probs and length trends. Fixes: balance lengths in data, raise β, lower LR, fewer steps, add an SFT term on chosen, try length-normalized methods (SimPO) or IPO, and evaluate with length-controlled judges.
After an RLHF run, reward is climbing fast but responses are full of flattery and filler. What do you do?
Classic reward hacking. Stop and roll back to an earlier checkpoint, increase the KL coefficient, add a length penalty, retrain or ensemble the reward model with examples of these exploits labelled as bad, monitor judge or human win rates rather than reward alone, and cap generation length.
Your fine-tuned open model now answers harmful requests it used to refuse. Why and how do you fix it?
Fine-tuning eroded safety alignment, which can happen even with benign data. Add safety and refusal demonstrations (and benign look-alike prompts to avoid over-refusal) to the data mix, lower the learning rate or use smaller LoRA rank, run a safety evaluation suite before release, and add inference-time guardrails such as input/output moderation.
You need 100 customer-specific variants of a model. How do you train and serve them?
Train one LoRA per customer on a shared base (same revision, same template) with a common pipeline and per-customer evaluation. Serve with multi-LoRA: one base in GPU memory, adapters paged in on demand, requests routed by adapter name and batched together. Version adapters, isolate customer data during training, and give the highest-traffic customers merged dedicated deployments if latency demands it.
Fine-tuned results vary a lot between runs. How do you make conclusions reliable?
Fix seeds and data order, but also run multiple seeds and report mean and variance; use larger evaluation sets and bootstrap confidence intervals; make sure early stopping and checkpoint selection are deterministic; and only claim a method is better when the gap exceeds seed noise. Small data (hundreds of examples) is especially noisy.
A teammate wants to train on outputs from a commercial API to build a competing model. What do you flag?
Terms of service of many providers prohibit using outputs to develop competing models; this is a legal and reputational risk. Also check licences of the base model and any datasets, data-privacy obligations for prompts sent to the API, and documentation of provenance. Alternatives: use teachers whose licences permit distillation, or human-written data.
You must deploy your fine-tuned 7B model on laptops without GPUs. What is the path?
Merge the adapter into a 16-bit base, convert to GGUF, quantize to a k-quant such as Q4_K_M (about 4–4.5 GB) or Q5_K_M for more quality, verify the embedded chat template, and run with a llama.cpp-based runtime. Re-run the regression suite on the quantized file, compare against 16-bit, and consider a smaller distilled model if latency is too high. See the Edge AI page for on-device constraints.
QLoRA training is much slower than you expected. What can you do?
QLoRA pays for dequantization and gradient checkpointing recomputation. Options: use bf16 LoRA if memory allows; increase micro-batch size (better GPU utilization) while keeping memory within limits; turn off checkpointing if memory permits; enable packing to cut padding; use FlashAttention; use optimized kernels (for example Unsloth- or Liger-style fused ops); and shorten maximum sequence length.
Your fine-tuned model's benchmark score looks suspiciously high. What do you check?
Contamination: overlap between training data (including synthetic data from a teacher that may have seen the benchmark) and the test set, using n-gram and embedding similarity. Check for duplicates across splits, grouped leakage (same documents in train and test), and whether prompts or answers were used to generate training data. Re-evaluate on a fresh, private test set.
Your LoRA with r=64 performs worse than r=16. Why?
If α was held fixed, the scale α/r dropped fourfold, lowering the effective learning rate; or the larger adapter overfit a small dataset. Re-run with α/r constant (or rsLoRA scaling), compare on validation metrics across seeds, and prefer the smaller rank if there is no real gain.
You fine-tuned a base (non-instruct) model on chat data and it outputs strange role markers. What went wrong?
The base model has no chat template, so you chose one; if its special tokens were not added to the tokenizer (or were added without resizing and training embeddings and the LM head), the model sees them as fragmented text and reproduces fragments. Add the tokens properly, train the new embedding rows (modules_to_save), mask only non-assistant turns, and use the same template at inference.
You want your model to reason better on math. SFT on solutions helped a little. What next?
Use RL with verifiable rewards: GRPO (or RLOO/PPO) on problems with checkable answers, a correctness reward plus a small format reward, prompts filtered to moderate difficulty so groups have reward variance, a KL penalty to the SFT model, and evaluation on held-out math sets. Alternatively or first, distil chain-of-thought traces from a stronger reasoning model (respecting its licence).