Generative AI & LLMs

Fine-Tuning & Alignment

Fine-tuning takes a pretrained language model and keeps training it on your own data so that it adopts a new behaviour, format, style or skill; alignment is the family of post-training methods (RLHF, DPO, GRPO and friends) that shape a model towards what humans actually prefer. This page covers the whole lifecycle, from deciding whether to fine-tune at all, through data, LoRA/QLoRA, memory math and preference optimization, to evaluating, merging, distilling and serving the result.

~130 min read 0 interview questions
In 30 seconds
  • Try prompting first, add RAG when the model lacks knowledge, and fine-tune when it lacks behaviour (format, tone, skill, latency/cost targets).
  • Supervised fine-tuning (SFT) is next-token prediction on curated prompt–response pairs, with the loss usually computed on the response tokens only.
  • Full fine-tuning with AdamW costs about 16 bytes per parameter before activations (≈112 GB for 7B); LoRA trains a low-rank update ΔW = (α/r)·BA on a frozen model, and QLoRA also stores the frozen model in 4-bit NF4, so a 7B model trains on one consumer GPU.
  • Alignment turns a helpful-looking model into a preferred one: classic RLHF is SFT → reward model → PPO with a KL penalty; DPO skips the reward model and RL loop; GRPO is the critic-free RL used for reasoning with verifiable rewards.
  • Data quality beats quantity: a few thousand clean, deduplicated, correctly templated examples usually beat a large, noisy dump.
  • The main risks are catastrophic forgetting, overfitting, safety regression, data leakage and licensing, so always compare against the base model on a regression suite.
  • To deploy, either merge the adapter into the base weights (no latency overhead) or serve many LoRAs over one shared base, then quantize (GPTQ/AWQ/GGUF) for cheap or on-device inference.

The big picture: where fine-tuning fits

A modern chat model is built in stages. Pretraining teaches a network general language and world knowledge by predicting the next token over trillions of tokens of text. The result, a base model, is a powerful autocomplete engine: ask it a question and it may continue with more questions, because that is what the internet looks like. Post-training then turns the base model into an assistant: supervised fine-tuning on demonstrations teaches it the question-answer format, and preference optimization (RLHF, DPO and relatives) teaches it which of several plausible answers humans actually prefer. When you "fine-tune" as a practitioner, you are adding one more post-training step on top of an existing base or instruct model, aimed at your own use case.

Analogy

Think of a medical graduate. Years of general schooling and medical school are pretraining: broad knowledge, expensive, done once. A residency in cardiology is fine-tuning: a much shorter, focused programme that turns the generalist into a specialist who follows the conventions of one department. Feedback from senior doctors on which diagnosis explanation was clearer is alignment: it does not add much new knowledge, it refines judgement and manner.

Mapping back: pretraining is the costly general education, SFT is the residency that teaches a specific way of working, and preference tuning is the senior-doctor feedback that sharpens which answers are preferred. Just like a resident, a fine-tuned model can also forget things from outside the specialty if it trains too narrowly.

                 THE LLM LIFECYCLE
 ┌──────────────┐   ┌──────────────┐   ┌──────────────────┐   ┌──────────────┐
 │ Pretraining  │──▶│ SFT /        │──▶│ Preference       │──▶│ Your domain  │
 │ next-token   │   │ instruction  │   │ optimization     │   │ fine-tune    │
 │ on ~10T tok  │   │ tuning       │   │ RLHF / DPO / RL  │   │ (LoRA/QLoRA) │
 └──────────────┘   └──────────────┘   └──────────────────┘   └──────────────┘
   base model         instruct model      chat/aligned model     specialised model
   (knowledge)        (follows format)    (preferred answers)    (your behaviour)
   months, $M+        hours-days           days                   minutes-hours

What fine-tuning actually changes

Fine-tuning runs ordinary gradient descent on the model's weights (all of them, or a small added set) using a loss computed on your examples. After training, the change lives in the weights: no extra prompt is needed at inference time. This is the key difference from prompting and in-context learning, where the weights are frozen and the model adapts only through the text you give it, and from retrieval-augmented generation (RAG), where facts are fetched at query time and pasted into the context.

Behaviour and format

Always answer in a given JSON schema, adopt a brand voice, follow a house style, stop being verbose. Fine-tuning is excellent at this.

Skills

Classify tickets, extract fields, write SQL for your schema, call your tools correctly, reason in a domain-specific way. Fine-tuning is good at this with enough examples.

Knowledge

Memorizing facts, especially ones that change, is where fine-tuning is weakest: it is slow to update, hard to cite, and prone to hallucinating near-miss facts. Use RAG for this.

Cost and latency

A small fine-tuned model can replace a large general model with a long prompt, cutting per-call cost and latency by an order of magnitude.

Interview angle A common opener is "What is the difference between pretraining, instruction tuning, fine-tuning and RLHF?" A strong answer names the objective and data of each stage: pretraining is self-supervised next-token prediction on raw text; instruction tuning is supervised next-token prediction on instruction–response pairs; task/domain fine-tuning is the same mechanism on narrow data; RLHF optimizes a learned reward from human preference comparisons under a KL constraint to stay close to the SFT model.
Common pitfall Believing that a model "learns" from the conversations you have with it. Inference never updates weights; only an explicit training run does. Likewise, "prompt tuning" in the casual sense of rewording prompts is not training, whereas the PEFT method called prompt tuning does train soft prompt vectors.

Prompt vs RAG vs fine-tune: making the decision

Fine-tuning is the most expensive lever in terms of data, engineering and maintenance, so it should be pulled deliberately. The standard progression is: start with a strong prompt (zero-shot, then few-shot), add tools or structured output when you need determinism or actions, add RAG when the model is missing knowledge, and fine-tune when the desired behaviour is still unreliable or too expensive to achieve through prompts. The rule of thumb is short: new knowledge → RAG; new behaviour → fine-tune; both → both.

Analogy

Imagine onboarding a capable new employee. Prompting is giving clear instructions in each email. RAG is handing them the company handbook to look things up as needed, which stays current as the handbook is edited. Fine-tuning is sending them on a training course so that the company's way of doing things becomes second nature, which is great for habits but useless for remembering a policy that changes next week.

Mapping back: instructions correspond to prompts (cheap, instant, but must be repeated every time), the handbook corresponds to a retrieval index (fresh and citable), and the training course corresponds to weight updates (durable habits, but frozen at training time).

NeedPromptingRAGFine-tuning
Fresh or frequently changing factsOnly if pasted in manuallyBest choice: update the indexPoor: facts are frozen, retraining needed
Citations and traceabilityNoYes, cite retrieved chunksNo
Consistent format, tone or styleWorks but drifts on long tailsDoes not helpBest choice
New skill (domain classification, extraction, tool calling)Few-shot may be enoughLimited helpStrong
Private data at scale (large corpus)Context window limitsDesigned for itContinued pretraining possible but costly
Lower latency and cost per callLong prompts cost tokensAdds retrieval latency and tokensSmall tuned model can replace a big one
Time to first resultMinutesDaysDays to weeks (data is the bottleneck)
Data neededA few examplesDocumentsHundreds to tens of thousands of labelled examples
MaintenanceEdit promptKeep index freshRetrain on base-model upgrades, re-evaluate
Hallucination controlModerateGood (grounding)Can increase if trained on facts it did not know

A decision procedure

  1. Define success and build an eval set first. Without 50–200 labelled test cases and a metric, you cannot tell whether any lever helped.
  2. Baseline with prompting. Zero-shot, then few-shot, then a structured system prompt. Record scores, cost and latency. See Prompt Engineering.
  3. Diagnose the failures. Are they missing facts (retrieval problem) or wrong behaviour despite having the facts (format, reasoning, tone, refusal)?
  4. Missing facts → RAG. Build retrieval and re-evaluate. See Retrieval-Augmented Generation.
  5. Behaviour still wrong, or too costly → fine-tune. Start with LoRA/QLoRA SFT on a few thousand high-quality examples; consider preference tuning if the issue is "which answer is better" rather than "what format".
  6. Combine. Many production systems fine-tune for style, format and tool use, and use RAG for facts. A retrieval-aware fine-tune (training on examples that include retrieved context) teaches the model to use context faithfully.

Good reasons to fine-tune

  • Strict output format that prompting breaks 2–5% of the time
  • Distilling a big model's behaviour into a small, cheap one
  • Domain language the base model handles poorly (clinical notes, legal clauses, internal codes)
  • Tool/function calling for a private API
  • Latency or on-device limits that force a small model
  • Removing a very long system prompt from every call

Bad reasons to fine-tune

  • "Teach it our product catalogue" that changes weekly
  • No evaluation set yet
  • Fewer than a hundred messy examples
  • Prompting has not been seriously tried
  • Hoping to fix factual hallucinations by training on facts
  • A closed model that only allows limited API fine-tuning, when prompting already works
Tip Fine-tuning and RAG are not rivals. A support bot can be fine-tuned on transcripts to learn tone, escalation rules and response format, while RAG supplies the current refund policy. Fine-tuning alone goes stale; RAG alone may sound generic.
Interview angle Expect "When would you choose fine-tuning over RAG?" Answer with the knowledge-vs-behaviour split, then add the operational factors: freshness, citations, cost per call, latency, data availability and maintenance. Bonus points for mentioning that fine-tuning on new facts can raise hallucination rates, because the model learns to produce confident answers in a style it cannot back up.
Common pitfall Comparing a fine-tuned model against a lazy prompt. Always compare against the best prompt you can write for the base model; many "fine-tuning wins" disappear against a good few-shot baseline.

Types of fine-tuning

"Fine-tuning" is an umbrella term. The variants differ in what data is used, what loss is optimized and which parameters are updated. The first two axes decide what the model learns; the third (full vs parameter-efficient) decides cost.

Analogy

Consider a chef. Reading hundreds of regional cookbooks cover to cover is continued pretraining. Practising exact recipes that customers order is instruction tuning. Specialising in only pastry is task-specific tuning. Having diners taste two versions of a dish and pick the better one is preference tuning. Whether the chef relearns every technique from scratch or just adds a few new habits is the full-vs-parameter-efficient choice.

Mapping back: raw domain text maps to continued pretraining, prompt–response pairs map to SFT, a narrow labelled task maps to task-specific tuning, and chosen-vs-rejected pairs map to preference optimization; the chef's retraining scope maps to updating all weights versus small adapters.

TypeDataLossTypical use
Continued pretraining (domain-adaptive pretraining)Unlabelled domain text: papers, code, contracts, other languagesNext-token loss on all tokensTeach vocabulary, style and knowledge of a domain or language; followed by SFT to restore chat ability
Supervised fine-tuning / instruction tuningInstruction (+ optional input) → response pairs, or multi-turn chatsNext-token loss, usually on response tokens onlyTurn a base model into an assistant, or teach a format/skill
Domain adaptationDomain-specific instructions (medical Q&A, legal drafting)SFT loss, sometimes after continued pretrainingSpecialist assistant for a vertical
Task-specific fine-tuningLabelled examples for one task (classification, NER, extraction, SQL)Cross-entropy on labels, via a classification head or generative targetsHigh accuracy on one narrow task, often with a small model
Preference tuningPrompt with chosen and rejected responses, or scalar rewardsRLHF (reward + PPO), DPO, ORPO, KTO, GRPOHelpfulness, harmlessness, style preferences, reasoning
Distillation fine-tuningTeacher model outputs (text or full probability distributions)SFT on teacher text, or KL to teacher logitsCompress a large model's behaviour into a small one
Long-context fine-tuningLong documents with adjusted positional scalingNext-token lossExtend context window (with RoPE scaling tricks)

Full fine-tuning vs parameter-efficient fine-tuning

Full fine-tuning

  • Every weight gets a gradient and optimizer state
  • Highest capacity; best for large distribution shifts (new language, continued pretraining)
  • Memory ≈ 16 bytes per parameter plus activations
  • Produces a full new copy of the model per task
  • Higher risk of catastrophic forgetting
  • Needs small learning rates (around 1e-5)

Parameter-efficient fine-tuning (PEFT)

  • Base weights frozen; train <1–2% extra parameters
  • Matches full fine-tuning on most SFT tasks when rank and target modules are adequate
  • Fits on one GPU; QLoRA fits 7B–8B in under 16 GB
  • Adapter files are megabytes; many tasks share one base
  • Forgets less, because the base is untouched
  • Tolerates larger learning rates (around 2e-4)

Instruction tuning in detail

Instruction tuning trains on examples that look like the requests a user would make. A single record typically has three fields, instruction, an optional input (extra context) and output (the target answer). Diversity of instructions matters as much as volume: models tuned on many different task types generalize to unseen instructions, which is why instruction tuning produced the jump from "autocomplete" to "assistant". Instruction tuning is different from prompt engineering: it changes weights during training, while prompt engineering changes only the input at inference.

{"instruction": "Classify the sentiment of the review.",
 "input": "The battery died after two days.",
 "output": "negative"}

{"instruction": "Explain what an ACE inhibitor does in one sentence.",
 "input": "",
 "output": "An ACE inhibitor lowers blood pressure by blocking the enzyme that makes angiotensin II, a hormone that narrows blood vessels."}
Interview angle "Is instruction tuning the same as fine-tuning?" It is a kind of fine-tuning. Fine-tuning is the general mechanism of further training; instruction tuning is fine-tuning on a broad mix of instruction–response pairs so the model follows natural-language commands; task fine-tuning targets one task; RLHF/DPO optimize for preferences rather than imitation.
Common pitfall Running SFT on raw domain documents and expecting a chat model. Unstructured text is continued-pretraining data; if you train a chat model on it without chat formatting, it drifts back towards autocomplete behaviour. Either format it into Q&A pairs or follow continued pretraining with an SFT stage.

Data preparation: the part that decides the outcome

In fine-tuning, the dataset is the specification. The model will imitate whatever it sees, including typos, inconsistent formats, wrong answers and a verbose tone, so most of the work in a successful project is data work. This section covers formats, chat templates, cleaning, deduplication, synthetic data, and how to split the data for honest evaluation.

Analogy

An apprentice copies the master's work exactly. Show them fifty beautifully finished pieces and they learn craftsmanship; show them five thousand rushed pieces with inconsistent measurements and they learn to be sloppy at scale. They also cannot learn from a piece whose instructions are written in a notation they have never seen.

Mapping back: the finished pieces are your training examples (quality matters more than count), the inconsistent measurements are noisy labels and mixed formats, and the unfamiliar notation is a prompt that does not match the model's chat template.

Common dataset formats

FormatShapeUsed for
Instruction (Alpaca-style)instruction, input, outputSingle-turn SFT
Conversational (messages)messages: [{role, content}, ...] with system/user/assistant (and tool) rolesMulti-turn SFT, tool calling; the modern default
Prompt–completionprompt, completionCompletion-only SFT, classification as text
Plain texttextContinued pretraining
Preference pairsprompt, chosen, rejectedReward models, DPO, ORPO, IPO, SimPO
Unpaired feedbackprompt, completion, label (good/bad)KTO
Prompts only (+ verifier)prompt and a reference answer or testOnline RL such as PPO or GRPO
{"messages": [
  {"role": "system", "content": "You are a concise support assistant for a telecom company."},
  {"role": "user", "content": "My bill is higher this month. Why?"},
  {"role": "assistant", "content": "Bills usually rise for three reasons: a promotion ended, you used extra data, or a one-off fee was added. Check the 'Charges' tab for the exact line item."}
]}

Chat templates: format must match the model

Every instruct model was trained with a specific chat template: special tokens that mark where each role's turn begins and ends. For example, one popular family wraps turns as <|im_start|>user ... <|im_end|>, another uses [INST] ... [/INST], and another uses header tokens such as <|start_header_id|>assistant<|end_header_id|>. If you fine-tune an instruct model with a different layout, you waste capacity re-teaching the format and often break its existing chat skills. In Hugging Face Transformers the tokenizer carries its template, and tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) renders it.

<|im_start|>system
You are a helpful medical assistant. Answer accurately and safely.<|im_end|>
<|im_start|>user
What is the first-line treatment for mild hypertension?<|im_end|>
<|im_start|>assistant
Lifestyle changes come first: reduce salt, ...<|im_end|>
  • add_generation_prompt=True appends the opening of the assistant turn, so the model knows it should answer next. Use it when rendering prompts for training-with-masking and for inference.
  • End-of-turn token. The target must end with the template's end-of-turn or EOS token (for example <|im_end|>). If you omit it, the model never learns when to stop and rambles until max_new_tokens.
  • No double special tokens. A rendered template often already contains BOS; tokenizing it again with add_special_tokens=True can prepend a second BOS.
  • Base models have no template. You may choose one (a common choice is to adopt the template of the matching instruct model), but then add any new special tokens to the tokenizer, resize embeddings and train the embedding rows too.

Loss masking: train on answers, not prompts

In SFT we usually compute the loss only on the assistant tokens (completion-only or "train on responses"). Prompt tokens still pass through the model as context, but their labels are set to -100 (the ignore index in PyTorch cross-entropy), or multiplied by a zero mask. Otherwise the model spends capacity learning to reproduce user prompts and system text, which is not the behaviour you want, and long prompts dominate the gradient. For multi-turn data, mask every user and system turn and supervise every assistant turn.

tokens : [sys][sys][user][user][user][asst-hdr] A1  A2  A3  <|im_end|> [pad][pad]
labels : -100 -100  -100  -100  -100    -100    A1  A2  A3  <|im_end|> -100 -100
                 context only (no loss)          loss here (incl. stop token)   ignored

Quality over quantity

A consistent finding across instruction-tuning work is that a small set of carefully curated examples (around a thousand in some well-known experiments) can produce a strong assistant from a good base model, while larger noisy sets add little. The base model already has the knowledge; SFT mostly teaches format and style, which is sometimes called the superficial alignment hypothesis. Rough sizes that work in practice:

GoalTypical size
Output format or style200–2,000 examples
Simple classification / extraction500–5,000 examples
Domain assistant (SFT)5,000–50,000 examples
General instruction tuning from a base model10,000–1,000,000 examples
Preference tuning (DPO)2,000–100,000 pairs
Continued pretrainingHundreds of millions to billions of tokens

Cleaning checklist

  1. Normalize the schema. One format, consistent roles, placeholder tokens like a literal <noinput> converted to an empty string, whitespace stripped.
  2. Filter bad examples. Empty or truncated answers, wrong language, broken markup, refusals where you want answers (or the reverse), answers that contradict the instruction.
  3. Deduplicate. Exact duplicates via hashing; near-duplicates via MinHash/LSH or embedding similarity. Duplicates overweight some behaviours and leak between train and test.
  4. Remove sensitive data. Scrub personal information, secrets and credentials; models can memorize and regurgitate rare strings.
  5. Check lengths. Compute token-length percentiles and set max_seq_length near the 95th percentile; drop or split outliers instead of silently truncating answers.
  6. Balance. Watch class balance and task mix; mix in some general-purpose data to reduce forgetting.
  7. Decontaminate. Remove any example that overlaps with your evaluation set or with public benchmarks you plan to report.
  8. Human spot-check. Read at least 100 random examples. Most dataset bugs are obvious to a human in ten minutes.

Synthetic data

Much modern fine-tuning data is generated by a stronger model: the teacher writes instructions (self-instruct style), answers them, rewrites human answers into a better format, or produces chosen/rejected pairs. Synthetic data is cheap and controllable but needs guardrails: seed it with diverse real prompts, filter with rules and an LLM judge, verify facts or run code/tests where possible, deduplicate aggressively (generators repeat themselves), and check the teacher's licence and terms of use, which often forbid training competing models. Beware model collapse: repeated training on model-generated text narrows diversity, so keep a share of real human data.

Train, validation and test splits

  • Shuffle once with a fixed seed, then slice train/validation/test (for example 80/10/10, or at small scale a few hundred / a few dozen / a few dozen as in the worked example later).
  • Split by group, not by row when examples are related (same customer, same document, same conversation), otherwise near-duplicates leak across splits.
  • Validation drives early stopping and hyperparameter choices; test is touched once at the end.
  • Keep a separate regression set of general tasks (math, coding, instruction following, safety prompts) to measure forgetting.

Packing and padding

Batches must be rectangular, so shorter sequences are padded. Padding wastes compute when lengths vary a lot. Packing concatenates several examples into one sequence of the maximum length, separated by EOS; with modern kernels, position IDs are reset per example and attention is blocked across boundaries so examples do not attend to each other. Packing can speed training several-fold on short-example datasets. Without packing, group similar lengths into the same batch to reduce padding.

Interview angle Interviewers like "How would you build a dataset for fine-tuning a support assistant?" A strong answer covers sourcing (real transcripts plus targeted synthetic), cleaning, PII removal, deduplication, formatting with the target model's chat template, completion-only loss, group-aware splits, a held-out test set and a general regression set, and a human review loop.
Common pitfall Setting the pad token equal to EOS and then building the attention mask or labels as input_ids != pad_token_id. This also masks the real end-of-sequence token, so the model never learns to stop. Build masks from sequence lengths instead, or use a dedicated pad token.

How supervised fine-tuning works under the hood

SFT uses exactly the same objective as pretraining: the model predicts each next token and pays a cross-entropy penalty for putting low probability on the true one. The only differences are the data (curated examples instead of raw web text) and the mask (loss on answer tokens only). Understanding the shift-by-one and the masking removes most confusion about "how the model learns to answer".

Analogy

Picture a language tutor who reads a model answer aloud one word at a time and, before each word, asks the student to guess it. The student hears the question in full but is only graded on guesses within the answer. Every wrong guess produces a correction proportional to how surprised the student was.

Mapping back: the question is the masked prompt, each guess is a next-token prediction, the grade is the cross-entropy loss on answer tokens, and the correction is the gradient update. Because the true previous word is always revealed before the next guess, this is teacher forcing.

LSFT(θ) = − (1 / Σt mt) · Σt mt · log pθ(xt | x<t) where x is the full tokenized sequence (prompt + answer), mt = 1 if token t belongs to the answer and 0 otherwise, and pθ is the model's softmax over the vocabulary. Dividing by the number of answer tokens gives a per-token mean.

The shift-by-one

A causal language model produces logits at every position; the logits at position t are a prediction for token t+1. So we compare logits[:, :-1] with input_ids[:, 1:] and shift the mask the same way. Hugging Face models do this internally when you pass labels; in a hand-written loop you do it yourself:

def batch_loss(model, input_ids, attn, comp_mask):
    """Mean next-token cross-entropy over answer tokens only."""
    logits = model(input_ids=input_ids, attention_mask=attn).logits
    shift_logits = logits[:, :-1, :].contiguous()          # predictions for t+1
    shift_labels = input_ids[:, 1:].contiguous()           # the true next tokens
    shift_mask   = comp_mask[:, 1:].contiguous().float()   # 1 on answer tokens

    per_tok = F.cross_entropy(shift_logits.view(-1, shift_logits.size(-1)),
                              shift_labels.view(-1), reduction="none")
    return (per_tok * shift_mask.view(-1)).sum() / shift_mask.sum().clamp(min=1.0)

Token-level vs sequence-level averaging

Averaging over all answer tokens in a batch weights long answers more heavily than short ones; averaging per sequence first weights every example equally. With gradient accumulation, the naive approach of averaging each micro-batch and then averaging micro-batches is subtly wrong when micro-batches contain different token counts; libraries now normalize by the total number of supervised tokens across the accumulated batch. For most SFT it is a minor effect, but it matters when lengths vary widely.

The training loop in one picture

for each epoch:
  for each micro-batch:
      forward  ──▶ masked loss / grad_accum ──▶ backward (grads accumulate)
      every grad_accum steps:
          clip_grad_norm(1.0) ──▶ optimizer.step() ──▶ scheduler.step() ──▶ zero_grad()
  evaluate on validation ──▶ keep best checkpoint ──▶ early-stop if val loss rises

Mixed precision and numerics

  • bf16 has the same exponent range as fp32, so it rarely overflows; it is the default on GPUs that support it (Ampere and newer).
  • fp16 has a narrow range and needs loss scaling; full fine-tuning in pure fp16 can overflow to inf/NaN. Older GPUs such as the T4 lack fast bf16, so a common pattern there is fp32 master weights for full fine-tuning, and fp16 compute for LoRA/QLoRA (where only small adapters are trained).
  • Master weights in fp32 keep tiny updates from being rounded away; mixed-precision training keeps an fp32 copy for the optimizer while computing in 16-bit.
Interview angle "Why mask the prompt tokens?" Because we want the model to learn the mapping from prompt to answer, not to model the distribution of prompts; including them wastes capacity, lets long prompts dominate the gradient and can teach the model to generate user-like text. Mention that for very short answers, some practitioners do keep a small weight on prompt tokens, and that it is an empirical choice.
Common pitfall Tokenizing the prompt and answer separately and concatenating the token lists. Tokenizers merge characters differently at the boundary, so the training tokens differ from what the model sees at inference. Tokenize the full string once and use offset mappings (or tokenize the prompt with the same settings and verify it is a prefix) to find where the answer begins.

Memory math for fine-tuning

Most practical fine-tuning questions reduce to "will it fit on my GPU?" GPU memory during training holds four things: the weights, the gradients, the optimizer states, and the activations saved for the backward pass (plus some framework overhead and temporary buffers). The first three scale with the number of trainable parameters; activations scale with batch size, sequence length, hidden size and number of layers.

Analogy

Think of renovating a large library. You need space for the books themselves (weights), a note on every book saying how to change it (gradients), and a running history for every book of past changes so edits are smooth (optimizer states). You also need scratch tables for the pages currently open (activations), and the number of tables grows with how many readers and how long each document is. Deciding to edit only a small index card per shelf instead of every book is LoRA; storing the books in a compressed format is QLoRA.

Mapping back: per-book notes and histories exist only for books you are allowed to edit, which is exactly why freezing the base model removes most of the gradient and optimizer memory, while compression shrinks the weight storage itself.

Bytes per parameter

ComponentFull FT, mixed precision + AdamWFull FT, pure fp32 + AdamWNotes
Weights2 (bf16)4Needed for forward/backward
Gradients2 (bf16)4Same size as trainable weights
fp32 master weights4(included above)Only in mixed precision
Adam first moment m448-bit Adam: 1 byte
Adam second moment v448-bit Adam: 1 byte
Total16 bytes16 bytesBefore activations
Mtrain ≈ Ntotal·bw + Ntrain·(bg + bopt) + Mact + Moverhead where Ntotal is all parameters, Ntrain the trainable ones, bw bytes per stored weight (2 for bf16, ≈0.5 for 4-bit), bg bytes per gradient, bopt bytes of optimizer state per trainable parameter (12 for mixed-precision AdamW including the master copy, 8 for fp32 Adam states only, 2 for 8-bit Adam), and Mact the saved activations.

Worked example: a 7B model

Take a 7B-parameter decoder with 32 layers, hidden size 4096, 32 attention heads and an MLP width of 11008 (a common 7B shape). Parameter count is about 6.7B; use 7B for round numbers.

MethodWeightsGrads + optimizerTotal before activationsFits on
Inference only, bf1614 GB014 GB (+ KV cache)One 24 GB GPU
Full FT, mixed precision AdamW14 GB7B × 14 B = 98 GB≈112 GBMultiple 80 GB GPUs with sharding (ZeRO-3/FSDP)
Full FT, 8-bit Adam + bf16 grads14 GB7B × (2 + 2 + 4 master) ≈ 56 GB≈70 GBOne 80 GB GPU, tight
LoRA r=16, all linear layers, bf16 base14 GB40M × 16 B ≈ 0.64 GB≈15 GBOne 24 GB GPU
QLoRA r=16, NF4 base with double quantization≈3.9 GB≈0.64 GB (less with paged 8-bit optimizer)≈4.5 GBOne 8–12 GB GPU (with checkpointing)

Where the 40M LoRA parameters come from: with rank r, an adapter on a din × dout linear layer adds r·(din + dout) parameters. Per layer: four attention projections of 4096×4096 give 4 × 16 × 8192 = 524,288; three MLP projections between 4096 and 11008 (gate, up, down) give 3 × 16 × 15104 = 724,992. That is ≈1.25M per layer, ×32 layers ≈ 40M parameters, or about 0.6% of the model. In bf16 the adapter file is about 80 MB. This 7B shape is multi-head attention; grouped-query models have thinner k_proj and v_proj, so the attention term is smaller. Embeddings and the LM head are not in this 40M unless you list them in modules_to_save.

Where the 3.9 GB NF4 figure comes from: 4 bits per weight is 0.5 bytes, so 7B weights take 3.5 GB. Blockwise quantization stores one scale per 64 weights; in fp32 that is 32/64 = 0.5 extra bits per weight, and double quantization reduces it to about 0.127 bits. Embeddings and the output head are often kept in 16-bit, which adds a few hundred MB.

Activations: the part people forget

For each transformer layer the backward pass needs intermediate tensors from the forward pass. A widely used estimate for 16-bit activations per layer, without recomputation, is roughly s·b·h·(34 + 5·a·s/h) bytes, where s is sequence length, b micro-batch size, h hidden size and a number of heads. The second term is the attention score matrix, which memory-efficient attention kernels (FlashAttention and SDPA) avoid storing.

Mact ≈ L · s · b · h · (34 + 5·a·s / h) bytes For L = 32, h = 4096, a = 32, s = 2048, b = 1: 34·s·b·h ≈ 0.29 GB and 5·a·s2·b ≈ 0.67 GB per layer, so ≈31 GB in total; with a fused attention kernel ≈9 GB; with gradient checkpointing only the layer inputs (2·s·b·h bytes ≈ 17 MB per layer) plus one layer's recomputation are held, i.e. about 1–2 GB.

These are estimates: exact numbers depend on MLP width, dropout masks and the framework. The important scaling rules are that activations grow linearly with batch size and layers, linearly with sequence length once fused attention is used (quadratically without it), and shrink dramatically with gradient checkpointing at a cost of roughly 20–35% more compute (one extra forward pass).

Levers to reduce memory, in the order people usually pull them

  1. Smaller micro-batch, more gradient accumulation. Same effective batch, far fewer activations.
  2. Gradient checkpointing. Recompute activations during backward; large savings for about a third more compute.
  3. Fused/flash attention. Removes the quadratic attention-matrix term.
  4. Shorter max_seq_length based on the data's length percentiles.
  5. LoRA instead of full fine-tuning: removes gradients and optimizer states for the frozen weights.
  6. QLoRA: 4-bit base weights, about a quarter of the bf16 footprint.
  7. 8-bit or paged optimizers to shrink optimizer states and survive spikes.
  8. Sharding and offload: ZeRO-2/3 or FSDP across GPUs, CPU/NVMe offload of optimizer states or parameters.
  9. Chunked loss / fused cross-entropy: the logits tensor (batch × seq × vocab) is huge for 150k-token vocabularies; computing the loss in chunks avoids materializing it all.
Tip The logits alone can be surprising: with a 152k vocabulary, a micro-batch of 4 sequences of 2048 tokens produces 4 × 2048 × 152,000 ≈ 1.25 billion logits, which is 2.5 GB in bf16 and 5 GB if upcast to fp32 for the loss. Large-vocabulary models often run out of memory at the loss, not in the transformer body.
Interview angle "How much memory to fully fine-tune a 7B model?" Walk through 16 bytes per parameter (2 weights + 2 grads + 4 master + 8 Adam) = 112 GB plus activations, then show how LoRA removes most of the gradient/optimizer term and QLoRA quarters the weights. Candidates who also mention activations, gradient checkpointing and logits memory stand out.
Common pitfall Estimating memory from the weights alone ("7B in bf16 is 14 GB, so it fits on a 24 GB card") and then running out of memory at the first backward pass. Training needs gradients, optimizer states and activations as well; inference needs the KV cache.

The PEFT family: training a tiny fraction of the model

Parameter-efficient fine-tuning (PEFT) freezes the pretrained weights and trains a small number of new or selected parameters. It works because the change needed to adapt a large model to a task appears to live in a low-dimensional space: research on "intrinsic dimension" showed that many tasks can be learned by optimizing only a few hundred or thousand directions in parameter space. PEFT methods differ in where the trainable parameters are placed.

Analogy

You bought a well-built house and want it to suit your family. You could rebuild every wall (full fine-tuning). Or you could add a small extension in the hallway (adapters), leave notes on the front door that everyone reads on the way in (prompt and prefix tuning), install dimmer switches on existing lights (IA3 scaling), or fit a thin overlay on existing walls that shifts them slightly in a few directions (LoRA). The house stays intact and you can remove the changes later.

Mapping back: the house is the frozen pretrained network, each renovation style is a PEFT method that inserts trainable parameters at a different place, and "removing the changes" is detaching the adapter to recover the original model.

MethodWhat is trainedTrainable paramsInference overheadNotes
Adapters (bottleneck)Small MLP (down-project, nonlinearity, up-project) inserted after attention and/or FFN, with a residual0.5–5%Extra sequential layers add latencyThe original PEFT method for transformers
Prefix tuningTrainable key/value vectors prepended at every layer~0.1%Uses context lengthOften reparameterized with an MLP during training for stability
Prompt tuning (soft prompts)A few dozen trainable embedding vectors prepended to the input only<0.01%Uses context lengthCompetitive only for very large models
P-tuningSoft prompt tokens produced by a small encoder, possibly at several layersSmallUses context lengthVariant of prompt tuning
BitFitOnly bias terms~0.1%NoneSurprisingly decent for encoders; weak for LLM SFT
IA3Learned vectors that rescale keys, values and FFN activations~0.01%None once mergedExtremely small
LoRALow-rank matrices A and B added to chosen linear layers0.1–2%None once mergedThe de facto standard
QLoRALoRA on a 4-bit quantized frozen baseSame as LoRADequantization cost if served in 4-bitLargest memory savings
DoRALoRA on direction plus a trained magnitude vectorLoRA + tinyNone once mergedOften closes the gap to full fine-tuning at low rank
Selective / partialOnly the top k layers, or layer norms, or embeddingsVariesNoneSimple baseline; freezes lower layers

Additive, sequential (adapters)

  • New modules in the forward path
  • Cannot be folded into existing weights
  • Adds latency at batch size 1

Reparameterization (LoRA, DoRA, IA3)

  • Learn a change to existing weights
  • Can be merged: W' = W + ΔW
  • Zero extra latency after merging
Interview angle "Compare adapters, prefix tuning and LoRA." Mention where parameters live, whether they can be merged (only reparameterization methods like LoRA and IA3 can), inference latency, context-length cost for prefix/prompt tuning, and that LoRA won in practice because it is simple, mergeable and works across model sizes.
Common pitfall Thinking PEFT saves compute in the forward pass. Activations still flow through the full frozen model, and the backward pass still computes gradients with respect to activations at every layer. PEFT mainly saves memory (no gradients or optimizer states for frozen weights) and storage, and makes the optimizer step cheap; per-step compute is only moderately lower than full fine-tuning.

LoRA: low-rank adaptation in depth

LoRA (Low-Rank Adaptation) is based on a hypothesis: the update a pretrained weight matrix needs during fine-tuning has low intrinsic rank. So instead of learning a full dout × din update ΔW, LoRA learns two thin matrices whose product has rank at most r, typically 8 to 64, while the original weight stays frozen.

Analogy

An audio engineer wants to change the sound of a full orchestra recording. Rather than editing every instrument track, she adds a small equalizer with a handful of sliders. Each slider moves a whole family of frequencies together, and a master volume knob controls how strongly the equalizer applies. With just a few sliders she can give the recording a noticeably different character.

Mapping back: the orchestra tracks are the frozen weight matrix W, the small set of sliders is the low-rank pair B and A (the rank r is the number of sliders), each slider's broad effect is a rank-one direction, and the master knob is the scaling factor α/r.

The math

h = W0x + ΔW x = W0x + (α / r) · B A x where W0 ∈ ℜdout×din is frozen, A ∈ ℜr×din (the "down" projection) and B ∈ ℜdout×r (the "up" projection) are trainable, r ≪ min(din, dout) is the rank, and α is a scaling hyperparameter.
  • Parameter count per adapted matrix: r·(din + dout) instead of din·dout. For a 4096×4096 projection with r = 16: 131,072 versus 16.8M, a 128× reduction.
  • Initialization: A is random (Kaiming/Gaussian), B is zero. So at step 0, BA = 0 and the adapted model is exactly the pretrained model; training starts from a known-good function and is stable.
  • Why not both zero? If both were zero, the gradient with respect to each would be zero (each gradient depends on the other matrix), so nothing would ever learn.
  • Scaling α/r: keeps the update size roughly constant when you change r, so you do not need to retune the learning rate for every rank. A common convention is α = r or α = 2r. rsLoRA argues the right scaling is α/√r, which keeps learning stable at high ranks.
  • Merging: after training, compute W' = W0 + (α/r)BA once. The merged model has exactly the original architecture and zero added latency.
  • Dropout: LoRA dropout (for example 0.05) is applied to the input of the adapter path only, as light regularization.
            x (d_in)
            │
   ┌────────┴─────────┐
   │                  │
┌──┴───┐          ┌───┴───┐
│  W0  │ frozen   │   A   │ r × d_in   (random init)
│d_out │          └───┬───┘
│× d_in│              │  r-dim bottleneck
└──┬───┘          ┌───┴───┐
   │              │   B   │ d_out × r  (zero init)
   │              └───┬───┘
   │                  │ × (α / r)
   └──────── + ───────┘
             │
             h (d_out)

Why it works so well

Empirically, fine-tuning updates of large pretrained models have rapidly decaying singular values: most of the change is captured by a few directions. LoRA exploits this, and it also acts as a regularizer: a low-rank update cannot rewrite the whole matrix, which is part of why LoRA forgets less than full fine-tuning (a finding often summarized as "LoRA learns less and forgets less"). The flip side is that for large distribution shifts, such as teaching a new language or continued pretraining on code, full fine-tuning or high-rank LoRA learns more.

Choosing rank, alpha and target modules

KnobTypical valuesGuidance
Rank r8, 16, 32, 648–16 is usually enough for format/style SFT; 32–128 for harder skills or larger shifts. Doubling r doubles adapter params; gains often flatten quickly.
Alpha αr or 2r (16, 32)Effective scale is α/r. Keep the ratio fixed when sweeping rank to isolate capacity from step size.
Target modulesAll linear layers: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_projThe original paper adapted only q and v; later work (including QLoRA) found adapting all linear layers matters more than raising rank. PEFT accepts target_modules="all-linear".
Dropout0.0–0.10.05 is common; use more on tiny datasets.
Bias"none"Training biases adds little.
Learning rate1e-4 to 3e-4 (2e-4 typical)About 10–20× higher than full fine-tuning; the frozen base and zero-initialized B make large steps safe.
modules_to_saveembed_tokens, lm_headOnly when you add new tokens (for example a chat template for a base model); those rows must be trained fully.

LoRA in code with PEFT

from peft import LoraConfig, get_peft_model, TaskType

lora_cfg = LoraConfig(
    task_type=TaskType.CAUSAL_LM,
    r=16,
    lora_alpha=32,               # scale = alpha / r = 2
    lora_dropout=0.05,
    bias="none",
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                    "gate_proj", "up_proj", "down_proj"],
)
model = get_peft_model(base_model, lora_cfg)
model.print_trainable_parameters()
# e.g. for a 0.5B model with all-linear r=16:
# trainable params: ~8.8M || all params: ~503M || trainable%: ~1.75

A from-scratch LoRA layer

import math, torch, torch.nn as nn

class LoRALinear(nn.Module):
    def __init__(self, base: nn.Linear, r=16, alpha=32, dropout=0.05):
        super().__init__()
        self.base = base
        for p in self.base.parameters():
            p.requires_grad = False                     # freeze W0 (and bias)
        self.A = nn.Parameter(torch.empty(r, base.in_features))
        self.B = nn.Parameter(torch.zeros(base.out_features, r))   # zero init
        nn.init.kaiming_uniform_(self.A, a=math.sqrt(5))
        self.scale = alpha / r
        self.drop = nn.Dropout(dropout)

    def forward(self, x):
        return self.base(x) + self.scale * (self.drop(x) @ self.A.T @ self.B.T)

    @torch.no_grad()
    def merge(self):
        self.base.weight += self.scale * (self.B @ self.A)   # W' = W0 + (alpha/r) BA
        return self.base

Variants worth knowing

DoRA

Weight-Decomposed LoRA splits each weight into a magnitude vector and a direction: W' = m · (W0 + BA) / ‖W0 + BA‖c (column-wise norm). LoRA updates the direction, and m is trained separately. This mimics full fine-tuning's learning pattern more closely and often helps at low rank, at some extra training cost. Mergeable after training.

rsLoRA

Rank-stabilized LoRA uses scale α/√r instead of α/r so that higher ranks actually learn faster rather than being damped. Useful when sweeping large ranks.

LoRA+

Uses a larger learning rate for B than for A (for example 16×), which improves convergence speed because the two matrices play asymmetric roles.

AdaLoRA

Allocates rank adaptively across layers using an SVD-like parameterization and prunes unimportant singular values during training.

PiSSA / LoftQ

Smarter initialization: PiSSA initializes A and B from the top singular vectors of W; LoftQ initializes adapters to compensate for quantization error in QLoRA.

VeRA and tied variants

Share frozen random A and B across layers and train only small scaling vectors, shrinking adapters by another 10× or more.

Interview angle Expect "Explain LoRA mathematically and why B is initialized to zero", "What does alpha do?", "How many parameters does LoRA add to a 4096×4096 layer at rank 8?" (8 × 8192 = 65,536) and "Does LoRA add inference latency?" (not after merging; a small overhead if kept separate). Strong candidates also know that targeting all linear layers usually matters more than increasing rank.
Common pitfall Increasing r without adjusting α. With α fixed, the scale α/r shrinks as r grows, so a higher rank can look worse simply because its effective learning rate dropped. Sweep with a fixed α/r ratio (or use rsLoRA), and compare validation metrics rather than training loss.

QLoRA: fine-tuning on a 4-bit base

LoRA removed gradients and optimizer states for frozen weights, but the frozen weights themselves still occupy 2 bytes per parameter in bf16. QLoRA stores the frozen base model in 4-bit and trains ordinary 16-bit LoRA adapters on top. During each forward and backward pass, weights are dequantized block by block to the compute dtype, used in the matrix multiply and discarded. Gradients flow through the dequantized weights to the adapters, but the 4-bit weights never change. QLoRA introduced three ideas that make this work without losing quality.

Analogy

A translator works from a huge reference encyclopedia but has a small desk. She keeps the encyclopedia as a compact pocket edition with abbreviated entries, expanding each entry to full text only while she reads it, and she writes her own corrections in a thin, full-quality notebook next to it. If the desk briefly overflows, she moves a few notebook pages to a shelf behind her and brings them back when needed.

Mapping back: the pocket edition is the NF4-quantized base, expanding an entry is on-the-fly dequantization to bf16/fp16, the full-quality notebook is the 16-bit LoRA adapters that are actually trained, and moving pages to the shelf is the paged optimizer spilling optimizer state to CPU memory during spikes.

1. 4-bit NormalFloat (NF4)

Pretrained weights are roughly normally distributed around zero. A uniform 4-bit grid (INT4) wastes levels in the sparse tails and has too few near zero, where most weights sit. NF4 places its 16 levels at the quantiles of a standard normal distribution, rescaled to [−1, 1], so each level is used about equally often: an information-theoretically near-optimal code for normally distributed data. It includes an exact zero.

Quantization is blockwise: weights are split into blocks of 64, each block is scaled by its absolute maximum (absmax) so values fall into [−1, 1], and each value is mapped to the nearest NF4 level. Blockwise scaling limits the damage an outlier can do to one block rather than a whole tensor.

qi = argmink | NF4k − wi / cb | ,   cb = maxj∈b |wj| ,   ŵi = cb · NF4qi where b is the block containing weight i (64 weights), cb is the block's absmax scale, NF4k are the 16 normal-quantile levels, qi is the stored 4-bit index and ŵi the dequantized value used in compute.

2. Double quantization

Each block of 64 weights needs its own fp32 scale: 32 bits / 64 weights = 0.5 extra bits per parameter, which is 0.44 GB for a 7B model. Double quantization quantizes those scales too: the fp32 constants are grouped into blocks of 256 and quantized to 8-bit with one fp32 second-level constant per group. The overhead falls to about 8/64 + 32/(64·256) ≈ 0.127 bits per parameter, saving roughly 0.37 bits per parameter (about 0.3 GB on 7B, several GB on 65B).

3. Paged optimizers

Memory use during training is spiky, particularly with long sequences and gradient checkpointing. Paged optimizers allocate optimizer states in NVIDIA unified memory, which automatically pages them to CPU RAM when the GPU is about to run out and pages them back when needed, turning a crash into a small slowdown. In practice you select optim="paged_adamw_8bit" or "paged_adamw_32bit".

Compute dtype and the full recipe

  • Storage dtype is NF4; compute dtype is bf16 (or fp16 on GPUs without bf16 support). Matmuls always happen in 16-bit.
  • prepare_model_for_kbit_training casts layer norms (and usually the output head) to fp32 for stability, enables gradient checkpointing, and makes input embeddings require gradients so gradients can flow through a frozen, checkpointed network into the adapters.
  • Adapters on all linear layers were key to QLoRA matching 16-bit full fine-tuning in the original experiments; the headline result was fine-tuning a 65B model on a single 48 GB GPU.
  • Speed: QLoRA is typically 20–40% slower per step than LoRA because of dequantization, a trade of time for memory.
import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
from peft import prepare_model_for_kbit_training, get_peft_model, LoraConfig

bnb_cfg = BitsAndBytesConfig(
    load_in_4bit=True,                      # store frozen weights in 4-bit
    bnb_4bit_quant_type="nf4",              # NormalFloat4 levels (vs "fp4")
    bnb_4bit_use_double_quant=True,         # quantize the per-block scales too
    bnb_4bit_compute_dtype=torch.bfloat16,  # float16 on T4-class GPUs
)
base = AutoModelForCausalLM.from_pretrained(
    "your-org/your-7b-model", quantization_config=bnb_cfg, device_map="auto")
base = prepare_model_for_kbit_training(base, use_gradient_checkpointing=True)
model = get_peft_model(base, LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,
                                        target_modules="all-linear",
                                        task_type="CAUSAL_LM"))
LoRAQLoRA
Frozen base precisionbf16/fp16 (2 bytes)NF4 (≈0.53 bytes with double quant)
Trainable parameters at the same rankIdenticalIdentical
Peak memory, 7B≈16–20 GB≈6–10 GB
Step speedFasterSlower (dequantization)
QualityReferenceUsually within noise of LoRA
MergingStraightforwardDequantize base to 16-bit first, then merge

For the broader quantization picture (PTQ, GPTQ, AWQ, INT8/INT4 inference, quantization-aware training and on-device formats) see Edge AI. Quantization-aware training combined with LoRA is also used to recover accuracy for aggressively quantized on-device models.

Interview angle A standard probe is "Explain the three innovations of QLoRA." Answer: NF4 (quantile-based 4-bit levels suited to normally distributed weights, blockwise absmax scaling), double quantization (quantize the scaling constants, about 0.37 bits per parameter saved), and paged optimizers (unified memory absorbs spikes). Then explain that gradients flow through the dequantized frozen weights into 16-bit adapters, and that compute is always in 16-bit.
Common pitfall Merging a LoRA adapter directly into 4-bit weights, or merging into a differently quantized copy than the one used in training. Adding a small high-precision update to 4-bit values and re-rounding destroys much of it. The safe path is: load the base in 16-bit, merge the adapter, then (optionally) re-quantize with a proper method such as GPTQ, AWQ or GGUF k-quants, and evaluate the result.

Hyperparameters that matter

Fine-tuning has fewer knobs than pretraining, but the defaults differ by method, and a wrong learning rate is the most common reason a run either does nothing or destroys the model. Treat the table below as starting points, then tune on validation loss and, more importantly, on task metrics.

Analogy

Tuning a fine-tuning run is like adjusting a car for a new track. The learning rate is how hard you press the accelerator; warmup is easing out of the pit lane before flooring it; the schedule is braking smoothly into the finish; batch size is how many laps you average before adjusting the setup; epochs are how many practice sessions you run before you start memorizing the track instead of learning to drive.

Mapping back: too much accelerator (learning rate) and the car spins off (divergence or forgetting), no warmup and the tires lose grip at the start (unstable early updates while Adam's statistics are still poor), and too many sessions produce a driver who only knows one track (overfitting).

HyperparameterFull fine-tuningLoRA / QLoRADPO and friends
Learning rate5e-6 to 2e-5 (smaller for bigger models)1e-4 to 3e-45e-7 to 5e-6 (full), ~5e-6 to 5e-5 (LoRA)
Epochs1–31–3 (up to 5 on tiny sets)1–3
Effective batch size32–512 sequences16–12832–128
Warmup3–10% of steps3–10%~10%
ScheduleCosine or linear decayCosine or linear decay (constant also works)Linear or cosine
Weight decay0–0.10–0.010
Gradient clipping1.00.3–1.01.0
OptimizerAdamW (fused), 8-bit AdamWAdamW, paged 8-bit AdamWAdamW
OtherGradient checkpointing, ZeRO/FSDPr 8–64, α = r or 2r, dropout 0.05β = 0.1 (range 0.01–0.5)

The rules behind the numbers

  • Learning rate is the most important knob. Full fine-tuning moves pretrained weights directly, so it needs tiny steps to avoid wrecking learned features. LoRA starts from a zero update on a frozen base and can take larger steps.
  • Effective batch size = micro-batch × gradient-accumulation steps × number of GPUs. When you increase it substantially, raise the learning rate a bit (square-root scaling is a safe heuristic for Adam).
  • Warmup protects the first steps, when Adam's second-moment estimates are poor and a few large updates can damage the model.
  • Epochs: SFT data is often seen only 1–3 times. Many passes over a small set drive the training loss towards zero while validation loss rises, the classic overfitting sign, and the model starts repeating training answers verbatim.
  • Sequence length: set from data percentiles; longer costs memory linearly (with fused attention) and wastes compute on padding if most examples are short.
  • Steps, not epochs, are what matter for compute planning: steps per epoch = ⌈N / effective batch⌉.
  • Seed and reproducibility: fix seeds for Python, NumPy and PyTorch; small-data fine-tuning can vary noticeably between seeds, so compare methods over several seeds if differences are small.
lr(t) = lrmax · t / Tw  (t < Tw);   lr(t) = lrmax · ½[1 + cos(π · (t − Tw) / (T − Tw))]  (t ≥ Tw) Linear warmup over Tw steps followed by cosine decay to zero at the final step T.

Reading the loss curves

PatternLikely causeAction
Train and validation both flat from the startLR too low, adapters not attached, all params frozen, mask all zerosCheck trainable-parameter count and mask; raise LR
Loss spikes or becomes NaNLR too high, fp16 overflow, bad example, no clippingLower LR, use bf16 or fp32, clip gradients, inspect batch
Train falls, validation falls then risesOverfittingEarly stop at best validation checkpoint, fewer epochs, more data, more dropout
Initial loss already very lowModel already knows the data, or leakageCheck for duplicates between train and validation
Staircase drops each epochMemorization of repeated examplesUsually a warning sign; reduce epochs
Validation loss lower but outputs worseLoss is a proxy; verbose or bland outputsEvaluate with generation-based metrics and a judge
Interview angle "Why does LoRA use a learning rate around 2e-4 while full fine-tuning uses about 1e-5?" Because the base is frozen and B starts at zero, LoRA's update starts small and is constrained to a low-rank subspace, so larger steps are safe; full fine-tuning modifies every pretrained weight directly, and large steps quickly overwrite general features (forgetting) or diverge.
Common pitfall Selecting the final checkpoint by default. The last epoch is often past the validation minimum. Save checkpoints each epoch (or every N steps), keep the one with the best validation score, and actually restore it before evaluation.

The tooling landscape and distributed training

The open-source stack for fine-tuning is layered. At the bottom are PyTorch and GPU kernels; above that are model libraries; then training libraries that implement SFT and preference losses; then distributed-training backends; and finally configuration-driven wrappers. Knowing which layer each tool occupies makes them easy to compare.

Analogy

Building a house involves raw materials, prefabricated components, specialist trades, heavy machinery for big jobs, and a general contractor who coordinates everything from a plan. You can hire a contractor with a standard plan or manage the trades yourself for full control.

Mapping back: PyTorch and kernels are the raw materials, Transformers and PEFT are the prefabricated components, TRL's trainers are the specialist trades, DeepSpeed and FSDP are the heavy machinery for multi-GPU jobs, and YAML-driven tools like Axolotl are the general contractor working from a plan.

ToolLayerWhat it provides
Hugging Face TransformersModelsModel classes, tokenizers with chat templates, Trainer, generation, quantization integrations
PEFTAdaptersLoRA, DoRA, IA3, prefix/prompt tuning, adapter saving/loading, merging, multi-adapter management
TRLTraining recipesSFTTrainer (packing, completion-only loss, chat formatting), RewardTrainer, DPOTrainer, ORPOTrainer, KTOTrainer, PPOTrainer, GRPOTrainer
bitsandbytesKernels8-bit and 4-bit (NF4/FP4) quantized linear layers, 8-bit and paged optimizers
AccelerateLaunch/distributionOne script that runs on CPU, one GPU, many GPUs or many nodes; mixed precision; wraps DeepSpeed and FSDP
DeepSpeedDistributionZeRO stages 1–3, CPU/NVMe offload, pipeline parallelism
PyTorch FSDPDistributionNative fully sharded data parallelism (parameters, gradients and optimizer states sharded)
UnslothOptimized trainerHand-written kernels and memory tricks for faster, lower-memory LoRA/QLoRA, mainly single-GPU; notebooks for popular models; GGUF export
AxolotlConfig wrapperYAML-configured SFT/DPO/QLoRA across many models and distributed backends
LLaMA-Factory, torchtuneConfig wrappers / recipesSimilar config-driven recipes; torchtune is PyTorch-native
FlashAttention, Liger kernelsKernelsMemory-efficient attention; fused RMSNorm/RoPE/cross-entropy to cut memory
Hosted fine-tuning APIsManaged serviceUpload JSONL, receive a tuned model endpoint; limited control over method and hyperparameters, no weight access

Distributed strategies

DDP (data parallel)

Each GPU holds a full copy of model, gradients and optimizer; batches are split and gradients all-reduced. Simple and fast, but the model must fit on each GPU.

ZeRO-1

Shard optimizer states across GPUs. Memory per GPU for the 12-byte optimizer part divides by the number of GPUs.

ZeRO-2

Also shard gradients. Good default for LoRA or mid-size full fine-tuning.

ZeRO-3 / FSDP full shard

Also shard parameters; each layer is gathered just-in-time for compute. Lets you fully fine-tune models far larger than one GPU, at the cost of more communication.

Offload

Move optimizer states (and even parameters) to CPU RAM or NVMe. Big memory win, significant slowdown.

Tensor / pipeline parallel

Split individual layers (tensor) or groups of layers (pipeline) across GPUs. Needed for very large models; more common in pretraining frameworks.

With ZeRO-3 on 8 GPUs, the 112 GB of weights, gradients and optimizer states for full 7B fine-tuning becomes about 14 GB per GPU, leaving room for activations on 40–80 GB cards.

Minimal SFT with TRL

from datasets import load_dataset
from trl import SFTTrainer, SFTConfig
from peft import LoraConfig

ds = load_dataset("json", data_files={"train": "train.jsonl", "eval": "val.jsonl"})

cfg = SFTConfig(
    output_dir="out/sft-lora",
    num_train_epochs=2,
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,       # effective batch 16 per GPU
    learning_rate=2e-4,
    lr_scheduler_type="cosine",
    warmup_ratio=0.05,
    max_grad_norm=1.0,
    bf16=True,
    gradient_checkpointing=True,
    max_length=1024,                      # named max_seq_length in older versions
    packing=False,
    assistant_only_loss=True,             # loss on assistant turns only (recent versions)
    eval_strategy="epoch",
    save_strategy="epoch",
    load_best_model_at_end=True,
    logging_steps=10,
)
trainer = SFTTrainer(
    model="your-org/your-8b-instruct",
    args=cfg,
    train_dataset=ds["train"],            # records with a "messages" field
    eval_dataset=ds["eval"],
    peft_config=LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,
                           target_modules="all-linear", task_type="CAUSAL_LM"),
)
trainer.train()
trainer.save_model("out/sft-lora/adapter")
Tip Library argument names change between versions (for example max_seq_length became max_length, and completion-only masking moved from a separate collator to config flags). Pin versions in requirements.txt and check the installed version's documentation rather than copying snippets blindly.
Interview angle "How would you fully fine-tune a 70B model?" Talk about sharding (ZeRO-3 or FSDP) across multiple nodes, bf16 mixed precision, gradient checkpointing, fused attention, possibly CPU offload, and the arithmetic: 70B × 16 bytes ≈ 1.1 TB of states, so at least 16–32 80 GB GPUs before activations. Then say that for most use cases you would try QLoRA on a single 80 GB GPU first.
Common pitfall Combining a 4-bit bitsandbytes model with ZeRO-3 or FSDP without checking support. Quantized weights and parameter sharding interact in version-specific ways; QLoRA on multiple GPUs usually works with plain DDP or specific supported FSDP settings. Verify on a small model before launching a long job.

Alignment with RLHF

After SFT, a model imitates demonstrations, but imitation has limits: there are many acceptable answers, demonstrators vary in quality, and it is much easier for people to compare two answers than to write a perfect one. Reinforcement Learning from Human Feedback (RLHF) uses those comparisons. It was the method behind the first widely used chat assistants, where annotators preferred a 1.3B aligned model over a 175B base model, showing that alignment can matter more than size for perceived helpfulness.

Analogy

Training a new chef for a restaurant: first she follows the recipe book (SFT). Then a food critic tastes many pairs of dishes and says which one is better; over time an assistant learns to predict the critic's verdict (reward model). The chef then experiments, getting quick scores from the assistant, while the head chef insists she must not stray too far from the house style (KL penalty), because the assistant can be fooled by, say, piling on salt.

Mapping back: the recipe book is the SFT model, the critic is human labellers, the assistant is the learned reward model, experimenting with scores is PPO optimization, the house-style rule is the KL penalty to the reference model, and fooling the assistant with salt is reward hacking.

The three-stage pipeline

  1. Supervised fine-tuning. Train the base model on human-written demonstrations of good responses. This teaches format and produces the reference policy πref.
  2. Reward model training. Sample several responses per prompt from the SFT model; humans rank them. Train a reward model (usually the SFT model with a scalar head replacing the vocabulary head) to give preferred responses higher scores.
  3. RL optimization (PPO). Generate responses with the policy, score them with the reward model, and update the policy to increase reward while a KL penalty keeps it close to πref.
  prompts ──▶ [SFT model] ──▶ y1, y2, ... ──▶ humans rank ──▶ (x, y_w, y_l) pairs
                                                                 │
                                                                 ▼
                                                       [Reward model r_φ(x, y)]
                                                                 │ score
  prompts ──▶ [Policy π_θ] ──▶ response y ──────────────────────▶│
                  ▲                                               ▼
                  │   PPO update  ◀──  reward = r_φ(x,y) − β·KL(π_θ ‖ π_ref)
                  └──────────────────  (value model estimates advantages)

Reward model: Bradley–Terry loss

P(yw ≻ yl | x) = σ( rφ(x, yw) − rφ(x, yl) ) ,   LRM = − E [ log σ( rφ(x, yw) − rφ(x, yl) ) ] where yw is the preferred (chosen) response, yl the rejected one, rφ the scalar reward model and σ the logistic sigmoid. Only reward differences matter, so rewards are often normalized to mean zero.

The RL objective

maxθ   Ex∼D, y∼πθ(·|x) [ rφ(x, y) ] − β · DKL( πθ(·|x) ‖ πref(·|x) ) where πθ is the policy being trained, πref the frozen SFT model and β the KL coefficient. The KL term prevents reward hacking, keeps language fluent and limits drift from the SFT distribution.

PPO optimizes this with a clipped surrogate objective. For each generated token it computes the probability ratio ρ = πθ(a|s) / πold(a|s) and an advantage estimate A (from a learned value model using generalized advantage estimation), then maximizes:

LPPO = E [ min( ρ·A, clip(ρ, 1 − ε, 1 + ε)·A ) ] with ε ≈ 0.2. This is an objective to maximize, not a loss to minimize. Clipping is applied to the ratio ρ, then multiplied by A; writing clip(ρA, …) is wrong. When A > 0 the clip removes the incentive to raise ρ above 1+ε; when A < 0 it removes the incentive to push ρ below 1−ε. The full PPO loss also subtracts a value-function term and often adds an entropy bonus.

Classic PPO-based RLHF juggles four models in memory: the policy, the frozen reference, the reward model and the value (critic) model. That is expensive and notoriously finicky (sensitive to KL coefficient, reward normalization, generation length and many implementation details), which motivated the simpler methods in the next section.

RLAIF and Constitutional AI

Human labels are slow and costly (large programmes take weeks to months). RLAIF (RL from AI Feedback) replaces human preference labels with judgements from a capable LLM prompted with guidelines. Constitutional AI is a structured version: a written list of principles (the "constitution") is used first in a supervised phase where the model critiques and revises its own harmful responses, and then in an RL phase where an AI judge chooses between responses according to the principles. It scales harmlessness training with far less human labelling of harmful content, and makes the target values explicit and auditable.

Known failure modes of RLHF

  • Reward hacking: the policy finds outputs the reward model overrates (overly long answers, flattering language, certain phrases). Mitigations: KL penalty, reward-model ensembles, length penalties, periodic reward-model retraining.
  • Sycophancy: agreeing with the user because labellers tend to prefer agreeable answers.
  • Verbosity bias: longer answers win comparisons, so the policy learns to pad.
  • Mode collapse / reduced diversity: aligned models become less creative and less calibrated than base models.
  • Alignment tax: small drops on some capability benchmarks after alignment; mixing pretraining-style loss into RL (as some pipelines do) reduces it.
Interview angle "Walk me through RLHF." Give the three stages, write the Bradley–Terry loss and the KL-regularized objective, name the four models in PPO, and explain the KL penalty's role (prevents reward hacking and drift). Then mention costs and instability and how DPO and GRPO address them.
Common pitfall Saying "RLHF teaches the model new knowledge." It mainly reshapes which of the model's existing behaviours are preferred: tone, helpfulness, refusals, format. The capabilities come from pretraining; alignment elicits and filters them.

Beyond PPO: DPO, ORPO, KTO, GRPO and friends

Since 2023 a wave of methods has simplified preference learning. They fall into two groups: offline, RL-free methods that learn directly from a fixed preference dataset with a classification-like loss (DPO, IPO, ORPO, KTO, SimPO), and online RL methods that still sample from the current policy but drop expensive components (GRPO, RLOO). Online methods with verifiable rewards (unit tests, exact answers) power modern reasoning models.

Analogy

Learning to write better essays. PPO is having a tutor who predicts a teacher's grades and coaching you draft by draft. DPO is being handed a stack of essay pairs already marked "this one is better" and directly studying what distinguishes them, with no tutor. GRPO is writing eight drafts of the same essay, having an automatic checker score them, and learning from which of your own drafts beat your average.

Mapping back: the tutor is the reward and value models, the marked pairs are an offline preference dataset used by DPO, and comparing drafts against their own average is GRPO's group-relative advantage, which replaces the learned critic.

DPO: Direct Preference Optimization

DPO's insight is that the KL-regularized RLHF objective has a closed-form optimal policy: π*(y|x) ∝ πref(y|x)·exp(r(x,y)/β). Rearranging, the reward can be expressed through the policy itself: r(x,y) = β·log(π(y|x)/πref(y|x)) + const. Substituting this "implicit reward" into the Bradley–Terry loss eliminates the reward model and the RL loop.

LDPO = − E(x, yw, yl) [ log σ( β · [ log (πθ(yw|x) / πref(yw|x)) − log (πθ(yl|x) / πref(yl|x)) ] ) ] where πθ is the model being trained, πref a frozen copy (usually the SFT model), log-probabilities are summed over response tokens, and β (typically 0.1) controls how far the policy may move from the reference: smaller β allows more deviation.
  • Needs: a preference dataset and two models (policy + frozen reference). With LoRA, the reference is just the base with adapters disabled, so only one copy is held in memory.
  • Pros: simple, stable, cheap, no sampling during training.
  • Cons: offline, so it cannot explore beyond the dataset; can overfit to preferences (driving both chosen and rejected likelihoods down); sensitive to how far the data is from the policy's own outputs; can increase verbosity.
  • Metrics to watch: reward accuracy (fraction of pairs where implicit reward of chosen > rejected), reward margins, and the log-probabilities of chosen responses (if they fall steadily, the model may be degrading).
from trl import DPOTrainer, DPOConfig
# dataset rows: {"prompt": [...messages...], "chosen": [...], "rejected": [...]}
cfg = DPOConfig(output_dir="out/dpo", beta=0.1, learning_rate=5e-6,
                per_device_train_batch_size=2, gradient_accumulation_steps=8,
                num_train_epochs=1, warmup_ratio=0.1, bf16=True,
                max_length=1024)
trainer = DPOTrainer(model=sft_model, ref_model=None,   # None + PEFT: base acts as reference
                     args=cfg, train_dataset=pref_ds,
                     processing_class=tokenizer, peft_config=lora_cfg)
trainer.train()

The rest of the family

MethodDataReference model?Key idea
IPOPairsYesReplaces the log-sigmoid with a squared loss towards a target margin, preventing DPO's tendency to push margins to infinity on deterministic preferences.
ORPOPairsNoCombines SFT and preference learning in one stage: SFT loss on chosen plus an odds-ratio term that favours the odds of chosen over rejected. No separate SFT stage and no reference model.
KTOUnpaired thumbs up/downYesInspired by prospect theory; learns from individual examples labelled desirable or undesirable, which is how production feedback is usually collected.
SimPOPairsNoUses length-normalized average log-probability as the implicit reward plus a target margin; reference-free and reduces length exploitation.
CPO, RRHF, SLiCPairs / rankingsVariesRanking or contrastive losses over candidate responses.
Online / iterative DPOFresh pairs from the current policy, labelled by a reward model or judgeYesReduces the off-policy gap of standard DPO by regenerating preference data each round.
RLOOPrompts + rewardYes (KL)REINFORCE with a leave-one-out baseline from k samples; no value model.
GRPOPrompts + reward functionYes (KL)Samples a group of G responses per prompt and uses group-normalized rewards as advantages; no critic.
ORPO:   L = LSFT(yw) − λ · log σ( log [oddsθ(yw|x) / oddsθ(yl|x)] ),   odds(y|x) = p(y|x) / (1 − p(y|x)) where p is the length-normalized sequence likelihood and λ weights the preference term.

GRPO and RL with verifiable rewards

Group Relative Policy Optimization keeps PPO's clipped objective and KL control but removes the value model. For each prompt, sample G responses (for example 8–16), score each with a reward function, and set each response's advantage to its reward standardized within the group. Responses better than their siblings are reinforced; worse ones are discouraged. This halves memory relative to PPO and works especially well when rewards are verifiable: a math answer matches the reference, code passes unit tests, output parses as valid JSON. Training on such rewards (often called RLVR) is how models learn long chain-of-thought reasoning, with behaviours like self-checking emerging from reward alone.

Ai = ( ri − mean(r1..G) ) / std(r1..G) ,   LGRPO = E [ (1/G) Σi (1/|yi|) Σt min( ρi,tAi, clip(ρi,t, 1−ε, 1+ε)Ai ) ] − β·DKL(πθ ‖ πref) where ri is the reward of the i-th sampled response, Ai is shared by every token in that response, and ρi,t is the per-token probability ratio against the sampling policy. If all G responses get the same reward, advantages are zero and the prompt teaches nothing, so prompt difficulty must be matched to the model. Some implementations add the KL term to the reward before standardizing instead of subtracting it as a separate loss.
from trl import GRPOTrainer, GRPOConfig

def correctness_reward(completions, answer, **kwargs):
    # 1.0 if the final line contains the reference answer, else 0.0
    return [1.0 if a in c[-1]["content"].splitlines()[-1] else 0.0
            for c, a in zip(completions, answer)]

def format_reward(completions, **kwargs):
    return [0.2 if "<answer>" in c[-1]["content"] else 0.0 for c in completions]

cfg = GRPOConfig(output_dir="out/grpo", num_generations=8, learning_rate=1e-6,
                 max_completion_length=512, beta=0.04, bf16=True)
trainer = GRPOTrainer(model=sft_model, reward_funcs=[correctness_reward, format_reward],
                      args=cfg, train_dataset=math_prompts)   # rows: prompt, answer
trainer.train()

GRPO versus PPO

This comparison is a favourite interview follow-up. Both maximize a clipped probability-ratio objective; they differ in how the advantage is estimated and what sits in GPU memory.

PPO (classic RLHF)GRPO
AdvantageLearned value (critic) + GAE, typically per tokenGroup-normalized reward: (r − mean)/std over G siblings; same A for every token in a response
Models in memoryPolicy, frozen reference, reward model, value model (four)Policy and frozen reference (two); the scorer is often a cheap verifier, not a neural RM
Samples per promptUsually one (or a few) on-policy rolloutA group of G (commonly 8–16)
Clip and KLYes: clip(ρ, 1−ε, 1+ε) and a KL penalty to πrefSame clip; KL either as a loss term or folded into the reward before grouping
Best fitSubjective quality where you already have (or will train) a reward modelVerifiable rewards: math answers, unit tests, schema checks (RLVR)
Typical failureReward hacking, critic instability, KL/length sensitivityZero-variance groups (all right or all wrong) give no gradient

RLOO is the close cousin: it also drops the critic, but uses a leave-one-out mean of the other samples as the baseline instead of a group z-score. Use PPO when you need a learned reward for fuzzy preferences; use GRPO/RLOO when a checker can score each sample.

Choosing a method

SituationReasonable choice
Have good demonstrations onlySFT
Have chosen/rejected pairs, want simplicitySFT then DPO (or ORPO in one stage)
Only thumbs-up/down logs from productionKTO
Answers can be checked automatically (math, code, schema)GRPO / RLOO with verifiable rewards
Subjective quality, large budget, strong RL teamReward model + PPO, or online DPO with a reward model
Need harmlessness at scale without human red-team labelsConstitutional AI / RLAIF
Interview angle Interviewers often ask you to derive or at least explain DPO: start from the KL-regularized objective, note its closed-form optimum, invert it to express reward in terms of the policy/reference log-ratio, and plug into Bradley–Terry; the partition function cancels because it depends only on x. Then compare: DPO is offline and cheap; PPO/GRPO are online and can explore; GRPO removes the critic by using group statistics.
Common pitfall Running DPO on a model that was never SFT-ed on data similar to the preference pairs, or on pairs generated by a very different model. DPO works best when chosen and rejected responses are plausible samples from the policy being trained; large distribution gaps lead to unstable training and falling likelihoods for both responses.

Evaluating a fine-tuned model

A lower loss is not the goal; better behaviour on the target task without breaking everything else is. Evaluation for fine-tuning therefore has two halves: did it learn the task? (held-out task metrics, human or judge ratings) and did it keep what it had? (general benchmarks, safety, regression tests). Always evaluate the base model with its best prompt on the same sets, so improvements are measured against a fair baseline. The broader methodology is covered in LLM Evaluation & Safety.

Analogy

A hospital hiring a newly specialised surgeon checks two things: can she perform the new procedure well (a practical exam on unseen cases), and has she still got the general skills every doctor needs (routine competency checks)? A great specialty score is worthless if she has forgotten basic hygiene.

Mapping back: unseen cases are the held-out test set, the practical exam is task metrics and judge ratings, and the routine competency checks are general benchmarks, safety prompts and regression tests that detect catastrophic forgetting.

Layers of evaluation

LayerWhat it measuresExamples
Training signalsOptimization healthTrain/validation loss, perplexity (eloss), gradient norm
Task metrics on held-out dataDid it learn the task?Accuracy, F1, exact match, JSON validity rate, pass@k for code, ROUGE-L, BERTScore
LLM-as-judgeOpen-ended qualityPairwise win rate vs base model, rubric scores for correctness, helpfulness, safety
Human evaluationGround truth for subjective qualityBlind pairwise comparisons by domain experts on a sample
General benchmarksForgettingMMLU-style knowledge, GSM8K-style math, HumanEval-style coding, IFEval-style instruction following, MT-Bench / Arena-style chat quality
Safety and red-teamingSafety regressionHarmful-request refusal rate, jailbreak suites, over-refusal on benign prompts, toxicity, bias probes
Regression testsProduct behaviourA fixed suite of real prompts with assertions (format, must-mention, must-not-say), run in CI for every new checkpoint
Online metricsReal impactA/B tests: resolution rate, thumbs-up rate, escalations, latency, cost

Lexical vs semantic metrics

ROUGE-L

  • Longest common subsequence of words between prediction and reference
  • Rewards saying the same words in the same order
  • Penalizes valid paraphrases
  • Cheap, deterministic

BERTScore-F1

  • Greedy matching of contextual token embeddings
  • Rewards similar meaning with different wording
  • Scores are compressed into a narrow high range; compare relatively
  • Can miss factual errors that use similar words

Reporting both gives a fuller picture: if ROUGE rises but BERTScore does not, the model may just be copying phrasing; if BERTScore rises with flat ROUGE, it paraphrases correctly. When references are themselves machine-generated or noisy, weight relative differences between methods more than absolute values.

LLM-as-judge done right

  • Use a pairwise format (model A vs model B) with a clear rubric; pairwise judgements are more reliable than absolute 1–10 scores.
  • Swap positions and average to cancel position bias; control for verbosity bias (judges favour longer answers) and self-preference (judges favour their own family's style).
  • Validate the judge against a few hundred human labels and report agreement.
  • Use a judge that is stronger than, and different from, the model being evaluated.

Evaluation hygiene

  1. Same prompts, same decoding for every model (for example greedy, same max_new_tokens), and decode only the newly generated tokens.
  2. Use the chat template at evaluation exactly as in training.
  3. Check contamination: n-gram or embedding overlap between training data and test sets.
  4. Error analysis: read failures, bucket them (too short, too long, hallucinated fact, missing caveat, wrong format), and compute simple statistics such as the prediction/reference length ratio and counts of empty or truncated outputs.
  5. Report cost alongside quality: trainable parameters, peak memory, training time, inference latency.
  6. Several seeds when differences are small, with confidence intervals (bootstrap over test examples).
Interview angle "How do you know your fine-tune worked?" A strong answer: a held-out test set from the same distribution, task metrics plus a validated LLM judge, comparison against the best-prompted base model, general and safety benchmarks to detect forgetting, a regression suite in CI, and finally an online A/B test. Mention that training loss is not a quality metric.
Common pitfall Evaluating only on the target task. Many fine-tunes look great on their own test set while silently losing instruction following, multilingual ability, or safety refusals. Always run the same regression and safety suite on the base and the fine-tuned model.

Risks and failure modes

Fine-tuning is powerful precisely because it changes the model, and every change can break something. The main risks are losing general capabilities, memorizing instead of generalizing, eroding safety behaviour, leaking data, and violating licences. Each has known mitigations.

Analogy

A pianist who practises only one concerto for months may play it flawlessly yet stumble on pieces she once knew, may start playing that concerto's phrasing everywhere, and may even forget the etiquette of performing. If she learned the piece from a pirated score, she may also be barred from performing it publicly.

Mapping back: stumbling on old pieces is catastrophic forgetting, over-applied phrasing is overfitting, forgotten etiquette is safety regression, and the pirated score is a licensing problem with training data or base-model terms.

Catastrophic forgetting

Gradient updates for the new task overwrite weights that encoded older skills. Signs: general benchmark drops, loss of multilingual ability, broken chat format, worse reasoning. Mitigations: LoRA instead of full fine-tuning, lower LR, fewer epochs, mix 5–20% general instruction data (replay), keep the chat template, regularize towards the original weights (L2-SP, EWC), or merge the fine-tuned model back towards the base.

Overfitting and memorization

Small data plus many epochs yields verbatim regurgitation, repetitive phrasing and poor generalization. Mitigations: early stopping on validation, more diverse data, dropout, lower rank, deduplication.

Safety regression

Research has shown that fine-tuning even on a handful of harmful examples, and sometimes on entirely benign data, can weaken a model's refusal behaviour. Mitigations: include safety examples in the mix, run refusal and jailbreak suites before release, add inference-time guardrails, keep adapters small.

Hallucination from new facts

Training on facts the model did not know teaches it to answer confidently beyond its knowledge. Keep facts in RAG; use fine-tuning for behaviour; include "I don't know" examples.

Data leakage and privacy

Personal information, secrets or customer data in training sets can be extracted later through prompting (membership-inference and extraction attacks). Scrub data, restrict access to tuned weights, consider differential privacy for sensitive data.

Evaluation leakage

Test examples or benchmark items present in training data inflate scores. Decontaminate and split by group.

Licence and terms issues

Base-model licences range from permissive (Apache 2.0, MIT) to custom community licences with usage restrictions and attribution requirements. Datasets have their own licences. Many API providers' terms prohibit using outputs to train competing models. Record provenance for every data source.

Reward hacking and sycophancy

In preference tuning, the model exploits weaknesses of the reward or judge (length, flattery). Monitor length, use KL control, diverse judges, and human spot checks.

Template and tokenizer mismatch

Training with one chat format and serving with another, or adding tokens without resizing embeddings, silently degrades quality. Test the exact serving path end to end.

Data poisoning and backdoors

Malicious examples can implant trigger phrases that cause specific behaviour. Vet data sources, especially scraped or crowd-sourced data, and scan third-party adapters before use.

A mitigation recipe for forgetting

  1. Measure first. Build a general regression suite and score the base model.
  2. Prefer PEFT. LoRA with moderate rank; the base weights stay intact and the adapter can be scaled down (for example multiply the LoRA scale by 0.5) if it is too strong.
  3. Mix data. Add general instruction and chat examples (replay) alongside the domain data.
  4. Train gently. Lower LR, 1–2 epochs, early stopping.
  5. Consider model merging. Interpolating fine-tuned and base weights often recovers general ability with a small loss of task gains.
Interview angle "What is catastrophic forgetting and how do you prevent it?" Define it (new-task gradients overwrite old capabilities because the same weights are shared), give detection (regression benchmarks against the base) and at least four mitigations (PEFT, replay/data mixing, lower LR and fewer epochs, regularization such as EWC, model merging). Mention the safety-regression finding as a special case.
Common pitfall Assuming a fine-tuned open-weights model inherits the base model's safety behaviour. Safety alignment is itself learned and can be undone by later training. Re-run safety evaluations on every fine-tuned release.

Merging adapters and merging models

"Merging" means two different things. Adapter merging folds a LoRA update into its base weights so the result is a normal model. Model merging combines several fine-tuned models (or adapters) that share a base into one model with the skills of each, using arithmetic on weights and no further training. Both are cheap and widely used.

Analogy

Think of transparent overlays on a map. Adapter merging is permanently printing one overlay onto the map so you no longer need to carry it separately. Model merging is taking overlays drawn by different people on copies of the same base map (one marks hiking trails, another marks restaurants) and combining them onto one sheet, resolving places where they disagree.

Mapping back: the base map is the shared pretrained model, each overlay is a weight delta (a task vector or LoRA update), printing is merge_and_unload, and resolving disagreements is what TIES and DARE do with conflicting parameter updates.

Adapter merging

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained("your-org/base-model",
                                            torch_dtype=torch.bfloat16)   # 16-bit, not 4-bit
model = PeftModel.from_pretrained(base, "out/sft-lora/adapter")
merged = model.merge_and_unload()           # W' = W0 + (alpha/r) * B @ A, adapters removed
merged.save_pretrained("out/merged", safe_serialization=True)
AutoTokenizer.from_pretrained("out/sft-lora/adapter").save_pretrained("out/merged")
  • Always merge into a 16-bit copy of the exact base used in training (same revision).
  • Save the tokenizer and chat template with the merged model.
  • PEFT can also combine several LoRAs into one adapter (add_weighted_adapter with methods such as linear, SVD, TIES or DARE), or keep many adapters loaded and switch with set_adapter.

Model merging methods

MethodHow it worksWhen to use
Linear / model soupWeighted average of weights of models fine-tuned from the same baseAveraging runs with different seeds or hyperparameters; often improves robustness
SLERPSpherical interpolation between two models' weights, preserving norm geometryBlending exactly two models smoothly
Task arithmeticTask vector τ = θft − θbase; add λτ to gain a skill, subtract to remove a behaviourCombining or negating skills
TIESTrim small deltas, elect a sign per parameter by majority mass, average only agreeing deltasMerging several task models with interference
DARERandomly drop a large fraction (for example 90%) of delta parameters and rescale the rest by 1/(1−p)Pre-processing before TIES or linear merges; reduces interference
Frankenmerges / passthroughStack layers from different models into a deeper modelExperimental; unpredictable
θmerged = θbase + Σk λk · τk ,   τk = θk − θbase Task arithmetic: each fine-tuned model k contributes its task vector τk with weight λk. Merging is only meaningful for models that share the same base and architecture.

Configuration-driven tools such as mergekit implement these methods. Merged models must be evaluated like any other: merges can silently lose instruction following or produce format mixtures.

Interview angle "You have separate LoRA adapters for SQL generation and for customer tone. How do you get both?" Options: serve both and route per request; train one adapter on combined data; or merge them (weighted combination, TIES/DARE) and evaluate each skill afterwards. A good answer notes that merging is fast but can interfere, while joint training is usually most reliable.
Common pitfall Trying to merge models fine-tuned from different bases, or from different versions of the same base. Weight arithmetic assumes the parameters live in the same coordinate system; otherwise the result is garbage.

Knowledge distillation

Distillation trains a small student model to mimic a large teacher. For LLMs it is the main way to obtain small models that behave far above their size: the student learns from the teacher's outputs (and ideally its full probability distribution), which carry much more information per example than one-hot labels.

Analogy

A master chess coach does not just tell a student the best move; she explains that two other moves were almost as good and one looked tempting but loses. The student who hears "move A 60%, move B 30%, move C 10%" learns the coach's judgement far faster than one who only hears "play A".

Mapping back: the coach is the teacher model, the percentages are its soft probability distribution over next tokens ("dark knowledge"), and learning from the full distribution rather than just the top choice is logit distillation, versus sequence-level distillation that trains on the teacher's final text.

Forms of distillation

FormSignalNeeds teacher logits?Notes
Sequence-level (black-box)Teacher-generated text used as SFT targetsNoWorks with API teachers; most synthetic-data fine-tuning is this
Logit / token-level (white-box)KL divergence between teacher and student next-token distributionsYes (same tokenizer, or alignment)Richer signal; common in open model families
On-policy distillation (GKD)Student generates, teacher scores each tokenYesFixes the train/inference mismatch of learning only on teacher text
Reasoning distillationTeacher's chain-of-thought traces as targetsNoTransfers long reasoning into small models effectively
Feature / hidden-stateMatch intermediate representationsYesUsed for encoders (for example compact BERT variants)
LKD = (1 − λ) · CE(y, ps) + λ · T2 · KL( softmax(zt/T) ‖ softmax(zs/T) ) where zt, zs are teacher and student logits, T is the temperature (higher T softens distributions and exposes relative preferences among wrong answers), λ mixes hard-label cross-entropy and distillation loss, and T2 keeps gradient magnitudes comparable across temperatures. Forward KL makes the student cover all teacher modes; reverse KL makes it focus on the main mode, often preferred for generation.
import torch.nn.functional as F

def kd_loss(student_logits, teacher_logits, labels, mask, T=2.0, lam=0.5):
    # logits: [B, S, V]; labels: [B, S]; mask: [B, S] (1 on answer tokens)
    ce = F.cross_entropy(student_logits.transpose(1, 2), labels, reduction="none")
    kl = F.kl_div(F.log_softmax(student_logits / T, -1),
                  F.log_softmax(teacher_logits / T, -1),
                  log_target=True, reduction="none").sum(-1) * (T * T)
    per_tok = (1 - lam) * ce + lam * kl
    return (per_tok * mask).sum() / mask.sum().clamp(min=1)
Interview angle "Distillation vs fine-tuning vs quantization?" Distillation creates a smaller architecture that imitates a teacher; fine-tuning adapts a model's behaviour; quantization keeps the same architecture but stores weights in fewer bits. They combine: distil into a small model, fine-tune it for the product, then quantize it for deployment.
Common pitfall Distilling from a commercial API's outputs without reading its terms. Many providers forbid using outputs to develop competing models. Model-extraction via repeated querying is also a recognized security threat to deployed models, which is why providers monitor for it.

Serving fine-tuned models

After training you have a base model plus an adapter (or a full checkpoint). Deployment choices trade off latency, memory, the number of variants you need, and target hardware. The main decisions are: merge or keep adapters separate, how to serve many variants, and how to quantize for cost or edge deployment.

Analogy

A camera shop can sell each customer a camera with a lens permanently glued on, or sell one camera body with a bag of interchangeable lenses. Glued lenses are simplest to use; interchangeable lenses let one body serve a portrait photographer in the morning and a wildlife photographer in the afternoon, at the cost of a moment to swap.

Mapping back: the glued lens is a merged model (no overhead, one task per copy), the camera body is the shared base model in GPU memory, the lenses are LoRA adapters of a few megabytes each, and the swap time is the small per-request cost of applying a different adapter.

Merged weights

  • Zero added latency; standard architecture
  • Works with any inference engine and any quantization tool
  • One full model copy per fine-tune (e.g. 14 GB each at 7B bf16)
  • Best for one or a few variants at high traffic

Separate adapters (multi-LoRA)

  • One base in memory, hundreds of adapters of tens of MB
  • Small compute overhead for the extra low-rank matmuls
  • Per-request adapter selection, per-tenant customization
  • Best for many low-traffic variants

Multi-LoRA serving

Modern inference engines batch requests for different adapters together. The base matmul is shared across the batch, and specialized kernels (segmented or gathered batched matmuls) apply each request's own A and B. Adapters live in a CPU/GPU cache and are paged in on demand. Engines such as vLLM (with LoRA enabled at startup), dedicated multi-LoRA servers, and research systems like S-LoRA and Punica follow this design; managed inference products expose the same idea. Constraints to know: a maximum rank per server, a maximum number of adapters resident on GPU at once, and all adapters must share the same base.

# serve a base model with adapters that clients select by name (vLLM CLI)
vllm serve your-org/base-8b-instruct \
    --enable-lora \
    --lora-modules support=./adapters/support sql=./adapters/sql \
    --max-lora-rank 64 --max-loras 8

# request the "sql" adapter through the OpenAI-compatible endpoint
curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" \
  -d '{"model": "sql", "messages": [{"role": "user", "content": "Top 5 customers by revenue?"}]}'

Quantizing after fine-tuning

Format / methodTypical bitsWhere it runsNotes
bf16/fp1616GPU serversReference quality
FP88Recent data-centre GPUsNear-lossless, about 2× memory saving
GPTQ4 (or 3/8)GPUPost-training quantization with second-order error correction using a calibration set
AWQ4GPUProtects salient weight channels based on activation statistics
bitsandbytes NF44GPUEasy, used in QLoRA; slower inference than GPTQ/AWQ kernels
GGUF (k-quants such as Q4_K_M, Q5_K_M, Q8_0)2–8CPU, Apple silicon, consumer GPUs, edgeSingle-file format for llama.cpp-family runtimes and local model runners

Use a calibration set that resembles your fine-tuning domain (not generic web text) when quantizing a specialised model, and re-run your regression suite: quantization errors hit rare, domain-specific behaviours first.

GGUF export for local and edge deployment

  1. Merge the adapter into a 16-bit base and save with the tokenizer.
  2. Convert to a 16-bit GGUF file with the llama.cpp conversion script.
  3. Quantize to the target type (Q4_K_M is a common balance of size and quality; Q8_0 for near-lossless).
  4. Run and evaluate with the same chat template; check that the template embedded in the GGUF metadata is correct.
# 1) merged model already saved in out/merged (see the merging section)
# 2) convert Hugging Face checkpoint to GGUF (16-bit)
python llama.cpp/convert_hf_to_gguf.py out/merged --outfile model-f16.gguf --outtype f16
# 3) quantize to 4-bit k-quant
./llama.cpp/build/bin/llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M
# 4) quick interactive test
./llama.cpp/build/bin/llama-cli -m model-Q4_K_M.gguf -cnv
# alternatively keep the adapter separate: convert it with convert_lora_to_gguf.py
# and load it at runtime next to the base GGUF (--lora adapter.gguf)

On-device stacks take this further: a shared base model ships once with the operating system or app, and small task-specific LoRA adapters (tens of MB) are hot-swapped per feature, while quantization-aware training plus LoRA recovers accuracy lost to aggressive 4-bit quantization. See Edge AI for runtimes, NPU constraints and memory budgets.

Interview angle "You have 200 enterprise customers, each wanting a customized model. How do you serve them?" Answer: one shared base, one LoRA per customer, multi-LoRA serving with adapter caching and batched heterogeneous-adapter kernels, per-tenant routing, and merged dedicated deployments only for the few highest-traffic customers. Mention isolation, versioning of adapters and a regression suite per adapter.
Common pitfall Serving with a different chat template, tokenizer or stop tokens than were used in training (for example a GGUF with a missing template, or a server that stops on a different EOS). Symptoms are rambling, repeated role headers, or answers that start with template debris.

End-to-end walkthrough: full fine-tuning vs LoRA vs QLoRA

This worked example puts everything together on a realistic small-scale study: adapt a 0.5B-parameter instruct model to answer medical instructions, training the same model on the same data three ways (full instruction fine-tuning, LoRA and QLoRA, the last two at ranks 8 and 16) on a single 16 GB T4-class GPU, then compare quality against cost. It uses a hand-written PyTorch loop rather than a high-level trainer so every mechanism is visible. Swap in a 7B–8B model and more data and the same code is a production-shaped QLoRA pipeline.

Analogy

It is a controlled cooking experiment: the same ingredients (data), the same oven (GPU) and the same recipe card (hyperparameters), with only one variable changed at a time, first how much of the dish you are allowed to modify, then whether the base stock is stored concentrated. Only then can you say which change caused a difference in taste or cost.

Mapping back: the ingredients are the fixed train/validation/test split, the variable being changed is the trainable-parameter scope (full, LoRA) and then the backbone precision (16/32-bit vs 4-bit), and "taste vs cost" is ROUGE-L/BERTScore against trainable %, peak memory and training time.

 data (instruction, input, output)
   │ clean ("<noinput>" → ""), shuffle(seed), split 300 / 30 / 30
   ▼
 chat template ──▶ tokenize once with offsets ──▶ completion mask (answer = 1)
   ▼
 ┌──────────────┬──────────────────────┬─────────────────────────────┐
 │ Full IFT     │ LoRA r=8, r=16       │ QLoRA r=8, r=16             │
 │ fp32, lr 1e-5│ fp32 base, lr 2e-4   │ NF4 base, fp16 compute,     │
 │ 100% trained │ adapters on 7 linears│ same adapters, lr 2e-4      │
 └──────┬───────┴──────────┬───────────┴──────────────┬──────────────┘
        └────────── shared loop: masked loss, accumulation, warmup+cosine,
                    clipping, per-epoch validation, keep best checkpoint
                                   ▼
          greedy generation on test ──▶ ROUGE-L + BERTScore ──▶ error analysis
                                   ▼
            cost table: % trainable, peak GPU memory, training time

Step 1: environment, seed and configuration

# pip install -q transformers datasets peft "bitsandbytes>=0.43" evaluate rouge_score bert_score
import gc, math, random, time
import numpy as np, pandas as pd, torch
import torch.nn.functional as F
from torch.utils.data import Dataset, DataLoader
from datasets import load_dataset
from transformers import (AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig,
                          get_cosine_schedule_with_warmup, set_seed)
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training, TaskType

SEED = 42
def seed_everything(seed=SEED):
    random.seed(seed); np.random.seed(seed)
    torch.manual_seed(seed); torch.cuda.manual_seed_all(seed); set_seed(seed)
seed_everything()
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"

MODEL_ID   = "Qwen/Qwen2.5-0.5B-Instruct"     # any small instruct model works
N_TRAIN, N_VAL, N_TEST = 300, 30, 30          # tiny on purpose: fits one session
MAX_LEN    = 256                              # chosen from the token-length p95
BATCH_SIZE, GRAD_ACCUM, EPOCHS = 4, 1, 3      # effective batch = 4
WARMUP_RATIO, MAX_GRAD_NORM = 0.05, 1.0
LR_FULL, LR_LORA = 1e-5, 2e-4                 # full FT needs a much smaller LR
LORA_DROPOUT = 0.05
TARGETS = ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]

Step 2: load, clean, split and inspect the data

The dataset has three fields: instruction, an optional input with extra context (stored as a literal <noinput> placeholder when absent) and output, a reference answer that was itself generated by an older LLM. That last point matters for evaluation: the references are a noisy proxy for truth, so relative comparisons between methods are more meaningful than absolute scores.

df = pd.DataFrame(load_dataset("your-org/medical-instructions")["train"])
df["input"] = df["input"].replace("<noinput>", "").fillna("")
df = df.sample(frac=1, random_state=SEED).reset_index(drop=True)   # shuffle once
train_df = df.iloc[:N_TRAIN]
val_df   = df.iloc[N_TRAIN:N_TRAIN + N_VAL]
test_df  = df.iloc[N_TRAIN + N_VAL:N_TRAIN + N_VAL + N_TEST]

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
if tokenizer.pad_token is None:
    tokenizer.pad_token = tokenizer.eos_token

lens = train_df.apply(lambda r: len(tokenizer(r["instruction"] + " " + r["input"] + " " + r["output"],
                                              add_special_tokens=False)["input_ids"]), axis=1)
print("rows with input:", (train_df["input"] != "").mean().round(2))
print(lens.describe(percentiles=[0.5, 0.95]))   # choose MAX_LEN near the p95
Tip The chat template itself adds roughly 30–40 tokens (system prompt and role markers), so set MAX_LEN to the p95 of content length plus template overhead, or you will truncate the longest answers, which are exactly the ones that teach the model to finish properly.

Step 3: chat-template prompts and an offset-based completion mask

SYSTEM_MSG = "You are a helpful medical assistant. Answer accurately and safely."

def build_prompt(instruction, input_text):
    user = f"{instruction}\n{input_text}" if input_text else instruction
    msgs = [{"role": "system", "content": SYSTEM_MSG},
            {"role": "user", "content": user}]
    # add_generation_prompt=True appends the assistant header so the model answers next
    return tokenizer.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)

class SFTDataset(Dataset):
    def __init__(self, frame):
        self.rows = [{"prompt": build_prompt(r["instruction"], r["input"]),
                      "answer": r["output"].strip()} for _, r in frame.iterrows()]
    def __len__(self): return len(self.rows)
    def __getitem__(self, i): return self.rows[i]

def tokenize_example(prompt, answer):
    """Tokenize prompt+answer ONCE; mark answer tokens using character offsets."""
    full = prompt + answer + tokenizer.eos_token          # end-of-turn so the model learns to stop
    enc = tokenizer(full, return_offsets_mapping=True, add_special_tokens=False,
                    truncation=True, max_length=MAX_LEN)
    start = len(prompt)                                    # first answer character
    comp_mask = [1 if s >= start else 0 for s, e in enc["offset_mapping"]]
    return enc["input_ids"], comp_mask

def collate(batch):
    ids, masks = zip(*(tokenize_example(b["prompt"], b["answer"]) for b in batch))
    L = max(len(x) for x in ids)
    input_ids = torch.full((len(ids), L), tokenizer.pad_token_id, dtype=torch.long)
    attn      = torch.zeros((len(ids), L), dtype=torch.long)
    comp_mask = torch.zeros((len(ids), L), dtype=torch.long)
    for i, (x, m) in enumerate(zip(ids, masks)):
        input_ids[i, :len(x)] = torch.tensor(x)
        attn[i, :len(x)] = 1                               # from length, NOT from pad id
        comp_mask[i, :len(m)] = torch.tensor(m)
    return input_ids, attn, comp_mask

train_loader = DataLoader(SFTDataset(train_df), batch_size=BATCH_SIZE, shuffle=True,  collate_fn=collate)
val_loader   = DataLoader(SFTDataset(val_df),   batch_size=BATCH_SIZE, shuffle=False, collate_fn=collate)

x, a, m = next(iter(train_loader))
print(x.shape, "supervised tokens:", int(m.sum()), "masked:", int((a - m).sum()))

Three design decisions are worth defending in an interview. The chat template reproduces the special tokens the instruct model was trained with. The completion mask restricts the loss to answer tokens. Tokenizing once with offsets avoids boundary drift: tokenizing the prompt separately and using its length can mis-align by a token when the tokenizer would merge characters across the boundary. Padding fills input_ids with the pad token but the masks with 0, so padded positions are neither attended to nor scored; and because the pad token equals EOS here, the attention mask is built from lengths so the real EOS is still attended and supervised.

Step 4: the shared training loop

def batch_loss(model, input_ids, attn, comp_mask):
    logits = model(input_ids=input_ids, attention_mask=attn).logits
    logits, labels = logits[:, :-1].float(), input_ids[:, 1:]
    mask = comp_mask[:, 1:].float()
    tok = F.cross_entropy(logits.reshape(-1, logits.size(-1)), labels.reshape(-1), reduction="none")
    return (tok * mask.reshape(-1)).sum() / mask.sum().clamp(min=1.0)

@torch.no_grad()
def eval_loss(model, loader):
    model.eval(); tot, n = 0.0, 0.0
    for x, a, m in loader:
        x, a, m = x.to(DEVICE), a.to(DEVICE), m.to(DEVICE)
        k = m[:, 1:].sum().item()
        tot += batch_loss(model, x, a, m).item() * k; n += k
    model.train()
    return tot / max(n, 1)

def gpu_mem_gb():
    return torch.cuda.max_memory_allocated() / 1e9 if torch.cuda.is_available() else 0.0

def free_memory():
    gc.collect(); torch.cuda.empty_cache(); torch.cuda.reset_peak_memory_stats()

def train_model(model, lr, tag, patience=1):
    seed_everything(); model.to(DEVICE)
    params = [p for p in model.parameters() if p.requires_grad]
    opt = torch.optim.AdamW(params, lr=lr)
    total = math.ceil(len(train_loader) / GRAD_ACCUM) * EPOCHS
    sched = get_cosine_schedule_with_warmup(opt, int(WARMUP_RATIO * total), total)
    hist, best_val, best_state, bad = {"train": [], "val": []}, float("inf"), None, 0
    t0 = time.time(); model.train(); opt.zero_grad()
    for ep in range(1, EPOCHS + 1):
        run = 0.0
        for i, (x, a, m) in enumerate(train_loader):
            x, a, m = x.to(DEVICE), a.to(DEVICE), m.to(DEVICE)
            loss = batch_loss(model, x, a, m) / GRAD_ACCUM
            loss.backward()
            if (i + 1) % GRAD_ACCUM == 0 or i + 1 == len(train_loader):
                torch.nn.utils.clip_grad_norm_(params, MAX_GRAD_NORM)
                opt.step(); sched.step(); opt.zero_grad()
            run += loss.item() * GRAD_ACCUM
        tr, va = run / len(train_loader), eval_loss(model, val_loader)
        hist["train"].append(tr); hist["val"].append(va)
        print(f"[{tag}] epoch {ep}: train {tr:.4f}  val {va:.4f}")
        if va < best_val:
            best_val, bad = va, 0
            best_state = {n: p.detach().cpu().clone()          # trainable tensors only
                          for n, p in model.named_parameters() if p.requires_grad}
        else:
            bad += 1
            if bad > patience: break                             # early stopping
    model.load_state_dict(best_state, strict=False)              # restore best checkpoint
    return model, hist, time.time() - t0, gpu_mem_gb()

def count_trainable(model):
    t = sum(p.numel() for p in model.parameters() if p.requires_grad)
    n = sum(p.numel() for p in model.parameters())
    return t, n, 100 * t / n

Every standard mechanism appears here: gradient accumulation (divide loss, step every GRAD_ACCUM micro-batches), linear warmup then cosine decay, gradient-norm clipping at 1.0, per-epoch validation, keeping and restoring the best checkpoint, and early stopping. Storing only trainable tensors keeps checkpoints tiny for LoRA; for full fine-tuning it is the whole model.

Step 5: run 1, full instruction fine-tuning

free_memory()
full = AutoModelForCausalLM.from_pretrained(MODEL_ID, torch_dtype=torch.float32)
print("trainable %:", round(count_trainable(full)[2], 2))       # 100.0
full, full_hist, full_time, full_mem = train_model(full, LR_FULL, "full-IFT")

The model is loaded in fp32 because pure fp16 full fine-tuning overflows on this model family and the T4 has no fast bf16; the tiny learning rate (1e-5) prevents the update from overwriting pretrained features. Memory: 0.49B parameters × 16 bytes ≈ 7.9 GB for weights, gradients and Adam states, plus activations, which makes full fine-tuning the most memory-hungry run even at this small scale.

Step 6: runs 2–3, LoRA at r = 8 and r = 16

def lora_cfg(r):
    return LoraConfig(task_type=TaskType.CAUSAL_LM, r=r, lora_alpha=2 * r,   # alpha/r fixed at 2
                      lora_dropout=LORA_DROPOUT, bias="none", target_modules=TARGETS)

lora_runs = {}
for r in (8, 16):
    free_memory(); seed_everything()
    base = AutoModelForCausalLM.from_pretrained(MODEL_ID, torch_dtype=torch.float32)
    model = get_peft_model(base, lora_cfg(r))
    pct = count_trainable(model)[2]                  # about 0.9% at r=8, 1.75% at r=16
    model, hist, t, mem = train_model(model, LR_LORA, f"LoRA r={r}")
    lora_runs[r] = dict(model=model, hist=hist, time=t, mem=mem, pct=pct)

Keeping α/r fixed isolates the effect of rank. The parameter percentage is higher than the "well under 1%" often quoted for 7B models because small models have narrow layers: for this architecture (hidden size 896, MLP width 4864, 24 layers, grouped-query attention with narrow key/value projections) rank 16 on all seven linear layers adds about 8.8M parameters to about 494M.

Step 7: runs 4–5, QLoRA at r = 8 and r = 16

bnb_cfg = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                             bnb_4bit_use_double_quant=True,
                             bnb_4bit_compute_dtype=torch.float16)   # T4: fp16 compute

qlora_runs = {}
for r in (8, 16):
    free_memory(); seed_everything()
    base = AutoModelForCausalLM.from_pretrained(MODEL_ID, quantization_config=bnb_cfg,
                                                torch_dtype=torch.float16)
    base = prepare_model_for_kbit_training(base)     # fp32 norms, checkpointing, input grads
    model = get_peft_model(base, lora_cfg(r))        # identical adapters to the LoRA runs
    pct = count_trainable(model)[2]
    model, hist, t, mem = train_model(model, LR_LORA, f"QLoRA r={r}")
    qlora_runs[r] = dict(model=model, hist=hist, time=t, mem=mem, pct=pct)

The only change from the LoRA runs is the backbone's precision, which makes this a clean controlled comparison. Trainable parameters match LoRA at the same rank exactly. Two small-model subtleties: first, the embedding matrix (vocabulary of about 152k × 896 ≈ 136M parameters) is not quantized, so on a 0.5B model the 4-bit saving is proportionally smaller than on 7B; second, gradient checkpointing (enabled by prepare_model_for_kbit_training) cuts activation memory but adds recomputation time, so the QLoRA runs are typically the slowest per epoch while using the least memory.

Step 8: generation-based evaluation

import evaluate
rouge, bertscore = evaluate.load("rouge"), evaluate.load("bertscore")

@torch.no_grad()
def generate_answers(model, frame, max_new_tokens=200):
    model.eval(); preds, refs = [], []
    for _, row in frame.iterrows():
        enc = tokenizer(build_prompt(row["instruction"], row["input"]), return_tensors="pt",
                        add_special_tokens=False).to(DEVICE)
        out = model.generate(**enc, max_new_tokens=max_new_tokens, do_sample=False,
                             pad_token_id=tokenizer.pad_token_id)
        preds.append(tokenizer.decode(out[0, enc["input_ids"].shape[1]:],     # new tokens only
                                      skip_special_tokens=True).strip())
        refs.append(row["output"].strip())
    model.train()
    return preds, refs

def score(preds, refs):
    rl = rouge.compute(predictions=preds, references=refs, use_stemmer=True)["rougeL"]
    bs = np.mean(bertscore.compute(predictions=preds, references=refs, lang="en")["f1"])
    return {"rougeL": round(float(rl), 4), "bertscore_f1": round(float(bs), 4)}

results = {}
runs = {"full_ift": dict(model=full, time=full_time, mem=full_mem, pct=100.0),
        **{f"lora_r{r}": v for r, v in lora_runs.items()},
        **{f"qlora_r{r}": v for r, v in qlora_runs.items()}}
for name, run in runs.items():
    p, refs = generate_answers(run["model"], test_df)
    results[name] = {**score(p, refs), "pct_trainable": round(run["pct"], 2),
                     "peak_mem_gb": round(run["mem"], 2), "train_time_s": round(run["time"], 1)}
print(pd.DataFrame(results).T)

Step 9: error analysis

best = max(results, key=lambda k: results[k]["rougeL"])
preds, refs = generate_answers(runs[best]["model"], test_df)
pl = np.array([len(tokenizer(p, add_special_tokens=False)["input_ids"]) for p in preds])
rl = np.array([len(tokenizer(r, add_special_tokens=False)["input_ids"]) for r in refs])
ratio = pl / np.maximum(rl, 1)
print("median length ratio:", np.median(ratio).round(2),
      "| empty:", int((pl == 0).sum()),
      "| much shorter (<0.5):", int((ratio < 0.5).sum()),
      "| much longer (>2):", int((ratio > 2).sum()))
for p, r in list(zip(preds, refs))[:3]:
    print("REF:", r[:300], "\nGEN:", p[:300], "\n" + "-" * 60)

What results to expect and how to read them

RunTrainable %Relative peak memoryRelative timeTypical quality
Full IFT (fp32)100HighestMediumSimilar to adapters; small data limits the gain from extra capacity
LoRA r=8≈0.9Much lowerFastestClose to full FT
LoRA r=16≈1.75Much lowerFastOften indistinguishable from r=8 on 300 examples
QLoRA r=8≈0.9LowestSlowestWithin noise of LoRA
QLoRA r=16≈1.75LowestSlowestWithin noise of LoRA

With only 300 examples and 3 epochs, differences between methods are usually small and can flip with the random seed, which is itself the key lesson: on a free GPU, LoRA or QLoRA gets essentially the same answer quality as full fine-tuning for a fraction of memory. Other typical observations: validation loss bottoms out after one or two epochs for full fine-tuning; lower training loss does not reliably predict better generated answers; the dominant failure modes are answers that are shorter or less specific than the references, missing safety caveats, and truncation when max_new_tokens is small. Reasonable next experiments are more data, a longer MAX_LEN, a larger model with QLoRA, or a light repetition penalty at decoding.

Step 10: save, merge and export

best_run = runs["qlora_r16"]["model"]
best_run.save_pretrained("out/qlora-r16-adapter")          # adapter only: a few tens of MB
tokenizer.save_pretrained("out/qlora-r16-adapter")

# merge into a 16-bit base (never into the 4-bit weights), then export or quantize
from peft import PeftModel
base16 = AutoModelForCausalLM.from_pretrained(MODEL_ID, torch_dtype=torch.float16)
merged = PeftModel.from_pretrained(base16, "out/qlora-r16-adapter").merge_and_unload()
merged.save_pretrained("out/merged-fp16"); tokenizer.save_pretrained("out/merged-fp16")
# next: convert to GGUF and quantize for local/edge use (see the serving section)
Interview angle Interviewers who see a project like this on a CV ask: why the chat template, why completion-only loss, why tokenize once with offsets, why fp32 for full fine-tuning on a T4, why a 20× larger LR for LoRA, what prepare_model_for_kbit_training does, why QLoRA has the same trainable-parameter count as LoRA, whether r=16 beat r=8, and why lower training loss did not mean better test answers. Have a one-sentence answer ready for each, backed by your own numbers.
Common pitfall Three bugs that appear constantly in hand-written SFT loops: forgetting to append the end-of-turn token (the model never stops), deriving the attention mask from input_ids != pad_token_id when pad equals EOS (the stop token is hidden), and saving the "best" state but never loading it back before evaluation. Also never treat a 0.5B model fine-tuned on machine-generated medical answers as clinically reliable; it is a methods exercise.

Quick revision

  • Fine-tuning continues gradient training of a pretrained model on your data; inference and prompting never change weights.
  • Rule of thumb: new knowledge calls for RAG, new behaviour calls for fine-tuning, and many systems use both.
  • Always baseline with the best prompt you can write and build an evaluation set before fine-tuning.
  • Continued pretraining uses raw domain text; SFT uses prompt–response pairs; preference tuning uses chosen/rejected pairs or rewards.
  • SFT is next-token cross-entropy with teacher forcing, usually on response tokens only (prompt labels set to -100).
  • Logits at position t predict token t+1, so shift logits and labels (and the mask) by one.
  • Use the model's own chat template via apply_chat_template; mismatched templates degrade quality.
  • Append the end-of-turn/EOS token to every target, or the model never learns to stop.
  • Tokenize prompt + answer once and use offsets to locate the answer boundary.
  • Quality beats quantity: a thousand clean, diverse examples can beat a hundred thousand noisy ones.
  • Deduplicate, scrub personal data, decontaminate against test sets, and split by group to avoid leakage.
  • Set maximum sequence length near the 95th percentile of tokenized length plus template overhead.
  • Packing concatenates short examples to cut padding waste; block cross-example attention when possible.
  • Full fine-tuning with mixed-precision AdamW needs about 16 bytes per parameter before activations: 112 GB for 7B.
  • Activations scale with batch × sequence × hidden × layers; gradient checkpointing trades roughly a third more compute for large savings.
  • Large vocabularies make the logits tensor a hidden memory hog; chunked or fused cross-entropy helps.
  • PEFT freezes the base and trains a small set of parameters: adapters, prefix/prompt tuning, IA3, BitFit, LoRA.
  • LoRA: h = W0x + (α/r)BAx, with A random and B zero so training starts from the pretrained model.
  • LoRA adds r·(din + dout) parameters per adapted matrix; rank 16 on all linear layers of a 7B model is about 40M parameters.
  • Targeting all linear layers usually matters more than increasing rank.
  • Keep α/r fixed when sweeping rank; rsLoRA uses α/√r for stability at high rank.
  • Merged LoRA adds zero inference latency; unmerged adapters add a small overhead but can be swapped.
  • DoRA separates magnitude and direction, training LoRA on the direction, and often helps at low rank.
  • QLoRA = NF4 4-bit frozen base + double quantization + paged optimizers + 16-bit LoRA adapters; compute is in bf16/fp16.
  • NF4 places 16 levels at normal-distribution quantiles with blockwise (64) absmax scaling.
  • Double quantization cuts scale-constant overhead from 0.5 to about 0.127 bits per parameter.
  • prepare_model_for_kbit_training upcasts norms, enables checkpointing and input gradients for k-bit training.
  • Typical LRs: full FT about 1e-5, LoRA about 2e-4, DPO about 5e-7 to 5e-6 (full) with β about 0.1.
  • 1–3 epochs, warmup 3–10%, cosine or linear decay, clipping at 1.0, and keep the best validation checkpoint.
  • RLHF = SFT, then a Bradley–Terry reward model from human comparisons, then PPO maximizing reward minus β·KL to the reference.
  • PPO-based RLHF keeps four models in memory: policy, reference, reward and value. The clip applies to ρ then multiplies by A; the objective is maximized.
  • DPO loss: −log σ(β [log(πθ(yw)/πref(yw)) − log(πθ(yl)/πref(yl))]); no reward model or sampling; smaller β allows more deviation from the reference.
  • ORPO and SimPO are reference-free; KTO uses unpaired thumbs-up/down data; IPO regularizes DPO's margins.
  • GRPO keeps PPO's clip but replaces the critic with group-standardized rewards over G samples; two models in memory instead of four; stalls when a group has zero reward variance.
  • RLAIF and Constitutional AI replace human preference labels with AI judgements guided by written principles.
  • Reward hacking, sycophancy and verbosity are classic preference-tuning failure modes; the KL penalty limits them.
  • Evaluate on held-out task data, with a validated LLM judge, on general benchmarks, safety suites and a regression suite, all against the base model.
  • Catastrophic forgetting is mitigated by PEFT, data replay, lower LR, fewer epochs, regularization and model merging.
  • Even benign fine-tuning can weaken safety refusals, so re-run safety evaluations on every release.
  • Merge adapters into a 16-bit base, never directly into 4-bit weights, then re-quantize if needed.
  • Model merging (linear, SLERP, task arithmetic, TIES, DARE) only works for models sharing a base.
  • Distillation trains a smaller student on teacher outputs or soft logits with temperature; check the teacher's terms of use.
  • Serve one merged model for high-traffic single tasks, or multi-LoRA over a shared base for many variants.
  • For local/edge deployment: merge, convert to GGUF, quantize (for example Q4_K_M), and re-evaluate with the right template.

Glossary

Adapter
A small trainable module inserted into a frozen network; in the narrow sense, a bottleneck MLP added after attention or feed-forward blocks.
AdamW
Adam optimizer with decoupled weight decay; stores two moment estimates per trainable parameter.
Alignment
Post-training that shapes a model's behaviour towards human intentions and values, such as helpfulness, honesty and harmlessness.
Alpha (LoRA)
Scaling hyperparameter; the LoRA update is multiplied by α/r.
Base model
A pretrained model before instruction tuning or alignment; it continues text rather than following instructions.
bitsandbytes
Library of 8-bit and 4-bit quantized layers and memory-efficient optimizers used by QLoRA.
Bradley–Terry model
Probability model for pairwise preferences: P(A preferred over B) = σ(rA − rB).
Catastrophic forgetting
Loss of previously learned abilities when a network is trained on new data.
Chat template
The exact special-token layout a chat model uses to mark system, user, assistant and tool turns.
Completion-only loss
Computing the training loss only on response tokens, masking prompt tokens.
Constitutional AI
Alignment approach in which a model critiques and revises outputs, and an AI judge ranks them, according to a written set of principles.
Continued pretraining
Further next-token training on large unlabelled domain text; also called domain-adaptive pretraining.
DARE
Model-merging pre-processing that randomly drops most delta parameters and rescales the remainder.
DeepSpeed ZeRO
Memory optimization that shards optimizer states (stage 1), gradients (stage 2) and parameters (stage 3) across GPUs.
Distillation
Training a smaller student model to imitate a larger teacher's outputs or probability distributions.
DoRA
Weight-decomposed LoRA: learns a magnitude vector and applies LoRA to the normalized direction of each weight.
Double quantization
Quantizing the per-block quantization constants themselves to save memory.
DPO
Direct Preference Optimization: aligns a model with preference pairs using a closed-form loss on policy-to-reference log-probability ratios.
Early stopping
Halting training when validation performance stops improving, keeping the best checkpoint.
Effective batch size
Micro-batch size × gradient-accumulation steps × number of data-parallel GPUs.
FSDP
Fully Sharded Data Parallel: PyTorch's native sharding of parameters, gradients and optimizer states.
GGUF
Single-file model format with metadata and many quantization types, used by llama.cpp-family runtimes.
Gradient accumulation
Summing gradients over several micro-batches before one optimizer step to simulate a larger batch.
Gradient checkpointing
Saving only some activations and recomputing the rest during backward to save memory.
GRPO
Group Relative Policy Optimization: PPO-style RL without a value model, using rewards standardized within a group of samples per prompt.
IA3
PEFT method that learns vectors rescaling keys, values and feed-forward activations.
Instruction tuning
Supervised fine-tuning on diverse instruction–response pairs so a model follows natural-language commands.
KL penalty
Term penalizing divergence of the trained policy from a reference model during preference optimization.
KTO
Preference method learning from unpaired examples labelled desirable or undesirable.
LLM-as-judge
Using a strong language model with a rubric to rate or compare outputs.
LoRA
Low-Rank Adaptation: trains a low-rank update BA added to frozen weight matrices.
Loss masking
Excluding tokens from the loss, typically by setting their labels to -100.
Merge and unload
Folding a LoRA update into the base weights and removing the adapter modules.
Mixed precision
Computing in 16-bit while keeping an fp32 master copy of weights for the optimizer.
Model merging
Combining weights of several fine-tuned models that share a base, without further training.
Multi-LoRA serving
Serving many LoRA adapters over one shared base model, batching requests for different adapters together.
NF4
4-bit NormalFloat: a 16-level code at normal-distribution quantiles, used for storing QLoRA base weights.
ORPO
Odds Ratio Preference Optimization: single-stage SFT plus odds-ratio preference loss, with no reference model.
Packing
Concatenating multiple short examples into one training sequence to reduce padding.
Paged optimizer
Optimizer whose states live in unified memory and page to CPU RAM during GPU memory spikes.
PEFT
Parameter-Efficient Fine-Tuning: methods that train a small fraction of parameters; also the name of the Hugging Face library.
Perplexity
Exponential of the average per-token cross-entropy; lower means the model finds the text less surprising.
PPO
Proximal Policy Optimization: RL algorithm using a clipped probability-ratio objective and a value model.
Prefix tuning
PEFT method that learns key/value vectors prepended at every attention layer.
Prompt tuning
PEFT method that learns a few soft embedding vectors prepended to the input.
QLoRA
LoRA fine-tuning on a base model stored in 4-bit NF4, with double quantization and paged optimizers.
Rank (LoRA)
Inner dimension r of the LoRA matrices; bounds the rank of the update and sets adapter size.
Reference model
Frozen copy of the starting policy used for the KL penalty or DPO log-ratios.
Replay
Mixing general or earlier-task data into fine-tuning to reduce forgetting.
Reward hacking
A policy exploiting flaws in a reward signal to score highly without truly improving.
Reward model
Model trained on human (or AI) comparisons to output a scalar score for a prompt–response pair.
RLAIF
Reinforcement Learning from AI Feedback: preference labels come from an AI judge instead of humans.
RLHF
Reinforcement Learning from Human Feedback: SFT, reward model from human preferences, then RL against that reward.
RLVR
RL with verifiable rewards: rewards from automatic checks such as correct final answers or passing tests.
rsLoRA
Rank-stabilized LoRA using scale α/√r.
SFT
Supervised Fine-Tuning: training on input–target demonstrations with next-token cross-entropy.
SFTTrainer
TRL's trainer for supervised fine-tuning, with chat formatting, packing and PEFT integration.
SimPO
Reference-free preference method using length-normalized log-probability as the implicit reward plus a margin.
Sycophancy
Tendency of a model to agree with or flatter the user rather than be accurate.
Synthetic data
Training examples generated or rewritten by a model rather than written by people.
Target modules
The layers (for example q_proj, v_proj, gate_proj) that receive LoRA adapters.
Task vector
Difference between fine-tuned and base weights, used for task arithmetic.
Teacher forcing
Training by feeding the true previous tokens rather than the model's own predictions.
TIES merging
Merging method that trims small deltas, elects a sign per parameter and averages agreeing deltas.
TRL
Hugging Face library of trainers for SFT, reward modelling, DPO, PPO, GRPO and related methods.
Warmup
Gradually increasing the learning rate from near zero over the first steps of training.

Interview questions

Fundamentals

What is fine-tuning an LLM?

Continuing to train a pretrained model with gradient descent on a smaller, task- or domain-specific dataset so that its weights (all of them, or a small added set) change to produce the desired behaviour. The result persists in the weights, unlike prompting, which only changes the input.

How is fine-tuning different from pretraining?

Pretraining is self-supervised next-token prediction on trillions of tokens of raw text, costs months and millions of dollars, and produces general knowledge. Fine-tuning starts from those weights, uses thousands to millions of curated examples, runs for minutes to days, and targets specific behaviour. The loss can be identical (next-token cross-entropy); data, scale, masking and learning rate differ.

What is the difference between a base model and an instruct (chat) model?

A base model is the output of pretraining: it continues text and may answer a question with more questions. An instruct model has additionally gone through SFT on instruction–response data and usually preference tuning, so it follows instructions, uses a chat template and refuses harmful requests.

When should you fine-tune instead of prompting or using RAG?

Fine-tune when the problem is behaviour: consistent format, tone, a specialised skill such as classification or tool calling, or when you need a small cheap model to replace a big one with a long prompt. Use RAG when the problem is knowledge, especially fresh, large or citable facts. Try prompting first because it is cheapest, and often combine fine-tuning (style) with RAG (facts).

What is instruction tuning?

Supervised fine-tuning on many diverse instruction–response pairs (instruction, optional input, output) so that the model learns to follow natural-language commands in general, including unseen ones. It is what turns a base model into an assistant.

Is instruction tuning the same as prompt engineering?

No. Instruction tuning happens at training time and changes weights. Prompt engineering happens at inference time and only changes the input text; the model is unchanged once the prompt is gone.

What is SFT?

Supervised Fine-Tuning: training on input–target demonstrations using next-token cross-entropy with teacher forcing, typically computing the loss only on the target (response) tokens.

What is RLHF in one paragraph?

Reinforcement Learning from Human Feedback is a post-training pipeline: first SFT on demonstrations; then humans compare pairs of model responses and a reward model is trained to predict their preferences; finally the model is optimized with RL (classically PPO) to maximize the reward while a KL penalty keeps it close to the SFT model. It aligns the model with what people prefer rather than just what demonstrators wrote.

What is PEFT and why is it useful?

Parameter-Efficient Fine-Tuning freezes the pretrained weights and trains a small number of extra or selected parameters (often under 1–2%). It cuts GPU memory (no gradients or optimizer states for frozen weights), produces tiny adapter files, lets many tasks share one base model, and reduces catastrophic forgetting, while matching full fine-tuning on most SFT tasks.

What is LoRA?

Low-Rank Adaptation keeps a weight W0 frozen and learns an update ΔW = (α/r)·BA, where A is r × din and B is dout × r with small rank r (such as 8–64). Only A and B are trained; after training they can be merged into W0 with no latency cost.

What is QLoRA?

LoRA on top of a base model stored in 4-bit NormalFloat (NF4) with double quantization, plus paged optimizers to absorb memory spikes. The frozen weights are dequantized to bf16/fp16 on the fly for each matmul; only the 16-bit adapters are trained. It lets a 7B–8B model be fine-tuned on a single 12–16 GB GPU.

What does the LoRA rank r control?

The inner dimension of A and B, which bounds the rank of the learned update and sets adapter size (r·(din+dout) parameters per matrix). Higher r means more capacity and memory; for most SFT tasks 8–32 is enough.

What does lora_alpha do?

It scales the update by α/r. It acts like a learning-rate multiplier for the adapter path and lets you change r without retuning the learning rate much. Common settings are α = r or α = 2r.

What are target modules in LoRA?

The layers that receive adapters, named by module, for example q_proj, k_proj, v_proj, o_proj in attention and gate_proj, up_proj, down_proj in the MLP. Adapting all linear layers generally works better than adapting only query and value projections.

What is a chat template and why does it matter for fine-tuning?

It is the exact special-token format a chat model uses to delimit system, user and assistant turns. Fine-tuning an instruct model with a different format wastes capacity re-learning structure and can break its chat behaviour; serving with a different template than training causes rambling or garbled output. Use tokenizer.apply_chat_template.

What is catastrophic forgetting?

The loss of previously learned capabilities when a network is trained on new data, because the same weights encode both old and new skills and new gradients overwrite the old solutions. In LLMs it shows up as worse general reasoning, instruction following, other languages or safety after narrow fine-tuning.

How much data do you need to fine-tune?

It depends on the goal: a format or style change can work with a few hundred to two thousand high-quality examples; simple classification with 500–5,000; a domain assistant with 5k–50k; preference tuning with a few thousand pairs or more. Quality, diversity and consistency matter more than raw count.

What is overfitting in fine-tuning and how do you spot it?

The model memorizes training examples instead of learning the general behaviour. Training loss keeps falling while validation loss rises; generations copy training answers verbatim or become repetitive. Fix with early stopping, fewer epochs, more or more diverse data, dropout and lower rank.

Why is the learning rate for LoRA higher than for full fine-tuning?

LoRA's base is frozen and B starts at zero, so the adapter begins as a no-op and its update is confined to a low-rank subspace; larger steps (about 2e-4) are safe. Full fine-tuning changes every pretrained weight directly, so it needs small steps (about 1e-5) to avoid destroying learned features or diverging.

What is an epoch, a step and gradient accumulation?

An epoch is one pass over the training set. A step is one optimizer update. Gradient accumulation sums gradients over several micro-batches before one step, so the effective batch size is micro-batch × accumulation steps × number of GPUs, without the activation memory of a large batch.

What is learning-rate warmup and why use it?

Increasing the learning rate linearly from near zero over the first few percent of steps. Early on, Adam's moment estimates are unreliable and a few large updates can damage the pretrained model; warmup keeps early updates small.

Can you fine-tune closed-source models?

Not in the full sense: you cannot access or update their weights. Some providers offer managed fine-tuning APIs for selected models, where you upload data and receive a private endpoint, with limited control over method and hyperparameters and no weight download. Open-weight models give full control.

What does merging a LoRA adapter mean?

Computing W' = W0 + (α/r)BA for every adapted layer and saving the result as an ordinary model, so no adapter modules remain and there is no inference overhead. In PEFT it is merge_and_unload().

What is DPO?

Direct Preference Optimization trains a model directly on chosen/rejected pairs with a logistic loss on the difference of policy-to-reference log-probability ratios, scaled by β. It achieves the RLHF objective without training a reward model or running RL, so it is simpler, cheaper and more stable.

What is a reward model?

A model (usually an LLM with a scalar output head) trained on preference comparisons to give higher scores to responses humans prefer. It serves as a learned proxy for human judgement during RL.

What is knowledge distillation?

Training a small student model to imitate a larger teacher, either on the teacher's generated text (sequence-level) or by matching its softened next-token distributions (logit-level, with a KL loss and temperature). It yields small models that perform far above their size.

Does fine-tuning add knowledge reliably?

Only weakly. Models absorb new facts slowly through fine-tuning, cannot cite them, and training on facts they do not know can increase hallucinations. Fine-tuning is best for behaviour and skills; retrieval is better for facts, especially changing ones.

Which libraries are commonly used to fine-tune open models?

Hugging Face Transformers (models, tokenizers, Trainer), PEFT (LoRA and other adapters), TRL (SFT, DPO, PPO, GRPO trainers), bitsandbytes (4/8-bit quantization and optimizers), Accelerate with DeepSpeed or FSDP for multi-GPU, and higher-level tools such as Unsloth (optimized single-GPU LoRA/QLoRA) and Axolotl or similar YAML-driven wrappers.

What is the difference between full fine-tuning and LoRA in terms of outputs saved?

Full fine-tuning saves a complete new model (for example 14 GB for 7B in bf16). LoRA saves only the adapter matrices, usually tens to a few hundred megabytes, which must be loaded on top of the exact same base model.

What is continued pretraining and when would you use it?

Further next-token training on large amounts of unlabelled domain text (papers, code, legal documents, a new language). Use it when the domain vocabulary and style are far from the base model's training data; follow it with SFT to restore instruction following.

Going deeper

Why do we compute the SFT loss only on response tokens?

We want to learn p(response | prompt), not to model prompts. Including prompt tokens wastes capacity, lets long prompts dominate the gradient, and can teach the model to produce user-like or system-prompt text. Prompt tokens still serve as context; only their labels are masked (-100). On very short-answer datasets some teams keep a small prompt-loss weight, but completion-only is the default.

Explain the shift-by-one in causal LM loss.

The logits at position t are the model's prediction for token t+1. So the loss compares logits[:, :-1] against input_ids[:, 1:], and the mask must be shifted the same way. Hugging Face models do this internally when you pass labels; custom loops must do it explicitly or they train the model to copy the current token.

Why tokenize prompt and answer together rather than separately?

Subword tokenizers can merge characters across the boundary differently when strings are joined, so separately tokenized pieces do not equal the tokenization of the full text the model sees at inference. Tokenize the concatenation once and use character offsets (return_offsets_mapping) to find which tokens belong to the answer.

Walk through the memory needed to fully fine-tune a 7B model.

With mixed-precision AdamW: 2 bytes bf16 weights + 2 bytes gradients + 4 bytes fp32 master weights + 8 bytes Adam moments = 16 bytes per parameter, so about 112 GB, plus activations (several GB to tens of GB depending on batch, sequence length, fused attention and checkpointing). It needs multiple 80 GB GPUs with ZeRO-3/FSDP, or aggressive tricks (8-bit Adam, offload) on one.

How much memory does QLoRA need for a 7B model?

About 3.5 GB for 4-bit weights plus roughly 0.1 GB of quantization constants with double quantization and a few hundred MB for unquantized embeddings/head; about 40M LoRA parameters (rank 16, all linear) cost around 0.6 GB with Adam states; activations with gradient checkpointing add 1–4 GB depending on sequence length. Total is roughly 6–10 GB, so a 12–16 GB GPU is comfortable.

Calculate LoRA parameters for a 4096 × 4096 projection at rank 8 and 16.

r·(din + dout) = 8 × 8192 = 65,536 at rank 8, and 131,072 at rank 16, compared with 16,777,216 for the full matrix (256× and 128× fewer).

Why is B initialized to zero and A randomly in LoRA?

With B = 0 the product BA is zero, so the adapted model starts exactly equal to the pretrained model, making early training stable. A must be random (non-zero); if both were zero, each matrix's gradient would be zero because it depends on the other, and nothing would learn.

Does LoRA reduce training compute as much as memory?

No. The forward pass still runs through the full model, and the backward pass still propagates activation gradients through every frozen layer; only the weight-gradient computation and optimizer update for frozen weights are skipped. Memory drops dramatically; step time drops modestly (and QLoRA is often slower than LoRA due to dequantization).

Explain NF4 quantization.

NormalFloat4 uses 16 levels placed at quantiles of a standard normal distribution normalized to [−1, 1], so that normally distributed weights use each level about equally (near-optimal for that distribution), with an exact zero. Weights are quantized in blocks of 64, each scaled by its absmax, which limits outlier damage.

What is double quantization and how much does it save?

Blockwise quantization stores one fp32 scale per 64 weights (0.5 bits per parameter). Double quantization quantizes those scales to 8-bit in groups of 256 with one fp32 second-level scale per group, reducing overhead to about 0.127 bits per parameter, saving roughly 0.37 bits per parameter (about 0.3 GB on a 7B model).

What are paged optimizers?

Optimizers whose state tensors are allocated in NVIDIA unified memory, so that when GPU memory spikes (for example on an unusually long batch) pages are automatically evicted to CPU RAM and brought back later. They turn out-of-memory crashes into small slowdowns.

What does prepare_model_for_kbit_training do?

For a quantized model it casts layer norms (and typically the output head) to fp32 for numerical stability, enables gradient checkpointing, and makes input embeddings require gradients (via a hook) so gradients can flow through the frozen, checkpointed network to the LoRA adapters.

Why is compute dtype fp16 rather than bf16 on some GPUs?

Older GPUs such as the T4 (Turing) lack native fast bf16 support, so compute uses fp16; on Ampere and newer, bf16 is preferred because its wider exponent range avoids overflow. Full fine-tuning in pure fp16 can overflow, which is why fp32 (or bf16 mixed precision) is used for it.

Why do QLoRA and LoRA have the same number of trainable parameters at the same rank?

The adapters are identical; QLoRA only changes the storage precision of the frozen backbone. Trainable parameter count depends on rank and target modules, not on how frozen weights are stored.

Compare adapters, prefix tuning, prompt tuning and LoRA.

Bottleneck adapters insert small MLPs into each block (sequential, not mergeable, adds latency). Prefix tuning learns key/value vectors at every layer; prompt tuning learns soft input embeddings only (both consume context length and prompt tuning needs very large models). LoRA learns low-rank weight updates that merge into existing weights, adding no latency, and works across sizes, which is why it dominates.

How do you pick the LoRA rank?

Start with 16 on all linear layers with α = 2r. Sweep 8/16/32/64 with α/r fixed and compare validation task metrics. Increase rank for large shifts (new language, complex skills, lots of data); keep it low for small datasets to reduce overfitting. Gains usually plateau quickly.

What are the three stages of classic RLHF and what models are involved?

SFT (produces the reference policy), reward-model training on pairwise human preferences (Bradley–Terry loss), and PPO optimization of the policy against the reward with a KL penalty. PPO needs the policy, a frozen reference, the reward model and a value model: four models in memory.

Write the reward model loss.

L = −E[log σ(r(x, yw) − r(x, yl))], the Bradley–Terry negative log-likelihood that the chosen response yw beats the rejected yl. Only reward differences matter, so rewards are shift-invariant and are often normalized.

What is the purpose of the KL penalty in RLHF?

It keeps the policy close to the SFT reference, which prevents reward hacking (exploiting reward-model errors in regions it never saw), preserves fluency and general capabilities, and stabilizes training. β controls the trade-off between reward and staying close.

How does DPO differ from PPO-based RLHF?

DPO is offline and RL-free: it uses a fixed preference dataset and a classification-style loss on log-probability ratios, needing only the policy and a frozen reference. PPO is online: it samples from the current policy, scores with a reward model, estimates advantages with a value model, and can explore beyond the dataset but is expensive and unstable.

Write the DPO loss.

LDPO = −E [ log σ( β ( log(πθ(yw|x)/πref(yw|x)) − log(πθ(yl|x)/πref(yl|x)) ) ) ]. Log-probabilities are summed over response tokens. Smaller β allows more deviation from the reference. Equivalent form: −log σ(β (log πθ(yw) − log πθ(yl) − log πref(yw) + log πref(yl))).

What does β mean in DPO?

It is the strength of the implicit KL constraint to the reference. Small β lets the policy move further from the reference (stronger preference fitting, more risk of degradation); large β keeps it close. Typical values are 0.05–0.5, with 0.1 a common default.

What are ORPO and KTO and when would you use them?

ORPO combines SFT and preference learning into one stage with an odds-ratio term and needs no reference model: useful when you want one cheap training run from a base or lightly tuned model. KTO learns from unpaired examples labelled good or bad, suitable for production thumbs-up/down logs where you rarely have paired comparisons.

What is GRPO?

Group Relative Policy Optimization samples a group of responses for each prompt, scores them, and uses each response's reward standardized within its group as the advantage, in a PPO-style clipped objective with a KL penalty. The advantage is sequence-level (one A per response) and is applied to every token; the probability ratio is still per token. It removes the value model, so typical memory is policy + reference rather than PPO's four models, and is well suited to verifiable rewards such as correct math answers or passing unit tests.

Compare GRPO and PPO.

Both maximize min(ρA, clip(ρ, 1−ε, 1+ε)A) with a KL penalty to a reference. PPO estimates A with a learned value model and GAE and, in RLHF, also keeps a reward model: four networks in memory, usually one rollout per prompt. GRPO drops the critic and sets A to the group z-score of G samples (8–16) of the same prompt; the scorer is often a deterministic verifier. Choose PPO when preferences are subjective and you already train a reward model; choose GRPO when answers can be checked automatically. GRPO's distinctive failure is a zero-variance group (all correct or all wrong), which contributes no gradient.

What is RLAIF / Constitutional AI?

RLAIF uses an AI model's judgements instead of human labels to create preference data. Constitutional AI structures this with a written set of principles: the model critiques and revises its own outputs against them (supervised phase), then an AI judge chooses between responses by the principles to train a preference model for RL. It scales harmlessness training and makes the values explicit.

How do you evaluate a fine-tuned model?

On a held-out test set with task metrics (accuracy, F1, exact match, format validity, ROUGE/BERTScore for free text), a validated LLM judge with pairwise comparisons against the base model, general benchmarks and safety suites to detect regression, a product regression suite in CI, and ultimately online A/B tests. Always compare against the best-prompted base model.

Why report both ROUGE-L and BERTScore?

ROUGE-L measures lexical overlap (longest common subsequence), rewarding the same wording; BERTScore measures embedding similarity, rewarding the same meaning with different wording. Together they distinguish copying from correct paraphrasing. Neither checks factual correctness reliably, so add judges or human review for high-stakes domains.

What is packing and what can go wrong with it?

Packing concatenates several examples into one full-length sequence to avoid padding waste. If attention is not blocked at example boundaries and position IDs are not reset, examples can attend to unrelated previous examples (cross-contamination), which slightly hurts quality; modern implementations handle this with per-example position IDs and variable-length attention kernels.

How do DeepSpeed ZeRO stages differ?

ZeRO-1 shards optimizer states across data-parallel GPUs; ZeRO-2 also shards gradients; ZeRO-3 also shards parameters, gathering each layer just in time. Higher stages save more memory per GPU at the cost of more communication. Offload variants move states to CPU or NVMe. FSDP is PyTorch's native equivalent of ZeRO-3.

What is gradient checkpointing and what does it cost?

Instead of storing every intermediate activation for the backward pass, store only checkpoints (such as each layer's input) and recompute the rest during backward. Activation memory drops from proportional to all intermediates to roughly one layer's worth plus the checkpoints, at the cost of roughly one extra forward pass (about 20–35% more compute).

How would you create synthetic fine-tuning data safely?

Seed a strong teacher with diverse real prompts, generate instructions and responses, filter with rules and an LLM judge, verify facts or run code/tests where possible, deduplicate heavily, balance topics, keep a share of human data to avoid collapse, and confirm that the teacher's licence and terms allow training on its outputs.

Advanced

Derive the DPO loss from the RLHF objective.

The KL-regularized objective maxπ E[r(x,y)] − β·KL(π ‖ πref) has the closed-form optimum π*(y|x) = πref(y|x)·exp(r(x,y)/β) / Z(x). Solving for the reward gives r(x,y) = β·log(π*(y|x)/πref(y|x)) + β·log Z(x). Substituting into the Bradley–Terry likelihood of the preference yw ≻ yl, the Z(x) terms cancel because they depend only on x, leaving L = −log σ(β[log(πθ(yw)/πref(yw)) − log(πθ(yl)/πref(yl))]). The policy itself parameterizes the reward: "your language model is secretly a reward model".

What does the DPO gradient do, intuitively?

It increases the log-probability of the chosen response and decreases that of the rejected one, weighted by σ(implicit reward of rejected − chosen), i.e. by how wrongly the current model ranks the pair. Pairs already ranked correctly by a wide margin contribute little; misranked pairs contribute most. This weighting is what prevents the degenerate behaviour of naive "unlikelihood" training.

What are known failure modes of DPO?

Both chosen and rejected likelihoods can fall (the model just widens the margin by making everything less likely), leading to degraded or out-of-distribution outputs; overfitting to deterministic preferences (motivating IPO); length exploitation, since longer chosen answers accumulate larger log-ratios (motivating length normalization in SimPO or length-controlled evaluation); and sensitivity to off-policy data generated by a different model. Remedies: SFT on chosen first, add an SFT/NLL term on chosen, tune β, use on-policy or iterative DPO.

Explain the PPO clipped objective and why clipping matters.

PPO maximizes E[min(ρA, clip(ρ, 1−ε, 1+ε)A)], where ρ is the new/old probability ratio for an action (token) and A the advantage. Clip the ratio, then multiply by A; do not write clip(ρA). When A > 0 the clip stops you rewarding ρ > 1+ε; when A < 0 it stops you rewarding ρ < 1−ε. Either way you lose the incentive to move the policy further in the improving direction, which is the trust region. Without clipping, a few high-|A| samples can cause large destructive updates. The full PPO loss also has a value-function term and often an entropy bonus.

How are advantages computed in PPO-RLHF versus GRPO?

In PPO-RLHF, the reward model score is given at the end of the sequence, a per-token KL penalty is added at each token, and a learned value model estimates expected return so generalized advantage estimation (GAE) produces per-token advantages. In GRPO, there is no value model: G responses are sampled per prompt and each response's advantage is its reward minus the group mean, divided by the group standard deviation, applied to all its tokens. RLOO similarly uses the leave-one-out mean of the other samples as a baseline.

Why can GRPO training stall, and how do you fix it?

If all responses in a group get the same reward (all correct or all wrong), standardized advantages are zero and the prompt contributes no gradient. Too-easy or too-hard prompt sets therefore stall learning. Fixes: curriculum or difficulty filtering to keep prompts where the model succeeds sometimes, larger group size, partial-credit rewards, and dropping zero-variance groups from the batch.

What is reward hacking and how do you detect and mitigate it?

The policy finds outputs the reward model overrates without being better: excessive length, flattering tone, repeated phrases, or for verifiable rewards, formatting tricks that fool the checker. Detect by tracking reward vs held-out human or judge preference (reward rising while true quality flattens), response length, KL from reference, and reading samples. Mitigate with KL penalties, reward-model ensembles and uncertainty, length penalties or normalization, robust verifiers, periodic reward-model retraining on new policy samples, and early stopping.

Why does fine-tuning on new factual knowledge increase hallucination?

Examples containing facts the model does not already know are learned slowly and, when learned, teach the model to produce confident answers not grounded in its internal knowledge. The model generalizes that behaviour to other questions it cannot answer. Fine-tuning mostly teaches how to use existing knowledge; new facts are better supplied via retrieval, and "I don't know" examples help calibration.

What is the superficial alignment hypothesis?

The idea that almost all knowledge and capability comes from pretraining, and alignment/SFT mainly teaches the format and style for interacting with users, which is why a small, very high-quality SFT set can produce a strong assistant. It implies that data quality and diversity matter far more than volume in SFT, and that SFT cannot compensate for a weak base model.

Compare LoRA and full fine-tuning in what they learn and forget.

Studies find that full fine-tuning learns larger-rank weight changes and achieves more on large distribution shifts (for example continued pretraining on code or math), while LoRA learns less but forgets less of the base model's general abilities, acting as a regularizer. For typical instruction tuning with modest data, LoRA on all linear layers with adequate rank performs on par with full fine-tuning.

Explain DoRA and why it can outperform LoRA.

DoRA decomposes each weight into a magnitude vector m (column norms) and a direction V/‖V‖. It trains m directly and applies LoRA to the direction: W' = m·(W0 + BA)/‖W0 + BA‖c. Analysis showed full fine-tuning tends to change magnitude and direction in a decoupled way while LoRA couples them; DoRA restores that flexibility, improving quality especially at low ranks, with extra compute for the norm and still mergeable after training.

Why does the α/r scaling slow learning at high rank, and what does rsLoRA change?

With α/r, the update's magnitude shrinks as r grows, and analysis of learning dynamics shows that the gradients' effect on the output scales poorly, so high-rank adapters effectively learn with a smaller step and fail to use their extra capacity. rsLoRA uses α/√r, which keeps the output update stable as rank increases, making higher ranks actually beneficial.

How do you estimate activation memory for a transformer, and what reduces it most?

A standard estimate for 16-bit activations per layer is s·b·h·(34 + 5·a·s/h) bytes (s sequence, b batch, h hidden, a heads). The 5·a·s2·b attention-score term vanishes with FlashAttention-style kernels; full gradient checkpointing reduces storage to about 2·s·b·h bytes per layer plus one layer's recompute. Micro-batch reduction scales linearly; sequence parallelism and activation offload help at extreme lengths.

How does multi-LoRA serving batch requests for different adapters efficiently?

The base model's matmul is computed once for the whole batch. For the low-rank path, requests are grouped by adapter and a segmented/gathered batched matmul kernel applies each request's own A and B (as in Punica's SGMV or S-LoRA). Adapters are stored in a unified memory pool with paging between CPU and GPU, alongside the KV cache. Limits include maximum rank, the number of GPU-resident adapters, and some overhead versus a merged model.

How do TIES and DARE improve model merging?

Naive averaging of task vectors suffers interference: redundant small changes add noise and opposite-sign changes cancel. TIES trims each task vector to its largest-magnitude entries, elects a sign per parameter by total magnitude, and averages only values agreeing with that sign. DARE randomly drops most delta entries (for example 90%) and rescales the rest by 1/(1−p), exploiting delta redundancy; it is often combined with TIES.

Why use temperature and the T2 factor in logit distillation? Forward or reverse KL?

Temperature T > 1 softens both distributions so the student learns the teacher's relative preferences among non-top tokens ("dark knowledge"). Gradients of the softened KL scale as 1/T2, so multiplying by T2 keeps their magnitude comparable to the hard-label term. Forward KL(teacher ‖ student) is mode-covering: the student spreads mass over all teacher modes and can produce odd samples; reverse KL is mode-seeking, focusing on high-probability teacher regions, often better for generation, and used in on-policy distillation.

Why can merging a LoRA into a 4-bit model harm quality, and what is the right procedure?

The adapter was trained against the dequantized 4-bit weights. Adding its small update to 4-bit values and re-rounding to the same 4-bit grid loses most of the update. Merging into the original 16-bit weights instead introduces a mismatch: the adapter compensated for quantization errors that are no longer there, but in practice that mismatch is small. Standard practice: load the base in 16-bit, merge, evaluate, then apply a fresh post-training quantization (GPTQ, AWQ, GGUF) with a domain calibration set and evaluate again. Methods like LoftQ initialize adapters to reduce this gap.

How does the gradient flow through a 4-bit frozen layer in QLoRA?

For y = Ŵx + s·BAx, backward needs ∂L/∂x = ŴT(∂L/∂y) + s·ATBT(∂L/∂y), which uses dequantized Ŵ in 16-bit, plus gradients for A and B. No gradient is computed or stored for Ŵ itself. The 4-bit weights are dequantized per block both in forward and in backward (and again during checkpoint recomputation), which is where the extra time goes.

How would you add new special tokens (for example tool-call tags) when fine-tuning?

Add them to the tokenizer, call model.resize_token_embeddings(len(tokenizer)), initialize new rows sensibly (for example the mean of existing embeddings rather than random), and make the embedding and output-head matrices trainable (with LoRA, list them in modules_to_save). Save the tokenizer with the adapter. Forgetting to train the new rows leaves the model unable to emit the new tokens.

What is the alignment tax and how can it be reduced?

The drop in some capabilities or benchmark scores after alignment (for example calibration, creativity or certain NLP tasks). Reductions: mixing pretraining-style next-token loss into RL updates, stronger KL constraints, model averaging between SFT and RL checkpoints, better preference data, and targeted capability data during alignment.

How does sequence-length normalization interact with preference losses?

DPO sums log-probabilities over tokens, so longer responses produce larger log-ratio magnitudes, giving systematic advantages to length and encouraging verbosity when chosen responses tend to be longer. SimPO and some DPO variants average per token and add a target margin; evaluations use length-controlled win rates. Data-side fixes include balancing lengths between chosen and rejected.

Explain on-policy versus off-policy preference data.

On-policy data consists of responses sampled from the current policy; off-policy data came from other models or earlier checkpoints. PPO and GRPO are on-policy by design. Offline DPO usually trains off-policy, and its effectiveness drops when the data is far from what the policy would generate, because it adjusts likelihoods of sequences the model would never produce. Iterative/online DPO regenerates and relabels pairs each round to stay on-policy.

What role does the reference model play in DPO when training with LoRA?

The reference provides log πref(y|x) for chosen and rejected responses. With a LoRA policy, the reference is simply the same model with adapters disabled, so no second copy is needed; alternatively, reference log-probabilities can be precomputed once over the dataset. The reference must be the model the policy started from (usually the SFT model), otherwise the implicit reward is miscalibrated.

How do you fine-tune for long context?

Adjust rotary position scaling (position interpolation, NTK-aware or YaRN scaling) to extend the positional range, then continue training on long documents so the model adapts; use FlashAttention, sequence packing with correct attention boundaries, gradient checkpointing and sequence/context parallelism to fit memory. Evaluate with long-context retrieval and reasoning tests, and check that short-context performance did not regress. Some LoRA variants additionally train embeddings and norms for long-context adaptation.

How can fine-tuning compromise safety even with benign data, and what would you do about it?

Safety behaviour is a learned, relatively shallow layer of the model; any gradient update that pushes it towards "always comply helpfully" can erode refusals, and research has shown degradation even on benign instruction data, with very few adversarial examples enough to largely remove guardrails. Mitigations: include safety demonstrations and refusals in the training mix, use PEFT with modest learning rates, evaluate harmful-request refusal and over-refusal before and after, add inference-time moderation, and restrict who can fine-tune hosted models.

Scenario & debugging

Your fine-tuned model is great at the new task but forgot general skills. What do you do?

This is catastrophic forgetting. First quantify it with a general regression suite against the base. Then: switch from full fine-tuning to LoRA (or lower the rank), reduce learning rate and epochs, mix 5–20% general instruction data into training (replay), make sure the original chat template is used, consider interpolating fine-tuned and base weights (or scaling down the LoRA contribution), and retrain with early stopping on a validation set that includes general tasks. If the product permits, serve the task adapter only for the task's traffic and the base model for everything else.

You hit OOM fine-tuning an 8B model on a 24 GB GPU. How do you fix it?

Do the math: bf16 weights alone are 16 GB, so full fine-tuning is impossible and even bf16 LoRA is tight. Use QLoRA (about 5 GB of weights), enable gradient checkpointing, micro-batch 1–2 with gradient accumulation for the effective batch, use FlashAttention/SDPA, cap max_seq_length at the data's p95, use a paged 8-bit optimizer, and consider a chunked/fused cross-entropy since large vocabularies make logits huge. Check that nothing else occupies the GPU (a leftover model from a previous run) and free memory between runs.

After fine-tuning, the model never stops generating and repeats itself. Why?

Almost always the training targets lacked an end-of-turn/EOS token, or the EOS token was masked out of the loss or attention (common when pad = EOS and masks are built from input_ids != pad_token_id), or the serving stop tokens differ from the template's end-of-turn. Fix the data to append the correct token, build masks from lengths, configure stop tokens and the generation config, and retrain.

The training loss does not decrease at all. What do you check?

Print trainable parameters (adapters may not be attached or everything is frozen); check that the completion mask is not all zeros (for example the prompt consumed the whole max_length due to truncation); confirm labels are shifted correctly; confirm the optimizer received the trainable parameters and the scheduler is not keeping LR at zero; check the learning rate is not tiny; and decode a batch to verify the text and mask are what you expect.

Loss becomes NaN a few hundred steps into training. What happened?

Likely fp16 overflow (especially full fine-tuning in fp16), a learning rate too high without warmup, missing gradient clipping, a zero-token mask producing division by zero, or a corrupted example. Switch to bf16 or fp32 master weights, add warmup and clip gradients at 1.0, clamp the loss denominator, lower LR, and bisect to the offending batch.

Training loss is near zero but test answers are worse than the base model. Explain.

Overfitting/memorization: too many epochs on a small dataset, possibly with duplicates. The model regurgitates training answers and loses generality. Use validation-based early stopping and restore the best checkpoint, reduce epochs, deduplicate, add data diversity, lower rank or LR, and evaluate with generation metrics rather than loss.

Validation loss improved but human raters prefer the base model. Why could that be?

Loss measures likelihood of references, not quality. If references are short, bland or machine-generated, the model learns to be bland; it may drop caveats, become terse or copy reference style. Evaluation prompts may differ from real traffic. Revisit data quality, use pairwise human or judge evaluations during model selection, and consider preference tuning to optimize what raters actually prefer.

The model works in your notebook but produces garbage when served. What do you check?

Template and tokenizer mismatch (serving applies a different chat template or none), missing special tokens, a different base revision than the adapter was trained on, adapter not actually loaded, merging into a quantized model, different stop tokens, or a quantized export that lost too much precision. Run the exact serving path on the evaluation set and diff outputs against the notebook.

You have 500 labelled examples for a classification task. Fine-tune or prompt?

Start with few-shot prompting and measure on a held-out slice. If accuracy or consistency is insufficient, or per-call cost matters, fine-tune a small model with LoRA (or a small encoder with a classification head) using 400 for training and 100 for validation/test with stratification. 500 clean examples are usually enough for a narrow classification task, and a fine-tuned small model is often both cheaper and more accurate.

A support bot must follow a fixed JSON format and know policies that change monthly. Design the solution.

RAG for the policies (update the index monthly, cite sources) and structured output/constrained decoding or a fine-tune for the format and tone. Fine-tune with LoRA on transcripts where the context includes retrieved passages, so the model learns to ground answers in them and emit the schema. Evaluate format validity, grounding and correctness; retrain only when behaviour changes, not when policies change.

DPO training shows rising reward accuracy but outputs become longer and worse. What is going on?

Likely length exploitation and likelihood collapse: chosen responses are longer, so the model learns that length increases the margin, while log-probabilities of both chosen and rejected fall. Check chosen log-probs and length trends. Fixes: balance lengths in data, raise β, lower LR, fewer steps, add an SFT term on chosen, try length-normalized methods (SimPO) or IPO, and evaluate with length-controlled judges.

After an RLHF run, reward is climbing fast but responses are full of flattery and filler. What do you do?

Classic reward hacking. Stop and roll back to an earlier checkpoint, increase the KL coefficient, add a length penalty, retrain or ensemble the reward model with examples of these exploits labelled as bad, monitor judge or human win rates rather than reward alone, and cap generation length.

Your fine-tuned open model now answers harmful requests it used to refuse. Why and how do you fix it?

Fine-tuning eroded safety alignment, which can happen even with benign data. Add safety and refusal demonstrations (and benign look-alike prompts to avoid over-refusal) to the data mix, lower the learning rate or use smaller LoRA rank, run a safety evaluation suite before release, and add inference-time guardrails such as input/output moderation.

You need 100 customer-specific variants of a model. How do you train and serve them?

Train one LoRA per customer on a shared base (same revision, same template) with a common pipeline and per-customer evaluation. Serve with multi-LoRA: one base in GPU memory, adapters paged in on demand, requests routed by adapter name and batched together. Version adapters, isolate customer data during training, and give the highest-traffic customers merged dedicated deployments if latency demands it.

Fine-tuned results vary a lot between runs. How do you make conclusions reliable?

Fix seeds and data order, but also run multiple seeds and report mean and variance; use larger evaluation sets and bootstrap confidence intervals; make sure early stopping and checkpoint selection are deterministic; and only claim a method is better when the gap exceeds seed noise. Small data (hundreds of examples) is especially noisy.

A teammate wants to train on outputs from a commercial API to build a competing model. What do you flag?

Terms of service of many providers prohibit using outputs to develop competing models; this is a legal and reputational risk. Also check licences of the base model and any datasets, data-privacy obligations for prompts sent to the API, and documentation of provenance. Alternatives: use teachers whose licences permit distillation, or human-written data.

You must deploy your fine-tuned 7B model on laptops without GPUs. What is the path?

Merge the adapter into a 16-bit base, convert to GGUF, quantize to a k-quant such as Q4_K_M (about 4–4.5 GB) or Q5_K_M for more quality, verify the embedded chat template, and run with a llama.cpp-based runtime. Re-run the regression suite on the quantized file, compare against 16-bit, and consider a smaller distilled model if latency is too high. See the Edge AI page for on-device constraints.

QLoRA training is much slower than you expected. What can you do?

QLoRA pays for dequantization and gradient checkpointing recomputation. Options: use bf16 LoRA if memory allows; increase micro-batch size (better GPU utilization) while keeping memory within limits; turn off checkpointing if memory permits; enable packing to cut padding; use FlashAttention; use optimized kernels (for example Unsloth- or Liger-style fused ops); and shorten maximum sequence length.

Your fine-tuned model's benchmark score looks suspiciously high. What do you check?

Contamination: overlap between training data (including synthetic data from a teacher that may have seen the benchmark) and the test set, using n-gram and embedding similarity. Check for duplicates across splits, grouped leakage (same documents in train and test), and whether prompts or answers were used to generate training data. Re-evaluate on a fresh, private test set.

Your LoRA with r=64 performs worse than r=16. Why?

If α was held fixed, the scale α/r dropped fourfold, lowering the effective learning rate; or the larger adapter overfit a small dataset. Re-run with α/r constant (or rsLoRA scaling), compare on validation metrics across seeds, and prefer the smaller rank if there is no real gain.

You fine-tuned a base (non-instruct) model on chat data and it outputs strange role markers. What went wrong?

The base model has no chat template, so you chose one; if its special tokens were not added to the tokenizer (or were added without resizing and training embeddings and the LM head), the model sees them as fragmented text and reproduces fragments. Add the tokens properly, train the new embedding rows (modules_to_save), mask only non-assistant turns, and use the same template at inference.

You want your model to reason better on math. SFT on solutions helped a little. What next?

Use RL with verifiable rewards: GRPO (or RLOO/PPO) on problems with checkable answers, a correctness reward plus a small format reward, prompts filtered to moderate difficulty so groups have reward variance, a KL penalty to the SFT model, and evaluation on held-out math sets. Alternatively or first, distil chain-of-thought traces from a stronger reasoning model (respecting its licence).