Deep Learning & Neural Networks
Deep learning trains stacks of simple, differentiable layers to turn raw data (pixels, audio, text) into useful predictions, learning the features on its own instead of relying on hand-crafted ones. This page walks from a single neuron to CNNs, LSTMs, attention, generative models and multi-GPU training, with worked numbers, PyTorch code and the questions interviewers actually ask.
- A neural network is a composition of layers: each computes a linear map (
Wx + b) followed by a non-linear activation. Without the non-linearity, any depth collapses into one linear layer. - Training is a loop: forward pass, loss, backpropagation (chain rule, output to input) to get gradients, optimizer step (SGD, Adam, AdamW). Repeat over mini-batches for many epochs.
- Deep nets became trainable thanks to ReLU-family activations, careful initialization (Xavier/He), normalization (BatchNorm/LayerNorm), residual connections, adaptive optimizers, big data and GPUs.
- Architectures encode assumptions: CNNs assume local, translation-invariant patterns (images); RNNs/LSTMs assume sequential order; attention lets every element look at every other element and leads straight to Transformers.
- Generalization is watched through the gap between training and validation loss; the fixes are regularization (weight decay, dropout, augmentation, early stopping) and more data.
- Debugging recipe: overfit one small batch first, watch loss curves and gradient norms, and treat NaN as a learning-rate, numerics or data problem until proven otherwise.
The big picture: what deep learning is and why it works now
Artificial intelligence is the broad goal of making machines act intelligently. Machine learning is the sub-field where behaviour is learned from data instead of being written as explicit rules. Deep learning is the part of machine learning that uses neural networks with many layers. The "deep" refers to the number of stacked layers, each one transforming the output of the previous one into a slightly more abstract representation.
The key difference from classical machine learning is where the features come from. A classical pipeline for spotting cats needs a human to design features (edge histograms, color statistics, texture descriptors), and a model such as logistic regression or a random forest learns on top of them. A deep network takes raw pixels and learns the features itself: early layers discover edges, middle layers combine edges into textures and parts, and late layers respond to whole concepts such as "cat face". This is called representation learning or end-to-end learning.
Classical machine learning is like hiring a chef and handing them ingredients that someone else has already chopped, measured and labelled; the chef only decides the final combination. Deep learning is like giving the chef the raw vegetables and letting them figure out how to chop, what to combine and in what order, all while tasting the dish (the loss) and adjusting. The pre-chopped ingredients are hand-engineered features; the chef learning the preparation steps is the network learning its own intermediate representations layer by layer.
A short history, and the three reasons it took off
Neural networks are old; what changed is that they finally became trainable at scale. The history is a sequence of booms and "AI winters" (periods when funding and interest dried up because promises outran results).
| Period | Milestone | Why it mattered |
|---|---|---|
| 1940s-1950s | Mathematical model of a neuron; the perceptron, a trainable linear classifier | First idea that a machine could learn weights from examples |
| Late 1960s-1970s | Proof that a single-layer perceptron cannot learn XOR; first AI winter | Showed that single layers are too weak; nobody knew how to train multi-layer nets yet |
| 1980s | Backpropagation popularized for multi-layer networks; expert systems boom | Gave an efficient way to compute gradients for every weight in a deep network |
| 1989-1998 | Convolutional networks read handwritten digits (LeNet); LSTM invented (1997) | Showed architectures that encode structure (space, time) work far better than plain MLPs |
| 1990s-2000s | Second winter for neural nets; SVMs and boosted trees dominate | Deep nets were hard to train (vanishing gradients), data was small, compute was slow |
| 2006-2010 | Layer-wise pre-training; large labelled datasets such as ImageNet (about 1.2 million images, 1,000 classes) | Showed deep nets could be trained and gave them enough data to shine |
| 2012 | AlexNet wins ImageNet by a huge margin using GPUs, ReLU and dropout | The moment the field switched to deep learning |
| 2014-2016 | GANs, sequence-to-sequence models, attention, Adam, BatchNorm, ResNet (152 layers) | Made very deep and generative models practical |
| 2017-2020 | Transformer ("attention is all you need"), BERT, GPT-2/3, denoising diffusion | One architecture for text, vision and audio; scaling laws; generative models of photographic quality |
| 2022 onward | Chat assistants, instruction tuning, RLHF/DPO, diffusion image generators, multimodal models | Deep learning becomes a general-purpose technology |
1. Data
The internet produced billions of images, documents and recordings. Deep nets have millions to trillions of parameters and need huge datasets to avoid memorizing. Self-supervised objectives (predict the next word, fill in the masked patch) unlocked unlabelled data, which is the vast majority of all data.
2. Compute
Neural nets are mostly matrix multiplications, which are embarrassingly parallel. GPUs (thousands of simple cores) and later TPUs and NPUs made training 10-100x faster than CPUs, turning month-long experiments into hours.
3. Algorithms
ReLU activations, Xavier/He initialization, BatchNorm and LayerNorm, residual connections, dropout, Adam, better learning-rate schedules and mixed precision each removed a barrier that made deep networks untrainable.
4. Software
Frameworks with automatic differentiation (PyTorch, TensorFlow, JAX) mean you write only the forward pass; gradients come for free. Open-source pretrained models made transfer learning the default.
When deep learning is (and is not) the right tool
Deep learning shines
- Unstructured, high-dimensional inputs: images, audio, video, free text
- Lots of data, or a pretrained model you can fine-tune
- Problems where the right features are hard to describe by hand
- Generation tasks: text, images, speech, code
Classical ML often wins
- Small or medium tabular datasets (gradient-boosted trees are very strong here)
- Strict interpretability or regulatory requirements
- Tight latency, memory or power budgets with simple signals
- Short texts where bag-of-words plus a linear model already scores well
A good real-world illustration: on a balanced movie-review sentiment dataset (50,000 reviews), a well-regularized linear SVM or logistic regression on bag-of-words features reached an F1 of about 0.875, while a vanilla RNN trained without gradient clipping stayed at chance level (F1 about 0.57) and a small attention model reached about 0.86. Deep learning is not automatically better; it needs the right architecture, tricks and data. That case study appears again in the RNN section.
Neurons, perceptrons and multi-layer networks
The building block is the artificial neuron, loosely inspired by biology. A biological neuron receives signals on its dendrites, integrates them in the cell body, and fires along its axon if the total exceeds a threshold; synapses strengthen or weaken with learning. The artificial version keeps only the arithmetic.
Think of a hiring manager scoring a candidate. Each input is a piece of evidence (years of experience, interview score, referral). The weights are how much the manager cares about each piece, and some can be negative (a red flag lowers the score). The bias is the manager's general mood: a strict manager starts with a negative bias and needs more evidence. The activation is the final decision rule, such as "hire if the score is above zero". Training is the manager revising how much each piece of evidence should count after seeing how past hires performed.
The perceptron and its limits
The perceptron is a single neuron with a step activation: output 1 if z > 0, else 0. Its decision boundary w·x + b = 0 is a straight line (a hyperplane in higher dimensions), so it can only solve linearly separable problems. AND and OR are separable; XOR (output 1 when exactly one input is 1) is not, because no single line separates (0,1) and (1,0) from (0,0) and (1,1). Adding a hidden layer fixes this: one hidden neuron can compute OR, another NAND, and the output neuron ANDs them together.
XOR is not linearly separable A 2-2-1 network solves it
x2 x1 ---+--> [h1 = OR ] ---+
1 | (1) (0) \ / \
| \/ [ y = AND(h1,h2) ]
| /\ /
0 | (0) (1) x2 ---+--> [h2 = NAND] ---+
+--------------------- x1
0 1 Hidden layer bends the space so
No single line separates the output layer can use a line.
the 1s from the 0s.
The multi-layer perceptron (MLP)
An MLP (also called a feedforward or fully connected network) stacks layers of neurons. Every neuron in one layer connects to every neuron in the next. Information flows one way, input to output, with no loops. A layer with n inputs and m outputs is just a matrix multiplication plus a bias vector, followed by an element-wise activation.
Input layer Hidden layer 1 Hidden layer 2 Output layer
(features) (4 units, ReLU) (3 units, ReLU) (softmax)
x1 o-----------> o -----------------> o ------------> o P(cat)
\ \ / / o o o P(dog)
x2 o---X-X-------> o -----------------> o ------------> o P(bird)
/ / \ \ o
x3 o-------------->
every arrow = one learnable weight; every unit also has a bias
Parameters: (3*4+4) + (4*3+3) + (3*3+3) = 16 + 15 + 12 = 43
Why non-linearity is essential
If every activation were the identity, two layers would compute W2(W1x + b1) + b2 = (W2W1)x + (W2b1 + b2), which is a single linear layer. A hundred linear layers are still one linear layer. The activation "bends" the space between layers, so stacked layers can carve curved decision boundaries and compose features. Real-world problems are not turned into linear ones; instead, the network builds a non-linear function out of many simple linear pieces joined by non-linearities. With ReLU, the network is literally a piecewise-linear function with a huge number of pieces.
Universal approximation, and why depth still matters
The universal approximation theorem says a network with one hidden layer and a suitable non-linear activation can approximate any continuous function on a bounded region to any accuracy, given enough hidden units. It is an existence result: it says nothing about how many units you need or whether gradient descent will find the weights. In practice, depth is far more efficient than width: deep networks reuse intermediate features (edges to parts to objects), and some functions need exponentially many units in a shallow network but only polynomially many in a deep one.
Worked example: one forward pass by hand
Take a 2-2-1 network with ReLU hidden units and a sigmoid output. Input x = [1, 2].
- Hidden neuron 1 weights [0.5, −0.2], bias 0.1: z1 = 0.5×1 + (−0.2)×2 + 0.1 = 0.2, so h1 = ReLU(0.2) = 0.2.
- Hidden neuron 2 weights [0.3, 0.4], bias −0.5: z2 = 0.3 + 0.8 − 0.5 = 0.6, so h2 = 0.6.
- Output neuron weights [1.0, −0.5], bias 0.1: z = 1.0×0.2 + (−0.5)×0.6 + 0.1 = 0.0.
- Output activation: ŷ = σ(0) = 1/(1 + e0) = 0.5. The model is exactly undecided; training will push the weights so that ŷ moves toward the true label.
The numbers inside hidden nodes (0.2, 0.6) are activations, not weights. Weights live on the connections. Weights can be any real numbers, do not need to sum to 1, and become decimals simply because they are updated continuously.
Matrix form and batching
Real code processes a mini-batch of B examples at once. Stack the inputs as rows of a matrix X of shape (B, n). Then a layer is H = f(X WT + b), where the bias is broadcast across rows. This is why GPUs matter: a whole batch through a layer is one large matrix multiplication.
import torch
import torch.nn as nn
class MLP(nn.Module):
def __init__(self, in_dim=20, hidden=64, n_classes=3, p_drop=0.1):
super().__init__()
self.net = nn.Sequential(
nn.Linear(in_dim, hidden), # W1: (64, 20), b1: (64,)
nn.ReLU(),
nn.Dropout(p_drop),
nn.Linear(hidden, hidden),
nn.ReLU(),
nn.Linear(hidden, n_classes) # raw scores (logits), no softmax here
)
def forward(self, x): # x: (batch, in_dim)
return self.net(x) # (batch, n_classes)
model = MLP()
x = torch.randn(32, 20) # a batch of 32 examples
logits = model(x)
print(logits.shape) # torch.Size([32, 3])
print(sum(p.numel() for p in model.parameters())) # 20*64+64 + 64*64+64 + 64*3+3 = 5,699
in × out + out for a dense layer, and always say that the theorem guarantees existence, not learnability or efficiency.Activation functions
An activation function is applied element-wise after each linear layer. It provides the non-linearity, and its derivative controls how well gradients flow backward. Choosing one is about three things: does it add non-linearity, does it keep gradients healthy, and is it cheap to compute.
Activations are like different kinds of valves on a water pipe. A sigmoid is a valve that is fully shut or fully open at the extremes and barely responds to extra pressure there, so turning the handle further does nothing (saturation). ReLU is a one-way check valve: no flow backward at all, full proportional flow forward. Leaky ReLU lets a trickle through backward so the valve never seizes. The pressure change you feel when you turn the handle is the derivative; if the valve stops responding, the learning signal (gradient) stops too.
The activation zoo
| Activation | Formula | Output range | Derivative | Pros | Cons | Typical use |
|---|---|---|---|---|---|---|
| Step | 1 if z > 0 else 0 | {0, 1} | 0 almost everywhere | Simple threshold | No gradient, so cannot be trained with backprop | Original perceptron only |
| Sigmoid | σ(z) = 1/(1 + e−z) | (0, 1) | σ(z)(1 − σ(z)), max 0.25 | Probability interpretation, smooth | Saturates (vanishing gradients), not zero-centered, exp is costly | Binary output layer, gates in LSTM/GRU |
| Tanh | (ez − e−z)/(ez + e−z) | (−1, 1) | 1 − tanh2(z), max 1 | Zero-centered, stronger gradient than sigmoid | Still saturates at extremes | RNN hidden state, LSTM cell output |
| ReLU | max(0, z) | [0, ∞) | 1 if z > 0, else 0 | Cheap, no saturation for z > 0, sparse activations | "Dying ReLU": units stuck at 0 never recover; not zero-centered | Default for CNNs and MLPs |
| Leaky ReLU / PReLU | z if z > 0 else αz (α ≈ 0.01, learned in PReLU) | (−∞, ∞) | 1 or α | Keeps a small gradient for negatives, no dead units | Extra hyperparameter; small gains in practice | GAN discriminators, some CNNs |
| ELU | z if z > 0 else α(ez − 1) | (−α, ∞) | 1 or ELU(z) + α | Smooth, mean activations closer to zero | Exp is costlier than ReLU | Some MLPs and CNNs |
| GELU | z · Φ(z), Φ = standard normal CDF | about (−0.17, ∞) | Smooth, non-monotonic near 0 | Smooth "probabilistic ReLU"; small negatives pass a little gradient | Slightly more compute | Transformers (BERT, GPT, ViT) |
| SiLU / Swish | z · σ(z) | about (−0.28, ∞) | Smooth | Self-gated, often beats ReLU on deep nets | Slightly more compute | EfficientNet; SwiGLU feed-forward blocks in modern LLMs |
| Softplus | ln(1 + ez) | (0, ∞) | σ(z) | Smooth ReLU, always positive | Slower | Predicting positive quantities (variances, rates) |
| Softmax | ezi / Σj ezj | (0, 1), sums to 1 | Jacobian: si(δij − sj) | Turns a score vector into a probability distribution | Operates on a whole vector; not used between hidden layers | Multi-class output; attention weights |
Worked numbers
- σ(0) = 0.5, σ(2) = 1/(1 + e−2) = 1/1.135 ≈ 0.881, σ(−4) ≈ 0.018. The derivative at z = 4 is 0.982 × 0.018 ≈ 0.018: almost no learning signal.
- ReLU(3) = 3, ReLU(−2) = 0. ReLU does not squash to [0, 1]; its output is unbounded above.
- GELU(−1) ≈ −0.159 while ReLU(−1) = 0; GELU(2) ≈ 1.955, nearly identical to ReLU for positive inputs.
- Softmax of logits [2, 1, 0]: e2 = 7.389, e1 = 2.718, e0 = 1, sum 11.107, giving [0.665, 0.245, 0.090]. Higher score, higher probability, and everything sums to 1.
Why ReLU is non-linear even though it looks linear
ReLU is linear on each side of zero, but the kink at zero breaks global linearity. Different neurons switch on for different inputs, so the network becomes a patchwork of many linear regions whose boundaries depend on the input. That patchwork can approximate any curve. ReLU is not differentiable exactly at 0; frameworks simply use a subgradient (0 or 1) there, and since landing exactly on 0 is rare with real-valued inputs, it causes no practical problem.
The dying ReLU problem
If a neuron's pre-activation becomes negative for every input (often after a large gradient step pushes its bias very negative), it outputs 0 and its gradient is 0, so its weights never update again. Fixes: Leaky ReLU/PReLU/ELU/GELU, He initialization, a lower learning rate, and normalization layers that keep pre-activations centered.
Quick decision guide
Hidden layers
- CNNs and MLPs: ReLU (or Leaky ReLU, SiLU)
- Transformers: GELU, or SiLU inside SwiGLU
- RNN/LSTM/GRU: tanh for states, sigmoid for gates
- GAN discriminators: Leaky ReLU
Output layer (depends on the task)
- Regression: none (identity)
- Bounded regression: tanh or sigmoid, rescaled
- Binary or multi-label classification: sigmoid per output
- Multi-class (exactly one label): softmax
import torch
import torch.nn.functional as F
z = torch.tensor([-2.0, -0.5, 0.0, 0.5, 2.0])
print(F.relu(z)) # tensor([0.0, 0.0, 0.0, 0.5, 2.0])
print(torch.sigmoid(z)) # tensor([0.119, 0.378, 0.500, 0.622, 0.881])
print(torch.tanh(z)) # tensor([-0.964, -0.462, 0.000, 0.462, 0.964])
print(F.gelu(z)) # small negative values survive
print(F.silu(z))
print(F.softmax(z, dim=0)) # sums to 1.0
nn.CrossEntropyLoss. That loss already applies log-softmax internally; applying softmax twice flattens the probabilities and slows or breaks training. Output raw logits during training and apply softmax only when you need probabilities at inference time. The same holds for nn.BCEWithLogitsLoss and sigmoid.Loss functions
The loss function turns "how wrong is the model?" into a single number that gradient descent can minimize. Strictly, a loss is the error on one example and the cost (or objective) is the average over the dataset or batch, often plus a regularization term. In practice and in libraries, "loss" is used for both, and nn.MSELoss() returns the batch mean by default.
A loss function is the scoring rule in a game. A darts scorer who squares the distance from the bullseye (MSE) punishes wild throws brutally, so players learn to avoid big misses. A scorer who counts plain distance (MAE) treats every centimetre equally and is not thrown off by one throw that hits the wall. A weather forecaster judged by cross-entropy is punished most when they say "0% chance of rain" and it pours: confident mistakes cost the most. The scoring rule you choose decides what the player (the model) learns to care about.
Regression losses
| Loss | Formula (per example) | Behaviour | Use when |
|---|---|---|---|
| MSE (L2) | (y − ŷ)2 | Quadratic: an error of 10 costs 100, an error of 50,000 costs 2.5 billion. Smooth, gradient shrinks near the minimum. Sensitive to outliers. Corresponds to maximum likelihood under Gaussian noise; predicts the conditional mean. | Default regression; large errors are genuinely worse |
| MAE (L1) | |y − ŷ| | Linear: robust to outliers; predicts the conditional median; gradient has constant magnitude, not smooth at 0 | Noisy targets with outliers |
| Huber (smooth L1) | ½e2 if |e| ≤ δ, else δ(|e| − ½δ) | Quadratic near zero, linear far away; best of both | Regression with some outliers; bounding-box regression |
| Quantile (pinball) | max(τe, (τ − 1)e), e = y − ŷ | Asymmetric: penalizes under- and over-prediction differently | Forecasting where over- and under-forecasting have different business costs; prediction intervals |
Business costs should drive the choice. If the true demand is 100 and model A predicts 120 while model B predicts 80, both have the same squared and absolute error, yet over-forecasting may mean wasted inventory while under-forecasting means lost sales. An asymmetric loss such as quantile loss encodes that preference directly.
Classification losses
Worked example. A spam email (y = 1). If the model predicts p = 0.9, BCE = −ln(0.9) ≈ 0.105 (small). If it predicts p = 0.1, BCE = −ln(0.1) ≈ 2.303. If it confidently says p = 0.05, BCE = −ln(0.05) ≈ 3.0, which is "very high". As p approaches 0 for a true positive, the loss goes to infinity. Correct and confident gives low loss; wrong and confident is punished hardest.
| Loss | Formula | Use when |
|---|---|---|
| Binary cross-entropy | see above | Binary classification, multi-label (one sigmoid per label) |
| Categorical cross-entropy | −log pcorrect | Multi-class, one label per example; paired with softmax. Labels are integers or one-hot vectors, never strings. |
| Hinge | max(0, 1 − y·s), y ∈ {−1, +1}, s = raw score | Max-margin classifiers (SVMs); zero loss once an example is correct with margin at least 1 |
| Focal | −α(1 − pt)γ log(pt), γ ≈ 2 | Heavy class imbalance (dense object detection, where most boxes are background). Down-weights easy, well-classified examples so hard ones dominate the gradient. |
| Label-smoothed CE | Replace the one-hot target with y' = (1 − ε)y + ε/K (PyTorch / Szegedy): true class (1 − ε) + ε/K, others ε/K. For ε = 0.1 and K = 10 this is 0.91 / 0.01, not 0.9 / 0.011. | Prevents over-confidence, improves calibration; standard in image classifiers and translation models. A close variant puts leftover mass only on the other classes: (1 − ε) and ε/(K − 1). |
| KL divergence | Σ p log(p/q) | Matching distributions: knowledge distillation, VAE regularizer, RLHF penalty |
| Contrastive / triplet / InfoNCE | Pull similar pairs together, push dissimilar apart (with a margin or softmax over negatives) | Embedding learning: face recognition, retrieval, CLIP-style image-text models |
| CTC | Sum over all valid alignments | Speech and handwriting recognition where input and output lengths differ |
Why cross-entropy and not MSE for classification?
- Probabilistic meaning: cross-entropy is the negative log-likelihood of a categorical (or Bernoulli) model, so minimizing it is maximum likelihood estimation.
- Better gradients: with a sigmoid or softmax output, the gradient of cross-entropy with respect to the logits is simply
p − y. With MSE, the gradient is multiplied by σ'(z), which is near zero when the model is confidently wrong, so learning stalls exactly when it is needed most. - Convexity in the last layer: logistic regression with cross-entropy is convex; with MSE it is not.
Entropy, cross-entropy and KL in one line each
- Entropy H(p) = −Σ p log p: how unpredictable a distribution is (a fair coin has 1 bit).
- Cross-entropy H(p, q) = −Σ p log q: the average cost of encoding outcomes from the true distribution p using the model's distribution q.
- KL divergence KL(p || q) = H(p, q) − H(p): the extra cost of using q instead of p. With one-hot labels H(p) = 0, so cross-entropy equals KL.
Calibration and confidence
A model's "confidence" is the probability it assigns to its predicted class. A model is well calibrated if predictions made with 80% confidence are right about 80% of the time. Accuracy measures correctness; calibration measures whether the probabilities can be trusted. Modern deep nets tend to be over-confident. Temperature scaling (dividing logits by a learned T > 1 on a validation set) fixes calibration without changing the ranking or accuracy. Label smoothing also helps. The decision threshold (0.5 by default) is separate from the loss and should be tuned on validation data for the precision/recall trade-off you need.
What is a "good" loss value?
There is no universal target. A loss value depends on the loss type, the number of classes and the noise in the data. Useful reference points: for K balanced classes, a model that predicts uniformly has cross-entropy ln(K) (0.693 for 2 classes, 2.303 for 10). Judge progress by: loss decreasing over time, training and validation loss staying close, and beating a sensible baseline. Do not train until the training loss reaches 0: real data contains noise, and driving training loss to zero usually means memorization.
import torch, torch.nn as nn
logits = torch.tensor([[2.0, 1.0, 0.0]]) # one example, 3 classes
target = torch.tensor([0]) # class index, not one-hot
print(nn.CrossEntropyLoss()(logits, target)) # -log(0.665) = 0.408
bin_logit = torch.tensor([-2.944]) # sigmoid(-2.944) = 0.05
print(nn.BCEWithLogitsLoss()(bin_logit, torch.tensor([1.0]))) # about 3.0
# class imbalance: weight the rare class more
loss_fn = nn.CrossEntropyLoss(weight=torch.tensor([1.0, 5.0, 1.0]),
label_smoothing=0.1)
nn.BCELoss on raw logits, or computing log(sigmoid(x)) yourself. Both are numerically unstable (log of 0 gives infinity, then NaN). Use the "with logits" versions, which fuse the activation and the log with the log-sum-exp trick.p − y is a strong signal.Gradient descent and backpropagation
Training means finding weights that minimize the loss. The loss surface of a neural network is a function of millions of weights, non-convex, full of saddle points and flat regions. We cannot solve for the minimum directly, so we walk downhill using the gradient: the vector of partial derivatives of the loss with respect to every weight. It points in the direction of steepest increase, so we step the opposite way.
Gradient descent is hiking down a mountain in thick fog. You cannot see the valley; you can only feel the slope under your feet, so you take a step in the steepest downhill direction and repeat. The step length is the learning rate: too short and you will be walking all night, too long and you overshoot and end up on the opposite slope, higher than before. You may stop in a small hollow (a local minimum) or on a flat plateau (a saddle point) rather than the deepest valley. Backpropagation is how you feel the slope for every one of millions of directions at once.
Batch, stochastic and mini-batch gradient descent
| Variant | Gradient computed on | Pros | Cons |
|---|---|---|---|
| Batch (full) GD | The entire training set per step | Exact gradient, smooth convergence | Very slow per step, does not fit in memory for big data |
| Stochastic GD | One example per step | Cheap steps, noise helps escape saddle points | Very noisy, poor hardware utilization |
| Mini-batch GD | B examples (typically 32-4,096) | Good GPU utilization, moderate noise that also regularizes | Batch size becomes a hyperparameter |
Terminology: an iteration (or step) is one parameter update on one mini-batch; an epoch is one full pass over the training set. With 50,000 examples and batch size 100, one epoch is 500 iterations. "SGD" in libraries usually means mini-batch SGD.
Backpropagation is the chain rule, applied efficiently
A network is a composition of functions, so the loss depends on an early weight through a chain of intermediate values. The chain rule says: if L depends on ŷ, which depends on z, which depends on w, then
- Forward pass Feed the input through the network, computing and caching every intermediate value (pre-activations z and activations a). Without a prediction there is no loss and no learning.
- Compute the loss Compare the prediction with the label.
- Backward pass Starting from ∂L/∂L = 1, walk the computation graph in reverse. At each node multiply the incoming (upstream) gradient by the node's local derivative and pass it to the node's inputs. Gradients arriving from multiple paths are summed.
- Update After all gradients are known, the optimizer updates every weight in all layers together. Gradients are computed output-first, but the updates happen simultaneously in one step.
Backprop and the optimizer are different things: backprop only computes gradients; the optimizer (SGD, Adam...) uses them to update weights. Both happen only during training, never during validation or inference.
Why backprop is fast
Computing each gradient separately by nudging one weight and re-running the network (finite differences) costs one forward pass per weight: O(W) passes for W weights, hopeless for a billion parameters. Backprop (reverse-mode automatic differentiation) gets all W gradients for roughly the cost of two to three forward passes. The price is memory: every intermediate activation from the forward pass must be stored for the backward pass, which is why activation memory often dominates training memory.
Worked example: one full training step by hand
A tiny network: one input, one ReLU hidden unit, one sigmoid output, binary cross-entropy loss. Input x = 2.0, label y = 1. Parameters w1 = 0.3, b1 = 0.1, w2 = 0.5, b2 = 0.2, learning rate η = 0.1.
x=2.0 --(w1=0.3, b1=0.1)--> z1=0.7 --ReLU--> h=0.7 --(w2=0.5, b2=0.2)--> z2=0.55 --sigmoid--> y_hat=0.634
|
L = -ln(0.634) = 0.456
| Step | Computation | Value |
|---|---|---|
| Forward: hidden pre-activation | z1 = w1x + b1 = 0.3×2 + 0.1 | 0.7 |
| Forward: hidden activation | h = ReLU(0.7) | 0.7 |
| Forward: output pre-activation | z2 = w2h + b2 = 0.5×0.7 + 0.2 | 0.55 |
| Forward: prediction | ŷ = σ(0.55) = 1/(1 + e−0.55) | 0.634 |
| Loss | L = −ln(0.634) | 0.456 |
| Backward: output error | ∂L/∂z2 = ŷ − y (sigmoid + BCE simplify) | −0.366 |
| Gradient for w2 | ∂L/∂w2 = ∂L/∂z2 · h = −0.366 × 0.7 | −0.256 |
| Gradient for b2 | ∂L/∂b2 = ∂L/∂z2 | −0.366 |
| Back through w2 | ∂L/∂h = ∂L/∂z2 · w2 = −0.366 × 0.5 | −0.183 |
| Back through ReLU | ∂L/∂z1 = ∂L/∂h · ReLU'(0.7) = −0.183 × 1 | −0.183 |
| Gradient for w1 | ∂L/∂w1 = ∂L/∂z1 · x = −0.183 × 2 | −0.366 |
| Gradient for b1 | ∂L/∂b1 = ∂L/∂z1 | −0.183 |
| Update | w2 = 0.5 + 0.0256; b2 = 0.2 + 0.0366; w1 = 0.3 + 0.0366; b1 = 0.1 + 0.0183 | 0.526, 0.237, 0.337, 0.118 |
| New forward pass | z1 = 0.792, h = 0.792, z2 = 0.653, ŷ = 0.658 | L = 0.419 (down from 0.456) |
All gradients were negative, so every parameter increased, pushing ŷ toward the target 1. Notice that the size of each weight's update is proportional to the activation feeding into it (w1 got a bigger gradient than b1 because x = 2): connections that contributed more to the error receive more "blame". If the hidden unit had been a sigmoid instead of a ReLU, the backward pass would have multiplied by σ'(z1) ≤ 0.25, which is the seed of the vanishing-gradient problem.
Autograd does this for you
import torch
x, y = torch.tensor(2.0), torch.tensor(1.0)
w1 = torch.tensor(0.3, requires_grad=True); b1 = torch.tensor(0.1, requires_grad=True)
w2 = torch.tensor(0.5, requires_grad=True); b2 = torch.tensor(0.2, requires_grad=True)
h = torch.relu(w1 * x + b1)
logit = w2 * h + b2
loss = torch.nn.functional.binary_cross_entropy_with_logits(logit, y)
loss.backward() # fills .grad on every leaf tensor
print(round(loss.item(), 3)) # 0.456
print(w1.grad, b1.grad, w2.grad, b2.grad) # -0.366, -0.183, -0.256, -0.366
with torch.no_grad(): # manual SGD step
for p in (w1, b1, w2, b2):
p -= 0.1 * p.grad
p.grad = None # gradients accumulate unless cleared
Local minima, saddle points and "good enough"
Neural network losses are non-convex, so there is no guarantee of reaching the global minimum. In high dimensions, most critical points are saddle points (downhill in some directions, uphill in others) rather than bad local minima, and research suggests that most local minima of large networks have loss close to the global minimum. Mini-batch noise and momentum help escape saddles. Success is judged not by the lowest training loss but by low, stable validation loss. Convex problems like linear regression have a single global minimum; deep learning aims for robust generalization, not exact optimality. Flat minima (wide valleys) tend to generalize better than sharp ones.
Gradient checking
When you implement a custom layer or backward function, compare the analytic gradient with a numerical one: (L(w + ε) − L(w − ε)) / 2ε with ε ≈ 1e−5, in double precision. A relative error below about 1e−6 means the gradient is correct. PyTorch provides torch.autograd.gradcheck.
optimizer.zero_grad(). PyTorch accumulates gradients into .grad by design (useful for gradient accumulation), so without clearing them each step adds to the previous step's gradients and the effective learning rate explodes.Optimizers: SGD, momentum, RMSProp, Adam, AdamW
Plain gradient descent uses one global learning rate and only the current gradient. That causes three problems: it zig-zags across narrow valleys (steep in one direction, shallow in another), it crawls across flat regions and saddle points, and it treats rarely-updated parameters the same as frequently-updated ones. Modern optimizers fix these by remembering past gradients (momentum) and by scaling the step per parameter (adaptive learning rates).
Plain SGD is a hiker who re-decides direction at every step based only on the ground under one foot, so on a narrow ravine they zig-zag from wall to wall. Momentum turns the hiker into a heavy ball rolling downhill: it builds up speed along the consistent direction and the side-to-side wobbles cancel out. RMSProp gives the hiker adaptive shoes that take small steps on steep, bumpy ground and longer strides on gentle slopes. Adam is the heavy ball wearing adaptive shoes. AdamW additionally shrinks the hiker's backpack (the weights) a little every step, independent of the terrain, which is proper weight decay.
The update rules
| Optimizer | Update (g = current gradient) | Idea | Typical settings |
|---|---|---|---|
| SGD | θ ← θ − ηg | Step against the gradient | η = 0.01-0.1 for CNNs with schedules |
| SGD + momentum | v ← μv + g; θ ← θ − ηv | Exponential moving average of gradients: accelerates in consistent directions, damps oscillation | μ = 0.9 |
| Nesterov momentum | Gradient evaluated at the look-ahead point θ − ημv | Corrects the step before overshooting; slightly better convergence | μ = 0.9, nesterov=True |
| AdaGrad | s ← s + g2; θ ← θ − ηg/(√s + ε) | Per-parameter rates; rare features get bigger steps | Good for sparse features; rate decays to zero over time |
| RMSProp | s ← ρs + (1 − ρ)g2; θ ← θ − ηg/(√s + ε) | AdaGrad with a moving average so the rate does not vanish | ρ = 0.9-0.99, η = 1e−3; popular for RNNs |
| Adam | m ← β1m + (1 − β1)g; v ← β2v + (1 − β2)g2; m̂ = m/(1 − β1t); v̂ = v/(1 − β2t); θ ← θ − ηm̂/(√v̂ + ε) | Momentum (first moment) plus RMSProp (second moment) with bias correction | η = 1e−3 (3e−4 is a common safe value), β1 = 0.9, β2 = 0.999, ε = 1e−8 |
| AdamW | Adam step, then θ ← θ − ηλθ (decoupled) | Weight decay applied directly to weights, not mixed into the adaptive gradient | λ = 0.01-0.1; default for Transformers and LLMs |
Why Adam needs bias correction
m and v start at zero, so in the first steps they are biased toward zero (after step 1, m = 0.1g with β1 = 0.9). Dividing by (1 − βt) rescales them so early estimates are unbiased; the correction fades as t grows. Without it, the first updates would be badly scaled, especially the second-moment estimate with β2 = 0.999.
Adam versus AdamW: the weight-decay subtlety
With plain SGD, adding an L2 penalty (λ/2)||θ||2 to the loss is exactly the same as weight decay. With Adam, it is not: the L2 gradient λθ gets divided by √v̂ like every other gradient, so weights with large historical gradients are barely regularized. AdamW decouples the decay, shrinking every weight by the same fraction each step. This gives more predictable regularization and better generalization, which is why AdamW is the standard for Transformers. Biases and normalization parameters are usually excluded from weight decay.
Which optimizer should you use?
Adam / AdamW
- Works well with little tuning; the default starting point
- Essential for Transformers, sparse gradients, embeddings
- Costs two extra state tensors per parameter (m and v): more memory
- Can generalize slightly worse than tuned SGD on some vision tasks
SGD + momentum
- Often reaches the best final accuracy on CNN image classification with a good schedule
- Only one extra state tensor (velocity)
- Needs more careful tuning of learning rate and schedule
- Slower in the early phase
Other names you may hear: LAMB/LARS (layer-wise adaptive rates for very large batch sizes), Adafactor (factorized second moments to save memory), Lion (sign-based update, less memory), Sophia and Shampoo (approximate second-order information). Second-order methods such as Newton's method use curvature (the Hessian) but are too expensive for millions of parameters.
import torch
model = torch.nn.Linear(10, 2)
sgd = torch.optim.SGD(model.parameters(), lr=0.1, momentum=0.9, nesterov=True, weight_decay=5e-4)
adam = torch.optim.Adam(model.parameters(), lr=3e-4)
adamw = torch.optim.AdamW(model.parameters(), lr=3e-4, betas=(0.9, 0.999), weight_decay=0.01)
# exclude biases and norm parameters from weight decay (common for Transformers)
decay, no_decay = [], []
for name, p in model.named_parameters():
(no_decay if p.ndim == 1 else decay).append(p)
opt = torch.optim.AdamW([{"params": decay, "weight_decay": 0.01},
{"params": no_decay, "weight_decay": 0.0}], lr=3e-4)
Learning rate, schedules and warmup
The learning rate controls step size. Too high: the loss oscillates, spikes or becomes NaN. Too low: training is painfully slow and may settle in a poor region. The best value also changes during training: early on you want big steps to make fast progress, later small steps to settle into a minimum. A schedule changes the learning rate over time.
Parking a car: you pull out of the driveway slowly while the engine warms up (warmup), drive fast on the highway to cover distance (peak learning rate), then slow down gradually as you approach the destination and creep into the parking spot (decay). Driving at highway speed into the parking spot overshoots it; creeping the whole way never gets you there. Warmup is the gentle start, cosine decay is the smooth braking, and the final small learning rate is the careful parking manoeuvre.
| Schedule | Shape | Where used |
|---|---|---|
| Constant | Flat | Baselines, quick experiments, some fine-tuning |
| Step decay | Divide by 10 every N epochs (for example at epochs 30, 60, 90) | Classic CNN recipes (ResNet on ImageNet) |
| Exponential decay | ηt = η0γt | Older recipes, RL |
| Reduce on plateau | Cut by a factor when validation loss stops improving for "patience" epochs | When you do not know the training length in advance |
| Cosine annealing | ηt = ηmin + ½(ηmax − ηmin)(1 + cos(πt/T)) | Default for modern vision and many other tasks |
| Linear warmup + cosine (or linear) decay | Ramp from ~0 to peak over the first 1-5% of steps, then decay | Every major Transformer and LLM |
| Warm restarts (SGDR) / cyclical | Periodically jump back up | Escaping sharp minima, snapshot ensembles |
| One-cycle | Warm up to a high peak, then anneal below the start | Fast CNN training ("super-convergence") |
| Inverse square root | η ∝ 1/√step after warmup | Original Transformer |
Why warmup exists
At step 0 the weights are random, the gradients can be large and poorly aligned, and Adam's second-moment estimates are based on very few samples, so they are noisy. A full learning rate at that moment can push the model into a bad region or cause immediate divergence (loss spikes to NaN). Warming up linearly over a few hundred to a few thousand steps lets the statistics stabilize. Warmup matters most for Transformers, large batch sizes and adaptive optimizers.
Finding a good learning rate
- LR range test Train for a few hundred steps while increasing the learning rate exponentially from 1e−7 to 1. Plot loss against learning rate (log scale).
- Pick a value about 10x below the point where the loss is falling fastest or just before it blows up.
- Search on a log scale If tuning by hand, try 1e−4, 3e−4, 1e−3, 3e−3 rather than evenly spaced values.
- Scale with batch size The linear scaling rule: when you multiply the batch size by k, multiply the SGD learning rate by roughly k (with warmup). For Adam, square-root scaling is a common heuristic.
- Fine-tuning Use a learning rate 10-100x smaller than for training from scratch (for example 1e−5 to 5e−5 for BERT-style models), so pretrained features are not destroyed.
import math, torch
opt = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=0.01)
total_steps, warmup = 10_000, 500
def lr_lambda(step):
if step < warmup:
return step / warmup
progress = (step - warmup) / max(1, total_steps - warmup)
return 0.5 * (1 + math.cos(math.pi * progress)) # multiplier on base lr
sched = torch.optim.lr_scheduler.LambdaLR(opt, lr_lambda)
# in the loop: loss.backward(); opt.step(); sched.step(); opt.zero_grad()
# built-in alternatives
# torch.optim.lr_scheduler.CosineAnnealingLR(opt, T_max=total_steps)
# torch.optim.lr_scheduler.OneCycleLR(opt, max_lr=1e-3, total_steps=total_steps)
# torch.optim.lr_scheduler.ReduceLROnPlateau(opt, factor=0.1, patience=3) # call .step(val_loss)
ReduceLROnPlateau needs the validation metric passed to step(). Log the learning rate every step so you can see the schedule you actually ran.Weight initialization
Before training, weights must be set to something. The choice matters more than it seems: bad initialization makes activations and gradients shrink or grow exponentially with depth, so the network either learns nothing or blows up on the first step.
Think of the game of telephone played by a line of people, where each person repeats the message at a volume set by their personal "gain". If everyone speaks a bit quieter than they heard, after 50 people the message is inaudible; if everyone speaks a bit louder, it becomes deafening noise. Good initialization picks each person's gain so the volume stays roughly constant down the line. And if every person were identical and started with the same gain, they would all behave the same forever, which is why we randomize.
Why not all zeros (or all the same value)?
If every weight in a layer starts equal, every neuron in that layer computes the same output, receives the same gradient and gets the same update. They stay identical forever: a layer of 512 units behaves like one unit. This is the symmetry problem. Random initialization breaks the symmetry so different neurons learn different features. Biases can safely start at zero because the random weights already break symmetry.
Why not just "small random numbers"?
The variance of a neuron's output is roughly fan_in × Var(w) × Var(x). If Var(w) is too small, the signal shrinks by a constant factor each layer and vanishes; too large and it explodes. The fix is to choose Var(w) so that the variance of activations (forward) and gradients (backward) stays roughly constant from layer to layer.
| Scheme | Variance of weights | Designed for | PyTorch |
|---|---|---|---|
| Xavier | Var(W) = 2 / (fan_in + fan_out); uniform limit √(6/(fan_in + fan_out)) | tanh, sigmoid, linear (symmetric activations) | nn.init.xavier_uniform_ |
| He | Var(W) = 2 / fan_in | ReLU family (ReLU zeroes half the inputs, so it needs twice the variance) | nn.init.kaiming_normal_(w, nonlinearity="relu") |
| Fan-in scaling | Var(W) = 1 / fan_in | SELU, linear | via normal_ |
| Orthogonal | Random orthogonal matrix, preserves norms exactly | RNN recurrent weights | nn.init.orthogonal_ |
| Small-std normal (0.02) plus scaled residual projections | std 0.02, output projections scaled by 1/√(2 × layers) | GPT-style Transformers | custom init |
Special cases worth knowing: initialize the LSTM forget-gate bias to 1 so the cell remembers by default early in training; initialize the last BatchNorm γ in each residual block to 0 so each block starts as an identity function; initialize a classifier's final bias to the log of class priors on imbalanced data so the initial loss is sensible.
import torch.nn as nn
def init_weights(m):
if isinstance(m, (nn.Linear, nn.Conv2d)):
nn.init.kaiming_normal_(m.weight, nonlinearity="relu")
if m.bias is not None:
nn.init.zeros_(m.bias)
model.apply(init_weights) # PyTorch layers already use a He-uniform variant by default
Overfitting, underfitting and regularization
The goal is not to fit the training data; it is to perform well on data the model has never seen (generalization). Deep nets have enough capacity to memorize random labels perfectly, so controlling overfitting is a core skill.
Three students prepare for an exam. The underfitter skimmed the chapter titles and does badly on both the practice questions and the real exam (high bias). The overfitter memorized the answers to every practice question word for word, scores 100% on practice and fails the real exam because the questions are phrased differently (high variance). The good student understood the concepts and scores well on both. Regularization is the teacher who forces students to explain answers in their own words, removes random pages from their notes (dropout), shuffles and rephrases the practice questions (data augmentation) and stops the cramming session when mock-exam scores stop improving (early stopping).
Diagnosing with training and validation loss
| Situation | Training loss | Validation loss | What to do |
|---|---|---|---|
| Underfitting (high bias) | High | High, close to training | Bigger model, train longer, better features or architecture, less regularization, higher learning rate or fix optimization bugs |
| Good fit | Low | Low, close to training | Ship it; check on the test set once |
| Overfitting (high variance) | Very low, still falling | Much higher, flat or rising | More data, augmentation, weight decay, dropout, early stopping, smaller model, pretrained weights |
Example: training loss 0.05 with validation loss 0.30 signals overfitting; 0.10 versus 0.12 is a good fit. There are no universal numeric thresholds; the gap and the trend matter. The expected test error decomposes as bias2 + variance + irreducible noise. Deep learning adds a twist called double descent: past the point where a model can exactly fit the training set, test error can go down again as models grow, which is why very large, well-regularized models generalize well.
loss |\ validation loss | \ ___....----'' | \ _.-'' | \ _.-' <- gap grows: overfitting | '.___...--'' best checkpoint (early stopping here) | '--..__ ^ | ''--..__|___ training loss +------------------------------------------------ epochs
The regularization toolbox
L2 / weight decay
Add λΣw2 to the loss (or shrink weights each step). Keeps weights small and the function smooth. Equivalent to a Gaussian prior on weights (MAP estimation). Use AdamW for adaptive optimizers.
L1
Add λΣ|w|. Drives many weights exactly to zero (sparsity, feature selection); corresponds to a Laplace prior. Rare in deep nets except for pruning.
Dropout
During training, zero each activation with probability p (0.1-0.5) and scale survivors by 1/(1 − p). Prevents co-adaptation of neurons; acts like averaging an ensemble of thinned networks. Disabled at inference.
Early stopping
Monitor validation loss, keep the best checkpoint, stop after "patience" epochs without improvement. Cheap and almost always worth it.
Data augmentation
Create label-preserving variants: flips, crops, rotations, color jitter, noise, mixup/cutmix for images; time-stretch and SpecAugment for audio; back-translation and token masking for text. Encodes invariances you know are true.
Label smoothing
Mix the one-hot with uniform: y' = (1 − ε)y + ε/K (PyTorch / Szegedy). For ε = 0.1 and K = 10 the true class is 0.91, not 0.9. Stops logits growing without bound and improves calibration.
More data and transfer learning
The strongest regularizer is more (diverse) data. The next best is starting from pretrained weights that already encode general features.
Architecture and noise
Smaller models, weight sharing (convolutions), BatchNorm's batch noise, stochastic depth (randomly skipping residual blocks), and small batch sizes all regularize implicitly.
How dropout works, precisely
Why it helps: no neuron can rely on a specific partner being present, so each learns features that are useful on their own; the network behaves like an ensemble of exponentially many sub-networks sharing weights. Where it is placed: after activations in fully connected layers, on embeddings and attention weights in Transformers (usually 0.1), and between (not inside) recurrent steps of RNNs. It is used less in modern CNNs, which rely on BatchNorm and augmentation. Monte Carlo dropout keeps dropout on at inference and averages several passes to estimate uncertainty.
import torch.nn as nn
model = nn.Sequential(
nn.Linear(128, 64), nn.ReLU(), nn.Dropout(p=0.5),
nn.Linear(64, 10)
)
model.train() # dropout active, BatchNorm uses batch statistics
model.eval() # dropout off, BatchNorm uses running statistics
# augmentation for images (torchvision)
from torchvision import transforms as T
train_tf = T.Compose([T.RandomResizedCrop(224), T.RandomHorizontalFlip(),
T.ColorJitter(0.2, 0.2, 0.2), T.ToTensor(),
T.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])])
val_tf = T.Compose([T.Resize(256), T.CenterCrop(224), T.ToTensor(),
T.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])])
Train, validation and test splits
Training data updates the weights. Validation data chooses hyperparameters, checkpoints and early stopping, and never updates weights. The test set is touched once at the end to report an unbiased estimate. Tuning on the test set leaks information and inflates results. For time series, split by time (never shuffle future into past); for users or patients, split by entity so the same person does not appear on both sides.
Normalization: BatchNorm, LayerNorm, RMSNorm
As weights change during training, the distribution of each layer's inputs keeps shifting and can drift toward huge or tiny values. Normalization layers re-standardize activations inside the network, which makes optimization much smoother, allows higher learning rates, and reduces sensitivity to initialization.
Picture a chain of rooms where each room's heater reacts to the temperature of air coming from the previous room. Without control, one room overheats, the next overreacts and freezes, and the whole building swings wildly. Normalization is a thermostat at every doorway that resets the air to a comfortable standard before it enters the next room, plus two dials (γ and β) that each room can adjust if it really prefers warmer or cooler air. BatchNorm sets the thermostat using the average of everyone passing through at the same moment (the batch); LayerNorm sets it using only the readings of one visitor (one example).
| Norm | Statistics computed over | Batch dependent? | Best for | Notes |
|---|---|---|---|---|
| BatchNorm | The batch (and spatial positions), separately for each channel | Yes | CNNs with batch size of at least ~16-32 | Keeps running mean/var for inference; breaks with tiny batches, variable-length sequences and autoregressive generation |
| LayerNorm | All features of one example (one token) | No | Transformers, RNNs | Same computation in training and inference; works with batch size 1 |
| RMSNorm | Features of one example, dividing by root-mean-square only (no mean subtraction, usually no β) | No | Modern LLMs (Llama family, T5) | Cheaper, similar quality |
| GroupNorm | Groups of channels within one example | No | Detection/segmentation with small batches | Between LayerNorm and InstanceNorm |
| InstanceNorm | Each channel of each example separately | No | Style transfer | Removes per-image contrast/style |
Tensor of activations: N (batch) x C (channels) x H*W (spatial)
BatchNorm LayerNorm InstanceNorm GroupNorm
normalize over normalize over normalize over normalize over
N, H, W per C C, H, W per N H, W per (N, C) H, W and a group
of channels per N
C -> C -> C -> C ->
N [#| | ] N [########] N [#| | ] N [###| ]
| [#| | ] | [ ] | [ | | ] | [ | ]
v [#| | ] v [ ] v [ | | ] v [ | ]
BatchNorm: training versus inference
running: μ ← (1 − m)μ + m μB, σ2 ← (1 − m)σ2 + m σunbiased2
eval: y = γ(x − μ)/√(σ2 + ε) + β m is the PyTorch
momentum (default 0.1). Normalization in the forward pass uses the biased batch variance (divide by n); the running variance is updated with the unbiased estimate (divide by n − 1). That mismatch is a common interview trap and can matter for tiny batches.During training, BatchNorm normalizes with the current mini-batch's mean and variance and updates those exponential running averages. During inference it uses the running averages, so a single example gets a deterministic output. This is why model.eval() matters: forgetting it at inference makes predictions depend on whatever else is in the batch; forgetting model.train() during training freezes the statistics. Fine-tuning with tiny batches often works better with BatchNorm layers frozen in eval mode.
Why normalization works
The original explanation was reducing "internal covariate shift" (changing input distributions to each layer). Later research found the main effect is a smoother loss landscape: gradients become more predictable, so larger learning rates are safe. Normalization also decouples the scale of weights from the function computed, and BatchNorm's batch noise acts as a mild regularizer. Rule of thumb: BatchNorm for CNNs, LayerNorm or RMSNorm for Transformers and RNNs, GroupNorm when batches are tiny.
Where to place it: pre-norm versus post-norm
In CNNs the classic order is Conv → BatchNorm → ReLU (the conv bias becomes redundant because BatchNorm subtracts the mean, so set bias=False). In Transformers, the original design applied LayerNorm after the residual addition (post-norm), which needs careful warmup. Modern models use pre-norm (x + Sublayer(LayerNorm(x))), which keeps a clean identity path through the residual stream and trains stably at great depth.
import torch.nn as nn
conv_block = nn.Sequential(
nn.Conv2d(64, 128, kernel_size=3, padding=1, bias=False), # bias redundant before BN
nn.BatchNorm2d(128),
nn.ReLU(inplace=True),
)
class PreNormBlock(nn.Module): # Transformer-style residual sublayer
def __init__(self, d, hidden):
super().__init__()
self.norm = nn.LayerNorm(d)
self.ff = nn.Sequential(nn.Linear(d, hidden), nn.GELU(), nn.Linear(hidden, d))
def forward(self, x): # x: (batch, seq, d)
return x + self.ff(self.norm(x))
Vanishing and exploding gradients, and residual connections
Backpropagation multiplies local derivatives layer by layer. In a network of depth N, the gradient reaching the first layer is roughly a product of N factors (weight matrices times activation derivatives). If those factors are typically below 1, the product shrinks exponentially (vanishing gradients): early layers stop learning. If they are above 1, it grows exponentially (exploding gradients): updates become huge, the loss spikes, and weights overflow to NaN.
Vanishing gradients are like a whisper passed down a long line of people: each repeats it a little quieter, and by the fiftieth person nothing is heard, so the people at the front never learn what went wrong. Exploding gradients are the opposite: each person shouts a bit louder than they heard, and soon it is deafening noise. A residual connection is an express intercom that runs from the back of the line straight to every person, so the message arrives clearly no matter how long the line is.
The arithmetic
The sigmoid's derivative is at most 0.25. Chaining 10 sigmoid layers multiplies the gradient by at most 0.2510 ≈ 9.5 × 10−7; 30 layers gives about 10−18. Conversely, a per-layer factor of 1.5 gives 1.520 ≈ 3,300 and 1.550 ≈ 6 × 108. This is a big part of why neural networks were stuck at a few layers for two decades. RNNs suffer worst because the same recurrent weight matrix is multiplied at every time step: over 400 time steps, even a factor of 0.99 per step leaves 0.018 of the signal.
The fixes that made deep learning work
| Fix | Addresses | How |
|---|---|---|
| ReLU-family activations | Vanishing | Derivative is exactly 1 for active units instead of at most 0.25 |
| Xavier / He initialization | Both | Keeps activation and gradient variance near 1 at the start |
| Normalization layers | Both | Re-scale activations at every layer, blocking compounding |
| Residual (skip) connections | Vanishing, degradation | Identity path lets gradients flow directly to early layers |
| Gradient clipping | Exploding | Rescale the gradient vector if its norm exceeds a threshold (commonly 1.0) |
| LSTM / GRU gates | Vanishing in RNNs | Additive cell-state updates controlled by a forget gate near 1 |
| Lower learning rate, warmup | Exploding | Smaller steps while statistics are unreliable |
| Auxiliary losses / deep supervision | Vanishing | Extra losses at intermediate layers (early Inception networks) |
Residual connections
Residual networks were motivated by the degradation problem: making a plain network deeper (say 56 layers instead of 20) made even the training error worse, which is not overfitting but an optimization failure. A deeper network should be at least as good, because extra layers could learn the identity, but learning an identity through a stack of non-linear layers is hard. With a skip connection, the block only needs to learn the residual F(x) = H(x) − x, and "do nothing" is simply F = 0. This allowed networks of 152 and even 1,000 layers. Residual connections are now everywhere: every Transformer block wraps attention and the feed-forward network in them. If the shapes differ (channels or spatial size change), the shortcut uses a 1×1 convolution with stride to match.
x ------------------------------+
| | identity shortcut
[conv 3x3, BN, ReLU] | (1x1 conv if shapes differ)
| |
[conv 3x3, BN] |
| |
+------------ (+) <-------------+
|
ReLU
|
y = ReLU(F(x) + x)
Gradient clipping
loss.backward()
total_norm = torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0) # returns the pre-clip norm
optimizer.step()
# log total_norm: persistent spikes point to a too-high LR or bad data
Norm clipping rescales the entire gradient vector when its L2 norm exceeds the threshold, preserving direction. Value clipping (clamping each element) distorts direction and is rarely preferred. Clipping is effectively mandatory for RNNs and standard for Transformers.
Convolutional neural networks (CNNs)
CNNs are designed for grid-structured data, where every value has a fixed position and well-defined neighbours: images (height × width × color channels), audio spectrograms, video frames, even some time series. An MLP on a 224×224 RGB image would need about 150,000 weights per hidden neuron and would treat a cat in the top-left corner as unrelated to the same cat in the bottom-right. CNNs build in two assumptions that fix both problems: locality (nearby pixels are related, so each unit looks only at a small patch) and translation equivariance (a pattern means the same thing wherever it appears, so the same filter is reused at every position).
Imagine searching a large dark photograph with a small flashlight that has a stencil in front of it, say a vertical bar. You slide the flashlight across every position and note how brightly the stencil pattern matches the patch underneath; the result is a map of where vertical edges are. A CNN layer carries dozens of flashlights with different stencils (filters), and it learns the stencil shapes itself. The next layer shines its flashlights on those maps, so it finds combinations of edges (corners, textures), and deeper layers find eyes, wheels and faces. Using the same stencil everywhere is parameter sharing; the small beam is the local receptive field.
The convolution operation
Worked example. A 5×5 grayscale image with a bright left region (value 10) and dark right region (0), and a 3×3 vertical-edge kernel:
Input (5x5) Kernel (3x3) Output (3x3), stride 1, no padding 10 10 10 0 0 1 0 -1 0 30 30 10 10 10 0 0 1 0 -1 0 30 30 10 10 10 0 0 * 1 0 -1 = 0 30 30 10 10 10 0 0 10 10 10 0 0 Top-left output: each row gives 10*1 + 10*0 + 10*(-1) = 0, so 0 in total (flat region). Next position: each row gives 10*1 + 10*0 + 0*(-1) = 10, three rows = 30 (edge found).
The output (a feature map) is large exactly where the edge is. Nobody designs this kernel in a CNN: all kernel values start random and are learned by backpropagation, just like MLP weights. Filters in early layers usually end up looking like edge and color detectors; the network discovers them because they reduce the loss.
Padding, stride and output size
- Padding adds a border (usually zeros) so the kernel can be centered on edge pixels. Without it, every convolution shrinks the map and border pixels are under-used. "Same" padding (P = (K − 1)/2 for odd K at stride 1) keeps the size; "valid" means no padding. Padding changes the output size, never the kernel size.
- Stride is how far the kernel jumps. Stride 2 halves the resolution (downsampling), captures coarser features and cuts computation, at the cost of fine detail.
- Dilation spaces out kernel taps to enlarge the receptive field without extra parameters (used in segmentation and audio models).
Channels and multiple filters
A color image has 3 input channels. Each filter spans all input channels (shape Cin × k × k) and produces one output channel by summing over them. A layer with Cout filters produces Cout feature maps stacked along depth, like RGB but with 32, 64 or 512 channels, each detecting a different pattern. In PyTorch, nn.Conv2d(16, 32, kernel_size=3, padding=1) means 16 input channels and 32 output channels (filters), not image sizes. Tensors are laid out as [batch, channels, height, width], for example [1, 3, 224, 224].
Pooling
Pooling summarizes small regions to shrink the feature maps. Max pooling keeps the strongest response in each window ("is this feature present somewhere here?"); average pooling keeps a smoother summary. A 2×2 window with stride 2 halves height and width. Example: the 4×4 map [[1,3,2,4],[5,6,1,2],[7,2,8,3],[4,5,9,1]] becomes [[6,4],[7,9]]. Pooling has no learnable parameters, reduces computation, and gives some tolerance to small shifts. The trade-off: it discards exact positions and fine detail, which is why segmentation and medical-imaging models use skip connections (U-Net) or fewer pooling steps, and many modern networks replace pooling with strided convolutions that learn their own downsampling. Global average pooling averages each channel over the whole map to one number, replacing large flatten-plus-dense heads and making the network work with any input size.
Receptive field
The receptive field of a unit is the region of the input that can influence it. Stacking n 3×3 convolutions at stride 1 gives a (2n + 1)×(2n + 1) receptive field: two layers see 5×5, three see 7×7. Two stacked 3×3 layers use 18 weights per channel pair versus 25 for one 5×5, and add an extra non-linearity; this is why small kernels dominate. Strides and pooling grow the receptive field multiplicatively, so deep layers see most of the image, which is how "semantic" layers can respond to whole objects.
The typical CNN pipeline
image [3x32x32] | Conv 3x3 (32 filters) + BN + ReLU -> [32x32x32] edges, colors | Conv 3x3 (32) + BN + ReLU, MaxPool 2 -> [32x16x16] | Conv 3x3 (64) + BN + ReLU, MaxPool 2 -> [64x8x8] textures, parts | Conv 3x3 (128) + BN + ReLU, MaxPool 2 -> [128x4x4] object parts | Global average pool -> [128] image embedding | Linear 128 -> 10 -> [10] logits classes v (softmax inside the loss during training) Spatial size shrinks, channel depth grows: less "where", more "what".
Flattening turns the final feature maps into a vector without losing information; each neuron of the following dense layer connects to every element. That vector (or the global-pooled one) is effectively an embedding of the image, task-specific unless the network was trained to produce general features. All parameters, convolution filters and dense weights alike, are trained together by one backward pass.
import torch
import torch.nn as nn
class SmallCNN(nn.Module):
def __init__(self, n_classes=10):
super().__init__()
def block(cin, cout, pool=True):
layers = [nn.Conv2d(cin, cout, 3, padding=1, bias=False),
nn.BatchNorm2d(cout), nn.ReLU(inplace=True)]
if pool:
layers.append(nn.MaxPool2d(2))
return layers
self.features = nn.Sequential(*block(3, 32, pool=False), *block(32, 32),
*block(32, 64), *block(64, 128))
self.head = nn.Sequential(nn.AdaptiveAvgPool2d(1), nn.Flatten(),
nn.Dropout(0.2), nn.Linear(128, n_classes))
def forward(self, x): # x: [B, 3, 32, 32]
return self.head(self.features(x)) # [B, n_classes] raw logits
x = torch.randn(8, 3, 32, 32)
print(SmallCNN()(x).shape) # torch.Size([8, 10])
Invariance, equivariance and beyond images
Convolution is translation equivariant (shift the input, the feature map shifts the same way). Pooling and global averaging add approximate translation invariance (the final prediction barely changes). It is not perfect: strides and pooling cause aliasing, so small shifts can change outputs, and CNNs are not naturally invariant to rotation or scale, which is why augmentation is used. CNNs also work on 1D data (audio, sensor signals, text via 1D convolutions), 3D data (video as frames × height × width, volumetric medical scans), and video with CNN-plus-RNN or CNN-plus-Transformer hybrids that combine spatial features per frame with temporal modelling across frames.
nn.AdaptiveAvgPool2d(1) before the head, or nn.LazyLinear, and print tensor shapes through the network once.Classic CNN architectures and transfer learning
The history of CNN architectures is a history of ideas that are still used everywhere: deeper stacks of small kernels, multi-branch blocks, 1×1 bottlenecks, residual shortcuts, and efficient depthwise convolutions. Knowing why each design appeared is more valuable in interviews than memorizing layer counts.
The architecture lineage is like the evolution of the car. The first models proved the concept (LeNet), then a much bigger engine won the race (AlexNet), engineers standardized parts (VGG's uniform 3×3 blocks), added multiple gears working in parallel (Inception), invented a transmission that lets power flow smoothly regardless of speed (ResNet's shortcuts), and finally optimized for fuel efficiency (MobileNet, EfficientNet). Transfer learning is buying a car with a proven engine and only replacing the body for your purpose instead of building an engine from scratch.
| Model (year) | Depth / size | Key idea | Legacy |
|---|---|---|---|
| LeNet-5 (1998) | ~7 layers, ~60K params | Conv → pool → conv → pool → dense for handwritten digits | The template for all CNNs |
| AlexNet (2012) | 8 layers, ~60M params | ReLU, dropout, data augmentation, training on two GPUs | Kicked off the deep learning era (top-5 ImageNet error ~16% vs ~26%) |
| VGG-16/19 (2014) | 16-19 layers, ~138M params | Only 3×3 convolutions, stacked; very uniform | Showed depth matters; popular feature extractor; heavy |
| GoogLeNet / Inception (2014) | 22 layers, ~7M params | Inception module: parallel 1×1, 3×3, 5×5 and pooling branches; 1×1 bottlenecks; global average pooling | Multi-scale features; efficiency through 1×1 convs |
| ResNet (2015) | 18-152 layers (ResNet-50: ~25M) | Residual shortcuts; bottleneck blocks (1×1 → 3×3 → 1×1); BatchNorm everywhere | Made very deep nets trainable; still the default backbone |
| DenseNet (2016) | 121-264 layers | Each layer receives feature maps of all previous layers (concatenation) | Feature reuse, parameter efficient |
| MobileNet v1-v3 (2017-2019) | ~2-5M params | Depthwise separable convolutions; inverted residuals with linear bottlenecks; hardware-aware search | Standard for phones and edge devices |
| EfficientNet (2019) | B0 (~5M) to B7 (~66M) | Compound scaling: grow depth, width and resolution together with fixed ratios; SiLU activations | Best accuracy per FLOP of its time |
| Vision Transformer (2020), ConvNeXt (2022) | Tens to hundreds of millions of params | ViT: split image into patches, treat them as tokens for a Transformer. ConvNeXt: a CNN modernized with Transformer-era design choices | Transformers and CNNs now compete; hybrids are common |
Key building blocks to explain
- 1×1 convolution: a per-pixel dense layer across channels. It mixes channel information and changes channel count cheaply, used as a bottleneck to reduce channels before an expensive 3×3 convolution.
- Depthwise separable convolution: split a standard convolution into a depthwise step (one k×k filter per input channel) and a pointwise 1×1 step. Remaining cost is about (1/Cout + 1/k2) of a dense convolution, roughly 8–9× cheaper for 3×3 kernels with wide output.
- Bottleneck block: 1×1 reduce → 3×3 → 1×1 expand, wrapped in a residual connection.
- Squeeze-and-excitation: global-pool each channel, pass through a tiny MLP, and rescale channels by the result: channel-wise attention.
- Compound scaling: scaling only depth or only width gives diminishing returns; scaling depth, width and resolution together is more efficient.
Beyond classification
The same backbones power other vision tasks: object detection (predict boxes and classes; two-stage R-CNN family versus single-stage YOLO/SSD/RetinaNet, which introduced focal loss), semantic segmentation (a class per pixel; encoder-decoder U-Net with skip connections, dilated convolutions), instance segmentation (Mask R-CNN), and pose estimation. For deployment on phones and embedded boards see Edge AI.
Transfer learning and fine-tuning
Features learned on a large dataset such as ImageNet transfer surprisingly well: early layers learn generic edges and textures, later layers more task-specific parts. Instead of training from scratch, start from pretrained weights and adapt them.
| Your situation | Strategy |
|---|---|
| Small dataset, similar to the pretraining data | Freeze the backbone, train only a new classification head (feature extraction) |
| Small dataset, different domain (medical, satellite) | Freeze early layers, fine-tune later layers and head with a small learning rate; heavy augmentation |
| Large dataset, similar | Fine-tune the whole network with a small learning rate |
| Large dataset, very different | Fine-tune everything, or train from scratch if compute allows; pretrained init still usually helps |
- Replace the head Swap the final layer for one with your number of classes.
- Match preprocessing Use the same input size and normalization statistics as pretraining.
- Train the head first With the backbone frozen, a few epochs at a normal learning rate so random head gradients do not wreck pretrained features.
- Unfreeze gradually Unfreeze deeper blocks with a 10-100x smaller learning rate; discriminative learning rates (smaller for earlier layers) help.
- Watch BatchNorm With small batches, keep BatchNorm layers in eval mode (frozen statistics).
import torch.nn as nn
from torchvision import models
model = models.resnet50(weights=models.ResNet50_Weights.IMAGENET1K_V2)
for p in model.parameters():
p.requires_grad = False # freeze backbone
model.fc = nn.Linear(model.fc.in_features, 5) # new head, trainable by default
# later: unfreeze the last stage with a smaller learning rate
for p in model.layer4.parameters():
p.requires_grad = True
optimizer = torch.optim.AdamW([
{"params": model.fc.parameters(), "lr": 1e-3},
{"params": model.layer4.parameters(), "lr": 1e-4},
], weight_decay=1e-4)
Recurrent networks: RNN, LSTM, GRU
Sequences such as text, speech, sensor readings and prices have an order, variable length, and dependencies that can span long distances ("The keys that I left on the kitchen table yesterday are missing"). An MLP takes a fixed-size input and has no notion of order; a recurrent neural network processes one element per time step while carrying a hidden state, a vector that summarizes everything seen so far.
A vanilla RNN reads a review word by word and keeps a single sticky note with its current understanding. After each word it rewrites the note, and because the note is small, older details get erased; after 200 words the opening is forgotten. An LSTM swaps the sticky note for a filing cabinet with three valves: one decides which old files to shred (forget gate), one decides which new notes to file (input gate), and one decides which files to pull out and show right now (output gate). Important words like "not" can be filed and kept for many steps. The sticky note is the hidden state; the filing cabinet is the LSTM cell state.
The vanilla RNN
Folded Unrolled through time (same weights at every step)
+------+ h0 -->[RNN]--h1-->[RNN]--h2-->[RNN]--h3-->[RNN]--h4
| v ^ | ^ | ^ |
x -->[ RNN ]--> y x1 y1 x2 y2 x3 y3 ...
^ |
+--h---+ An RNN has one (or a few) layers, run across many time
steps; it is deep in time, not necessarily in layers.
Input-output patterns
| Pattern | Example |
|---|---|
| Many-to-one | Sentiment of a review, classifying a sensor window, spam detection |
| One-to-many | Image captioning (image embedding in, word sequence out), music generation |
| Many-to-many, aligned | Part-of-speech tagging, named-entity recognition, frame labelling (one output per input) |
| Many-to-many, unaligned (seq2seq) | Translation, summarization, speech-to-text (lengths differ) |
For tagging, the output at step t conditions on inputs x1..xt (the current input included). For generation or forecasting, the model predicts step t from x1..xt−1, then feeds its own prediction back in.
Backpropagation through time (BPTT) and its problem
Training unrolls the network over the sequence, computes the loss (summed over time steps or at the end), and backpropagates through the unrolled graph. Weights are usually updated once per sequence or batch, not after every word. The number of backward steps equals the sequence length; truncated BPTT limits it to, say, 100 steps for long streams. Because the gradient flowing from step t back to step k contains the product of (t − k) Jacobians involving Whh and tanh', it vanishes when the largest singular value of Whh is below about 1 and explodes when it is above. In practice vanilla RNNs struggle to learn dependencies beyond 10-20 steps.
LSTM: long short-term memory
c̃t = tanh(Wc[ht−1, xt] + bc) ct = ft ⊙ ct−1 + it ⊙ c̃t ht = ot ⊙ tanh(ct) [h, x] is concatenation (joining vectors end to end, not multiplication); ⊙ is element-wise (Hadamard) multiplication. Gates use sigmoid so each entry is a 0-to-1 "valve"; the candidate c̃ uses tanh to propose values in (−1, 1). Biases are ordinary learnable parameters and can take any real value.
| Component | Role | Example ("I lived in Delhi, now I live in Bangalore") |
|---|---|---|
| Forget gate f | How much of each old cell-state entry to keep | On reaching "now", reduce the stored "Delhi" information |
| Input gate i | How much of the candidate to write | Write "Bangalore" strongly |
| Candidate c̃ | Proposed new information from the current input and previous hidden state | An encoding of "Bangalore" |
| Cell state c | Long-term memory, updated additively | Now holds "current city = Bangalore" |
| Output gate o | How much of the memory to expose as the hidden state (used for predictions and the next step) | Expose the city when a later word needs it |
All gates are computed in parallel from the same inputs at each step; there is no "input first, forget second" ordering. The forget gate is not redundant with the input gate: the input gate only filters what enters now and has no control over what is already stored, so without a forget gate the memory would accumulate stale information forever. If the forget gate wrongly deletes something needed later, the loss increases and backpropagation adjusts its weights.
Why LSTMs help with vanishing gradients: the cell-state update is additive. The gradient of ct with respect to ct−1 is just ft (element-wise), not a product with a weight matrix and a squashing derivative. When the forget gate stays near 1, gradients flow back many steps almost unchanged, the same idea as a residual connection but through time. It reduces, but does not eliminate, vanishing gradients.
Parameter count: an LSTM layer has 4 gate-sized weight blocks: 4 × (h × (h + x) + h). With input size 100 and hidden size 128: 4 × (128 × 228 + 128) = 117,248 parameters. (PyTorch keeps two bias vectors per gate, adding another 4h.)
GRU: gated recurrent unit
LSTM
- Separate cell state and hidden state, three gates
- Slightly more expressive; can be better on very long dependencies
- About 33% more parameters and compute than a GRU
GRU
- Two gates, single state
- Faster, fewer parameters, good on smaller datasets
- Often matches LSTM in practice; try it first when efficiency matters
Bidirectional and stacked RNNs
A standard LSTM reads left to right, so the representation of a word knows only the past. A bidirectional RNN runs a second RNN right to left and concatenates both hidden states at each position, so each word sees context from both sides ("the film was not as bad as critics claimed"). Each direction still needs its own memory and gates. Bidirectional models suit tasks where the whole input is available (classification, tagging, the encoder of a translator) but not real-time generation, which cannot see the future. Stacked RNNs feed one layer's hidden-state sequence into another for more capacity; 2-4 layers with dropout between them is typical.
import torch
import torch.nn as nn
from torch.nn.utils.rnn import pack_padded_sequence
class SentimentLSTM(nn.Module):
def __init__(self, vocab_size, emb_dim=128, hidden=128, n_classes=2, pad_idx=0):
super().__init__()
self.emb = nn.Embedding(vocab_size, emb_dim, padding_idx=pad_idx)
self.lstm = nn.LSTM(emb_dim, hidden, num_layers=2, batch_first=True,
bidirectional=True, dropout=0.3)
self.head = nn.Linear(2 * hidden, n_classes)
def forward(self, tokens, lengths): # tokens: [B, T] int ids
x = self.emb(tokens) # [B, T, emb_dim]
packed = pack_padded_sequence(x, lengths.cpu(), batch_first=True,
enforce_sorted=False) # skip padding
_, (h_n, _) = self.lstm(packed) # h_n: [layers*2, B, hidden]
h = torch.cat([h_n[-2], h_n[-1]], dim=1) # last layer, both directions
return self.head(h) # [B, n_classes]
# in the training loop, always clip for RNNs:
# torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
Case study: sentiment on long movie reviews
A useful real-world lesson comes from training several models on 50,000 movie reviews (balanced, about 230 words each, truncated to 400 tokens):
| Model | Test F1 | What happened |
|---|---|---|
| Linear SVM / logistic regression on bag-of-words | 0.875 | Strong regularized linear baselines; hard to beat on this task |
| Vanilla RNN, random embeddings | 0.568 | Loss stuck around 0.70 (chance level, since ln 2 = 0.693) for 10 epochs: vanishing gradients over 400 steps, no clipping |
| LSTM, random embeddings | 0.528 | Training loss fell (0.69 to 0.57) but predictions were skewed toward one class; lower training loss did not mean better F1 |
| BiLSTM + pretrained 300-d word vectors | 0.811 | Pretrained embeddings and bidirectionality added about 24 F1 points, but training loss collapsed to 0.05: overfitting |
| Attention-based classifier | 0.863 | Every token attends to every other: no long sequential chain, no vanishing gradient |
Takeaways: a strong classical baseline is essential; gradient clipping and shorter sequences are non-optional for RNNs on long text; pretrained embeddings matter enormously; and low training loss is necessary but not sufficient.
Where RNNs still make sense
Transformers dominate NLP because they parallelize across the sequence and capture long-range dependencies directly. RNNs remain useful for streaming and low-latency tasks (live speech recognition, keyword spotting, sensor processing on microcontrollers), where inputs arrive one step at a time and the model needs constant memory per step rather than a growing attention cache. Modern state-space models and linear-attention RNN variants revive this idea at scale. RNNs are architectures, not learning paradigms: they can be used supervised, self-supervised or unsupervised. Non-sequential data (customer age and salary) should not be forced into an RNN; sequences with missing values need imputation or masking first.
pack_padded_sequence, or gather the hidden state at each sequence's true last index, and set padding_idx in the embedding.Sequence-to-sequence models and the attention idea
Translation and summarization map a sequence to another sequence of a different length. The encoder-decoder (seq2seq) design solves this with two RNNs: the encoder reads the input and compresses it into a vector; the decoder generates the output one token at a time, conditioned on that vector and on the tokens generated so far.
Plain seq2seq is like an interpreter who must listen to an entire speech, memorize it as a single mental summary, and only then translate without taking notes. For a short sentence that works; for a long speech details are lost. Attention gives the interpreter a full transcript and a highlighter: while producing each translated word, they glance back at the transcript and highlight the few source words most relevant right now. The single mental summary is the fixed context vector; the transcript is the set of all encoder hidden states; the highlighter strengths are the attention weights.
Encoder (reads source) Decoder (writes target)
"the" "cat" "sat" <s> "le" "chat"
| | | | | |
[enc]->[enc]->[enc]--- context vector c --->[dec]-->[dec]-->[dec]--> ...
h1 h2 h3 | | |
"le" "chat" "s'est"
Bottleneck: the whole source squeezed into one fixed-size vector c.
With attention: at every decoder step t
score(s_t, h_i) for every source position i -> softmax -> weights a_ti
context c_t = sum_i a_ti * h_i (a different context for every output word)
Training and decoding
- Teacher forcing: during training the decoder receives the true previous token rather than its own prediction. It is fast and stable but creates exposure bias: at inference the model sees its own (possibly wrong) outputs, a distribution it never trained on. Scheduled sampling mixes the two.
- Decoding: greedy (pick the most likely token), beam search (keep the k best partial sequences), or sampling with temperature, top-k or top-p. These strategies carry over to modern language models; see Transformers & LLMs.
- Rare words: a fixed vocabulary maps unseen names and numbers to an unknown token, so the model cannot reproduce them. Pointer-generator networks add a copy mechanism that can copy source tokens directly; subword tokenization (BPE, WordPiece) largely solves the problem in modern systems.
Attention, in general form
Attention solved the bottleneck (the decoder can access every encoder state), gave direct short gradient paths between any output and any input position, and produced interpretable alignments (which source word each target word used). The next conceptual leap was to ask: if attention is this good, do we need recurrence at all? Self-attention lets every position of one sequence attend to every other position of the same sequence, in parallel, with no recurrence. Stacking self-attention with feed-forward layers, residual connections, LayerNorm and positional information gives the Transformer, the basis of BERT, GPT and every modern large language model.
| Property | RNN | CNN (1D) | Self-attention |
|---|---|---|---|
| Operations along the sequence | Sequential, O(n) steps | Parallel | Parallel |
| Path length between distant tokens | O(n) | O(log n) or O(n/k) | O(1) |
| Compute per layer | O(n · d2) | O(k · n · d2) | O(n2 · d) |
| Memory at inference per new token | Constant state | Window | Grows with context (KV cache) |
Embeddings: learned dense representations
Neural networks need numbers. Categorical things (words, product IDs, users, tokens) have no natural numeric form. An embedding maps each item to a dense vector of, say, 64-4,096 real numbers, learned so that items that behave similarly end up close together in that vector space. Embeddings are how deep learning represents words, tokens, images, users and documents, and they power semantic search, recommendations and retrieval-augmented generation.
One-hot encoding is like giving every person in a city a unique locker number: it identifies them perfectly but says nothing about who they are, and locker 1,001 is no more similar to locker 1,002 than to locker 50,000. An embedding is like placing everyone on a map according to their interests, so neighbours share hobbies and you can say "go from 'chef' in the direction of 'baking' and you reach 'pastry chef'". The locker number is the one-hot vector; the map coordinates are the embedding; the directions on the map are meaningful relationships learned from data.
From one-hot to dense vectors
One-hot / bag-of-words (sparse)
- Vector length = vocabulary size (e.g. 50,000), a single 1
- Every pair of words is orthogonal: cosine("cat", "kitten") = 0
- Huge, mostly zero inputs; no notion of similarity
- Bag-of-words loses order: "dog bites man" = "man bites dog"
- Fine when the number of categories is small
Embeddings (dense)
- Short vectors (e.g. 300), every value used
- Geometry encodes meaning: cosine("cat", "kitten") high
- Compact, fast, similarity search is natural
- Learned from data, either for a task or with self-supervision
- Standard for large vocabularies and IDs
An embedding layer is mathematically a one-hot vector multiplied by a weight matrix E of shape (vocabulary size × dimension), which simply selects one row, so frameworks implement it as a table lookup. The rows start random and are trained by backpropagation like any other weights. The dimension (300 was a popular choice for word vectors, a few hundred to a few thousand for modern models) is a hyperparameter: too small cannot capture enough relationships, too large wastes compute and can overfit.
Distributional semantics and word2vec
"You shall know a word by the company it keeps": words appearing in similar contexts have similar meanings. Word2vec turns this into a self-supervised prediction task on raw text, with no labels needed.
- CBOW (continuous bag of words): predict the centre word from the surrounding context words within a window (order ignored). Faster; good for frequent words.
- Skip-gram: predict each context word from the centre word. Slower; better for rare words and fine relationships. You choose one of the two, not both.
- Every word takes a turn as the centre word as the window slides over the corpus.
- Negative sampling: instead of a softmax over the whole vocabulary, train a binary classifier to distinguish real (centre, context) pairs from random "negative" pairs; much cheaper.
- Window size: small windows (2-5) capture syntactic similarity; larger (5-10+) capture topical similarity. Tune it as a hyperparameter.
The trained input weight matrix is the embedding table. The famous result is that directions carry meaning: vec("king") − vec("man") + vec("woman") ≈ vec("queen"). GloVe learns similar vectors from global co-occurrence counts; fastText adds character n-grams so it can build vectors for misspelled or unseen words.
Static versus contextual embeddings
Word2vec gives one vector per word regardless of context, so "bank" in "river bank" and "bank account" share a vector: polysemy is not handled, and words not seen in training have no vector at all (out-of-vocabulary, which is a representation limit, not overfitting). Contextual embeddings compute a different vector for each occurrence based on the whole sentence. ELMo did this with deep bidirectional LSTMs (slower, heavier); BERT and GPT-style Transformers do it with self-attention, combining token embeddings with positional information. Chat assistants use exactly these contextual, Transformer-based dense representations over subword tokens.
Embeddings beyond words
- Images: the pooled output of a CNN or vision Transformer; trained contrastively with text (CLIP-style) so images and captions share one space.
- Sentences and documents: sentence-embedding models trained with contrastive losses; the backbone of semantic search and RAG.
- Users and items: recommender systems learn an embedding per user and per item; the dot product predicts affinity (matrix factorization is the linear special case).
- Categorical features in tabular models: embedding tables for high-cardinality columns (zip code, product ID) instead of huge one-hot vectors.
import torch, torch.nn as nn, torch.nn.functional as F
emb = nn.Embedding(num_embeddings=10_000, embedding_dim=300, padding_idx=0)
ids = torch.tensor([[12, 845, 3, 0, 0]]) # a padded sentence of token ids
vectors = emb(ids) # [1, 5, 300]
# load pretrained vectors and optionally freeze them
pretrained = torch.randn(10_000, 300) # stand-in for word2vec/fastText
emb = nn.Embedding.from_pretrained(pretrained, freeze=True, padding_idx=0)
a, b = emb(torch.tensor(12)), emb(torch.tensor(845))
print(F.cosine_similarity(a, b, dim=0))
Generative models: autoencoders, VAEs, GANs and diffusion
Discriminative models learn p(label | input): "is this a cat?". Generative models learn the data distribution p(x) itself, so they can produce new samples: images, audio, molecules, text. Four families matter for interviews, and the modern language models on the Transformers page are a fifth (autoregressive models that generate one token at a time).
An autoencoder is a student who summarizes a book into a one-page note and then rewrites the book from the note alone. A VAE is the same student, but the notes must follow a standard format so any randomly filled-in note still produces a sensible book. A GAN is an art forger and a detective locked in a contest: the forger improves until the detective cannot tell fakes from originals. A diffusion model is a restorer who learned to clean a painting covered in dust one light pass at a time, and who can then "restore" a canvas of pure dust into a brand-new painting.
Autoencoders
An autoencoder has an encoder that compresses the input x into a low-dimensional latent code z, and a decoder that reconstructs x̂ from z. It is trained to minimize reconstruction loss (MSE or BCE) and needs no labels. The narrow bottleneck forces it to keep only the most important structure, like a non-linear PCA.
x (784 pixels) -> [Encoder: 784 -> 256 -> 32] -> z (32 numbers) -> [Decoder: 32 -> 256 -> 784] -> x_hat
loss = || x - x_hat ||^2
- Uses: dimensionality reduction, pretraining, anomaly detection (train on normal data; high reconstruction error flags anomalies), and compression.
- Denoising autoencoder: corrupt the input and reconstruct the clean version, which learns robust features (a precursor of masked-prediction pretraining).
- Sparse autoencoder: penalize latent activations to learn interpretable features.
- Limitation for generation: a plain autoencoder's latent space has holes; decoding a random z gives garbage.
Variational autoencoders (VAEs)
A VAE makes the latent space smooth and sampleable. The encoder outputs the parameters of a distribution, a mean μ(x) and standard deviation σ(x), and a latent z is sampled from it. The loss adds a term pulling each encoded distribution toward a standard normal prior.
The reparameterization trick: sampling is not differentiable, so write z = μ + σ ⊙ ε with ε ~ N(0, I) drawn outside the network. Now gradients flow through μ and σ. VAEs give smooth interpolation and a principled likelihood, but their samples tend to be blurry. Their encoders are used as the compressor in latent diffusion: an image of 512×512×3 becomes a 64×64×4 latent before diffusion runs, which is much cheaper.
class VAE(nn.Module):
def __init__(self, d_in=784, d_lat=16):
super().__init__()
self.enc = nn.Sequential(nn.Linear(d_in, 256), nn.ReLU())
self.mu, self.logvar = nn.Linear(256, d_lat), nn.Linear(256, d_lat)
self.dec = nn.Sequential(nn.Linear(d_lat, 256), nn.ReLU(), nn.Linear(256, d_in))
def forward(self, x):
h = self.enc(x)
mu, logvar = self.mu(h), self.logvar(h)
z = mu + torch.exp(0.5 * logvar) * torch.randn_like(mu) # reparameterization
return self.dec(z), mu, logvar
def vae_loss(x_logits, x, mu, logvar):
rec = F.binary_cross_entropy_with_logits(x_logits, x, reduction="sum")
kl = -0.5 * torch.sum(1 + logvar - mu.pow(2) - logvar.exp())
return (rec + kl) / x.size(0)
Generative adversarial networks (GANs)
A generator G maps random noise z to a fake sample; a discriminator D outputs the probability that a sample is real. They are trained in alternation as a two-player minimax game:
| Failure mode | Symptom | Mitigations |
|---|---|---|
| Mode collapse | Generator produces a few near-identical outputs | Minibatch discrimination, Wasserstein loss, unrolled updates, diversity terms |
| Discriminator too strong | Generator gradients vanish, no progress | Balance learning rates (TTUR), label smoothing, noise on inputs, spectral normalization |
| Non-convergence | Losses oscillate; no monotonic loss to watch | WGAN-GP (Earth-Mover distance with gradient penalty), careful architectures (DCGAN rules), evaluate with FID not loss |
GANs produce sharp samples and are fast at inference (one forward pass), which keeps them popular for super-resolution, face editing, style transfer and real-time applications. For general image generation, diffusion models have largely replaced them because they are more stable and diverse. Sample quality is measured with FID (distance between feature statistics of real and generated images; lower is better) and Inception Score.
Diffusion models (intuition)
Diffusion models learn to reverse a gradual noising process.
- Forward process (fixed, no learning) Add a little Gaussian noise to an image over T steps (for example 1,000) until it is pure noise. A closed form lets you jump to any step directly: xt = √ᾱt x0 + √(1 − ᾱt) ε.
- Training Pick a random image, a random step t and random noise ε; create xt; train a network (a U-Net or a diffusion Transformer) to predict the noise ε from (xt, t). The loss is plain MSE between true and predicted noise.
- Sampling Start from pure noise and repeatedly subtract the predicted noise (plus a little fresh noise), step by step, until a clean image emerges. Faster samplers (DDIM and successors) need 20-50 steps; distilled and consistency models need 1-4.
- Conditioning Feed a text embedding through cross-attention so the denoiser follows a prompt. Classifier-free guidance runs the model with and without the prompt and extrapolates toward the prompted prediction; a guidance scale of about 7 trades diversity for prompt faithfulness.
- Latent diffusion Run the whole process in a VAE's compressed latent space instead of pixel space, which cuts compute by an order of magnitude or more.
| Family | Sample quality | Diversity | Sampling speed | Training stability | Likelihood |
|---|---|---|---|---|---|
| VAE | Blurry | Good | Fast (one pass) | Stable | Lower bound |
| GAN | Sharp | Risk of mode collapse | Fast (one pass) | Fragile | None |
| Diffusion | Excellent | Excellent | Slow (many steps), improving | Stable (simple MSE) | Bound available |
| Autoregressive (Transformers) | Excellent for discrete data such as text | Good | Sequential, one token per step | Stable | Exact |
Training at scale: batches, GPUs, mixed precision and parallelism
Modern models are limited by memory and throughput as much as by ideas. This section covers what uses GPU memory, how to fit bigger models and batches, and how training is spread across many GPUs.
Training a big model is like running a large restaurant kitchen. A CPU is one master chef doing everything in sequence; a GPU is a brigade of thousands of line cooks each doing a simple task in parallel, which is perfect when the recipe is "chop ten thousand onions" (matrix multiplication). Mixed precision is using quick rough measurements for most steps and precise scales only where it matters. Gradient accumulation is cooking a banquet in several smaller rounds when the pot is too small. Data parallelism is opening identical kitchens that each cook part of the orders and then agree on the recipe changes; model parallelism is splitting one giant recipe across kitchens because no single kitchen has room for all the equipment.
Why GPUs
Neural network compute is dominated by dense matrix multiplications and convolutions, which decompose into millions of independent multiply-adds. CPUs have a few powerful cores optimized for sequential, branchy code; GPUs have thousands of simpler cores plus specialized tensor cores for low-precision matrix math and very high memory bandwidth. The practical rule: keep the GPU fed (fast data loading, large enough batches) and avoid moving tensors between CPU and GPU inside the loop.
What uses memory during training
| Item | Size | Notes |
|---|---|---|
| Weights | 4 bytes/param in FP32, 2 in FP16/BF16 | A 7B-parameter model is 14 GB in BF16 just for weights |
| Gradients | Same as weights | One per trainable parameter |
| Optimizer state | Adam: 8 bytes/param (m and v in FP32); SGD with momentum: 4 | Plus an FP32 master copy of weights (4 bytes) under mixed precision |
| Activations | Proportional to batch size × sequence length or image size × depth × width | Often the largest item; kept for the backward pass |
Rule of thumb for mixed-precision Adam: about 16 bytes per parameter before activations (2 weights + 2 gradients + 4 master weights + 8 Adam state). A 1B-parameter model needs about 16 GB just for these, which explains why fine-tuning large models needs sharding or parameter-efficient methods such as LoRA.
Batch size
- Larger batches give more accurate gradient estimates and better hardware utilization, but each example contributes less and very large batches can generalize worse unless the learning rate is scaled and warmed up.
- Smaller batches add gradient noise, which regularizes and can help escape sharp minima, but underuse the hardware.
- BatchNorm needs reasonably large per-GPU batches; LayerNorm does not care.
- Practical approach: pick the largest batch that fits (or a known-good value), scale the learning rate accordingly, and use gradient accumulation to reach a target effective batch size.
Mixed precision
| Format | Bits (exponent / mantissa) | Range and precision | Use |
|---|---|---|---|
| FP32 | 32 (8 / 23) | Wide range, about 7 decimal digits | Safe default; master weights, reductions |
| FP16 | 16 (5 / 10) | Narrow range (max about 65,504); small gradients underflow to 0 | Fast on tensor cores; needs loss scaling |
| BF16 | 16 (8 / 7) | Same range as FP32, less precision | Preferred for training on modern GPUs/TPUs; no loss scaling needed |
| FP8, INT8, INT4 | 8 or fewer | Very limited | Cutting-edge training (FP8), and inference quantization |
Automatic mixed precision runs matrix multiplications and convolutions in FP16/BF16 (roughly 2x faster, half the activation memory) while keeping numerically sensitive operations (softmax, normalization, loss, weight updates) in FP32. With FP16, loss scaling multiplies the loss by a large factor before backward so small gradients do not underflow, then unscales before the optimizer step, skipping steps whose gradients overflow to infinity.
scaler = torch.amp.GradScaler("cuda") # only needed for float16
for x, y in loader:
x, y = x.to(device, non_blocking=True), y.to(device, non_blocking=True)
optimizer.zero_grad(set_to_none=True)
with torch.autocast(device_type="cuda", dtype=torch.float16):
loss = loss_fn(model(x), y)
scaler.scale(loss).backward()
scaler.unscale_(optimizer) # so clipping sees true gradient values
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
scaler.step(optimizer) # skips the step if grads are inf/NaN
scaler.update()
# with bfloat16: use dtype=torch.bfloat16 and drop the scaler entirely
Gradient accumulation
If a batch of 256 does not fit, run 8 micro-batches of 32, call backward on each (gradients add up in .grad), and step the optimizer once. Divide each micro-batch loss by the number of accumulation steps so the gradient equals the average over the effective batch. The math matches a true batch of 256, except for BatchNorm statistics, which still see only 32 examples.
accum = 8
optimizer.zero_grad(set_to_none=True)
for i, (x, y) in enumerate(loader):
loss = loss_fn(model(x.to(device)), y.to(device)) / accum
loss.backward()
if (i + 1) % accum == 0:
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
optimizer.step()
scheduler.step()
optimizer.zero_grad(set_to_none=True)
Activation (gradient) checkpointing
Instead of storing every activation for the backward pass, store only some (for example at block boundaries) and recompute the rest during backward. This trades roughly 30% extra compute for a large activation-memory reduction, allowing longer sequences, bigger batches or deeper models (torch.utils.checkpoint).
Distributed training
| Strategy | What is split | Communication | When |
|---|---|---|---|
| Data parallel (DDP) | The batch: each GPU holds a full model copy and processes a different slice of data | All-reduce (average) gradients every step, overlapped with backward | Model fits on one GPU; the default way to scale |
| Sharded data parallel (ZeRO, FSDP) | Optimizer state, gradients and even weights are sharded across GPUs; gathered just in time | More traffic than DDP (all-gather + reduce-scatter) | Model plus optimizer state does not fit on one GPU |
| Tensor (intra-layer) parallel | Individual weight matrices split across GPUs | Heavy, needs fast links within a node | Very large layers (big LLMs) |
| Pipeline parallel | Consecutive groups of layers on different GPUs; micro-batches flow through like an assembly line | Activations passed between stages; idle "bubbles" at start and end | Very deep models across nodes |
| 3D / hybrid | Data + tensor + pipeline combined | Tuned to the cluster topology | Training frontier models on thousands of GPUs |
Data parallel (DDP) Pipeline parallel (4 stages)
GPU0: full model <- batch slice 0 GPU0: layers 1-8 ---> activations
GPU1: full model <- batch slice 1 GPU1: layers 9-16 --->
GPU2: full model <- batch slice 2 GPU2: layers 17-24 --->
GPU3: full model <- batch slice 3 GPU3: layers 25-32 ---> loss
\___ all-reduce gradients ___/ micro-batches keep all stages busy
every GPU applies the same update
With DDP, use a DistributedSampler so each process sees a different shard of the data, and remember the effective batch size is per-GPU batch × number of GPUs (× accumulation steps). Throughput is often limited by the data pipeline (decoding images, tokenizing text); use multiple DataLoader workers, pinned memory and pre-processed data formats.
Debugging training: loss curves and a systematic checklist
Neural networks fail silently. Code with a bug often still runs, and the loss still goes down a bit, just not as far as it should. Experienced practitioners follow a routine: start simple, verify each piece, and change one thing at a time.
Debugging a network is like diagnosing a car that "does not drive well". A good mechanic does not replace the engine first; they check the fuel (data), the ignition (can it start at all: overfit one batch), the dashboard gauges (loss curves, gradient norms, learning rate) and then road-test in stages. Overfitting a single batch is turning the key in the garage: if the engine cannot even idle, there is no point taking it on the highway.
Step 1: overfit a tiny batch
Take one batch of 2-32 examples, turn off regularization (dropout, weight decay, augmentation) and train on it repeatedly. A correct model and pipeline should drive the loss to nearly zero and reach 100% accuracy on that batch within a few hundred steps. If it cannot, the problem is a bug, not a lack of data or capacity: wrong labels paired with inputs, a loss applied to the wrong tensors, a softmax applied twice, a detached graph, gradients not flowing to some layers, or a learning rate far off.
xb, yb = next(iter(train_loader))
xb, yb = xb[:16].to(device), yb[:16].to(device)
model.train()
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
for step in range(300):
opt.zero_grad()
loss = loss_fn(model(xb), yb)
loss.backward()
opt.step()
if step % 50 == 0:
print(step, round(loss.item(), 4)) # should approach 0
Step 2: read the curves
| What you see | Likely cause | What to try |
|---|---|---|
| Loss flat from the start, near ln(K) | Learning rate far too low or too high, dead units, labels shuffled relative to inputs, frozen parameters, vanishing gradients | LR range test; check that gradients are non-zero per layer; verify data-label pairing visually |
| Loss decreases then explodes or becomes NaN | Learning rate too high, no warmup, exploding gradients, log(0) or division by zero, FP16 overflow, bad input values | Lower LR, add warmup, clip gradients, use stable loss functions, check inputs for NaN/inf, use BF16 or a GradScaler |
| Loss oscillates wildly | Learning rate too high, batch too small, bad data shuffling (sorted by class) | Lower LR, bigger batch, shuffle properly |
| Training loss falls, validation loss rises | Overfitting (or train/validation distribution mismatch) | Early stopping, augmentation, weight decay, dropout, more data, smaller model |
| Both losses high and flat | Underfitting, model too small, too much regularization, bug | Bigger model, train longer, reduce regularization, overfit-one-batch test |
| Validation loss lower than training loss | Dropout/augmentation active only in training, training loss averaged over an epoch while weights improved, easier validation set, or leakage | Usually fine; check for leakage and duplicated examples |
| Validation accuracy stuck at the majority-class rate | Class imbalance; model predicts one class | Class weights, focal loss, resampling, better metrics (F1, PR-AUC) |
| Loss plateaus, then drops suddenly after an LR decay | Normal behaviour with step schedules | Consider cosine or tune decay points |
| Periodic spikes every epoch | Last partial batch, a bad example, data order effects | drop_last, inspect the batches at the spikes |
Step 3: the full checklist
- Look at the data Visualize inputs after preprocessing and augmentation, with their labels. Check shapes, dtypes, value ranges, class balance, duplicates between splits.
- Establish baselines Majority class, a linear model, a published or pretrained model. Know what "good" looks like.
- Check the initial loss About ln(K) for K balanced classes; much larger means bad initialization or output scaling.
- Overfit one batch Proves the model, loss and optimizer are wired correctly.
- Monitor gradients Per-layer gradient norms; zero means no flow (detached tensors, dead units, frozen layers); huge means explosion.
- Monitor activations Fraction of dead ReLUs, saturated sigmoids, activation statistics per layer.
- Check modes
model.train()for training,model.eval()plustorch.no_grad()for evaluation. - Scale up gradually Add data, capacity and regularization one step at a time, and keep a log of experiments with fixed random seeds.
NaN hunting
- Check the data first:
torch.isfinite(x).all()on inputs and targets. - Use numerically stable losses (
CrossEntropyLoss,BCEWithLogitsLoss), add ε inside logs, square roots and divisions. - Lower the learning rate and add warmup; clip gradients; log the gradient norm to find the step where it spikes.
- In FP16, use a GradScaler or switch to BF16.
torch.autograd.set_detect_anomaly(True)pinpoints the operation that produced a NaN in backward (slow; debug only).
A complete PyTorch training loop
Every PyTorch training loop, from a tiny classifier to large language model fine-tuning, follows the same five-beat rhythm per batch: zero the gradients, forward, loss, backward, step. The rest is bookkeeping: data loading, device placement, evaluation, schedules, mixed precision, early stopping and checkpointing. The example below trains the CNN from earlier on CIFAR-10 (32×32 color images, 10 classes) and is a solid template for most supervised tasks.
A training loop is a practice routine for a musician. Each practice round they clear their head of the last attempt (zero_grad), play the piece (forward), listen to the recording and count mistakes (loss), figure out which fingers caused which mistakes (backward), and adjust their technique (optimizer step). Every so often they perform for a teacher without adjusting anything mid-performance (evaluation in eval mode, no gradients), keep a recording of their best performance so far (checkpointing), and stop practising a piece once the teacher's scores stop improving (early stopping).
import math, time, torch, torch.nn as nn
from torch.utils.data import DataLoader, random_split
from torchvision import datasets, transforms as T
torch.manual_seed(42)
device = "cuda" if torch.cuda.is_available() else "cpu"
# ---------- data ----------
mean, std = (0.4914, 0.4822, 0.4465), (0.2470, 0.2435, 0.2616)
train_tf = T.Compose([T.RandomCrop(32, padding=4), T.RandomHorizontalFlip(),
T.ToTensor(), T.Normalize(mean, std)])
eval_tf = T.Compose([T.ToTensor(), T.Normalize(mean, std)])
full_train = datasets.CIFAR10("data", train=True, download=True, transform=train_tf)
train_set, val_idx = random_split(full_train, [45_000, 5_000],
generator=torch.Generator().manual_seed(0))
val_set = torch.utils.data.Subset(
datasets.CIFAR10("data", train=True, transform=eval_tf), val_idx.indices) # no augmentation
test_set = datasets.CIFAR10("data", train=False, download=True, transform=eval_tf)
train_loader = DataLoader(train_set, batch_size=128, shuffle=True, num_workers=4,
pin_memory=True, drop_last=True)
val_loader = DataLoader(val_set, batch_size=256, num_workers=4, pin_memory=True)
test_loader = DataLoader(test_set, batch_size=256, num_workers=4, pin_memory=True)
# ---------- model, loss, optimizer, schedule ----------
model = SmallCNN(n_classes=10).to(device) # defined in the CNN section
loss_fn = nn.CrossEntropyLoss(label_smoothing=0.1)
optimizer = torch.optim.AdamW(model.parameters(), lr=3e-3, weight_decay=5e-4)
epochs = 30
total_steps = epochs * len(train_loader)
warmup = len(train_loader) # one epoch of warmup
def lr_lambda(step):
if step < warmup:
return (step + 1) / warmup
return 0.5 * (1 + math.cos(math.pi * (step - warmup) / (total_steps - warmup)))
scheduler = torch.optim.lr_scheduler.LambdaLR(optimizer, lr_lambda)
use_amp = device == "cuda"
amp_dtype = torch.bfloat16 if use_amp and torch.cuda.is_bf16_supported() else torch.float16
scaler = torch.amp.GradScaler("cuda", enabled=use_amp and amp_dtype == torch.float16)
# ---------- one epoch of training ----------
def train_one_epoch():
model.train() # dropout on, BN uses batch stats
total, correct, running = 0, 0, 0.0
for x, y in train_loader:
x, y = x.to(device, non_blocking=True), y.to(device, non_blocking=True)
optimizer.zero_grad(set_to_none=True) # 1. clear old gradients
with torch.autocast(device_type="cuda", dtype=amp_dtype, enabled=use_amp):
logits = model(x) # 2. forward
loss = loss_fn(logits, y) # 3. loss
scaler.scale(loss).backward() # 4. backward
scaler.unscale_(optimizer)
grad_norm = nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
scaler.step(optimizer) # 5. update
scaler.update()
scheduler.step() # per-step schedule
running += loss.item() * y.size(0)
correct += (logits.argmax(1) == y).sum().item()
total += y.size(0)
return running / total, correct / total, grad_norm.item()
# ---------- evaluation ----------
@torch.no_grad() # no graph, less memory, faster
def evaluate(loader):
model.eval() # dropout off, BN running stats
total, correct, running = 0, 0, 0.0
for x, y in loader:
x, y = x.to(device), y.to(device)
logits = model(x)
running += nn.functional.cross_entropy(logits, y, reduction="sum").item()
correct += (logits.argmax(1) == y).sum().item()
total += y.size(0)
return running / total, correct / total
# ---------- main loop with early stopping and checkpointing ----------
best_val, patience, bad_epochs = float("inf"), 5, 0
for epoch in range(epochs):
t0 = time.time()
tr_loss, tr_acc, gnorm = train_one_epoch()
va_loss, va_acc = evaluate(val_loader)
print(f"epoch {epoch:02d} | train {tr_loss:.3f}/{tr_acc:.3f} | "
f"val {va_loss:.3f}/{va_acc:.3f} | lr {scheduler.get_last_lr()[0]:.2e} | "
f"grad {gnorm:.2f} | {time.time() - t0:.0f}s")
if va_loss < best_val - 1e-4:
best_val, bad_epochs = va_loss, 0
torch.save({"model": model.state_dict(), "optimizer": optimizer.state_dict(),
"scheduler": scheduler.state_dict(), "epoch": epoch}, "best.pt")
else:
bad_epochs += 1
if bad_epochs >= patience:
print("early stopping")
break
# ---------- final test, once ----------
model.load_state_dict(torch.load("best.pt", map_location=device)["model"])
test_loss, test_acc = evaluate(test_loader)
print(f"test loss {test_loss:.3f} test acc {test_acc:.3f}")
# ---------- inference on new data ----------
model.eval()
with torch.no_grad():
probs = torch.softmax(model(x_new.to(device)), dim=1) # probabilities only at inference
pred = probs.argmax(1)
What each piece is for
| Piece | Why it is there |
|---|---|
Dataset / DataLoader | Batching, shuffling (training only), parallel loading with workers, pinned memory for faster host-to-GPU copies |
| Separate transforms | Augment the training split only; validation and test use deterministic preprocessing |
model.train() / model.eval() | Switch dropout and BatchNorm behaviour |
torch.no_grad() | Skip building the autograd graph during evaluation: saves memory and time |
zero_grad(set_to_none=True) | Gradients accumulate by default; setting to None is slightly faster than zero-filling |
| Autocast + GradScaler | Mixed precision speed and memory savings with FP32 safety where needed |
| Gradient clipping | Protection against rare exploding steps; its norm is also a health metric |
| Warmup + cosine schedule | Stable start, smooth convergence |
| Checkpoint dict | Save model, optimizer and scheduler states so training can resume exactly |
| Early stopping on validation loss | Regularization and compute savings; the test set is used once at the very end |
Custom datasets and saving for deployment
from torch.utils.data import Dataset
class TabularDataset(Dataset):
def __init__(self, X, y):
self.X = torch.as_tensor(X, dtype=torch.float32)
self.y = torch.as_tensor(y, dtype=torch.long)
def __len__(self):
return len(self.y)
def __getitem__(self, i):
return self.X[i], self.y[i]
# export for serving or mobile runtimes
model.eval()
example = torch.randn(1, 3, 32, 32, device=device)
torch.onnx.export(model, example, "model.onnx", input_names=["image"],
output_names=["logits"], dynamic_axes={"image": {0: "batch"}})
Keras (TensorFlow) wraps the same loop in model.compile(optimizer, loss, metrics) and model.fit(x, y, validation_data=..., callbacks=[EarlyStopping(monitor="val_loss", patience=5, restore_best_weights=True)]). The concepts are identical; PyTorch just makes the loop explicit. For deploying the trained model on phones and embedded hardware (quantization, runtimes, NPUs) see Edge AI.
running += loss) instead of loss.item(). That keeps every batch's computation graph alive and leaks GPU memory until the process crashes. Use .item() or .detach() for anything you log.no_grad, put the model and data on the same device, output logits with CrossEntropyLoss, and explain how you would add mixed precision, gradient accumulation, clipping and checkpointing.Quick revision
- Deep learning is machine learning with multi-layer neural networks that learn their own features from raw data (representation learning).
- It took off because of large datasets, GPUs, and algorithmic fixes: ReLU, good initialization, normalization, residual connections, dropout and Adam.
- A neuron computes z = w·x + b and outputs f(z); weights live on connections, activations live in nodes.
- A single perceptron only separates linearly separable data; XOR needs a hidden layer.
- Without non-linear activations, any stack of layers collapses into one linear layer.
- The universal approximation theorem guarantees existence of a one-hidden-layer approximator, not that training finds it; depth is more parameter-efficient.
- Dense layer parameters = inputs × outputs + outputs.
- Sigmoid saturates and its derivative is at most 0.25; tanh is zero-centered; ReLU has derivative 1 for positive inputs but can die.
- GELU and SiLU are smooth ReLU-like activations used in Transformers and modern LLMs.
- Output activations follow the task: identity for regression, sigmoid for binary or multi-label, softmax for single-label multi-class.
- Loss is per example, cost is the average (plus regularization); libraries call both "loss".
- MSE penalizes large errors quadratically and is outlier-sensitive; MAE is robust; Huber mixes both.
- Cross-entropy is negative log-likelihood; confident wrong predictions are punished hardest (−ln 0.05 ≈ 3.0).
- Softmax plus cross-entropy has the gradient p − y with respect to logits.
- Feed raw logits to CrossEntropyLoss and BCEWithLogitsLoss; never softmax twice.
- Focal loss and class weights handle imbalance; label smoothing and temperature scaling improve calibration.
- Gradient descent: θ ← θ − η∇L; mini-batch SGD is the practical default.
- An epoch is one pass over the data; an iteration is one update on one mini-batch.
- Backpropagation is reverse-mode chain rule; it computes all gradients in about the cost of one or two extra forward passes but must store activations.
- Backprop computes gradients; the optimizer applies them; both happen only in training.
- Momentum smooths and accelerates; RMSProp adapts per-parameter step sizes; Adam combines both with bias correction.
- AdamW decouples weight decay from the adaptive update and is the default for Transformers.
- Learning rate is the most important hyperparameter; too high diverges, too low crawls.
- Warmup stabilizes early training; cosine decay is the common schedule; scale the learning rate with batch size.
- Zero or equal initialization leaves neurons symmetric; Xavier suits tanh/sigmoid, He (variance 2/fan_in) suits ReLU.
- Overfitting: low training loss, high validation loss; underfitting: both high.
- Regularizers: weight decay, dropout, early stopping, data augmentation, label smoothing, more data, transfer learning.
- Inverted dropout scales kept units by 1/(1 − p) in training and is disabled at inference.
- BatchNorm normalizes over the batch per channel and uses running statistics at inference; the forward pass uses biased batch variance, the running variance is updated unbiased; LayerNorm normalizes each example and suits Transformers; RMSNorm drops the mean.
- Label smoothing in PyTorch is y' = (1 − ε)y + ε/K, not (1 − ε) and ε/(K − 1).
- Vanishing or exploding gradients come from multiplying many factors below or above 1 through depth or time.
- Residual connections (y = F(x) + x) give gradients a direct path and solved the degradation problem. The residual stream is the vector every block reads and adds to.
- Pre-LN (x + F(Norm(x))) keeps a clean residual path and is the modern default; Post-LN (Norm(x + F(x))) needs warmup.
- SwiGLU is W2(SiLU(W1x) ⊙ W3x); hidden width ≈ 8d/3 keeps the parameter count equal to a 4d MLP.
- Clip gradients by norm (typically 1.0) for RNNs and Transformers.
- CNNs exploit locality and parameter sharing; conv output size = ⌊(W + 2P − K)/S⌋ + 1.
- Conv parameters = (Cin × k × k + 1) × Cout, independent of image size. Dilated output size uses effective kernel 1 + D(K − 1).
- Pooling downsamples and adds shift tolerance but discards fine positional detail.
- Stacked 3×3 convolutions grow the receptive field cheaply; 1×1 convolutions mix channels.
- Architecture lineage: LeNet, AlexNet, VGG, Inception, ResNet, DenseNet, MobileNet, EfficientNet, ViT.
- Transfer learning: replace the head, freeze then gradually unfreeze, use a much smaller learning rate.
- RNNs share weights across time and carry a hidden state; BPTT unrolls them and suffers vanishing gradients.
- LSTM gates (forget, input, output) control an additively updated cell state; GRU merges them into update and reset gates.
- Bidirectional RNNs see both past and future context and cannot be used for streaming generation.
- Seq2seq compresses the source into a context vector; attention removes this bottleneck and leads to Transformers.
- Embeddings are learned dense vectors; word2vec (CBOW, skip-gram) learns them from context; contextual models give per-occurrence vectors.
- Autoencoders compress and reconstruct; VAEs add a KL term and the reparameterization trick to make the latent space sampleable.
- GANs train a generator against a discriminator and risk mode collapse; diffusion models learn to predict added noise and denoise step by step.
- Mixed-precision Adam needs about 16 bytes per parameter before activations; BF16 avoids FP16's loss scaling.
- Gradient accumulation simulates big batches; activation checkpointing trades compute for memory.
- DDP replicates the model and all-reduces gradients; FSDP/ZeRO shard states; tensor and pipeline parallelism split the model.
- Debug by overfitting one small batch, checking the initial loss (about ln K), and watching loss, learning rate and gradient norms.
- The PyTorch rhythm per batch: zero_grad, forward, loss, backward, step; use train/eval modes and no_grad for evaluation.
Glossary
- Activation
- The output value of a neuron after its activation function; also used for the function itself.
- Adam
- Adaptive optimizer combining momentum with per-parameter scaling by the running root-mean-square of gradients, plus bias correction.
- AdamW
- Adam with weight decay applied directly to weights instead of through the gradient.
- Attention
- A mechanism that computes a weighted average of values, with weights from query-key similarity, letting a model focus on relevant positions.
- Autoencoder
- Network trained to reconstruct its input through a low-dimensional bottleneck.
- Autograd
- Automatic differentiation: the framework records operations and computes gradients via the chain rule.
- Backpropagation
- Algorithm that computes the gradient of the loss with respect to every parameter by applying the chain rule from output to input.
- Batch normalization
- Normalizes each channel using mini-batch statistics, then applies learned scale and shift; uses running statistics at inference.
- Batch size
- Number of examples processed before one parameter update.
- BF16
- 16-bit floating-point format with FP32's exponent range and reduced precision; preferred for mixed-precision training.
- Bias (parameter)
- Learnable offset added to a neuron's weighted sum, shifting its activation threshold.
- Bidirectional RNN
- Two RNNs reading a sequence in opposite directions, with states concatenated per position.
- BPTT
- Backpropagation through time: backprop on an RNN unrolled over the sequence.
- Calibration
- How well predicted probabilities match observed frequencies of being correct.
- CBOW
- Continuous bag of words: word2vec variant that predicts a centre word from its context words.
- Cell state
- The LSTM's long-term memory vector, updated additively under gate control.
- Convolution
- Sliding a small learnable kernel over an input and computing a dot product at each position.
- Cosine similarity
- Dot product of two vectors divided by the product of their lengths; measures directional similarity.
- Cross-entropy
- Loss equal to the negative log-probability assigned to the correct class; negative log-likelihood.
- Data augmentation
- Creating label-preserving variations of training examples to improve generalization.
- Data parallelism
- Each device holds a model copy and processes a different data slice; gradients are averaged across devices.
- Depthwise separable convolution
- A per-channel spatial convolution followed by a 1×1 convolution; far cheaper than a standard convolution.
- Diffusion model
- Generative model trained to reverse a gradual noising process by predicting the added noise.
- Dropout
- Regularizer that randomly zeroes activations during training and is disabled at inference.
- Dying ReLU
- A ReLU unit whose input is always negative, so it outputs zero and receives no gradient.
- Early stopping
- Halting training when validation performance stops improving, keeping the best checkpoint.
- Embedding
- A learned dense vector representing a discrete item such as a word, token, user or product.
- Encoder-decoder
- Architecture where an encoder builds a representation of the input and a decoder generates the output from it.
- Epoch
- One complete pass through the training dataset.
- Exploding gradient
- Gradients growing exponentially through depth or time, causing unstable updates or NaN.
- Feature map
- The 2D output of applying one convolutional filter across an input.
- Fine-tuning
- Continuing training of a pretrained model on a new task, usually with a small learning rate.
- Focal loss
- Cross-entropy scaled by (1 − pt)γ to focus training on hard, misclassified examples.
- FSDP / ZeRO
- Sharded data parallelism that splits parameters, gradients and optimizer state across devices.
- GAN
- Generative adversarial network: a generator and a discriminator trained against each other.
- GELU
- Smooth activation x·Φ(x), standard in Transformers.
- Gradient
- Vector of partial derivatives of the loss with respect to the parameters; points toward steepest increase.
- Gradient accumulation
- Summing gradients over several micro-batches before one optimizer step to simulate a larger batch.
- Gradient clipping
- Rescaling gradients whose norm exceeds a threshold to prevent explosions.
- GRU
- Gated recurrent unit: an RNN cell with update and reset gates and a single state.
- He initialization
- Random weights with variance 2/fan_in, designed for ReLU networks.
- Hidden state
- The vector an RNN carries between time steps, summarizing the sequence so far.
- Hyperparameter
- A setting chosen before training (learning rate, batch size, depth) rather than learned.
- Label smoothing
- Training toward softened targets y' = (1 − ε)y + ε/K instead of a hard one-hot, to reduce over-confidence.
- Layer normalization
- Normalizes across the features of each example independently of the batch.
- Learning rate
- Step size multiplying the gradient in each parameter update.
- Logits
- Raw, unnormalized scores output by a model before sigmoid or softmax.
- LSTM
- Long short-term memory: an RNN cell with forget, input and output gates and a cell state.
- Mixed precision
- Training with 16-bit arithmetic for speed and memory while keeping sensitive parts in FP32.
- MLP
- Multi-layer perceptron: stacked fully connected layers with non-linear activations.
- Mode collapse
- GAN failure where the generator produces only a few kinds of outputs.
- Momentum
- Optimizer technique that accumulates a moving average of gradients to accelerate and smooth updates.
- Overfitting
- Fitting training data (including noise) so closely that performance on new data suffers.
- Padding
- Adding border values (usually zeros) around an input so convolutions can cover edges and control output size.
- Parameter sharing
- Reusing the same weights at multiple positions or time steps, as in CNNs and RNNs.
- Perceptron
- A single neuron with a step activation; a linear binary classifier.
- Pooling
- Downsampling by taking the max or average over small windows.
- Pre-LN / Post-LN
- Where normalization sits relative to the residual add: inside the branch (Pre-LN, x + F(Norm(x))) or on the residual path (Post-LN, Norm(x + F(x))).
- Receptive field
- The region of the input that can influence a given unit.
- ReLU
- Rectified linear unit, max(0, x); the default hidden activation for CNNs and MLPs.
- Reparameterization trick
- Writing a sample as z = μ + σε so gradients can flow through a random sampling step.
- Residual connection
- A shortcut that adds a block's input to its output (y = F(x) + x).
- Residual stream
- The per-example (or per-token) vector that every residual block reads from and writes an additive update into.
- RMSNorm
- Normalization dividing by the root-mean-square of features without subtracting the mean.
- RNN
- Recurrent neural network: processes sequences step by step with a hidden state and shared weights.
- Saddle point
- A point with zero gradient that is a minimum in some directions and a maximum in others.
- Seq2seq
- Encoder-decoder model mapping an input sequence to an output sequence of possibly different length.
- Sigmoid
- Activation 1/(1 + e−x) mapping to (0, 1); used for binary outputs and gates.
- Skip-gram
- Word2vec variant that predicts context words from a centre word.
- Softmax
- Function that turns a vector of logits into a probability distribution summing to 1.
- Stride
- Number of positions a convolution or pooling window moves at each step.
- SwiGLU
- Gated feed-forward activation W2(SiLU(W1x) ⊙ W3x); hidden width is set near 8d/3 to keep the parameter count equal to a 4d MLP.
- Teacher forcing
- Feeding the true previous token to a decoder during training.
- Transfer learning
- Reusing a model trained on one task as the starting point for another.
- Underfitting
- A model too simple or undertrained to capture the patterns, with high training and validation error.
- Universal approximation theorem
- Result that a sufficiently wide one-hidden-layer network can approximate any continuous function on a bounded domain.
- VAE
- Variational autoencoder: an autoencoder with a probabilistic latent space regularized toward a standard normal.
- Vanishing gradient
- Gradients shrinking exponentially toward early layers or time steps, stalling their learning.
- Warmup
- Gradually increasing the learning rate at the start of training for stability.
- Weight decay
- Shrinking weights toward zero each step; equivalent to L2 regularization for plain SGD.
- Word2vec
- Self-supervised method that learns word embeddings by predicting words from context or context from words.
- Xavier initialization
- Random weights with variance 2/(fan_in + fan_out), designed for tanh and sigmoid networks.
Interview questions
Fundamentals
What is deep learning and how does it differ from classical machine learning?
Deep learning is a subset of machine learning that uses neural networks with many layers. The key difference is feature learning: classical ML typically relies on hand-engineered features fed into a model such as logistic regression or gradient-boosted trees, whereas deep networks learn hierarchical features directly from raw data (pixels to edges to parts to objects). Deep learning excels on unstructured data (images, audio, text) with lots of data and compute; classical ML often wins on small tabular datasets and when interpretability matters.
Why did deep learning take off around 2012 when neural networks had existed for decades?
Three things converged: (1) data, with large labelled datasets such as ImageNet and later web-scale unlabelled data; (2) compute, with GPUs making the matrix math 10-100x faster; (3) algorithms that fixed trainability: ReLU, dropout, better initialization, BatchNorm, residual connections and adaptive optimizers. Frameworks with automatic differentiation lowered the engineering barrier. AlexNet's large win on ImageNet in 2012 demonstrated the combination.
What is an artificial neuron? What do w, b and h mean?
A neuron computes a weighted sum of its inputs plus a bias, z = w·x + b, then applies an activation function, h = f(z). The weights w are learnable parameters on the connections that determine how strongly (and in which direction) each input influences the neuron, analogous to synapse strength. The bias b shifts the threshold. h is the neuron's activation (output), which is passed to the next layer. For example, if z = 0.4 then h = σ(0.4) ≈ 0.599: the number shown inside a node is its activation, not a weight.
What is a perceptron and what is its main limitation?
A perceptron is a single neuron with a step activation that outputs 0 or 1. It learns a linear decision boundary, so it can only solve linearly separable problems. It cannot learn XOR, because no single straight line separates XOR's positive and negative points. Adding a hidden layer (a multi-layer perceptron) solves this.
What is a multi-layer perceptron (feedforward network)?
An MLP is a stack of fully connected layers: input, one or more hidden layers, and an output layer. Each layer computes f(Wx + b). Information flows in one direction only (no loops or feedback), which is why it is also called a feedforward network. It works well for tabular data and as a building block (the feed-forward sublayer in Transformers) but has no built-in notion of spatial or sequential structure.
Why do neural networks need non-linear activation functions?
Without them, every layer is linear and a composition of linear maps is still linear: W2(W1x) = (W2W1)x. Any depth would collapse into a single linear layer, unable to learn curved decision boundaries. Activations introduce the "bends" that let stacked layers represent complex non-linear functions. Bounding outputs (as sigmoid does) is a secondary benefit useful mainly for probability outputs.
What is the difference between linear and non-linear relationships?
A linear relationship has a constant rate of change (y = 2x + 3: each unit of x adds 2 to y; its graph is a straight line, or a plane or hyperplane in higher dimensions). A non-linear one does not (y = x2, y = sin x). Not every non-linear function can be converted exactly into a linear one; neural networks instead build non-linear functions by combining linear transformations with non-linear activations.
What is the forward pass and why is it needed?
The forward pass feeds the input through the network layer by layer (linear transform, then activation) to produce a prediction ŷ. It is needed because the loss compares ŷ with the true label; without a prediction there is no loss, no gradient and no learning. The forward pass also caches intermediate values that backpropagation reuses.
What is a loss function and why do we need one?
A loss function measures how wrong the model's prediction is, as a single number. We need it to (1) quantify error and (2) provide the signal that gradient descent minimizes: predict, compute loss, compute gradients, adjust weights to reduce the loss, repeat. It acts as the feedback that tells the model how to learn from mistakes.
Are the loss function and the cost function the same thing?
Strictly, no: the loss is the error for one example, and the cost is the average (or sum) of losses over the dataset or batch, possibly plus a regularization term. In practice and in libraries the terms are used interchangeably; a formula with (1/n)Σ is technically the cost, but it is still commonly called "the loss" because it is the objective minimized during training.
Which loss should you use for regression and which for classification?
- Regression: MSE by default; MAE or Huber if there are outliers; quantile loss for asymmetric costs.
- Binary classification (and multi-label): binary cross-entropy with a sigmoid per output.
- Multi-class, one label per example: categorical cross-entropy with softmax.
- Imbalanced classes: weighted cross-entropy or focal loss.
Beyond the problem type, the business cost of different mistakes should drive the choice.
Why does MSE penalize large errors so heavily?
Because the error is squared. An error of 10 contributes 100, while an error of 50,000 contributes 2.5 billion. Large mistakes dominate the loss and the gradient, so the model focuses on fixing them first. This is desirable when big errors are genuinely worse, but it makes MSE sensitive to outliers. Squaring also makes the loss smooth and easy to optimize.
What is cross-entropy in simple terms?
Cross-entropy measures how good the predicted probabilities are compared with the true class: it is the negative log of the probability assigned to the correct class. True label 1 with prediction 0.9 gives −ln 0.9 ≈ 0.105 (low); prediction 0.1 gives 2.303 (high). Correct and confident means low loss; wrong and confident means very high loss. Formally, it is the negative log-likelihood of the labels under the model.
Why is there a negative sign in cross-entropy?
Probabilities are between 0 and 1, so their logarithms are negative (or zero). The minus sign makes the loss positive, and it turns "maximize the log-likelihood of the correct class" into "minimize the negative log-likelihood", which fits gradient descent's convention of minimizing.
How is a binary cross-entropy of about 3.0 obtained for a spam email predicted at 0.05?
With y = 1 and p = 0.05: BCE = −[1·ln(0.05) + 0·ln(0.95)] = −ln(0.05) ≈ 2.996. The second term vanishes because (1 − y) = 0. The value is high because the model assigned only 5% probability to the true class.
What is softmax, and is it a loss function?
Softmax is an activation that converts a vector of raw scores (logits) into probabilities that are positive and sum to 1: softmax(z)i = ezi / Σj ezj. For [2, 1, 0] it gives about [0.665, 0.245, 0.090]. It is computed, not predefined, and it is not a loss: softmax produces probabilities and categorical cross-entropy computes the error from them. They are typically used together (and fused inside CrossEntropyLoss).
Why do we use a sigmoid at the output for binary classification?
The last layer outputs an unbounded real number (a logit). Sigmoid maps it smoothly into (0, 1), so it can be read as a probability (0.9 = 90% spam) and plugged into binary cross-entropy, which is defined on probabilities. It is differentiable, so gradients flow. Sigmoid on top of the linear part: z = w·h + b, then ŷ = 1/(1 + e−z).
How does categorical cross-entropy handle class names such as "cat" or "dog"?
It never sees strings. Classes are encoded as integers (cat = 0, dog = 1, fish = 2) or one-hot vectors ([1, 0, 0] for cat). The model's softmax output, such as [0.7, 0.2, 0.1], is compared with the target, and the loss reduces to −log of the probability at the true index (−ln 0.7 ≈ 0.357 here). In PyTorch, CrossEntropyLoss takes integer class indices.
What is ReLU and why is it so popular?
ReLU(x) = max(0, x): it passes positive values unchanged and sets negatives to 0 (ReLU(3) = 3, ReLU(−2) = 0). It does not squash outputs into [0, 1]. It is popular because it is extremely cheap, its gradient is 1 for positive inputs (no saturation, so it mitigates vanishing gradients), and it produces sparse activations. Its weakness is "dying" units that get stuck at zero.
Compare sigmoid, tanh and ReLU.
| Sigmoid | Tanh | ReLU | |
|---|---|---|---|
| Range | (0, 1) | (−1, 1) | [0, ∞) |
| Max derivative | 0.25 | 1 | 1 (for x > 0) |
| Zero-centered | No | Yes | No |
| Saturates | Both ends | Both ends | Only for x < 0 (zero gradient) |
| Use | Binary outputs, gates | RNN states | Hidden layers of CNNs/MLPs |
What does "differentiable" mean here, and why does it matter?
A function is differentiable if it is smooth enough that we can compute its derivative (gradient) with respect to its inputs or parameters. Backpropagation needs the derivative of every operation to propagate the error signal, so losses and activations must be differentiable almost everywhere. The step function has zero derivative everywhere it is defined, so it cannot be trained with gradient descent; ReLU is non-differentiable only at a single point, which is handled with a subgradient.
What is gradient descent?
An iterative optimization algorithm that updates parameters in the direction that reduces the loss: θ ← θ − η∇L. The gradient gives the direction of steepest increase, so we step the opposite way, with the learning rate η controlling the step size. Libraries apply it (or variants such as Adam) automatically during training; you do not run it by hand.
What is the learning rate and how does it affect training?
It is the step size of each update. If the gradient suggests a weight should change by 10, a learning rate of 0.1 moves it by 1 and 1.0 moves it by 10. Too high: overshooting, oscillation or divergence (NaN). Too low: very slow learning, possibly stuck on plateaus. Choosing it balances convergence speed and stability; it is usually the most important hyperparameter.
What are epochs, batch size and iterations?
An epoch is one full pass over the training data. Batch size is the number of examples processed before one weight update. An iteration (step) is one such update. With 10,000 examples and batch size 100, one epoch has 100 iterations. Training typically runs many epochs so the model gradually improves.
What is the difference between parameters and hyperparameters? Are weights hyperparameters?
Parameters (weights and biases) are learned automatically by gradient descent. Hyperparameters are chosen before training and control how learning happens: learning rate, batch size, epochs, number of layers and units, dropout rate, weight decay, kernel size. Weights are not hyperparameters. Hyperparameters are also not derived from features; features are the input data. Tune hyperparameters on validation data.
How are weights initialized, and do they stay the same during training?
They start as small random values drawn from a carefully scaled distribution (Xavier for tanh/sigmoid, He for ReLU), not derived from correlations in the data. They do not stay the same: every training step updates them based on the gradient of the loss. They are not adjusted randomly; the gradient determines the direction. Over many epochs they converge to values that capture patterns in the data.
Do the weights of a neuron need to sum to 1, or be decimals?
No to both. Weights can be any real numbers, positive or negative, with no constraint on their sum; each is learned independently. They usually look like decimals simply because they are initialized as small random values and updated in small continuous steps. (Attention weights and softmax outputs do sum to 1, but those are activations, not learned weights.)
What is backpropagation in one paragraph?
Backpropagation computes the gradient of the loss with respect to every weight by applying the chain rule backward from the output: each layer multiplies the gradient coming from above by its local derivative and passes it down. It answers "if I nudge this weight, how much does the error change?" for all weights in one backward sweep. The optimizer then uses those gradients to update the weights.
In which phase do backpropagation and gradient descent happen?
Only during training. In each training step: forward pass, loss on training data, backpropagation, weight update. Validation and test use only forward passes with frozen weights (in PyTorch: model.eval() and torch.no_grad()) to measure generalization.
What are training loss and validation loss, and how do they differ?
Training loss is computed on training data and is used to update weights; it measures how well the model fits what it sees. Validation loss is computed on held-out data with no weight updates; it estimates performance on unseen data and guides hyperparameter choices and early stopping.
How do you tell whether a model is underfitting, overfitting or fitting well?
- Underfitting: training and validation loss both high.
- Good fit: both low and close (for example 0.10 vs 0.12).
- Overfitting: training loss very low, validation loss much higher or rising (for example 0.05 vs 0.30).
There are no universal numeric thresholds; the gap and the trend between the two curves matter, compared against a baseline.
Is there an ideal loss value that means the model is good?
No. The value depends on the loss type, number of classes and noise. Judge by: the loss decreasing over time, training and validation loss staying close, and beating a baseline. A useful reference is the uniform-guess cross-entropy ln(K), 0.693 for two classes.
Should we keep training until the training MSE reaches zero?
No. Real data contains noise, so zero training error means the model is memorizing noise and will likely overfit. Train until validation performance stops improving (early stopping) and keep the best checkpoint.
If the loss on training data reaches 0, is that always overfitting?
Not necessarily. Zero training loss means the model fits the training set perfectly; it is overfitting only if validation or test performance is much worse. A near-zero loss on validation data would be ideal generalization, not overfitting. In practice zero training loss on noisy data is a warning sign worth checking.
What is overfitting and how do you reduce it?
Overfitting is when the model learns the training data, including noise, too closely and fails to generalize. Remedies: more or augmented data, weight decay (L2), dropout, early stopping, a smaller model, label smoothing, transfer learning from pretrained weights, and cross-validation for model selection.
What are regularization and dropout?
Regularization is any technique that improves generalization at the cost of fitting training data less tightly. Classic regularization adds a penalty to the loss (L2: λΣw2, L1: λΣ|w|) to discourage large weights. Dropout randomly zeroes a fraction of activations during each training step so the network cannot rely on specific neurons and learns redundant, robust features; it is turned off at inference.
What is a CNN and what kind of data is it for?
A convolutional neural network applies small learnable filters across grid-structured data (data arranged on a regular grid with defined neighbours: images, spectrograms, video frames). It exploits spatial locality and parameter sharing, so it needs far fewer parameters than a dense network and recognizes a pattern wherever it appears. Typical structure: convolution + ReLU (+ BatchNorm) blocks, pooling or strided convolutions, then a classifier head.
What is a filter (kernel) in a CNN, and who decides its values?
A filter is a small matrix of learnable weights (for example 3×3 × input channels) that slides over the input computing a dot product with each patch; the result is a feature map showing where that pattern occurs. Nobody sets the values: they are initialized randomly and learned by backpropagation, exactly like other weights. Different tasks lead to different learned filters (edges and textures for general images, lesion-like patterns for medical images). The kernel size and the number of filters are hyperparameters.
What is padding and why use it?
Padding adds a border of extra pixels (usually zeros) around the input before convolution. It lets the filter process edge pixels properly and controls output size: a 5×5 input with a 3×3 filter gives 3×3 without padding but stays 5×5 with padding 1 ("same" padding). It changes the output size, never the kernel size.
What is pooling and why do we downsample?
Pooling summarizes small windows, for example taking the maximum of each 2×2 block with stride 2, halving height and width. We downsample to reduce computation and memory, drop redundant detail (neighbouring pixels are similar), increase the receptive field, and gain tolerance to small shifts: what matters is that a feature is present, not its exact pixel position. The trade-off is lost fine detail.
What is an RNN and what is a hidden state?
A recurrent neural network processes a sequence one element at a time, reusing the same weights at each step. The hidden state is its memory: a vector updated at every step from the current input and the previous hidden state (ht = tanh(Wxxt + Whht−1 + b)), summarizing everything seen so far, like the running understanding you keep while reading a sentence. It is computed, not set manually.
What is the difference between a CNN and an RNN?
A CNN is built for spatial, grid data: it looks at local patches with shared filters and processes all positions in parallel, focusing on where patterns occur. An RNN is built for sequential data: it processes elements step by step, carrying a hidden state that remembers what came before, focusing on order and context. They can be combined, for example a CNN for per-frame features and an RNN across frames for video.
What is an embedding?
A dense, learned vector representing a discrete item (word, token, product, user) such that similar items are close in the vector space. It replaces sparse one-hot vectors, which are huge and treat every pair of items as unrelated. Embeddings are learned by backpropagation, either for a specific task or with self-supervised objectives such as word2vec, and enable similarity search, clustering and transfer learning.
Why are GPUs preferred over CPUs for deep learning?
Neural networks are dominated by matrix multiplications, which consist of millions of independent multiply-adds. CPUs have a few powerful cores optimized for sequential logic; GPUs have thousands of simpler cores and tensor cores that execute these operations in parallel, with much higher memory bandwidth. Training that takes weeks on CPUs can take hours on GPUs.
Going deeper
Walk through a full training step for a 2-layer network, including the backward pass.
Example: x = 2, hidden ReLU with w1 = 0.3, b1 = 0.1; sigmoid output with w2 = 0.5, b2 = 0.2; label y = 1; BCE loss; η = 0.1.
- Forward: z1 = 0.7, h = 0.7, z2 = 0.55, ŷ = 0.634, L = −ln 0.634 = 0.456.
- Backward: ∂L/∂z2 = ŷ − y = −0.366; ∂L/∂w2 = −0.366 × 0.7 = −0.256; ∂L/∂h = −0.366 × 0.5 = −0.183; ReLU passes it (z1 > 0); ∂L/∂w1 = −0.183 × 2 = −0.366.
- Update: w2 = 0.526, w1 = 0.337, biases similarly; the new loss is 0.419.
Key points: cache forward values, multiply upstream gradient by local derivative, and each weight's gradient is proportional to the activation feeding it.
Why is the derivative of the sigmoid output with BCE simply ŷ − y?
∂L/∂ŷ = −y/ŷ + (1 − y)/(1 − ŷ) and ∂ŷ/∂z = ŷ(1 − ŷ) (since σ' = σ(1 − σ)). Multiplying: −y(1 − ŷ) + (1 − y)ŷ = ŷ − y. The same holds for softmax plus categorical cross-entropy (gradient p − y). This clean gradient never saturates when the model is confidently wrong, which is why cross-entropy trains better than MSE for classification.
Why is cross-entropy preferred over MSE for classification?
- It is the maximum-likelihood objective for categorical outputs.
- Its gradient with respect to logits is p − y; with MSE the gradient is multiplied by σ'(z), which is tiny when the output is saturated and confidently wrong, so learning stalls.
- It heavily penalizes confident mistakes, encouraging meaningful probabilities.
- For logistic regression it gives a convex problem; MSE does not.
How do you choose a loss when the business penalizes over-forecasting and under-forecasting differently?
Use an asymmetric loss. If actual demand is 100, predicting 120 and 80 have identical MSE and MAE, but over-forecasting may cost inventory while under-forecasting loses sales. Quantile (pinball) loss with τ > 0.5 penalizes under-prediction more (the model learns a higher quantile); a custom weighted loss (different multipliers for positive and negative errors) also works. MSE penalizes big errors with no direction preference; MAE treats all errors equally.
Is regularization (lambda) part of binary cross-entropy?
No. BCE is the data loss. Lambda is the regularization strength, a hyperparameter for a separate penalty term added on top: total = BCE + λΣw2 (or implemented as weight decay in the optimizer). It is applied to control overfitting and is independent of which data loss you use.
What is calibration, and how does it differ from accuracy?
Accuracy measures how often predictions are correct. Calibration measures whether predicted probabilities match reality: among predictions made with 70% confidence, about 70% should be right. A model can be accurate but over-confident. Calibration matters when decisions depend on the probability (risk scoring, medical triage, thresholds). Fix it with temperature scaling on a validation set, label smoothing, or isotonic/Platt scaling; these change probabilities but not rankings.
Where is the confidence represented in BCE, and where is the classification threshold set?
Confidence is the predicted probability of the positive class (the sigmoid output); the log term makes confident errors expensive. The threshold is not part of BCE; it is applied to the probability at decision time, 0.5 by default, and should be tuned on validation data (for example 0.3 for high recall in screening, 0.7 for high precision in fraud alerts) using precision-recall or ROC analysis.
Which models can output predicted probabilities?
Logistic regression (sigmoid output), neural networks with sigmoid or softmax outputs, Naive Bayes, and tree ensembles (class frequencies in leaves or votes). SVMs output margins, not probabilities, unless calibrated (Platt scaling). Not all "probabilities" are well calibrated; check with a reliability diagram.
Why can ReLU introduce non-linearity when it looks linear for positive inputs?
ReLU is piecewise linear: 0 for negative inputs, identity for positive. The kink at zero breaks global linearity. Different neurons switch on or off for different inputs, so the network becomes a combination of many linear regions whose boundaries depend on the input. Enough regions can approximate any curve. If every unit were purely linear, layers would collapse into one linear map.
ReLU is not differentiable at zero. How does backpropagation deal with that?
Frameworks use a subgradient at exactly 0 (usually 0, sometimes 1). With real-valued inputs, landing exactly on 0 is extremely rare, and everywhere else ReLU has a well-defined derivative (0 or 1), so gradient descent works smoothly in practice.
What is the dying ReLU problem and how do you fix it?
If a neuron's pre-activation is negative for every input (often after a large update pushes its bias very negative), ReLU outputs 0 and its gradient is 0, so it never updates again and effectively dies. Fixes: Leaky ReLU, PReLU, ELU or GELU (small gradient for negatives); He initialization; lower learning rate; normalization layers; monitoring the fraction of dead units.
Why do Transformers use GELU instead of ReLU?
GELU (x·Φ(x)) is smooth and lets small negative values pass with a small non-zero gradient (GELU(−1) ≈ −0.16), avoiding ReLU's hard cutoff and dead units. For positive inputs it behaves almost like ReLU. Empirically it trains slightly better in large Transformers; many recent LLMs use SiLU inside SwiGLU gated feed-forward blocks for the same reasons.
Softmax or sigmoid for a multi-label problem (an image can contain both a cat and a dog)?
Sigmoid on each output with binary cross-entropy per label. Softmax forces probabilities to compete and sum to 1, which is right only when exactly one class is true. With independent sigmoids, both "cat" and "dog" can be near 1. Tune a threshold per label.
Explain momentum, RMSProp and Adam, and when to use each.
- Momentum keeps a moving average of gradients (velocity), accelerating along consistent directions and damping oscillations.
- RMSProp divides each parameter's step by the running root-mean-square of its gradients, so steep directions get smaller steps and flat ones larger.
- Adam combines both, with bias correction for the zero-initialized averages. Default choice for most problems and essential for Transformers.
SGD with momentum and a good schedule can generalize slightly better for CNN image classification. Full-batch GD uses all data per step; SGD uses one example; mini-batch (most common) uses a small batch.
Why does Adam need bias correction?
Its moment estimates m and v start at zero. Early on they are biased toward zero (after one step with β1 = 0.9, m = 0.1g). Dividing by (1 − βt) makes them unbiased estimates of the gradient mean and uncentered variance. The correction matters most for v, whose β2 = 0.999 means it would otherwise stay tiny for thousands of steps, producing overly large early updates.
What is the difference between L2 regularization and weight decay, and why AdamW?
For plain SGD they are equivalent: the L2 penalty's gradient λw shrinks weights proportionally each step. In Adam, the L2 gradient is divided by the adaptive denominator √v̂, so parameters with large historical gradients receive almost no regularization. AdamW applies weight decay directly to the weights (w ← w − ηλw) separately from the adaptive update, giving uniform, predictable regularization and better generalization.
How do you choose a learning rate, and does it adjust automatically as the slope flattens?
It does not adjust automatically in vanilla gradient descent: the gradient sets the direction, the learning rate is a fixed multiplier. Choose it with an LR range test (increase exponentially, pick about 10x below where the loss falls fastest), search on a log scale (1e−4, 3e−4, 1e−3...), then apply a schedule (warmup plus cosine or step decay). Adaptive optimizers scale per-parameter steps but still need a base learning rate.
What is learning-rate warmup and why is it needed?
Warmup linearly increases the learning rate from near zero to its peak over the first few hundred or thousand steps. At the start, weights are random, gradients are large and poorly aligned, and Adam's variance estimates are noisy, so a full learning rate can cause immediate divergence. Warmup is standard for Transformers, large batches and adaptive optimizers.
Why can't all weights be initialized to zero?
All neurons in a layer would compute the same output, receive the same gradient and receive the same update, so they would remain identical forever (the symmetry problem): a wide layer would behave like one neuron. Random initialization breaks symmetry. Biases can start at zero because the random weights already differentiate neurons.
Explain Xavier and He initialization and why they differ by a factor of 2.
Both choose weight variance so that activation and gradient variance stay roughly constant across layers. Xavier uses Var = 2/(fan_in + fan_out), derived for symmetric activations such as tanh. He uses Var = 2/fan_in for ReLU: because ReLU zeroes about half its inputs, it halves the variance, so doubling the weight variance compensates. Wrong scaling makes signals vanish or explode with depth.
How does dropout work, and what changes at inference?
During training each activation is zeroed with probability p and survivors are scaled by 1/(1 − p) (inverted dropout) so the expected value stays the same. A fresh mask is sampled every step. At inference dropout is disabled and no scaling is needed. It prevents co-adaptation and approximates averaging many sub-networks. In PyTorch, model.train() and model.eval() switch this behaviour.
Compare L1 and L2 regularization.
L2 (λΣw2) shrinks all weights smoothly toward zero, keeps every feature, is differentiable everywhere and corresponds to a Gaussian prior. L1 (λΣ|w|) drives many weights exactly to zero (its constraint region is a diamond whose corners lie on the axes), performing feature selection, and corresponds to a Laplace prior; it is not differentiable at zero. Deep learning mostly uses L2/weight decay.
What is early stopping and how is it implemented?
Monitor validation loss (or metric) each epoch, save a checkpoint whenever it improves, and stop when it has not improved for a "patience" number of epochs (for example 5); then restore the best checkpoint. Example: validation loss 0.12 at epoch 5, 0.10 at epoch 10, 0.15 at epoch 15, so the epoch-10 weights are kept. It limits effective training time and is equivalent to a form of regularization.
How does BatchNorm work, and how does it behave differently in training and inference?
For each channel it subtracts the mini-batch mean, divides by the mini-batch standard deviation, then applies learned scale γ and shift β. In training it uses batch statistics and updates exponential running averages, μ ← (1 − m)μ + m μB (PyTorch momentum m = 0.1). The forward pass uses the biased batch variance (divide by n); the running variance is updated with the unbiased estimate (divide by n − 1). In inference it uses only the running averages, so outputs are deterministic and independent of other examples in the batch. It enables higher learning rates, reduces sensitivity to initialization and adds mild regularization. Forgetting model.eval() is the classic bug that makes predictions depend on who else is in the batch.
Why do Transformers use LayerNorm instead of BatchNorm?
LayerNorm normalizes each token's features independently of other examples, so it works with any batch size (including 1), variable sequence lengths and padding, and it behaves identically in training and autoregressive inference. BatchNorm's statistics would mix tokens across examples and positions, are noisy with small batches and break during token-by-token generation. Many recent LLMs use RMSNorm, a cheaper variant.
What causes vanishing gradients and how can they be fixed?
Backprop multiplies derivatives through every layer (or time step). If those factors are below 1 (sigmoid's derivative is at most 0.25, so 10 layers give at most 10−6), the gradient reaching early layers becomes negligible and they stop learning. Fixes: ReLU-family activations, Xavier/He initialization, normalization layers, residual connections, LSTM/GRU gates for sequences, and shorter paths such as attention.
What are exploding gradients and how do you handle them?
Gradients growing exponentially through depth or time (per-layer factor 1.5 gives 1.520 ≈ 3,300), causing huge updates, loss spikes and NaN. Common in RNNs. Remedies: gradient clipping by norm (for example 1.0), lower learning rate with warmup, proper initialization, normalization layers, and gated or residual architectures. Monitor the gradient norm to detect it.
What is ResNet and why does it work?
ResNet adds identity shortcuts around blocks: y = F(x) + x. It addressed the degradation problem, where deeper plain networks had higher training error than shallower ones. The block only needs to learn a residual, and learning "do nothing" is easy (F = 0). The identity path gives gradients a direct route back (the Jacobian contains +I), enabling networks of 100+ layers. Residual connections are now standard in almost every deep architecture, including Transformers.
How do you calculate the output size of a convolution? Give examples.
Output = ⌊(W + 2P − K)/S⌋ + 1. A 5×5 input with a 3×3 kernel, no padding, stride 1 gives 3×3 (three positions per row). A 32×32 input with a 5×5 kernel gives 28×28. A 224×224 input with a 7×7 kernel, stride 2, padding 3 gives 112×112. A 4×4 map with 2×2 pooling, stride 2 gives 2×2; with a 3×3 window and stride 1 it also gives 2×2.
How many parameters does a convolutional layer have?
(Cin × k × k + 1) × Cout, including one bias per filter. Conv2d(3, 64, 3): (27 + 1) × 64 = 1,792. Conv2d(64, 128, 3): (576 + 1) × 128 = 73,856. The count does not depend on the image size, unlike a dense layer.
What is parameter sharing and why does it matter?
Using the same weights at multiple locations: in a CNN the same filter is applied at every spatial position; in an RNN the same matrices at every time step. It drastically reduces parameters (a 3×3×3 filter bank with 64 filters has 1,728 weights versus millions for a dense layer), improves generalization, and builds in translation equivariance: a feature detector works anywhere in the image.
What do the arguments of nn.Conv2d(16, 32, kernel_size=3, padding=1) mean?
16 input channels (feature maps arriving from the previous layer), 32 output channels (the number of filters, hence feature maps produced), 3×3 kernels, and padding of 1 so the spatial size is preserved at stride 1. It has (16×9 + 1)×32 = 4,640 parameters. The numbers are not image sizes.
Max pooling versus average pooling: when do you use each? Is pooling the same as dimensionality reduction?
Max pooling keeps the strongest activation, good for detecting whether a feature is present (classification). Average pooling gives a smoother summary; global average pooling is common before the classifier head. Median pooling is rare. Pooling is similar to dimensionality reduction in that it shrinks data, but it is local, fixed and spatial (4×4 to 2×2), whereas methods like PCA are global, learned projections of the feature dimension (100 features to 10).
Why are several Conv + ReLU layers stacked, and how do conv and fully connected weights get trained?
Each layer builds on the previous: early layers detect edges, middle layers textures and parts, deeper ones objects, while the receptive field grows. ReLU between them prevents the stack from collapsing into one linear operation. All parameters (filters and dense weights) are trained simultaneously: one forward pass, one loss, one backward pass computes gradients for every layer, and the optimizer updates them together.
What happens when feature maps are flattened, and is the result an embedding?
Flattening reshapes the C×H×W feature tensor into a 1D vector without losing information; each neuron of the next dense layer then connects to every element, with a weight matrix of size (neurons × vector length) that is randomly initialized and learned. That vector (or a global-pooled version or a hidden layer after it) is effectively an image embedding. Unlike word2vec, it is task-dependent unless the network was trained to produce general-purpose features.
Explain the LSTM gates.
- Forget gate f = σ(Wf[h, x] + bf): what fraction of each cell-state entry to keep.
- Input gate i: how much of the candidate to write.
- Candidate c̃ = tanh(Wc[h, x] + bc): proposed new information.
- Cell update c = f ⊙ cprev + i ⊙ c̃.
- Output gate o: how much of tanh(c) to expose as the hidden state h, which is used for predictions.
All gates are computed in parallel at each step; [h, x] is concatenation and ⊙ is element-wise multiplication.
Why does an LSTM need a forget gate if the input gate already filters information?
The input gate only controls what enters at the current step; it has no control over what is already stored. Without a forget gate, the cell state would keep accumulating information, including outdated or irrelevant details ("Delhi" after the text moves on to "Bangalore"). The forget gate lets the model selectively clear memory. If it wrongly forgets something important, the loss increases and backpropagation adjusts it.
How does the LSTM reduce vanishing gradients? Can it still suffer from them?
The cell state is updated additively, so ∂ct/∂ct−1 = ft (element-wise) rather than a product involving a weight matrix and a squashing derivative. When the forget gate stays near 1, gradients flow across many steps almost unchanged. It reduces the problem greatly but does not eliminate it on very long sequences, and it does not prevent exploding gradients, so clipping is still used.
What is a GRU and how does it compare with an LSTM?
A GRU has an update gate (how much of the state to replace, merging forget and input) and a reset gate (how much past state to use when proposing a candidate), with no separate cell state. It has about 25% fewer parameters than an LSTM of the same size, is faster and often performs similarly, especially on smaller datasets. LSTMs can be more expressive for very long dependencies. GRU is not universally better; try both.
Are RNN weights different for each word in the sequence?
No. The same Wx and Wh are reused at every time step; a 5-word sentence uses one set of weights five times. That is why the number of parameters does not depend on sequence length. An RNN has one or a few layers but runs across many time steps; time steps are not separate layers.
When are RNN weights updated: after every word or after the sequence?
Typically after processing the whole sequence (or a batch of sequences): the forward pass runs all time steps, the loss is computed, backpropagation through time computes gradients, and one update is applied. Truncated BPTT updates after chunks of k steps for long streams, for efficiency.
What is a bidirectional LSTM and why is memory still needed in each direction?
It runs one LSTM left-to-right and another right-to-left and concatenates their states at each position, so each token's representation uses both past and future context. Each direction still needs its cell state and gates to decide what to keep over long distances; reading backward only provides future context, not memory management. BiLSTMs cannot be used for streaming generation because they need the full input.
What is word2vec? Compare CBOW and skip-gram.
Word2vec learns word embeddings from raw text by prediction within a context window. CBOW predicts the centre word from the averaged context words ("I am ___ AI" to "learning"); it is faster and good for frequent words. Skip-gram predicts context words from the centre word; slower, better for rare words. They are alternative formulations, not used together. Training uses negative sampling instead of a full softmax. It is self-supervised learning, not reinforcement learning.
Word2vec starts from one-hot vectors. How does it capture relationships?
The one-hot vector only identifies the word; multiplying it by the input weight matrix selects one row, the word's embedding. Training adjusts these rows so that words appearing in similar contexts produce similar predictions, which pulls their vectors together. The relationships live in the learned weights, not in the one-hot codes. Initial values are random; the final positive and negative numbers reflect learned patterns, not human-interpretable meanings.
How is the embedding dimension (for example 300) chosen?
It is a hyperparameter chosen empirically. Classic word vectors found 100-300 dimensions large enough to capture semantic relationships and small enough to be efficient. Too small cannot encode enough structure; too large wastes compute and may overfit. Modern Transformer models use 768 to several thousand dimensions. Tune it on downstream performance.
How is cosine similarity computed and interpreted?
cos(a, b) = (a·b)/(||a|| ||b||): dot product divided by the product of the vector lengths. It measures the angle, not magnitude: 1 means same direction (very similar), 0 unrelated, −1 opposite. Example: [1, 2] and [2, 4] give 1.0. One-hot vectors of different words always give 0, which is why they cannot express similarity.
What is transfer learning and how would you fine-tune a pretrained CNN?
Transfer learning reuses a model trained on a large dataset as the starting point for a new task. Steps: replace the classification head, match the pretrained preprocessing, freeze the backbone and train the head, then unfreeze later blocks with a 10-100x smaller learning rate (discriminative rates), use augmentation and early stopping, and keep BatchNorm in eval mode for small batches. With little data, feature extraction alone often suffices.
What is the difference between an autoencoder and a VAE?
An autoencoder maps each input to a single latent point and reconstructs it; its latent space can have gaps, so random codes decode poorly. A VAE encodes a distribution (μ, σ), samples from it via the reparameterization trick, and adds a KL term pulling the distributions toward N(0, I), making the latent space smooth and sampleable for generation. VAEs trade some sharpness (blurrier outputs) for a principled generative model.
How does a GAN work?
A generator maps random noise to samples; a discriminator classifies samples as real or fake. They are trained alternately: the discriminator to tell real from fake, the generator to fool it (in practice maximizing log D(G(z))). At equilibrium the generator matches the data distribution and the discriminator outputs 0.5. GANs produce sharp images and fast sampling but are unstable to train and prone to mode collapse.
What is mixed-precision training?
Running most computation (matrix multiplications, convolutions) in 16-bit floats for about 2x speed and half the activation memory, while keeping sensitive operations (softmax, normalization, loss, weight updates, a master copy of weights) in FP32. FP16 has a narrow range, so it needs dynamic loss scaling to avoid gradient underflow; BF16 has FP32's range and usually needs no scaling. In PyTorch: torch.autocast plus GradScaler for FP16.
What is gradient accumulation and when do you use it?
Running several micro-batches, calling backward on each so gradients sum in .grad, and stepping the optimizer once. It simulates a large batch when memory is limited (8 micro-batches of 32 behave like a batch of 256). Divide each loss by the number of accumulation steps. BatchNorm still sees only the micro-batch size.
Write the standard PyTorch training step and explain each line.
model.train()
for x, y in loader:
x, y = x.to(device), y.to(device)
optimizer.zero_grad() # clear accumulated gradients
logits = model(x) # forward pass (raw scores)
loss = loss_fn(logits, y) # e.g. CrossEntropyLoss on logits
loss.backward() # backprop: fill .grad for all parameters
optimizer.step() # update weights using the gradientsEvaluation uses model.eval() and with torch.no_grad():. Softmax is applied inside CrossEntropyLoss during training and explicitly only at inference when probabilities are needed.
Where does the softmax happen if forward() returns raw logits?
During training, inside the loss: nn.CrossEntropyLoss applies log-softmax internally and combines it with negative log-likelihood, which is more numerically stable than applying softmax yourself. During inference, apply torch.softmax(logits, dim=1) if you need probabilities; for the predicted class alone, argmax of the logits is enough because softmax preserves order.
How do stride and padding change what a CNN learns?
Larger strides shrink feature maps faster, reducing computation and capturing coarser, more global features but potentially skipping fine details. Smaller strides keep spatial detail at higher cost. Padding preserves spatial size and ensures border pixels are covered as often as central ones; without it maps shrink every layer and edge information is under-represented.
Advanced
What does the universal approximation theorem guarantee, and why do we still use deep networks?
It guarantees that a network with one hidden layer and a non-polynomial activation can approximate any continuous function on a compact domain arbitrarily well, given enough units. It does not say how many units are needed, whether gradient descent will find the weights, or whether the result generalizes. Depth is exponentially more efficient for many compositional functions, reuses intermediate features, and empirically generalizes better; that is why deep rather than wide networks are used.
Why is backpropagation efficient compared with numerical differentiation, and what is its cost?
Finite differences need one or two forward passes per parameter, O(W) passes for W parameters. Reverse-mode automatic differentiation computes the gradient with respect to all parameters in one backward pass costing about 2x the forward pass, because the loss is a scalar and intermediate results are reused. The cost is memory: activations from the forward pass must be stored (or recomputed with checkpointing). Forward-mode AD is efficient in the opposite case (few inputs, many outputs).
The loss surface has many local minima. How can we be sure we have found the global minimum?
We cannot, and we do not need to. Deep network losses are non-convex; in high dimensions most critical points are saddle points, and many local minima have loss close to the global one. SGD noise, momentum, random initialization, learning-rate schedules and restarts help find good regions. Success is judged by validation performance: flat, wide minima tend to generalize better than sharp ones. Convex problems (linear regression) have a single global minimum, but deep learning targets robust generalization rather than exact optimality.
Why does large-batch training sometimes generalize worse, and how is it fixed?
Large batches reduce gradient noise, which can lead to sharper minima, and give fewer updates per epoch. Fixes: scale the learning rate with batch size (linear rule for SGD, square-root heuristics for Adam), use warmup, train for enough steps, use layer-wise adaptive optimizers (LARS, LAMB) at very large scale, and add regularization. With these, batches of tens of thousands train well.
Derive why He initialization uses variance 2/fan_in.
For z = Σ wixi with independent zero-mean weights, Var(z) = fan_in · Var(w) · E[x2]. If x = ReLU(zprev) with zprev symmetric around zero, E[x2] = ½Var(zprev). Keeping Var(z) = Var(zprev) requires fan_in · Var(w) · ½ = 1, so Var(w) = 2/fan_in. Xavier balances the forward (fan_in) and backward (fan_out) conditions for symmetric activations, giving 2/(fan_in + fan_out).
Why does BatchNorm help optimization? Is "internal covariate shift" the right explanation?
The original motivation was reducing internal covariate shift (changing layer-input distributions). Later experiments showed that injecting shift after BatchNorm barely hurt, and that its main effect is smoothing the loss landscape (more Lipschitz loss and gradients), which permits larger learning rates. It also makes the function invariant to weight scale, so gradient descent effectively adapts its step size, and batch noise regularizes. Its drawbacks are batch-size dependence and train/inference discrepancy.
Pre-norm versus post-norm Transformers: what is the difference and why does it matter?
Post-norm applies LayerNorm after the residual addition: LN(x + F(x)). Gradients to early layers pass through many LayerNorms, which makes deep post-norm models unstable without careful warmup. Pre-norm applies it inside the branch: x + F(LN(x)), leaving an unnormalized identity path (the residual stream) that carries gradients cleanly; it trains stably at depth and is used by most modern LLMs, sometimes with a final LayerNorm. Post-norm can reach slightly better quality when it trains successfully.
Explain the degradation problem and why residual networks behave like ensembles.
Degradation: deeper plain networks had higher training error than shallower ones, an optimization failure rather than overfitting. Residual connections make identity mappings trivial, so extra depth cannot hurt. Unrolling a residual network shows it is a collection of paths of different lengths; experiments deleting single blocks at test time show only mild degradation, and most gradient flows through relatively short paths, so ResNets behave partly like ensembles of shallower networks.
Compute the cost saving of depthwise separable convolutions.
A standard k×k convolution costs H·W·Cin·Cout·k2 multiply-adds. Depthwise (H·W·Cin·k2) plus pointwise (H·W·Cin·Cout) gives a ratio of 1/Cout + 1/k2. For k = 3 and Cout = 256, that is about 0.115, roughly 8.7x fewer operations and parameters, with a small accuracy cost. This is the basis of MobileNet and EfficientNet.
What is the receptive field and how do you compute it for stacked layers?
The input region that influences a unit. For stacked layers, rl = rl−1 + (kl − 1) × jl−1, where j is the cumulative stride (product of previous strides). Three 3×3 stride-1 convolutions give 7×7. Strides and pooling multiply growth, and dilated convolutions expand it without extra parameters. The effective receptive field (where influence is concentrated) is smaller than the theoretical one and roughly Gaussian.
Are CNNs truly translation invariant?
Convolution is translation equivariant; pooling and global averaging add approximate invariance. It is not perfect: strided convolutions and pooling cause aliasing, so shifting an image by one pixel can change outputs noticeably; padding leaks absolute position information; and CNNs are not inherently invariant to rotation, scale or viewpoint. Anti-aliased downsampling (blur before subsampling), augmentation and equivariant architectures help.
Explain backpropagation through time and why vanilla RNNs struggle with long dependencies, mathematically.
Unroll the RNN over T steps and backprop through the unrolled graph. The gradient from step t to step k includes ∏i=k+1..t ∂hi/∂hi−1 = ∏ diag(tanh'(·)) Whh. If the largest singular value of Whh times the maximal tanh' is below 1, the product decays exponentially in (t − k); above 1 it can explode. Hence dependencies beyond tens of steps are hard to learn. Truncated BPTT caps the backward length for efficiency, which further limits learnable dependency length.
Why is the LSTM forget-gate bias often initialized to 1?
With zero bias the forget gate starts near σ(0) = 0.5, so the cell state halves every step and early gradients vanish over time. A bias of 1 (gate ≈ 0.73) or more makes the LSTM remember by default at the start, letting gradients flow across many steps until the model learns what to forget. It is a simple change that improves learning of long-range dependencies.
Count the parameters of an LSTM layer.
Four blocks (forget, input, output, candidate), each with a weight matrix of size h × (h + x) and a bias of size h: 4(h(h + x) + h). For x = 100, h = 128: 4 × (128 × 228 + 128) = 117,248. PyTorch stores separate input and hidden biases, adding 4h = 512. A GRU has 3 blocks; a bidirectional layer doubles the count.
What is teacher forcing, and what problem does it create?
During training, the decoder receives the ground-truth previous token instead of its own prediction, which speeds convergence and stabilizes training. It causes exposure bias: at inference the model conditions on its own outputs, which may contain errors it never saw during training, so errors compound. Mitigations include scheduled sampling, sequence-level training objectives, and in large language models the sheer scale and diversity of pretraining data.
Why does seq2seq performance drop on rare or unknown words, and do pointer-generator networks help?
A fixed vocabulary maps rare or unseen words to an unknown token, so the model loses their identity and cannot reproduce names, numbers or technical terms. Pointer-generator networks combine generating from the vocabulary with copying tokens directly from the source via attention, controlled by a learned switch probability; they clearly outperform plain attention seq2seq on summarization with many named entities. Subword tokenization (BPE, WordPiece) and Transformers address the same issue in modern systems.
What problem did attention solve, and why did it lead to Transformers?
Seq2seq squeezed the entire source into one fixed-size vector, losing information on long inputs. Attention computes a new context vector per output step as a weighted sum of all encoder states, with weights from query-key similarity, providing direct O(1) paths between any positions and interpretable alignments. Since attention alone provided the long-range connectivity, recurrence could be dropped: self-attention processes all tokens in parallel, enabling efficient training on GPUs and better scaling, which is the Transformer.
In a world dominated by Transformers, where do RNNs or LSTMs still have practical advantages?
Streaming and real-time settings where inputs arrive step by step and latency and memory must be constant: on-device speech recognition and keyword spotting, sensor and time-series processing on microcontrollers, and some control tasks. An RNN keeps a fixed-size state, while a Transformer's attention cache grows with context and attention costs grow quadratically. Recurrent-style state-space and linear-attention models bring this advantage to large-scale modelling.
How do static embeddings, ELMo and Transformer embeddings differ?
Static embeddings (word2vec, GloVe, fastText) assign one vector per word type regardless of context, so polysemous words like "bank" get a single vector, and pure word-level models have no vector for unseen words. ELMo produces contextual embeddings from a deep bidirectional LSTM language model, combining layers, at higher computational cost. Transformer models (BERT, GPT) produce contextual embeddings via self-attention over subword tokens, capturing both-side context (BERT) or left context (GPT) and scaling much better. Regularize heavy contextual models with dropout, weight decay and early stopping.
What is negative sampling and why is it needed?
The full softmax over a vocabulary of V words costs O(V) per training pair. Negative sampling replaces it with binary logistic classification: raise the score of the true (centre, context) pair and lower the scores of k randomly sampled "negative" words (k ≈ 5-20), drawn from the unigram distribution raised to the 3/4 power. Cost becomes O(k). The same idea underlies contrastive losses such as InfoNCE.
Explain the VAE loss and the reparameterization trick.
The VAE maximizes the evidence lower bound: Eq(z|x)[log p(x|z)] − KL(q(z|x) || p(z)). The first term is reconstruction; the second keeps the approximate posterior close to the prior N(0, I), with a closed form ½Σ(μ2 + σ2 − log σ2 − 1) for Gaussians. Sampling z ~ N(μ, σ2) is not differentiable with respect to μ and σ, so write z = μ + σε with ε ~ N(0, I), moving the randomness outside the computation graph so gradients flow.
Why are GANs hard to train, and what does the Wasserstein GAN change?
GAN training is a two-player game, not a minimization, so there is no monotonic objective: oscillation, mode collapse and vanishing generator gradients (when the discriminator wins easily, since the Jensen-Shannon divergence saturates when distributions do not overlap) are common. WGAN replaces the discriminator with a critic estimating the Earth-Mover (Wasserstein-1) distance under a Lipschitz constraint (weight clipping, or better a gradient penalty or spectral normalization), which gives useful gradients even when distributions are disjoint and a loss that correlates with sample quality.
What does a diffusion model learn, and how is text conditioning added?
The forward process mixes data with Gaussian noise over T steps, with a closed form xt = √ᾱtx0 + √(1 − ᾱt)ε. A network (U-Net or diffusion Transformer) learns ε̂(xt, t) with an MSE loss. Sampling starts from noise and iteratively removes predicted noise. Text conditioning injects prompt embeddings via cross-attention; classifier-free guidance trains with the prompt randomly dropped and at inference extrapolates ε̂uncond + w(ε̂cond − ε̂uncond). Latent diffusion runs in a VAE's latent space for efficiency.
Estimate the GPU memory needed to fully fine-tune a 7B-parameter model with AdamW in mixed precision.
About 16 bytes per parameter before activations: BF16 weights (2) + BF16 gradients (2) + FP32 master weights (4) + Adam m and v (8). 7B × 16 bytes ≈ 112 GB, plus activations that depend on batch size and sequence length. That exceeds a single 80 GB GPU, so options are FSDP/ZeRO sharding across GPUs, CPU offloading, activation checkpointing, 8-bit optimizers, or parameter-efficient fine-tuning (LoRA/QLoRA), which trains only small adapter matrices.
Compare DDP, FSDP/ZeRO, tensor parallelism and pipeline parallelism.
- DDP: full replica per GPU, different data shards, gradients all-reduced each step. Simple and efficient if the model fits.
- FSDP/ZeRO: shards optimizer state (stage 1), gradients (2) and parameters (3) across GPUs, gathering weights just in time. Fits much larger models at extra communication cost.
- Tensor parallelism: splits individual matrix multiplications across GPUs; very communication-heavy, used within a node with fast interconnects.
- Pipeline parallelism: places consecutive layers on different GPUs and streams micro-batches; suffers pipeline bubbles.
Frontier models combine all of them (3D parallelism) matched to the cluster topology.
FP16 versus BF16: why does BF16 usually not need loss scaling?
FP16 has 5 exponent bits (maximum about 65,504, and small gradients below about 6e−8 underflow to zero), so gradients must be scaled up before backward and unscaled afterwards. BF16 has 8 exponent bits, the same range as FP32, so underflow and overflow are rare; it sacrifices mantissa precision (7 bits), which neural networks tolerate well. Accumulations and weight updates are still done in FP32.
What is activation checkpointing, and what does it trade?
Instead of storing all intermediate activations for the backward pass, store only some (for example at the boundaries of each Transformer block) and recompute the others during backward. With checkpoints every √n layers, activation memory drops from O(n) to about O(√n) for roughly one extra forward pass (about 30% more compute). It enables longer sequences and larger batches.
Is there a principled way to predict whether an architectural change will help, before running huge experiments?
Partly. Reason about inductive bias and gradient flow first: does the change preserve a clean gradient path, keep activation variance stable, avoid information bottlenecks, and give the model a way to discard stale information (removing an LSTM's forget gate, for example, would let memory clutter accumulate)? Then run small controlled ablations: small models, synthetic tasks that isolate the capability (copying, long-range recall), and diagnostic signals such as training and validation loss, gradient norms, activation statistics, gate saturation and attention entropy. Scaling-law fits across several small sizes help predict whether a gain persists at scale. Theory (universal approximation, locality, equivariance) motivates architectures, but many details are validated empirically.
What is double descent?
Classical theory predicts a U-shaped test-error curve as model size grows. In deep learning, test error often rises near the interpolation threshold (where the model can just fit the training data exactly) and then falls again as models become much larger: a second descent. It also appears in epochs and dataset size. It helps explain why heavily over-parameterized networks, with implicit regularization from SGD and explicit regularization, generalize well.
What is knowledge distillation?
Training a small student model to match the softened output distribution of a large teacher, using a KL-divergence loss on softmax outputs with temperature T > 1 (often combined with the usual label loss). Soft targets carry "dark knowledge" about class similarities (a "3" that looks like an "8"), so students learn more per example than from hard labels. It is widely used to compress models for deployment.
How does label smoothing affect training and calibration?
PyTorch and the original paper mix the one-hot target with the uniform distribution: y' = (1 − ε)y + ε/K, so the true class gets (1 − ε) + ε/K and every other class gets ε/K (ε ≈ 0.1). A close variant uses (1 − ε) and ε/(K − 1) instead. Either way the optimal logits stay finite, which prevents unbounded logit growth and over-confidence and usually improves calibration and generalization. Side effects: it can hurt knowledge distillation (teacher outputs lose information about class similarity) and slightly compresses the representation clusters.
Scenario & debugging
Your training loss suddenly becomes NaN. What do you do?
- Find the step: log loss and gradient norm; look for a spike just before the NaN.
- Check the data:
torch.isfiniteon inputs and labels of the offending batch; look for division by zero in preprocessing, empty images, extreme values. - Lower the learning rate and add warmup; enable gradient clipping (max norm 1.0).
- Use stable losses (
CrossEntropyLoss,BCEWithLogitsLoss); add ε to logs, square roots and divisions in custom code. - In FP16, use a GradScaler or switch to BF16.
- Use
torch.autograd.set_detect_anomaly(True)to locate the operation producing NaN.
Training loss keeps decreasing, but validation loss starts rising after epoch 8. What is happening and what do you try?
Classic overfitting (unless the validation set comes from a different distribution). Immediate action: early stopping at the best validation epoch. Then: data augmentation, weight decay, dropout, label smoothing, a smaller or pretrained model, more data, and checking for train/validation mismatch or duplicates. If validation accuracy is still improving while its loss rises, the model is becoming over-confident on the errors it makes; calibration or label smoothing helps.
The loss does not decrease at all from the first epoch. How do you debug?
- Check the initial loss is about ln(K); verify inputs are paired with the right labels by visualizing a batch.
- Try to overfit 16 examples with regularization off; if it fails, it is a bug.
- Confirm gradients are non-zero for every layer (frozen parameters,
.detach(), integer tensors, missingrequires_grad, wrong parameters passed to the optimizer). - Run an LR range test; the rate may be far too low or so high that units died.
- Check that
optimizer.zero_grad(),loss.backward()andoptimizer.step()are all called, in that order. - Check input normalization and loss/activation pairing (no double softmax).
Your RNN's loss stays around 0.69 for all epochs on a binary task with long sequences. What is wrong?
0.693 = ln 2, the loss of guessing 50/50: the model learns nothing. Likely vanishing (or exploding) gradients over hundreds of time steps in a vanilla RNN. Fixes: gradient clipping, switch to LSTM/GRU, truncate or chunk sequences, pack padded sequences so padding is ignored, use pretrained embeddings, tune the learning rate, and consider attention or a 1D-CNN. Also verify labels are not shuffled relative to inputs.
Accuracy is 98% on a fraud dataset, yet the business says the model is useless. Why?
If 98% of transactions are legitimate, predicting "not fraud" for everything scores 98%. Accuracy is misleading under imbalance. Evaluate precision, recall, F1 and PR-AUC for the fraud class, and look at the confusion matrix. Fix with class-weighted or focal loss, resampling, threshold tuning for the business's precision/recall trade-off, and more fraud examples or features.
You have only 800 labelled medical images. How do you build a classifier?
Start from a pretrained backbone (ImageNet or, better, a medical-domain model); replace the head; train the head with the backbone frozen, then unfreeze later blocks with a small learning rate. Use strong but label-preserving augmentation, class weights, early stopping and cross-validation. Split by patient to avoid leakage. Consider self-supervised pretraining on unlabelled scans, and report sensitivity and specificity rather than accuracy. Be careful with pooling that could discard tiny lesions; use higher input resolution if needed.
A model trained on one GPU was moved to 8 GPUs with DDP and now converges worse. Why?
The effective batch size grew 8x, so the same learning rate gives far fewer, relatively smaller updates per epoch, and the warmup and schedule lengths are now wrong in steps. Scale the learning rate (linear for SGD, more conservative for Adam), rescale warmup and total steps, and check each process uses a DistributedSampler with set_epoch so data is not duplicated. With BatchNorm, per-GPU batches may be small; consider SyncBatchNorm.
Validation loss is consistently lower than training loss. Is something wrong?
Often not. Dropout and augmentation are active only during training, making training loss higher; training loss is averaged over the epoch while the weights improved, whereas validation is measured at the end; label smoothing inflates training loss. Investigate if the gap is large: the validation set may be easier, or there may be leakage (duplicates between splits, features derived from labels).
Predictions change depending on which other examples are in the batch at inference time. What is the bug?
The model is in training mode at inference, so BatchNorm uses the current batch's statistics (and dropout is random). Call model.eval() before inference and wrap it in torch.no_grad(). If the problem persists, check for other batch-dependent operations such as normalizing inputs with per-batch statistics.
GPU memory runs out during training. What are your options, in order?
- Make sure you log
loss.item(), not the loss tensor, and usetorch.no_grad()in evaluation (a common leak). - Reduce batch size and use gradient accumulation to keep the effective batch.
- Enable mixed precision (BF16/FP16).
- Activation checkpointing.
- Shorter sequences or smaller images; a smaller model.
- Memory-efficient optimizers (8-bit Adam) or sharding (FSDP/ZeRO), CPU offloading.
- For fine-tuning large models: LoRA or QLoRA.
Training is slow and GPU utilization shows about 30%. How do you speed it up?
The GPU is starved, usually by the input pipeline. Increase DataLoader workers, enable pin_memory and non_blocking copies, pre-process or cache data (decoded images, tokenized text), avoid CPU-GPU synchronizations in the loop (.item() every step, printing tensors), increase batch size, use mixed precision and torch.compile, and profile with the PyTorch profiler to find the bottleneck.
Loss curves look fine, but the model performs badly in production. What could be going on?
Distribution shift or pipeline mismatch: different preprocessing (normalization, resizing, tokenization) between training and serving, data leakage that inflated offline metrics, a validation set that does not represent production, or drift after deployment. Verify preprocessing parity with a golden test set, evaluate on recent production samples, monitor input feature statistics and prediction distributions, compare predictions with delayed ground truth, and set up retraining triggers.
A deployed model's accuracy degrades over months. How do you detect and address it?
Monitor continuously: compare predictions with actual outcomes when labels arrive, track metrics over time, and watch for data drift (input distribution changes, measured with statistics such as population stability index or KL divergence) and concept drift (the relationship between inputs and labels changes). Address with scheduled or triggered retraining on recent data, human review of low-confidence cases, and shadow deployment of new versions before switching traffic.
Your CNN misses very small objects (such as tiny tumours or distant pedestrians). What would you change?
Aggressive pooling and striding discard small details. Use higher input resolution, fewer or later downsampling steps, dilated convolutions, multi-scale features (feature pyramid networks, U-Net skip connections), tiling large images, anchor/scale settings suited to small objects in detectors, focal loss for the heavy background imbalance, and augmentation that preserves small objects.
After adding more layers to a plain CNN, even training accuracy got worse. Why, and what is the fix?
This is the degradation problem: an optimization difficulty in deep plain networks (vanishing gradients, poor conditioning), not overfitting, since training accuracy dropped too. Add residual connections, BatchNorm, He initialization, and a suitable learning-rate schedule with warmup.
Loss oscillates wildly and never settles. What do you try?
Lower the learning rate or add a decay schedule; increase batch size to reduce gradient noise; check that data is shuffled (batches sorted by class cause oscillation); add gradient clipping; check for label noise or corrupted examples; add momentum or switch to Adam if using plain SGD.
A large fraction of ReLU units output zero for every input. What happened and how do you fix it?
Dying ReLUs, usually caused by a too-high learning rate or a large update pushing biases negative, or poor initialization. Lower the learning rate, use He initialization, add BatchNorm before activations, switch to Leaky ReLU, GELU or ELU, and monitor the dead fraction per layer during training.
Fine-tuning a pretrained model made it worse than the frozen baseline. Why?
Most often the learning rate was too high, destroying pretrained features (catastrophic forgetting), or a random new head sent large gradients into the backbone. Train the head first with the backbone frozen, then unfreeze gradually with a 10-100x smaller learning rate, use warmup, keep BatchNorm frozen with small batches, and verify you used the same preprocessing and normalization as pretraining.
The model is confident (0.99) on inputs that are nothing like the training data. How do you handle this?
Softmax probabilities are not a reliable uncertainty measure, especially out of distribution. Options: temperature scaling for in-distribution calibration; out-of-distribution detection (energy scores, maximum logit, feature-space distance such as Mahalanobis); ensembles or Monte Carlo dropout for uncertainty; an explicit "unknown" class trained with outlier examples; and abstaining or routing to a human below a confidence threshold.
Your GAN keeps generating nearly identical images. What is it and how do you fix it?
Mode collapse: the generator found a few outputs that fool the discriminator. Try a Wasserstein loss with gradient penalty, spectral normalization, minibatch-discrimination or diversity terms, balanced learning rates for generator and discriminator (TTUR), label smoothing and instance noise, and a larger batch. Monitor diversity with FID and sample grids rather than losses. If stability remains poor, a diffusion model may be a better fit.
Training loss drops to nearly zero on a BiLSTM, but test F1 is much worse than a simple logistic regression. What do you conclude?
The BiLSTM overfits (and may be over-predicting one class), and the task may be well served by bag-of-words features. Add dropout (including on embeddings), weight decay, early stopping, smaller hidden sizes, pretrained frozen embeddings, and check the precision/recall balance and threshold. Always keep the classical baseline: if it is equal or better, prefer it for cost and interpretability, or move to a pretrained Transformer fine-tuned with care.
A teammate reports great test accuracy after tuning hyperparameters for many rounds on the test set. What is the problem?
Test-set leakage through model selection: the reported number is optimistically biased because the test set was used to make decisions. Hyperparameters must be tuned on a validation set (or with cross-validation), and the test set evaluated once at the end. Recover by collecting a fresh test set or using a nested cross-validation estimate.
You need to deploy an image classifier on a phone with a 20 ms latency budget. What do you change?
Choose an efficient architecture (MobileNet or EfficientNet-Lite family, depthwise separable convolutions), reduce input resolution, distil from a larger teacher, quantize to INT8 (post-training or quantization-aware training), prune if helpful, export to a mobile runtime and run on the NPU/GPU, and benchmark on the actual target devices, including thermal throttling. See Edge AI for the full pipeline.
You must choose between a CNN, an RNN and a Transformer for a new problem. How do you decide?
Match the inductive bias to the data and constraints: CNNs for grid data with local patterns (images, spectrograms) and efficiency; RNNs/GRUs for streaming sequences with tight memory and latency limits; Transformers when long-range dependencies matter, data and compute are plentiful, or strong pretrained models exist (text, increasingly vision and audio). For tabular data, start with gradient-boosted trees. Consider pretrained availability, data size, latency and deployment target, and always benchmark against a simple baseline.
Your validation metric jumps around a lot between epochs. How do you make decisions reliably?
The validation set is probably too small or training too noisy. Enlarge the validation set, use cross-validation or multiple seeds and report the mean and standard deviation, smooth the metric (exponential moving average of weights also stabilizes it), lower the learning rate late in training, and base early stopping on a smoothed metric with adequate patience.
Your model works on your laptop but gives different results each run. How do you make experiments reproducible?
Fix seeds for Python, NumPy and PyTorch (torch.manual_seed, including CUDA), seed the DataLoader workers and data splits, enable deterministic algorithms (torch.use_deterministic_algorithms(True), disable cuDNN benchmark mode), pin library versions, and log configurations and data versions. Accept that some GPU operations remain non-deterministic; compare results across several seeds instead of one run.
The gradient norm periodically spikes by 100x, followed by loss spikes. What do you do?
Inspect the batches at the spikes for bad or extreme examples (corrupted files, mislabeled data, extremely long sequences). Enable or tighten gradient clipping, lower the peak learning rate or lengthen warmup, check Adam's ε and β2 (a lower β2 such as 0.95 is common for large models), and, for large models, skip updates when the norm is anomalous and restart from a checkpoint before the spike.
Training accuracy is stuck at exactly the majority-class rate. What is happening?
The model predicts one class for everything. Causes: severe class imbalance with an unweighted loss, a learning rate too high (units died) or too low, a bug where the output does not depend on the input (features zeroed, wrong tensor fed), or labels misaligned with inputs. Check that predictions vary across inputs, overfit a small balanced batch, use class weights or focal loss, and inspect the input pipeline.