Python & Data Tools for ML
The practical toolkit behind almost every machine-learning project: the Python language features you will read and write daily, the NumPy and pandas data stack, plotting and exploratory analysis, scikit-learn for classical models, and PyTorch for neural networks. Read top to bottom and you will be able to take a raw CSV to an evaluated, saved model, and defend every line of it in an interview.
- Python is the glue; the heavy maths runs in compiled C, Fortran or CUDA code underneath NumPy, pandas, scikit-learn and PyTorch. Your job is to keep work inside those fast libraries and out of Python loops.
- NumPy gives you typed n-dimensional arrays, vectorized operations and broadcasting. Knowing shapes, axes and views versus copies prevents most bugs.
- pandas adds labelled tables: select with
loc/iloc, summarise withgroupby, combine withmerge, and watch dtypes and memory. - scikit-learn has one API for everything:
fit,transform,predict. Wrap preprocessing and model in aPipelineso cross-validation never leaks test information. - PyTorch is NumPy on GPUs plus automatic gradients. Every training loop is zero_grad, forward, loss, backward, step; switch to
model.eval()andno_gradfor inference. - Interviewers probe mutability, generators, decorators, the GIL, broadcasting, groupby, data leakage, metric choice, and debugging stories like out-of-memory errors or NaN losses.
The big picture: why Python runs ML
Python itself is a slow, interpreted, dynamically typed language. It became the language of machine learning anyway because it is easy to read and because its scientific libraries push the expensive work into compiled code. When you write X @ W in NumPy, one Python instruction triggers millions of multiply-adds inside an optimized linear-algebra library (BLAS). When you call model(x) in PyTorch on a GPU, Python only launches kernels; the GPU does the arithmetic.
Think of a film director and a studio crew. The director (Python) is not the fastest at building sets or lighting scenes, but gives short, clear instructions. The crew (C, Fortran and CUDA libraries) does the physical heavy lifting at great speed. A good director says "light the whole stage" once instead of pointing at each bulb in turn. In code terms: one vectorized call on a whole array beats a Python loop over each element, because the loop makes the slow director do the crew's job.
The layer cake
+------------------------------------------------------------------+ | Your notebook / script / service (Python: glue, control flow) | +------------------------------------------------------------------+ | scikit-learn PyTorch Hugging Face LLM SDKs | modelling +------------------------------------------------------------------+ | pandas (tables) matplotlib / seaborn (plots) SciPy | data + stats +------------------------------------------------------------------+ | NumPy ndarray (typed, contiguous memory, vectorized ufuncs) | array core +------------------------------------------------------------------+ | C / C++ / Fortran / CUDA : BLAS, LAPACK, cuBLAS, cuDNN, Arrow | speed lives here +------------------------------------------------------------------+
The standard ML workflow in Python
- Load Read CSV, Parquet, SQL or JSON into a pandas DataFrame.
- Explore (EDA) Check shape, dtypes, missing values, distributions, class balance, correlations and leakage risks.
- Split Hold out a test set first (stratified for classification) so it stays untouched.
- Preprocess Impute, encode categoricals, scale numerics, engineer features, all inside a pipeline fitted on training data only.
- Train and tune Fit models, compare with cross-validation, tune hyperparameters with a search.
- Evaluate Use metrics that match the business cost (recall for missed cancers, precision for spam, RMSE for prices), inspect errors, tune the decision threshold on validation data.
- Ship Save the whole pipeline, version data and code, monitor in production.
NumPy
Fast typed arrays and maths. The data format every other library speaks.
pandas
Labelled tables for loading, cleaning, joining and summarising data.
matplotlib / seaborn
Plots for exploration and reporting.
scikit-learn
Classical ML: preprocessing, models, validation and metrics with one consistent API.
PyTorch
Tensors on GPU plus automatic differentiation for deep learning.
Hugging Face
Pretrained transformer models and tokenizers in a few lines.
for i in range(len(...)) over data, ask whether a NumPy, pandas or PyTorch operation can express the same thing on the whole array.Environments, packaging and Jupyter
Every project should live in its own isolated environment so that one project's numpy 1.26 cannot break another project's numpy 2.x. An environment is simply a folder containing a Python interpreter plus the packages installed for it.
A virtual environment is like a separate toolbox per job site. The plumber's box and the electrician's box each hold the exact tool versions that job needs, so borrowing a wrench for one job never removes it from the other. Mapping back: each project gets its own .venv folder with its own package versions, and a lock file (requirements.txt) is the inventory list that lets anyone rebuild the same toolbox.
venv + pip (built in)
# create and activate an isolated environment
python -m venv .venv
.venv\Scripts\activate # Windows PowerShell / cmd
source .venv/bin/activate # macOS / Linux
python -m pip install --upgrade pip
pip install numpy pandas scikit-learn matplotlib seaborn torch
pip freeze > requirements.txt # pin exact versions
pip install -r requirements.txt # recreate elsewhere
deactivate
The options compared
| Tool | What it is | When to use |
|---|---|---|
venv + pip | Standard library environments plus the default installer | Default choice for pure-Python projects |
conda / mamba | Environment and package manager that also installs non-Python binaries (CUDA, MKL, compilers) | Heavy scientific stacks, GPU toolkits, mixed-language dependencies |
uv, poetry, pip-tools | Faster resolvers and lock-file managers | Reproducible builds for teams and CI |
| Docker | Whole operating-system image with Python inside | Deployment, exact reproducibility including system libraries |
Packaging your own code
Modern projects describe themselves in a pyproject.toml file (name, version, dependencies, build backend). A folder with an __init__.py is a package; any .py file is a module. Installing your project in editable mode (pip install -e .) lets notebooks import your code while you keep editing it.
# pyproject.toml (minimal)
[project]
name = "churn_model"
version = "0.1.0"
requires-python = ">=3.10"
dependencies = ["pandas>=2.0", "scikit-learn>=1.4"]
[build-system]
requires = ["setuptools>=68"]
build-backend = "setuptools.build_meta"
The guard if __name__ == "__main__": makes a file behave as both an importable module and a runnable script: the block runs only when the file is executed directly, not when it is imported. It is also required on Windows for multiprocessing and PyTorch DataLoader workers, because child processes re-import the main file.
Jupyter essentials
A notebook talks to a kernel, a live Python process that keeps variables in memory between cells. That is convenient for exploration and dangerous for reproducibility, because the order you ran cells in is invisible to a reader.
%pip install seaborn # installs into the kernel's own environment
%timeit np.sort(arr) # micro-benchmark a line
%%time # time a whole cell (must be the first line)
%load_ext autoreload
%autoreload 2 # re-import edited .py modules automatically
%matplotlib inline # render plots below cells
df.head() # the last expression in a cell is displayed
%pip install over !pip install, since the shell version may install into a different interpreter than the running kernel.Core Python: types, mutability and containers
In Python, every value is an object with a type, and variables are just names bound to objects. Assignment never copies data; it attaches another name to the same object. Whether that shared object can be changed in place (mutability) explains a large share of Python bugs and interview questions.
Variables are sticky notes, not boxes. Writing b = a sticks a second note on the same whiteboard. If the whiteboard is erasable (a list or dict), writing on it through either note changes what both notes point to. If it is laminated (an int, str or tuple), "changing" it really means making a new laminated card and moving one sticky note to it. Mapping back: mutable objects are shared and changed in place; immutable objects are replaced, so other names are unaffected.
Built-in types at a glance
| Type | Example | Mutable? | Typical ML use |
|---|---|---|---|
int, float, complex | 10, 3e-4 | No | Hyperparameters, counters |
bool | True | No | Flags; a subclass of int (True + True == 2) |
str | "relu" | No | Text data, config keys |
bytes | b"\x00" | No | Raw files, network payloads |
tuple | (32, 3, 224, 224) | No | Shapes, fixed records, dict keys |
list | [0.9, 0.8] | Yes | Ordered collections, loss history |
dict | {"lr": 1e-3} | Yes | Configs, JSON, lookups (insertion-ordered since 3.7) |
set / frozenset | {"cat", "dog"} | Yes / No | Vocabularies, de-duplication, fast membership |
NoneType | None | No | "No value" sentinel |
Names, identity and equality
a = [1, 2]
b = a # same object, two names
b.append(3)
print(a) # [1, 2, 3] -- a changed too
print(a is b) # True (same identity)
print(a == [1, 2, 3], a is [1, 2, 3]) # True False (equal value, different object)
x = 10
y = x
y += 1 # ints are immutable: y now names a NEW object
print(x, y) # 10 11
== compares values (calls __eq__); is compares identity (same object in memory). Use is only for singletons: if x is None. CPython caches small integers (-5 to 256) and some strings, so is may accidentally appear to work on them; never rely on it.
Shallow versus deep copies
import copy
grid = [[0, 0], [0, 0]]
shallow = grid.copy() # new outer list, SAME inner lists
deep = copy.deepcopy(grid) # everything duplicated
grid[0][0] = 9
print(shallow[0][0], deep[0][0]) # 9 0
bad = [[0] * 2] * 2 # two references to ONE inner list
bad[0][0] = 1
print(bad) # [[1, 0], [1, 0]]
good = [[0] * 2 for _ in range(2)]
Container cost cheat sheet
| Operation | list | dict / set | collections.deque |
|---|---|---|---|
| Append at end | O(1) amortized | O(1) average insert | O(1) |
| Insert / pop at front | O(n) | n/a | O(1) |
| Index by position | O(1) | n/a | O(n) in the middle |
Membership x in c | O(n) | O(1) average | O(n) |
| Sort | O(n log n), stable | n/a | n/a |
Dict keys and set members must be hashable: immutable and with a stable __hash__. That is why a tuple can be a dict key but a list cannot. See Data Structures & Algorithms for how hash tables achieve O(1) average lookups.
Strings and f-strings
name, lr, acc = "resnet", 3e-4, 0.91234
print(f"{name:>8} lr={lr:.1e} acc={acc:.2%}") # resnet lr=3.0e-04 acc=91.23%
print(f"{acc=:.3f}") # acc=0.912 (self-documenting, 3.8+)
words = "the cat sat".split() # ['the', 'cat', 'sat']
print(" ".join(words).upper()) # THE CAT SAT
# strings are immutable: building with += in a loop is O(n^2); collect then join
Truthiness and handy built-ins
Empty containers, 0, 0.0, "" and None are falsy; everything else is truthy. But note that NumPy arrays and pandas Series refuse to be used in an if directly (they raise "truth value is ambiguous"); use .any(), .all() or .empty.
from collections import Counter, defaultdict
labels = ["cat", "dog", "cat", "bird", "cat"]
print(Counter(labels).most_common(2)) # [('cat', 3), ('dog', 1)]
by_len = defaultdict(list)
for w in labels:
by_len[len(w)].append(w)
print(dict(by_len)) # {3: ['cat', 'dog', 'cat', 'cat'], 4: ['bird']}
for i, (w, n) in enumerate(zip(["a", "b"], [1, 2])):
print(i, w, n) # 0 a 1 / 1 b 2
print(sorted(labels, key=len, reverse=True)[:2]) # ['bird', 'cat']
Floating point
print(0.1 + 0.2 == 0.3) # False: binary floats cannot represent 0.1 exactly
import math
print(math.isclose(0.1 + 0.2, 0.3)) # True -- compare with a tolerance (np.isclose for arrays)
print(float("nan") == float("nan")) # False: NaN is not equal to anything, even itself
def f(x, acc=[]) creates the list once, when the function is defined, so every call shares it: f(1) then f(2) returns [1, 2]. Use acc=None and create a new list inside. The same trap appears with dict defaults in config helpers.is versus ==, shallow versus deep copy, and the output of the [[0]*2]*2 puzzle. Strong candidates explain the "names bound to objects" model instead of memorising individual cases.Functions, comprehensions, closures and decorators
Functions in Python are ordinary objects: you can pass them as arguments, return them from other functions, and store them in dicts. That single fact enables callbacks (learning-rate schedulers), closures (factories that remember settings) and decorators (wrappers that add timing, caching or retries).
A decorator is gift wrapping. The present inside (your function) is unchanged, but the wrapping adds something on the outside: a label with the time it took, a guarantee to retry if it breaks, or a cache that remembers previous answers. A closure is a backpack a function carries: variables from the place it was created travel with it. Mapping back: @timer returns a new function that calls the original and adds behaviour around it, and that wrapper remembers the original through a closure.
Arguments: positional, keyword, defaults, *args and **kwargs
def build_mlp(in_dim, *hidden, activation="relu", dropout=0.0, **extra):
print(in_dim, hidden, activation, dropout, extra)
build_mlp(30, 64, 32, dropout=0.1, bias=False)
# 30 (64, 32) relu 0.1 {'bias': False}
def scale(x, /, factor=2, *, clip=None): # x positional-only, clip keyword-only
y = x * factor
return min(y, clip) if clip is not None else y
print(scale(5, clip=8)) # 8
params = {"activation": "gelu", "dropout": 0.2}
build_mlp(10, 16, **params) # unpack a dict into keyword arguments
shape = (2, 3)
rows, cols = shape # tuple unpacking
first, *rest = [1, 2, 3, 4] # first=1, rest=[2, 3, 4]
*args collects extra positional arguments into a tuple; **kwargs collects extra keyword arguments into a dict. At a call site, * and ** do the reverse: they spread a sequence or dict into arguments. Library code uses this heavily to forward options, for example model(**tokenizer_output) in Hugging Face.
Comprehensions
nums = range(10)
squares = [n * n for n in nums if n % 2 == 0] # [0, 4, 16, 36, 64]
tok2id = {t: i for i, t in enumerate(["the", "cat"])} # {'the': 0, 'cat': 1}
id2tok = {i: t for t, i in tok2id.items()} # invert a mapping
uniq_lens = {len(w) for w in ["a", "bb", "cc"]} # {1, 2}
total = sum(n * n for n in nums) # generator expression: no list built
matrix = [[r * c for c in range(3)] for r in range(2)] # [[0, 0, 0], [0, 1, 2]]
flat = [x for row in matrix for x in row] # loops read left to right
Comprehensions are faster than equivalent append loops (less interpreter overhead) and clearer when short. If a comprehension needs nested conditions or side effects, use a normal loop. For numeric work on large data, neither beats a NumPy vectorized operation.
Lambda, map, filter, sorted with key
runs = [{"name": "a", "f1": 0.81}, {"name": "b", "f1": 0.87}]
best = max(runs, key=lambda r: r["f1"]) # {'name': 'b', 'f1': 0.87}
ranked = sorted(runs, key=lambda r: -r["f1"])
lengths = list(map(len, ["hi", "there"])) # [2, 5]
long_words = list(filter(lambda w: len(w) > 2, ["a", "the", "cat"])) # ['the', 'cat']
Scope (LEGB) and closures
Python resolves a name by searching Local, then Enclosing function, then Global (module), then Built-in scope. Assigning to a name inside a function makes it local unless declared global or nonlocal.
def make_scheduler(base_lr, decay):
step = 0
def next_lr():
nonlocal step # modify the enclosing variable
lr = base_lr * decay ** step
step += 1
return lr
return next_lr # the inner function "closes over" base_lr, decay, step
sched = make_scheduler(0.1, 0.5)
print(sched(), sched(), sched()) # 0.1 0.05 0.025
fns = [lambda: i for i in range(3)] then fns[0]() returns 2, not 0, because every lambda looks up i when called, after the loop finished. Fix it by binding at definition time: lambda i=i: i.Decorators
import time, functools
def timer(func):
@functools.wraps(func) # keep the original name and docstring
def wrapper(*args, **kwargs):
start = time.perf_counter()
result = func(*args, **kwargs)
print(f"{func.__name__} took {time.perf_counter() - start:.3f}s")
return result
return wrapper
def retry(times=3, exceptions=(ConnectionError,)): # decorator factory: takes arguments
def decorator(func):
@functools.wraps(func)
def wrapper(*args, **kwargs):
for attempt in range(1, times + 1):
try:
return func(*args, **kwargs)
except exceptions:
if attempt == times:
raise
time.sleep(2 ** attempt) # exponential backoff
return wrapper
return decorator
@timer
def train_epoch():
time.sleep(0.2)
train_epoch() # train_epoch took 0.200s
@functools.lru_cache(maxsize=None) # memoize pure functions
def fib(n):
return n if n < 2 else fib(n - 1) + fib(n - 2)
print(fib(80)) # 23416728348467685, instantly
@timer above a definition is shorthand for train_epoch = timer(train_epoch). Built-in decorators you will meet constantly: @property, @staticmethod, @classmethod, @dataclass, @functools.lru_cache, @torch.no_grad(), and @app.get(...) in web frameworks.
*args, **kwargs forwarding, returning the result, functools.wraps, and whether you can extend it to a decorator that takes arguments (three nested levels). Follow-ups: what is a closure, what does nonlocal do, and why does lru_cache require hashable arguments.Iterators and generators: streaming data lazily
An iterable is anything you can loop over (list, dict, file, range). An iterator is the object that actually produces the values one at a time and remembers where it is. A generator is the easiest way to write an iterator: a function containing yield. Generators compute values on demand, so they can stream a 50 GB log file or an infinite data source using almost no memory. This is exactly how data loaders feed mini-batches to a model.
A list is a printed book: every page exists up front and you can flip to page 300. A generator is a storyteller: each time you say "next", they tell you one more sentence and then pause, remembering exactly where they were. They never need the whole story written down. Mapping back: next() resumes the generator function until its next yield, the paused frame keeps its local variables, and when the story ends Python raises StopIteration, which a for loop handles silently.
The iterator protocol
nums = [10, 20]
it = iter(nums) # calls nums.__iter__()
print(next(it)) # 10 (calls it.__next__())
print(next(it)) # 20
# next(it) now raises StopIteration -- a for loop catches this for you
class Countdown:
def __init__(self, start):
self.current = start
def __iter__(self):
return self
def __next__(self):
if self.current <= 0:
raise StopIteration
self.current -= 1
return self.current + 1
print(list(Countdown(3))) # [3, 2, 1]
Generators
def countdown(start):
while start > 0:
yield start # pause here, hand out a value
start -= 1
gen = countdown(3)
print(next(gen), list(gen)) # 3 [2, 1]
print(list(gen)) # [] -- a generator is exhausted after one pass
def batches(items, size):
"""Yield successive mini-batches without building them all at once."""
for i in range(0, len(items), size):
yield items[i:i + size]
print(list(batches(list(range(7)), 3))) # [[0, 1, 2], [3, 4, 5], [6]]
def read_jsonl(path):
import json
with open(path, encoding="utf-8") as f:
for line in f: # files are lazy iterators over lines
yield json.loads(line)
# pipelines of generators: nothing runs until something consumes them
# valid = (r for r in read_jsonl("events.jsonl") if r.get("label") is not None)
Generator expressions and memory
import sys
squares_list = [n * n for n in range(1_000_000)]
squares_gen = (n * n for n in range(1_000_000))
print(sys.getsizeof(squares_list) > 8_000_000) # True: about 8 MB of pointers
print(sys.getsizeof(squares_gen) < 500) # True: a few hundred bytes, whatever n is
itertools and yield from
import itertools as it
print(list(it.islice(it.count(0, 5), 4))) # [0, 5, 10, 15]
print(list(it.chain([1, 2], [3]))) # [1, 2, 3]
grid = list(it.product([0.01, 0.1], ["relu", "gelu"])) # hyperparameter grid
print(len(grid)) # 4
print(list(it.accumulate([1, 2, 3, 4]))) # [1, 3, 6, 10] running sum
print(list(it.batched(range(5), 2))) # [(0, 1), (2, 3), (4,)] (3.12+)
def flatten(nested):
for item in nested:
if isinstance(item, list):
yield from flatten(item) # delegate to a sub-generator
else:
yield item
print(list(flatten([1, [2, [3, 4]], 5]))) # [1, 2, 3, 4, 5]
List / list comprehension
- All values computed and stored immediately
- Random access,
len(), reusable many times - Memory grows with n
Generator / generator expression
- Values computed on demand
- Single pass, no
len(), no indexing - Constant memory; can be infinite
total = sum(gen); count = len(list(gen)) gives a count of zero because the first pass exhausted it. Either materialise it once into a list, or recreate the generator. The same bug appears when a PyTorch-style data iterator is reused without being re-created each epoch.yield do", and "how would you process a file larger than RAM". A good answer names the protocol (__iter__, __next__, StopIteration), explains that a generator's frame is suspended rather than returned, and proposes streaming with a generator or pandas chunksize.Classes, OOP and dunder methods
Object-oriented code bundles data (attributes) with behaviour (methods). You need it in ML because frameworks are built on it: every PyTorch model is a class that inherits from nn.Module, every custom dataset implements __len__ and __getitem__, and every scikit-learn estimator follows a class-based contract.
A class is a cookie cutter and objects are the cookies. The cutter defines the shape (attributes and methods); each cookie has its own decorations (instance values). Dunder ("double underscore") methods are the standard sockets on an appliance: if your device has the right plug shape, it works with every outlet in the house. Mapping back: implementing __len__ lets len(obj) work, __getitem__ lets obj[i] work, and __call__ lets obj(x) work, so your class plugs into Python syntax and libraries that expect those sockets.
A class from scratch
class Neuron:
activation = "relu" # class attribute: shared by all instances
def __init__(self, weights, bias=0.0): # initializer, runs on Neuron(...)
self.weights = list(weights) # instance attributes
self.bias = bias
def forward(self, inputs):
z = sum(w * x for w, x in zip(self.weights, inputs)) + self.bias
return max(0.0, z)
def __call__(self, inputs): # makes instances callable like functions
return self.forward(inputs)
def __repr__(self): # unambiguous developer view
return f"Neuron(weights={self.weights}, bias={self.bias})"
n = Neuron([0.5, -0.2], bias=0.1)
print(n) # Neuron(weights=[0.5, -0.2], bias=0.1)
print(n([1.0, 2.0])) # 0.19999999999999998 (0.2 up to float rounding)
self is simply the instance, passed automatically as the first argument when you call n.forward(...). __init__ is not a constructor in the strict sense (that is __new__); it initialises an already-created object.
Inheritance and super()
class Model:
def __init__(self, name):
self.name = name
def predict(self, x):
raise NotImplementedError
class Constant(Model):
def __init__(self, name, value):
super().__init__(name) # run the parent's initializer
self.value = value
def predict(self, x): # override (polymorphism)
return [self.value] * len(x)
print(Constant("baseline", 1).predict([7, 8, 9])) # [1, 1, 1]
print(isinstance(Constant("b", 0), Model)) # True
With multiple inheritance, Python follows the method resolution order (MRO, the C3 linearization, viewable via Cls.__mro__); super() means "next class in the MRO", not simply "my parent".
The dunder methods worth knowing
| Method | Triggered by | ML example |
|---|---|---|
__init__ | Cls(...) | Define layers in nn.Module |
__repr__ / __str__ | repr(x), printing / str(x) | Readable model summaries |
__len__ | len(x) | Dataset size for a DataLoader |
__getitem__ | x[i], slicing, iteration fallback | Return one (features, label) sample |
__iter__ / __next__ | for loops | Streaming datasets |
__call__ | x(...) | model(x) runs hooks and then forward |
__eq__ / __hash__ | ==, dict keys, sets | Caching configs |
__add__, __matmul__ | +, @ | How tensors overload operators |
__enter__ / __exit__ | with blocks | torch.no_grad(), file handles |
__contains__ | in | Vocabulary membership |
class Vocab:
def __init__(self, tokens):
self.itos = sorted(set(tokens))
self.stoi = {t: i for i, t in enumerate(self.itos)}
def __len__(self):
return len(self.itos)
def __getitem__(self, token):
return self.stoi.get(token, -1) # -1 for unknown tokens
def __contains__(self, token):
return token in self.stoi
v = Vocab("the cat sat on the mat".split())
print(len(v), v["cat"], v["dog"], "mat" in v) # 5 0 -1 True
property, classmethod, staticmethod
class Temperature:
def __init__(self, celsius):
self.celsius = celsius
@property
def fahrenheit(self): # computed attribute, accessed without ()
return self.celsius * 9 / 5 + 32
@classmethod
def from_fahrenheit(cls, f): # alternative constructor, receives the class
return cls((f - 32) * 5 / 9)
@staticmethod
def is_valid(c): # plain function namespaced in the class
return c >= -273.15
t = Temperature.from_fahrenheit(212)
print(t.celsius, t.fahrenheit, Temperature.is_valid(-300)) # 100.0 212.0 False
Python has no truly private attributes. By convention _name means "internal, do not touch"; __name triggers name mangling to _ClassName__name to avoid clashes in subclasses.
class Run: history = [] shares one list across every instance, so appending to run1.history also shows up in run2.history. Put per-instance state in __init__ as self.history = []. In PyTorch, forgetting super().__init__() in an nn.Module subclass causes an error as soon as you assign a layer, because the module's registries do not exist yet.self; __str__ versus __repr__; class versus instance attributes; classmethod versus staticmethod; what the MRO is; why you call model(x) and not model.forward(x) in PyTorch (because __call__ runs registered hooks around forward). Explaining composition over deep inheritance is a plus.Robust Python: typing, dataclasses, exceptions and context managers
Research code often starts as a notebook, but production ML code must be readable, checkable and safe when things fail. Four features do most of that work: type hints (document and statically check intent), dataclasses (tidy configuration objects), exceptions (explicit failure handling) and context managers (guaranteed cleanup).
Type hints are labels on the shelves of a warehouse: nothing stops you putting soup on the "screws" shelf, but inspectors (a type checker such as mypy) will flag it before customers notice. A context manager is a hotel key card that automatically deactivates at checkout, no matter whether you leave calmly or run out during a fire alarm. Mapping back: hints are not enforced at runtime but are checked by tools, and a with block guarantees its cleanup code (__exit__) runs even when an exception is raised inside.
Type hints
from typing import Optional, Union, Literal, Callable, Iterator, TypedDict, Protocol
import numpy as np
def top_k(scores: list[float], k: int = 5) -> list[int]:
return sorted(range(len(scores)), key=scores.__getitem__, reverse=True)[:k]
def load(path: str, sep: Optional[str] = None) -> "pd.DataFrame": ...
Metric = Literal["accuracy", "f1", "recall"] # only these strings allowed
Scorer = Callable[[np.ndarray, np.ndarray], float] # a function type
class Sample(TypedDict): # typed shape for a dict (e.g. JSON records)
text: str
label: int
class SupportsPredict(Protocol): # structural typing: "anything with .predict"
def predict(self, X: np.ndarray) -> np.ndarray: ...
print(top_k([0.1, 0.9, 0.4], k=2)) # [1, 2]
Since Python 3.10 you can write int | None instead of Optional[int]. Hints are ignored at runtime by the interpreter; libraries such as Pydantic and FastAPI read them to validate data, which is why LLM structured-output code relies on them (see LLM APIs).
Dataclasses for configs and records
from dataclasses import dataclass, field, asdict
@dataclass(frozen=True) # frozen: fields cannot be reassigned
class TrainConfig:
lr: float = 3e-4
epochs: int = 10
layers: list[int] = field(default_factory=lambda: [64, 32]) # safe mutable default
seed: int = 42
cfg = TrainConfig(lr=1e-3)
print(cfg) # TrainConfig(lr=0.001, epochs=10, layers=[64, 32], seed=42)
print(asdict(cfg)["epochs"], cfg == TrainConfig(lr=1e-3)) # 10 True
# cfg.lr = 0.1 -> raises FrozenInstanceError
@dataclass generates __init__, __repr__ and __eq__ from annotated fields. Use slots=True (3.10+) to save memory for millions of small objects. Choose Pydantic instead when you need runtime validation of untrusted input such as API payloads or LLM output.
Exceptions
import json
class DataValidationError(ValueError):
"""Raised when an input file fails schema checks."""
def load_config(path):
try:
with open(path, encoding="utf-8") as f:
cfg = json.load(f)
except FileNotFoundError:
return {"lr": 1e-3} # sensible default
except json.JSONDecodeError as e:
raise DataValidationError(f"bad JSON in {path}") from e # keep the cause
else:
return cfg # runs only if no exception
finally:
print("config load attempted") # always runs
- Catch the most specific exception you can handle; never write a bare
except:, which also swallowsKeyboardInterrupt. raise ... from echains exceptions so the traceback shows the root cause.- Python style is EAFP ("easier to ask forgiveness than permission"): try the operation and catch failure, rather than checking every precondition first.
- Use
assertfor internal sanity checks (shapes during development) but not for validating user input, because asserts are removed when Python runs with-O.
Context managers
import time
from contextlib import contextmanager
@contextmanager
def timed(label):
start = time.perf_counter()
try:
yield # the body of the with-block runs here
finally:
print(f"{label}: {time.perf_counter() - start:.2f}s")
with timed("feature engineering"):
time.sleep(0.1) # feature engineering: 0.10s
class Timer: # the same idea, class-based
def __enter__(self):
self.start = time.perf_counter()
return self
def __exit__(self, exc_type, exc, tb):
self.elapsed = time.perf_counter() - self.start
return False # False: do not swallow exceptions
with Timer() as t:
sum(range(10**6))
print(t.elapsed < 1) # True
ML examples: with open(...), with torch.no_grad():, with torch.autocast("cuda"):, with pd.option_context("display.max_rows", 5):, with mlflow.start_run():.
except Exception: pass around a batch hides corrupted data and makes the model silently train on less data. Log the exception with context (batch index, file name), count failures, and fail loudly past a threshold.with statement work?" Expect to describe __enter__/__exit__ and that __exit__ receives the exception and can suppress it by returning True. Also common: else versus finally in try blocks, dataclass versus namedtuple versus Pydantic, and whether type hints are enforced (no, unless a tool or library checks them).Concurrency: the GIL, threads, processes and asyncio
CPython has a Global Interpreter Lock (GIL): only one thread executes Python bytecode at a time within a process. Threads therefore do not speed up pure-Python number crunching, but they do help when threads spend most of their time waiting (network, disk), because waiting threads release the GIL. Heavy NumPy, pandas and PyTorch operations also release the GIL while their C code runs. For CPU-bound pure-Python work, use multiple processes, each with its own interpreter and GIL.
Picture a kitchen with one chef's knife (the GIL). Several cooks (threads) can share the kitchen, but only the one holding the knife can chop. If most of the work is waiting for water to boil (I/O), sharing one knife is fine and the kitchen stays busy. If the work is all chopping (CPU), extra cooks just queue for the knife, so you need more kitchens (processes). asyncio is one very organised cook who starts every pot boiling and checks each when it is ready. Mapping back: threads for I/O-bound work, processes for CPU-bound Python work, and asyncio for thousands of concurrent network calls such as LLM API requests.
| Approach | Parallel CPU? | Best for | Costs |
|---|---|---|---|
threading / ThreadPoolExecutor | No for Python code (GIL), yes inside C code that releases it | Downloading files, database or API calls, reading many files | Race conditions; shared state needs locks |
multiprocessing / ProcessPoolExecutor | Yes | CPU-heavy Python: feature engineering, parsing, simulations | Process start-up, pickling data between processes, more memory |
asyncio | No (single thread) | Very many concurrent I/O tasks: LLM calls, web scraping, services | Needs async-aware libraries; one blocking call stalls everything |
| Vectorized libraries / GPU | Yes, internally | Numerical work on arrays | Must reshape the problem into array operations |
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor
import math, time
def fetch(url): # I/O-bound: mostly waiting
time.sleep(0.5)
return len(url)
def heavy(n): # CPU-bound pure Python
return sum(math.isqrt(i) for i in range(n))
if __name__ == "__main__": # required for processes on Windows and macOS
urls = [f"doc{i}" for i in range(8)]
with ThreadPoolExecutor(max_workers=8) as ex:
sizes = list(ex.map(fetch, urls)) # about 0.5 s total instead of 4 s
with ProcessPoolExecutor() as ex:
totals = list(ex.map(heavy, [2_000_000] * 4)) # uses several CPU cores
asyncio for many API calls
import asyncio
async def call_llm(prompt, sem):
async with sem: # cap concurrency (rate limits)
await asyncio.sleep(0.2) # stands in for an async HTTP call
return prompt.upper()
async def main(prompts):
sem = asyncio.Semaphore(5)
return await asyncio.gather(*(call_llm(p, sem) for p in prompts))
print(asyncio.run(main(["hi", "hello", "hey"]))) # ['HI', 'HELLO', 'HEY']
In Jupyter an event loop is already running, so use await main(prompts) directly in a cell instead of asyncio.run.
time.sleep or a synchronous HTTP client) inside an async def freezes the whole event loop, so you get no concurrency at all. Use the async client or push the blocking call to a thread with asyncio.to_thread. With multiprocessing, forgetting the __main__ guard on Windows spawns processes recursively.DataLoader(num_workers=4) uses worker processes to decode and augment data in parallel.NumPy: arrays, vectorization and broadcasting
A NumPy ndarray is a block of memory holding many values of one type, plus metadata describing its shape and how to step through memory (strides). Because every element has the same type and sits next to its neighbours, NumPy can hand the whole block to optimized C loops and CPU vector instructions. A Python list, by contrast, is an array of pointers to separate objects scattered around memory, each carrying its own type information.
A Python list is a shelf of assorted parcels, each with its own label that must be read before you can open it. A NumPy array is an egg carton: identical slots, identical contents, so a machine can process the whole carton in one sweep. Broadcasting is like a stamp that says "add one pinch of salt to every egg in the row": you describe the operation once for a smaller shape and NumPy repeats it across the larger one without making copies. Mapping back: homogeneous dtype and contiguous memory enable vectorized ufuncs, and broadcasting stretches size-1 dimensions virtually instead of physically duplicating data.
Creating arrays and reading their metadata
import numpy as np
a = np.array([1, 2, 3, 4])
print(a.dtype, a.shape, a.ndim, a.itemsize, a.nbytes) # int64 (4,) 1 8 32
print(np.array([1, 2.5]).dtype) # float64 (upcast to a common type)
print(np.array([1, "a"]).dtype) # <U21 (everything became a string!)
np.zeros((2, 3)); np.ones(4); np.full((2, 2), 7); np.eye(2)
print(np.arange(0, 1, 0.25)) # [0. 0.25 0.5 0.75] (stop excluded)
print(np.linspace(0, 1, 5)) # [0. 0.25 0.5 0.75 1. ] (stop included)
x = np.arange(12).reshape(3, 4) # reshape keeps the same data
print(x)
# [[ 0 1 2 3]
# [ 4 5 6 7]
# [ 8 9 10 11]]
dtypes that matter in ML
| dtype | Bytes | Use |
|---|---|---|
float64 | 8 | NumPy and pandas default; scientific precision |
float32 | 4 | Deep learning default; halves memory |
float16 / bfloat16 | 2 | Mixed-precision training and inference (bfloat16 lives in PyTorch, not core NumPy) |
int64, int32, int8, uint8 | 8 / 4 / 1 / 1 | Labels, token IDs, image pixels (0 to 255), quantized weights |
bool | 1 | Masks |
img = np.array([200, 100], dtype=np.uint8)
print(img + img) # [144 200]: 400 wrapped around modulo 256!
print(img.astype(np.int32) + img) # [400 200] cast before arithmetic
print(np.array([127], dtype=np.int8) + 1) # [-128] silent overflow in arrays
X32 = np.random.default_rng(0).normal(size=(1000, 100)).astype(np.float32) # half the RAM of float64
Vectorization and ufuncs
A ufunc (universal function) applies an operation element-wise in compiled code: arithmetic operators, np.exp, np.log, np.maximum, comparisons. Replacing a Python loop with one ufunc call typically speeds things up 10 to 100 times.
x = np.random.default_rng(0).normal(size=1_000_000)
# slow: interpreted loop, one Python float object per element
out = [max(0.0, v) for v in x]
# fast: one call into C
relu = np.maximum(0, x)
sigmoid = 1 / (1 + np.exp(-x))
clipped = np.clip(x, -1, 1)
labels = np.where(x > 0, 1, 0) # vectorized if/else
# in a notebook:
# %timeit [max(0.0, v) for v in x] # on the order of 100 ms
# %timeit np.maximum(0, x) # on the order of 1 ms
Indexing: basic, boolean and fancy
x = np.arange(12).reshape(3, 4)
print(x[1, 2]) # 6 row 1, column 2
print(x[:, 1]) # [1 5 9] every row, column 1
print(x[1:, ::2]) # [[ 4 6] [ 8 10]] rows 1.., every 2nd column
print(x[-1]) # [ 8 9 10 11] last row
print(x[x % 2 == 0]) # [ 0 2 4 6 8 10] boolean mask -> 1-D result
print(x[[0, 2], [1, 3]]) # [ 1 11] fancy: picks (0,1) and (2,3)
print(x[[0, 2]]) # rows 0 and 2
mask = (x > 2) & (x < 8) # combine masks with & | ~ and parentheses, not and/or
Views versus copies
Basic slicing (start:stop:step) returns a view: a new array object that shares the original memory. Fancy (integer-list) and boolean indexing return copies. reshape returns a view when possible; .copy() always copies.
b = np.arange(6)
v = b[1:4] # view
v[0] = 99
print(b) # [ 0 99 2 3 4 5] original changed!
c = b[[1, 2]] # copy (fancy indexing)
c[0] = -1
print(b) # [ 0 99 2 3 4 5] unchanged
print(np.shares_memory(b, v), np.shares_memory(b, c)) # True False
safe = b[1:4].copy() # explicit copy when you plan to modify
Axes and aggregation
The axis argument names the dimension that gets collapsed. For a table of shape (rows, columns), axis=0 aggregates down the rows and returns one value per column; axis=1 aggregates across columns and returns one value per row.
axis=1 -->
+----+----+----+----+
axis=0 | 0 | 1 | 2 | 3 | sum(axis=1) -> 6
| +----+----+----+----+
v | 4 | 5 | 6 | 7 | 22
+----+----+----+----+
| 8 | 9 | 10 | 11 | 38
+----+----+----+----+
sum(axis=0) -> 12 15 18 21
x = np.arange(12).reshape(3, 4)
print(x.sum(axis=0)) # [12 15 18 21] per column
print(x.sum(axis=1)) # [ 6 22 38] per row
print(x.mean(axis=1, keepdims=True).shape) # (3, 1) keeps a broadcastable shape
print(np.argmax([[1, 9], [8, 2]], axis=1)) # [1 0] index of the max in each row
print(np.argsort([30, 10, 20])) # [1 2 0] indices that would sort
vals, counts = np.unique([3, 1, 3, 2], return_counts=True) # [1 2 3] [1 1 2]
Broadcasting
When two arrays have different shapes, NumPy compares their shapes from the right. Two dimensions are compatible if they are equal or one of them is 1 (a missing leading dimension counts as 1). Size-1 dimensions are stretched virtually to match.
A (8, 1, 6, 1) B (7, 1, 5) Result (8, 7, 6, 5) X (100, 3) data: 100 samples, 3 features mean (3,) per-feature mean X - mean (100, 3) centred data, no loop, no copy of mean
M = np.array([[1, 2, 3], [4, 5, 6]])
print(M - M.mean(axis=0)) # centre each column
# [[-1.5 -1.5 -1.5]
# [ 1.5 1.5 1.5]]
col = np.array([[10], [20]]) # shape (2, 1)
row = np.array([1, 2, 3]) # shape (3,)
print(col + row) # shape (2, 3): [[11 12 13] [21 22 23]]
# standardize features (what StandardScaler does)
X = np.random.default_rng(0).normal(5, 2, size=(100, 3))
Xs = (X - X.mean(axis=0)) / X.std(axis=0)
# pairwise squared distances between points (n, d) and centroids (k, d)
P, C = np.random.rand(5, 2), np.random.rand(3, 2)
D = ((P[:, None, :] - C[None, :, :]) ** 2).sum(axis=-1) # (5, 1, 2) - (1, 3, 2) -> (5, 3)
np.ones(3) + np.ones(4)
# ValueError: operands could not be broadcast together with shapes (3,) (4,)
Use arr[:, None] or np.newaxis or np.expand_dims to add a size-1 axis deliberately; use squeeze to remove them.
Combining and reshaping
X = np.array([[1., 2.], [3., 4.]])
np.concatenate([X, X], axis=0).shape # (4, 2) join along an existing axis
np.stack([X, X]).shape # (2, 2, 2) create a new axis
np.hstack([X, X]).shape # (2, 4)
X.reshape(-1) # flatten (view if possible); -1 means "infer"
X.ravel(); X.flatten() # ravel: view if possible; flatten: always a copy
X.T # transpose is a view (not contiguous)
Linear algebra
X = np.array([[1., 2.], [3., 4.]])
print(X @ X) # matrix product [[ 7. 10.] [15. 22.]]
print(X * X) # element-wise [[ 1. 4.] [ 9. 16.]]
A = np.array([[3., 1.], [1., 2.]]); b = np.array([9., 8.])
print(np.linalg.solve(A, b)) # [2. 3.] prefer solve over inv(A) @ b
print(np.linalg.norm([3, 4])) # 5.0
print(np.linalg.eigh(A)[0].round(3)) # [1.382 3.618] eigenvalues of a symmetric matrix
U, S, Vt = np.linalg.svd(np.random.rand(4, 3), full_matrices=False) # basis of PCA
w, *_ = np.linalg.lstsq(np.c_[np.ones(4), [1, 2, 3, 4]], [3, 5, 7, 9], rcond=None)
print(w.round(3)) # [1. 2.] intercept 1, slope 2 (y = 1 + 2x)
# one dense neural-network layer
Xb = np.random.default_rng(0).normal(size=(4, 3)) # 4 samples, 3 features
W = np.random.default_rng(1).normal(size=(3, 2)); bias = np.array([0.1, -0.2])
Z = Xb @ W + bias # (4,3) @ (3,2) = (4,2); bias broadcasts across rows
Statistics and a numerically stable softmax
d = [22, 18, 14, 10, 15, 20, 25, 12]
print(np.mean(d), np.median(d)) # 17.0 16.5
print(np.var(d), np.var(d, ddof=1).round(3)) # 23.25 26.571 population vs sample
print(np.std(d).round(3)) # 4.822
print(np.percentile([15, 18, 67, 100, 21, 50], [25, 50, 75])) # [18.75 35.5 62.75]
print(np.quantile([15, 18, 67, 100, 21, 50], 1/3)) # 20.0 (linear interpolation)
print(np.nan == np.nan, np.mean([1, np.nan, 3]), np.nanmean([1, np.nan, 3])) # False nan 2.0
def softmax(z, axis=-1):
z = z - z.max(axis=axis, keepdims=True) # subtracting the max prevents overflow
e = np.exp(z)
return e / e.sum(axis=axis, keepdims=True)
print(softmax(np.array([1000., 1001., 1002.])).round(3)) # [0.09 0.245 0.665]
Note that NumPy's var and std default to the population formula (ddof=0) while pandas defaults to the sample formula (ddof=1), so the same column can give slightly different answers in the two libraries. The statistics themselves are covered in Math & Statistics for ML.
Random numbers done right
rng = np.random.default_rng(42) # modern Generator API, local and seedable
rng.normal(0, 1, size=(2, 3)) # Gaussian
rng.uniform(0, 1, size=5)
rng.integers(0, 10, size=5) # high end excluded
rng.choice(["a", "b", "c"], size=4, p=[0.5, 0.3, 0.2])
idx = rng.permutation(100) # shuffled indices for a manual split
train_idx, test_idx = idx[:80], idx[80:]
Prefer default_rng over the legacy global np.random.seed: a generator object can be passed around explicitly, so two parts of a program do not silently share and disturb one global random state.
(5,) minus a column of shape (5, 1) broadcasts to a (5, 5) matrix rather than raising an error, which silently produces a wrong loss. Print shapes at every step, and add assert y_pred.shape == y_true.shape in metric code. Other classics: using and/or on arrays (use &, | with parentheses), integer overflow in uint8 images, and modifying a slice that turned out to be a view.axis=0 means, * versus @, and "implement softmax, standardization, cosine similarity or pairwise distances without loops". Numerical stability (subtract the max before exp) is a frequent follow-up.pandas: working with tables
pandas wraps NumPy arrays with labels. A Series is a one-dimensional array with an index (row labels). A DataFrame is a collection of Series that share the same index, one per column, and each column can have its own dtype. Almost every operation aligns on labels, which is powerful (joins and arithmetic just work) and occasionally surprising (mismatched indexes produce NaN).
A DataFrame is a spreadsheet with superpowers. Each column is a typed list, the row numbers on the left are the index, and every operation you would do with filters, pivot tables and VLOOKUP has a one-line equivalent. groupby is sorting receipts into envelopes by store, then totalling each envelope. Mapping back: filtering is a boolean mask over rows, pivot tables are pivot_table, VLOOKUP is merge, and "split, apply, combine" is groupby().agg().
Series and DataFrame basics
import pandas as pd, numpy as np
s = pd.Series([10, 20, 30], index=["a", "b", "c"])
print(s["b"], s.iloc[-1], s[s > 15].index.tolist()) # 20 30 ['b', 'c']
df = pd.DataFrame({
"name": ["Asha", "Ben", "Chen", "Dia"],
"age": [34, np.nan, 29, 41],
"dept": ["ml", "web", "ml", None],
"salary": [120, 80, 110, 95],
})
df.shape # (4, 4)
df.dtypes # name: str (object before pandas 3), age: float64, dept: str, salary: int64
df.head(); df.tail(2); df.sample(2, random_state=0)
df.info() # dtypes, non-null counts, memory
df.describe() # count, mean, std, min, quartiles, max for numeric columns
df.columns.tolist(); df.index
Notice that age became float64: a classic NumPy-backed integer column cannot hold NaN, so pandas upcasts it. Use the nullable dtype "Int64" (capital I) if you need integers with missing values; it displays missing entries as <NA>.
Loading and saving data
df = pd.read_csv("data.csv",
usecols=["id", "ts", "amount", "city"], # read only what you need
dtype={"city": "category", "amount": "float32"},
parse_dates=["ts"],
na_values=["", "NA", "?"])
pd.read_parquet("data.parquet") # columnar, typed, compressed: best for big data
pd.read_json("records.jsonl", lines=True)
pd.read_excel("sheet.xlsx", sheet_name="Q1")
pd.read_sql("SELECT * FROM orders", connection)
for chunk in pd.read_csv("huge.csv", chunksize=100_000): # stream a file bigger than RAM
process(chunk)
df.to_parquet("clean.parquet", index=False)
df.to_csv("out.csv", index=False) # index=False avoids an extra unnamed column
Selecting: [], loc and iloc
| Syntax | Selects by | Slice end | Example |
|---|---|---|---|
df["col"] / df[["a", "b"]] | Column name(s) | n/a | A Series / a DataFrame |
df.loc[rows, cols] | Labels and boolean masks | Inclusive | df.loc[df.age > 30, ["name", "salary"]] |
df.iloc[rows, cols] | Integer positions | Exclusive | df.iloc[:10, 0:3] |
df.at[r, c] / df.iat[i, j] | Single scalar, fast | n/a | df.at[0, "name"] |
print(df.loc[0:2, "salary"].tolist()) # [120, 80, 110] labels 0..2 INCLUDED
print(df.iloc[0:2]["salary"].tolist()) # [120, 80] positions 0,1 (2 excluded)
print(df.loc[df.age > 30, ["name", "salary"]])
# name salary
# 0 Asha 120
# 3 Dia 95
Filtering and creating columns
high = df[(df.salary > 90) & (df.dept == "ml")] # & | ~ with parentheses
df[df.dept.isin(["ml", "data"])]
df[df.name.str.startswith("A")]
df.query("salary > 90 and dept == 'ml'") # readable string filter
df["salary_k"] = df.salary * 1000 # vectorized new column
df = df.assign(ratio=lambda t: t.salary / t.salary.max()) # chain-friendly
df["band"] = np.where(df.salary >= 100, "high", "normal")
df["tier"] = np.select([df.salary < 90, df.salary < 110], ["low", "mid"], default="top")
df = df.rename(columns={"salary": "pay"}).drop(columns=["salary_k"])
df.sort_values(["dept", "pay"], ascending=[True, False])
df.nlargest(2, "pay")
groupby: split, apply, combine
rows split by key apply (sum) combine Pune 250 --> Pune: 250, 150 --> 400 --\ Delhi 400 --> Delhi: 400, 100 --> 500 ----> Delhi 500 | Mumbai 300 | Pune 400 Pune 150 Mumbai: 300 --> 300 --/ Mumbai 300 Delhi 100
sales = pd.DataFrame({
"city": ["Pune", "Delhi", "Pune", "Mumbai", "Delhi"],
"sales": [250, 400, 150, 300, 100],
"units": [5, 8, 3, 6, 2],
})
print(sales.groupby("city")["sales"].sum())
# city
# Delhi 500
# Mumbai 300
# Pune 400
# Name: sales, dtype: int64
print(sales.groupby("city").agg(total=("sales", "sum"), avg_units=("units", "mean")))
# total avg_units
# city
# Delhi 500 5.0
# Mumbai 300 6.0
# Pune 400 4.0
# transform returns a result aligned to the ORIGINAL rows (same length)
sales["share"] = sales.sales / sales.groupby("city")["sales"].transform("sum")
print(sales.share.tolist()) # [0.625, 0.8, 0.375, 1.0, 0.2]
sales.groupby("city", as_index=False)["sales"].sum() # key stays a column
sales.groupby("city").filter(lambda g: g.sales.sum() > 350) # keep whole groups
sales.groupby(["city", "units"]).size() # multi-key counts
sales.city.value_counts(normalize=True) # class proportions
| Method | Output shape | Use for |
|---|---|---|
agg | One row per group | Summaries: sum, mean, count, custom reductions |
transform | Same length as input | Group-relative features: share of group, z-score within group, fill with group mean |
filter | Subset of original rows | Drop small or irrelevant groups |
apply | Anything | Flexible but slow; last resort |
Combining tables: merge, join, concat
left = pd.DataFrame({"k": [1, 2, 3], "a": ["x", "y", "z"]})
right = pd.DataFrame({"k": [2, 3, 4], "b": [True, False, True]})
print(pd.merge(left, right, on="k", how="left", indicator=True))
# k a b _merge
# 0 1 x NaN left_only
# 1 2 y True both
# 2 3 z False both
pd.merge(orders, customers, on="customer_id", how="left",
validate="many_to_one") # raises if customers has duplicate keys
pd.merge(a, b, left_on="uid", right_on="user_id", suffixes=("_a", "_b"))
a.join(b, how="inner") # join on the index
pd.concat([jan, feb], ignore_index=True) # stack rows (same columns)
pd.concat([X_num, X_cat], axis=1) # side by side (aligns on index!)
how= | Keeps | SQL equivalent |
|---|---|---|
inner (default) | Keys present in both | INNER JOIN |
left | All left keys; unmatched right columns become NaN | LEFT JOIN |
right | All right keys | RIGHT JOIN |
outer | Union of keys | FULL OUTER JOIN |
cross | Every combination | CROSS JOIN |
Reshaping: pivot, pivot_table, melt, crosstab
wide = pd.DataFrame({"id": [1, 2], "jan": [10, 20], "feb": [11, 21]})
long = wide.melt(id_vars="id", var_name="month", value_name="sales") # wide -> long
print(long)
# id month sales
# 0 1 jan 10
# 1 2 jan 20
# 2 1 feb 11
# 3 2 feb 21
back = long.pivot(index="id", columns="month", values="sales") # long -> wide
rev = pd.DataFrame({"region": ["N", "N", "S", "S"], "q": ["Q1", "Q2", "Q1", "Q2"], "rev": [10, 15, 7, 9]})
print(rev.pivot_table(index="region", columns="q", values="rev", aggfunc="sum", margins=True))
# q Q1 Q2 All
# region
# N 10 15 25
# S 7 9 16
# All 17 24 41
pd.crosstab(rev.region, rev.q) # frequency table of two categoricals
pivot fails if an index/column pair appears twice; pivot_table aggregates duplicates instead. stack/unstack move levels between the row index and the columns.
Missing values
df.isna().sum() # missing count per column
df.isna().mean().sort_values() # missing fraction, worst last
df.dropna() # drop rows with ANY missing value
df.dropna(subset=["age"]) # only rows missing age
df.dropna(axis=1, thresh=int(0.7 * len(df))) # keep columns at least 70% filled
df.fillna({"age": df.age.median(), "dept": "unknown"})
df.age.fillna(df.groupby("dept").age.transform("median")) # group-aware imputation
ts.ffill(); ts.bfill(); ts.interpolate() # time series: forward-fill, back-fill, interpolate
df["age_missing"] = df.age.isna().astype(int) # missingness can itself be informative
For modelling, prefer imputing inside a scikit-learn pipeline (SimpleImputer) so the fill values are learned from the training fold only. Filling the whole DataFrame with the global median before splitting is a small but real form of leakage.
apply versus vectorized operations
| Approach | Speed | Example |
|---|---|---|
| Vectorized column arithmetic | Fastest | df.a * df.b, np.log1p(df.x) |
| Built-in accessors | Fast | df.s.str.lower(), df.t.dt.hour |
np.where / np.select / map with a dict | Fast | df.code.map({"A": 1, "B": 2}) |
Series.apply(func) | Slow (Python per element) | Complex custom logic |
df.apply(func, axis=1) | Slower (builds a Series per row) | Row logic spanning many columns |
iterrows() | Slowest; also loses dtypes | Almost never |
# slow
df["bmi"] = df.apply(lambda r: r.weight / r.height ** 2, axis=1)
# fast (often 100x or more on large frames)
df["bmi"] = df.weight / df.height ** 2
# if you truly need a loop, itertuples is much faster than iterrows
for row in df.itertuples(index=False):
pass
Strings, dates and times
names = pd.Series(["a b", "c"])
print(names.str.split().str.len().tolist(), names.str.upper().tolist()) # [2, 1] ['A B', 'C']
df.email.str.extract(r"@(\w+)\.") # regex capture group to a column
df.text.str.contains("refund", case=False, na=False)
ev = pd.DataFrame({
"ts": pd.to_datetime(["2026-01-01 09:00", "2026-01-01 17:00", "2026-01-02 10:00", "2026-01-04 12:00"]),
"clicks": [3, 5, 2, 7]})
ev["dow"] = ev.ts.dt.day_name(); ev["hour"] = ev.ts.dt.hour # Thursday 9, Thursday 17, ...
print(ev.set_index("ts").resample("D")["clicks"].sum().tolist()) # [8, 2, 0, 7] gap day filled with 0
print(ev.clicks.rolling(2).mean().tolist()) # [nan, 4.0, 3.5, 4.5]
print(ev.clicks.shift(1).tolist()) # [nan, 3.0, 5.0, 2.0] lag feature
print(ev.clicks.diff().tolist()) # [nan, 2.0, -3.0, 5.0]
pd.to_datetime(raw, format="%d/%m/%Y", errors="coerce") # explicit format; bad values -> NaT
ev.ts.dt.tz_localize("UTC").dt.tz_convert("Asia/Kolkata")
center=True) or a shift(-1) uses the future and leaks the target into features. Sort by time and by entity before computing lags, and compute them per entity with groupby("user").x.shift(1).Binning and categoricals
ages = [5, 17, 30, 64, 80]
print(pd.cut(ages, bins=[0, 18, 60, 120], labels=["child", "adult", "senior"]).tolist())
# ['child', 'child', 'adult', 'senior', 'senior'] fixed edges, (0, 18] includes 18
print(pd.qcut(pd.Series([1, 2, 3, 4, 100]), 2, labels=["lo", "hi"]).tolist())
# ['lo', 'lo', 'lo', 'hi', 'hi'] equal-frequency bins
c = pd.Series(["lo", "hi", "lo"], dtype="category")
print(c.cat.categories.tolist(), c.cat.codes.tolist()) # ['hi', 'lo'] [1, 0, 1]
size = pd.Categorical(["M", "S", "L"], categories=["S", "M", "L"], ordered=True) # ordered: S < M < L
pd.get_dummies(df.dept, prefix="dept", dtype=int) # one-hot for quick analysis
A category column stores each distinct value once plus a small integer code per row, which slashes memory for repetitive strings (city, product, country) and speeds up groupby.
Memory optimisation
big = pd.DataFrame({"i": np.arange(1000, dtype="int64"), "f": np.random.rand(1000)})
print(big.memory_usage(deep=True).sum()) # 16132 bytes (index + 8 + 8 bytes per row)
big["i"] = pd.to_numeric(big.i, downcast="integer") # int16 is enough for 0..999
big["f"] = big.f.astype("float32")
print(big.memory_usage(deep=True).sum()) # 6132 bytes, about 62% less
def shrink(df):
for col in df.select_dtypes("integer"):
df[col] = pd.to_numeric(df[col], downcast="integer")
for col in df.select_dtypes("float"):
df[col] = pd.to_numeric(df[col], downcast="float")
for col in df.select_dtypes(["object", "string"]):
if df[col].nunique() / len(df) < 0.5: # repetitive text -> category
df[col] = df[col].astype("category")
return df
- Measure
df.info(memory_usage="deep")anddf.memory_usage(deep=True). - Read less
usecols,nrowsfor prototyping, filter early. - Type tightly Downcast numbers,
categoryfor low-cardinality strings, Arrow-backed strings. - Store columnar Parquet keeps dtypes and lets you read only some columns.
- Stream or scale out
chunksize, or move to Polars, DuckDB, Dask or Spark when data truly exceeds one machine's memory.
Copies, views and chained assignment
# WRONG: chained indexing. The first [] may return a copy, so the write is lost.
df[df.salary > 90]["salary"] = 0 # warning; original df unchanged
# RIGHT: one .loc call selects rows and column together
df.loc[df.salary > 90, "salary"] = 0
# taking a subset you intend to modify independently
sub = df[df.salary > 90].copy()
Older pandas emits SettingWithCopyWarning for ambiguous cases. With Copy-on-Write (the default behaviour from pandas 3.0), any derived object behaves like a copy, chained assignment never modifies the original, and pandas warns with a ChainedAssignmentError. The fix is the same in every version: one .loc assignment.
pd.concat([X_num, X_cat], axis=1) after one frame was filtered or shuffled aligns rows by index label, filling mismatches with NaN and silently doubling the row count. Reset indexes with reset_index(drop=True) when you mean "by position". Also watch for many-to-many merges that explode row counts; use validate= and compare len() before and after.loc versus iloc (labels, inclusive end, versus positions, exclusive end); merge types; agg versus transform; handling missing values; why apply is slow; how to reduce memory; SettingWithCopy. Live tasks: "top 3 products by revenue per region", "fraction of missing values per column", "day-over-day change per user", "pivot this long table". Show the idiomatic vectorized answer, then mention edge cases (ties, NaN, duplicates).# "top 2 cities by total sales" and "top 1 row per group"
sales.groupby("city")["sales"].sum().nlargest(2)
sales.sort_values("sales", ascending=False).groupby("city").head(1)
Visualization: matplotlib and seaborn
matplotlib is the low-level plotting engine: a Figure is the canvas, and each Axes is one plot area on it with its own x-axis, y-axis, title and legend. seaborn sits on top and draws statistical plots straight from DataFrames with sensible defaults, colouring by a column with a single hue= argument. Pick the plot from the question you are asking, not from habit.
matplotlib is a set of drawing instruments and a sheet of paper: total control, but you draw every line yourself. seaborn is a template book of standard charts: open to "compare distributions by group", hand it your table, and it draws the chart properly. Mapping back: fig, ax = plt.subplots() gives you the paper and a drawing area, and seaborn functions accept ax= so you can place their charts onto your own layout and still tweak them with matplotlib calls.
The object-oriented matplotlib pattern
import numpy as np, matplotlib.pyplot as plt
x = np.linspace(-6, 6, 200)
fig, axes = plt.subplots(1, 2, figsize=(10, 3.5)) # 1 row, 2 plots
axes[0].plot(x, 1 / (1 + np.exp(-x)), label="sigmoid")
axes[0].plot(x, np.tanh(x), label="tanh", linestyle="--")
axes[0].axhline(0, color="gray", lw=0.8)
axes[0].set(title="Activations", xlabel="z", ylabel="output")
axes[0].legend()
axes[1].hist(np.random.default_rng(0).normal(size=1000), bins=30, alpha=0.7)
axes[1].set_title("Histogram")
fig.tight_layout()
fig.savefig("plots.png", dpi=150) # or plt.show() in scripts
seaborn one-liners
import seaborn as sns
sns.histplot(df, x="mean radius", hue="diagnosis", kde=True) # distribution by class
sns.boxplot(df, x="diagnosis", y="mean area") # spread and outliers
sns.scatterplot(df, x="mean radius", y="mean texture", hue="diagnosis")
sns.countplot(df, x="diagnosis") # class balance
sns.heatmap(df.corr(numeric_only=True), cmap="coolwarm", center=0) # correlations
sns.pairplot(df[cols + ["diagnosis"]], hue="diagnosis") # all pairs (small cols only)
sns.lineplot(history, x="epoch", y="loss", hue="split") # learning curves
Which plot answers which question
| Question | Plot | Watch out for |
|---|---|---|
| What does one numeric variable look like? | Histogram, KDE, box plot, violin | Bin count changes the story; use log scale for heavy tails |
| How balanced are my classes or categories? | Count / bar plot | Sort bars; show percentages |
| Do two numeric variables move together? | Scatter (hexbin or alpha for many points) | Overplotting hides density |
| How does a numeric variable differ across groups? | Box, violin, strip, grouped histogram | Unequal group sizes |
| How does something change over time? | Line plot | Irregular timestamps; resample first |
| Which features are correlated? | Correlation heatmap | Pearson only captures linear relationships |
| How is my model doing? | Confusion matrix, ROC, precision-recall curve, calibration plot, residual plot, learning curve | ROC looks optimistic under heavy imbalance; prefer PR curves |
| What drives the model? | Horizontal bar of coefficients or importances, partial dependence | Impurity importances favour high-cardinality features |
Model diagnostic plots
from sklearn.metrics import ConfusionMatrixDisplay, RocCurveDisplay, PrecisionRecallDisplay
ConfusionMatrixDisplay.from_estimator(model, X_test, y_test)
RocCurveDisplay.from_estimator(model, X_test, y_test)
PrecisionRecallDisplay.from_estimator(model, X_test, y_test)
# learning curves from a PyTorch or Keras history
plt.plot(train_losses, label="train"); plt.plot(val_losses, label="val"); plt.legend()
# val loss rising while train keeps falling = overfitting: stop earlier or regularize
plt.* interface with the object-oriented ax.* interface in one figure, so labels land on the wrong subplot. Pick fig, ax = plt.subplots() and stick with ax. Other mistakes: truncated y-axes that exaggerate differences, pie charts with many slices, and pairplots on 30 columns that take minutes and show nothing readable.Exploratory data analysis (EDA) workflow
EDA is the structured interrogation of a dataset before modelling. Its goals are to understand what each column means, find data-quality problems, spot relationships worth modelling, and catch leakage before it inflates your metrics. Good EDA is a checklist you run every time, then a few targeted deep dives.
EDA is a pre-flight inspection. A pilot walks around the aircraft checking tyres, fuel, and flaps before take-off, following the same checklist every flight, because problems are cheap on the ground and catastrophic in the air. Mapping back: checking shape, types, missing values, duplicates, distributions, class balance and leakage before training is cheap; discovering them after a model is deployed is expensive.
- Understand the question What is the target, what is the unit of prediction (patient, transaction, day), and what data is available at prediction time?
- Structure
shape,dtypes,head(),info(). Are numbers stored as strings? Are dates parsed? - Quality Missing values per column, duplicates (
df.duplicated().sum()), impossible values (negative ages), inconsistent categories ("NY" versus "New York"). - Univariate
describe(), histograms,value_counts(); note skew, outliers and scale differences. - Target Class balance for classification; distribution and skew for regression.
- Bivariate Each feature against the target: grouped means, box plots, correlation with the target.
- Multivariate Correlation heatmap, multicollinearity, a quick PCA scatter.
- Leakage hunt Features that are near-perfect predictors, IDs correlated with the target, columns created after the outcome, duplicate entities across train and test.
- Decide Write down the preprocessing each finding implies (impute, log-transform, encode, drop, stratify, group split).
A reusable EDA snippet
def quick_eda(df, target):
print("shape:", df.shape)
print("duplicates:", df.duplicated().sum())
report = pd.DataFrame({
"dtype": df.dtypes.astype(str),
"missing_%": (df.isna().mean() * 100).round(1),
"unique": df.nunique(),
})
print(report.sort_values("missing_%", ascending=False).head(15))
print(df[target].value_counts(normalize=True).round(3))
num = df.select_dtypes("number").drop(columns=[target], errors="ignore")
print(num.describe().T[["mean", "std", "min", "max"]].round(2).head(10))
print("skew (top):", num.skew().abs().sort_values(ascending=False).head(5).round(2).to_dict())
corr = num.corrwith(df[target]).abs().sort_values(ascending=False)
print("most target-correlated:", corr.head(5).round(3).to_dict())
Descriptive statistics in pandas
scores = pd.Series([22, 18, 14, 10, 15, 20, 25, 12])
scores.mean(), scores.median(), scores.mode().tolist() # 17.0 16.5 [10, 12, 14, ...] (all tie)
scores.var(), scores.std() # 26.571..., 5.154... (sample, ddof=1 by default)
scores.var(ddof=0) # 23.25 population variance, matches NumPy's default
scores.quantile([0.25, 0.5, 0.75]) # quartiles
iqr = scores.quantile(0.75) - scores.quantile(0.25)
(scores - scores.mean()).abs().mean() # mean absolute deviation: 4.25
scores.max() - scores.min() # range: 15
The mean is pulled by outliers while the median is not, so a mean far above the median signals right skew (incomes, prices, counts). The IQR rule flags outliers below Q1 − 1.5 × IQR or above Q3 + 1.5 × IQR, which is exactly what box-plot whiskers show.
Outliers: detect, then decide
q1, q3 = df.amount.quantile([0.25, 0.75])
iqr = q3 - q1
outliers = df[(df.amount < q1 - 1.5 * iqr) | (df.amount > q3 + 1.5 * iqr)]
z = (df.amount - df.amount.mean()) / df.amount.std() # |z| > 3 rule for roughly normal data
df["amount_log"] = np.log1p(df.amount) # compress a heavy right tail
df["amount_clip"] = df.amount.clip(upper=df.amount.quantile(0.99)) # winsorize
Do not delete outliers automatically. A 10-million transaction may be an error or exactly the fraud you want to catch. Investigate, then choose to fix, cap, transform, or keep and use a robust model or loss.
scikit-learn: the estimator API, pipelines and evaluation
scikit-learn succeeds because every object follows the same small contract. An estimator learns from data with fit(X, y). A transformer also has transform(X) (and the shortcut fit_transform). A predictor has predict(X), and classifiers usually add predict_proba(X) and score(X, y). Hyperparameters are set in the constructor; learned values get a trailing underscore (coef_, mean_, classes_). Once you know this contract, every model in the library is familiar.
A pipeline is a factory assembly line. At commissioning time (fit) each station calibrates its tools on sample parts: the imputer measures typical values, the scaler measures means and spreads, the model learns its weights. In production (predict) new parts travel through the same calibrated stations in the same order, and nobody recalibrates using the customer's parts. Mapping back: fit learns parameters only from training data, transform/predict apply them unchanged to new data, and wrapping all stations in one Pipeline guarantees that cross-validation recalibrates them on each training fold only.
The contract in five lines
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
scaler = StandardScaler() # hyperparameters go in the constructor
X_train_s = scaler.fit_transform(X_train) # learns mean_ and scale_ from TRAIN only
X_test_s = scaler.transform(X_test) # re-uses train statistics: no leakage
clf = LogisticRegression(max_iter=1000).fit(X_train_s, y_train)
clf.predict(X_test_s); clf.predict_proba(X_test_s)[:, 1]; clf.score(X_test_s, y_test)
| Method | Who has it | What it does |
|---|---|---|
fit(X, y) | Every estimator | Learn parameters; returns self so calls can chain |
transform(X) | Transformers (scalers, encoders, PCA, imputers) | Apply learned parameters to data |
fit_transform(X) | Transformers | Both at once; use only on training data |
predict(X) | Models | Class labels or regression values |
predict_proba(X) / decision_function(X) | Most classifiers | Probabilities per class / raw scores (margins) |
score(X, y) | Models | Accuracy for classifiers, R² for regressors |
get_params() / set_params() | Every estimator | How grid search reads and changes hyperparameters |
Preprocessing toolbox
| Transformer | Does | Use when |
|---|---|---|
StandardScaler | Subtract mean, divide by standard deviation | Linear models, SVM, kNN, PCA, neural nets |
MinMaxScaler | Rescale to [0, 1] | Bounded inputs, images, some neural nets |
RobustScaler | Subtract median, divide by IQR | Features with strong outliers |
PowerTransformer, np.log1p | Reduce skew | Heavy-tailed amounts, counts |
SimpleImputer, KNNImputer | Fill missing values (mean, median, most frequent, constant) | Any column with gaps; add add_indicator=True to keep a "was missing" flag |
OneHotEncoder | One binary column per category | Nominal categories (city, colour); use handle_unknown="ignore" |
OrdinalEncoder | Map categories to integers | Ordered categories, or tree models |
TargetEncoder | Replace a category by a cross-fitted mean target | High-cardinality categories (zip codes) |
PolynomialFeatures | Add squares and interactions | Let linear models fit curves |
TfidfVectorizer | Text to sparse weighted word counts | Classic text classification and retrieval baselines |
Distance- and gradient-based models care about feature scale; tree-based models (decision trees, random forests, gradient boosting) do not. On the breast-cancer data, scaling moves kNN from 0.912 to 0.956 test accuracy and an RBF SVM from 0.904 to 0.974, while an unscaled random forest already reaches 0.965.
Pipelines and ColumnTransformer
import numpy as np, pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.linear_model import LogisticRegression
df = pd.DataFrame({
"age": [25, np.nan, 47, 35, 52, 23],
"income": [40, 60, np.nan, 80, 120, 30],
"city": ["Pune", "Delhi", "Pune", np.nan, "Mumbai", "Delhi"],
"bought": [0, 1, 1, 0, 1, 0],
})
X, y = df.drop(columns="bought"), df.bought
numeric, categorical = ["age", "income"], ["city"]
preprocess = ColumnTransformer([
("num", Pipeline([("impute", SimpleImputer(strategy="median")),
("scale", StandardScaler())]), numeric),
("cat", Pipeline([("impute", SimpleImputer(strategy="most_frequent")),
("onehot", OneHotEncoder(handle_unknown="ignore"))]), categorical),
])
model = Pipeline([("pre", preprocess), ("clf", LogisticRegression())])
model.fit(X, y)
print(model.named_steps["pre"].get_feature_names_out())
# ['num__age' 'num__income' 'cat__city_Delhi' 'cat__city_Mumbai' 'cat__city_Pune']
print(model.predict(pd.DataFrame({"age": [30], "income": [70], "city": ["Chennai"]})))
# [0] unseen city "Chennai" is encoded as all zeros instead of crashing
Parameters of steps inside a pipeline are addressed as step__param, with nesting as deep as needed: clf__C, pre__num__impute__strategy. make_pipeline and make_column_transformer build the same objects with auto-generated step names. Calling set_output(transform="pandas") makes transformers return DataFrames with readable column names.
Splitting data correctly
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42) # stratify keeps class ratios
| Splitter | Use when |
|---|---|
KFold / StratifiedKFold | Independent rows; stratified for classification |
GroupKFold / StratifiedGroupKFold | Several rows per patient, user or device; keeps each group on one side |
TimeSeriesSplit | Temporal data; always trains on the past and validates on the future |
RepeatedStratifiedKFold | Small datasets where one CV run is noisy |
5-fold cross-validation (training data only; test set stays locked away) fold 1: [VAL ][train][train][train][train] fold 2: [train][VAL ][train][train][train] fold 3: [train][train][VAL ][train][train] fold 4: [train][train][train][VAL ][train] fold 5: [train][train][train][train][VAL ] score = mean of 5 validation scores (report the std too)
Cross-validation
from sklearn.model_selection import StratifiedKFold, cross_val_score, cross_validate
from sklearn.pipeline import make_pipeline
from sklearn.datasets import load_breast_cancer
X, y = load_breast_cancer(return_X_y=True)
y = 1 - y # make malignant the positive class (1)
pipe = make_pipeline(StandardScaler(), LogisticRegression(max_iter=5000))
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)
res = cross_validate(pipe, X, y, cv=cv, scoring=["accuracy", "recall", "roc_auc"])
print({k: round(v.mean(), 3) for k, v in res.items() if k.startswith("test")})
# {'test_accuracy': 0.979, 'test_recall': 0.958, 'test_roc_auc': 0.995}
Why the pipeline matters: a leakage demonstration
from sklearn.feature_selection import SelectKBest, f_classif
rng = np.random.default_rng(0)
Xr = rng.normal(size=(100, 10_000)) # pure noise features
yr = rng.integers(0, 2, 100) # random labels: true accuracy is 50%
# WRONG: select the 20 "best" features using ALL rows, then cross-validate
leaky = SelectKBest(f_classif, k=20).fit_transform(Xr, yr)
print(cross_val_score(LogisticRegression(max_iter=1000), leaky, yr, cv=5).mean()) # 0.87 (!)
# RIGHT: selection happens inside each training fold
honest = make_pipeline(SelectKBest(f_classif, k=20), LogisticRegression(max_iter=1000))
print(cross_val_score(honest, Xr, yr, cv=5).mean()) # 0.48
The leaky version reports 87% accuracy on random noise, because the feature selector already "saw" the validation labels. Any step that learns from data, including scaling, imputation, feature selection, target encoding, oversampling (SMOTE) and PCA, must live inside the pipeline. Cross-validation estimates how well that whole procedure generalises; a bootstrap of a single test-set metric instead puts an error bar on a frozen model's score. Use CV to choose models, and a paired bootstrap (or McNemar) on the sealed test set to compare two finished models.
Hyperparameter search
from sklearn.model_selection import GridSearchCV, RandomizedSearchCV
from scipy.stats import loguniform
param_grid = {"logisticregression__C": [0.01, 0.1, 1, 10, 100],
"logisticregression__class_weight": [None, "balanced"]}
grid = GridSearchCV(pipe, param_grid, cv=cv, scoring="recall", n_jobs=-1)
grid.fit(X_train, y_train)
grid.best_params_, grid.best_score_ # best combination and its mean CV score
best = grid.best_estimator_ # already refit on all training data
pd.DataFrame(grid.cv_results_)[["params", "mean_test_score", "std_test_score"]]
rand = RandomizedSearchCV(pipe, {"logisticregression__C": loguniform(1e-3, 1e3)},
n_iter=30, cv=cv, scoring="roc_auc", random_state=0, n_jobs=-1)
- Grid search tries every combination: cost multiplies with each parameter added. Randomized search samples a fixed budget and usually finds good regions faster when only a few parameters matter.
HalvingGridSearchCVand libraries such as Optuna (Bayesian optimisation) spend more budget on promising candidates.best_score_is optimistically biased because it was used for selection. Report final performance on the untouched test set, or use nested cross-validation (a search inside each outer fold) when the dataset is small.- With multiple metrics, pass
scoring={"rec": "recall", "prec": "precision"}and choose which one torefit.
Metrics: pick the one that matches the cost of errors
| Task | Metric | Choose when |
|---|---|---|
| Classification | Accuracy | Balanced classes, equal error costs |
| Recall (sensitivity) | Missing a positive is costly: cancer, fraud, safety | |
| Precision | False alarms are costly: spam filtering, alerts that page a human | |
| F1, macro-F1 | Need a balance; macro-F1 treats every class equally in multiclass problems | |
| ROC-AUC | Threshold-free ranking quality; can look optimistic with rare positives | |
| PR-AUC (average precision) | Rare positive class | |
| Log loss, Brier score | You need well-calibrated probabilities | |
| Regression | MAE | Robust to outliers, same units as the target |
| RMSE | Large errors are especially bad | |
| R² | Fraction of variance explained, relative to predicting the mean | |
| MAPE | Relative error; breaks when true values are near zero |
from sklearn.metrics import (accuracy_score, precision_score, recall_score, f1_score,
confusion_matrix, classification_report, roc_auc_score, average_precision_score,
mean_absolute_error, mean_squared_error, r2_score)
y_pred = best.predict(X_test)
y_proba = best.predict_proba(X_test)[:, 1] # probability of class 1
print(confusion_matrix(y_test, y_pred))
print(classification_report(y_test, y_pred, digits=3))
roc_auc_score(y_test, y_proba) # needs scores, not hard labels
f1_score(y_true_multi, y_pred_multi, average="macro")
Imbalance and decision thresholds
class_weight="balanced"re-weights the loss so the rare class counts more; often the simplest effective fix.- The default 0.5 threshold is arbitrary. Choose a threshold on validation data (or out-of-fold predictions) with
precision_recall_curve, according to the business constraint such as "recall at least 0.95". Newer versions also provideTunedThresholdClassifierCVfor this. - Resampling (under-sampling, SMOTE from the imbalanced-learn library) must be applied only to training folds, inside a pipeline.
- Always compare against a trivial baseline:
DummyClassifier()scores 0.627 on the breast-cancer data just by predicting the majority class.
A model zoo with one API
from sklearn.neighbors import KNeighborsClassifier
from sklearn.svm import SVC
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier, HistGradientBoostingClassifier
models = {
"logreg": make_pipeline(StandardScaler(), LogisticRegression(max_iter=5000)),
"knn": make_pipeline(StandardScaler(), KNeighborsClassifier(n_neighbors=5)),
"svm": make_pipeline(StandardScaler(), SVC(probability=True)),
"tree": DecisionTreeClassifier(max_depth=4, random_state=0),
"forest": RandomForestClassifier(n_estimators=300, random_state=0, n_jobs=-1),
"gbm": HistGradientBoostingClassifier(random_state=0), # handles NaN natively
}
for name, m in models.items():
s = cross_val_score(m, X_train, y_train, cv=cv, scoring="recall")
print(f"{name:7s} recall {s.mean():.3f} +/- {s.std():.3f}")
Algorithm details (how trees split, what the SVM margin is, why kNN suffers in high dimensions) are covered in Machine Learning. Unsupervised tools follow the same API: PCA(n_components=2).fit_transform(X), KMeans(n_clusters=3, n_init=10).fit_predict(X).
Inspecting and saving models
coefs = pd.Series(best[-1].coef_[0], index=feature_names).sort_values(key=abs) # linear models
importances = forest.feature_importances_ # trees
from sklearn.inspection import permutation_importance
pi = permutation_importance(best, X_test, y_test, scoring="recall", n_repeats=10, random_state=0)
import joblib
joblib.dump(best, "model.joblib") # save the WHOLE pipeline, preprocessing included
model = joblib.load("model.joblib") # same library versions required; never load untrusted files
model.predict(new_rows)
fit_transform on the test set, or scaling before splitting; (2) tuning hyperparameters or the threshold on the test set, so the "test" score is no longer unbiased; (3) saving only the model and re-implementing preprocessing by hand at serving time, which drifts. Also note that recall_score uses pos_label=1; in the raw breast-cancer data 0 means malignant, so remap with y = 1 - y or pass pos_label=0.fit versus transform versus fit_transform; what data leakage is and how pipelines prevent it; why stratify; k-fold versus group versus time-series splits; grid versus random search; precision versus recall trade-offs with a concrete business example; ROC-AUC versus PR-AUC under imbalance; how you would pick a threshold; which models need scaling. Strong candidates always mention a baseline and a held-out test set that is touched once.PyTorch fundamentals: tensors, autograd and the training loop
PyTorch rests on three ideas. Tensors are NumPy-like arrays that can live on a GPU. Autograd records the operations you perform on tensors that require gradients and, when you call backward(), applies the chain rule backwards through that record to compute every gradient. Modules (nn.Module) package parameters and a forward computation into reusable, nestable building blocks. Around these sit Dataset/DataLoader for feeding data and optimizers for updating weights.
Autograd is a hiker dropping breadcrumbs. On the way forward, every operation leaves a crumb recording where it came from. When you call backward(), the hiker walks the trail in reverse, and at each crumb works out how much a small step there would have changed the final loss. The optimizer is the guide who then nudges every parameter a little downhill. Mapping back: the crumbs are the dynamic computation graph (grad_fn), the reverse walk is backpropagation, gradients land in each parameter's .grad, and optimizer.step() updates the weights.
Tensors
import torch, numpy as np
x = torch.tensor([[1.0, 2.0], [3.0, 4.0]])
print(x.shape, x.dtype, x.device) # torch.Size([2, 2]) torch.float32 cpu
print(torch.tensor([1, 2]).dtype) # torch.int64 -- the default for class labels
print(x @ x, x * x) # matrix product / element-wise, same as NumPy
print(x.sum(dim=0), x.mean(dim=1)) # tensor([4., 6.]) tensor([1.5000, 3.5000]) dim = axis
torch.zeros(2, 3); torch.randn(4, 3); torch.arange(6); torch.linspace(0, 1, 5)
# reshaping
x.unsqueeze(0).shape # torch.Size([1, 2, 2]) add a batch dimension
x.permute(1, 0); x.T # reorder dimensions
x.T.reshape(4) # tensor([1., 3., 2., 4.]) copies if needed
# x.T.view(4) -> RuntimeError: view needs contiguous memory; use .contiguous() or reshape
# NumPy bridge: shares memory on CPU
a = np.ones(2)
t = torch.from_numpy(a) # float64 tensor sharing a's buffer
a[0] = 5
print(t) # tensor([5., 1.], dtype=torch.float64)
back = t.numpy() # GPU tensors need .cpu() first; grad tensors need .detach()
Devices
device = ("cuda" if torch.cuda.is_available()
else "mps" if torch.backends.mps.is_available() # Apple silicon
else "cpu")
model = model.to(device) # moves parameters in place
xb, yb = xb.to(device), yb.to(device) # tensors return a NEW tensor; reassign!
# every tensor in one operation must be on the same device, or you get
# "Expected all tensors to be on the same device"
Autograd
x = torch.tensor(2.0, requires_grad=True)
y = x ** 3 + 2 * x # dy/dx = 3x^2 + 2
y.backward()
print(x.grad) # tensor(14.)
w = torch.tensor(1.0, requires_grad=True)
for _ in range(2):
loss = (w * 3) ** 2 # d/dw = 18w = 18
loss.backward()
print(w.grad) # tensor(18.) then tensor(36.) -- gradients ACCUMULATE
w.grad.zero_() # this is why training loops call optimizer.zero_grad()
with torch.no_grad(): # no graph recorded: inference and manual updates
w -= 0.1 * 18
frozen = y.detach() # same data, cut from the graph (requires_grad=False)
Gradients accumulate by design, so that you can sum gradients over several small batches before one update (gradient accumulation). The price is that you must clear them each step.
Building models with nn.Module
import torch.nn as nn, torch.nn.functional as F
class MLP(nn.Module):
def __init__(self, d_in, d_hidden=32, n_classes=2, p_drop=0.2):
super().__init__() # must come first
self.net = nn.Sequential(
nn.Linear(d_in, d_hidden), # weight (32, d_in), bias (32,)
nn.ReLU(),
nn.Dropout(p_drop),
nn.Linear(d_hidden, n_classes),
)
def forward(self, x):
return self.net(x) # raw logits: no softmax here
model = MLP(30)
print(sum(p.numel() for p in model.parameters())) # 1058 = 30*32+32 + 32*2+2
print(model(torch.randn(5, 30)).shape) # torch.Size([5, 2])
for name, p in model.named_parameters():
print(name, tuple(p.shape)) # net.0.weight (32, 30) ...
Assigning an nn.Module or nn.Parameter as an attribute registers it automatically, so model.parameters(), .to(device) and state_dict() all find it. A plain Python list of layers is not registered; use nn.ModuleList or nn.ModuleDict.
Losses and output layers
| Task | Model output | Loss | Target dtype and shape |
|---|---|---|---|
| Multiclass | Logits, shape (N, C) | nn.CrossEntropyLoss (applies log-softmax internally) | long, shape (N,), values 0..C-1 |
| Binary / multi-label | One logit per label, shape (N,) or (N, L) | nn.BCEWithLogitsLoss (applies sigmoid internally, numerically stable) | float, same shape as output |
| Regression | Values, shape (N,) or (N, 1) | nn.MSELoss, nn.L1Loss, nn.HuberLoss | float, same shape as output |
logits = torch.tensor([[2.0, 0.5, -1.0]])
print(nn.CrossEntropyLoss()(logits, torch.tensor([0])).item()) # 0.2413... = -log softmax(logits)[0]
probs = F.softmax(logits, dim=1) # only for reporting or inference, never before CrossEntropyLoss
Datasets and DataLoaders
from torch.utils.data import Dataset, DataLoader, TensorDataset
class TabularDataset(Dataset):
def __init__(self, X, y):
self.X = torch.as_tensor(X, dtype=torch.float32)
self.y = torch.as_tensor(y, dtype=torch.long)
def __len__(self):
return len(self.y)
def __getitem__(self, i):
return self.X[i], self.y[i] # one sample; DataLoader stacks them into batches
train_dl = DataLoader(TabularDataset(X_tr, y_tr), batch_size=32, shuffle=True,
num_workers=0, pin_memory=torch.cuda.is_available())
val_dl = DataLoader(TabularDataset(X_va, y_va), batch_size=256, shuffle=False)
xb, yb = next(iter(train_dl))
print(xb.shape, yb.shape) # torch.Size([32, 30]) torch.Size([32])
num_workers > 0 starts worker processes that prepare batches in parallel, which matters when each item needs disk reads or image decoding. For tensors already in memory, 0 is often fastest. A custom collate_fn handles variable-length samples such as padded text.
The training loop
for each epoch:
model.train()
for each batch:
optimizer.zero_grad() 1. clear old gradients
logits = model(xb) 2. forward pass (builds graph)
loss = loss_fn(logits, yb) 3. scalar loss
loss.backward() 4. backprop: fill p.grad
optimizer.step() 5. update weights
model.eval(); with torch.no_grad(): validate, log, checkpoint best
import torch, torch.nn as nn
from torch.utils.data import TensorDataset, DataLoader
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
torch.manual_seed(0)
X, y = load_breast_cancer(return_X_y=True); y = 1 - y # 1 = malignant
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)
sc = StandardScaler().fit(X_tr); X_tr, X_te = sc.transform(X_tr), sc.transform(X_te)
train_ds = TensorDataset(torch.tensor(X_tr, dtype=torch.float32), torch.tensor(y_tr, dtype=torch.long))
test_ds = TensorDataset(torch.tensor(X_te, dtype=torch.float32), torch.tensor(y_te, dtype=torch.long))
train_dl = DataLoader(train_ds, batch_size=32, shuffle=True)
test_dl = DataLoader(test_ds, batch_size=256)
device = "cuda" if torch.cuda.is_available() else "cpu"
model = MLP(30).to(device)
opt = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=1e-4)
loss_fn = nn.CrossEntropyLoss()
for epoch in range(1, 31):
model.train(); total = 0.0
for xb, yb in train_dl:
xb, yb = xb.to(device), yb.to(device)
opt.zero_grad()
loss = loss_fn(model(xb), yb)
loss.backward()
opt.step()
total += loss.item() * xb.size(0) # .item() -> Python float, frees the graph
if epoch % 10 == 0:
model.eval(); correct = 0
with torch.no_grad():
for xb, yb in test_dl:
correct += (model(xb.to(device)).argmax(1) == yb.to(device)).sum().item()
print(f"epoch {epoch:2d} train_loss {total/len(train_ds):.4f} test_acc {correct/len(test_ds):.3f}")
# epoch 10 train_loss 0.0947 test_acc 0.974
# epoch 20 train_loss 0.0661 test_acc 0.974
# epoch 30 train_loss 0.0536 test_acc 0.974 (exact numbers vary by version and hardware)
In real projects, evaluate on a separate validation split for early stopping and model selection, and keep the test set for one final measurement, exactly as with scikit-learn.
train() versus eval(), no_grad versus inference_mode
model.train() / model.eval()
- Switch layer behaviour
- Dropout: active in train, identity in eval
- BatchNorm: batch statistics in train, stored running statistics in eval
- Does not turn off gradient tracking
torch.no_grad() / torch.inference_mode()
- Switch gradient tracking
- No graph built: less memory, faster
inference_modeis stricter and slightly faster- Does not change dropout or BatchNorm
For inference you need both: model.eval() and with torch.no_grad():.
Saving and loading
torch.save({"model": model.state_dict(), "optimizer": opt.state_dict(),
"epoch": epoch, "config": {"d_in": 30}}, "checkpoint.pt")
ckpt = torch.load("checkpoint.pt", map_location="cpu", weights_only=False) # False: file also has epoch/config; only load trusted files
model = MLP(**ckpt["config"])
model.load_state_dict(ckpt["model"])
model.eval()
Save the state_dict (a dict of tensors) rather than pickling the whole model object, which ties the file to your exact class definitions and file paths. Recent PyTorch versions default torch.load(..., weights_only=True), which refuses non-tensor objects; a training checkpoint that also stores epoch or a config dict needs weights_only=False and must come from a trusted source. The safetensors format stores tensors without pickle and avoids that risk.
Useful extras
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0) # tame exploding gradients
sched = torch.optim.lr_scheduler.CosineAnnealingLR(opt, T_max=30) # call sched.step() each epoch
with torch.autocast(device_type="cuda", dtype=torch.bfloat16): # mixed precision on GPU
loss = loss_fn(model(xb), yb)
model = torch.compile(model) # graph compilation (PyTorch 2.x)
torch.manual_seed(0); np.random.seed(0); random.seed(0) # reproducibility
zero_grad() (gradients accumulate), forgetting model.eval() (dropout stays on, results fluctuate), applying softmax before CrossEntropyLoss (double softmax, poor training), float labels for CrossEntropyLoss or long labels for BCEWithLogitsLoss, int64 inputs to a Linear layer ("mat1 and mat2 must have the same dtype, but got Long and Float"), x.to(device) without reassigning, and accumulating total += loss instead of loss.item(), which keeps every graph alive and leaks GPU memory.eval() and no_grad(); why gradients accumulate; view versus reshape; detach(); what requires_grad does; how to freeze layers for fine-tuning (p.requires_grad = False and pass only trainable parameters to the optimizer); state_dict saving; and how you would debug a loss that does not decrease (overfit one small batch first). Deeper theory is in Deep Learning.Hugging Face basics: pipelines, tokenizers and Auto classes
The Hugging Face transformers library gives you thousands of pretrained models behind a uniform interface. Three layers of abstraction cover almost every use: pipeline() for a task in one line, AutoTokenizer to turn text into token IDs, and AutoModel... classes to load the right architecture for a checkpoint by name. The theory of what these models do (attention, tokenization, decoding) lives in Transformers & LLMs; this section is the Python mechanics.
A pipeline is a coffee machine with a "cappuccino" button: beans in, drink out, no need to know the grind size. The tokenizer is the grinder that turns beans (raw text) into a consistent powder (token IDs) that only fits this machine's filter. The Auto classes are a universal adapter that looks at the model's label and snaps in the right internal parts. Mapping back: pipeline("sentiment-analysis") hides tokenization, model forward pass and post-processing; using the tokenizer and model directly gives you control over batching, truncation and devices, and the tokenizer must always match its model.
from transformers import pipeline
clf = pipeline("sentiment-analysis") # downloads a default model on first use
print(clf(["Great battery life!", "Screen cracked in a week."]))
# [{'label': 'POSITIVE', 'score': 0.99...}, {'label': 'NEGATIVE', 'score': 0.99...}]
# other tasks: "zero-shot-classification", "ner", "question-answering",
# "summarization", "text-generation", "feature-extraction"
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
name = "distilbert-base-uncased-finetuned-sst-2-english"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForSequenceClassification.from_pretrained(name).eval()
batch = tok(["Great battery life!", "Screen cracked in a week."],
padding=True, truncation=True, max_length=128, return_tensors="pt")
print(batch.keys()) # input_ids, attention_mask
with torch.no_grad():
logits = model(**batch).logits # shape (2, 2)
print(logits.softmax(-1).argmax(-1).tolist(), model.config.id2label)
| Class | Returns | Typical use |
|---|---|---|
AutoModel | Hidden states (embeddings) | Feature extraction, custom heads |
AutoModelForSequenceClassification | Class logits per text | Sentiment, intent, topic |
AutoModelForTokenClassification | Logits per token | Named-entity recognition |
AutoModelForCausalLM | Next-token logits, .generate() | Text generation (GPT-style) |
AutoModelForSeq2SeqLM | Encoder-decoder generation | Translation, summarization |
- The
attention_maskmarks real tokens (1) versus padding (0) so padded positions are ignored. - Models are ordinary PyTorch
nn.Modules:.to(device),.eval(),no_gradand custom training loops all work. TheTrainerclass and thedatasetslibrary wrap the loop and data handling for fine-tuning. - Load large models with
torch_dtype=torch.bfloat16anddevice_map="auto"to save memory; parameter-efficient fine-tuning (LoRA via thepeftlibrary) trains small adapter matrices instead of all weights. - In restricted networks, download models through your organisation's approved mirror and load them from a local folder path with
from_pretrained("path/to/model").
truncation=True and crashing on inputs longer than the model's maximum length, and running a large model on CPU in float32 when a smaller distilled model would do.AutoModel and AutoModelForCausalLM, how to batch inference efficiently, and how you would fine-tune a classifier on a small labelled set (pretrained encoder plus classification head, low learning rate, few epochs, early stopping, or LoRA for large models).Calling LLM APIs from Python
Most LLM application code is ordinary Python: build a list of messages, send an HTTP request through an SDK, parse the response, handle errors and rate limits, and validate structured output. Many providers and self-hosted servers expose an OpenAI-compatible interface, so one client pattern works across them. Prompting, structured outputs, tool calling and evaluation are covered in depth in LLM APIs.
Calling an LLM API is like ordering at a busy restaurant by phone. You give your standing instructions (system message) and the order (user message), you are charged by the size of the order (tokens), the line is sometimes busy (rate limits), so you wait and redial politely (retry with backoff), and you check the delivery against the receipt before paying (validate the structured output). Mapping back: messages carry role-tagged text, usage fields report tokens, HTTP 429 errors signal rate limits, and schema validation catches malformed answers.
import os
from openai import OpenAI # the same client works with OpenAI-compatible servers
from pydantic import BaseModel, ValidationError
client = OpenAI(api_key=os.environ["LLM_API_KEY"], # never hard-code keys
base_url=os.environ.get("LLM_BASE_URL")) # your approved endpoint
resp = client.chat.completions.create(
model=os.environ.get("LLM_MODEL", "gpt-4o-mini"),
messages=[{"role": "system", "content": "You are a concise tutor."},
{"role": "user", "content": "Explain overfitting in one sentence."}],
temperature=0, max_tokens=80)
print(resp.choices[0].message.content, resp.usage.total_tokens)
class Review(BaseModel):
sentiment: str
rating: int
raw = '{"sentiment": "positive", "rating": 5}' # e.g. a JSON-mode response
try:
review = Review.model_validate_json(raw)
except ValidationError as e:
print("retry or fall back:", e)
- Keep keys in environment variables or a secrets manager; never commit them or print them in notebooks.
- Wrap calls with timeouts and exponential-backoff retries on rate-limit and transient server errors.
- Use
asynciowith a semaphore (see the concurrency section) to evaluate hundreds of prompts in parallel within rate limits. - Cache responses keyed by (model, prompt, parameters) during development to save cost and make runs repeatable.
- Log token usage and latency per call; cost scales with input plus output tokens.
eval, a shell or a SQL query. Also avoid sending confidential data to endpoints your organisation has not approved.Performance and best practices
Speed work follows one rule: measure first, then fix the biggest bottleneck. In data code the usual culprits are Python-level loops over rows, wrong dtypes that bloat memory, re-reading or re-computing the same data, and moving data between CPU and GPU or between processes more than necessary.
Optimising without profiling is like trying to shorten a commute by buying faster shoes when the real delay is a 40-minute wait at one traffic light. A profiler is the GPS log that shows exactly where the minutes go. Mapping back: cProfile, line_profiler and %timeit reveal the hot function or line, and fixing that one spot (usually by vectorizing it) beats micro-tuning everything else.
Measure
# notebook
%timeit df.a * df.b # repeated timing of one line
%prun train_features(df) # function-level profile
%load_ext line_profiler
%lprun -f build_features build_features(df) # line-by-line (needs line_profiler)
%memit df = pd.read_csv("big.csv") # peak memory (needs memory_profiler)
# script
import cProfile, pstats, tracemalloc
cProfile.run("main()", "prof.out")
pstats.Stats("prof.out").sort_stats("cumulative").print_stats(10)
tracemalloc.start(); run(); print(tracemalloc.get_traced_memory()) # (current, peak) bytes
# command line: python -m cProfile -s cumtime train.py ; py-spy for sampling live processes
The speed ladder (fastest first)
- Avoid the work Filter early, read fewer columns, cache intermediate results (Parquet,
joblib.Memory,lru_cache). - Vectorize NumPy ufuncs, pandas column operations,
np.where/np.select,groupbybuilt-ins such as"sum"rather than a Python lambda. - Use the right structure Set or dict for membership tests instead of a list;
collections.dequefor queues; sorted arrays withnp.searchsorted. - Compile hot loops Numba
@njitor Cython for loops that genuinely cannot be vectorized. - Parallelize
n_jobs=-1in scikit-learn,joblib.Parallel, process pools, DataLoader workers. - Change engine or hardware Polars or DuckDB for large tables, Dask or Spark for multi-machine data, GPU for large matrix work.
import numpy as np
from numba import njit
@njit # compiled to machine code on first call
def count_runs(x):
runs = 0
for i in range(1, x.size): # sequential dependency: hard to vectorize
if x[i] != x[i - 1]:
runs += 1
return runs
from joblib import Parallel, delayed
results = Parallel(n_jobs=-1)(delayed(score_file)(p) for p in paths)
Memory habits
- Use
float32for model inputs andcategoryfor repetitive strings; check withmemory_usage(deep=True). - Process in chunks or with generators when data exceeds RAM; write intermediate results to Parquet.
- Delete large temporaries (
del df_raw) and callgc.collect()in long notebooks; restart the kernel when in doubt. - Prefer in-place NumPy operations for huge arrays (
np.multiply(a, 2, out=a),a *= 2) to avoid temporary copies. - On GPU: smaller batches, mixed precision, gradient accumulation,
torch.no_grad()for evaluation, and.item()when logging scalars.
Code-quality habits that pay off in ML
Configuration
Keep hyperparameters in a dataclass or YAML file, not scattered literals. Log the config with every run.
Seeds and versions
Seed Python, NumPy and PyTorch; pin library versions; record the data snapshot.
Tests
Unit-test feature functions with tiny DataFrames (pytest); assert shapes, dtypes and no NaN after preprocessing.
Logging
Use logging instead of print in pipelines; track experiments with a tracking tool.
Style
PEP 8 naming, type hints, docstrings, a formatter and linter (for example black and ruff).
Notebooks to modules
Explore in notebooks, then move reusable code into a package with a command-line entry point.
apply adds pickling overhead and memory copies; vectorizing that one line usually gives a far larger speed-up on a single core. Also beware nested parallelism: n_jobs=-1 in both a grid search and the model inside it oversubscribes the CPU.apply/iterrows with vectorized operations or merges; fix dtypes; read only needed columns from Parquet; cache intermediate steps; then parallelize or switch engines if still needed. Quantify each step with before-and-after timings.Worked example: breast-cancer diagnosis end to end
This project ties every section together on the classic breast-cancer dataset: 569 tumours, 30 numeric features computed from cell-nucleus images, and a binary label. Missing a malignant tumour (a false negative) is far worse than an unnecessary follow-up (a false positive), so the primary metric is recall on the malignant class, with a sanity floor on precision. Every output below comes from running the code with scikit-learn 1.9 and pandas 3.0; small numeric differences across versions are normal.
Think of the project as a clinical trial. You recruit patients (load data), run baseline health checks (EDA), seal an envelope of patients nobody may look at until the end (the test set), try treatments on the rest with careful rotation so each is judged fairly (cross-validation), pick the best and set its dose (tuning and threshold), and only then open the sealed envelope once to report results. Mapping back: the sealed envelope is X_test, rotation is StratifiedKFold, the dose is the decision threshold, and opening the envelope twice would make the reported result untrustworthy.
Step 1: load and relabel
import numpy as np, pandas as pd
from sklearn.datasets import load_breast_cancer
data = load_breast_cancer(as_frame=True)
df = data.frame.copy()
df["target"] = 1 - df["target"] # raw data: 0 = malignant; now 1 = malignant (positive class)
print(df.shape) # (569, 31)
print(df["target"].value_counts())
# target
# 0 357 benign
# 1 212 malignant (about 37%: mildly imbalanced)
print(df.isna().sum().sum()) # 0 no missing values
Step 2: exploratory analysis
print(df[["mean radius", "mean area", "mean smoothness"]].describe().round(2))
# mean radius mean area mean smoothness
# count 569.00 569.00 569.00
# mean 14.13 654.89 0.10
# std 3.52 351.91 0.01
# min 6.98 143.50 0.05
# ... (quartile rows omitted)
# max 28.11 2501.00 0.16
# scales differ by 4 orders of magnitude -> scale for linear models, SVM, kNN
print(df.groupby("target")[["mean radius", "worst concave points"]].mean().round(3))
# mean radius worst concave points
# target
# 0 12.147 0.074
# 1 17.463 0.182 malignant tumours are larger and more irregular
import matplotlib.pyplot as plt, seaborn as sns
fig, axes = plt.subplots(1, 3, figsize=(13, 3.5))
sns.countplot(df, x="target", ax=axes[0])
sns.histplot(df, x="mean radius", hue="target", kde=True, ax=axes[1])
sns.heatmap(df.iloc[:, :10].corr(), cmap="coolwarm", center=0, ax=axes[2])
fig.tight_layout()
# the heatmap shows radius, perimeter and area almost perfectly correlated (they are
# geometrically linked): harmless for trees, handled by regularization in logistic regression
Step 3: split before anything learns from data
from sklearn.model_selection import train_test_split
X, y = df.drop(columns="target"), df["target"]
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42)
print(X_train.shape, X_test.shape, round(y_train.mean(), 3), round(y_test.mean(), 3))
# (455, 30) (114, 30) 0.374 0.368 stratification preserved the class ratio
Step 4: pipeline and cross-validated baseline
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, cross_val_score
pipe = Pipeline([("scale", StandardScaler()),
("clf", LogisticRegression(max_iter=5000))])
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(pipe, X_train, y_train, cv=cv, scoring="recall")
print(scores.round(3), round(scores.mean(), 3), round(scores.std(), 3))
# [0.882 0.971 1. 0.941 0.971] 0.953 0.04 fold-to-fold spread matters on small data
Step 5: tune hyperparameters on training data only
from sklearn.model_selection import GridSearchCV
grid = GridSearchCV(pipe,
{"clf__C": [0.01, 0.1, 1, 10, 100],
"clf__class_weight": [None, "balanced"]},
cv=cv, scoring="recall", n_jobs=-1)
grid.fit(X_train, y_train)
print(grid.best_params_, round(grid.best_score_, 3))
# {'clf__C': 1, 'clf__class_weight': 'balanced'} 0.959
Step 6: choose the decision threshold from out-of-fold predictions
from sklearn.model_selection import cross_val_predict
from sklearn.metrics import precision_recall_curve
best = grid.best_estimator_
oof = cross_val_predict(best, X_train, y_train, cv=cv, method="predict_proba")[:, 1]
prec, rec, thr = precision_recall_curve(y_train, oof)
ok = np.where(rec[:-1] >= 0.98)[0] # thresholds meeting the recall target
threshold = thr[ok[-1]] # the highest such threshold keeps precision best
print(round(threshold, 3), round(rec[ok[-1]], 3), round(prec[ok[-1]], 3))
# 0.284 0.982 0.923 on training folds: 98% recall at 92% precision
Step 7: open the sealed envelope once
from sklearn.metrics import classification_report, confusion_matrix, roc_auc_score, average_precision_score
proba = best.predict_proba(X_test)[:, 1]
y_pred = (proba >= threshold).astype(int)
print(confusion_matrix(y_test, y_pred))
# [[70 2] 70 benign correct, 2 false alarms
# [ 2 40]] 2 missed malignant, 40 caught
print(classification_report(y_test, y_pred, target_names=["benign", "malignant"], digits=3))
# precision recall f1-score support
# benign 0.972 0.972 0.972 72
# malignant 0.952 0.952 0.952 42
# accuracy 0.965 114
print(round(roc_auc_score(y_test, proba), 3), round(average_precision_score(y_test, proba), 3))
# 0.995 0.993
Test recall (0.952) is a little below the 0.98 seen on training folds. With only 42 malignant test cases, one tumour moves recall by 2.4 percentage points, so this gap is within noise; the honest report is "about 95 to 98% recall with roughly 95% precision", together with the fold standard deviation. If the clinical requirement is stricter, lower the threshold further and accept more false alarms, deciding on validation data, never by peeking at the test set again.
Step 8: interpret and save
coefs = pd.Series(best.named_steps["clf"].coef_[0], index=X.columns)
print(coefs.sort_values(key=abs, ascending=False).head(5).round(3))
# worst texture 1.466
# radius error 1.309
# worst symmetry 1.077
# mean concave points 0.970
# compactness error -0.961 coefficients are per standard deviation, because inputs were scaled
import joblib
joblib.dump({"model": best, "threshold": float(threshold), "features": list(X.columns)},
"breast_cancer_v1.joblib")
Because many features are strongly correlated, individual coefficients can shift sign or size between fits; use permutation importance or a regularized, reduced feature set before drawing medical conclusions. A PyTorch MLP on the same split (see the PyTorch section) reaches the same 0.974 accuracy at the default threshold, which is a useful lesson: on small tabular data, a well-tuned linear model is a very strong baseline.
scoring="recall" optimise recall for benign tumours instead.Quick revision
- Python is the orchestration layer; speed comes from keeping work inside compiled libraries (NumPy, BLAS, cuDNN) rather than Python loops.
- Use one isolated environment per project (venv or conda) with pinned versions; use
%pipinside notebooks. - Restart the kernel and run all cells before sharing a notebook, to expose hidden state.
- Variables are names bound to objects; assignment never copies.
- Lists, dicts and sets are mutable; ints, floats, strings, tuples and frozensets are immutable.
==compares values,iscompares identity; useisonly withNoneand other singletons.- Never use a mutable default argument; default to
Noneand create the object inside. - Dict and set lookups are O(1) on average; list membership is O(n). Keys must be hashable.
*argscollects extra positional arguments into a tuple;**kwargscollects extra keyword arguments into a dict.- A closure remembers variables from its enclosing scope; closures bind late, so loop variables are read at call time.
- A decorator is a function that takes a function and returns a wrapped one; always use
functools.wraps. - Generators (
yield) produce values lazily, use constant memory, and can be consumed only once. - Dunder methods (
__len__,__getitem__,__call__,__enter__) plug your classes into Python syntax and into frameworks. - Type hints are not enforced at runtime; tools such as mypy and libraries such as Pydantic read them.
withguarantees cleanup via__exit__, even when an exception is raised.- The GIL allows one thread to run Python bytecode at a time: threads or asyncio for I/O-bound work, processes for CPU-bound Python work.
- A full PyTorch checkpoint (state_dict plus epoch or config) needs
torch.load(..., weights_only=False)and must be trusted; tensor-only loads can keep the safer default. - NumPy arrays are homogeneous, typed and contiguous, which enables vectorized C loops.
- Broadcasting aligns shapes from the right; dimensions must match or be 1.
- Basic slicing returns a view; fancy and boolean indexing return copies.
axis=0collapses rows (one result per column);axis=1collapses columns (one result per row).*is element-wise,@is matrix multiplication.- Subtract the maximum before exponentiating in softmax for numerical stability.
- NumPy's
var/stddefault toddof=0; pandas defaults toddof=1. locselects by label with an inclusive end;ilocselects by position with an exclusive end.- Combine pandas masks with
&,|,~and parentheses. groupby().agg()returns one row per group;transform()returns a result aligned to the original rows.- Validate merges (
validate=, row counts) to catch duplicate keys that explode joins. - Vectorized column operations beat
apply, which beatsiterrows. - Assign with a single
df.loc[mask, col] = value; chained assignment does not modify the original. - Shrink pandas memory with downcasting,
categorydtype,usecols, Parquet and chunking. - EDA checklist: structure, quality, univariate, target, bivariate, multivariate, leakage, decisions.
- scikit-learn contract:
fitlearns,transform/predictapply; learned attributes end with an underscore. - Put every data-dependent step inside a
Pipelineso cross-validation cannot leak. - Use CV to choose models; use a paired bootstrap or McNemar on a sealed test set to compare two finished models.
ColumnTransformerapplies different preprocessing to numeric and categorical columns.- Stratify classification splits; use group splits for repeated entities and time-series splits for temporal data.
- Tune on training data with CV; touch the test set once at the end.
- Choose metrics by error cost: recall when misses are costly, precision when false alarms are costly, PR-AUC for rare positives.
- Pick the decision threshold on validation or out-of-fold predictions, not the test set.
- Scale features for linear models, SVM, kNN, PCA and neural networks; trees do not need it.
- PyTorch loop:
zero_grad, forward, loss,backward,step. model.eval()changes dropout and BatchNorm behaviour;torch.no_grad()stops gradient tracking; inference needs both.CrossEntropyLosstakes raw logits and integer class labels;BCEWithLogitsLosstakes logits and float targets.- Save
state_dicts;weights_only=Trueis the safe default, but a checkpoint that also stores epoch or config needsweights_only=Falsefrom a trusted file. - Hugging Face:
pipelinefor quick tasks,AutoTokenizerplusAutoModelFor...for control; tokenizer and model must match. - Profile before optimising; vectorize first, compile or parallelize second, change engine last.
Glossary
- apply
- pandas method that runs a Python function per element, row or group; flexible but slow compared with vectorized operations.
- asyncio
- Python's standard library for single-threaded concurrency using
async/await, ideal for many simultaneous network calls. - Autograd
- PyTorch's automatic differentiation engine, which records operations and computes gradients with the chain rule on
backward(). - Axis
- A dimension of an array; the
axisargument names the dimension an aggregation collapses. - Broadcasting
- NumPy and PyTorch rules for combining arrays of different shapes by virtually stretching size-1 dimensions.
- Categorical dtype
- pandas type storing each distinct value once plus integer codes per row, saving memory for repetitive values.
- Closure
- A function that remembers variables from the scope in which it was defined.
- ColumnTransformer
- scikit-learn tool that applies different transformers to different column subsets and concatenates the results.
- Context manager
- Object used with
withthat sets something up in__enter__and guarantees cleanup in__exit__. - Copy-on-Write
- pandas behaviour (default from version 3.0) where derived objects act as copies, so modifying one never silently changes another.
- Cross-validation
- Repeatedly splitting training data into training and validation folds to estimate performance more reliably than a single split.
- Data leakage
- Information unavailable at prediction time (often from the test set or the target) sneaking into training, inflating reported performance.
- Dataclass
- Class decorated with
@dataclass, which auto-generates__init__,__repr__and__eq__from annotated fields. - DataFrame
- pandas two-dimensional labelled table made of columns (Series) sharing one index.
- DataLoader
- PyTorch utility that batches, shuffles and optionally loads data in parallel worker processes from a Dataset.
- Decorator
- A callable that takes a function and returns a modified or wrapped function, applied with
@name. - dtype
- The data type of array or column elements, such as
float32,int64orcategory. - Dunder method
- A "double underscore" special method like
__len__that Python calls to implement built-in syntax and functions. - EDA
- Exploratory data analysis: systematic inspection of data quality, distributions and relationships before modelling.
- Estimator
- Any scikit-learn object that learns from data via
fit. - Fancy indexing
- Indexing an array with lists or arrays of integers; always returns a copy.
- Generator
- A function using
yieldthat produces values lazily, one at a time, suspending between values. - GIL
- Global Interpreter Lock: a CPython mutex allowing only one thread to execute Python bytecode at a time per process.
- GridSearchCV
- scikit-learn search that cross-validates every combination of a hyperparameter grid and refits the best.
- groupby
- pandas split-apply-combine operation that groups rows by keys and aggregates, transforms or filters each group.
- Hashable
- An object with a stable hash and equality, which can therefore be a dict key or set member.
- iloc
- pandas position-based indexer; slice ends are exclusive.
- Iterator
- Object implementing
__next__(and__iter__) that yields successive values and raisesStopIterationwhen done. - Jupyter kernel
- The live Python process behind a notebook that keeps state between cell executions.
- loc
- pandas label-based indexer that also accepts boolean masks; slice ends are inclusive.
- Logits
- Raw, unnormalised model outputs before softmax or sigmoid.
- melt
- pandas operation that reshapes a wide table into a long one.
- merge
- pandas SQL-style join of two DataFrames on key columns.
- MRO
- Method resolution order: the sequence Python searches through base classes to find a method.
- Mutable
- Able to be changed in place after creation (lists, dicts, sets, arrays).
- ndarray
- NumPy's n-dimensional, homogeneous, typed array.
- nn.Module
- Base class for PyTorch layers and models, which registers parameters and submodules and defines
forward. - no_grad
- PyTorch context that disables gradient tracking, used for inference and manual weight updates.
- One-hot encoding
- Representing a category as a vector with a single 1 in the column for that category.
- Parquet
- Compressed, columnar file format that preserves dtypes and allows reading only selected columns.
- Pipeline
- scikit-learn object chaining transformers and a final estimator so they are fitted and applied together.
- pivot_table
- pandas reshaping that aggregates values into a grid of index versus column categories.
- requires_grad
- PyTorch tensor flag indicating that operations on it should be recorded for gradient computation.
- Series
- pandas one-dimensional labelled array.
- state_dict
- Dictionary mapping a PyTorch model's or optimizer's parameter names to tensors; the recommended save format.
- Stratification
- Splitting data so each part keeps the same class proportions as the whole.
- Tensor
- PyTorch's n-dimensional array, which can live on CPU or GPU and can track gradients.
- Tokenizer
- Component that converts text into integer token IDs (and back) for a specific language model.
- transform (groupby)
- pandas group operation that returns a result aligned to the original rows, such as each row's share of its group total.
- Transformer (scikit-learn)
- An estimator with a
transformmethod, such as a scaler, encoder or imputer. - ufunc
- NumPy universal function that operates element-wise on arrays in compiled code.
- Vectorization
- Expressing computation as whole-array operations instead of explicit Python loops.
- View
- An array object that shares memory with another array, created for example by basic slicing.
- Virtual environment
- An isolated folder with its own Python interpreter and packages for one project.
Interview questions
Fundamentals
What is the difference between a list and a tuple, and when would you use each?
Both are ordered sequences. A list is mutable (append, remove, assign items); a tuple is immutable. Because tuples cannot change, they are hashable when their contents are hashable, so they can be dict keys or set members, and they signal "fixed record" intent, such as a tensor shape (32, 3, 224, 224) or a function returning several values. Tuples are also slightly smaller and faster to create. Use a list for collections that grow or change, such as a loss history.
Which Python built-in types are mutable and which are immutable?
Mutable: list, dict, set, bytearray, and most user-defined objects (also NumPy arrays and DataFrames). Immutable: int, float, complex, bool, str, bytes, tuple, frozenset, None. Note that a tuple containing a list is itself immutable (you cannot swap the list out) but the inner list can still change, and such a tuple is not hashable.
What is the difference between is and ==?
== checks value equality by calling __eq__. is checks identity: whether both names refer to the very same object. [1, 2] == [1, 2] is True but [1, 2] is [1, 2] is False. Use is for singletons like None, True, False. CPython caches small integers and some strings, so is sometimes appears to work on them, but relying on that is a bug.
What does this print, and why? a = [1, 2]; b = a; b.append(3); print(a)
It prints [1, 2, 3]. Assignment binds a second name to the same list object; it does not copy. Appending through b mutates the one shared list. To get an independent list use a.copy(), list(a) or a[:] (shallow), or copy.deepcopy(a) for nested structures.
Explain shallow copy versus deep copy.
A shallow copy creates a new outer container but fills it with references to the same inner objects; a deep copy recursively copies everything.
import copy
grid = [[0, 0], [0, 0]]
s = grid.copy(); d = copy.deepcopy(grid)
grid[0][0] = 9
print(s[0][0], d[0][0]) # 9 0With pandas, df.copy() is deep by default for the data; with NumPy, arr.copy() duplicates the buffer.
What is the mutable default argument trap?
Default values are evaluated once, when the def statement runs, not on each call. A mutable default is therefore shared across calls:
def add(x, acc=[]):
acc.append(x); return acc
add(1); print(add(2)) # [1, 2]
def add(x, acc=None): # fix
if acc is None:
acc = []
acc.append(x); return accWhat are *args and **kwargs?
In a function definition, *args gathers extra positional arguments into a tuple and **kwargs gathers extra keyword arguments into a dict. At a call site, *seq and **mapping unpack them back into arguments. They are used to write flexible wrappers and to forward options, for example a decorator's wrapper(*args, **kwargs) or model(**tokenizer_output). The names are convention; the stars are what matter.
What is a list comprehension, and when should you not use one?
A compact expression that builds a list from an iterable with optional filtering: [x * x for x in nums if x % 2 == 0]. It is usually faster and clearer than a loop with append. Avoid it when the logic needs several nested conditions or side effects (printing, writing files), when you do not need the list (use a generator expression, for example inside sum()), or when the data is numeric and large (use NumPy vectorization).
What is the difference between a list comprehension and a generator expression?
[x * x for x in data] builds the whole list in memory immediately. (x * x for x in data) creates a lazy generator that computes each value when requested, using constant memory, but it can be iterated only once and has no len() or indexing. For a million items the list takes about 8 MB while the generator object takes around 200 bytes.
What is a lambda function?
An anonymous single-expression function: lambda r: r["f1"]. Common as a key= for sorted/max, in pandas assign, or as a quick callback. It cannot contain statements. For anything longer, or anything you need to test or pickle (multiprocessing cannot pickle lambdas), define a named function.
How do you handle exceptions in Python? What do else and finally do?
Use try/except SpecificError. The else block runs only if no exception occurred (keep the code that may fail in try small, and put follow-up code in else). finally always runs, for cleanup. Catch specific exceptions, never a bare except:, and re-raise with raise NewError(...) from e to keep the cause. Define custom exceptions by subclassing Exception or a more specific built-in.
What is a virtual environment and why use one?
An isolated directory containing its own interpreter link and installed packages, created with python -m venv .venv (or conda). It prevents version conflicts between projects, makes dependencies explicit (pip freeze > requirements.txt), and allows reproducing the environment on another machine or in CI. conda additionally manages non-Python binaries such as CUDA libraries.
What does if __name__ == "__main__": do?
When a file is run directly, its __name__ is "__main__"; when imported, it is the module name. The guard lets a file be both an importable module and a script. It is essential on Windows and macOS for multiprocessing and PyTorch DataLoader workers, which start new processes that re-import the main module; without the guard the process-creating code would run again in every child.
What is NumPy and why is it faster than Python lists?
NumPy provides the ndarray: a contiguous block of same-typed values plus shape and stride metadata. Operations run in compiled C loops (often using SIMD instructions and optimized BLAS) with no per-element type checks or object overhead. A Python list stores pointers to separately allocated objects, and each operation goes through the interpreter. Vectorized NumPy code is typically 10 to 100 times faster and uses far less memory.
What do shape, ndim, dtype and size tell you about an array?
shape is the length along each dimension, such as (3, 4); ndim is the number of dimensions (2); dtype is the element type (int64, float32); size is the total number of elements (12). itemsize is bytes per element and nbytes is total bytes. Printing shapes is the first debugging step for most ML bugs.
What is the difference between * and @ for NumPy arrays?
* is element-wise multiplication (shapes must match or broadcast). @ (or np.matmul/np.dot for 2-D) is matrix multiplication: (n, k) @ (k, m) gives (n, m). For [[1,2],[3,4]], X * X is [[1,4],[9,16]] and X @ X is [[7,10],[15,22]].
What does axis=0 versus axis=1 mean?
The axis is the dimension that gets collapsed. For a 2-D array of shape (rows, columns), sum(axis=0) adds down the rows and returns one value per column; sum(axis=1) adds across the columns and returns one value per row. In pandas the same logic applies: df.mean() (axis 0) gives column means, and df.drop(columns=...) is the readable alternative to axis=1.
What is the difference between a pandas Series and a DataFrame?
A Series is a one-dimensional labelled array with one dtype. A DataFrame is a two-dimensional table: an ordered collection of Series (columns) sharing one row index, where each column may have a different dtype. Selecting one column with df["col"] returns a Series; df[["col"]] returns a one-column DataFrame.
What is the difference between loc and iloc?
loc selects by labels (index values and column names) and boolean masks, and its slices include the end label. iloc selects by integer position, and its slices exclude the end, like Python. With the default RangeIndex, df.loc[0:2] returns three rows and df.iloc[0:2] returns two. After filtering or sorting, labels and positions no longer coincide, which is where bugs appear.
How do you check for and handle missing values in pandas?
Detect with df.isna().sum() or df.isna().mean() for fractions. Options: drop rows or columns (dropna, with subset or thresh), fill with a constant, mean, median or mode (fillna), fill within groups (groupby().transform("median")), forward-fill for time series, or model-based imputation. Adding a "was missing" indicator often helps. For ML, do the imputation inside a scikit-learn pipeline so statistics are learned from training folds only.
How do you filter rows by multiple conditions in pandas?
Build boolean masks and combine them with & (and), | (or), ~ (not), wrapping each condition in parentheses because these operators bind tighter than comparisons:
df[(df.salary > 90) & (df.dept == "ml")]
df[df.city.isin(["Pune", "Delhi"]) & ~df.name.str.startswith("A")]
df.query("salary > 90 and dept == 'ml'")Using Python's and/or raises "The truth value of a Series is ambiguous".
What does groupby do?
It implements split-apply-combine: split rows into groups by one or more keys, apply a function to each group (aggregate, transform, or filter), and combine the results. Example: sales.groupby("city")["sales"].sum() gives total sales per city. Named aggregation, agg(total=("sales", "sum"), avg=("units", "mean")), gives clean column names.
What is the difference between merge, join and concat?
pd.merge is a SQL-style join on key columns with how = inner, left, right, outer or cross. DataFrame.join is a convenience method that joins on the index by default. pd.concat stacks objects along an axis: rows (axis=0, same columns) or side by side (axis=1, aligned on index), with no key matching beyond index alignment.
What is scikit-learn's estimator API?
Every estimator is configured through its constructor and learns with fit(X, y), which returns self. Transformers add transform and fit_transform; predictors add predict, and most classifiers add predict_proba or decision_function; models have score. Learned attributes end in an underscore (coef_, mean_). get_params/set_params make every estimator tunable by generic tools like GridSearchCV.
What is the difference between fit, transform and fit_transform?
fit learns parameters from data (a scaler learns means and standard deviations). transform applies those learned parameters to data. fit_transform does both in one step on the same data, sometimes more efficiently. Use fit_transform on training data only; call transform on validation, test and production data so they are processed with training statistics.
Why do we split data into training and test sets, and what does stratify do?
The test set estimates performance on unseen data; if the model or any preprocessing sees it, the estimate becomes optimistic. train_test_split(X, y, test_size=0.2, stratify=y, random_state=42) keeps the class proportions equal in both parts, which matters for imbalanced or small datasets (for example, 37% malignant in both the training and test sets). random_state makes the split reproducible.
Why is feature scaling needed, and for which models?
Models based on distances (kNN, k-means, SVM with RBF kernel), on gradients (logistic regression, neural networks) or on variance (PCA) are dominated by large-scale features if features are not scaled; optimizers also converge faster on standardized inputs. Tree-based models split on thresholds one feature at a time, so they are scale-invariant. On the breast-cancer data, scaling lifts kNN from about 0.91 to 0.96 accuracy.
What is a PyTorch tensor and how does it differ from a NumPy array?
Both are n-dimensional typed arrays with similar operations. Tensors additionally can live on GPUs or other accelerators (.to("cuda")) and can record operations for automatic differentiation (requires_grad=True). Defaults differ: PyTorch uses float32 while NumPy uses float64. torch.from_numpy and .numpy() share memory on CPU.
What are the five steps of a PyTorch training iteration?
optimizer.zero_grad() to clear old gradients; logits = model(xb) forward pass; loss = loss_fn(logits, yb); loss.backward() to compute gradients; optimizer.step() to update parameters. Wrap the epoch with model.train(), and validation with model.eval() plus torch.no_grad().
What is Jupyter, and what are its main risks?
An interactive notebook environment where code cells run in a persistent kernel, mixing code, output, plots and prose; ideal for exploration and teaching. Risks: hidden state and out-of-order execution (results depend on which cells ran when), poor version-control diffs, and code that is hard to test or reuse. Mitigate with "restart and run all", moving reusable code into modules, and tools that strip outputs or pair notebooks with scripts.
How would you compute the mean, median, mode, variance and standard deviation in Python?
By hand you can use sum(x)/len(x), sort for the median, and a counting dict for the mode, but in practice use the libraries:
import numpy as np, pandas as pd, statistics
d = [22, 18, 14, 10, 15, 20, 25, 12]
np.mean(d), np.median(d) # 17.0 16.5
statistics.multimode([1, 2, 2, 3]) # [2]
np.var(d), np.std(d) # 23.25 4.82 (population, ddof=0)
pd.Series(d).var() # 26.57 (sample, ddof=1)
np.percentile(d, [25, 75]) # quartilesState which variance formula you use; the sample version divides by n − 1.
Going deeper
What is a decorator? Write one that times a function.
A decorator is a callable that takes a function and returns a new function, usually wrapping the original with extra behaviour. @timer above def f means f = timer(f).
import time, functools
def timer(func):
@functools.wraps(func)
def wrapper(*args, **kwargs):
t0 = time.perf_counter()
try:
return func(*args, **kwargs)
finally:
print(f"{func.__name__}: {time.perf_counter() - t0:.3f}s")
return wrapperfunctools.wraps preserves the name and docstring; *args, **kwargs forwards any signature; the result is returned.
How do you write a decorator that takes arguments?
Add one more level: a factory that receives the arguments and returns the actual decorator.
def retry(times=3):
def decorator(func):
@functools.wraps(func)
def wrapper(*args, **kwargs):
for i in range(times):
try:
return func(*args, **kwargs)
except ConnectionError:
if i == times - 1:
raise
return wrapper
return decorator
@retry(times=5)
def fetch(): ...What is a closure, and what is the late-binding gotcha?
A closure is an inner function that keeps access to variables of its enclosing function after that function has returned (for example, a learning-rate scheduler remembering base_lr). Closures look up variables when called, not when defined, so [lambda: i for i in range(3)] gives three functions that all return 2. Fix by binding a default: lambda i=i: i, or use functools.partial. Use nonlocal to reassign an enclosing variable.
Explain the LEGB scope rule.
Python resolves names by searching Local (current function), Enclosing (outer functions), Global (module), then Built-in scope. Assigning to a name inside a function makes it local for the whole function, which is why reading a global and then assigning to it in the same function raises UnboundLocalError. Use global or nonlocal to rebind outer names, though passing values explicitly is usually cleaner.
What is the difference between an iterable, an iterator and a generator?
An iterable has __iter__ returning an iterator (lists, dicts, files). An iterator has __next__, returns successive values, raises StopIteration when done, and is itself iterable. A generator is an iterator created by a function containing yield (or a generator expression); Python saves its frame between values. A list can be iterated many times; an iterator or generator is exhausted after one pass.
Write a generator that yields mini-batches from a dataset.
import numpy as np
def minibatches(X, y, batch_size, shuffle=True, seed=0):
idx = np.arange(len(X))
if shuffle:
np.random.default_rng(seed).shuffle(idx)
for start in range(0, len(X), batch_size):
b = idx[start:start + batch_size]
yield X[b], y[b]
for xb, yb in minibatches(np.arange(10).reshape(5, 2), np.arange(5), 2, shuffle=False):
print(xb.shape) # (2, 2) (2, 2) (1, 2)It holds only one batch at a time, handles the final partial batch, and re-shuffles per epoch if you pass a new seed.
What are __str__ and __repr__, and how do they differ?
__repr__ is the unambiguous developer representation, ideally valid code to recreate the object (Neuron(weights=[0.5], bias=0.1)); it is used by the REPL, containers and debuggers. __str__ is the friendly end-user representation used by print and str(); if absent, Python falls back to __repr__. Implement at least __repr__.
What is the difference between @staticmethod, @classmethod and an instance method?
An instance method receives the instance (self) and can use its state. A classmethod receives the class (cls), commonly used for alternative constructors such as Config.from_yaml(path) or AutoModel.from_pretrained(name), and works correctly for subclasses. A staticmethod receives neither; it is a plain function namespaced inside the class for organisation.
What does super() do, and what is the MRO?
super() returns a proxy that delegates to the next class in the method resolution order, so super().__init__() runs the parent initializer. The MRO (C3 linearization, visible in Cls.__mro__) defines a consistent order across multiple inheritance so each base is visited once. In PyTorch, super().__init__() in an nn.Module subclass sets up parameter registries and must be called before assigning layers.
What is a context manager, and how do you write one?
An object used with with whose __enter__ runs at the start and __exit__ at the end, even if an exception occurs; __exit__ receives exception details and can suppress the exception by returning True. Write one as a class, or with @contextlib.contextmanager around a generator that does setup, yields, then cleans up in finally. Examples: files, locks, torch.no_grad(), temporary pandas options.
When would you use a dataclass, a namedtuple, a dict or a Pydantic model?
A dict for loose, dynamic data such as parsed JSON. A namedtuple for lightweight immutable records with positional access. A @dataclass for typed, readable configuration or record objects with defaults, optional immutability (frozen=True) and generated methods, without runtime validation. A Pydantic model when data crosses a trust boundary (API requests, LLM output) and must be validated and converted at runtime.
What is the GIL, and how does it affect ML code?
The Global Interpreter Lock lets only one thread execute Python bytecode at a time in a CPython process, which simplifies memory management (reference counting). Consequences: threads do not speed up CPU-bound pure-Python code, but they do help I/O-bound code because waiting threads release the GIL. NumPy, pandas, scikit-learn and PyTorch release the GIL inside heavy C routines and often multithread internally. For CPU-bound Python, use processes (multiprocessing, joblib, DataLoader workers). An optional free-threaded CPython build exists in recent versions but is not yet the default.
Threads, processes or asyncio: which would you choose for (a) 5,000 API calls, (b) parsing 200 large log files with pure Python, (c) matrix multiplication?
(a) asyncio with an async HTTP client and a semaphore to respect rate limits, or a thread pool if only a synchronous client exists. (b) A process pool, since parsing is CPU-bound Python and processes bypass the GIL. (c) Neither: call NumPy or PyTorch, which already use optimized multithreaded BLAS or the GPU.
State NumPy's broadcasting rules and give the result shape of (8, 1, 6, 1) combined with (7, 1, 5).
Align shapes from the right; prepend 1s to the shorter shape; each dimension pair must be equal or contain a 1; the result takes the larger size. (8, 1, 6, 1) and (1, 7, 1, 5) give (8, 7, 6, 5). Shapes (3,) and (4,) are incompatible and raise "operands could not be broadcast together".
What is the difference between a view and a copy in NumPy? How can you tell?
A view shares memory with the original, so writes through it change the original; basic slicing, reshape (when possible), transpose and ravel (when possible) return views. Fancy indexing with integer lists, boolean masks, flatten and .copy() return copies. Check with np.shares_memory(a, b) or b.base is a.
b = np.arange(6); v = b[1:4]; v[0] = 99
print(b) # [ 0 99 2 3 4 5]Implement a numerically stable softmax in NumPy that works on a batch.
def softmax(z, axis=-1):
z = z - z.max(axis=axis, keepdims=True)
e = np.exp(z)
return e / e.sum(axis=axis, keepdims=True)
softmax(np.array([1000., 1001., 1002.])) # [0.09 0.245 0.665] (naive exp overflows to inf)Subtracting the maximum does not change the result (it cancels in the ratio) but keeps exponents at or below zero. keepdims=True keeps shapes broadcastable for 2-D batches.
Standardize each column of a matrix and compute pairwise Euclidean distances without loops.
Xs = (X - X.mean(axis=0)) / X.std(axis=0) # broadcasting (n, d) with (d,)
# pairwise distances between A (n, d) and B (m, d)
D = np.sqrt(((A[:, None, :] - B[None, :, :]) ** 2).sum(-1)) # (n, m); memory n*m*d
# memory-lean version using |a-b|^2 = |a|^2 + |b|^2 - 2ab
D2 = (A**2).sum(1)[:, None] + (B**2).sum(1)[None, :] - 2 * A @ B.T
D = np.sqrt(np.maximum(D2, 0)) # clip tiny negatives from roundingCompute cosine similarity between a query vector and every row of a matrix in NumPy.
def cosine_sim(q, M):
q = q / np.linalg.norm(q)
M = M / np.linalg.norm(M, axis=1, keepdims=True)
return M @ q # shape (n,)
top3 = np.argsort(-cosine_sim(q, M))[:3] # or np.argpartition for large nThis is the core of embedding search in retrieval systems. Guard against zero-norm rows by adding a small epsilon.
What is the difference between agg, transform and apply after a groupby?
agg reduces each group to one row (sum, mean, custom). transform returns a result the same length as the input, aligned to the original rows, ideal for features like "share of group total" or "fill with group median". apply accepts any function returning a scalar, Series or DataFrame; it is the most flexible and the slowest. Prefer built-in string names ("sum", "mean") which run in optimized code.
df["share"] = df.sales / df.groupby("city").sales.transform("sum")How do you get the top N rows per group in pandas?
# top 2 products by revenue within each region
(df.sort_values("revenue", ascending=False)
.groupby("region")
.head(2))
# or with ranking, keeping ties explicit
df[df.groupby("region").revenue.rank(method="first", ascending=False) <= 2]
# single best row per group
df.loc[df.groupby("region").revenue.idxmax()]Explain pivot, pivot_table and melt.
melt converts wide to long: identifier columns stay, other columns become (variable, value) pairs, the format most plotting and modelling tools prefer. pivot converts long to wide and fails if an index/column pair is duplicated. pivot_table also goes long to wide but aggregates duplicates with aggfunc and can add totals (margins=True), like a spreadsheet pivot table.
Why is df.apply(..., axis=1) slow, and what do you use instead?
It calls a Python function once per row and constructs a Series object for each row, so it runs at interpreter speed with heavy object overhead. Replace with vectorized column arithmetic (df.a / df.b ** 2), np.where/np.select for conditionals, .map(dict) for lookups, .str/.dt accessors, or a merge for table lookups. If a loop is truly unavoidable, use itertuples, or compile the logic with Numba.
What causes SettingWithCopyWarning, and how does Copy-on-Write change things?
Chained indexing such as df[df.x > 0]["y"] = 1 first creates an intermediate object that might be a view or a copy, then assigns into it, so the original may or may not change. Older pandas warns. Under Copy-on-Write (the default behaviour from pandas 3.0), every derived object behaves as a copy, so chained assignment never updates the original and pandas raises a chained-assignment warning. The fix in all versions: df.loc[df.x > 0, "y"] = 1, and .copy() when you deliberately want an independent subset.
How do you reduce the memory usage of a large DataFrame?
- Measure with
df.memory_usage(deep=True). - Load only needed columns (
usecols) and specifydtypeat read time. - Downcast numerics:
pd.to_numeric(s, downcast="integer"),float64tofloat32. - Convert low-cardinality strings to
category. - Store as Parquet; read in chunks; filter early; delete temporaries.
A 1,000-row frame with an int64 and a float64 column drops from about 16 KB to 6 KB after downcasting to int16 and float32.
How do you create lag and rolling features for time-series data without leakage?
df = df.sort_values(["user", "date"])
g = df.groupby("user").amount
df["lag_1"] = g.shift(1) # yesterday's value
df["roll_7"] = g.transform(lambda s: s.shift(1).rolling(7).mean()) # past 7 days, excluding today
df["pct_change"] = g.pct_change()Shift before rolling so the current row's target period is excluded, compute per entity, never use center=True or negative shifts, and validate with TimeSeriesSplit.
What is a scikit-learn Pipeline, and why is it important?
A Pipeline chains transformers and a final estimator into one estimator. Calling fit fits each step on the output of the previous one; predict passes new data through the fitted steps. Benefits: no leakage in cross-validation (each fold refits preprocessing on its training part), one object to tune (step__param), one object to save and deploy, and less glue code.
How would you preprocess a dataset with numeric and categorical columns in scikit-learn?
pre = ColumnTransformer([
("num", make_pipeline(SimpleImputer(strategy="median"), StandardScaler()), num_cols),
("cat", make_pipeline(SimpleImputer(strategy="most_frequent"),
OneHotEncoder(handle_unknown="ignore")), cat_cols),
])
model = make_pipeline(pre, LogisticRegression(max_iter=1000))handle_unknown="ignore" keeps prediction working when a new category appears. For high-cardinality categories consider TargetEncoder or hashing; for tree models OrdinalEncoder is often enough.
What is cross-validation, and how do you choose the splitter?
k-fold cross-validation trains on k − 1 folds and validates on the remaining fold, k times, and averages the scores, giving a lower-variance estimate than a single split and a spread (std) to judge stability. Use StratifiedKFold for classification, GroupKFold when multiple rows belong to the same patient or user, TimeSeriesSplit for temporal data, and repeated CV for very small datasets.
GridSearchCV versus RandomizedSearchCV: how do they work and which do you choose?
Grid search evaluates every combination of listed values with cross-validation, then refits the best on the full training data (best_estimator_). Randomized search samples a fixed number of combinations from distributions (for example loguniform(1e-3, 1e3) for C). Random search is more efficient when only a few hyperparameters matter or the space is large; grid search is fine for small, discrete grids. Bayesian tools such as Optuna and successive-halving searches go further.
Precision versus recall: define them and give a case where each matters most.
Precision = TP / (TP + FP): of predicted positives, how many are right. Recall = TP / (TP + FN): of actual positives, how many were found. Recall matters most when misses are costly (cancer screening, fraud, safety defects). Precision matters most when false alarms are costly (spam filters deleting real mail, alerts that wake an engineer). F1 is their harmonic mean. Moving the decision threshold trades one against the other.
ROC-AUC versus PR-AUC: when is each appropriate?
ROC-AUC measures how well scores rank positives above negatives across all thresholds, plotting true-positive rate against false-positive rate; it is insensitive to class balance, which can make it look excellent when positives are rare. PR-AUC (average precision) plots precision against recall and focuses on the positive class, so it better reflects performance under heavy imbalance. The random baseline for ROC-AUC is 0.5; for PR-AUC it equals the positive rate.
What is the difference between model.eval() and torch.no_grad()?
model.eval() sets a flag on every module that changes layer behaviour: dropout stops dropping, and BatchNorm uses running statistics instead of batch statistics. It does not stop gradient tracking. torch.no_grad() (or the stricter inference_mode()) stops building the autograd graph, saving memory and time, but does not change layer behaviour. Inference needs both; remember model.train() before the next training epoch.
Why do you need optimizer.zero_grad()?
PyTorch accumulates gradients into .grad on every backward() call rather than overwriting them. That enables gradient accumulation across micro-batches, but in a normal loop, without zeroing, each step would use the sum of all previous gradients and training would diverge. zero_grad(set_to_none=True), the default in recent versions, sets gradients to None for a small speed and memory gain.
How do you write a custom PyTorch Dataset?
Subclass torch.utils.data.Dataset and implement __len__ (number of samples) and __getitem__(i) (return one sample, typically a (features, label) tuple of tensors). Do expensive loading lazily in __getitem__ for large data (read an image file per index), and let DataLoader handle batching, shuffling and parallel workers. For streams, subclass IterableDataset and implement __iter__.
Which loss function and target format do you use for binary, multiclass and multi-label classification in PyTorch?
Multiclass: nn.CrossEntropyLoss on raw logits of shape (N, C) with long targets of shape (N,). Binary: either one logit with nn.BCEWithLogitsLoss and float targets, or two logits with cross-entropy. Multi-label: one logit per label with BCEWithLogitsLoss and float 0/1 targets of shape (N, L). Never apply softmax or sigmoid before these losses; they include it in a numerically stable way.
What does a Hugging Face tokenizer return, and why do you need an attention mask?
Calling the tokenizer returns a dict-like object with input_ids (integer token IDs, including special tokens like the start and separator tokens), attention_mask (1 for real tokens, 0 for padding), and for some models token_type_ids. When a batch is padded to equal length, the mask tells the model to ignore padding positions so they do not affect attention or pooled outputs. Use padding=True, truncation=True, return_tensors="pt".
Advanced
How does CPython manage memory?
Every object carries a reference count; when it drops to zero the object is freed immediately. Reference counting cannot free cycles (A refers to B, B refers to A), so a generational cyclic garbage collector (gc module) periodically finds and frees unreachable cycles. Small objects come from a dedicated allocator (pymalloc) that keeps memory pools, which is why a Python process may not return memory to the operating system after freeing objects. Large NumPy and PyTorch buffers are allocated separately; GPU memory in PyTorch goes through a caching allocator, so nvidia-smi can show memory as used even after tensors are freed (torch.cuda.empty_cache() releases the cache).
How are Python dicts implemented, and why are lookups O(1) on average?
A dict is a hash table: the key's hash selects a slot, and collisions are resolved by open addressing (probing other slots). Since Python 3.6 the layout is a compact array of entries in insertion order plus a sparse index table, which saves memory and makes iteration order equal insertion order (guaranteed from 3.7). Average lookup, insertion and deletion are O(1); the worst case (many collisions) is O(n). The table resizes as it fills, which makes inserts amortized O(1). Keys must be hashable and must not change their hash while stored.
What are __slots__, and when are they useful?
Declaring __slots__ = ("x", "y") in a class replaces the per-instance __dict__ with fixed storage for those attributes. It reduces memory per object substantially and speeds attribute access slightly, and it prevents creating new attributes by typo. Useful when you create millions of small objects (tokens, graph nodes); @dataclass(slots=True) does it for you. Downsides: no dynamic attributes, and care is needed with inheritance.
What is the descriptor protocol, and how does @property use it?
A descriptor is an object defining __get__, and optionally __set__/__delete__, stored as a class attribute. When you access obj.attr, Python finds the descriptor on the class and calls its __get__(obj, type) instead of returning it directly. property is a built-in descriptor wrapping getter, setter and deleter functions; functions themselves are descriptors, which is how methods get bound to self. ORMs, Pydantic fields and functools.cached_property use the same mechanism.
How does asyncio work under the hood?
An event loop runs in one thread and manages coroutines (async def functions). When a coroutine hits await on something not ready (a socket read), it yields control back to the loop, which registers interest with the operating system's I/O notification mechanism and runs another ready task. When the I/O completes, the waiting coroutine is resumed. Concurrency is therefore cooperative: only one piece of Python runs at a time, and any blocking call without await stalls every task. asyncio.gather runs many awaitables concurrently; semaphores cap concurrency; asyncio.to_thread offloads blocking code.
What are strides, and how do they make transposes and sliding windows free?
Strides are the number of bytes to jump in memory to move one step along each dimension. A C-ordered int64 array of shape (3, 4) has strides (32, 8). Its transpose simply swaps them to (8, 32) without moving any data, which is why .T is a view but not contiguous. Tricks like sliding_window_view(a, 3) build overlapping windows by reusing memory with carefully chosen strides:
from numpy.lib.stride_tricks import sliding_window_view
sliding_window_view(np.arange(6), 3) # [[0 1 2] [1 2 3] [2 3 4] [3 4 5]], no copySuch views are read-only by default; writing to overlapping windows would be confusing.
What is np.einsum, and give three examples.
Einstein summation expresses multiply-and-sum operations with index labels: repeated indices are multiplied, indices absent from the output are summed.
np.einsum("ij,jk->ik", A, B) # matrix multiply, same as A @ B
np.einsum("ii->", M) # trace
np.einsum("ij->j", A) # column sums
np.einsum("bhqd,bhkd->bhqk", Q, K) # batched attention scores per headIt is expressive and avoids intermediate transposes; optimize=True picks an efficient contraction order for multi-operand expressions.
What numerical-precision issues matter in ML code?
float32has about 7 significant digits:np.float32(16777216) + 1is still 16777216, so summing millions of small values can lose accuracy; accumulate in float64 or use pairwise summation (NumPy'ssumalready does).expoverflows for large inputs: use max-subtraction in softmax,logsumexp,log1p, and fused losses likeBCEWithLogitsLoss.float16has a small range, so gradients can underflow to zero; mixed-precision training uses loss scaling, whilebfloat16keeps float32's range with less precision.- Never test floats with
==; usenp.isclose.
What is nested cross-validation, and when do you need it?
An outer CV loop estimates generalisation; inside each outer training split, an inner CV loop performs hyperparameter search. Because the outer validation fold never influenced the tuning, the outer score is an unbiased estimate of the whole "tune and train" procedure. You need it for small datasets where there is no room for a separate test set, or when comparing modelling procedures rigorously.
inner = GridSearchCV(pipe, grid, cv=StratifiedKFold(5, shuffle=True, random_state=0))
outer_scores = cross_val_score(inner, X, y, cv=StratifiedKFold(5, shuffle=True, random_state=1))How do you write a custom scikit-learn transformer that works in pipelines and grid search?
from sklearn.base import BaseEstimator, TransformerMixin
class Log1p(BaseEstimator, TransformerMixin):
def __init__(self, cols=None): # store params unchanged; no logic here
self.cols = cols
def fit(self, X, y=None):
self.n_features_in_ = X.shape[1] # learned state ends with an underscore
return self
def transform(self, X):
X = np.asarray(X, dtype=float).copy()
idx = self.cols if self.cols is not None else slice(None)
X[:, idx] = np.log1p(X[:, idx])
return XBaseEstimator provides get_params/set_params from the __init__ signature (so __init__ must only store arguments); TransformerMixin provides fit_transform. For stateless functions, FunctionTransformer(np.log1p) is simpler.
Why can target encoding leak, and how do you prevent it?
Target encoding replaces a category with the mean target of rows in that category. If a row's own target contributes to its encoding, rare categories essentially encode the label, and the model learns a shortcut that disappears on new data. Prevent it with cross-fitting (compute each row's encoding from other folds), smoothing toward the global mean for rare categories, and fitting the encoder inside the CV pipeline. scikit-learn's TargetEncoder performs internal cross-fitting in fit_transform.
What is probability calibration, and how do you check and fix it?
A classifier is calibrated if, among predictions of 0.8, about 80% are positive. Many models are not (SVM scores, boosted trees, heavily regularized or class-weighted models). Check with a reliability diagram (CalibrationDisplay) and the Brier score or log loss. Fix with CalibratedClassifierCV using sigmoid (Platt) scaling for small data or isotonic regression for larger data, fitted on data not used for training. Calibration matters when probabilities drive decisions, such as expected-cost thresholds or risk scores.
Implement logistic regression with gradient descent in NumPy.
def train_logreg(X, y, lr=0.1, epochs=500, l2=0.0):
n, d = X.shape
w, b = np.zeros(d), 0.0
for _ in range(epochs):
p = 1 / (1 + np.exp(-(X @ w + b))) # predicted probabilities
grad_w = X.T @ (p - y) / n + l2 * w # gradient of mean log loss (+ L2)
grad_b = (p - y).mean()
w -= lr * grad_w
b -= lr * grad_b
return w, b
# on standardized breast-cancer data: training accuracy about 0.986Key points: vectorized gradient X.T @ (p - y) / n, standardize inputs first, and for stability compute the loss with np.logaddexp if you log it.
Implement k-means in NumPy.
def kmeans(X, k, iters=100, seed=0):
rng = np.random.default_rng(seed)
C = X[rng.choice(len(X), k, replace=False)] # init from data points
for _ in range(iters):
labels = ((X[:, None, :] - C[None]) ** 2).sum(-1).argmin(1) # assign
newC = np.array([X[labels == j].mean(0) if np.any(labels == j) else C[j]
for j in range(k)]) # update
if np.allclose(newC, C):
break
C = newC
return C, labelsMention k-means++ initialization, multiple restarts (n_init), handling empty clusters, and scaling features first.
How does PyTorch's autograd graph work? What are leaf tensors and retain_graph?
PyTorch builds the graph dynamically during the forward pass: each result tensor stores a grad_fn pointing to the operation that created it and its inputs. Leaf tensors are those created directly by the user (such as parameters) with requires_grad=True; by default only leaves keep .grad after backward() (call retain_grad() on intermediates to keep theirs). After backward() the graph's saved buffers are freed to save memory, so calling backward twice through the same graph fails unless the first call used retain_graph=True. Because the graph is rebuilt every iteration, ordinary Python control flow (if, loops) works in models.
How does mixed-precision training work in PyTorch?
Inside torch.autocast(device_type="cuda", dtype=torch.float16 or torch.bfloat16), eligible operations such as matrix multiplications run in half precision for speed and memory savings, while precision-sensitive operations (reductions, softmax, losses) stay in float32. With float16, small gradients can underflow, so a torch.amp.GradScaler multiplies the loss before backward, unscales before the step, and skips steps with infinite gradients. bfloat16 has float32's exponent range and usually needs no scaler. Master weights stay float32.
DataParallel versus DistributedDataParallel: what is the difference?
DataParallel runs in one process: it splits each batch across GPUs, replicates the model every step and gathers outputs on one GPU, which is limited by the GIL and an overloaded main GPU. DistributedDataParallel runs one process per GPU (possibly across machines); each has its own model copy and data shard (DistributedSampler), and gradients are averaged with all-reduce during backward. DDP is faster and the recommended approach; for models too large for one GPU, sharded approaches such as FSDP split parameters, gradients and optimizer states.
How do you freeze layers and use different learning rates for parts of a model?
for p in model.backbone.parameters():
p.requires_grad = False # frozen: no gradients computed
opt = torch.optim.AdamW([
{"params": model.head.parameters(), "lr": 1e-3},
{"params": [p for p in model.encoder.parameters() if p.requires_grad], "lr": 1e-5},
], weight_decay=0.01)Keep frozen BatchNorm layers in eval mode if you do not want their running statistics updated, since requires_grad=False does not stop those updates. Check the trainable parameter count before training.
How do you make PyTorch experiments reproducible?
Seed every generator (random.seed, np.random.seed or explicit generators, torch.manual_seed, which also seeds CUDA); give the DataLoader a seeded generator and a worker_init_fn for workers; set torch.use_deterministic_algorithms(True) and torch.backends.cudnn.benchmark = False; pin library versions and hardware. Some GPU operations remain non-deterministic or become slower in deterministic mode, so bitwise reproducibility across different GPUs or versions is not guaranteed; report the mean and spread over several seeds.
What are the options for serializing models, and what are their risks?
scikit-learn: joblib/pickle, which requires the same library versions and can execute arbitrary code when loading, so only load trusted files; alternatives include skops for safer loading, or ONNX export for cross-language serving. PyTorch: save the state_dict with torch.save. Recent versions default to weights_only=True, which is safe but rejects a checkpoint dict that also stores epoch or config; use weights_only=False only for trusted full checkpoints. Prefer safetensors for tensor-only storage without pickle, and TorchScript, torch.export or ONNX for deployment without Python class definitions. Always store preprocessing, thresholds, feature lists and version metadata with the model.
What is a MultiIndex in pandas, and when is it useful?
A hierarchical index with several levels, for example (store, date). It results naturally from groupby on multiple keys or pivot_table. It enables selecting by partial keys (df.loc["Pune"], df.xs("2026-01-01", level="date")) and reshaping with stack/unstack. Many people call reset_index() right away because flat columns are easier to merge and export; know both.
When would you use Polars, DuckDB, Dask or Spark instead of pandas?
pandas is ideal when data fits comfortably in memory (a rule of thumb is a few times smaller than RAM). Polars is a multithreaded, Arrow-based DataFrame library with lazy query optimisation, much faster for large single-machine workloads. DuckDB runs fast analytical SQL directly on Parquet or DataFrames with out-of-core execution. Dask parallelises pandas-like code across cores or a cluster. Spark is for very large, distributed data in a cluster. Choose the simplest tool that fits the data size and team skills.
How would you implement scaled dot-product attention in NumPy and check its shapes?
def attention(Q, K, V, mask=None):
d_k = Q.shape[-1]
scores = Q @ K.swapaxes(-1, -2) / np.sqrt(d_k) # (..., q, k)
if mask is not None:
scores = np.where(mask, scores, -1e9) # block masked positions
w = softmax(scores, axis=-1) # rows sum to 1
return w @ V, w # (..., q, d_v)Using swapaxes(-1, -2) instead of .T keeps it correct for batched inputs. The theory behind the formula is in Transformers & LLMs.
Scenario & debugging
Your pandas job runs out of memory while loading and processing a 20 GB CSV. What do you do?
- Load only needed columns (
usecols) with explicit compact dtypes (float32, small ints,categoryfor repetitive strings) and parse dates once. - Prototype on
nrows=100_000to learn dtypes and memory per row. - Stream with
chunksize, reduce each chunk (filter, aggregate), and combine the small results. - Convert once to Parquet, partitioned if useful; later reads are faster and can select columns and row groups.
- Avoid intermediate copies: drop temporaries, avoid wide
applyresults, and check for accidental many-to-many merges. - If still too big, use DuckDB or Polars (lazy, out-of-core), or Dask/Spark.
Your training loss becomes NaN after a few hundred steps. How do you debug it?
- Check the data:
np.isnan(X).any(), infinities, division by zero in feature engineering, unscaled huge values. - Lower the learning rate; add warm-up; clip gradients (
clip_grad_norm_). - Look for numerically unsafe operations:
log(0),sqrtof negatives, manual softmax or sigmoid followed by log; useCrossEntropyLoss/BCEWithLogitsLosson logits and add epsilons. - With float16 mixed precision, use a GradScaler or switch to bfloat16.
- Locate the first bad operation with
torch.autograd.set_detect_anomaly(True)and by logging gradient norms per layer. - Check labels are in range (for example, class index equal to
n_classes).
A fraud model shows 99.5% accuracy but the business says it catches nothing. What happened?
With about 0.5% fraud, predicting "not fraud" for everything gives 99.5% accuracy. Accuracy is the wrong metric. Evaluate recall, precision and PR-AUC on the fraud class, compare against a DummyClassifier baseline, use class weighting or resampling inside the pipeline, and pick a decision threshold on validation data that meets the business's recall target at an acceptable alert volume. Report the confusion matrix in counts the business understands ("we catch 70 of 100 frauds with 300 alerts a day").
Cross-validation says 0.95 AUC, but the model performs at 0.70 in production. What are the likely causes?
- Leakage: preprocessing fitted before splitting, features computed with future information, target-derived columns, or the same entity in train and validation folds (need
GroupKFold). - Wrong validation scheme for time: random folds on temporal data; use
TimeSeriesSplitor an out-of-time holdout. - Training-serving skew: preprocessing re-implemented differently in production, different category spellings, default values for missing fields.
- Data drift: the population or behaviour changed since training.
Investigate by replaying production inputs through the saved pipeline, comparing feature distributions between training and production, and re-validating with a time-based split.
You get "CUDA out of memory" during training. What are your options?
- Reduce the batch size, and use gradient accumulation to keep the effective batch size.
- Use mixed precision (
autocastwith bfloat16 or float16). - Make sure evaluation runs under
torch.no_grad()and that you logloss.item(), not the tensor, so old graphs are freed. - Delete large unused tensors, avoid keeping predictions on the GPU, and move metrics to CPU.
- Use gradient checkpointing (recompute activations in backward), shorter sequence lengths, a smaller model, or parameter-efficient fine-tuning such as LoRA.
- Check for other processes on the GPU with
nvidia-smi, and inspecttorch.cuda.memory_summary().
Your PyTorch training loss does not decrease at all. What do you check?
- Try to overfit a single small batch; if the model cannot drive the loss near zero, the bug is in the code, not in the data volume.
- Verify the loop:
zero_grad,backward,stepall called; the optimizer receivedmodel.parameters()(not an empty or stale list); parameters actually haverequires_grad=True. - Check that labels match inputs after shuffling and that label encoding is correct.
- Check the loss setup: raw logits into
CrossEntropyLoss, correct target dtype and shape (no accidental broadcasting from (N,) versus (N, 1)). - Tune the learning rate (too low: no movement; too high: bouncing); check input scaling.
- Inspect gradient norms per layer for vanishing or exploding gradients.
Evaluating the same trained model twice on the same validation set gives different accuracies. Why?
Almost certainly model.eval() was not called, so dropout is still randomly zeroing activations (and BatchNorm is using batch statistics and updating its running averages). Call model.eval() before evaluation and model.train() afterwards. Other causes: random test-time augmentation, a shuffled DataLoader combined with a metric that depends on order, or non-deterministic GPU kernels (tiny differences only).
After a merge, your DataFrame has more rows than either input. What happened and how do you prevent it?
The join key is duplicated on both sides (a many-to-many relationship), so each matching pair produces a row. Check key uniqueness with df.key.is_unique or df.key.duplicated().sum(), de-duplicate or aggregate the lookup table first, and pass validate="many_to_one" (or "one_to_one") to merge so pandas raises an error. Also check for missing or inconsistent key types (int versus string IDs) that cause silent non-matches, using indicator=True.
A notebook runs perfectly for you but fails for a colleague. How do you fix it?
Likely causes: hidden kernel state (variables from deleted or re-ordered cells), different package versions, absolute local file paths, missing environment variables, or unseeded randomness. Restart the kernel and run all cells from top to bottom; pin dependencies in a lock file or environment file; use relative paths or a config file; seed random generators; and move reusable logic into a module that can be tested.
A feature-engineering script with df.apply(axis=1) takes two hours. How do you speed it up?
- Profile (
%prun,line_profiler) to confirm where time goes. - Rewrite row-wise logic as column operations: arithmetic,
np.where/np.select,.strand.dtaccessors,mapwith dicts. - Replace per-row lookups with a single
merge; replace per-group Python functions with built-ingroupbyaggregations ortransform. - Fix dtypes (numbers stored as strings slow everything).
- If a sequential loop is unavoidable, compile it with Numba; then parallelize across partitions if still needed.
Such rewrites commonly give speed-ups of 50 to 500 times.
GPU utilisation is only 20% during training. What is the bottleneck and how do you fix it?
Usually the input pipeline: the GPU waits for the CPU to load, decode and augment data. Increase num_workers, set pin_memory=True and non_blocking=True on transfers, use persistent_workers=True and prefetching, pre-process data offline into a fast format, and increase batch size if memory allows. Also avoid synchronising calls in the loop (.item() or printing every step, moving tensors to CPU). Confirm with the PyTorch profiler.
Your DataLoader with num_workers=4 hangs or crashes on Windows. Why?
Windows starts worker processes with the "spawn" method, which re-imports the main script. Without an if __name__ == "__main__": guard, the training code re-executes in every worker. Datasets defined inside a notebook or using lambdas may also fail to pickle. Fix: put the training entry point under the guard, define Dataset classes and collate functions at module level in a .py file, or use num_workers=0 in notebooks.
Every training run gives different results. How do you make them consistent, and is perfect reproducibility realistic?
Fix all seeds (Python, NumPy, PyTorch and random_state in scikit-learn splitters and models), seed DataLoader shuffling, enable deterministic algorithms, and pin library versions. Record the data snapshot and configuration. Exact bitwise results across different hardware, drivers or library versions are not guaranteed, so report the mean and standard deviation across several seeds, which is more honest than one lucky run.
A model loaded in production gives different predictions than it did in the notebook. What could be wrong?
- Preprocessing mismatch: only the model was saved, and scaling or encoding was re-implemented differently. Save the entire pipeline.
- Column order or names differ; select features by the saved feature list.
- Library version differences for pickled objects.
- For PyTorch:
model.eval()not called, a different tokenizer, a different dtype, or a missing custom threshold. - Unseen categories or missing values handled by defaults.
Write a regression test that feeds fixed sample inputs and compares outputs to stored expected predictions.
A categorical feature has 50,000 distinct values (for example, merchant ID). How do you encode it?
One-hot encoding creates 50,000 sparse columns, which works for linear models with sparse matrices but is heavy and overfits rare levels. Options: group rare levels into "other" (OneHotEncoder(min_frequency=..., max_categories=...)), frequency or count encoding, cross-fitted target encoding (TargetEncoder), feature hashing, native categorical support in gradient-boosting libraries, or learned embeddings in a neural network. Handle unseen IDs at prediction time explicitly.
A decision tree scores 100% on training data and 70% on test. What do you do?
This is overfitting: an unconstrained tree memorises the training set. Regularize with max_depth, min_samples_leaf, min_samples_split or cost-complexity pruning (ccp_alpha), tuned by cross-validation; or switch to an ensemble (random forest, gradient boosting) that averages many trees. Plot training versus validation score against depth to see the sweet spot between underfitting (depth 1) and overfitting (no limit).
Your classifier predicts only the majority class. What are the possible fixes?
Check the probability outputs first: the model may rank well (good AUC) while the 0.5 threshold is simply too high for a rare class, in which case lower the threshold using validation data. Otherwise use class_weight="balanced" or a weighted loss, resample training folds, ensure features are informative and scaled, and verify that labels were not corrupted by a mapping bug. Evaluate with recall, precision and PR-AUC rather than accuracy.
Dates in your dataset such as "03/04/2026" are being parsed inconsistently. How do you handle it?
Ambiguous formats are parsed month-first by default: pd.to_datetime("03/04/2026") gives 4 March, while dayfirst=True gives 3 April. Always pass an explicit format="%d/%m/%Y" when you know it, use errors="coerce" to turn bad values into NaT and count them, and store dates in ISO 8601 (YYYY-MM-DD) or as typed columns in Parquet. Be explicit about time zones (tz_localize, tz_convert).
You need to run 10,000 prompts through an LLM API for an evaluation. How do you do it efficiently and safely?
- Use an async client with
asyncio.gatherand a semaphore sized to the rate limit (or a thread pool with a sync client). - Retry transient failures and rate-limit responses with exponential backoff and jitter; set timeouts.
- Write results incrementally (JSONL) keyed by prompt ID so a crash can resume without repeating work; cache by (model, prompt, parameters).
- Use
temperature=0for reproducible scoring, validate structured output with Pydantic, and log token usage for cost. - Use only approved endpoints and keep keys in environment variables.
GPU or CPU memory keeps growing every epoch in your PyTorch script. What is leaking?
The classic cause is accumulating tensors that still hold their autograd graphs, for example total_loss += loss or appending outputs to a list for later metrics. Use loss.item() and outputs.detach().cpu(). Other causes: evaluation without no_grad, storing whole batches in a history list, and growing caches in custom datasets. Monitor with torch.cuda.memory_allocated() per epoch to confirm the fix.
A medical dataset has several scans per patient, and cross-validation looks too good. Why, and how do you fix it?
With random row-level splitting, scans from the same patient appear in both training and validation folds, so the model can recognise the patient rather than learn the disease; performance on new patients will be lower. Split by patient with GroupKFold or StratifiedGroupKFold(groups=patient_id), and make the final test set contain entirely unseen patients. The same applies to users, devices, sessions and near-duplicate documents.
A clinical team requires recall of at least 0.95 for malignant cases with precision of at least 0.60. How do you deliver a model that meets this?
- Relabel so malignant is the positive class, split with stratification, and keep a locked test set.
- Build a pipeline (scaling plus model), tune with
scoring="recall"or average precision, and considerclass_weight="balanced". - Get out-of-fold probabilities on the training data (
cross_val_predict(..., method="predict_proba")), compute the precision-recall curve, and choose the highest threshold with recall at or above the target plus a safety margin, checking that precision stays above 0.60. - Evaluate once on the test set at that threshold and report recall with a confidence interval; with about 40 positive test cases, one miss moves recall by roughly 2.5 points.
- Document the threshold with the model, and monitor recall after deployment.
A filter df[df.ratio == 0.3] returns no rows even though you can see 0.3 in the data. Why?
Floating-point values that display as 0.3 are often 0.30000000000000004 or similar after arithmetic, so exact equality fails. Use a tolerance, df[np.isclose(df.ratio, 0.3)], or round deliberately before comparing, or store exact quantities as integers (cents rather than currency units) or decimals. Also note that NaN == NaN is False; use isna().
A teammate asks you to hand over your trained model for a batch-scoring job. What do you deliver?
A single artifact containing the whole fitted pipeline (preprocessing plus model), the decision threshold, the expected input schema (column names, dtypes, allowed categories), and metadata: library versions, training data snapshot and date, metrics with confidence ranges, and the random seed. Add a small scoring function or command-line script, a sample input with expected outputs as a regression test, and a pinned environment file. For PyTorch, deliver the state_dict plus model code and config, or an exported format such as ONNX.