AI Foundations

Python & Data Tools for ML

The practical toolkit behind almost every machine-learning project: the Python language features you will read and write daily, the NumPy and pandas data stack, plotting and exploratory analysis, scikit-learn for classical models, and PyTorch for neural networks. Read top to bottom and you will be able to take a raw CSV to an evaluated, saved model, and defend every line of it in an interview.

~135 min read 0 interview questions
In 30 seconds
  • Python is the glue; the heavy maths runs in compiled C, Fortran or CUDA code underneath NumPy, pandas, scikit-learn and PyTorch. Your job is to keep work inside those fast libraries and out of Python loops.
  • NumPy gives you typed n-dimensional arrays, vectorized operations and broadcasting. Knowing shapes, axes and views versus copies prevents most bugs.
  • pandas adds labelled tables: select with loc/iloc, summarise with groupby, combine with merge, and watch dtypes and memory.
  • scikit-learn has one API for everything: fit, transform, predict. Wrap preprocessing and model in a Pipeline so cross-validation never leaks test information.
  • PyTorch is NumPy on GPUs plus automatic gradients. Every training loop is zero_grad, forward, loss, backward, step; switch to model.eval() and no_grad for inference.
  • Interviewers probe mutability, generators, decorators, the GIL, broadcasting, groupby, data leakage, metric choice, and debugging stories like out-of-memory errors or NaN losses.

The big picture: why Python runs ML

Python itself is a slow, interpreted, dynamically typed language. It became the language of machine learning anyway because it is easy to read and because its scientific libraries push the expensive work into compiled code. When you write X @ W in NumPy, one Python instruction triggers millions of multiply-adds inside an optimized linear-algebra library (BLAS). When you call model(x) in PyTorch on a GPU, Python only launches kernels; the GPU does the arithmetic.

Analogy

Think of a film director and a studio crew. The director (Python) is not the fastest at building sets or lighting scenes, but gives short, clear instructions. The crew (C, Fortran and CUDA libraries) does the physical heavy lifting at great speed. A good director says "light the whole stage" once instead of pointing at each bulb in turn. In code terms: one vectorized call on a whole array beats a Python loop over each element, because the loop makes the slow director do the crew's job.

The layer cake

+------------------------------------------------------------------+
|  Your notebook / script / service  (Python: glue, control flow)  |
+------------------------------------------------------------------+
|  scikit-learn   PyTorch   Hugging Face   LLM SDKs                |  modelling
+------------------------------------------------------------------+
|  pandas (tables)   matplotlib / seaborn (plots)   SciPy          |  data + stats
+------------------------------------------------------------------+
|  NumPy ndarray  (typed, contiguous memory, vectorized ufuncs)    |  array core
+------------------------------------------------------------------+
|  C / C++ / Fortran / CUDA : BLAS, LAPACK, cuBLAS, cuDNN, Arrow   |  speed lives here
+------------------------------------------------------------------+

The standard ML workflow in Python

  1. Load Read CSV, Parquet, SQL or JSON into a pandas DataFrame.
  2. Explore (EDA) Check shape, dtypes, missing values, distributions, class balance, correlations and leakage risks.
  3. Split Hold out a test set first (stratified for classification) so it stays untouched.
  4. Preprocess Impute, encode categoricals, scale numerics, engineer features, all inside a pipeline fitted on training data only.
  5. Train and tune Fit models, compare with cross-validation, tune hyperparameters with a search.
  6. Evaluate Use metrics that match the business cost (recall for missed cancers, precision for spam, RMSE for prices), inspect errors, tune the decision threshold on validation data.
  7. Ship Save the whole pipeline, version data and code, monitor in production.

NumPy

Fast typed arrays and maths. The data format every other library speaks.

pandas

Labelled tables for loading, cleaning, joining and summarising data.

matplotlib / seaborn

Plots for exploration and reporting.

scikit-learn

Classical ML: preprocessing, models, validation and metrics with one consistent API.

PyTorch

Tensors on GPU plus automatic differentiation for deep learning.

Hugging Face

Pretrained transformer models and tokenizers in a few lines.

Common pitfall Writing ML code "Python style" with explicit loops over rows, pixels or tokens. It works on a toy sample and then takes hours on real data. Whenever you write for i in range(len(...)) over data, ask whether a NumPy, pandas or PyTorch operation can express the same thing on the whole array.
Interview angle "Python is slow, so why is it used for ML?" A strong answer: Python is an orchestration layer; hot loops run in compiled libraries (BLAS, cuDNN), so the interpreter overhead is paid once per array operation, not per element. Add that the ecosystem, readability and notebook workflow matter more for research speed than raw interpreter speed, and that when Python does become the bottleneck you vectorize, use Numba or Cython, or move work to the GPU.
Tip For the maths behind these tools (vectors, gradients, probability, metrics), see Math & Statistics for ML. For the algorithms themselves, see Machine Learning and Deep Learning.

Environments, packaging and Jupyter

Every project should live in its own isolated environment so that one project's numpy 1.26 cannot break another project's numpy 2.x. An environment is simply a folder containing a Python interpreter plus the packages installed for it.

Analogy

A virtual environment is like a separate toolbox per job site. The plumber's box and the electrician's box each hold the exact tool versions that job needs, so borrowing a wrench for one job never removes it from the other. Mapping back: each project gets its own .venv folder with its own package versions, and a lock file (requirements.txt) is the inventory list that lets anyone rebuild the same toolbox.

venv + pip (built in)

# create and activate an isolated environment
python -m venv .venv
.venv\Scripts\activate            # Windows PowerShell / cmd
source .venv/bin/activate         # macOS / Linux

python -m pip install --upgrade pip
pip install numpy pandas scikit-learn matplotlib seaborn torch
pip freeze > requirements.txt     # pin exact versions
pip install -r requirements.txt   # recreate elsewhere
deactivate

The options compared

ToolWhat it isWhen to use
venv + pipStandard library environments plus the default installerDefault choice for pure-Python projects
conda / mambaEnvironment and package manager that also installs non-Python binaries (CUDA, MKL, compilers)Heavy scientific stacks, GPU toolkits, mixed-language dependencies
uv, poetry, pip-toolsFaster resolvers and lock-file managersReproducible builds for teams and CI
DockerWhole operating-system image with Python insideDeployment, exact reproducibility including system libraries

Packaging your own code

Modern projects describe themselves in a pyproject.toml file (name, version, dependencies, build backend). A folder with an __init__.py is a package; any .py file is a module. Installing your project in editable mode (pip install -e .) lets notebooks import your code while you keep editing it.

# pyproject.toml (minimal)
[project]
name = "churn_model"
version = "0.1.0"
requires-python = ">=3.10"
dependencies = ["pandas>=2.0", "scikit-learn>=1.4"]

[build-system]
requires = ["setuptools>=68"]
build-backend = "setuptools.build_meta"

The guard if __name__ == "__main__": makes a file behave as both an importable module and a runnable script: the block runs only when the file is executed directly, not when it is imported. It is also required on Windows for multiprocessing and PyTorch DataLoader workers, because child processes re-import the main file.

Jupyter essentials

A notebook talks to a kernel, a live Python process that keeps variables in memory between cells. That is convenient for exploration and dangerous for reproducibility, because the order you ran cells in is invisible to a reader.

%pip install seaborn        # installs into the kernel's own environment
%timeit np.sort(arr)        # micro-benchmark a line
%%time                      # time a whole cell (must be the first line)
%load_ext autoreload
%autoreload 2               # re-import edited .py modules automatically
%matplotlib inline          # render plots below cells
df.head()                   # the last expression in a cell is displayed
Common pitfall Hidden state: you delete a cell that defined a variable, but the variable still lives in the kernel, so the notebook "works" for you and crashes for everyone else. Before sharing, always use Restart Kernel and Run All. Also prefer %pip install over !pip install, since the shell version may install into a different interpreter than the running kernel.
Interview angle Expect "how do you make an ML project reproducible?" Mention: isolated environment with pinned versions (lock file), fixed random seeds, data versioning, configuration files instead of hard-coded constants, code moved from notebooks into modules with tests, and a clean restart-and-run-all or a scripted pipeline. Bonus: record library versions and hardware alongside model artifacts.

Core Python: types, mutability and containers

In Python, every value is an object with a type, and variables are just names bound to objects. Assignment never copies data; it attaches another name to the same object. Whether that shared object can be changed in place (mutability) explains a large share of Python bugs and interview questions.

Analogy

Variables are sticky notes, not boxes. Writing b = a sticks a second note on the same whiteboard. If the whiteboard is erasable (a list or dict), writing on it through either note changes what both notes point to. If it is laminated (an int, str or tuple), "changing" it really means making a new laminated card and moving one sticky note to it. Mapping back: mutable objects are shared and changed in place; immutable objects are replaced, so other names are unaffected.

Built-in types at a glance

TypeExampleMutable?Typical ML use
int, float, complex10, 3e-4NoHyperparameters, counters
boolTrueNoFlags; a subclass of int (True + True == 2)
str"relu"NoText data, config keys
bytesb"\x00"NoRaw files, network payloads
tuple(32, 3, 224, 224)NoShapes, fixed records, dict keys
list[0.9, 0.8]YesOrdered collections, loss history
dict{"lr": 1e-3}YesConfigs, JSON, lookups (insertion-ordered since 3.7)
set / frozenset{"cat", "dog"}Yes / NoVocabularies, de-duplication, fast membership
NoneTypeNoneNo"No value" sentinel

Names, identity and equality

a = [1, 2]
b = a              # same object, two names
b.append(3)
print(a)           # [1, 2, 3]  -- a changed too
print(a is b)      # True   (same identity)
print(a == [1, 2, 3], a is [1, 2, 3])   # True False  (equal value, different object)

x = 10
y = x
y += 1             # ints are immutable: y now names a NEW object
print(x, y)        # 10 11

== compares values (calls __eq__); is compares identity (same object in memory). Use is only for singletons: if x is None. CPython caches small integers (-5 to 256) and some strings, so is may accidentally appear to work on them; never rely on it.

Shallow versus deep copies

import copy
grid = [[0, 0], [0, 0]]
shallow = grid.copy()          # new outer list, SAME inner lists
deep = copy.deepcopy(grid)     # everything duplicated
grid[0][0] = 9
print(shallow[0][0], deep[0][0])   # 9 0

bad = [[0] * 2] * 2            # two references to ONE inner list
bad[0][0] = 1
print(bad)                     # [[1, 0], [1, 0]]
good = [[0] * 2 for _ in range(2)]

Container cost cheat sheet

Operationlistdict / setcollections.deque
Append at endO(1) amortizedO(1) average insertO(1)
Insert / pop at frontO(n)n/aO(1)
Index by positionO(1)n/aO(n) in the middle
Membership x in cO(n)O(1) averageO(n)
SortO(n log n), stablen/an/a

Dict keys and set members must be hashable: immutable and with a stable __hash__. That is why a tuple can be a dict key but a list cannot. See Data Structures & Algorithms for how hash tables achieve O(1) average lookups.

Strings and f-strings

name, lr, acc = "resnet", 3e-4, 0.91234
print(f"{name:>8} lr={lr:.1e} acc={acc:.2%}")   #   resnet lr=3.0e-04 acc=91.23%
print(f"{acc=:.3f}")                            # acc=0.912   (self-documenting, 3.8+)
words = "the cat sat".split()                   # ['the', 'cat', 'sat']
print(" ".join(words).upper())                  # THE CAT SAT
# strings are immutable: building with += in a loop is O(n^2); collect then join

Truthiness and handy built-ins

Empty containers, 0, 0.0, "" and None are falsy; everything else is truthy. But note that NumPy arrays and pandas Series refuse to be used in an if directly (they raise "truth value is ambiguous"); use .any(), .all() or .empty.

from collections import Counter, defaultdict
labels = ["cat", "dog", "cat", "bird", "cat"]
print(Counter(labels).most_common(2))     # [('cat', 3), ('dog', 1)]
by_len = defaultdict(list)
for w in labels:
    by_len[len(w)].append(w)
print(dict(by_len))                       # {3: ['cat', 'dog', 'cat', 'cat'], 4: ['bird']}
for i, (w, n) in enumerate(zip(["a", "b"], [1, 2])):
    print(i, w, n)                        # 0 a 1 / 1 b 2
print(sorted(labels, key=len, reverse=True)[:2])   # ['bird', 'cat']

Floating point

print(0.1 + 0.2 == 0.3)          # False: binary floats cannot represent 0.1 exactly
import math
print(math.isclose(0.1 + 0.2, 0.3))   # True -- compare with a tolerance (np.isclose for arrays)
print(float("nan") == float("nan"))    # False: NaN is not equal to anything, even itself
Common pitfall Mutable default arguments. def f(x, acc=[]) creates the list once, when the function is defined, so every call shares it: f(1) then f(2) returns [1, 2]. Use acc=None and create a new list inside. The same trap appears with dict defaults in config helpers.
Interview angle Classic probes: difference between list and tuple (mutability, hashability, intent, slightly less memory), why a list cannot be a dict key, is versus ==, shallow versus deep copy, and the output of the [[0]*2]*2 puzzle. Strong candidates explain the "names bound to objects" model instead of memorising individual cases.

Functions, comprehensions, closures and decorators

Functions in Python are ordinary objects: you can pass them as arguments, return them from other functions, and store them in dicts. That single fact enables callbacks (learning-rate schedulers), closures (factories that remember settings) and decorators (wrappers that add timing, caching or retries).

Analogy

A decorator is gift wrapping. The present inside (your function) is unchanged, but the wrapping adds something on the outside: a label with the time it took, a guarantee to retry if it breaks, or a cache that remembers previous answers. A closure is a backpack a function carries: variables from the place it was created travel with it. Mapping back: @timer returns a new function that calls the original and adds behaviour around it, and that wrapper remembers the original through a closure.

Arguments: positional, keyword, defaults, *args and **kwargs

def build_mlp(in_dim, *hidden, activation="relu", dropout=0.0, **extra):
    print(in_dim, hidden, activation, dropout, extra)

build_mlp(30, 64, 32, dropout=0.1, bias=False)
# 30 (64, 32) relu 0.1 {'bias': False}

def scale(x, /, factor=2, *, clip=None):   # x positional-only, clip keyword-only
    y = x * factor
    return min(y, clip) if clip is not None else y
print(scale(5, clip=8))        # 8

params = {"activation": "gelu", "dropout": 0.2}
build_mlp(10, 16, **params)    # unpack a dict into keyword arguments
shape = (2, 3)
rows, cols = shape             # tuple unpacking
first, *rest = [1, 2, 3, 4]    # first=1, rest=[2, 3, 4]

*args collects extra positional arguments into a tuple; **kwargs collects extra keyword arguments into a dict. At a call site, * and ** do the reverse: they spread a sequence or dict into arguments. Library code uses this heavily to forward options, for example model(**tokenizer_output) in Hugging Face.

Comprehensions

nums = range(10)
squares = [n * n for n in nums if n % 2 == 0]        # [0, 4, 16, 36, 64]
tok2id = {t: i for i, t in enumerate(["the", "cat"])} # {'the': 0, 'cat': 1}
id2tok = {i: t for t, i in tok2id.items()}            # invert a mapping
uniq_lens = {len(w) for w in ["a", "bb", "cc"]}       # {1, 2}
total = sum(n * n for n in nums)                      # generator expression: no list built
matrix = [[r * c for c in range(3)] for r in range(2)]  # [[0, 0, 0], [0, 1, 2]]
flat = [x for row in matrix for x in row]            # loops read left to right

Comprehensions are faster than equivalent append loops (less interpreter overhead) and clearer when short. If a comprehension needs nested conditions or side effects, use a normal loop. For numeric work on large data, neither beats a NumPy vectorized operation.

Lambda, map, filter, sorted with key

runs = [{"name": "a", "f1": 0.81}, {"name": "b", "f1": 0.87}]
best = max(runs, key=lambda r: r["f1"])        # {'name': 'b', 'f1': 0.87}
ranked = sorted(runs, key=lambda r: -r["f1"])
lengths = list(map(len, ["hi", "there"]))      # [2, 5]
long_words = list(filter(lambda w: len(w) > 2, ["a", "the", "cat"]))  # ['the', 'cat']

Scope (LEGB) and closures

Python resolves a name by searching Local, then Enclosing function, then Global (module), then Built-in scope. Assigning to a name inside a function makes it local unless declared global or nonlocal.

def make_scheduler(base_lr, decay):
    step = 0
    def next_lr():
        nonlocal step              # modify the enclosing variable
        lr = base_lr * decay ** step
        step += 1
        return lr
    return next_lr                 # the inner function "closes over" base_lr, decay, step

sched = make_scheduler(0.1, 0.5)
print(sched(), sched(), sched())   # 0.1 0.05 0.025
Common pitfall Late binding in closures. fns = [lambda: i for i in range(3)] then fns[0]() returns 2, not 0, because every lambda looks up i when called, after the loop finished. Fix it by binding at definition time: lambda i=i: i.

Decorators

import time, functools

def timer(func):
    @functools.wraps(func)                 # keep the original name and docstring
    def wrapper(*args, **kwargs):
        start = time.perf_counter()
        result = func(*args, **kwargs)
        print(f"{func.__name__} took {time.perf_counter() - start:.3f}s")
        return result
    return wrapper

def retry(times=3, exceptions=(ConnectionError,)):   # decorator factory: takes arguments
    def decorator(func):
        @functools.wraps(func)
        def wrapper(*args, **kwargs):
            for attempt in range(1, times + 1):
                try:
                    return func(*args, **kwargs)
                except exceptions:
                    if attempt == times:
                        raise
                    time.sleep(2 ** attempt)       # exponential backoff
        return wrapper
    return decorator

@timer
def train_epoch():
    time.sleep(0.2)
train_epoch()                                  # train_epoch took 0.200s

@functools.lru_cache(maxsize=None)             # memoize pure functions
def fib(n):
    return n if n < 2 else fib(n - 1) + fib(n - 2)
print(fib(80))                                 # 23416728348467685, instantly

@timer above a definition is shorthand for train_epoch = timer(train_epoch). Built-in decorators you will meet constantly: @property, @staticmethod, @classmethod, @dataclass, @functools.lru_cache, @torch.no_grad(), and @app.get(...) in web frameworks.

Interview angle "Write a decorator that times a function" is a very common live-coding task. Interviewers look for *args, **kwargs forwarding, returning the result, functools.wraps, and whether you can extend it to a decorator that takes arguments (three nested levels). Follow-ups: what is a closure, what does nonlocal do, and why does lru_cache require hashable arguments.

Iterators and generators: streaming data lazily

An iterable is anything you can loop over (list, dict, file, range). An iterator is the object that actually produces the values one at a time and remembers where it is. A generator is the easiest way to write an iterator: a function containing yield. Generators compute values on demand, so they can stream a 50 GB log file or an infinite data source using almost no memory. This is exactly how data loaders feed mini-batches to a model.

Analogy

A list is a printed book: every page exists up front and you can flip to page 300. A generator is a storyteller: each time you say "next", they tell you one more sentence and then pause, remembering exactly where they were. They never need the whole story written down. Mapping back: next() resumes the generator function until its next yield, the paused frame keeps its local variables, and when the story ends Python raises StopIteration, which a for loop handles silently.

The iterator protocol

nums = [10, 20]
it = iter(nums)          # calls nums.__iter__()
print(next(it))          # 10   (calls it.__next__())
print(next(it))          # 20
# next(it) now raises StopIteration -- a for loop catches this for you

class Countdown:
    def __init__(self, start):
        self.current = start
    def __iter__(self):
        return self
    def __next__(self):
        if self.current <= 0:
            raise StopIteration
        self.current -= 1
        return self.current + 1

print(list(Countdown(3)))   # [3, 2, 1]

Generators

def countdown(start):
    while start > 0:
        yield start          # pause here, hand out a value
        start -= 1

gen = countdown(3)
print(next(gen), list(gen))  # 3 [2, 1]
print(list(gen))             # []  -- a generator is exhausted after one pass

def batches(items, size):
    """Yield successive mini-batches without building them all at once."""
    for i in range(0, len(items), size):
        yield items[i:i + size]

print(list(batches(list(range(7)), 3)))   # [[0, 1, 2], [3, 4, 5], [6]]

def read_jsonl(path):
    import json
    with open(path, encoding="utf-8") as f:
        for line in f:                    # files are lazy iterators over lines
            yield json.loads(line)

# pipelines of generators: nothing runs until something consumes them
# valid = (r for r in read_jsonl("events.jsonl") if r.get("label") is not None)

Generator expressions and memory

import sys
squares_list = [n * n for n in range(1_000_000)]
squares_gen = (n * n for n in range(1_000_000))
print(sys.getsizeof(squares_list) > 8_000_000)   # True: about 8 MB of pointers
print(sys.getsizeof(squares_gen) < 500)          # True: a few hundred bytes, whatever n is

itertools and yield from

import itertools as it
print(list(it.islice(it.count(0, 5), 4)))              # [0, 5, 10, 15]
print(list(it.chain([1, 2], [3])))                     # [1, 2, 3]
grid = list(it.product([0.01, 0.1], ["relu", "gelu"]))  # hyperparameter grid
print(len(grid))                                       # 4
print(list(it.accumulate([1, 2, 3, 4])))               # [1, 3, 6, 10]  running sum
print(list(it.batched(range(5), 2)))                   # [(0, 1), (2, 3), (4,)]  (3.12+)

def flatten(nested):
    for item in nested:
        if isinstance(item, list):
            yield from flatten(item)     # delegate to a sub-generator
        else:
            yield item
print(list(flatten([1, [2, [3, 4]], 5])))   # [1, 2, 3, 4, 5]

List / list comprehension

  • All values computed and stored immediately
  • Random access, len(), reusable many times
  • Memory grows with n

Generator / generator expression

  • Values computed on demand
  • Single pass, no len(), no indexing
  • Constant memory; can be infinite
Common pitfall Consuming a generator twice. Code such as total = sum(gen); count = len(list(gen)) gives a count of zero because the first pass exhausted it. Either materialise it once into a list, or recreate the generator. The same bug appears when a PyTorch-style data iterator is reused without being re-created each epoch.
Interview angle Expect "what is the difference between an iterator and an iterable", "what does yield do", and "how would you process a file larger than RAM". A good answer names the protocol (__iter__, __next__, StopIteration), explains that a generator's frame is suspended rather than returned, and proposes streaming with a generator or pandas chunksize.

Classes, OOP and dunder methods

Object-oriented code bundles data (attributes) with behaviour (methods). You need it in ML because frameworks are built on it: every PyTorch model is a class that inherits from nn.Module, every custom dataset implements __len__ and __getitem__, and every scikit-learn estimator follows a class-based contract.

Analogy

A class is a cookie cutter and objects are the cookies. The cutter defines the shape (attributes and methods); each cookie has its own decorations (instance values). Dunder ("double underscore") methods are the standard sockets on an appliance: if your device has the right plug shape, it works with every outlet in the house. Mapping back: implementing __len__ lets len(obj) work, __getitem__ lets obj[i] work, and __call__ lets obj(x) work, so your class plugs into Python syntax and libraries that expect those sockets.

A class from scratch

class Neuron:
    activation = "relu"                     # class attribute: shared by all instances

    def __init__(self, weights, bias=0.0):  # initializer, runs on Neuron(...)
        self.weights = list(weights)        # instance attributes
        self.bias = bias

    def forward(self, inputs):
        z = sum(w * x for w, x in zip(self.weights, inputs)) + self.bias
        return max(0.0, z)

    def __call__(self, inputs):             # makes instances callable like functions
        return self.forward(inputs)

    def __repr__(self):                     # unambiguous developer view
        return f"Neuron(weights={self.weights}, bias={self.bias})"

n = Neuron([0.5, -0.2], bias=0.1)
print(n)                 # Neuron(weights=[0.5, -0.2], bias=0.1)
print(n([1.0, 2.0]))     # 0.19999999999999998  (0.2 up to float rounding)

self is simply the instance, passed automatically as the first argument when you call n.forward(...). __init__ is not a constructor in the strict sense (that is __new__); it initialises an already-created object.

Inheritance and super()

class Model:
    def __init__(self, name):
        self.name = name
    def predict(self, x):
        raise NotImplementedError

class Constant(Model):
    def __init__(self, name, value):
        super().__init__(name)      # run the parent's initializer
        self.value = value
    def predict(self, x):           # override (polymorphism)
        return [self.value] * len(x)

print(Constant("baseline", 1).predict([7, 8, 9]))   # [1, 1, 1]
print(isinstance(Constant("b", 0), Model))         # True

With multiple inheritance, Python follows the method resolution order (MRO, the C3 linearization, viewable via Cls.__mro__); super() means "next class in the MRO", not simply "my parent".

The dunder methods worth knowing

MethodTriggered byML example
__init__Cls(...)Define layers in nn.Module
__repr__ / __str__repr(x), printing / str(x)Readable model summaries
__len__len(x)Dataset size for a DataLoader
__getitem__x[i], slicing, iteration fallbackReturn one (features, label) sample
__iter__ / __next__for loopsStreaming datasets
__call__x(...)model(x) runs hooks and then forward
__eq__ / __hash__==, dict keys, setsCaching configs
__add__, __matmul__+, @How tensors overload operators
__enter__ / __exit__with blockstorch.no_grad(), file handles
__contains__inVocabulary membership
class Vocab:
    def __init__(self, tokens):
        self.itos = sorted(set(tokens))
        self.stoi = {t: i for i, t in enumerate(self.itos)}
    def __len__(self):
        return len(self.itos)
    def __getitem__(self, token):
        return self.stoi.get(token, -1)     # -1 for unknown tokens
    def __contains__(self, token):
        return token in self.stoi

v = Vocab("the cat sat on the mat".split())
print(len(v), v["cat"], v["dog"], "mat" in v)   # 5 0 -1 True

property, classmethod, staticmethod

class Temperature:
    def __init__(self, celsius):
        self.celsius = celsius

    @property
    def fahrenheit(self):                 # computed attribute, accessed without ()
        return self.celsius * 9 / 5 + 32

    @classmethod
    def from_fahrenheit(cls, f):          # alternative constructor, receives the class
        return cls((f - 32) * 5 / 9)

    @staticmethod
    def is_valid(c):                      # plain function namespaced in the class
        return c >= -273.15

t = Temperature.from_fahrenheit(212)
print(t.celsius, t.fahrenheit, Temperature.is_valid(-300))   # 100.0 212.0 False

Python has no truly private attributes. By convention _name means "internal, do not touch"; __name triggers name mangling to _ClassName__name to avoid clashes in subclasses.

Common pitfall Mutable class attributes. class Run: history = [] shares one list across every instance, so appending to run1.history also shows up in run2.history. Put per-instance state in __init__ as self.history = []. In PyTorch, forgetting super().__init__() in an nn.Module subclass causes an error as soon as you assign a layer, because the module's registries do not exist yet.
Interview angle Typical questions: what is self; __str__ versus __repr__; class versus instance attributes; classmethod versus staticmethod; what the MRO is; why you call model(x) and not model.forward(x) in PyTorch (because __call__ runs registered hooks around forward). Explaining composition over deep inheritance is a plus.

Robust Python: typing, dataclasses, exceptions and context managers

Research code often starts as a notebook, but production ML code must be readable, checkable and safe when things fail. Four features do most of that work: type hints (document and statically check intent), dataclasses (tidy configuration objects), exceptions (explicit failure handling) and context managers (guaranteed cleanup).

Analogy

Type hints are labels on the shelves of a warehouse: nothing stops you putting soup on the "screws" shelf, but inspectors (a type checker such as mypy) will flag it before customers notice. A context manager is a hotel key card that automatically deactivates at checkout, no matter whether you leave calmly or run out during a fire alarm. Mapping back: hints are not enforced at runtime but are checked by tools, and a with block guarantees its cleanup code (__exit__) runs even when an exception is raised inside.

Type hints

from typing import Optional, Union, Literal, Callable, Iterator, TypedDict, Protocol
import numpy as np

def top_k(scores: list[float], k: int = 5) -> list[int]:
    return sorted(range(len(scores)), key=scores.__getitem__, reverse=True)[:k]

def load(path: str, sep: Optional[str] = None) -> "pd.DataFrame": ...
Metric = Literal["accuracy", "f1", "recall"]        # only these strings allowed
Scorer = Callable[[np.ndarray, np.ndarray], float]  # a function type

class Sample(TypedDict):          # typed shape for a dict (e.g. JSON records)
    text: str
    label: int

class SupportsPredict(Protocol):  # structural typing: "anything with .predict"
    def predict(self, X: np.ndarray) -> np.ndarray: ...

print(top_k([0.1, 0.9, 0.4], k=2))   # [1, 2]

Since Python 3.10 you can write int | None instead of Optional[int]. Hints are ignored at runtime by the interpreter; libraries such as Pydantic and FastAPI read them to validate data, which is why LLM structured-output code relies on them (see LLM APIs).

Dataclasses for configs and records

from dataclasses import dataclass, field, asdict

@dataclass(frozen=True)              # frozen: fields cannot be reassigned
class TrainConfig:
    lr: float = 3e-4
    epochs: int = 10
    layers: list[int] = field(default_factory=lambda: [64, 32])  # safe mutable default
    seed: int = 42

cfg = TrainConfig(lr=1e-3)
print(cfg)           # TrainConfig(lr=0.001, epochs=10, layers=[64, 32], seed=42)
print(asdict(cfg)["epochs"], cfg == TrainConfig(lr=1e-3))   # 10 True
# cfg.lr = 0.1  -> raises FrozenInstanceError

@dataclass generates __init__, __repr__ and __eq__ from annotated fields. Use slots=True (3.10+) to save memory for millions of small objects. Choose Pydantic instead when you need runtime validation of untrusted input such as API payloads or LLM output.

Exceptions

import json

class DataValidationError(ValueError):
    """Raised when an input file fails schema checks."""

def load_config(path):
    try:
        with open(path, encoding="utf-8") as f:
            cfg = json.load(f)
    except FileNotFoundError:
        return {"lr": 1e-3}                         # sensible default
    except json.JSONDecodeError as e:
        raise DataValidationError(f"bad JSON in {path}") from e   # keep the cause
    else:
        return cfg                                  # runs only if no exception
    finally:
        print("config load attempted")              # always runs
  • Catch the most specific exception you can handle; never write a bare except:, which also swallows KeyboardInterrupt.
  • raise ... from e chains exceptions so the traceback shows the root cause.
  • Python style is EAFP ("easier to ask forgiveness than permission"): try the operation and catch failure, rather than checking every precondition first.
  • Use assert for internal sanity checks (shapes during development) but not for validating user input, because asserts are removed when Python runs with -O.

Context managers

import time
from contextlib import contextmanager

@contextmanager
def timed(label):
    start = time.perf_counter()
    try:
        yield                              # the body of the with-block runs here
    finally:
        print(f"{label}: {time.perf_counter() - start:.2f}s")

with timed("feature engineering"):
    time.sleep(0.1)                        # feature engineering: 0.10s

class Timer:                               # the same idea, class-based
    def __enter__(self):
        self.start = time.perf_counter()
        return self
    def __exit__(self, exc_type, exc, tb):
        self.elapsed = time.perf_counter() - self.start
        return False                       # False: do not swallow exceptions

with Timer() as t:
    sum(range(10**6))
print(t.elapsed < 1)                       # True

ML examples: with open(...), with torch.no_grad():, with torch.autocast("cuda"):, with pd.option_context("display.max_rows", 5):, with mlflow.start_run():.

Common pitfall Swallowing errors inside a long training job: except Exception: pass around a batch hides corrupted data and makes the model silently train on less data. Log the exception with context (batch index, file name), count failures, and fail loudly past a threshold.
Interview angle "How does a with statement work?" Expect to describe __enter__/__exit__ and that __exit__ receives the exception and can suppress it by returning True. Also common: else versus finally in try blocks, dataclass versus namedtuple versus Pydantic, and whether type hints are enforced (no, unless a tool or library checks them).

Concurrency: the GIL, threads, processes and asyncio

CPython has a Global Interpreter Lock (GIL): only one thread executes Python bytecode at a time within a process. Threads therefore do not speed up pure-Python number crunching, but they do help when threads spend most of their time waiting (network, disk), because waiting threads release the GIL. Heavy NumPy, pandas and PyTorch operations also release the GIL while their C code runs. For CPU-bound pure-Python work, use multiple processes, each with its own interpreter and GIL.

Analogy

Picture a kitchen with one chef's knife (the GIL). Several cooks (threads) can share the kitchen, but only the one holding the knife can chop. If most of the work is waiting for water to boil (I/O), sharing one knife is fine and the kitchen stays busy. If the work is all chopping (CPU), extra cooks just queue for the knife, so you need more kitchens (processes). asyncio is one very organised cook who starts every pot boiling and checks each when it is ready. Mapping back: threads for I/O-bound work, processes for CPU-bound Python work, and asyncio for thousands of concurrent network calls such as LLM API requests.

ApproachParallel CPU?Best forCosts
threading / ThreadPoolExecutorNo for Python code (GIL), yes inside C code that releases itDownloading files, database or API calls, reading many filesRace conditions; shared state needs locks
multiprocessing / ProcessPoolExecutorYesCPU-heavy Python: feature engineering, parsing, simulationsProcess start-up, pickling data between processes, more memory
asyncioNo (single thread)Very many concurrent I/O tasks: LLM calls, web scraping, servicesNeeds async-aware libraries; one blocking call stalls everything
Vectorized libraries / GPUYes, internallyNumerical work on arraysMust reshape the problem into array operations
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor
import math, time

def fetch(url):                    # I/O-bound: mostly waiting
    time.sleep(0.5)
    return len(url)

def heavy(n):                      # CPU-bound pure Python
    return sum(math.isqrt(i) for i in range(n))

if __name__ == "__main__":         # required for processes on Windows and macOS
    urls = [f"doc{i}" for i in range(8)]
    with ThreadPoolExecutor(max_workers=8) as ex:
        sizes = list(ex.map(fetch, urls))       # about 0.5 s total instead of 4 s
    with ProcessPoolExecutor() as ex:
        totals = list(ex.map(heavy, [2_000_000] * 4))   # uses several CPU cores

asyncio for many API calls

import asyncio

async def call_llm(prompt, sem):
    async with sem:                         # cap concurrency (rate limits)
        await asyncio.sleep(0.2)            # stands in for an async HTTP call
        return prompt.upper()

async def main(prompts):
    sem = asyncio.Semaphore(5)
    return await asyncio.gather(*(call_llm(p, sem) for p in prompts))

print(asyncio.run(main(["hi", "hello", "hey"])))   # ['HI', 'HELLO', 'HEY']

In Jupyter an event loop is already running, so use await main(prompts) directly in a cell instead of asyncio.run.

Note Recent CPython versions (3.13 onwards) offer an optional "free-threaded" build without the GIL. It is opt-in and many extension libraries are still catching up, so in interviews describe the GIL as the default behaviour and mention free-threading as the direction of travel.
Common pitfall Calling a blocking function (such as time.sleep or a synchronous HTTP client) inside an async def freezes the whole event loop, so you get no concurrency at all. Use the async client or push the blocking call to a thread with asyncio.to_thread. With multiprocessing, forgetting the __main__ guard on Windows spawns processes recursively.
Interview angle "What is the GIL and how do you get parallelism in Python?" A complete answer: the GIL serialises bytecode execution per process; it protects reference counting; I/O and many C extensions release it; use processes for CPU-bound Python, threads or asyncio for I/O-bound, and vectorized or GPU libraries for numeric work. Tie it to ML: PyTorch DataLoader(num_workers=4) uses worker processes to decode and augment data in parallel.

NumPy: arrays, vectorization and broadcasting

A NumPy ndarray is a block of memory holding many values of one type, plus metadata describing its shape and how to step through memory (strides). Because every element has the same type and sits next to its neighbours, NumPy can hand the whole block to optimized C loops and CPU vector instructions. A Python list, by contrast, is an array of pointers to separate objects scattered around memory, each carrying its own type information.

Analogy

A Python list is a shelf of assorted parcels, each with its own label that must be read before you can open it. A NumPy array is an egg carton: identical slots, identical contents, so a machine can process the whole carton in one sweep. Broadcasting is like a stamp that says "add one pinch of salt to every egg in the row": you describe the operation once for a smaller shape and NumPy repeats it across the larger one without making copies. Mapping back: homogeneous dtype and contiguous memory enable vectorized ufuncs, and broadcasting stretches size-1 dimensions virtually instead of physically duplicating data.

Creating arrays and reading their metadata

import numpy as np
a = np.array([1, 2, 3, 4])
print(a.dtype, a.shape, a.ndim, a.itemsize, a.nbytes)   # int64 (4,) 1 8 32
print(np.array([1, 2.5]).dtype)        # float64  (upcast to a common type)
print(np.array([1, "a"]).dtype)        # <U21     (everything became a string!)

np.zeros((2, 3)); np.ones(4); np.full((2, 2), 7); np.eye(2)
print(np.arange(0, 1, 0.25))           # [0.   0.25 0.5  0.75]  (stop excluded)
print(np.linspace(0, 1, 5))            # [0.   0.25 0.5  0.75 1.  ]  (stop included)
x = np.arange(12).reshape(3, 4)        # reshape keeps the same data
print(x)
# [[ 0  1  2  3]
#  [ 4  5  6  7]
#  [ 8  9 10 11]]

dtypes that matter in ML

dtypeBytesUse
float648NumPy and pandas default; scientific precision
float324Deep learning default; halves memory
float16 / bfloat162Mixed-precision training and inference (bfloat16 lives in PyTorch, not core NumPy)
int64, int32, int8, uint88 / 4 / 1 / 1Labels, token IDs, image pixels (0 to 255), quantized weights
bool1Masks
img = np.array([200, 100], dtype=np.uint8)
print(img + img)                         # [144 200]: 400 wrapped around modulo 256!
print(img.astype(np.int32) + img)        # [400 200]  cast before arithmetic
print(np.array([127], dtype=np.int8) + 1)   # [-128]  silent overflow in arrays
X32 = np.random.default_rng(0).normal(size=(1000, 100)).astype(np.float32)   # half the RAM of float64

Vectorization and ufuncs

A ufunc (universal function) applies an operation element-wise in compiled code: arithmetic operators, np.exp, np.log, np.maximum, comparisons. Replacing a Python loop with one ufunc call typically speeds things up 10 to 100 times.

x = np.random.default_rng(0).normal(size=1_000_000)

# slow: interpreted loop, one Python float object per element
out = [max(0.0, v) for v in x]

# fast: one call into C
relu = np.maximum(0, x)
sigmoid = 1 / (1 + np.exp(-x))
clipped = np.clip(x, -1, 1)
labels = np.where(x > 0, 1, 0)          # vectorized if/else

# in a notebook:
# %timeit [max(0.0, v) for v in x]       # on the order of 100 ms
# %timeit np.maximum(0, x)               # on the order of 1 ms

Indexing: basic, boolean and fancy

x = np.arange(12).reshape(3, 4)
print(x[1, 2])          # 6            row 1, column 2
print(x[:, 1])          # [1 5 9]      every row, column 1
print(x[1:, ::2])       # [[ 4  6] [ 8 10]]  rows 1.., every 2nd column
print(x[-1])            # [ 8  9 10 11] last row
print(x[x % 2 == 0])    # [ 0  2  4  6  8 10]  boolean mask -> 1-D result
print(x[[0, 2], [1, 3]])  # [ 1 11]  fancy: picks (0,1) and (2,3)
print(x[[0, 2]])        # rows 0 and 2
mask = (x > 2) & (x < 8)   # combine masks with & | ~ and parentheses, not and/or

Views versus copies

Basic slicing (start:stop:step) returns a view: a new array object that shares the original memory. Fancy (integer-list) and boolean indexing return copies. reshape returns a view when possible; .copy() always copies.

b = np.arange(6)
v = b[1:4]              # view
v[0] = 99
print(b)                # [ 0 99  2  3  4  5]   original changed!
c = b[[1, 2]]           # copy (fancy indexing)
c[0] = -1
print(b)                # [ 0 99  2  3  4  5]   unchanged
print(np.shares_memory(b, v), np.shares_memory(b, c))   # True False
safe = b[1:4].copy()    # explicit copy when you plan to modify

Axes and aggregation

The axis argument names the dimension that gets collapsed. For a table of shape (rows, columns), axis=0 aggregates down the rows and returns one value per column; axis=1 aggregates across columns and returns one value per row.

          axis=1  -->
        +----+----+----+----+
axis=0  |  0 |  1 |  2 |  3 |   sum(axis=1) ->  6
  |     +----+----+----+----+
  v     |  4 |  5 |  6 |  7 |                  22
        +----+----+----+----+
        |  8 |  9 | 10 | 11 |                  38
        +----+----+----+----+
sum(axis=0) -> 12   15   18   21
x = np.arange(12).reshape(3, 4)
print(x.sum(axis=0))                       # [12 15 18 21]  per column
print(x.sum(axis=1))                       # [ 6 22 38]     per row
print(x.mean(axis=1, keepdims=True).shape) # (3, 1)  keeps a broadcastable shape
print(np.argmax([[1, 9], [8, 2]], axis=1)) # [1 0]   index of the max in each row
print(np.argsort([30, 10, 20]))            # [1 2 0] indices that would sort
vals, counts = np.unique([3, 1, 3, 2], return_counts=True)   # [1 2 3] [1 1 2]

Broadcasting

When two arrays have different shapes, NumPy compares their shapes from the right. Two dimensions are compatible if they are equal or one of them is 1 (a missing leading dimension counts as 1). Size-1 dimensions are stretched virtually to match.

A      (8, 1, 6, 1)
B         (7, 1, 5)
Result (8, 7, 6, 5)

X      (100, 3)        data: 100 samples, 3 features
mean        (3,)        per-feature mean
X - mean  (100, 3)      centred data, no loop, no copy of mean
M = np.array([[1, 2, 3], [4, 5, 6]])
print(M - M.mean(axis=0))          # centre each column
# [[-1.5 -1.5 -1.5]
#  [ 1.5  1.5  1.5]]
col = np.array([[10], [20]])       # shape (2, 1)
row = np.array([1, 2, 3])          # shape (3,)
print(col + row)                   # shape (2, 3): [[11 12 13] [21 22 23]]

# standardize features (what StandardScaler does)
X = np.random.default_rng(0).normal(5, 2, size=(100, 3))
Xs = (X - X.mean(axis=0)) / X.std(axis=0)

# pairwise squared distances between points (n, d) and centroids (k, d)
P, C = np.random.rand(5, 2), np.random.rand(3, 2)
D = ((P[:, None, :] - C[None, :, :]) ** 2).sum(axis=-1)   # (5, 1, 2) - (1, 3, 2) -> (5, 3)

np.ones(3) + np.ones(4)
# ValueError: operands could not be broadcast together with shapes (3,) (4,)

Use arr[:, None] or np.newaxis or np.expand_dims to add a size-1 axis deliberately; use squeeze to remove them.

Combining and reshaping

X = np.array([[1., 2.], [3., 4.]])
np.concatenate([X, X], axis=0).shape   # (4, 2)  join along an existing axis
np.stack([X, X]).shape                 # (2, 2, 2)  create a new axis
np.hstack([X, X]).shape                # (2, 4)
X.reshape(-1)                          # flatten (view if possible); -1 means "infer"
X.ravel(); X.flatten()                 # ravel: view if possible; flatten: always a copy
X.T                                    # transpose is a view (not contiguous)

Linear algebra

X = np.array([[1., 2.], [3., 4.]])
print(X @ X)             # matrix product  [[ 7. 10.] [15. 22.]]
print(X * X)             # element-wise    [[ 1.  4.] [ 9. 16.]]

A = np.array([[3., 1.], [1., 2.]]); b = np.array([9., 8.])
print(np.linalg.solve(A, b))            # [2. 3.]   prefer solve over inv(A) @ b
print(np.linalg.norm([3, 4]))           # 5.0
print(np.linalg.eigh(A)[0].round(3))    # [1.382 3.618]  eigenvalues of a symmetric matrix
U, S, Vt = np.linalg.svd(np.random.rand(4, 3), full_matrices=False)  # basis of PCA
w, *_ = np.linalg.lstsq(np.c_[np.ones(4), [1, 2, 3, 4]], [3, 5, 7, 9], rcond=None)
print(w.round(3))                       # [1. 2.]   intercept 1, slope 2 (y = 1 + 2x)

# one dense neural-network layer
Xb = np.random.default_rng(0).normal(size=(4, 3))   # 4 samples, 3 features
W = np.random.default_rng(1).normal(size=(3, 2)); bias = np.array([0.1, -0.2])
Z = Xb @ W + bias        # (4,3) @ (3,2) = (4,2); bias broadcasts across rows

Statistics and a numerically stable softmax

d = [22, 18, 14, 10, 15, 20, 25, 12]
print(np.mean(d), np.median(d))                  # 17.0 16.5
print(np.var(d), np.var(d, ddof=1).round(3))     # 23.25 26.571  population vs sample
print(np.std(d).round(3))                        # 4.822
print(np.percentile([15, 18, 67, 100, 21, 50], [25, 50, 75]))   # [18.75 35.5  62.75]
print(np.quantile([15, 18, 67, 100, 21, 50], 1/3))              # 20.0 (linear interpolation)
print(np.nan == np.nan, np.mean([1, np.nan, 3]), np.nanmean([1, np.nan, 3]))   # False nan 2.0

def softmax(z, axis=-1):
    z = z - z.max(axis=axis, keepdims=True)      # subtracting the max prevents overflow
    e = np.exp(z)
    return e / e.sum(axis=axis, keepdims=True)
print(softmax(np.array([1000., 1001., 1002.])).round(3))   # [0.09  0.245 0.665]

Note that NumPy's var and std default to the population formula (ddof=0) while pandas defaults to the sample formula (ddof=1), so the same column can give slightly different answers in the two libraries. The statistics themselves are covered in Math & Statistics for ML.

Random numbers done right

rng = np.random.default_rng(42)        # modern Generator API, local and seedable
rng.normal(0, 1, size=(2, 3))          # Gaussian
rng.uniform(0, 1, size=5)
rng.integers(0, 10, size=5)            # high end excluded
rng.choice(["a", "b", "c"], size=4, p=[0.5, 0.3, 0.2])
idx = rng.permutation(100)             # shuffled indices for a manual split
train_idx, test_idx = idx[:80], idx[80:]

Prefer default_rng over the legacy global np.random.seed: a generator object can be passed around explicitly, so two parts of a program do not silently share and disturb one global random state.

Common pitfall Shape (n,) versus (n, 1). A 1-D array of shape (5,) minus a column of shape (5, 1) broadcasts to a (5, 5) matrix rather than raising an error, which silently produces a wrong loss. Print shapes at every step, and add assert y_pred.shape == y_true.shape in metric code. Other classics: using and/or on arrays (use &, | with parentheses), integer overflow in uint8 images, and modifying a slice that turned out to be a view.
Interview angle Expect: why NumPy is faster than lists (contiguous typed memory, C loops, SIMD, no per-element type checks), state the broadcasting rule and predict a result shape, view versus copy, what axis=0 means, * versus @, and "implement softmax, standardization, cosine similarity or pairwise distances without loops". Numerical stability (subtract the max before exp) is a frequent follow-up.

pandas: working with tables

pandas wraps NumPy arrays with labels. A Series is a one-dimensional array with an index (row labels). A DataFrame is a collection of Series that share the same index, one per column, and each column can have its own dtype. Almost every operation aligns on labels, which is powerful (joins and arithmetic just work) and occasionally surprising (mismatched indexes produce NaN).

Analogy

A DataFrame is a spreadsheet with superpowers. Each column is a typed list, the row numbers on the left are the index, and every operation you would do with filters, pivot tables and VLOOKUP has a one-line equivalent. groupby is sorting receipts into envelopes by store, then totalling each envelope. Mapping back: filtering is a boolean mask over rows, pivot tables are pivot_table, VLOOKUP is merge, and "split, apply, combine" is groupby().agg().

Series and DataFrame basics

import pandas as pd, numpy as np
s = pd.Series([10, 20, 30], index=["a", "b", "c"])
print(s["b"], s.iloc[-1], s[s > 15].index.tolist())   # 20 30 ['b', 'c']

df = pd.DataFrame({
    "name":   ["Asha", "Ben", "Chen", "Dia"],
    "age":    [34, np.nan, 29, 41],
    "dept":   ["ml", "web", "ml", None],
    "salary": [120, 80, 110, 95],
})
df.shape            # (4, 4)
df.dtypes           # name: str (object before pandas 3), age: float64, dept: str, salary: int64
df.head(); df.tail(2); df.sample(2, random_state=0)
df.info()           # dtypes, non-null counts, memory
df.describe()       # count, mean, std, min, quartiles, max for numeric columns
df.columns.tolist(); df.index

Notice that age became float64: a classic NumPy-backed integer column cannot hold NaN, so pandas upcasts it. Use the nullable dtype "Int64" (capital I) if you need integers with missing values; it displays missing entries as <NA>.

Loading and saving data

df = pd.read_csv("data.csv",
                 usecols=["id", "ts", "amount", "city"],   # read only what you need
                 dtype={"city": "category", "amount": "float32"},
                 parse_dates=["ts"],
                 na_values=["", "NA", "?"])
pd.read_parquet("data.parquet")        # columnar, typed, compressed: best for big data
pd.read_json("records.jsonl", lines=True)
pd.read_excel("sheet.xlsx", sheet_name="Q1")
pd.read_sql("SELECT * FROM orders", connection)
for chunk in pd.read_csv("huge.csv", chunksize=100_000):   # stream a file bigger than RAM
    process(chunk)
df.to_parquet("clean.parquet", index=False)
df.to_csv("out.csv", index=False)      # index=False avoids an extra unnamed column

Selecting: [], loc and iloc

SyntaxSelects bySlice endExample
df["col"] / df[["a", "b"]]Column name(s)n/aA Series / a DataFrame
df.loc[rows, cols]Labels and boolean masksInclusivedf.loc[df.age > 30, ["name", "salary"]]
df.iloc[rows, cols]Integer positionsExclusivedf.iloc[:10, 0:3]
df.at[r, c] / df.iat[i, j]Single scalar, fastn/adf.at[0, "name"]
print(df.loc[0:2, "salary"].tolist())    # [120, 80, 110]  labels 0..2 INCLUDED
print(df.iloc[0:2]["salary"].tolist())   # [120, 80]       positions 0,1 (2 excluded)
print(df.loc[df.age > 30, ["name", "salary"]])
#    name  salary
# 0  Asha     120
# 3   Dia      95

Filtering and creating columns

high = df[(df.salary > 90) & (df.dept == "ml")]       # & | ~ with parentheses
df[df.dept.isin(["ml", "data"])]
df[df.name.str.startswith("A")]
df.query("salary > 90 and dept == 'ml'")               # readable string filter
df["salary_k"] = df.salary * 1000                      # vectorized new column
df = df.assign(ratio=lambda t: t.salary / t.salary.max())  # chain-friendly
df["band"] = np.where(df.salary >= 100, "high", "normal")
df["tier"] = np.select([df.salary < 90, df.salary < 110], ["low", "mid"], default="top")
df = df.rename(columns={"salary": "pay"}).drop(columns=["salary_k"])
df.sort_values(["dept", "pay"], ascending=[True, False])
df.nlargest(2, "pay")

groupby: split, apply, combine

rows            split by key        apply (sum)      combine
Pune   250  -->  Pune:  250, 150  -->  400    --\
Delhi  400  -->  Delhi: 400, 100  -->  500    ---->  Delhi 500 | Mumbai 300 | Pune 400
Pune   150       Mumbai: 300      -->  300    --/
Mumbai 300
Delhi  100
sales = pd.DataFrame({
    "city":  ["Pune", "Delhi", "Pune", "Mumbai", "Delhi"],
    "sales": [250, 400, 150, 300, 100],
    "units": [5, 8, 3, 6, 2],
})
print(sales.groupby("city")["sales"].sum())
# city
# Delhi     500
# Mumbai    300
# Pune      400
# Name: sales, dtype: int64

print(sales.groupby("city").agg(total=("sales", "sum"), avg_units=("units", "mean")))
#         total  avg_units
# city
# Delhi     500        5.0
# Mumbai    300        6.0
# Pune      400        4.0

# transform returns a result aligned to the ORIGINAL rows (same length)
sales["share"] = sales.sales / sales.groupby("city")["sales"].transform("sum")
print(sales.share.tolist())      # [0.625, 0.8, 0.375, 1.0, 0.2]

sales.groupby("city", as_index=False)["sales"].sum()       # key stays a column
sales.groupby("city").filter(lambda g: g.sales.sum() > 350)  # keep whole groups
sales.groupby(["city", "units"]).size()                       # multi-key counts
sales.city.value_counts(normalize=True)                       # class proportions
MethodOutput shapeUse for
aggOne row per groupSummaries: sum, mean, count, custom reductions
transformSame length as inputGroup-relative features: share of group, z-score within group, fill with group mean
filterSubset of original rowsDrop small or irrelevant groups
applyAnythingFlexible but slow; last resort

Combining tables: merge, join, concat

left = pd.DataFrame({"k": [1, 2, 3], "a": ["x", "y", "z"]})
right = pd.DataFrame({"k": [2, 3, 4], "b": [True, False, True]})
print(pd.merge(left, right, on="k", how="left", indicator=True))
#    k  a      b     _merge
# 0  1  x    NaN  left_only
# 1  2  y   True       both
# 2  3  z  False       both

pd.merge(orders, customers, on="customer_id", how="left",
         validate="many_to_one")        # raises if customers has duplicate keys
pd.merge(a, b, left_on="uid", right_on="user_id", suffixes=("_a", "_b"))
a.join(b, how="inner")                  # join on the index
pd.concat([jan, feb], ignore_index=True)  # stack rows (same columns)
pd.concat([X_num, X_cat], axis=1)       # side by side (aligns on index!)
how=KeepsSQL equivalent
inner (default)Keys present in bothINNER JOIN
leftAll left keys; unmatched right columns become NaNLEFT JOIN
rightAll right keysRIGHT JOIN
outerUnion of keysFULL OUTER JOIN
crossEvery combinationCROSS JOIN

Reshaping: pivot, pivot_table, melt, crosstab

wide = pd.DataFrame({"id": [1, 2], "jan": [10, 20], "feb": [11, 21]})
long = wide.melt(id_vars="id", var_name="month", value_name="sales")   # wide -> long
print(long)
#    id month  sales
# 0   1   jan     10
# 1   2   jan     20
# 2   1   feb     11
# 3   2   feb     21
back = long.pivot(index="id", columns="month", values="sales")         # long -> wide

rev = pd.DataFrame({"region": ["N", "N", "S", "S"], "q": ["Q1", "Q2", "Q1", "Q2"], "rev": [10, 15, 7, 9]})
print(rev.pivot_table(index="region", columns="q", values="rev", aggfunc="sum", margins=True))
# q       Q1  Q2  All
# region
# N       10  15   25
# S        7   9   16
# All     17  24   41
pd.crosstab(rev.region, rev.q)          # frequency table of two categoricals

pivot fails if an index/column pair appears twice; pivot_table aggregates duplicates instead. stack/unstack move levels between the row index and the columns.

Missing values

df.isna().sum()                     # missing count per column
df.isna().mean().sort_values()      # missing fraction, worst last
df.dropna()                         # drop rows with ANY missing value
df.dropna(subset=["age"])           # only rows missing age
df.dropna(axis=1, thresh=int(0.7 * len(df)))   # keep columns at least 70% filled
df.fillna({"age": df.age.median(), "dept": "unknown"})
df.age.fillna(df.groupby("dept").age.transform("median"))   # group-aware imputation
ts.ffill(); ts.bfill(); ts.interpolate()   # time series: forward-fill, back-fill, interpolate
df["age_missing"] = df.age.isna().astype(int)   # missingness can itself be informative

For modelling, prefer imputing inside a scikit-learn pipeline (SimpleImputer) so the fill values are learned from the training fold only. Filling the whole DataFrame with the global median before splitting is a small but real form of leakage.

apply versus vectorized operations

ApproachSpeedExample
Vectorized column arithmeticFastestdf.a * df.b, np.log1p(df.x)
Built-in accessorsFastdf.s.str.lower(), df.t.dt.hour
np.where / np.select / map with a dictFastdf.code.map({"A": 1, "B": 2})
Series.apply(func)Slow (Python per element)Complex custom logic
df.apply(func, axis=1)Slower (builds a Series per row)Row logic spanning many columns
iterrows()Slowest; also loses dtypesAlmost never
# slow
df["bmi"] = df.apply(lambda r: r.weight / r.height ** 2, axis=1)
# fast (often 100x or more on large frames)
df["bmi"] = df.weight / df.height ** 2
# if you truly need a loop, itertuples is much faster than iterrows
for row in df.itertuples(index=False):
    pass

Strings, dates and times

names = pd.Series(["a b", "c"])
print(names.str.split().str.len().tolist(), names.str.upper().tolist())   # [2, 1] ['A B', 'C']
df.email.str.extract(r"@(\w+)\.")        # regex capture group to a column
df.text.str.contains("refund", case=False, na=False)

ev = pd.DataFrame({
    "ts": pd.to_datetime(["2026-01-01 09:00", "2026-01-01 17:00", "2026-01-02 10:00", "2026-01-04 12:00"]),
    "clicks": [3, 5, 2, 7]})
ev["dow"] = ev.ts.dt.day_name(); ev["hour"] = ev.ts.dt.hour   # Thursday 9, Thursday 17, ...
print(ev.set_index("ts").resample("D")["clicks"].sum().tolist())   # [8, 2, 0, 7]  gap day filled with 0
print(ev.clicks.rolling(2).mean().tolist())    # [nan, 4.0, 3.5, 4.5]
print(ev.clicks.shift(1).tolist())             # [nan, 3.0, 5.0, 2.0]  lag feature
print(ev.clicks.diff().tolist())               # [nan, 2.0, -3.0, 5.0]
pd.to_datetime(raw, format="%d/%m/%Y", errors="coerce")   # explicit format; bad values -> NaT
ev.ts.dt.tz_localize("UTC").dt.tz_convert("Asia/Kolkata")
Common pitfall Rolling and lag features must only look backwards. A centred rolling window (center=True) or a shift(-1) uses the future and leaks the target into features. Sort by time and by entity before computing lags, and compute them per entity with groupby("user").x.shift(1).

Binning and categoricals

ages = [5, 17, 30, 64, 80]
print(pd.cut(ages, bins=[0, 18, 60, 120], labels=["child", "adult", "senior"]).tolist())
# ['child', 'child', 'adult', 'senior', 'senior']   fixed edges, (0, 18] includes 18
print(pd.qcut(pd.Series([1, 2, 3, 4, 100]), 2, labels=["lo", "hi"]).tolist())
# ['lo', 'lo', 'lo', 'hi', 'hi']                    equal-frequency bins

c = pd.Series(["lo", "hi", "lo"], dtype="category")
print(c.cat.categories.tolist(), c.cat.codes.tolist())   # ['hi', 'lo'] [1, 0, 1]
size = pd.Categorical(["M", "S", "L"], categories=["S", "M", "L"], ordered=True)   # ordered: S < M < L
pd.get_dummies(df.dept, prefix="dept", dtype=int)        # one-hot for quick analysis

A category column stores each distinct value once plus a small integer code per row, which slashes memory for repetitive strings (city, product, country) and speeds up groupby.

Memory optimisation

big = pd.DataFrame({"i": np.arange(1000, dtype="int64"), "f": np.random.rand(1000)})
print(big.memory_usage(deep=True).sum())    # 16132 bytes (index + 8 + 8 bytes per row)
big["i"] = pd.to_numeric(big.i, downcast="integer")   # int16 is enough for 0..999
big["f"] = big.f.astype("float32")
print(big.memory_usage(deep=True).sum())    # 6132 bytes, about 62% less

def shrink(df):
    for col in df.select_dtypes("integer"):
        df[col] = pd.to_numeric(df[col], downcast="integer")
    for col in df.select_dtypes("float"):
        df[col] = pd.to_numeric(df[col], downcast="float")
    for col in df.select_dtypes(["object", "string"]):
        if df[col].nunique() / len(df) < 0.5:      # repetitive text -> category
            df[col] = df[col].astype("category")
    return df
  1. Measure df.info(memory_usage="deep") and df.memory_usage(deep=True).
  2. Read less usecols, nrows for prototyping, filter early.
  3. Type tightly Downcast numbers, category for low-cardinality strings, Arrow-backed strings.
  4. Store columnar Parquet keeps dtypes and lets you read only some columns.
  5. Stream or scale out chunksize, or move to Polars, DuckDB, Dask or Spark when data truly exceeds one machine's memory.

Copies, views and chained assignment

# WRONG: chained indexing. The first [] may return a copy, so the write is lost.
df[df.salary > 90]["salary"] = 0      # warning; original df unchanged
# RIGHT: one .loc call selects rows and column together
df.loc[df.salary > 90, "salary"] = 0
# taking a subset you intend to modify independently
sub = df[df.salary > 90].copy()

Older pandas emits SettingWithCopyWarning for ambiguous cases. With Copy-on-Write (the default behaviour from pandas 3.0), any derived object behaves like a copy, chained assignment never modifies the original, and pandas warns with a ChainedAssignmentError. The fix is the same in every version: one .loc assignment.

Common pitfall Index alignment surprises. pd.concat([X_num, X_cat], axis=1) after one frame was filtered or shuffled aligns rows by index label, filling mismatches with NaN and silently doubling the row count. Reset indexes with reset_index(drop=True) when you mean "by position". Also watch for many-to-many merges that explode row counts; use validate= and compare len() before and after.
Interview angle Common prompts: loc versus iloc (labels, inclusive end, versus positions, exclusive end); merge types; agg versus transform; handling missing values; why apply is slow; how to reduce memory; SettingWithCopy. Live tasks: "top 3 products by revenue per region", "fraction of missing values per column", "day-over-day change per user", "pivot this long table". Show the idiomatic vectorized answer, then mention edge cases (ties, NaN, duplicates).
# "top 2 cities by total sales" and "top 1 row per group"
sales.groupby("city")["sales"].sum().nlargest(2)
sales.sort_values("sales", ascending=False).groupby("city").head(1)

Visualization: matplotlib and seaborn

matplotlib is the low-level plotting engine: a Figure is the canvas, and each Axes is one plot area on it with its own x-axis, y-axis, title and legend. seaborn sits on top and draws statistical plots straight from DataFrames with sensible defaults, colouring by a column with a single hue= argument. Pick the plot from the question you are asking, not from habit.

Analogy

matplotlib is a set of drawing instruments and a sheet of paper: total control, but you draw every line yourself. seaborn is a template book of standard charts: open to "compare distributions by group", hand it your table, and it draws the chart properly. Mapping back: fig, ax = plt.subplots() gives you the paper and a drawing area, and seaborn functions accept ax= so you can place their charts onto your own layout and still tweak them with matplotlib calls.

The object-oriented matplotlib pattern

import numpy as np, matplotlib.pyplot as plt
x = np.linspace(-6, 6, 200)
fig, axes = plt.subplots(1, 2, figsize=(10, 3.5))       # 1 row, 2 plots
axes[0].plot(x, 1 / (1 + np.exp(-x)), label="sigmoid")
axes[0].plot(x, np.tanh(x), label="tanh", linestyle="--")
axes[0].axhline(0, color="gray", lw=0.8)
axes[0].set(title="Activations", xlabel="z", ylabel="output")
axes[0].legend()
axes[1].hist(np.random.default_rng(0).normal(size=1000), bins=30, alpha=0.7)
axes[1].set_title("Histogram")
fig.tight_layout()
fig.savefig("plots.png", dpi=150)                       # or plt.show() in scripts

seaborn one-liners

import seaborn as sns
sns.histplot(df, x="mean radius", hue="diagnosis", kde=True)   # distribution by class
sns.boxplot(df, x="diagnosis", y="mean area")                  # spread and outliers
sns.scatterplot(df, x="mean radius", y="mean texture", hue="diagnosis")
sns.countplot(df, x="diagnosis")                               # class balance
sns.heatmap(df.corr(numeric_only=True), cmap="coolwarm", center=0)   # correlations
sns.pairplot(df[cols + ["diagnosis"]], hue="diagnosis")        # all pairs (small cols only)
sns.lineplot(history, x="epoch", y="loss", hue="split")        # learning curves

Which plot answers which question

QuestionPlotWatch out for
What does one numeric variable look like?Histogram, KDE, box plot, violinBin count changes the story; use log scale for heavy tails
How balanced are my classes or categories?Count / bar plotSort bars; show percentages
Do two numeric variables move together?Scatter (hexbin or alpha for many points)Overplotting hides density
How does a numeric variable differ across groups?Box, violin, strip, grouped histogramUnequal group sizes
How does something change over time?Line plotIrregular timestamps; resample first
Which features are correlated?Correlation heatmapPearson only captures linear relationships
How is my model doing?Confusion matrix, ROC, precision-recall curve, calibration plot, residual plot, learning curveROC looks optimistic under heavy imbalance; prefer PR curves
What drives the model?Horizontal bar of coefficients or importances, partial dependenceImpurity importances favour high-cardinality features

Model diagnostic plots

from sklearn.metrics import ConfusionMatrixDisplay, RocCurveDisplay, PrecisionRecallDisplay
ConfusionMatrixDisplay.from_estimator(model, X_test, y_test)
RocCurveDisplay.from_estimator(model, X_test, y_test)
PrecisionRecallDisplay.from_estimator(model, X_test, y_test)

# learning curves from a PyTorch or Keras history
plt.plot(train_losses, label="train"); plt.plot(val_losses, label="val"); plt.legend()
# val loss rising while train keeps falling = overfitting: stop earlier or regularize
Common pitfall Mixing the stateful plt.* interface with the object-oriented ax.* interface in one figure, so labels land on the wrong subplot. Pick fig, ax = plt.subplots() and stick with ax. Other mistakes: truncated y-axes that exaggerate differences, pie charts with many slices, and pairplots on 30 columns that take minutes and show nothing readable.
Interview angle Interviewers rarely ask for plotting syntax; they ask "which plot would you use to check X" and "what would you look for". Good answers link the plot to a decision: skewed histogram leads to a log transform; class-count plot shows imbalance and leads to stratification and PR metrics; a correlation heatmap reveals redundant features or a suspiciously perfect correlation that signals leakage.

Exploratory data analysis (EDA) workflow

EDA is the structured interrogation of a dataset before modelling. Its goals are to understand what each column means, find data-quality problems, spot relationships worth modelling, and catch leakage before it inflates your metrics. Good EDA is a checklist you run every time, then a few targeted deep dives.

Analogy

EDA is a pre-flight inspection. A pilot walks around the aircraft checking tyres, fuel, and flaps before take-off, following the same checklist every flight, because problems are cheap on the ground and catastrophic in the air. Mapping back: checking shape, types, missing values, duplicates, distributions, class balance and leakage before training is cheap; discovering them after a model is deployed is expensive.

  1. Understand the question What is the target, what is the unit of prediction (patient, transaction, day), and what data is available at prediction time?
  2. Structure shape, dtypes, head(), info(). Are numbers stored as strings? Are dates parsed?
  3. Quality Missing values per column, duplicates (df.duplicated().sum()), impossible values (negative ages), inconsistent categories ("NY" versus "New York").
  4. Univariate describe(), histograms, value_counts(); note skew, outliers and scale differences.
  5. Target Class balance for classification; distribution and skew for regression.
  6. Bivariate Each feature against the target: grouped means, box plots, correlation with the target.
  7. Multivariate Correlation heatmap, multicollinearity, a quick PCA scatter.
  8. Leakage hunt Features that are near-perfect predictors, IDs correlated with the target, columns created after the outcome, duplicate entities across train and test.
  9. Decide Write down the preprocessing each finding implies (impute, log-transform, encode, drop, stratify, group split).

A reusable EDA snippet

def quick_eda(df, target):
    print("shape:", df.shape)
    print("duplicates:", df.duplicated().sum())
    report = pd.DataFrame({
        "dtype": df.dtypes.astype(str),
        "missing_%": (df.isna().mean() * 100).round(1),
        "unique": df.nunique(),
    })
    print(report.sort_values("missing_%", ascending=False).head(15))
    print(df[target].value_counts(normalize=True).round(3))
    num = df.select_dtypes("number").drop(columns=[target], errors="ignore")
    print(num.describe().T[["mean", "std", "min", "max"]].round(2).head(10))
    print("skew (top):", num.skew().abs().sort_values(ascending=False).head(5).round(2).to_dict())
    corr = num.corrwith(df[target]).abs().sort_values(ascending=False)
    print("most target-correlated:", corr.head(5).round(3).to_dict())

Descriptive statistics in pandas

scores = pd.Series([22, 18, 14, 10, 15, 20, 25, 12])
scores.mean(), scores.median(), scores.mode().tolist()   # 17.0 16.5 [10, 12, 14, ...] (all tie)
scores.var(), scores.std()          # 26.571..., 5.154...  (sample, ddof=1 by default)
scores.var(ddof=0)                  # 23.25  population variance, matches NumPy's default
scores.quantile([0.25, 0.5, 0.75])  # quartiles
iqr = scores.quantile(0.75) - scores.quantile(0.25)
(scores - scores.mean()).abs().mean()   # mean absolute deviation: 4.25
scores.max() - scores.min()         # range: 15

The mean is pulled by outliers while the median is not, so a mean far above the median signals right skew (incomes, prices, counts). The IQR rule flags outliers below Q1 − 1.5 × IQR or above Q3 + 1.5 × IQR, which is exactly what box-plot whiskers show.

Outliers: detect, then decide

q1, q3 = df.amount.quantile([0.25, 0.75])
iqr = q3 - q1
outliers = df[(df.amount < q1 - 1.5 * iqr) | (df.amount > q3 + 1.5 * iqr)]
z = (df.amount - df.amount.mean()) / df.amount.std()     # |z| > 3 rule for roughly normal data
df["amount_log"] = np.log1p(df.amount)                    # compress a heavy right tail
df["amount_clip"] = df.amount.clip(upper=df.amount.quantile(0.99))   # winsorize

Do not delete outliers automatically. A 10-million transaction may be an error or exactly the fraud you want to catch. Investigate, then choose to fix, cap, transform, or keep and use a robust model or loss.

Common pitfall Doing EDA-driven feature selection on the full dataset, including the test rows, and then reporting test performance. Choosing features by their correlation with the target on all data leaks test information. Explore freely, but make target-dependent decisions (feature selection, target encoding, threshold choice) on training data or inside cross-validation.
Interview angle "You get a new dataset. Walk me through your first hour." Structure beats detail: clarify the target and prediction time, check structure and quality, look at the target distribution, univariate and bivariate views, hunt for leakage, then state concrete preprocessing decisions. Mention a quick baseline (majority class or linear model) to anchor later results.

scikit-learn: the estimator API, pipelines and evaluation

scikit-learn succeeds because every object follows the same small contract. An estimator learns from data with fit(X, y). A transformer also has transform(X) (and the shortcut fit_transform). A predictor has predict(X), and classifiers usually add predict_proba(X) and score(X, y). Hyperparameters are set in the constructor; learned values get a trailing underscore (coef_, mean_, classes_). Once you know this contract, every model in the library is familiar.

Analogy

A pipeline is a factory assembly line. At commissioning time (fit) each station calibrates its tools on sample parts: the imputer measures typical values, the scaler measures means and spreads, the model learns its weights. In production (predict) new parts travel through the same calibrated stations in the same order, and nobody recalibrates using the customer's parts. Mapping back: fit learns parameters only from training data, transform/predict apply them unchanged to new data, and wrapping all stations in one Pipeline guarantees that cross-validation recalibrates them on each training fold only.

The contract in five lines

from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

scaler = StandardScaler()                 # hyperparameters go in the constructor
X_train_s = scaler.fit_transform(X_train) # learns mean_ and scale_ from TRAIN only
X_test_s = scaler.transform(X_test)       # re-uses train statistics: no leakage
clf = LogisticRegression(max_iter=1000).fit(X_train_s, y_train)
clf.predict(X_test_s); clf.predict_proba(X_test_s)[:, 1]; clf.score(X_test_s, y_test)
MethodWho has itWhat it does
fit(X, y)Every estimatorLearn parameters; returns self so calls can chain
transform(X)Transformers (scalers, encoders, PCA, imputers)Apply learned parameters to data
fit_transform(X)TransformersBoth at once; use only on training data
predict(X)ModelsClass labels or regression values
predict_proba(X) / decision_function(X)Most classifiersProbabilities per class / raw scores (margins)
score(X, y)ModelsAccuracy for classifiers, R² for regressors
get_params() / set_params()Every estimatorHow grid search reads and changes hyperparameters

Preprocessing toolbox

TransformerDoesUse when
StandardScalerSubtract mean, divide by standard deviationLinear models, SVM, kNN, PCA, neural nets
MinMaxScalerRescale to [0, 1]Bounded inputs, images, some neural nets
RobustScalerSubtract median, divide by IQRFeatures with strong outliers
PowerTransformer, np.log1pReduce skewHeavy-tailed amounts, counts
SimpleImputer, KNNImputerFill missing values (mean, median, most frequent, constant)Any column with gaps; add add_indicator=True to keep a "was missing" flag
OneHotEncoderOne binary column per categoryNominal categories (city, colour); use handle_unknown="ignore"
OrdinalEncoderMap categories to integersOrdered categories, or tree models
TargetEncoderReplace a category by a cross-fitted mean targetHigh-cardinality categories (zip codes)
PolynomialFeaturesAdd squares and interactionsLet linear models fit curves
TfidfVectorizerText to sparse weighted word countsClassic text classification and retrieval baselines

Distance- and gradient-based models care about feature scale; tree-based models (decision trees, random forests, gradient boosting) do not. On the breast-cancer data, scaling moves kNN from 0.912 to 0.956 test accuracy and an RBF SVM from 0.904 to 0.974, while an unscaled random forest already reaches 0.965.

Pipelines and ColumnTransformer

import numpy as np, pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.linear_model import LogisticRegression

df = pd.DataFrame({
    "age":    [25, np.nan, 47, 35, 52, 23],
    "income": [40, 60, np.nan, 80, 120, 30],
    "city":   ["Pune", "Delhi", "Pune", np.nan, "Mumbai", "Delhi"],
    "bought": [0, 1, 1, 0, 1, 0],
})
X, y = df.drop(columns="bought"), df.bought
numeric, categorical = ["age", "income"], ["city"]

preprocess = ColumnTransformer([
    ("num", Pipeline([("impute", SimpleImputer(strategy="median")),
                      ("scale", StandardScaler())]), numeric),
    ("cat", Pipeline([("impute", SimpleImputer(strategy="most_frequent")),
                      ("onehot", OneHotEncoder(handle_unknown="ignore"))]), categorical),
])
model = Pipeline([("pre", preprocess), ("clf", LogisticRegression())])
model.fit(X, y)
print(model.named_steps["pre"].get_feature_names_out())
# ['num__age' 'num__income' 'cat__city_Delhi' 'cat__city_Mumbai' 'cat__city_Pune']
print(model.predict(pd.DataFrame({"age": [30], "income": [70], "city": ["Chennai"]})))
# [0]   unseen city "Chennai" is encoded as all zeros instead of crashing

Parameters of steps inside a pipeline are addressed as step__param, with nesting as deep as needed: clf__C, pre__num__impute__strategy. make_pipeline and make_column_transformer build the same objects with auto-generated step names. Calling set_output(transform="pandas") makes transformers return DataFrames with readable column names.

Splitting data correctly

from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, stratify=y, random_state=42)   # stratify keeps class ratios
SplitterUse when
KFold / StratifiedKFoldIndependent rows; stratified for classification
GroupKFold / StratifiedGroupKFoldSeveral rows per patient, user or device; keeps each group on one side
TimeSeriesSplitTemporal data; always trains on the past and validates on the future
RepeatedStratifiedKFoldSmall datasets where one CV run is noisy
5-fold cross-validation (training data only; test set stays locked away)
fold 1:  [VAL ][train][train][train][train]
fold 2:  [train][VAL ][train][train][train]
fold 3:  [train][train][VAL ][train][train]
fold 4:  [train][train][train][VAL ][train]
fold 5:  [train][train][train][train][VAL ]
score = mean of 5 validation scores (report the std too)

Cross-validation

from sklearn.model_selection import StratifiedKFold, cross_val_score, cross_validate
from sklearn.pipeline import make_pipeline
from sklearn.datasets import load_breast_cancer
X, y = load_breast_cancer(return_X_y=True)
y = 1 - y                                     # make malignant the positive class (1)
pipe = make_pipeline(StandardScaler(), LogisticRegression(max_iter=5000))
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)
res = cross_validate(pipe, X, y, cv=cv, scoring=["accuracy", "recall", "roc_auc"])
print({k: round(v.mean(), 3) for k, v in res.items() if k.startswith("test")})
# {'test_accuracy': 0.979, 'test_recall': 0.958, 'test_roc_auc': 0.995}

Why the pipeline matters: a leakage demonstration

from sklearn.feature_selection import SelectKBest, f_classif
rng = np.random.default_rng(0)
Xr = rng.normal(size=(100, 10_000))          # pure noise features
yr = rng.integers(0, 2, 100)                 # random labels: true accuracy is 50%

# WRONG: select the 20 "best" features using ALL rows, then cross-validate
leaky = SelectKBest(f_classif, k=20).fit_transform(Xr, yr)
print(cross_val_score(LogisticRegression(max_iter=1000), leaky, yr, cv=5).mean())   # 0.87 (!)

# RIGHT: selection happens inside each training fold
honest = make_pipeline(SelectKBest(f_classif, k=20), LogisticRegression(max_iter=1000))
print(cross_val_score(honest, Xr, yr, cv=5).mean())                                 # 0.48

The leaky version reports 87% accuracy on random noise, because the feature selector already "saw" the validation labels. Any step that learns from data, including scaling, imputation, feature selection, target encoding, oversampling (SMOTE) and PCA, must live inside the pipeline. Cross-validation estimates how well that whole procedure generalises; a bootstrap of a single test-set metric instead puts an error bar on a frozen model's score. Use CV to choose models, and a paired bootstrap (or McNemar) on the sealed test set to compare two finished models.

Hyperparameter search

from sklearn.model_selection import GridSearchCV, RandomizedSearchCV
from scipy.stats import loguniform

param_grid = {"logisticregression__C": [0.01, 0.1, 1, 10, 100],
              "logisticregression__class_weight": [None, "balanced"]}
grid = GridSearchCV(pipe, param_grid, cv=cv, scoring="recall", n_jobs=-1)
grid.fit(X_train, y_train)
grid.best_params_, grid.best_score_       # best combination and its mean CV score
best = grid.best_estimator_               # already refit on all training data
pd.DataFrame(grid.cv_results_)[["params", "mean_test_score", "std_test_score"]]

rand = RandomizedSearchCV(pipe, {"logisticregression__C": loguniform(1e-3, 1e3)},
                          n_iter=30, cv=cv, scoring="roc_auc", random_state=0, n_jobs=-1)
  • Grid search tries every combination: cost multiplies with each parameter added. Randomized search samples a fixed budget and usually finds good regions faster when only a few parameters matter.
  • HalvingGridSearchCV and libraries such as Optuna (Bayesian optimisation) spend more budget on promising candidates.
  • best_score_ is optimistically biased because it was used for selection. Report final performance on the untouched test set, or use nested cross-validation (a search inside each outer fold) when the dataset is small.
  • With multiple metrics, pass scoring={"rec": "recall", "prec": "precision"} and choose which one to refit.

Metrics: pick the one that matches the cost of errors

Precision = TP / (TP + FP)    Recall = TP / (TP + FN)    F1 = 2 · P · R / (P + R) TP true positives, FP false positives, FN false negatives. In scikit-learn's confusion matrix rows are true labels and columns are predictions: [[TN, FP], [FN, TP]] for labels 0 and 1.
TaskMetricChoose when
ClassificationAccuracyBalanced classes, equal error costs
Recall (sensitivity)Missing a positive is costly: cancer, fraud, safety
PrecisionFalse alarms are costly: spam filtering, alerts that page a human
F1, macro-F1Need a balance; macro-F1 treats every class equally in multiclass problems
ROC-AUCThreshold-free ranking quality; can look optimistic with rare positives
PR-AUC (average precision)Rare positive class
Log loss, Brier scoreYou need well-calibrated probabilities
RegressionMAERobust to outliers, same units as the target
RMSELarge errors are especially bad
R²Fraction of variance explained, relative to predicting the mean
MAPERelative error; breaks when true values are near zero
from sklearn.metrics import (accuracy_score, precision_score, recall_score, f1_score,
    confusion_matrix, classification_report, roc_auc_score, average_precision_score,
    mean_absolute_error, mean_squared_error, r2_score)
y_pred = best.predict(X_test)
y_proba = best.predict_proba(X_test)[:, 1]          # probability of class 1
print(confusion_matrix(y_test, y_pred))
print(classification_report(y_test, y_pred, digits=3))
roc_auc_score(y_test, y_proba)                      # needs scores, not hard labels
f1_score(y_true_multi, y_pred_multi, average="macro")

Imbalance and decision thresholds

  • class_weight="balanced" re-weights the loss so the rare class counts more; often the simplest effective fix.
  • The default 0.5 threshold is arbitrary. Choose a threshold on validation data (or out-of-fold predictions) with precision_recall_curve, according to the business constraint such as "recall at least 0.95". Newer versions also provide TunedThresholdClassifierCV for this.
  • Resampling (under-sampling, SMOTE from the imbalanced-learn library) must be applied only to training folds, inside a pipeline.
  • Always compare against a trivial baseline: DummyClassifier() scores 0.627 on the breast-cancer data just by predicting the majority class.

A model zoo with one API

from sklearn.neighbors import KNeighborsClassifier
from sklearn.svm import SVC
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier, HistGradientBoostingClassifier
models = {
    "logreg": make_pipeline(StandardScaler(), LogisticRegression(max_iter=5000)),
    "knn":    make_pipeline(StandardScaler(), KNeighborsClassifier(n_neighbors=5)),
    "svm":    make_pipeline(StandardScaler(), SVC(probability=True)),
    "tree":   DecisionTreeClassifier(max_depth=4, random_state=0),
    "forest": RandomForestClassifier(n_estimators=300, random_state=0, n_jobs=-1),
    "gbm":    HistGradientBoostingClassifier(random_state=0),   # handles NaN natively
}
for name, m in models.items():
    s = cross_val_score(m, X_train, y_train, cv=cv, scoring="recall")
    print(f"{name:7s} recall {s.mean():.3f} +/- {s.std():.3f}")

Algorithm details (how trees split, what the SVM margin is, why kNN suffers in high dimensions) are covered in Machine Learning. Unsupervised tools follow the same API: PCA(n_components=2).fit_transform(X), KMeans(n_clusters=3, n_init=10).fit_predict(X).

Inspecting and saving models

coefs = pd.Series(best[-1].coef_[0], index=feature_names).sort_values(key=abs)   # linear models
importances = forest.feature_importances_                                         # trees
from sklearn.inspection import permutation_importance
pi = permutation_importance(best, X_test, y_test, scoring="recall", n_repeats=10, random_state=0)

import joblib
joblib.dump(best, "model.joblib")       # save the WHOLE pipeline, preprocessing included
model = joblib.load("model.joblib")     # same library versions required; never load untrusted files
model.predict(new_rows)
Common pitfall The top three scikit-learn mistakes: (1) calling fit_transform on the test set, or scaling before splitting; (2) tuning hyperparameters or the threshold on the test set, so the "test" score is no longer unbiased; (3) saving only the model and re-implementing preprocessing by hand at serving time, which drifts. Also note that recall_score uses pos_label=1; in the raw breast-cancer data 0 means malignant, so remap with y = 1 - y or pass pos_label=0.
Interview angle Expect: fit versus transform versus fit_transform; what data leakage is and how pipelines prevent it; why stratify; k-fold versus group versus time-series splits; grid versus random search; precision versus recall trade-offs with a concrete business example; ROC-AUC versus PR-AUC under imbalance; how you would pick a threshold; which models need scaling. Strong candidates always mention a baseline and a held-out test set that is touched once.

PyTorch fundamentals: tensors, autograd and the training loop

PyTorch rests on three ideas. Tensors are NumPy-like arrays that can live on a GPU. Autograd records the operations you perform on tensors that require gradients and, when you call backward(), applies the chain rule backwards through that record to compute every gradient. Modules (nn.Module) package parameters and a forward computation into reusable, nestable building blocks. Around these sit Dataset/DataLoader for feeding data and optimizers for updating weights.

Analogy

Autograd is a hiker dropping breadcrumbs. On the way forward, every operation leaves a crumb recording where it came from. When you call backward(), the hiker walks the trail in reverse, and at each crumb works out how much a small step there would have changed the final loss. The optimizer is the guide who then nudges every parameter a little downhill. Mapping back: the crumbs are the dynamic computation graph (grad_fn), the reverse walk is backpropagation, gradients land in each parameter's .grad, and optimizer.step() updates the weights.

Tensors

import torch, numpy as np
x = torch.tensor([[1.0, 2.0], [3.0, 4.0]])
print(x.shape, x.dtype, x.device)     # torch.Size([2, 2]) torch.float32 cpu
print(torch.tensor([1, 2]).dtype)     # torch.int64 -- the default for class labels
print(x @ x, x * x)                   # matrix product / element-wise, same as NumPy
print(x.sum(dim=0), x.mean(dim=1))    # tensor([4., 6.]) tensor([1.5000, 3.5000])  dim = axis
torch.zeros(2, 3); torch.randn(4, 3); torch.arange(6); torch.linspace(0, 1, 5)

# reshaping
x.unsqueeze(0).shape                  # torch.Size([1, 2, 2])  add a batch dimension
x.permute(1, 0); x.T                  # reorder dimensions
x.T.reshape(4)                        # tensor([1., 3., 2., 4.])  copies if needed
# x.T.view(4) -> RuntimeError: view needs contiguous memory; use .contiguous() or reshape

# NumPy bridge: shares memory on CPU
a = np.ones(2)
t = torch.from_numpy(a)               # float64 tensor sharing a's buffer
a[0] = 5
print(t)                              # tensor([5., 1.], dtype=torch.float64)
back = t.numpy()                      # GPU tensors need .cpu() first; grad tensors need .detach()

Devices

device = ("cuda" if torch.cuda.is_available()
          else "mps" if torch.backends.mps.is_available()   # Apple silicon
          else "cpu")
model = model.to(device)                  # moves parameters in place
xb, yb = xb.to(device), yb.to(device)     # tensors return a NEW tensor; reassign!
# every tensor in one operation must be on the same device, or you get
# "Expected all tensors to be on the same device"

Autograd

x = torch.tensor(2.0, requires_grad=True)
y = x ** 3 + 2 * x            # dy/dx = 3x^2 + 2
y.backward()
print(x.grad)                 # tensor(14.)

w = torch.tensor(1.0, requires_grad=True)
for _ in range(2):
    loss = (w * 3) ** 2       # d/dw = 18w = 18
    loss.backward()
    print(w.grad)             # tensor(18.) then tensor(36.)  -- gradients ACCUMULATE
w.grad.zero_()                # this is why training loops call optimizer.zero_grad()

with torch.no_grad():         # no graph recorded: inference and manual updates
    w -= 0.1 * 18
frozen = y.detach()           # same data, cut from the graph (requires_grad=False)

Gradients accumulate by design, so that you can sum gradients over several small batches before one update (gradient accumulation). The price is that you must clear them each step.

Building models with nn.Module

import torch.nn as nn, torch.nn.functional as F

class MLP(nn.Module):
    def __init__(self, d_in, d_hidden=32, n_classes=2, p_drop=0.2):
        super().__init__()                      # must come first
        self.net = nn.Sequential(
            nn.Linear(d_in, d_hidden),          # weight (32, d_in), bias (32,)
            nn.ReLU(),
            nn.Dropout(p_drop),
            nn.Linear(d_hidden, n_classes),
        )
    def forward(self, x):
        return self.net(x)                      # raw logits: no softmax here

model = MLP(30)
print(sum(p.numel() for p in model.parameters()))   # 1058 = 30*32+32 + 32*2+2
print(model(torch.randn(5, 30)).shape)               # torch.Size([5, 2])
for name, p in model.named_parameters():
    print(name, tuple(p.shape))                      # net.0.weight (32, 30) ...

Assigning an nn.Module or nn.Parameter as an attribute registers it automatically, so model.parameters(), .to(device) and state_dict() all find it. A plain Python list of layers is not registered; use nn.ModuleList or nn.ModuleDict.

Losses and output layers

TaskModel outputLossTarget dtype and shape
MulticlassLogits, shape (N, C)nn.CrossEntropyLoss (applies log-softmax internally)long, shape (N,), values 0..C-1
Binary / multi-labelOne logit per label, shape (N,) or (N, L)nn.BCEWithLogitsLoss (applies sigmoid internally, numerically stable)float, same shape as output
RegressionValues, shape (N,) or (N, 1)nn.MSELoss, nn.L1Loss, nn.HuberLossfloat, same shape as output
logits = torch.tensor([[2.0, 0.5, -1.0]])
print(nn.CrossEntropyLoss()(logits, torch.tensor([0])).item())   # 0.2413...  = -log softmax(logits)[0]
probs = F.softmax(logits, dim=1)       # only for reporting or inference, never before CrossEntropyLoss

Datasets and DataLoaders

from torch.utils.data import Dataset, DataLoader, TensorDataset

class TabularDataset(Dataset):
    def __init__(self, X, y):
        self.X = torch.as_tensor(X, dtype=torch.float32)
        self.y = torch.as_tensor(y, dtype=torch.long)
    def __len__(self):
        return len(self.y)
    def __getitem__(self, i):
        return self.X[i], self.y[i]           # one sample; DataLoader stacks them into batches

train_dl = DataLoader(TabularDataset(X_tr, y_tr), batch_size=32, shuffle=True,
                      num_workers=0, pin_memory=torch.cuda.is_available())
val_dl = DataLoader(TabularDataset(X_va, y_va), batch_size=256, shuffle=False)
xb, yb = next(iter(train_dl))
print(xb.shape, yb.shape)                     # torch.Size([32, 30]) torch.Size([32])

num_workers > 0 starts worker processes that prepare batches in parallel, which matters when each item needs disk reads or image decoding. For tensors already in memory, 0 is often fastest. A custom collate_fn handles variable-length samples such as padded text.

The training loop

for each epoch:
    model.train()
    for each batch:
        optimizer.zero_grad()      1. clear old gradients
        logits = model(xb)         2. forward pass (builds graph)
        loss = loss_fn(logits, yb) 3. scalar loss
        loss.backward()            4. backprop: fill p.grad
        optimizer.step()           5. update weights
    model.eval(); with torch.no_grad(): validate, log, checkpoint best
import torch, torch.nn as nn
from torch.utils.data import TensorDataset, DataLoader
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler

torch.manual_seed(0)
X, y = load_breast_cancer(return_X_y=True); y = 1 - y           # 1 = malignant
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)
sc = StandardScaler().fit(X_tr); X_tr, X_te = sc.transform(X_tr), sc.transform(X_te)

train_ds = TensorDataset(torch.tensor(X_tr, dtype=torch.float32), torch.tensor(y_tr, dtype=torch.long))
test_ds = TensorDataset(torch.tensor(X_te, dtype=torch.float32), torch.tensor(y_te, dtype=torch.long))
train_dl = DataLoader(train_ds, batch_size=32, shuffle=True)
test_dl = DataLoader(test_ds, batch_size=256)

device = "cuda" if torch.cuda.is_available() else "cpu"
model = MLP(30).to(device)
opt = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=1e-4)
loss_fn = nn.CrossEntropyLoss()

for epoch in range(1, 31):
    model.train(); total = 0.0
    for xb, yb in train_dl:
        xb, yb = xb.to(device), yb.to(device)
        opt.zero_grad()
        loss = loss_fn(model(xb), yb)
        loss.backward()
        opt.step()
        total += loss.item() * xb.size(0)      # .item() -> Python float, frees the graph
    if epoch % 10 == 0:
        model.eval(); correct = 0
        with torch.no_grad():
            for xb, yb in test_dl:
                correct += (model(xb.to(device)).argmax(1) == yb.to(device)).sum().item()
        print(f"epoch {epoch:2d}  train_loss {total/len(train_ds):.4f}  test_acc {correct/len(test_ds):.3f}")
# epoch 10  train_loss 0.0947  test_acc 0.974
# epoch 20  train_loss 0.0661  test_acc 0.974
# epoch 30  train_loss 0.0536  test_acc 0.974   (exact numbers vary by version and hardware)

In real projects, evaluate on a separate validation split for early stopping and model selection, and keep the test set for one final measurement, exactly as with scikit-learn.

train() versus eval(), no_grad versus inference_mode

model.train() / model.eval()

  • Switch layer behaviour
  • Dropout: active in train, identity in eval
  • BatchNorm: batch statistics in train, stored running statistics in eval
  • Does not turn off gradient tracking

torch.no_grad() / torch.inference_mode()

  • Switch gradient tracking
  • No graph built: less memory, faster
  • inference_mode is stricter and slightly faster
  • Does not change dropout or BatchNorm

For inference you need both: model.eval() and with torch.no_grad():.

Saving and loading

torch.save({"model": model.state_dict(), "optimizer": opt.state_dict(),
            "epoch": epoch, "config": {"d_in": 30}}, "checkpoint.pt")

ckpt = torch.load("checkpoint.pt", map_location="cpu", weights_only=False)  # False: file also has epoch/config; only load trusted files
model = MLP(**ckpt["config"])
model.load_state_dict(ckpt["model"])
model.eval()

Save the state_dict (a dict of tensors) rather than pickling the whole model object, which ties the file to your exact class definitions and file paths. Recent PyTorch versions default torch.load(..., weights_only=True), which refuses non-tensor objects; a training checkpoint that also stores epoch or a config dict needs weights_only=False and must come from a trusted source. The safetensors format stores tensors without pickle and avoids that risk.

Useful extras

torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)   # tame exploding gradients
sched = torch.optim.lr_scheduler.CosineAnnealingLR(opt, T_max=30)   # call sched.step() each epoch
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):       # mixed precision on GPU
    loss = loss_fn(model(xb), yb)
model = torch.compile(model)                                         # graph compilation (PyTorch 2.x)
torch.manual_seed(0); np.random.seed(0); random.seed(0)              # reproducibility
Common pitfall The PyTorch bug parade: forgetting zero_grad() (gradients accumulate), forgetting model.eval() (dropout stays on, results fluctuate), applying softmax before CrossEntropyLoss (double softmax, poor training), float labels for CrossEntropyLoss or long labels for BCEWithLogitsLoss, int64 inputs to a Linear layer ("mat1 and mat2 must have the same dtype, but got Long and Float"), x.to(device) without reassigning, and accumulating total += loss instead of loss.item(), which keeps every graph alive and leaks GPU memory.
Interview angle Expect to write the training loop from memory and explain each line; the difference between eval() and no_grad(); why gradients accumulate; view versus reshape; detach(); what requires_grad does; how to freeze layers for fine-tuning (p.requires_grad = False and pass only trainable parameters to the optimizer); state_dict saving; and how you would debug a loss that does not decrease (overfit one small batch first). Deeper theory is in Deep Learning.

Hugging Face basics: pipelines, tokenizers and Auto classes

The Hugging Face transformers library gives you thousands of pretrained models behind a uniform interface. Three layers of abstraction cover almost every use: pipeline() for a task in one line, AutoTokenizer to turn text into token IDs, and AutoModel... classes to load the right architecture for a checkpoint by name. The theory of what these models do (attention, tokenization, decoding) lives in Transformers & LLMs; this section is the Python mechanics.

Analogy

A pipeline is a coffee machine with a "cappuccino" button: beans in, drink out, no need to know the grind size. The tokenizer is the grinder that turns beans (raw text) into a consistent powder (token IDs) that only fits this machine's filter. The Auto classes are a universal adapter that looks at the model's label and snaps in the right internal parts. Mapping back: pipeline("sentiment-analysis") hides tokenization, model forward pass and post-processing; using the tokenizer and model directly gives you control over batching, truncation and devices, and the tokenizer must always match its model.

from transformers import pipeline
clf = pipeline("sentiment-analysis")                  # downloads a default model on first use
print(clf(["Great battery life!", "Screen cracked in a week."]))
# [{'label': 'POSITIVE', 'score': 0.99...}, {'label': 'NEGATIVE', 'score': 0.99...}]
# other tasks: "zero-shot-classification", "ner", "question-answering",
#              "summarization", "text-generation", "feature-extraction"

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
name = "distilbert-base-uncased-finetuned-sst-2-english"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForSequenceClassification.from_pretrained(name).eval()

batch = tok(["Great battery life!", "Screen cracked in a week."],
            padding=True, truncation=True, max_length=128, return_tensors="pt")
print(batch.keys())            # input_ids, attention_mask
with torch.no_grad():
    logits = model(**batch).logits          # shape (2, 2)
print(logits.softmax(-1).argmax(-1).tolist(), model.config.id2label)
ClassReturnsTypical use
AutoModelHidden states (embeddings)Feature extraction, custom heads
AutoModelForSequenceClassificationClass logits per textSentiment, intent, topic
AutoModelForTokenClassificationLogits per tokenNamed-entity recognition
AutoModelForCausalLMNext-token logits, .generate()Text generation (GPT-style)
AutoModelForSeq2SeqLMEncoder-decoder generationTranslation, summarization
  • The attention_mask marks real tokens (1) versus padding (0) so padded positions are ignored.
  • Models are ordinary PyTorch nn.Modules: .to(device), .eval(), no_grad and custom training loops all work. The Trainer class and the datasets library wrap the loop and data handling for fine-tuning.
  • Load large models with torch_dtype=torch.bfloat16 and device_map="auto" to save memory; parameter-efficient fine-tuning (LoRA via the peft library) trains small adapter matrices instead of all weights.
  • In restricted networks, download models through your organisation's approved mirror and load them from a local folder path with from_pretrained("path/to/model").
Common pitfall Mismatched tokenizer and model (for example, tokenizing with one checkpoint and running another) produces garbage without any error. Also: forgetting truncation=True and crashing on inputs longer than the model's maximum length, and running a large model on CPU in float32 when a smaller distilled model would do.
Interview angle Questions here are practical: what does a tokenizer return, why padding and attention masks, the difference between AutoModel and AutoModelForCausalLM, how to batch inference efficiently, and how you would fine-tune a classifier on a small labelled set (pretrained encoder plus classification head, low learning rate, few epochs, early stopping, or LoRA for large models).

Calling LLM APIs from Python

Most LLM application code is ordinary Python: build a list of messages, send an HTTP request through an SDK, parse the response, handle errors and rate limits, and validate structured output. Many providers and self-hosted servers expose an OpenAI-compatible interface, so one client pattern works across them. Prompting, structured outputs, tool calling and evaluation are covered in depth in LLM APIs.

Analogy

Calling an LLM API is like ordering at a busy restaurant by phone. You give your standing instructions (system message) and the order (user message), you are charged by the size of the order (tokens), the line is sometimes busy (rate limits), so you wait and redial politely (retry with backoff), and you check the delivery against the receipt before paying (validate the structured output). Mapping back: messages carry role-tagged text, usage fields report tokens, HTTP 429 errors signal rate limits, and schema validation catches malformed answers.

import os
from openai import OpenAI            # the same client works with OpenAI-compatible servers
from pydantic import BaseModel, ValidationError

client = OpenAI(api_key=os.environ["LLM_API_KEY"],          # never hard-code keys
                base_url=os.environ.get("LLM_BASE_URL"))    # your approved endpoint

resp = client.chat.completions.create(
    model=os.environ.get("LLM_MODEL", "gpt-4o-mini"),
    messages=[{"role": "system", "content": "You are a concise tutor."},
              {"role": "user", "content": "Explain overfitting in one sentence."}],
    temperature=0, max_tokens=80)
print(resp.choices[0].message.content, resp.usage.total_tokens)

class Review(BaseModel):
    sentiment: str
    rating: int

raw = '{"sentiment": "positive", "rating": 5}'     # e.g. a JSON-mode response
try:
    review = Review.model_validate_json(raw)
except ValidationError as e:
    print("retry or fall back:", e)
  • Keep keys in environment variables or a secrets manager; never commit them or print them in notebooks.
  • Wrap calls with timeouts and exponential-backoff retries on rate-limit and transient server errors.
  • Use asyncio with a semaphore (see the concurrency section) to evaluate hundreds of prompts in parallel within rate limits.
  • Cache responses keyed by (model, prompt, parameters) during development to save cost and make runs repeatable.
  • Log token usage and latency per call; cost scales with input plus output tokens.
Common pitfall Trusting model output as if it were typed data. Always parse and validate (Pydantic), handle the failure path, and never pass model output directly to eval, a shell or a SQL query. Also avoid sending confidential data to endpoints your organisation has not approved.
Interview angle Python-side questions: how you would run 10,000 prompts efficiently (async plus semaphore, batching, retries, caching, checkpointing progress), how you handle rate limits, how you guarantee valid JSON, and how you store secrets.

Performance and best practices

Speed work follows one rule: measure first, then fix the biggest bottleneck. In data code the usual culprits are Python-level loops over rows, wrong dtypes that bloat memory, re-reading or re-computing the same data, and moving data between CPU and GPU or between processes more than necessary.

Analogy

Optimising without profiling is like trying to shorten a commute by buying faster shoes when the real delay is a 40-minute wait at one traffic light. A profiler is the GPS log that shows exactly where the minutes go. Mapping back: cProfile, line_profiler and %timeit reveal the hot function or line, and fixing that one spot (usually by vectorizing it) beats micro-tuning everything else.

Measure

# notebook
%timeit df.a * df.b                    # repeated timing of one line
%prun train_features(df)               # function-level profile
%load_ext line_profiler
%lprun -f build_features build_features(df)   # line-by-line (needs line_profiler)
%memit df = pd.read_csv("big.csv")     # peak memory (needs memory_profiler)

# script
import cProfile, pstats, tracemalloc
cProfile.run("main()", "prof.out")
pstats.Stats("prof.out").sort_stats("cumulative").print_stats(10)
tracemalloc.start(); run(); print(tracemalloc.get_traced_memory())   # (current, peak) bytes
# command line: python -m cProfile -s cumtime train.py ; py-spy for sampling live processes

The speed ladder (fastest first)

  1. Avoid the work Filter early, read fewer columns, cache intermediate results (Parquet, joblib.Memory, lru_cache).
  2. Vectorize NumPy ufuncs, pandas column operations, np.where/np.select, groupby built-ins such as "sum" rather than a Python lambda.
  3. Use the right structure Set or dict for membership tests instead of a list; collections.deque for queues; sorted arrays with np.searchsorted.
  4. Compile hot loops Numba @njit or Cython for loops that genuinely cannot be vectorized.
  5. Parallelize n_jobs=-1 in scikit-learn, joblib.Parallel, process pools, DataLoader workers.
  6. Change engine or hardware Polars or DuckDB for large tables, Dask or Spark for multi-machine data, GPU for large matrix work.
import numpy as np
from numba import njit

@njit                                    # compiled to machine code on first call
def count_runs(x):
    runs = 0
    for i in range(1, x.size):          # sequential dependency: hard to vectorize
        if x[i] != x[i - 1]:
            runs += 1
    return runs

from joblib import Parallel, delayed
results = Parallel(n_jobs=-1)(delayed(score_file)(p) for p in paths)

Memory habits

  • Use float32 for model inputs and category for repetitive strings; check with memory_usage(deep=True).
  • Process in chunks or with generators when data exceeds RAM; write intermediate results to Parquet.
  • Delete large temporaries (del df_raw) and call gc.collect() in long notebooks; restart the kernel when in doubt.
  • Prefer in-place NumPy operations for huge arrays (np.multiply(a, 2, out=a), a *= 2) to avoid temporary copies.
  • On GPU: smaller batches, mixed precision, gradient accumulation, torch.no_grad() for evaluation, and .item() when logging scalars.

Code-quality habits that pay off in ML

Configuration

Keep hyperparameters in a dataclass or YAML file, not scattered literals. Log the config with every run.

Seeds and versions

Seed Python, NumPy and PyTorch; pin library versions; record the data snapshot.

Tests

Unit-test feature functions with tiny DataFrames (pytest); assert shapes, dtypes and no NaN after preprocessing.

Logging

Use logging instead of print in pipelines; track experiments with a tracking tool.

Style

PEP 8 naming, type hints, docstrings, a formatter and linter (for example black and ruff).

Notebooks to modules

Explore in notebooks, then move reusable code into a package with a command-line entry point.

Common pitfall Premature parallelism. Spinning up a process pool for a job dominated by one slow pandas apply adds pickling overhead and memory copies; vectorizing that one line usually gives a far larger speed-up on a single core. Also beware nested parallelism: n_jobs=-1 in both a grid search and the model inside it oversubscribes the CPU.
Interview angle "This pandas script takes two hours. What do you do?" Profile to find the hot spot; replace apply/iterrows with vectorized operations or merges; fix dtypes; read only needed columns from Parquet; cache intermediate steps; then parallelize or switch engines if still needed. Quantify each step with before-and-after timings.

Worked example: breast-cancer diagnosis end to end

This project ties every section together on the classic breast-cancer dataset: 569 tumours, 30 numeric features computed from cell-nucleus images, and a binary label. Missing a malignant tumour (a false negative) is far worse than an unnecessary follow-up (a false positive), so the primary metric is recall on the malignant class, with a sanity floor on precision. Every output below comes from running the code with scikit-learn 1.9 and pandas 3.0; small numeric differences across versions are normal.

Analogy

Think of the project as a clinical trial. You recruit patients (load data), run baseline health checks (EDA), seal an envelope of patients nobody may look at until the end (the test set), try treatments on the rest with careful rotation so each is judged fairly (cross-validation), pick the best and set its dose (tuning and threshold), and only then open the sealed envelope once to report results. Mapping back: the sealed envelope is X_test, rotation is StratifiedKFold, the dose is the decision threshold, and opening the envelope twice would make the reported result untrustworthy.

Step 1: load and relabel

import numpy as np, pandas as pd
from sklearn.datasets import load_breast_cancer

data = load_breast_cancer(as_frame=True)
df = data.frame.copy()
df["target"] = 1 - df["target"]        # raw data: 0 = malignant; now 1 = malignant (positive class)
print(df.shape)                        # (569, 31)
print(df["target"].value_counts())
# target
# 0    357      benign
# 1    212      malignant  (about 37%: mildly imbalanced)
print(df.isna().sum().sum())           # 0  no missing values

Step 2: exploratory analysis

print(df[["mean radius", "mean area", "mean smoothness"]].describe().round(2))
#        mean radius  mean area  mean smoothness
# count       569.00     569.00           569.00
# mean         14.13     654.89             0.10
# std           3.52     351.91             0.01
# min           6.98     143.50             0.05
# ...          (quartile rows omitted)
# max          28.11    2501.00             0.16
# scales differ by 4 orders of magnitude -> scale for linear models, SVM, kNN

print(df.groupby("target")[["mean radius", "worst concave points"]].mean().round(3))
#         mean radius  worst concave points
# target
# 0            12.147                 0.074
# 1            17.463                 0.182       malignant tumours are larger and more irregular

import matplotlib.pyplot as plt, seaborn as sns
fig, axes = plt.subplots(1, 3, figsize=(13, 3.5))
sns.countplot(df, x="target", ax=axes[0])
sns.histplot(df, x="mean radius", hue="target", kde=True, ax=axes[1])
sns.heatmap(df.iloc[:, :10].corr(), cmap="coolwarm", center=0, ax=axes[2])
fig.tight_layout()
# the heatmap shows radius, perimeter and area almost perfectly correlated (they are
# geometrically linked): harmless for trees, handled by regularization in logistic regression

Step 3: split before anything learns from data

from sklearn.model_selection import train_test_split
X, y = df.drop(columns="target"), df["target"]
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, stratify=y, random_state=42)
print(X_train.shape, X_test.shape, round(y_train.mean(), 3), round(y_test.mean(), 3))
# (455, 30) (114, 30) 0.374 0.368      stratification preserved the class ratio

Step 4: pipeline and cross-validated baseline

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, cross_val_score

pipe = Pipeline([("scale", StandardScaler()),
                 ("clf", LogisticRegression(max_iter=5000))])
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(pipe, X_train, y_train, cv=cv, scoring="recall")
print(scores.round(3), round(scores.mean(), 3), round(scores.std(), 3))
# [0.882 0.971 1.    0.941 0.971] 0.953 0.04     fold-to-fold spread matters on small data

Step 5: tune hyperparameters on training data only

from sklearn.model_selection import GridSearchCV
grid = GridSearchCV(pipe,
                    {"clf__C": [0.01, 0.1, 1, 10, 100],
                     "clf__class_weight": [None, "balanced"]},
                    cv=cv, scoring="recall", n_jobs=-1)
grid.fit(X_train, y_train)
print(grid.best_params_, round(grid.best_score_, 3))
# {'clf__C': 1, 'clf__class_weight': 'balanced'} 0.959

Step 6: choose the decision threshold from out-of-fold predictions

from sklearn.model_selection import cross_val_predict
from sklearn.metrics import precision_recall_curve
best = grid.best_estimator_
oof = cross_val_predict(best, X_train, y_train, cv=cv, method="predict_proba")[:, 1]
prec, rec, thr = precision_recall_curve(y_train, oof)
ok = np.where(rec[:-1] >= 0.98)[0]           # thresholds meeting the recall target
threshold = thr[ok[-1]]                       # the highest such threshold keeps precision best
print(round(threshold, 3), round(rec[ok[-1]], 3), round(prec[ok[-1]], 3))
# 0.284 0.982 0.923      on training folds: 98% recall at 92% precision

Step 7: open the sealed envelope once

from sklearn.metrics import classification_report, confusion_matrix, roc_auc_score, average_precision_score
proba = best.predict_proba(X_test)[:, 1]
y_pred = (proba >= threshold).astype(int)
print(confusion_matrix(y_test, y_pred))
# [[70  2]      70 benign correct, 2 false alarms
#  [ 2 40]]      2 missed malignant, 40 caught
print(classification_report(y_test, y_pred, target_names=["benign", "malignant"], digits=3))
#               precision    recall  f1-score   support
#       benign      0.972     0.972     0.972        72
#    malignant      0.952     0.952     0.952        42
#     accuracy                          0.965       114
print(round(roc_auc_score(y_test, proba), 3), round(average_precision_score(y_test, proba), 3))
# 0.995 0.993

Test recall (0.952) is a little below the 0.98 seen on training folds. With only 42 malignant test cases, one tumour moves recall by 2.4 percentage points, so this gap is within noise; the honest report is "about 95 to 98% recall with roughly 95% precision", together with the fold standard deviation. If the clinical requirement is stricter, lower the threshold further and accept more false alarms, deciding on validation data, never by peeking at the test set again.

Step 8: interpret and save

coefs = pd.Series(best.named_steps["clf"].coef_[0], index=X.columns)
print(coefs.sort_values(key=abs, ascending=False).head(5).round(3))
# worst texture          1.466
# radius error           1.309
# worst symmetry         1.077
# mean concave points    0.970
# compactness error     -0.961       coefficients are per standard deviation, because inputs were scaled
import joblib
joblib.dump({"model": best, "threshold": float(threshold), "features": list(X.columns)},
            "breast_cancer_v1.joblib")

Because many features are strongly correlated, individual coefficients can shift sign or size between fits; use permutation importance or a regularized, reduced feature set before drawing medical conclusions. A PyTorch MLP on the same split (see the PyTorch section) reaches the same 0.974 accuracy at the default threshold, which is a useful lesson: on small tabular data, a well-tuned linear model is a very strong baseline.

Common pitfall Three mistakes this example avoids: scaling before splitting, optimising accuracy when the cost of errors is asymmetric, and picking the threshold by trying values on the test set. A fourth, subtler one: forgetting the relabelling, which makes scoring="recall" optimise recall for benign tumours instead.
Interview angle Being able to narrate this pipeline, with the reasons behind stratification, scaling inside the pipeline, recall as metric, out-of-fold threshold selection, and a single test evaluation, is often worth more than knowing exotic models. Expect follow-ups such as "how would this change with 1% positives?" (PR-AUC, stronger class weighting, more careful thresholds) or "multiple scans per patient?" (group-aware splitting).

Quick revision

  • Python is the orchestration layer; speed comes from keeping work inside compiled libraries (NumPy, BLAS, cuDNN) rather than Python loops.
  • Use one isolated environment per project (venv or conda) with pinned versions; use %pip inside notebooks.
  • Restart the kernel and run all cells before sharing a notebook, to expose hidden state.
  • Variables are names bound to objects; assignment never copies.
  • Lists, dicts and sets are mutable; ints, floats, strings, tuples and frozensets are immutable.
  • == compares values, is compares identity; use is only with None and other singletons.
  • Never use a mutable default argument; default to None and create the object inside.
  • Dict and set lookups are O(1) on average; list membership is O(n). Keys must be hashable.
  • *args collects extra positional arguments into a tuple; **kwargs collects extra keyword arguments into a dict.
  • A closure remembers variables from its enclosing scope; closures bind late, so loop variables are read at call time.
  • A decorator is a function that takes a function and returns a wrapped one; always use functools.wraps.
  • Generators (yield) produce values lazily, use constant memory, and can be consumed only once.
  • Dunder methods (__len__, __getitem__, __call__, __enter__) plug your classes into Python syntax and into frameworks.
  • Type hints are not enforced at runtime; tools such as mypy and libraries such as Pydantic read them.
  • with guarantees cleanup via __exit__, even when an exception is raised.
  • The GIL allows one thread to run Python bytecode at a time: threads or asyncio for I/O-bound work, processes for CPU-bound Python work.
  • A full PyTorch checkpoint (state_dict plus epoch or config) needs torch.load(..., weights_only=False) and must be trusted; tensor-only loads can keep the safer default.
  • NumPy arrays are homogeneous, typed and contiguous, which enables vectorized C loops.
  • Broadcasting aligns shapes from the right; dimensions must match or be 1.
  • Basic slicing returns a view; fancy and boolean indexing return copies.
  • axis=0 collapses rows (one result per column); axis=1 collapses columns (one result per row).
  • * is element-wise, @ is matrix multiplication.
  • Subtract the maximum before exponentiating in softmax for numerical stability.
  • NumPy's var/std default to ddof=0; pandas defaults to ddof=1.
  • loc selects by label with an inclusive end; iloc selects by position with an exclusive end.
  • Combine pandas masks with &, |, ~ and parentheses.
  • groupby().agg() returns one row per group; transform() returns a result aligned to the original rows.
  • Validate merges (validate=, row counts) to catch duplicate keys that explode joins.
  • Vectorized column operations beat apply, which beats iterrows.
  • Assign with a single df.loc[mask, col] = value; chained assignment does not modify the original.
  • Shrink pandas memory with downcasting, category dtype, usecols, Parquet and chunking.
  • EDA checklist: structure, quality, univariate, target, bivariate, multivariate, leakage, decisions.
  • scikit-learn contract: fit learns, transform/predict apply; learned attributes end with an underscore.
  • Put every data-dependent step inside a Pipeline so cross-validation cannot leak.
  • Use CV to choose models; use a paired bootstrap or McNemar on a sealed test set to compare two finished models.
  • ColumnTransformer applies different preprocessing to numeric and categorical columns.
  • Stratify classification splits; use group splits for repeated entities and time-series splits for temporal data.
  • Tune on training data with CV; touch the test set once at the end.
  • Choose metrics by error cost: recall when misses are costly, precision when false alarms are costly, PR-AUC for rare positives.
  • Pick the decision threshold on validation or out-of-fold predictions, not the test set.
  • Scale features for linear models, SVM, kNN, PCA and neural networks; trees do not need it.
  • PyTorch loop: zero_grad, forward, loss, backward, step.
  • model.eval() changes dropout and BatchNorm behaviour; torch.no_grad() stops gradient tracking; inference needs both.
  • CrossEntropyLoss takes raw logits and integer class labels; BCEWithLogitsLoss takes logits and float targets.
  • Save state_dicts; weights_only=True is the safe default, but a checkpoint that also stores epoch or config needs weights_only=False from a trusted file.
  • Hugging Face: pipeline for quick tasks, AutoTokenizer plus AutoModelFor... for control; tokenizer and model must match.
  • Profile before optimising; vectorize first, compile or parallelize second, change engine last.

Glossary

apply
pandas method that runs a Python function per element, row or group; flexible but slow compared with vectorized operations.
asyncio
Python's standard library for single-threaded concurrency using async/await, ideal for many simultaneous network calls.
Autograd
PyTorch's automatic differentiation engine, which records operations and computes gradients with the chain rule on backward().
Axis
A dimension of an array; the axis argument names the dimension an aggregation collapses.
Broadcasting
NumPy and PyTorch rules for combining arrays of different shapes by virtually stretching size-1 dimensions.
Categorical dtype
pandas type storing each distinct value once plus integer codes per row, saving memory for repetitive values.
Closure
A function that remembers variables from the scope in which it was defined.
ColumnTransformer
scikit-learn tool that applies different transformers to different column subsets and concatenates the results.
Context manager
Object used with with that sets something up in __enter__ and guarantees cleanup in __exit__.
Copy-on-Write
pandas behaviour (default from version 3.0) where derived objects act as copies, so modifying one never silently changes another.
Cross-validation
Repeatedly splitting training data into training and validation folds to estimate performance more reliably than a single split.
Data leakage
Information unavailable at prediction time (often from the test set or the target) sneaking into training, inflating reported performance.
Dataclass
Class decorated with @dataclass, which auto-generates __init__, __repr__ and __eq__ from annotated fields.
DataFrame
pandas two-dimensional labelled table made of columns (Series) sharing one index.
DataLoader
PyTorch utility that batches, shuffles and optionally loads data in parallel worker processes from a Dataset.
Decorator
A callable that takes a function and returns a modified or wrapped function, applied with @name.
dtype
The data type of array or column elements, such as float32, int64 or category.
Dunder method
A "double underscore" special method like __len__ that Python calls to implement built-in syntax and functions.
EDA
Exploratory data analysis: systematic inspection of data quality, distributions and relationships before modelling.
Estimator
Any scikit-learn object that learns from data via fit.
Fancy indexing
Indexing an array with lists or arrays of integers; always returns a copy.
Generator
A function using yield that produces values lazily, one at a time, suspending between values.
GIL
Global Interpreter Lock: a CPython mutex allowing only one thread to execute Python bytecode at a time per process.
GridSearchCV
scikit-learn search that cross-validates every combination of a hyperparameter grid and refits the best.
groupby
pandas split-apply-combine operation that groups rows by keys and aggregates, transforms or filters each group.
Hashable
An object with a stable hash and equality, which can therefore be a dict key or set member.
iloc
pandas position-based indexer; slice ends are exclusive.
Iterator
Object implementing __next__ (and __iter__) that yields successive values and raises StopIteration when done.
Jupyter kernel
The live Python process behind a notebook that keeps state between cell executions.
loc
pandas label-based indexer that also accepts boolean masks; slice ends are inclusive.
Logits
Raw, unnormalised model outputs before softmax or sigmoid.
melt
pandas operation that reshapes a wide table into a long one.
merge
pandas SQL-style join of two DataFrames on key columns.
MRO
Method resolution order: the sequence Python searches through base classes to find a method.
Mutable
Able to be changed in place after creation (lists, dicts, sets, arrays).
ndarray
NumPy's n-dimensional, homogeneous, typed array.
nn.Module
Base class for PyTorch layers and models, which registers parameters and submodules and defines forward.
no_grad
PyTorch context that disables gradient tracking, used for inference and manual weight updates.
One-hot encoding
Representing a category as a vector with a single 1 in the column for that category.
Parquet
Compressed, columnar file format that preserves dtypes and allows reading only selected columns.
Pipeline
scikit-learn object chaining transformers and a final estimator so they are fitted and applied together.
pivot_table
pandas reshaping that aggregates values into a grid of index versus column categories.
requires_grad
PyTorch tensor flag indicating that operations on it should be recorded for gradient computation.
Series
pandas one-dimensional labelled array.
state_dict
Dictionary mapping a PyTorch model's or optimizer's parameter names to tensors; the recommended save format.
Stratification
Splitting data so each part keeps the same class proportions as the whole.
Tensor
PyTorch's n-dimensional array, which can live on CPU or GPU and can track gradients.
Tokenizer
Component that converts text into integer token IDs (and back) for a specific language model.
transform (groupby)
pandas group operation that returns a result aligned to the original rows, such as each row's share of its group total.
Transformer (scikit-learn)
An estimator with a transform method, such as a scaler, encoder or imputer.
ufunc
NumPy universal function that operates element-wise on arrays in compiled code.
Vectorization
Expressing computation as whole-array operations instead of explicit Python loops.
View
An array object that shares memory with another array, created for example by basic slicing.
Virtual environment
An isolated folder with its own Python interpreter and packages for one project.

Interview questions

Fundamentals

What is the difference between a list and a tuple, and when would you use each?

Both are ordered sequences. A list is mutable (append, remove, assign items); a tuple is immutable. Because tuples cannot change, they are hashable when their contents are hashable, so they can be dict keys or set members, and they signal "fixed record" intent, such as a tensor shape (32, 3, 224, 224) or a function returning several values. Tuples are also slightly smaller and faster to create. Use a list for collections that grow or change, such as a loss history.

Which Python built-in types are mutable and which are immutable?

Mutable: list, dict, set, bytearray, and most user-defined objects (also NumPy arrays and DataFrames). Immutable: int, float, complex, bool, str, bytes, tuple, frozenset, None. Note that a tuple containing a list is itself immutable (you cannot swap the list out) but the inner list can still change, and such a tuple is not hashable.

What is the difference between is and ==?

== checks value equality by calling __eq__. is checks identity: whether both names refer to the very same object. [1, 2] == [1, 2] is True but [1, 2] is [1, 2] is False. Use is for singletons like None, True, False. CPython caches small integers and some strings, so is sometimes appears to work on them, but relying on that is a bug.

What does this print, and why? a = [1, 2]; b = a; b.append(3); print(a)

It prints [1, 2, 3]. Assignment binds a second name to the same list object; it does not copy. Appending through b mutates the one shared list. To get an independent list use a.copy(), list(a) or a[:] (shallow), or copy.deepcopy(a) for nested structures.

Explain shallow copy versus deep copy.

A shallow copy creates a new outer container but fills it with references to the same inner objects; a deep copy recursively copies everything.

import copy
grid = [[0, 0], [0, 0]]
s = grid.copy(); d = copy.deepcopy(grid)
grid[0][0] = 9
print(s[0][0], d[0][0])   # 9 0

With pandas, df.copy() is deep by default for the data; with NumPy, arr.copy() duplicates the buffer.

What is the mutable default argument trap?

Default values are evaluated once, when the def statement runs, not on each call. A mutable default is therefore shared across calls:

def add(x, acc=[]):
    acc.append(x); return acc
add(1); print(add(2))    # [1, 2]

def add(x, acc=None):    # fix
    if acc is None:
        acc = []
    acc.append(x); return acc
What are *args and **kwargs?

In a function definition, *args gathers extra positional arguments into a tuple and **kwargs gathers extra keyword arguments into a dict. At a call site, *seq and **mapping unpack them back into arguments. They are used to write flexible wrappers and to forward options, for example a decorator's wrapper(*args, **kwargs) or model(**tokenizer_output). The names are convention; the stars are what matter.

What is a list comprehension, and when should you not use one?

A compact expression that builds a list from an iterable with optional filtering: [x * x for x in nums if x % 2 == 0]. It is usually faster and clearer than a loop with append. Avoid it when the logic needs several nested conditions or side effects (printing, writing files), when you do not need the list (use a generator expression, for example inside sum()), or when the data is numeric and large (use NumPy vectorization).

What is the difference between a list comprehension and a generator expression?

[x * x for x in data] builds the whole list in memory immediately. (x * x for x in data) creates a lazy generator that computes each value when requested, using constant memory, but it can be iterated only once and has no len() or indexing. For a million items the list takes about 8 MB while the generator object takes around 200 bytes.

What is a lambda function?

An anonymous single-expression function: lambda r: r["f1"]. Common as a key= for sorted/max, in pandas assign, or as a quick callback. It cannot contain statements. For anything longer, or anything you need to test or pickle (multiprocessing cannot pickle lambdas), define a named function.

How do you handle exceptions in Python? What do else and finally do?

Use try/except SpecificError. The else block runs only if no exception occurred (keep the code that may fail in try small, and put follow-up code in else). finally always runs, for cleanup. Catch specific exceptions, never a bare except:, and re-raise with raise NewError(...) from e to keep the cause. Define custom exceptions by subclassing Exception or a more specific built-in.

What is a virtual environment and why use one?

An isolated directory containing its own interpreter link and installed packages, created with python -m venv .venv (or conda). It prevents version conflicts between projects, makes dependencies explicit (pip freeze > requirements.txt), and allows reproducing the environment on another machine or in CI. conda additionally manages non-Python binaries such as CUDA libraries.

What does if __name__ == "__main__": do?

When a file is run directly, its __name__ is "__main__"; when imported, it is the module name. The guard lets a file be both an importable module and a script. It is essential on Windows and macOS for multiprocessing and PyTorch DataLoader workers, which start new processes that re-import the main module; without the guard the process-creating code would run again in every child.

What is NumPy and why is it faster than Python lists?

NumPy provides the ndarray: a contiguous block of same-typed values plus shape and stride metadata. Operations run in compiled C loops (often using SIMD instructions and optimized BLAS) with no per-element type checks or object overhead. A Python list stores pointers to separately allocated objects, and each operation goes through the interpreter. Vectorized NumPy code is typically 10 to 100 times faster and uses far less memory.

What do shape, ndim, dtype and size tell you about an array?

shape is the length along each dimension, such as (3, 4); ndim is the number of dimensions (2); dtype is the element type (int64, float32); size is the total number of elements (12). itemsize is bytes per element and nbytes is total bytes. Printing shapes is the first debugging step for most ML bugs.

What is the difference between * and @ for NumPy arrays?

* is element-wise multiplication (shapes must match or broadcast). @ (or np.matmul/np.dot for 2-D) is matrix multiplication: (n, k) @ (k, m) gives (n, m). For [[1,2],[3,4]], X * X is [[1,4],[9,16]] and X @ X is [[7,10],[15,22]].

What does axis=0 versus axis=1 mean?

The axis is the dimension that gets collapsed. For a 2-D array of shape (rows, columns), sum(axis=0) adds down the rows and returns one value per column; sum(axis=1) adds across the columns and returns one value per row. In pandas the same logic applies: df.mean() (axis 0) gives column means, and df.drop(columns=...) is the readable alternative to axis=1.

What is the difference between a pandas Series and a DataFrame?

A Series is a one-dimensional labelled array with one dtype. A DataFrame is a two-dimensional table: an ordered collection of Series (columns) sharing one row index, where each column may have a different dtype. Selecting one column with df["col"] returns a Series; df[["col"]] returns a one-column DataFrame.

What is the difference between loc and iloc?

loc selects by labels (index values and column names) and boolean masks, and its slices include the end label. iloc selects by integer position, and its slices exclude the end, like Python. With the default RangeIndex, df.loc[0:2] returns three rows and df.iloc[0:2] returns two. After filtering or sorting, labels and positions no longer coincide, which is where bugs appear.

How do you check for and handle missing values in pandas?

Detect with df.isna().sum() or df.isna().mean() for fractions. Options: drop rows or columns (dropna, with subset or thresh), fill with a constant, mean, median or mode (fillna), fill within groups (groupby().transform("median")), forward-fill for time series, or model-based imputation. Adding a "was missing" indicator often helps. For ML, do the imputation inside a scikit-learn pipeline so statistics are learned from training folds only.

How do you filter rows by multiple conditions in pandas?

Build boolean masks and combine them with & (and), | (or), ~ (not), wrapping each condition in parentheses because these operators bind tighter than comparisons:

df[(df.salary > 90) & (df.dept == "ml")]
df[df.city.isin(["Pune", "Delhi"]) & ~df.name.str.startswith("A")]
df.query("salary > 90 and dept == 'ml'")

Using Python's and/or raises "The truth value of a Series is ambiguous".

What does groupby do?

It implements split-apply-combine: split rows into groups by one or more keys, apply a function to each group (aggregate, transform, or filter), and combine the results. Example: sales.groupby("city")["sales"].sum() gives total sales per city. Named aggregation, agg(total=("sales", "sum"), avg=("units", "mean")), gives clean column names.

What is the difference between merge, join and concat?

pd.merge is a SQL-style join on key columns with how = inner, left, right, outer or cross. DataFrame.join is a convenience method that joins on the index by default. pd.concat stacks objects along an axis: rows (axis=0, same columns) or side by side (axis=1, aligned on index), with no key matching beyond index alignment.

What is scikit-learn's estimator API?

Every estimator is configured through its constructor and learns with fit(X, y), which returns self. Transformers add transform and fit_transform; predictors add predict, and most classifiers add predict_proba or decision_function; models have score. Learned attributes end in an underscore (coef_, mean_). get_params/set_params make every estimator tunable by generic tools like GridSearchCV.

What is the difference between fit, transform and fit_transform?

fit learns parameters from data (a scaler learns means and standard deviations). transform applies those learned parameters to data. fit_transform does both in one step on the same data, sometimes more efficiently. Use fit_transform on training data only; call transform on validation, test and production data so they are processed with training statistics.

Why do we split data into training and test sets, and what does stratify do?

The test set estimates performance on unseen data; if the model or any preprocessing sees it, the estimate becomes optimistic. train_test_split(X, y, test_size=0.2, stratify=y, random_state=42) keeps the class proportions equal in both parts, which matters for imbalanced or small datasets (for example, 37% malignant in both the training and test sets). random_state makes the split reproducible.

Why is feature scaling needed, and for which models?

Models based on distances (kNN, k-means, SVM with RBF kernel), on gradients (logistic regression, neural networks) or on variance (PCA) are dominated by large-scale features if features are not scaled; optimizers also converge faster on standardized inputs. Tree-based models split on thresholds one feature at a time, so they are scale-invariant. On the breast-cancer data, scaling lifts kNN from about 0.91 to 0.96 accuracy.

What is a PyTorch tensor and how does it differ from a NumPy array?

Both are n-dimensional typed arrays with similar operations. Tensors additionally can live on GPUs or other accelerators (.to("cuda")) and can record operations for automatic differentiation (requires_grad=True). Defaults differ: PyTorch uses float32 while NumPy uses float64. torch.from_numpy and .numpy() share memory on CPU.

What are the five steps of a PyTorch training iteration?

optimizer.zero_grad() to clear old gradients; logits = model(xb) forward pass; loss = loss_fn(logits, yb); loss.backward() to compute gradients; optimizer.step() to update parameters. Wrap the epoch with model.train(), and validation with model.eval() plus torch.no_grad().

What is Jupyter, and what are its main risks?

An interactive notebook environment where code cells run in a persistent kernel, mixing code, output, plots and prose; ideal for exploration and teaching. Risks: hidden state and out-of-order execution (results depend on which cells ran when), poor version-control diffs, and code that is hard to test or reuse. Mitigate with "restart and run all", moving reusable code into modules, and tools that strip outputs or pair notebooks with scripts.

How would you compute the mean, median, mode, variance and standard deviation in Python?

By hand you can use sum(x)/len(x), sort for the median, and a counting dict for the mode, but in practice use the libraries:

import numpy as np, pandas as pd, statistics
d = [22, 18, 14, 10, 15, 20, 25, 12]
np.mean(d), np.median(d)            # 17.0 16.5
statistics.multimode([1, 2, 2, 3])  # [2]
np.var(d), np.std(d)                # 23.25 4.82  (population, ddof=0)
pd.Series(d).var()                  # 26.57       (sample, ddof=1)
np.percentile(d, [25, 75])          # quartiles

State which variance formula you use; the sample version divides by n − 1.

Going deeper

What is a decorator? Write one that times a function.

A decorator is a callable that takes a function and returns a new function, usually wrapping the original with extra behaviour. @timer above def f means f = timer(f).

import time, functools
def timer(func):
    @functools.wraps(func)
    def wrapper(*args, **kwargs):
        t0 = time.perf_counter()
        try:
            return func(*args, **kwargs)
        finally:
            print(f"{func.__name__}: {time.perf_counter() - t0:.3f}s")
    return wrapper

functools.wraps preserves the name and docstring; *args, **kwargs forwards any signature; the result is returned.

How do you write a decorator that takes arguments?

Add one more level: a factory that receives the arguments and returns the actual decorator.

def retry(times=3):
    def decorator(func):
        @functools.wraps(func)
        def wrapper(*args, **kwargs):
            for i in range(times):
                try:
                    return func(*args, **kwargs)
                except ConnectionError:
                    if i == times - 1:
                        raise
        return wrapper
    return decorator

@retry(times=5)
def fetch(): ...
What is a closure, and what is the late-binding gotcha?

A closure is an inner function that keeps access to variables of its enclosing function after that function has returned (for example, a learning-rate scheduler remembering base_lr). Closures look up variables when called, not when defined, so [lambda: i for i in range(3)] gives three functions that all return 2. Fix by binding a default: lambda i=i: i, or use functools.partial. Use nonlocal to reassign an enclosing variable.

Explain the LEGB scope rule.

Python resolves names by searching Local (current function), Enclosing (outer functions), Global (module), then Built-in scope. Assigning to a name inside a function makes it local for the whole function, which is why reading a global and then assigning to it in the same function raises UnboundLocalError. Use global or nonlocal to rebind outer names, though passing values explicitly is usually cleaner.

What is the difference between an iterable, an iterator and a generator?

An iterable has __iter__ returning an iterator (lists, dicts, files). An iterator has __next__, returns successive values, raises StopIteration when done, and is itself iterable. A generator is an iterator created by a function containing yield (or a generator expression); Python saves its frame between values. A list can be iterated many times; an iterator or generator is exhausted after one pass.

Write a generator that yields mini-batches from a dataset.
import numpy as np
def minibatches(X, y, batch_size, shuffle=True, seed=0):
    idx = np.arange(len(X))
    if shuffle:
        np.random.default_rng(seed).shuffle(idx)
    for start in range(0, len(X), batch_size):
        b = idx[start:start + batch_size]
        yield X[b], y[b]

for xb, yb in minibatches(np.arange(10).reshape(5, 2), np.arange(5), 2, shuffle=False):
    print(xb.shape)    # (2, 2) (2, 2) (1, 2)

It holds only one batch at a time, handles the final partial batch, and re-shuffles per epoch if you pass a new seed.

What are __str__ and __repr__, and how do they differ?

__repr__ is the unambiguous developer representation, ideally valid code to recreate the object (Neuron(weights=[0.5], bias=0.1)); it is used by the REPL, containers and debuggers. __str__ is the friendly end-user representation used by print and str(); if absent, Python falls back to __repr__. Implement at least __repr__.

What is the difference between @staticmethod, @classmethod and an instance method?

An instance method receives the instance (self) and can use its state. A classmethod receives the class (cls), commonly used for alternative constructors such as Config.from_yaml(path) or AutoModel.from_pretrained(name), and works correctly for subclasses. A staticmethod receives neither; it is a plain function namespaced inside the class for organisation.

What does super() do, and what is the MRO?

super() returns a proxy that delegates to the next class in the method resolution order, so super().__init__() runs the parent initializer. The MRO (C3 linearization, visible in Cls.__mro__) defines a consistent order across multiple inheritance so each base is visited once. In PyTorch, super().__init__() in an nn.Module subclass sets up parameter registries and must be called before assigning layers.

What is a context manager, and how do you write one?

An object used with with whose __enter__ runs at the start and __exit__ at the end, even if an exception occurs; __exit__ receives exception details and can suppress the exception by returning True. Write one as a class, or with @contextlib.contextmanager around a generator that does setup, yields, then cleans up in finally. Examples: files, locks, torch.no_grad(), temporary pandas options.

When would you use a dataclass, a namedtuple, a dict or a Pydantic model?

A dict for loose, dynamic data such as parsed JSON. A namedtuple for lightweight immutable records with positional access. A @dataclass for typed, readable configuration or record objects with defaults, optional immutability (frozen=True) and generated methods, without runtime validation. A Pydantic model when data crosses a trust boundary (API requests, LLM output) and must be validated and converted at runtime.

What is the GIL, and how does it affect ML code?

The Global Interpreter Lock lets only one thread execute Python bytecode at a time in a CPython process, which simplifies memory management (reference counting). Consequences: threads do not speed up CPU-bound pure-Python code, but they do help I/O-bound code because waiting threads release the GIL. NumPy, pandas, scikit-learn and PyTorch release the GIL inside heavy C routines and often multithread internally. For CPU-bound Python, use processes (multiprocessing, joblib, DataLoader workers). An optional free-threaded CPython build exists in recent versions but is not yet the default.

Threads, processes or asyncio: which would you choose for (a) 5,000 API calls, (b) parsing 200 large log files with pure Python, (c) matrix multiplication?

(a) asyncio with an async HTTP client and a semaphore to respect rate limits, or a thread pool if only a synchronous client exists. (b) A process pool, since parsing is CPU-bound Python and processes bypass the GIL. (c) Neither: call NumPy or PyTorch, which already use optimized multithreaded BLAS or the GPU.

State NumPy's broadcasting rules and give the result shape of (8, 1, 6, 1) combined with (7, 1, 5).

Align shapes from the right; prepend 1s to the shorter shape; each dimension pair must be equal or contain a 1; the result takes the larger size. (8, 1, 6, 1) and (1, 7, 1, 5) give (8, 7, 6, 5). Shapes (3,) and (4,) are incompatible and raise "operands could not be broadcast together".

What is the difference between a view and a copy in NumPy? How can you tell?

A view shares memory with the original, so writes through it change the original; basic slicing, reshape (when possible), transpose and ravel (when possible) return views. Fancy indexing with integer lists, boolean masks, flatten and .copy() return copies. Check with np.shares_memory(a, b) or b.base is a.

b = np.arange(6); v = b[1:4]; v[0] = 99
print(b)          # [ 0 99  2  3  4  5]
Implement a numerically stable softmax in NumPy that works on a batch.
def softmax(z, axis=-1):
    z = z - z.max(axis=axis, keepdims=True)
    e = np.exp(z)
    return e / e.sum(axis=axis, keepdims=True)
softmax(np.array([1000., 1001., 1002.]))   # [0.09 0.245 0.665]  (naive exp overflows to inf)

Subtracting the maximum does not change the result (it cancels in the ratio) but keeps exponents at or below zero. keepdims=True keeps shapes broadcastable for 2-D batches.

Standardize each column of a matrix and compute pairwise Euclidean distances without loops.
Xs = (X - X.mean(axis=0)) / X.std(axis=0)        # broadcasting (n, d) with (d,)

# pairwise distances between A (n, d) and B (m, d)
D = np.sqrt(((A[:, None, :] - B[None, :, :]) ** 2).sum(-1))       # (n, m); memory n*m*d
# memory-lean version using |a-b|^2 = |a|^2 + |b|^2 - 2ab
D2 = (A**2).sum(1)[:, None] + (B**2).sum(1)[None, :] - 2 * A @ B.T
D = np.sqrt(np.maximum(D2, 0))                   # clip tiny negatives from rounding
Compute cosine similarity between a query vector and every row of a matrix in NumPy.
def cosine_sim(q, M):
    q = q / np.linalg.norm(q)
    M = M / np.linalg.norm(M, axis=1, keepdims=True)
    return M @ q                                 # shape (n,)
top3 = np.argsort(-cosine_sim(q, M))[:3]         # or np.argpartition for large n

This is the core of embedding search in retrieval systems. Guard against zero-norm rows by adding a small epsilon.

What is the difference between agg, transform and apply after a groupby?

agg reduces each group to one row (sum, mean, custom). transform returns a result the same length as the input, aligned to the original rows, ideal for features like "share of group total" or "fill with group median". apply accepts any function returning a scalar, Series or DataFrame; it is the most flexible and the slowest. Prefer built-in string names ("sum", "mean") which run in optimized code.

df["share"] = df.sales / df.groupby("city").sales.transform("sum")
How do you get the top N rows per group in pandas?
# top 2 products by revenue within each region
(df.sort_values("revenue", ascending=False)
   .groupby("region")
   .head(2))
# or with ranking, keeping ties explicit
df[df.groupby("region").revenue.rank(method="first", ascending=False) <= 2]
# single best row per group
df.loc[df.groupby("region").revenue.idxmax()]
Explain pivot, pivot_table and melt.

melt converts wide to long: identifier columns stay, other columns become (variable, value) pairs, the format most plotting and modelling tools prefer. pivot converts long to wide and fails if an index/column pair is duplicated. pivot_table also goes long to wide but aggregates duplicates with aggfunc and can add totals (margins=True), like a spreadsheet pivot table.

Why is df.apply(..., axis=1) slow, and what do you use instead?

It calls a Python function once per row and constructs a Series object for each row, so it runs at interpreter speed with heavy object overhead. Replace with vectorized column arithmetic (df.a / df.b ** 2), np.where/np.select for conditionals, .map(dict) for lookups, .str/.dt accessors, or a merge for table lookups. If a loop is truly unavoidable, use itertuples, or compile the logic with Numba.

What causes SettingWithCopyWarning, and how does Copy-on-Write change things?

Chained indexing such as df[df.x > 0]["y"] = 1 first creates an intermediate object that might be a view or a copy, then assigns into it, so the original may or may not change. Older pandas warns. Under Copy-on-Write (the default behaviour from pandas 3.0), every derived object behaves as a copy, so chained assignment never updates the original and pandas raises a chained-assignment warning. The fix in all versions: df.loc[df.x > 0, "y"] = 1, and .copy() when you deliberately want an independent subset.

How do you reduce the memory usage of a large DataFrame?
  • Measure with df.memory_usage(deep=True).
  • Load only needed columns (usecols) and specify dtype at read time.
  • Downcast numerics: pd.to_numeric(s, downcast="integer"), float64 to float32.
  • Convert low-cardinality strings to category.
  • Store as Parquet; read in chunks; filter early; delete temporaries.

A 1,000-row frame with an int64 and a float64 column drops from about 16 KB to 6 KB after downcasting to int16 and float32.

How do you create lag and rolling features for time-series data without leakage?
df = df.sort_values(["user", "date"])
g = df.groupby("user").amount
df["lag_1"] = g.shift(1)                                  # yesterday's value
df["roll_7"] = g.transform(lambda s: s.shift(1).rolling(7).mean())   # past 7 days, excluding today
df["pct_change"] = g.pct_change()

Shift before rolling so the current row's target period is excluded, compute per entity, never use center=True or negative shifts, and validate with TimeSeriesSplit.

What is a scikit-learn Pipeline, and why is it important?

A Pipeline chains transformers and a final estimator into one estimator. Calling fit fits each step on the output of the previous one; predict passes new data through the fitted steps. Benefits: no leakage in cross-validation (each fold refits preprocessing on its training part), one object to tune (step__param), one object to save and deploy, and less glue code.

How would you preprocess a dataset with numeric and categorical columns in scikit-learn?
pre = ColumnTransformer([
    ("num", make_pipeline(SimpleImputer(strategy="median"), StandardScaler()), num_cols),
    ("cat", make_pipeline(SimpleImputer(strategy="most_frequent"),
                          OneHotEncoder(handle_unknown="ignore")), cat_cols),
])
model = make_pipeline(pre, LogisticRegression(max_iter=1000))

handle_unknown="ignore" keeps prediction working when a new category appears. For high-cardinality categories consider TargetEncoder or hashing; for tree models OrdinalEncoder is often enough.

What is cross-validation, and how do you choose the splitter?

k-fold cross-validation trains on k − 1 folds and validates on the remaining fold, k times, and averages the scores, giving a lower-variance estimate than a single split and a spread (std) to judge stability. Use StratifiedKFold for classification, GroupKFold when multiple rows belong to the same patient or user, TimeSeriesSplit for temporal data, and repeated CV for very small datasets.

GridSearchCV versus RandomizedSearchCV: how do they work and which do you choose?

Grid search evaluates every combination of listed values with cross-validation, then refits the best on the full training data (best_estimator_). Randomized search samples a fixed number of combinations from distributions (for example loguniform(1e-3, 1e3) for C). Random search is more efficient when only a few hyperparameters matter or the space is large; grid search is fine for small, discrete grids. Bayesian tools such as Optuna and successive-halving searches go further.

Precision versus recall: define them and give a case where each matters most.

Precision = TP / (TP + FP): of predicted positives, how many are right. Recall = TP / (TP + FN): of actual positives, how many were found. Recall matters most when misses are costly (cancer screening, fraud, safety defects). Precision matters most when false alarms are costly (spam filters deleting real mail, alerts that wake an engineer). F1 is their harmonic mean. Moving the decision threshold trades one against the other.

ROC-AUC versus PR-AUC: when is each appropriate?

ROC-AUC measures how well scores rank positives above negatives across all thresholds, plotting true-positive rate against false-positive rate; it is insensitive to class balance, which can make it look excellent when positives are rare. PR-AUC (average precision) plots precision against recall and focuses on the positive class, so it better reflects performance under heavy imbalance. The random baseline for ROC-AUC is 0.5; for PR-AUC it equals the positive rate.

What is the difference between model.eval() and torch.no_grad()?

model.eval() sets a flag on every module that changes layer behaviour: dropout stops dropping, and BatchNorm uses running statistics instead of batch statistics. It does not stop gradient tracking. torch.no_grad() (or the stricter inference_mode()) stops building the autograd graph, saving memory and time, but does not change layer behaviour. Inference needs both; remember model.train() before the next training epoch.

Why do you need optimizer.zero_grad()?

PyTorch accumulates gradients into .grad on every backward() call rather than overwriting them. That enables gradient accumulation across micro-batches, but in a normal loop, without zeroing, each step would use the sum of all previous gradients and training would diverge. zero_grad(set_to_none=True), the default in recent versions, sets gradients to None for a small speed and memory gain.

How do you write a custom PyTorch Dataset?

Subclass torch.utils.data.Dataset and implement __len__ (number of samples) and __getitem__(i) (return one sample, typically a (features, label) tuple of tensors). Do expensive loading lazily in __getitem__ for large data (read an image file per index), and let DataLoader handle batching, shuffling and parallel workers. For streams, subclass IterableDataset and implement __iter__.

Which loss function and target format do you use for binary, multiclass and multi-label classification in PyTorch?

Multiclass: nn.CrossEntropyLoss on raw logits of shape (N, C) with long targets of shape (N,). Binary: either one logit with nn.BCEWithLogitsLoss and float targets, or two logits with cross-entropy. Multi-label: one logit per label with BCEWithLogitsLoss and float 0/1 targets of shape (N, L). Never apply softmax or sigmoid before these losses; they include it in a numerically stable way.

What does a Hugging Face tokenizer return, and why do you need an attention mask?

Calling the tokenizer returns a dict-like object with input_ids (integer token IDs, including special tokens like the start and separator tokens), attention_mask (1 for real tokens, 0 for padding), and for some models token_type_ids. When a batch is padded to equal length, the mask tells the model to ignore padding positions so they do not affect attention or pooled outputs. Use padding=True, truncation=True, return_tensors="pt".

Advanced

How does CPython manage memory?

Every object carries a reference count; when it drops to zero the object is freed immediately. Reference counting cannot free cycles (A refers to B, B refers to A), so a generational cyclic garbage collector (gc module) periodically finds and frees unreachable cycles. Small objects come from a dedicated allocator (pymalloc) that keeps memory pools, which is why a Python process may not return memory to the operating system after freeing objects. Large NumPy and PyTorch buffers are allocated separately; GPU memory in PyTorch goes through a caching allocator, so nvidia-smi can show memory as used even after tensors are freed (torch.cuda.empty_cache() releases the cache).

How are Python dicts implemented, and why are lookups O(1) on average?

A dict is a hash table: the key's hash selects a slot, and collisions are resolved by open addressing (probing other slots). Since Python 3.6 the layout is a compact array of entries in insertion order plus a sparse index table, which saves memory and makes iteration order equal insertion order (guaranteed from 3.7). Average lookup, insertion and deletion are O(1); the worst case (many collisions) is O(n). The table resizes as it fills, which makes inserts amortized O(1). Keys must be hashable and must not change their hash while stored.

What are __slots__, and when are they useful?

Declaring __slots__ = ("x", "y") in a class replaces the per-instance __dict__ with fixed storage for those attributes. It reduces memory per object substantially and speeds attribute access slightly, and it prevents creating new attributes by typo. Useful when you create millions of small objects (tokens, graph nodes); @dataclass(slots=True) does it for you. Downsides: no dynamic attributes, and care is needed with inheritance.

What is the descriptor protocol, and how does @property use it?

A descriptor is an object defining __get__, and optionally __set__/__delete__, stored as a class attribute. When you access obj.attr, Python finds the descriptor on the class and calls its __get__(obj, type) instead of returning it directly. property is a built-in descriptor wrapping getter, setter and deleter functions; functions themselves are descriptors, which is how methods get bound to self. ORMs, Pydantic fields and functools.cached_property use the same mechanism.

How does asyncio work under the hood?

An event loop runs in one thread and manages coroutines (async def functions). When a coroutine hits await on something not ready (a socket read), it yields control back to the loop, which registers interest with the operating system's I/O notification mechanism and runs another ready task. When the I/O completes, the waiting coroutine is resumed. Concurrency is therefore cooperative: only one piece of Python runs at a time, and any blocking call without await stalls every task. asyncio.gather runs many awaitables concurrently; semaphores cap concurrency; asyncio.to_thread offloads blocking code.

What are strides, and how do they make transposes and sliding windows free?

Strides are the number of bytes to jump in memory to move one step along each dimension. A C-ordered int64 array of shape (3, 4) has strides (32, 8). Its transpose simply swaps them to (8, 32) without moving any data, which is why .T is a view but not contiguous. Tricks like sliding_window_view(a, 3) build overlapping windows by reusing memory with carefully chosen strides:

from numpy.lib.stride_tricks import sliding_window_view
sliding_window_view(np.arange(6), 3)   # [[0 1 2] [1 2 3] [2 3 4] [3 4 5]], no copy

Such views are read-only by default; writing to overlapping windows would be confusing.

What is np.einsum, and give three examples.

Einstein summation expresses multiply-and-sum operations with index labels: repeated indices are multiplied, indices absent from the output are summed.

np.einsum("ij,jk->ik", A, B)          # matrix multiply, same as A @ B
np.einsum("ii->", M)                  # trace
np.einsum("ij->j", A)                 # column sums
np.einsum("bhqd,bhkd->bhqk", Q, K)    # batched attention scores per head

It is expressive and avoids intermediate transposes; optimize=True picks an efficient contraction order for multi-operand expressions.

What numerical-precision issues matter in ML code?
  • float32 has about 7 significant digits: np.float32(16777216) + 1 is still 16777216, so summing millions of small values can lose accuracy; accumulate in float64 or use pairwise summation (NumPy's sum already does).
  • exp overflows for large inputs: use max-subtraction in softmax, logsumexp, log1p, and fused losses like BCEWithLogitsLoss.
  • float16 has a small range, so gradients can underflow to zero; mixed-precision training uses loss scaling, while bfloat16 keeps float32's range with less precision.
  • Never test floats with ==; use np.isclose.
What is nested cross-validation, and when do you need it?

An outer CV loop estimates generalisation; inside each outer training split, an inner CV loop performs hyperparameter search. Because the outer validation fold never influenced the tuning, the outer score is an unbiased estimate of the whole "tune and train" procedure. You need it for small datasets where there is no room for a separate test set, or when comparing modelling procedures rigorously.

inner = GridSearchCV(pipe, grid, cv=StratifiedKFold(5, shuffle=True, random_state=0))
outer_scores = cross_val_score(inner, X, y, cv=StratifiedKFold(5, shuffle=True, random_state=1))
How do you write a custom scikit-learn transformer that works in pipelines and grid search?
from sklearn.base import BaseEstimator, TransformerMixin
class Log1p(BaseEstimator, TransformerMixin):
    def __init__(self, cols=None):       # store params unchanged; no logic here
        self.cols = cols
    def fit(self, X, y=None):
        self.n_features_in_ = X.shape[1] # learned state ends with an underscore
        return self
    def transform(self, X):
        X = np.asarray(X, dtype=float).copy()
        idx = self.cols if self.cols is not None else slice(None)
        X[:, idx] = np.log1p(X[:, idx])
        return X

BaseEstimator provides get_params/set_params from the __init__ signature (so __init__ must only store arguments); TransformerMixin provides fit_transform. For stateless functions, FunctionTransformer(np.log1p) is simpler.

Why can target encoding leak, and how do you prevent it?

Target encoding replaces a category with the mean target of rows in that category. If a row's own target contributes to its encoding, rare categories essentially encode the label, and the model learns a shortcut that disappears on new data. Prevent it with cross-fitting (compute each row's encoding from other folds), smoothing toward the global mean for rare categories, and fitting the encoder inside the CV pipeline. scikit-learn's TargetEncoder performs internal cross-fitting in fit_transform.

What is probability calibration, and how do you check and fix it?

A classifier is calibrated if, among predictions of 0.8, about 80% are positive. Many models are not (SVM scores, boosted trees, heavily regularized or class-weighted models). Check with a reliability diagram (CalibrationDisplay) and the Brier score or log loss. Fix with CalibratedClassifierCV using sigmoid (Platt) scaling for small data or isotonic regression for larger data, fitted on data not used for training. Calibration matters when probabilities drive decisions, such as expected-cost thresholds or risk scores.

Implement logistic regression with gradient descent in NumPy.
def train_logreg(X, y, lr=0.1, epochs=500, l2=0.0):
    n, d = X.shape
    w, b = np.zeros(d), 0.0
    for _ in range(epochs):
        p = 1 / (1 + np.exp(-(X @ w + b)))        # predicted probabilities
        grad_w = X.T @ (p - y) / n + l2 * w       # gradient of mean log loss (+ L2)
        grad_b = (p - y).mean()
        w -= lr * grad_w
        b -= lr * grad_b
    return w, b
# on standardized breast-cancer data: training accuracy about 0.986

Key points: vectorized gradient X.T @ (p - y) / n, standardize inputs first, and for stability compute the loss with np.logaddexp if you log it.

Implement k-means in NumPy.
def kmeans(X, k, iters=100, seed=0):
    rng = np.random.default_rng(seed)
    C = X[rng.choice(len(X), k, replace=False)]            # init from data points
    for _ in range(iters):
        labels = ((X[:, None, :] - C[None]) ** 2).sum(-1).argmin(1)   # assign
        newC = np.array([X[labels == j].mean(0) if np.any(labels == j) else C[j]
                         for j in range(k)])                          # update
        if np.allclose(newC, C):
            break
        C = newC
    return C, labels

Mention k-means++ initialization, multiple restarts (n_init), handling empty clusters, and scaling features first.

How does PyTorch's autograd graph work? What are leaf tensors and retain_graph?

PyTorch builds the graph dynamically during the forward pass: each result tensor stores a grad_fn pointing to the operation that created it and its inputs. Leaf tensors are those created directly by the user (such as parameters) with requires_grad=True; by default only leaves keep .grad after backward() (call retain_grad() on intermediates to keep theirs). After backward() the graph's saved buffers are freed to save memory, so calling backward twice through the same graph fails unless the first call used retain_graph=True. Because the graph is rebuilt every iteration, ordinary Python control flow (if, loops) works in models.

How does mixed-precision training work in PyTorch?

Inside torch.autocast(device_type="cuda", dtype=torch.float16 or torch.bfloat16), eligible operations such as matrix multiplications run in half precision for speed and memory savings, while precision-sensitive operations (reductions, softmax, losses) stay in float32. With float16, small gradients can underflow, so a torch.amp.GradScaler multiplies the loss before backward, unscales before the step, and skips steps with infinite gradients. bfloat16 has float32's exponent range and usually needs no scaler. Master weights stay float32.

DataParallel versus DistributedDataParallel: what is the difference?

DataParallel runs in one process: it splits each batch across GPUs, replicates the model every step and gathers outputs on one GPU, which is limited by the GIL and an overloaded main GPU. DistributedDataParallel runs one process per GPU (possibly across machines); each has its own model copy and data shard (DistributedSampler), and gradients are averaged with all-reduce during backward. DDP is faster and the recommended approach; for models too large for one GPU, sharded approaches such as FSDP split parameters, gradients and optimizer states.

How do you freeze layers and use different learning rates for parts of a model?
for p in model.backbone.parameters():
    p.requires_grad = False                      # frozen: no gradients computed
opt = torch.optim.AdamW([
    {"params": model.head.parameters(), "lr": 1e-3},
    {"params": [p for p in model.encoder.parameters() if p.requires_grad], "lr": 1e-5},
], weight_decay=0.01)

Keep frozen BatchNorm layers in eval mode if you do not want their running statistics updated, since requires_grad=False does not stop those updates. Check the trainable parameter count before training.

How do you make PyTorch experiments reproducible?

Seed every generator (random.seed, np.random.seed or explicit generators, torch.manual_seed, which also seeds CUDA); give the DataLoader a seeded generator and a worker_init_fn for workers; set torch.use_deterministic_algorithms(True) and torch.backends.cudnn.benchmark = False; pin library versions and hardware. Some GPU operations remain non-deterministic or become slower in deterministic mode, so bitwise reproducibility across different GPUs or versions is not guaranteed; report the mean and spread over several seeds.

What are the options for serializing models, and what are their risks?

scikit-learn: joblib/pickle, which requires the same library versions and can execute arbitrary code when loading, so only load trusted files; alternatives include skops for safer loading, or ONNX export for cross-language serving. PyTorch: save the state_dict with torch.save. Recent versions default to weights_only=True, which is safe but rejects a checkpoint dict that also stores epoch or config; use weights_only=False only for trusted full checkpoints. Prefer safetensors for tensor-only storage without pickle, and TorchScript, torch.export or ONNX for deployment without Python class definitions. Always store preprocessing, thresholds, feature lists and version metadata with the model.

What is a MultiIndex in pandas, and when is it useful?

A hierarchical index with several levels, for example (store, date). It results naturally from groupby on multiple keys or pivot_table. It enables selecting by partial keys (df.loc["Pune"], df.xs("2026-01-01", level="date")) and reshaping with stack/unstack. Many people call reset_index() right away because flat columns are easier to merge and export; know both.

When would you use Polars, DuckDB, Dask or Spark instead of pandas?

pandas is ideal when data fits comfortably in memory (a rule of thumb is a few times smaller than RAM). Polars is a multithreaded, Arrow-based DataFrame library with lazy query optimisation, much faster for large single-machine workloads. DuckDB runs fast analytical SQL directly on Parquet or DataFrames with out-of-core execution. Dask parallelises pandas-like code across cores or a cluster. Spark is for very large, distributed data in a cluster. Choose the simplest tool that fits the data size and team skills.

How would you implement scaled dot-product attention in NumPy and check its shapes?
def attention(Q, K, V, mask=None):
    d_k = Q.shape[-1]
    scores = Q @ K.swapaxes(-1, -2) / np.sqrt(d_k)       # (..., q, k)
    if mask is not None:
        scores = np.where(mask, scores, -1e9)             # block masked positions
    w = softmax(scores, axis=-1)                          # rows sum to 1
    return w @ V, w                                       # (..., q, d_v)

Using swapaxes(-1, -2) instead of .T keeps it correct for batched inputs. The theory behind the formula is in Transformers & LLMs.

Scenario & debugging

Your pandas job runs out of memory while loading and processing a 20 GB CSV. What do you do?
  1. Load only needed columns (usecols) with explicit compact dtypes (float32, small ints, category for repetitive strings) and parse dates once.
  2. Prototype on nrows=100_000 to learn dtypes and memory per row.
  3. Stream with chunksize, reduce each chunk (filter, aggregate), and combine the small results.
  4. Convert once to Parquet, partitioned if useful; later reads are faster and can select columns and row groups.
  5. Avoid intermediate copies: drop temporaries, avoid wide apply results, and check for accidental many-to-many merges.
  6. If still too big, use DuckDB or Polars (lazy, out-of-core), or Dask/Spark.
Your training loss becomes NaN after a few hundred steps. How do you debug it?
  • Check the data: np.isnan(X).any(), infinities, division by zero in feature engineering, unscaled huge values.
  • Lower the learning rate; add warm-up; clip gradients (clip_grad_norm_).
  • Look for numerically unsafe operations: log(0), sqrt of negatives, manual softmax or sigmoid followed by log; use CrossEntropyLoss/BCEWithLogitsLoss on logits and add epsilons.
  • With float16 mixed precision, use a GradScaler or switch to bfloat16.
  • Locate the first bad operation with torch.autograd.set_detect_anomaly(True) and by logging gradient norms per layer.
  • Check labels are in range (for example, class index equal to n_classes).
A fraud model shows 99.5% accuracy but the business says it catches nothing. What happened?

With about 0.5% fraud, predicting "not fraud" for everything gives 99.5% accuracy. Accuracy is the wrong metric. Evaluate recall, precision and PR-AUC on the fraud class, compare against a DummyClassifier baseline, use class weighting or resampling inside the pipeline, and pick a decision threshold on validation data that meets the business's recall target at an acceptable alert volume. Report the confusion matrix in counts the business understands ("we catch 70 of 100 frauds with 300 alerts a day").

Cross-validation says 0.95 AUC, but the model performs at 0.70 in production. What are the likely causes?
  • Leakage: preprocessing fitted before splitting, features computed with future information, target-derived columns, or the same entity in train and validation folds (need GroupKFold).
  • Wrong validation scheme for time: random folds on temporal data; use TimeSeriesSplit or an out-of-time holdout.
  • Training-serving skew: preprocessing re-implemented differently in production, different category spellings, default values for missing fields.
  • Data drift: the population or behaviour changed since training.

Investigate by replaying production inputs through the saved pipeline, comparing feature distributions between training and production, and re-validating with a time-based split.

You get "CUDA out of memory" during training. What are your options?
  1. Reduce the batch size, and use gradient accumulation to keep the effective batch size.
  2. Use mixed precision (autocast with bfloat16 or float16).
  3. Make sure evaluation runs under torch.no_grad() and that you log loss.item(), not the tensor, so old graphs are freed.
  4. Delete large unused tensors, avoid keeping predictions on the GPU, and move metrics to CPU.
  5. Use gradient checkpointing (recompute activations in backward), shorter sequence lengths, a smaller model, or parameter-efficient fine-tuning such as LoRA.
  6. Check for other processes on the GPU with nvidia-smi, and inspect torch.cuda.memory_summary().
Your PyTorch training loss does not decrease at all. What do you check?
  • Try to overfit a single small batch; if the model cannot drive the loss near zero, the bug is in the code, not in the data volume.
  • Verify the loop: zero_grad, backward, step all called; the optimizer received model.parameters() (not an empty or stale list); parameters actually have requires_grad=True.
  • Check that labels match inputs after shuffling and that label encoding is correct.
  • Check the loss setup: raw logits into CrossEntropyLoss, correct target dtype and shape (no accidental broadcasting from (N,) versus (N, 1)).
  • Tune the learning rate (too low: no movement; too high: bouncing); check input scaling.
  • Inspect gradient norms per layer for vanishing or exploding gradients.
Evaluating the same trained model twice on the same validation set gives different accuracies. Why?

Almost certainly model.eval() was not called, so dropout is still randomly zeroing activations (and BatchNorm is using batch statistics and updating its running averages). Call model.eval() before evaluation and model.train() afterwards. Other causes: random test-time augmentation, a shuffled DataLoader combined with a metric that depends on order, or non-deterministic GPU kernels (tiny differences only).

After a merge, your DataFrame has more rows than either input. What happened and how do you prevent it?

The join key is duplicated on both sides (a many-to-many relationship), so each matching pair produces a row. Check key uniqueness with df.key.is_unique or df.key.duplicated().sum(), de-duplicate or aggregate the lookup table first, and pass validate="many_to_one" (or "one_to_one") to merge so pandas raises an error. Also check for missing or inconsistent key types (int versus string IDs) that cause silent non-matches, using indicator=True.

A notebook runs perfectly for you but fails for a colleague. How do you fix it?

Likely causes: hidden kernel state (variables from deleted or re-ordered cells), different package versions, absolute local file paths, missing environment variables, or unseeded randomness. Restart the kernel and run all cells from top to bottom; pin dependencies in a lock file or environment file; use relative paths or a config file; seed random generators; and move reusable logic into a module that can be tested.

A feature-engineering script with df.apply(axis=1) takes two hours. How do you speed it up?
  1. Profile (%prun, line_profiler) to confirm where time goes.
  2. Rewrite row-wise logic as column operations: arithmetic, np.where/np.select, .str and .dt accessors, map with dicts.
  3. Replace per-row lookups with a single merge; replace per-group Python functions with built-in groupby aggregations or transform.
  4. Fix dtypes (numbers stored as strings slow everything).
  5. If a sequential loop is unavoidable, compile it with Numba; then parallelize across partitions if still needed.

Such rewrites commonly give speed-ups of 50 to 500 times.

GPU utilisation is only 20% during training. What is the bottleneck and how do you fix it?

Usually the input pipeline: the GPU waits for the CPU to load, decode and augment data. Increase num_workers, set pin_memory=True and non_blocking=True on transfers, use persistent_workers=True and prefetching, pre-process data offline into a fast format, and increase batch size if memory allows. Also avoid synchronising calls in the loop (.item() or printing every step, moving tensors to CPU). Confirm with the PyTorch profiler.

Your DataLoader with num_workers=4 hangs or crashes on Windows. Why?

Windows starts worker processes with the "spawn" method, which re-imports the main script. Without an if __name__ == "__main__": guard, the training code re-executes in every worker. Datasets defined inside a notebook or using lambdas may also fail to pickle. Fix: put the training entry point under the guard, define Dataset classes and collate functions at module level in a .py file, or use num_workers=0 in notebooks.

Every training run gives different results. How do you make them consistent, and is perfect reproducibility realistic?

Fix all seeds (Python, NumPy, PyTorch and random_state in scikit-learn splitters and models), seed DataLoader shuffling, enable deterministic algorithms, and pin library versions. Record the data snapshot and configuration. Exact bitwise results across different hardware, drivers or library versions are not guaranteed, so report the mean and standard deviation across several seeds, which is more honest than one lucky run.

A model loaded in production gives different predictions than it did in the notebook. What could be wrong?
  • Preprocessing mismatch: only the model was saved, and scaling or encoding was re-implemented differently. Save the entire pipeline.
  • Column order or names differ; select features by the saved feature list.
  • Library version differences for pickled objects.
  • For PyTorch: model.eval() not called, a different tokenizer, a different dtype, or a missing custom threshold.
  • Unseen categories or missing values handled by defaults.

Write a regression test that feeds fixed sample inputs and compares outputs to stored expected predictions.

A categorical feature has 50,000 distinct values (for example, merchant ID). How do you encode it?

One-hot encoding creates 50,000 sparse columns, which works for linear models with sparse matrices but is heavy and overfits rare levels. Options: group rare levels into "other" (OneHotEncoder(min_frequency=..., max_categories=...)), frequency or count encoding, cross-fitted target encoding (TargetEncoder), feature hashing, native categorical support in gradient-boosting libraries, or learned embeddings in a neural network. Handle unseen IDs at prediction time explicitly.

A decision tree scores 100% on training data and 70% on test. What do you do?

This is overfitting: an unconstrained tree memorises the training set. Regularize with max_depth, min_samples_leaf, min_samples_split or cost-complexity pruning (ccp_alpha), tuned by cross-validation; or switch to an ensemble (random forest, gradient boosting) that averages many trees. Plot training versus validation score against depth to see the sweet spot between underfitting (depth 1) and overfitting (no limit).

Your classifier predicts only the majority class. What are the possible fixes?

Check the probability outputs first: the model may rank well (good AUC) while the 0.5 threshold is simply too high for a rare class, in which case lower the threshold using validation data. Otherwise use class_weight="balanced" or a weighted loss, resample training folds, ensure features are informative and scaled, and verify that labels were not corrupted by a mapping bug. Evaluate with recall, precision and PR-AUC rather than accuracy.

Dates in your dataset such as "03/04/2026" are being parsed inconsistently. How do you handle it?

Ambiguous formats are parsed month-first by default: pd.to_datetime("03/04/2026") gives 4 March, while dayfirst=True gives 3 April. Always pass an explicit format="%d/%m/%Y" when you know it, use errors="coerce" to turn bad values into NaT and count them, and store dates in ISO 8601 (YYYY-MM-DD) or as typed columns in Parquet. Be explicit about time zones (tz_localize, tz_convert).

You need to run 10,000 prompts through an LLM API for an evaluation. How do you do it efficiently and safely?
  • Use an async client with asyncio.gather and a semaphore sized to the rate limit (or a thread pool with a sync client).
  • Retry transient failures and rate-limit responses with exponential backoff and jitter; set timeouts.
  • Write results incrementally (JSONL) keyed by prompt ID so a crash can resume without repeating work; cache by (model, prompt, parameters).
  • Use temperature=0 for reproducible scoring, validate structured output with Pydantic, and log token usage for cost.
  • Use only approved endpoints and keep keys in environment variables.
GPU or CPU memory keeps growing every epoch in your PyTorch script. What is leaking?

The classic cause is accumulating tensors that still hold their autograd graphs, for example total_loss += loss or appending outputs to a list for later metrics. Use loss.item() and outputs.detach().cpu(). Other causes: evaluation without no_grad, storing whole batches in a history list, and growing caches in custom datasets. Monitor with torch.cuda.memory_allocated() per epoch to confirm the fix.

A medical dataset has several scans per patient, and cross-validation looks too good. Why, and how do you fix it?

With random row-level splitting, scans from the same patient appear in both training and validation folds, so the model can recognise the patient rather than learn the disease; performance on new patients will be lower. Split by patient with GroupKFold or StratifiedGroupKFold(groups=patient_id), and make the final test set contain entirely unseen patients. The same applies to users, devices, sessions and near-duplicate documents.

A clinical team requires recall of at least 0.95 for malignant cases with precision of at least 0.60. How do you deliver a model that meets this?
  1. Relabel so malignant is the positive class, split with stratification, and keep a locked test set.
  2. Build a pipeline (scaling plus model), tune with scoring="recall" or average precision, and consider class_weight="balanced".
  3. Get out-of-fold probabilities on the training data (cross_val_predict(..., method="predict_proba")), compute the precision-recall curve, and choose the highest threshold with recall at or above the target plus a safety margin, checking that precision stays above 0.60.
  4. Evaluate once on the test set at that threshold and report recall with a confidence interval; with about 40 positive test cases, one miss moves recall by roughly 2.5 points.
  5. Document the threshold with the model, and monitor recall after deployment.
A filter df[df.ratio == 0.3] returns no rows even though you can see 0.3 in the data. Why?

Floating-point values that display as 0.3 are often 0.30000000000000004 or similar after arithmetic, so exact equality fails. Use a tolerance, df[np.isclose(df.ratio, 0.3)], or round deliberately before comparing, or store exact quantities as integers (cents rather than currency units) or decimals. Also note that NaN == NaN is False; use isna().

A teammate asks you to hand over your trained model for a batch-scoring job. What do you deliver?

A single artifact containing the whole fitted pipeline (preprocessing plus model), the decision threshold, the expected input schema (column names, dtypes, allowed categories), and metadata: library versions, training data snapshot and date, metrics with confidence ranges, and the random seed. Add a small scoring function or command-line script, a sample input with expected outputs as a regression test, and a pinned environment file. For PyTorch, deliver the state_dict plus model code and config, or an exported format such as ONNX.