AI Foundations

NLP Fundamentals

Natural language processing is the craft of turning messy human text into numbers a machine can learn from, and turning model outputs back into useful language. This page walks the full path from text cleaning and TF-IDF through word embeddings, subword tokenizers, language models and sequence models to evaluation metrics, a complete eight-model sentiment case study, and how classic NLP tasks live on inside LLM systems.

~130 min read 0 interview questions
In 30 seconds
  • Every NLP system is a pipeline: text → tokens → numbers (vectors) → model → prediction or generated text. Most bugs live in the first two arrows.
  • Sparse representations (one-hot, bag-of-words, TF-IDF) are cheap, interpretable and still very strong baselines; they ignore order and meaning similarity.
  • Dense embeddings (word2vec, GloVe, FastText) put words with similar contexts close together; contextual embeddings (ELMo, BERT) give a word a different vector in every sentence.
  • Modern models use subword tokenizers (BPE, WordPiece, Unigram) so that no word is ever truly out of vocabulary.
  • RNN → LSTM → seq2seq + attention → Transformer is the story of fixing long-range memory and enabling parallel training.
  • Pick the metric for the task: F1 for classification and NER, BLEU for translation, ROUGE for summaries, perplexity for language models, BERTScore or LLM-as-judge for paraphrase-heavy outputs.
  • In the eight-model IMDB sentiment study, linear SVM and logistic regression on binary bag-of-words (F1 0.875) beat every neural model; baselines matter.

What NLP is and why language is hard

Natural Language Processing (NLP) is the field that builds systems which read, understand, transform or generate human language. It covers Natural Language Understanding (NLU: classify, extract, answer) and Natural Language Generation (NLG: summarize, translate, converse). Search engines, spam filters, autocomplete, voice assistants, machine translation and every chatbot you have used are NLP systems.

The core problem is simple to state: computers do arithmetic on numbers, but language arrives as a stream of characters. Everything in NLP is a strategy for building a numeric representation of text that keeps as much meaning as possible while staying cheap enough to compute.

Analogy

Think of a foreign tourist trying to follow a fast-talking local. First they split the speech into words they can look up (tokenization), then they check a phrasebook for what each word means (embeddings), then they use the surrounding sentence to guess which meaning is intended (context modeling), and finally they respond. The tourist is the NLP pipeline: the phrasebook is the vocabulary and embedding table, and the ability to use surrounding words is what sequence models and attention provide.

Why language is hard for machines

Ambiguity

"I saw her duck" (a bird or a movement?). "Bank" (river or money?). One surface form, many meanings. Lexical, syntactic and referential ambiguity all appear in normal text.

Compositionality and order

"Dog bites man" and "man bites dog" contain the same words. "Not bad" is positive; "bad" is negative. Meaning depends on structure, not just on the word set.

Long-range dependencies

"The keys that the man who lives next door left on the table are missing." The verb agrees with a word seven tokens earlier.

Variation and noise

Typos, slang, emojis, hashtags, code-mixing ("delivery bahut late thi"), transcription errors from speech. Real text is nothing like textbook text.

World knowledge and pragmatics

Sarcasm ("Great, another delay"), idioms, implied intent ("It is cold in here" = please close the window). Understanding requires knowledge outside the text.

Zipf's law and sparsity

A handful of words are extremely frequent; most words are rare. Any vocabulary leaves a long tail of rare or unseen words, so models must generalize to things they barely saw.

The four eras of NLP

EraMain ideaTypical methodsWeakness
Rules (1950s–1980s)Hand-written grammars and pattern rulesRegex, parsers, expert systemsBrittle, does not scale, endless edge cases
Statistical (1990s–2012)Learn probabilities from countsn-gram LMs, Naive Bayes, HMM, CRF, TF-IDF + SVMHeavy feature engineering, no notion of word similarity
Neural (2013–2017)Learn dense features end to endword2vec, GloVe, RNN, LSTM, GRU, seq2seq + attentionSequential, slow to train, limited memory
Pretrained Transformers (2018–now)Pretrain once on huge corpora, adapt many timesBERT, GPT, T5, Llama; fine-tuning, prompting, RAGCompute cost, hallucination, opaque reasoning

The canonical NLP pipeline

raw text
   │  normalize (Unicode, case, whitespace, HTML)
   ▼
clean text
   │  tokenize (words / subwords / characters)
   ▼
tokens ──▶ ids (vocabulary lookup)
   │  represent (one-hot, TF-IDF, embeddings, contextual encoder)
   ▼
vectors / matrix  [seq_len × dim]
   │  model (linear classifier, CRF, LSTM, Transformer)
   ▼
output (label, tag per token, span, generated text)
   │  post-process + evaluate (F1, BLEU, ROUGE, perplexity)
   ▼
product decision
Interview angle A common opener is "walk me through how you would build a text classifier." Strong candidates narrate the pipeline above, start with a TF-IDF + logistic regression baseline, name the metric up front (usually macro-F1 if classes are imbalanced), then justify when they would move to a fine-tuned encoder or an LLM. Jumping straight to "use GPT" without a baseline or metric is a red flag.
Common pitfall Treating NLP as "just call a model." Most production failures come from the data side: encoding bugs, tokenization mismatch between training and serving, label noise, and domain shift (training on movie reviews, serving on tweets).

Text preprocessing

Preprocessing converts raw strings into a consistent form so that the same meaning maps to the same tokens. How much you do depends on the model: classical models (bag-of-words, TF-IDF) benefit from aggressive cleaning; modern pretrained Transformers want the text nearly raw and come with their own tokenizer that must be used exactly as trained.

Analogy

Preprocessing is like washing and chopping vegetables before cooking. A stew (bag-of-words model) is fine with everything diced into uniform cubes. A sushi chef (a pretrained Transformer) wants the fish exactly as it was prepared during training; over-chopping ruins it. Mapping back: the cleaning steps are the knife work, and the model decides how much knife work is helpful.

Normalization

  • Unicode normalization (NFC/NFKC): make visually identical characters identical bytes, e.g. "é" as one code point vs "e" + combining accent; full-width digits to ASCII.
  • Case folding: lowercase to merge "Apple"/"apple". Loses information (Apple the company vs apple the fruit), so NER and cased models often keep case.
  • Noise removal: strip HTML tags (IMDB reviews contain <br />), URLs, markup; or replace them with placeholders like <URL>, <NUM>, <USER>.
  • Contractions and spelling: "don't" → "do not", "gr8" → "great". Useful for noisy social text.
  • Emoji and emoticon handling: often carry sentiment; map to words (":)" → "smile") instead of deleting.

Tokenization

Tokenization splits text into units (tokens) that become vocabulary entries. Three granularities exist:

LevelExample for "unhappiness!"ProsCons
Word["unhappiness", "!"]Intuitive, short sequencesHuge vocabulary, out-of-vocabulary (OOV) words, poor for morphology-rich languages
Character["u","n","h",...,"!"]Tiny vocabulary, no OOVVery long sequences, each unit carries little meaning
Subword["un", "happi", "ness", "!"]Balanced vocabulary (30k–200k), no true OOV, shares morphologyTokens may not align with linguistic units; tokenizer must match the model

Word tokenization is not just split(): punctuation ("end." vs "end"), contractions ("can't"), hyphens ("state-of-the-art"), numbers ("3.14"), and languages without spaces (Chinese, Japanese, Thai) all need rules or learned segmenters. Sentence segmentation ("Dr. Smith arrived. He sat.") is its own small problem.

Stop words

Stop words are very frequent function words ("the", "is", "of", "and") that carry little topical content. Removing them shrinks bag-of-words features and helps topic models and keyword search. But be careful: "not", "no", "never" are in many default stop lists, and deleting them destroys sentiment ("not good" becomes "good"). Modern neural models should not have stop words removed at all.

Stemming vs lemmatization

Stemming

  • Rule-based suffix chopping (Porter, Snowball, Lancaster).
  • "studies" → "studi", "running" → "run", "better" → "better".
  • Fast, no dictionary, language-specific rules.
  • Output may not be a real word; can over-stem ("university", "universe" → "univers") or under-stem.

Lemmatization

  • Dictionary + morphology lookup to the canonical form (lemma).
  • "studies" → "study", "ran" → "run", "better" (adjective) → "good".
  • Needs the part of speech for accuracy ("meeting" noun vs verb).
  • Slower, but outputs real words; better for interpretable features.
Analogy

Stemming is a hedge trimmer: fast, cuts every branch to the same height, sometimes lops off something useful. Lemmatization is a gardener who knows each plant species and prunes it correctly. Mapping back: the trimmer is the suffix-stripping rule set, and the gardener's plant knowledge is the dictionary plus part-of-speech information.

Other preprocessing steps

  • Part-of-speech tags as features (needed for accurate lemmatization and for rule-based extraction).
  • Negation handling for classical sentiment: prefix tokens after "not" until punctuation, e.g. "not good movie" → "not NOT_good NOT_movie".
  • Deduplication and language identification before training; near-duplicate reviews inflate test scores.
  • PII masking: replace emails, phone numbers, account numbers with typed placeholders before text leaves a secure boundary.

Python: a classical preprocessing pipeline

import re, unicodedata
import nltk
from nltk.corpus import stopwords
from nltk.stem import PorterStemmer, WordNetLemmatizer
# nltk.download("stopwords"); nltk.download("wordnet"); nltk.download("punkt")

STOP = set(stopwords.words("english")) - {"not", "no", "nor", "never"}  # keep negations!
stem = PorterStemmer()
lemma = WordNetLemmatizer()

def clean(text: str) -> str:
    text = unicodedata.normalize("NFKC", text)
    text = re.sub(r"<[^>]+>", " ", text)          # strip HTML like <br />
    text = re.sub(r"https?://\S+", " <URL> ", text)
    text = text.lower()
    text = re.sub(r"[^a-z0-9<>'\s]", " ", text)
    return re.sub(r"\s+", " ", text).strip()

def tokens(text: str, mode="lemma"):
    toks = [t for t in nltk.word_tokenize(clean(text)) if t not in STOP]
    if mode == "stem":
        return [stem.stem(t) for t in toks]
    return [lemma.lemmatize(t) for t in toks]

print(tokens("The movies were NOT bad<br /><br />Studies show... https://x.y"))
# ['movie', 'not', 'bad', 'study', 'show', 'url']
Tip Fit every preprocessing decision inside the training pipeline (for example, a scikit-learn Pipeline) so the exact same function runs at serving time. Train/serve skew in preprocessing is one of the most common silent failures in NLP.
Common pitfall Running classic cleaning (lowercasing, stop-word removal, stemming) before feeding a pretrained BERT or GPT model. These models were trained on raw text with their own tokenizer; stripping casing, punctuation and function words removes signal they rely on and usually lowers accuracy.
Interview angle Expect "stemming vs lemmatization, and when would you use neither?" The complete answer: stemming is fast and crude, lemmatization is accurate and needs part of speech, and with subword-tokenized Transformers you use neither because the tokenizer and the model already handle morphology. Bonus points for mentioning that stop-word lists remove negations.

Text representation: one-hot, bag-of-words, n-grams, TF-IDF

Once text is tokenized, each document must become a fixed-length numeric vector (for classical models) or a sequence of vectors (for neural models). The simplest family of representations is sparse: vectors as long as the vocabulary, mostly zeros.

Analogy

Picture a library with 50,000 books and a checklist for each visitor: "did they borrow this book, yes or no?" Two visitors who borrowed mostly the same books look identical, even if one loved a book and the other returned it in disgust, and the list never records the order in which books were borrowed. Mapping back: each book is a vocabulary word, the checklist is a one-hot/binary bag-of-words vector, and the missing order and feelings are exactly what these representations throw away.

One-hot word vectors

Give each word in a vocabulary of size V an index. The word's vector has length V with a single 1 at its index. With V = 8: cat = [1,0,0,0,0,0,0,0], king = [0,0,1,0,0,0,0,0], queen = [0,0,0,1,0,0,0,0].

cos(ei, ej) = (ei · ej) / (‖ei‖ ‖ej‖) = 0 / (1 · 1) = 0  for all i ≠ j Two distinct one-hot vectors never share a non-zero position, so their dot product is 0. Every word is equally unrelated to every other word: "cat" is as far from "kitten" as from "algebra".

Problems: no notion of similarity, vectors as long as the vocabulary (100k+ dimensions, 99.99% zeros), and no representation for words unseen in training (OOV).

Bag-of-words (BoW)

A document vector is the sum (or binary OR) of its word one-hots: position j holds the count of word j (count BoW) or 1 if present (binary BoW). Order is discarded, so "dog bites man" equals "man bites dog".

Docthecatsatonmatdoglog
D1 "the cat sat on the mat"2111100
D2 "the dog sat on the log"2011011

n-grams

An n-gram is a run of n consecutive tokens. Adding bigrams ("not bad", "very good") and trigrams restores some local order and negation, at the cost of a much larger, sparser feature space. Character n-grams (e.g. 3–5 characters) are robust to typos and morphology and work well for language ID and noisy text.

TF-IDF

TF-IDF (term frequency × inverse document frequency) reweights counts so that words frequent in this document but rare across the corpus get high scores, and words that appear everywhere ("the") get low scores.

tfidf(t, d) = tf(t, d) × idf(t),    idf(t) = log( N / df(t) ) tf(t, d) = count of term t in document d (often divided by document length or log-scaled as 1 + log tf); N = number of documents; df(t) = number of documents containing t. scikit-learn's default uses a smoothed idf = ln((1 + N)/(1 + df)) + 1 and then L2-normalizes each row.

Worked TF-IDF example

Corpus of N = 3 documents: D1 "the cat sat on the mat", D2 "the dog sat on the log", D3 "cats and dogs". Use tf = raw count / document length and idf = ln(N / df) (natural log).

Termdfidf = ln(3/df)tf in D1tf-idf in D1
the2ln 1.5 = 0.4052/6 = 0.3330.135
cat1ln 3 = 1.0991/6 = 0.1670.183
sat20.4050.1670.068
on20.4050.1670.068
mat11.0990.1670.183

Reading the table: "cat" and "mat" appear once each but get the highest weight because they are unique to D1. "the" appears twice but is discounted because it is in two of three documents. A word present in every document gets idf = ln 1 = 0 and vanishes entirely. Notice also that "cat" and "cats" are different features unless you stem or lemmatize.

Measuring similarity between text vectors

MetricFormulaUse when
Cosine similaritya·b / (‖a‖‖b‖)Default for TF-IDF and embeddings: compares direction, ignores length, so a 100-word and a 1000-word document on the same topic still score high.
Dot productΣ aibiEqual to cosine when vectors are unit-normalized; used by many embedding models and vector databases for speed.
Euclidean (L2)√Σ(ai−bi)2Magnitude matters; on normalized vectors it is a monotone function of cosine: ‖a−b‖2 = 2 − 2cos.
Manhattan (L1)Σ|ai−bi|Robust to single large differences; occasionally used on sparse features.
Jaccard|A ∩ B| / |A ∪ B|Sets of tokens or shingles; near-duplicate detection (with MinHash).
Edit (Levenshtein) distancemin insertions/deletions/substitutionsSpelling correction, fuzzy entity matching ("Jon" vs "John").

Python: BoW, n-grams, TF-IDF and cosine search

from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

docs = ["the cat sat on the mat",
        "the dog sat on the log",
        "diffusion models generate images"]

bow = CountVectorizer(ngram_range=(1, 2))        # unigrams + bigrams
Xb = bow.fit_transform(docs)
print(Xb.shape, bow.get_feature_names_out()[:6])

tfidf = TfidfVectorizer(sublinear_tf=True, min_df=1)
X = tfidf.fit_transform(docs)                   # sparse (3 docs x vocab), rows L2-normalised
q = tfidf.transform(["cat on a mat"])
print(cosine_similarity(q, X).round(2))         # doc 0 scores highest, doc 2 scores 0

BM25: TF-IDF's search-engine cousin

BM25 is the ranking function behind most keyword search. It improves TF-IDF with term-frequency saturation (the 10th occurrence of a word adds less than the 2nd) and document-length normalization. Defaults k1 ≈ 1.2–1.5 and b = 0.75. It remains the lexical half of hybrid retrieval in RAG systems (see RAG).

BM25(q, d) = Σt∈q idf(t) · tf(t,d)(k1+1) / ( tf(t,d) + k1(1 − b + b·|d|/avgdl) ) |d| is document length and avgdl the average length; b controls length normalization, k1 controls how fast term frequency saturates.
Tip A TF-IDF (word 1–2 grams, sublinear tf, min_df 2–5) + logistic regression or linear SVM model trains in seconds, is explainable through its per-word weights, and is shockingly hard to beat on topic and sentiment classification. Always build it first.
Common pitfall Fitting the vectorizer (vocabulary and idf) on the full dataset before the train/test split. The idf statistics then leak test-set information. Fit on train only, then transform validation and test.
Interview angle Interviewers often ask you to compute TF-IDF by hand for two or three tiny documents, or ask "why does a word that appears in every document get zero weight?" Also expect "cosine vs Euclidean for text": cosine is length-invariant, and on L2-normalized vectors the two give the same ranking.

Word embeddings: word2vec, GloVe, FastText

A word embedding is a dense, learned vector (typically 100–300 dimensions for classic methods, 768–4096 inside modern Transformers) in which geometry encodes meaning: similar words point in similar directions. Compared with one-hot, a 300-dimensional embedding is roughly 170× more compact than a 50,000-dimensional one-hot vector, and every dimension is active.

Analogy

An embedding is a GPS coordinate for meaning. Latitude and longitude place a city on a map so nearby cities have nearby numbers; an embedding places a word in a 300-dimensional "meaning map" so "cat" lands near "dog" and "kitten" but far from "car". Mapping back: each coordinate axis is an embedding dimension, the distance between cities is cosine distance, and the map itself is learned from how words are used.

The distributional hypothesis

"You shall know a word by the company it keeps." Words that occur in similar contexts tend to have similar meanings. "King" and "queen" both appear near "throne", "rules", "kingdom", "royal", so any method that predicts context will be pushed to give them similar vectors. No dictionary, thesaurus, grammar or spelling information is used; the signal is purely co-occurrence.

word2vec

word2vec trains a shallow network with one hidden (projection) layer on a fake prediction task. The learned weight matrix, not the predictions, is the product: row i of the input matrix is the embedding of word i. Multiplying a one-hot vector by this matrix simply selects that row (an embedding lookup).

CBOW (Continuous Bag of Words)

  • Input: the surrounding context words (averaged). Output: predict the center word.
  • "the cat ___ on the" → "sat".
  • Faster; smooths over context; good for frequent words.
  • Order within the window is ignored (hence "bag").

Skip-gram

  • Input: the center word. Output: predict each context word.
  • "sat" → "the", "cat", "on", "the".
  • Slower; creates more training pairs per word, so rare words get more updates; better for rare words and small corpora.
  • The default choice in most tutorials.
Skip-gram objective: maximize (1/T) Σt Σ−c≤j≤c, j≠0 log P(wt+j | wt),    P(o | c) = exp(uo·vc) / Σw∈V exp(uw·vc) vc is the "input" vector of center word c, uo the "output" vector of context word o, c the window size. The softmax denominator sums over the whole vocabulary, which is too expensive, so it is replaced by negative sampling or hierarchical softmax.

Negative sampling

Instead of a V-way softmax, turn each (center, context) pair into a binary classification: real pair (label 1) vs k random "negative" words drawn from the corpus (label 0). k = 5–20 for small data, 2–5 for large data. Negatives are sampled from the unigram distribution raised to the 3/4 power, which boosts rare words slightly.

log σ(uo·vc) + Σi=1..k log σ(−uni·vc),    ni ~ P(w) ∝ count(w)3/4 σ is the sigmoid. The model pulls true context vectors toward the center vector and pushes a few random words away; cost per update drops from O(V) to O(k).

Other word2vec tricks: subsampling frequent words (randomly drop "the", "a" with probability growing with frequency), window size as a hyperparameter (small windows of 2–5 capture syntactic similarity such as "walk"~"run"; large windows of 5–10+ capture topical similarity such as "walk"~"park"), and dimension 100–300 chosen empirically as a balance of expressiveness and cost.

Analogy

Negative sampling is like learning faces at a party by being told "these two people came together" plus "these five random guests did not." You never compare everyone with everyone; a few contrasting examples are enough. Mapping back: the couple is the true (center, context) pair, the random guests are the sampled negatives, and skipping the full guest list is skipping the full-vocabulary softmax.

GloVe

GloVe (Global Vectors) builds a global word-word co-occurrence matrix X, where Xij counts how often word j appears in the context of word i, then fits vectors so their dot product approximates the log co-occurrence. It combines the global statistics of matrix-factorization methods (like LSA) with the linear-structure benefits of word2vec.

J = Σi,j f(Xij) ( wi·w̃j + bi + b̃j − log Xij )2 f is a weighting function that caps the influence of very frequent pairs and ignores zero counts. The final embedding is usually w + w̃.

FastText: subword embeddings

FastText represents each word as the sum of vectors for its character n-grams (typically 3–6 characters) plus the whole word. "where" with n = 3 becomes <wh, whe, her, ere, re> plus <where>, with < and > marking boundaries. Benefits:

  • OOV words still get a vector from their n-grams ("unfriendliness" shares pieces with "friendly", "unhappiness").
  • Morphologically rich languages (Finnish, Turkish, Hindi) and typos are handled much better.
  • Pretrained 300-dimensional vectors exist for 150+ languages, which is why the IMDB case study below used them.

Embedding arithmetic

vec(king) − vec(man) + vec(woman) ≈ vec(queen)     vec(Paris) − vec(France) + vec(Germany) ≈ vec(Berlin) Consistent relationships (gender, capital-of, verb tense) show up as roughly constant offset directions. The answer is found by nearest-neighbour search with cosine similarity, excluding the three input words (otherwise "king" itself is often closest).

Typical cosine similarities in a good embedding space: cat–kitten ≈ 0.9, king–queen ≈ 0.85, happy–joyful ≈ 0.85, cat–algebra ≈ 0.1. Note that antonyms are often close ("hot"/"cold", "good"/"bad") because they occur in identical contexts; embeddings capture relatedness, not truth or polarity.

Static vs contextual embeddings

word2vec, GloVe and FastText are static: one vector per word type regardless of sentence. "Bank" in "withdraw cash at the bank" and "the boat drifted to the river bank" gets the identical vector, an average of all its senses. Training with CBOW does not change this; the context is used during training, but the output is still one vector per word.

Contextual embeddings compute a vector per token occurrence from the whole sentence. ELMo (2018) used a deep bidirectional LSTM language model and combined its layers; the two "bank" vectors come out clearly different (cosine around 0.5). BERT and every modern Transformer do the same with self-attention, and are now the default. The cost: an encoder forward pass per sentence instead of a table lookup.

Static (word2vec, GloVe, FastText)

  • Lookup table, microseconds per word.
  • One sense per word; polysemy collapsed.
  • Tiny memory, runs anywhere.
  • Good features for small models and quick similarity.

Contextual (ELMo, BERT, LLM hidden states)

  • Needs a neural forward pass.
  • Different vector for each sense and context.
  • Much stronger downstream accuracy.
  • Standard for any serious NLP system today.

Python: training and querying word2vec / FastText

from gensim.models import Word2Vec, FastText

sentences = [["the", "king", "rules", "the", "kingdom"],
             ["the", "queen", "rules", "the", "kingdom"],
             ["a", "man", "walks"], ["a", "woman", "walks"]]   # use millions of sentences in practice

w2v = Word2Vec(sentences, vector_size=100, window=5, min_count=1,
               sg=1,            # 1 = skip-gram, 0 = CBOW
               negative=5,      # negative samples
               sample=1e-3,     # subsample frequent words
               epochs=50)
print(w2v.wv.most_similar("king", topn=3))
print(w2v.wv.most_similar(positive=["king", "woman"], negative=["man"], topn=1))

ft = FastText(sentences, vector_size=100, window=5, min_count=1, min_n=3, max_n=6)
print(ft.wv["kingdoms"][:5])     # OOV word still gets a vector from its n-grams
# w2v.wv["kingdoms"] would raise KeyError: word2vec has no subword fallback

Using pretrained embeddings in a neural model

import torch, torch.nn as nn
import numpy as np

def build_embedding(vocab, kv, dim=300, freeze=False):
    """vocab: dict word->id; kv: gensim KeyedVectors (e.g. FastText)."""
    W = np.random.normal(0, 0.1, (len(vocab), dim)).astype("float32")
    W[0] = 0.0                                      # id 0 = <pad>
    hits = 0
    for w, i in vocab.items():
        if w in kv:
            W[i] = kv[w]; hits += 1
    print(f"coverage: {hits / len(vocab):.1%}")      # report how many words were found
    return nn.Embedding.from_pretrained(torch.tensor(W), freeze=freeze, padding_idx=0)
Common pitfall Saying that word2vec "understands" meaning or that king–queen similarity comes from spelling or grammar. It comes only from shared contexts. Also, embeddings absorb corpus biases (for example occupation–gender associations), which then flow into downstream decisions; audit before using them in hiring, lending or moderation.
Interview angle Classic probes: "CBOW vs skip-gram, which is better for rare words and why?" (skip-gram, because each occurrence of a rare center word produces several training pairs), "why negative sampling?" (avoid the O(V) softmax), "how does FastText handle OOV?" (character n-grams), and "how does a word get different vectors in different sentences?" (contextual encoders such as ELMo or BERT). Be ready to explain that one-hot times the embedding matrix is just a row lookup.

Sentence and document embeddings

Many tasks (semantic search, clustering, deduplication, RAG retrieval) need one vector for a whole sentence, paragraph or document so that texts with similar meaning score high on cosine similarity even when they share no words: "The cat sits on the mat" vs "A kitten rests on the rug" should be close; "Quarterly revenue grew 12%" should be far.

Analogy

Averaging word vectors is like describing a movie by blending the colour of every frame into one swatch: you get the overall mood but lose the plot. A trained sentence encoder is like a film critic writing a one-line summary: it decides what matters. Mapping back: the frames are token vectors, the blended swatch is mean pooling of static vectors, and the critic is a model such as SBERT trained specifically to make summaries of similar texts land close together.

From simple to strong

  1. Average (mean) pooling of static vectors Average the word2vec/FastText vectors of all tokens. Surprisingly decent baseline; ignores order and is dominated by frequent words.
  2. TF-IDF-weighted average / SIF Weight each word vector by its idf (or by a/(a + p(w))), then remove the common principal component. Reduces the influence of "the", "is".
  3. Doc2Vec (paragraph vectors) Learn a vector per document jointly with word vectors. Needs inference-time training for new documents; mostly historical now.
  4. Raw BERT [CLS] or mean of BERT token outputs Contextual but not trained for similarity; raw BERT sentence vectors are often worse than averaged GloVe for semantic similarity because the space is anisotropic (all vectors crowd into a narrow cone).
  5. SBERT (Sentence-BERT) A siamese/bi-encoder: the same BERT encodes each sentence, mean pooling gives the sentence vector, and the network is fine-tuned on sentence pairs (natural language inference, paraphrase) so that cosine similarity reflects meaning.
  6. Modern embedding models Contrastively trained on hundreds of millions of (query, relevant passage) pairs with in-batch negatives and hard negatives: E5, BGE, GTE, Nomic, Jina, Cohere and OpenAI embedding APIs. Often support instructions/prefixes ("query: ", "passage: "), long inputs (8k+ tokens), multilingual text and Matryoshka truncation (use the first 256 of 1024 dimensions with a small quality loss).
InfoNCE loss: L = − log [ exp(sim(q, p+)/τ) / ( exp(sim(q, p+)/τ) + Σj exp(sim(q, pj−)/τ) ) ] q is the query embedding, p+ its positive passage, p− the negatives (other passages in the batch plus mined hard negatives), τ a temperature (around 0.01–0.1). The loss pulls matching pairs together and pushes mismatches apart.

Bi-encoder vs cross-encoder

Bi-encoder (SBERT, embedding APIs)

  • Encodes query and document separately into vectors.
  • Documents can be pre-embedded and indexed (vector database, approximate nearest neighbour search).
  • Millisecond search over millions of documents.
  • Less precise: no token-level interaction between query and document.

Cross-encoder (re-ranker)

  • Feeds "[query] [SEP] [document]" jointly into one Transformer and outputs a relevance score.
  • Full attention between query and document tokens, so much more accurate.
  • Cannot pre-compute; one forward pass per pair.
  • Used to re-rank the top 20–100 bi-encoder results.

Python: sentence embeddings and semantic search

from sentence_transformers import SentenceTransformer, CrossEncoder, util

bi = SentenceTransformer("all-MiniLM-L6-v2")          # 384-dim, fast
docs = ["The cat sits on the mat.",
        "A kitten rests on the rug.",
        "Quarterly revenue grew 12%."]
D = bi.encode(docs, normalize_embeddings=True)
q = bi.encode("where is the small cat sleeping?", normalize_embeddings=True)
scores = util.cos_sim(q, D)[0]
print(scores)                                         # kitten/cat high, revenue near 0

top = scores.argsort(descending=True)[:2]
ce = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
print(ce.predict([("where is the small cat sleeping?", docs[i]) for i in top]))
Tip Choose embedding models by evaluating on your data (recall@k on a few hundred labelled query–document pairs), not only on public leaderboards. Always use the same model, same prefix convention and same normalization for queries and documents; re-embed the whole corpus when you switch models, because vectors from different models are not comparable.
Common pitfall Using an embedding model with a 512-token limit on 5,000-token documents: the text beyond the limit is silently truncated and never influences the vector. Chunk long documents first.
Interview angle "Why not just use BERT's [CLS] vector for similarity?" is a favourite. Answer: it was never trained for similarity (next-sentence prediction and masked language modeling objectives), the space is anisotropic, and SBERT-style contrastive fine-tuning fixes this. Follow-ups: bi-encoder vs cross-encoder trade-off, and what hard negatives are.

Subword tokenization: BPE, WordPiece, Unigram, SentencePiece

Word-level vocabularies hit a wall: they either explode in size or map rare and new words to a single [UNK] token. Character-level models avoid OOV but produce very long sequences. Subword tokenization is the compromise used by essentially every modern language model: frequent words stay whole ("the", "movie"), rare words split into reusable pieces ("tokenization" → "token" + "ization").

Analogy

Subword tokenization is like LEGO. A toy company could mould a unique plastic piece for every possible object (word vocabulary), or ship only 1×1 bricks (characters). Instead it makes a few thousand common shapes that combine into anything, with big pre-built pieces for things people build all the time. Mapping back: the brick catalogue is the tokenizer vocabulary, the common big pieces are frequent whole words, and building a new object from standard bricks is encoding an unseen word from known subwords.

Byte-Pair Encoding (BPE)

BPE starts from characters and repeatedly merges the most frequent adjacent pair into a new symbol, recording each merge rule in order, until the vocabulary reaches a target size (for example 32k, 50k, 100k or 200k). At inference, a new word is split into characters and the learned merges are replayed in the same order. Used by GPT-2/3/4 (byte-level BPE via tiktoken), Llama 3, RoBERTa, and many others.

Worked BPE example

Training corpus word counts: low: 5, lower: 2, newest: 6, widest: 3. Append an end-of-word marker _ so suffixes stay distinct.

Start (characters):
  l o w _        ×5
  l o w e r _    ×2
  n e w e s t _  ×6
  w i d e s t _  ×3

Count adjacent pairs (weighted by word frequency):
  e s : 6 + 3 = 9     s t : 6 + 3 = 9     t _ : 9
  l o : 5 + 2 = 7     o w : 7             w e : 2 + 6 = 8
  n e : 6   e w : 6   w _ : 5   e r : 2   r _ : 2   w i : 3  i d : 3  d e : 3

Merge 1: (e, s) → "es"    (tie with s t / t _ broken by first seen)
  n e w es t _ ×6      w i d es t _ ×3
Merge 2: (es, t) → "est"  count 9
  n e w est _ ×6       w i d est _ ×3
Merge 3: (est, _) → "est_" count 9
Merge 4: (l, o) → "lo"    count 7
Merge 5: (lo, w) → "low"  count 7
  low _ ×5   low e r _ ×2
Merge 6: (n, e) → "ne"    count 6 ... then (ne, w) → "new", (new, est_) → "newest_"

Vocabulary now: base characters + {es, est, est_, lo, low, ne, new, newest_, ...}

Encode unseen word "lowest":  l o w e s t _
  apply merges in order → low est_     → tokens: ["low", "est_"]

The unseen word "lowest" is represented perfectly by two meaningful pieces, which is the whole point. Vocabulary size is the main hyperparameter: bigger vocabularies produce shorter sequences (cheaper attention, more text per context window) but larger embedding and output matrices and more rarely-trained tokens.

Byte-level BPE

Instead of starting from Unicode characters (there are over 150,000), byte-level BPE starts from the 256 possible bytes of UTF-8. Every string, including emoji, any script and binary junk, can be encoded, so there is literally no unknown token. GPT-2 onward use this. The cost: characters from non-Latin scripts take 2–4 bytes and often end up as more tokens.

WordPiece

Used by BERT, DistilBERT, ELECTRA. Same bottom-up merging idea, but a pair is chosen by how much it increases training-data likelihood, approximately count(ab) / (count(a) × count(b)), rather than by raw count. That prefers merging pieces that are rarely seen apart. Continuation pieces are marked with ##: "playing" → ["play", "##ing"]. At inference it uses greedy longest-match-first from the left. Unknown characters become [UNK].

Unigram language model tokenization

Works top-down: start with a large candidate vocabulary, fit a unigram language model where each token has a probability, and iteratively remove the tokens whose removal hurts corpus likelihood the least, until the target size is reached. Encoding picks the most probable segmentation (Viterbi), and training can sample alternative segmentations (subword regularization), which makes models more robust to noise. Used by T5, ALBERT, XLNet, mBART, many multilingual models.

SentencePiece

SentencePiece is a library and a design choice rather than a separate algorithm: it trains BPE or Unigram directly on raw text, treating the space as a normal symbol (shown as ▁, "lower one-eighth block"). No language-specific pre-tokenizer is needed, which is essential for Chinese, Japanese and Thai, and decoding is exactly reversible. Llama 1/2, T5, mT5, XLM-R and Gemma use SentencePiece.

AlgorithmDirectionSelection criterionMarker styleUsed by
BPEBottom-up mergesMost frequent pairġ (GPT-2 space) or end-of-wordGPT family, RoBERTa, Llama 3
WordPieceBottom-up mergesPair that most increases likelihood## continuationBERT, DistilBERT
Unigram LMTop-down pruningLeast likelihood loss when removed▁ word startT5, ALBERT, XLNet
SentencePieceWrapper for BPE/UnigramRaw text, whitespace as a symbol▁Llama 1/2, mT5, XLM-R

Which to pick in an interview: BPE is deterministic (replay merges in order) and is the GPT-family default. Unigram is probabilistic: Viterbi picks the most likely segmentation and training can sample alternatives (subword regularization), which helps noisy and multilingual text (T5, ALBERT). WordPiece is BERT's cousin of BPE that scores merges by likelihood gain rather than raw count. SentencePiece is the library that trains BPE or Unigram on raw bytes/text with space as a symbol.

Special tokens and tokenizer outputs

  • BERT-style: [CLS] at the start (its output is used for classification), [SEP] between/after segments, [PAD], [MASK], [UNK].
  • GPT-style: <|endoftext|>; chat models add role markers such as <|im_start|> or [INST]. A chat template must be applied exactly as in training.
  • A Hugging Face tokenizer returns input_ids, attention_mask (1 = real token, 0 = padding) and, for pair tasks, token_type_ids. Use return_offsets_mapping=True to map tokens back to character spans (essential for NER and extraction).

Python: seeing tokenizers in action

import tiktoken
from transformers import AutoTokenizer

enc = tiktoken.get_encoding("cl100k_base")               # GPT-3.5/4 tokenizer
ids = enc.encode("Retrieval-augmented generation rocks!")
print(len(ids), [enc.decode([i]) for i in ids])
# 7 ['Retrieval', '-', 'augmented', ' generation', ' rocks', '!'] (illustrative)

bert = AutoTokenizer.from_pretrained("bert-base-uncased")
out = bert("Tokenization is unbelievably useful", return_offsets_mapping=True)
print(bert.convert_ids_to_tokens(out["input_ids"]))
# ['[CLS]', 'token', '##ization', 'is', 'unbelievably', 'useful', '[SEP]']
print(out["offset_mapping"])                            # (start, end) char spans per token

# Training your own BPE tokenizer
from tokenizers import Tokenizer, models, trainers, pre_tokenizers
tok = Tokenizer(models.BPE(unk_token="[UNK]"))
tok.pre_tokenizer = pre_tokenizers.ByteLevel()
trainer = trainers.BpeTrainer(vocab_size=8000, special_tokens=["[UNK]", "[PAD]"])
tok.train(files=["corpus.txt"], trainer=trainer)

Why tokenization matters for LLM behaviour and cost

  • Cost and context: APIs bill per token and context windows are measured in tokens. English averages about 1.3 tokens per word (roughly 4 characters per token); Hindi, Tamil, Arabic or Thai text can take 2–5× more tokens for the same content, so they cost more and fit less in the window.
  • Arithmetic and spelling: numbers split unpredictably ("12345" might become "123" + "45"), and the model never sees individual letters of a multi-character token, which is why "count the r's in strawberry" was hard.
  • Leading spaces matter: " hello" and "hello" are different tokens in GPT tokenizers; prompt formatting changes the token sequence.
  • Glitch tokens: tokens present in the vocabulary but almost absent from training data can produce bizarre behaviour.
  • Domain vocabulary: code, chemistry, or legal text can fragment heavily; adding domain tokens requires resizing and training the embedding matrix.
Common pitfall Mixing tokenizers: using a different tokenizer (or a different version, casing setting or chat template) at inference than the model was trained with. The token ids then point to the wrong embedding rows and output quality collapses without any error message.
Common pitfall Claiming subword tokenization "solves" OOV. It removes the unknown token problem, but a rare word split into many tiny pieces still has a weak, poorly-trained representation. It bridges sparse and dense models very effectively; it does not grant understanding of never-seen concepts.
Interview angle Expect to run BPE merges by hand on a toy corpus, to contrast BPE, WordPiece and Unigram in one sentence each, and to explain why non-English users pay more per API call. Senior-level follow-ups: effect of vocabulary size on sequence length and embedding parameters, why byte-level BPE has no UNK, and how to extend a tokenizer for a new domain or language.

Language modeling basics: n-gram LMs and perplexity

A language model (LM) assigns a probability to a sequence of tokens, or equivalently predicts the next token given the previous ones. It is the engine behind autocomplete, speech recognition rescoring, spelling correction and, scaled up enormously, every LLM (see Transformers & LLMs).

P(w1, …, wT) = ∏t=1..T P(wt | w1, …, wt−1) The chain rule of probability: any joint probability over a sequence factorizes exactly into next-token predictions. Models differ only in how they approximate each conditional.
Analogy

An n-gram model is a phone keyboard's predictive text that only looks at your last one or two words: after "see you", it suggests "soon" because that is what usually came next in the past. Mapping back: the last n−1 words are the context window, "usually came next" is a count ratio from the training corpus, and the keyboard's inability to remember what you said three messages ago is the Markov assumption.

n-gram language models

The Markov assumption truncates history to the previous n−1 tokens. Probabilities are maximum-likelihood count ratios:

P(wt | wt−1) = count(wt−1, wt) / count(wt−1)    (bigram) Trigram: P(wt | wt−2, wt−1) = count(wt−2 wt−1 wt) / count(wt−2 wt−1). Sentence start and end markers <s> and </s> are added.

Worked example. Corpus: "<s> I like cats </s>", "<s> I like dogs </s>", "<s> you like cats </s>". Then P(I | <s>) = 2/3, P(like | I) = 2/2 = 1, P(cats | like) = 2/3, P(</s> | cats) = 1. So P("I like cats") = 2/3 × 1 × 2/3 × 1 = 4/9 ≈ 0.44. But P("you like dogs") = 1/3 × 1 × 1/3 × 1 = 1/9, and any sentence containing an unseen bigram such as "cats like" gets probability exactly 0.

Smoothing: handling unseen n-grams

  • Add-one (Laplace): (count + 1) / (count(context) + V). Simple, but moves too much probability to unseen events.
  • Add-k: use a small k like 0.01 instead of 1.
  • Backoff and interpolation: if a trigram is unseen, fall back to the bigram, then unigram; or mix them with weights λ.
  • Kneser-Ney: the best classical method; the lower-order distribution uses "in how many different contexts does this word appear" rather than raw frequency ("Francisco" is frequent but almost only after "San").

Limits of n-gram LMs: storage grows explosively with n, they cannot share statistics between similar words ("cat" and "kitten" are unrelated symbols), and they cannot see beyond n−1 tokens. Neural LMs (first feed-forward, then RNN, then Transformer) solve all three by working with embeddings and learned context.

Perplexity

Perplexity (PPL) measures how surprised a model is by held-out text. It is the exponential of the average negative log-likelihood per token (the cross-entropy):

PPL = exp( −(1/T) Σt=1..T log P(wt | w<t) ) = P(w1…wT)−1/T Lower is better. Intuition: a perplexity of k means the model is, on average, as uncertain as if choosing uniformly among k tokens. A uniform model over a vocabulary V has PPL = V.

Example: if a model assigns probabilities 0.5, 0.25, 0.5, 0.25 to four test tokens, the average log-probability is (ln 0.5 + ln 0.25 + ln 0.5 + ln 0.25)/4 = −1.04, so PPL = e1.04 ≈ 2.83: the model is as uncertain as a choice among about 2.8 options per token.

import math, torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("gpt2")
lm = AutoModelForCausalLM.from_pretrained("gpt2").eval()

def perplexity(text):
    ids = tok(text, return_tensors="pt").input_ids
    with torch.no_grad():
        loss = lm(ids, labels=ids).loss        # mean cross-entropy over predicted tokens
    return math.exp(loss.item())

print(perplexity("The cat sat on the mat."))    # low (fluent)
print(perplexity("Mat the on sat cat the."))    # much higher (scrambled)
Common pitfall Comparing perplexities of models with different tokenizers or vocabularies. Perplexity is per token, and a tokenizer that produces more, smaller tokens changes the number. Compare bits-per-character or bits-per-byte instead, or keep the tokenizer fixed. Also, low perplexity does not mean factual, helpful or safe output.
Interview angle Typical questions: "Define perplexity and interpret a value of 20", "Why do n-gram models need smoothing?", "What is the relationship between perplexity and cross-entropy loss?" (PPL = eloss when loss is in nats). Strong answers note that LLM pretraining is exactly next-token prediction with this cross-entropy.

Sequence models: RNN, LSTM, GRU, seq2seq, attention and the move to Transformers

Bag-of-words discards order. Sequence models read tokens in order and maintain a running summary, so "not bad" can differ from "bad". This section is the conceptual bridge from classic NLP to Transformers; the network mechanics are covered in depth on the Deep Learning page.

Recurrent neural networks (RNN)

Analogy

A vanilla RNN reads a review word by word while keeping one sticky note of "what I think so far." Every new word forces a rewrite of the same small note, so after 230 words the opening sentence is barely legible. Mapping back: the sticky note is the hidden state ht, rewriting it is the recurrent update, and the fading of early notes is the vanishing gradient.

ht = tanh(Wh ht−1 + Wx xt + b),    y = softmax(Wy hT) xt is the embedding of token t, ht the hidden state. The same weights are reused at every time step. For classification the final state hT (or a pooled combination of all states) feeds the classifier.

Backpropagation through time (BPTT) multiplies the gradient by the Jacobian ∂ht/∂ht−1 = diag(1 − ht2) Wh once per step. Over hundreds of steps the product shrinks exponentially if its norm is below 1 (vanishing gradient: early words stop influencing learning) or blows up if above 1 (exploding gradient: NaN losses). If each step scales the gradient by about 0.99, after 100 steps 0.99100 ≈ 0.37; with a factor of 0.9, 0.9100 ≈ 0.00003.

LSTM

Analogy

An LSTM replaces the sticky note with a filing cabinet and three valves: one decides which old folders to shred, one decides which new pages to file, one decides what to show your boss right now. You can file away "not" and keep it for the next ten words. Mapping back: the cabinet is the cell state Ct, the valves are the forget, input and output gates, and the boss sees the hidden state ht.

ft = σ(Wf[ht−1, xt] + bf)   it = σ(Wi[ht−1, xt] + bi)   ot = σ(Wo[ht−1, xt] + bo) C̃t = tanh(Wc[ht−1, xt] + bc)    Ct = ft ⊙ Ct−1 + it ⊙ C̃t    ht = ot ⊙ tanh(Ct) ⊙ is element-wise multiplication. Because Ct is updated additively, ∂Ct/∂Ct−1 = ft; when the forget gate stays near 1 the gradient flows back almost unchanged (the "constant error carousel"). Initializing bf to 1 helps early training.

Tiny worked example (scalar, all weights 0.5, biases 0, ht−1 = 0.5, xt = 1.0, Ct−1 = 0.3): every gate input is 0.5 × 1.5 = 0.75, so f = i = o = σ(0.75) = 0.679; C̃ = tanh(0.75) = 0.635; Ct = 0.679 × 0.3 + 0.679 × 0.635 = 0.204 + 0.431 = 0.635; ht = 0.679 × tanh(0.635) = 0.679 × 0.561 = 0.381.

A good example of the forget gate firing: in "Alice finished her talk. Bob then started his", the gender/subject information about Alice should be dropped when a new subject appears.

GRU

The Gated Recurrent Unit merges the cell and hidden state and uses two gates: an update gate z (how much of the old state to keep, like forget + input tied together) and a reset gate r (how much of the past to use when proposing the new state). About 25% fewer parameters than an LSTM, often similar accuracy, faster to train.

zt = σ(Wz[ht−1, xt])   rt = σ(Wr[ht−1, xt])   h̃t = tanh(W[rt ⊙ ht−1, xt])   ht = (1 − zt) ⊙ ht−1 + zt ⊙ h̃t

Bidirectional and stacked RNNs

A bidirectional RNN/LSTM runs one pass left-to-right and one right-to-left and concatenates the states, so each position sees both past and future context. Ideal for tagging (NER, POS) and classification where the full input is available; impossible for left-to-right generation. Stacking 2–3 layers adds abstraction. The combination (BiLSTM + CRF) was the state of the art for NER before BERT.

Sequence-to-sequence (encoder–decoder)

For tasks where input and output lengths differ (translation, summarization, question answering), an encoder RNN reads the source into a context vector and a decoder RNN generates the target token by token, feeding back its own previous output. During training the decoder receives the true previous token (teacher forcing); at inference it receives its own prediction, which causes exposure bias (errors compound).

The bottleneck: the entire source sentence must be squeezed into one fixed-size vector. Translation quality dropped sharply for sentences longer than about 20–30 words.

Attention

Analogy

Translating a sentence without attention is like memorizing a whole paragraph and then reciting it in another language with the book closed. Attention lets the translator keep the book open and glance back at the relevant source words for each word they write. Mapping back: the open book is the full set of encoder states, the glance is the attention weight distribution, and what the translator reads is the context vector.

At each decoding step, the decoder state (the query) is compared with every encoder state (the keys); a softmax turns the scores into weights; the context is the weighted sum of encoder states (the values). Bahdanau (additive) attention scores with a small MLP; Luong (multiplicative) attention uses a dot product. Transformers use scaled dot-product attention:

Attention(Q, K, V) = softmax( Q KT / √dk ) V Q: what each position is looking for; K: what each position offers; V: the content passed on. Dividing by √dk keeps dot products from growing with dimension, so softmax does not saturate into one-hot and gradients stay healthy. Each row of the weight matrix sums to 1.

Self-attention applies the same idea within one sequence: every token attends to every other token, so "it" in "The animal didn't cross the street because it was tired" can attend directly to "animal". Cross-attention lets decoder tokens attend to encoder tokens (the bridge between source and target in translation).

Why Transformers replaced RNNs

Limitation of RNNsHow the Transformer addresses it
Sequential computation, cannot parallelize across time stepsAll positions processed at once with matrix multiplications; massive GPU parallelism
Vanishing/exploding gradients limit effective memoryAny two tokens are connected by a path of length 1 through attention; residual connections + layer norm stabilize depth
Fixed-size hidden state bottleneckEvery token keeps its own representation; nothing is compressed into one vector
Long-range dependencies fadeDirect attention to any earlier token regardless of distance

The trade-offs: attention is O(n2) in sequence length, and because self-attention is permutation-invariant, word order must be injected with positional encodings (sinusoidal, learned, or rotary). Three families emerged: encoder-only (BERT: understanding, classification, NER, embeddings), decoder-only (GPT, Llama: generation, chat), encoder-decoder (T5, BART: translation, summarization).

Python: an LSTM text classifier with the fixes that matter

import torch, torch.nn as nn
from torch.nn.utils.rnn import pack_padded_sequence

class BiLSTMClassifier(nn.Module):
    def __init__(self, vocab_size, emb_dim=300, hidden=128, n_classes=2, emb=None):
        super().__init__()
        self.emb = emb or nn.Embedding(vocab_size, emb_dim, padding_idx=0)
        self.lstm = nn.LSTM(emb_dim, hidden, num_layers=1, batch_first=True,
                            bidirectional=True)
        self.drop = nn.Dropout(0.3)
        self.fc = nn.Linear(2 * hidden, n_classes)

    def forward(self, ids, lengths):
        x = self.drop(self.emb(ids))
        packed = pack_padded_sequence(x, lengths.cpu(), batch_first=True,
                                      enforce_sorted=False)     # ignore padding
        _, (h, _) = self.lstm(packed)
        h = torch.cat([h[-2], h[-1]], dim=1)                     # last fwd + last bwd state
        return self.fc(self.drop(h))                             # raw logits

model = BiLSTMClassifier(vocab_size=20000)
opt = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=1e-2)
loss_fn = nn.CrossEntropyLoss()

def train_step(ids, lengths, y):
    model.train(); opt.zero_grad()
    loss = loss_fn(model(ids, lengths), y)
    loss.backward()
    torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)      # essential for RNNs
    opt.step()
    return loss.item()
Common pitfall Training an RNN on 400-token padded sequences without gradient clipping, without packing (so the "final" hidden state is computed after hundreds of padding tokens), and with a learning rate tuned for a CNN. Loss stays flat near ln 2 ≈ 0.693 for binary classification, meaning the model is guessing. The IMDB case study below shows exactly this failure.
Interview angle Be ready to: derive why gradients vanish in a vanilla RNN, explain how the LSTM cell state's additive update fixes it, compare LSTM and GRU, explain teacher forcing and exposure bias, and list the four RNN limitations that motivated Transformers. A common trick question: "Can you use a bidirectional LSTM for text generation?" (No, the backward pass would need future tokens.)

Classic NLP tasks

Almost every NLP problem reduces to one of a small number of output shapes. Recognizing the shape tells you the model head, the loss and the metric.

Output shapeTasksTypical headMetric
One label per textClassification, sentiment, intent, spam, toxicityLinear/softmax on pooled vectorAccuracy, macro-F1, AUC
One label per tokenNER, POS tagging, chunking, slot fillingToken classifier, often + CRFEntity-level F1, token accuracy
Span in the inputExtractive QA, keyphrase extractionStart/end position logitsExact match, token F1
StructureDependency/constituency parsing, relation extractionPairwise scorer, graph decoderUAS/LAS, relation F1
New textTranslation, summarization, generative QA, dialogueDecoder / seq2seq LMBLEU, ROUGE, BERTScore, human or LLM judge
Latent groupingTopic modeling, clusteringLDA, NMF, embeddings + k-means/HDBSCANCoherence, human inspection

Text classification and sentiment analysis

Analogy

A logistic-regression sentiment model is a balance scale: words like "masterpiece", "brilliant" and "loved" are weights on the positive pan, "boring", "waste" and "terrible" on the negative pan, and the heavier side wins. Mapping back: each word's learned coefficient is its weight, the sum over words present in the review is the logit, and the tipping point is the decision threshold.

Sentiment analysis classifies opinion polarity (positive/negative/neutral) or intensity (1–5 stars). Variants:

  • Document-level: one label per review.
  • Sentence-level: useful when reviews mix opinions.
  • Aspect-based sentiment analysis (ABSA): extract (aspect, opinion, polarity) triples. "Battery life is great but the camera is blurry" → (battery life, great, positive), (camera, blurry, negative). Needed whenever sentiment must drive action.
  • Emotion detection: anger, joy, fear, sadness; stance detection: for/against a target.

Hard cases: negation ("not bad"), contrast ("I expected to hate it, but..."), sarcasm ("Oh great, it broke on day one"), comparative sentences, domain shift ("unpredictable" is positive for a thriller plot, negative for a car's brakes). Approaches in order of power: lexicons (VADER, SentiWordNet) → TF-IDF + linear model → LSTM/CNN with pretrained embeddings → fine-tuned encoder (BERT, RoBERTa, DeBERTa) → zero/few-shot LLM prompting.

Named Entity Recognition (NER)

NER finds spans that name things and assigns a type: PERSON, ORG, LOC, DATE, MONEY, plus domain types (DRUG, SYMPTOM, ACCOUNT_ID, CLAUSE_TYPE). It is framed as token classification with a tagging scheme:

Token:   Maria   Lopez   visited  New    York   on  Monday
BIO:     B-PER   I-PER   O        B-LOC  I-LOC  O   B-DATE
BIOES:   B-PER   E-PER   O        B-LOC  E-LOC  O   S-DATE
(B = begin, I = inside, O = outside, E = end, S = single-token entity)

Models: CRF on hand features → BiLSTM-CRF → fine-tuned BERT token classifier (the default) → LLM extraction with a JSON schema for zero-shot or rare types. The CRF layer learns valid tag transitions (I-PER cannot follow B-LOC). With subword tokenizers, label only the first subword of each word (set others to -100 so the loss ignores them) and use offset mappings to recover character spans.

Encoder-only models are the natural fit for NER: each token needs full bidirectional context, the output is aligned one-to-one with the input, and inference is cheap. A decoder-only LLM can do it via prompting, but costs more and needs span grounding.

Part-of-speech (POS) tagging

Assign grammatical categories (noun, verb, adjective...; Universal Dependencies has 17 tags, Penn Treebank 45). Classic models: Hidden Markov Model with Viterbi decoding (emission P(word | tag) × transition P(tag | previous tag)), then CRFs, then BiLSTMs, now transformers at around 97–98% accuracy for English. Uses: lemmatization, rule-based extraction ("adjective + noun" phrases for aspect mining), feature engineering, linguistic analysis.

Parsing

  • Constituency parsing builds a phrase-structure tree (S → NP VP) from a grammar.
  • Dependency parsing links each word to its head with a labelled relation (nsubj, dobj, amod). "The camera takes blurry photos": takes → nsubj → camera, takes → dobj → photos, photos → amod → blurry. Great for extracting who-did-what-to-whom and linking opinion words to aspects. Measured by UAS (unlabelled attachment score) and LAS (labelled).
  • Other structure tasks: coreference resolution (linking "she" to "Dr. Rao"), relation extraction (founded_by(Company, Person)), semantic role labeling (agent, patient, instrument).

Topic modeling (LDA)

Analogy

LDA imagines every document was written by a cook drawing from several spice jars: first choose a recipe mix (40% "delivery", 60% "billing"), then for each word pick a jar according to the mix and draw a word from that jar. Given only the finished dishes, LDA works backwards to guess what is in each jar and what mix each dish used. Mapping back: jars are topics (distributions over words), the recipe mix is the document's topic distribution, and working backwards is Bayesian inference.

Latent Dirichlet Allocation is a generative probabilistic model: each document has a topic mixture θd ~ Dirichlet(α), each topic k has a word distribution φk ~ Dirichlet(β), and each word is generated by picking a topic z ~ θd then a word w ~ φz. Inference by Gibbs sampling or variational Bayes. You choose the number of topics K; evaluate with topic coherence (for example Cv or NPMI) and human inspection. Alternatives: NMF on TF-IDF (fast, often cleaner topics), and BERTopic (sentence embeddings → UMAP → HDBSCAN clusters → class-based TF-IDF to name clusters), which handles short text like tweets and chat messages far better than LDA.

from sklearn.feature_extraction.text import CountVectorizer
from sklearn.decomposition import LatentDirichletAllocation

cv = CountVectorizer(stop_words="english", min_df=5, max_df=0.5)
X = cv.fit_transform(complaints)                      # LDA wants raw counts, not TF-IDF
lda = LatentDirichletAllocation(n_components=8, random_state=0).fit(X)
vocab = cv.get_feature_names_out()
for k, comp in enumerate(lda.components_):
    print(k, [vocab[i] for i in comp.argsort()[-8:][::-1]])
doc_topics = lda.transform(X)                         # per-document topic mixture

Summarization

  • Extractive: select the most important sentences verbatim (TextRank: a PageRank graph over sentence similarities; or a classifier per sentence). Faithful by construction, can be choppy.
  • Abstractive: generate new sentences (BART, PEGASUS, T5, LLMs). Fluent and concise, but can hallucinate facts not in the source. Check faithfulness with entailment models or QA-based checks, not just ROUGE.
  • Long documents: chunk and summarize hierarchically (map-reduce), or use long-context models; watch for "lost in the middle" effects.

Machine translation

Evolution: rule-based → phrase-based statistical MT (alignment models plus an n-gram LM) → RNN seq2seq with attention → Transformer encoder–decoder (the original 2017 Transformer was built for translation) → massively multilingual models (NLLB, M2M-100, mBART) and LLMs. Decoding uses beam search (keep the top-B partial translations at each step, typical B = 4–5) with length normalization. Evaluated with BLEU, chrF, COMET and human judgement.

Question answering

  • Extractive QA (SQuAD-style): given question + passage, predict start and end token of the answer span with an encoder. Metrics: exact match and token-level F1.
  • Open-domain QA: retrieve relevant passages first (BM25 or dense retrieval), then read. This retriever–reader pattern is the ancestor of RAG.
  • Generative / closed-book QA: an LLM answers from parametric memory; fluent but hallucination-prone, hence retrieval grounding (see RAG).
  • Related: natural language inference (entailment / contradiction / neutral), which powers zero-shot classification and faithfulness checking.

Other tasks worth naming

Intent detection + slot filling

Voice assistants and chatbots: "set the temperature to 22" → intent=set_climate, slots {temp: 22}. Joint models classify the utterance and tag tokens at once.

Entity linking

Map a mention ("Apple", "Jaguar") to a unique entry in a knowledge base or ontology (a company ID, an ICD code). Needs candidate generation plus disambiguation using context.

Keyword/keyphrase extraction

TF-IDF top terms, RAKE, YAKE, KeyBERT (embedding similarity between document and candidate phrases).

Text similarity and deduplication

Embeddings + cosine for paraphrase; MinHash + Jaccard on shingles for near-duplicates at web scale.

Spelling and grammar correction

Edit distance + language model scoring classically; seq2seq or LLM rewriting today.

Speech-adjacent NLP

Automatic speech recognition output is lowercase, unpunctuated and error-prone; add punctuation restoration and design for noisy input.

Python: modern task heads with Hugging Face pipelines

from transformers import pipeline

clf = pipeline("sentiment-analysis")                      # fine-tuned encoder
print(clf("The plot was predictable, but the acting was superb."))

ner = pipeline("ner", model="dslim/bert-base-NER", aggregation_strategy="simple")
print(ner("Priya joined Infosys in Bengaluru in March 2021."))
# [{'entity_group': 'PER', 'word': 'Priya', ...}, {'entity_group': 'ORG', ...}, {'entity_group': 'LOC', ...}]

zs = pipeline("zero-shot-classification", model="facebook/bart-large-mnli")
print(zs("My parcel arrived two weeks late.",
         candidate_labels=["delivery", "product quality", "billing"]))

qa = pipeline("question-answering")
print(qa(question="Where did Priya join?", context="Priya joined Infosys in Bengaluru."))

summ = pipeline("summarization", model="facebook/bart-large-cnn")
print(summ(long_article, max_length=80, min_length=30)[0]["summary_text"])
Common pitfall Evaluating NER with token-level accuracy. Most tokens are "O", so a model that predicts "O" everywhere scores 85–90% accuracy while finding zero entities. Use entity-level (exact span + type) precision, recall and F1, for example with seqeval.
Interview angle Interviewers like to name a task and ask you to identify the output shape, architecture, loss and metric. For example: NER → token classification, encoder + optional CRF, cross-entropy per token, entity-level F1. Also common: "extractive vs abstractive summarization trade-offs", "how does LDA work and how do you choose K?", and "why is aspect-based sentiment more useful than document sentiment for a business?"

Evaluating NLP systems

Choosing the metric is a design decision, not an afterthought. Classification and extraction have a single correct answer, so precision/recall-style metrics apply. Generation has many acceptable answers, so metrics compare against references (BLEU, ROUGE), use embeddings (BERTScore), or use judges (humans, LLMs).

Analogy

Grading a translation with BLEU is like grading an essay by counting how many phrases match the teacher's model answer: cheap and consistent, but it punishes a brilliant essay that uses different words. BERTScore is a teacher who accepts synonyms; an LLM judge or human grader reads for meaning and quality. Mapping back: phrase matches are n-gram overlaps, synonym tolerance is embedding similarity, and reading for meaning is rubric-based judgement.

Classification and extraction metrics

Precision = TP / (TP + FP)    Recall = TP / (TP + FN)    F1 = 2 · P · R / (P + R) Precision: of what I flagged, how much was right. Recall: of what was there, how much I found. F1 is the harmonic mean, so it is low if either is low.
  • Accuracy is fine only for balanced classes (IMDB is 50/50, so random guessing scores 50%).
  • Macro-F1 averages per-class F1 equally (protects minority classes); micro-F1 pools all decisions (dominated by frequent classes); weighted-F1 weights by support.
  • ROC-AUC / PR-AUC: threshold-free ranking quality; PR-AUC is more informative when positives are rare (toxicity, fraud).
  • Entity-level F1 for NER: a predicted entity counts only if both span boundaries and type match exactly (strict) or overlap (lenient).
  • Exact match (EM) and token F1 for extractive QA.
  • Cohen's kappa for inter-annotator agreement: if humans agree only 70% of the time, do not expect a model to reach 95%.

BLEU (translation, precision-oriented)

BLEU = BP · exp( Σn=1..4 wn log pn ),    BP = 1 if c > r, else e(1 − r/c) pn = modified (clipped) n-gram precision: each candidate n-gram counts at most as many times as it appears in the reference; wn = 1/4; c = candidate length, r = reference length. The brevity penalty BP stops a model from scoring well by outputting only a few safe words.

Worked example. Reference "the cat is on the mat" (r = 6), candidate "the cat is on mat" (c = 5).

  • Unigrams: the, cat, is, on, mat all appear in the reference → p1 = 5/5 = 1.0.
  • Bigrams: "the cat", "cat is", "is on" match; "on mat" does not → p2 = 3/4 = 0.75.
  • Trigrams: "the cat is", "cat is on" match; "is on mat" does not → p3 = 2/3. 4-grams: "the cat is on" matches, "cat is on mat" does not → p4 = 1/2.
  • BP = e1 − 6/5 = e−0.2 = 0.819.
  • BLEU = 0.819 × (1.0 × 0.75 × 0.667 × 0.5)1/4 = 0.819 × 0.250.25 = 0.819 × 0.707 ≈ 0.58.

Why clipping matters: the candidate "the the the the" would get unclipped p1 = 4/4; clipped, "the" counts at most twice (its count in the reference), giving 2/4. BLEU is designed for corpus-level scoring; sentence-level BLEU is noisy and needs smoothing. Use sacrebleu for reproducible, standardized tokenization.

ROUGE (summarization, recall-oriented)

  • ROUGE-N = overlapping n-grams / n-grams in the reference (recall; precision and F variants are also reported).
  • ROUGE-L uses the longest common subsequence (in-order, not necessarily contiguous), so it rewards correct word order without requiring exact phrase matches.
  • Same example: unigram overlap 5 of the reference's 6 → ROUGE-1 recall = 0.83; LCS("the cat is on the mat", "the cat is on mat") = 5 → ROUGE-L recall = 5/6 = 0.83; ROUGE-2 recall = 3/5 = 0.6.

BLEU asks "how much of my output is correct?"; ROUGE asks "how much of the reference did I cover?", which is the right question for summaries.

METEOR, chrF and learned metrics

  • METEOR: aligns candidate and reference words using exact matches, stems and synonyms (WordNet), computes a recall-weighted F-mean, then applies a fragmentation penalty for scrambled order. Correlates better with humans than BLEU at the sentence level.
  • chrF: character n-gram F-score; robust for morphologically rich languages.
  • BERTScore: embed every token of candidate and reference with a contextual model, greedily match each token to its most similar counterpart by cosine, and compute precision, recall and F1 over those similarities. Gives credit for "large" vs "big". Optional idf weighting and baseline rescaling.
  • COMET, BLEURT: neural models trained to predict human quality ratings; the current standard for machine translation evaluation.
  • LLM-as-a-judge: prompt a strong model with a rubric to score or compare outputs. Scales human-like judgement for open-ended text; check for position bias, verbosity bias and self-preference, and calibrate against human labels.
  • Faithfulness metrics for summaries and RAG answers: NLI-based entailment (does the source entail each claim?), QA-based checks (QAGS, QuestEval), and citation precision.
TaskPrimary metricsWatch out for
Sentiment / classificationAccuracy (balanced), macro-F1, PR-AUCClass imbalance, threshold choice, label noise
NER / extractionEntity-level P/R/F1 (strict)Boundary errors, nested entities
TranslationBLEU, chrF, COMETTokenization of the metric, single reference
SummarizationROUGE-1/2/L, BERTScore, faithfulnessROUGE rewards copying, ignores hallucination
Language modelingPerplexity, bits per byteTokenizer-dependent
Open-ended generation / chatHuman eval, LLM-as-judge, win rateJudge bias, benchmark contamination
RetrievalRecall@k, MRR, nDCGMissing relevance labels
import evaluate
from sklearn.metrics import classification_report
from seqeval.metrics import classification_report as ner_report

print(classification_report(y_true, y_pred, digits=3))           # P/R/F1 per class + macro

print(ner_report([["B-PER", "I-PER", "O"]], [["B-PER", "O", "O"]]))  # entity-level

bleu = evaluate.load("sacrebleu")
print(bleu.compute(predictions=["the cat is on mat"],
                   references=[["the cat is on the mat"]])["score"])
rouge = evaluate.load("rouge")
print(rouge.compute(predictions=["the cat is on mat"], references=["the cat is on the mat"]))
bs = evaluate.load("bertscore")
print(bs.compute(predictions=["a big dog"], references=["a large dog"], lang="en")["f1"])
Common pitfall Reporting a high ROUGE score as proof that a summarizer is good. ROUGE rewards lexical overlap; a summary that copies the reference's phrasing but inverts a number or adds an invented fact can still score high. Pair overlap metrics with faithfulness checks and human review.
Interview angle Expect "BLEU vs ROUGE: which is precision and which is recall, and why?", "what does the brevity penalty do?", "why is BERTScore better for paraphrases?", and "how would you evaluate an LLM summarizer in production?" A strong answer combines automatic metrics, a faithfulness check, a small human-labelled gold set, and online signals (edits, thumbs down, escalation rate).

End-to-end case study: IMDB sentiment with eight models

This case study compares four classical and four neural models on the same binary sentiment task. It is valuable precisely because the result is counter-intuitive: the simple models won. Every lesson here generalizes to real projects.

Analogy

A movie studio hands you 50,000 audience feedback forms, half "loved it" and half "hated it," and asks you to build a machine that sorts new forms. You try a word-counting clerk, a detective, a line-drawer, a panel of judges, and four increasingly sophisticated readers. Mapping back: each clerk or reader is one model, the forms are IMDB reviews, and the sorting accuracy on a fresh stack of forms is the test F1.

The dataset

PropertyValue
Total reviews50,000 (25,000 train, 25,000 test)
Classes2 (positive, negative), exactly 50/50
Average lengthAbout 230 words; sequences truncated/padded to 400 tokens for neural models
NoiseHTML line breaks (<br />), long reviews, sarcasm, plot summaries mixed with opinion

Because classes are balanced, random guessing gives about 50% accuracy and an F1 near 0.5–0.67 depending on which class is guessed. Any model below roughly 0.55 F1 has learned essentially nothing.

Representation for the classical models: binary one-hot bag-of-words

Each review became a binary vector over the vocabulary (up to 50,000 words): 1 if the word appears, 0 otherwise. Frequency, rarity (no idf) and order are all ignored. "The film is not bad" and "the film is bad, not good" produce nearly identical vectors. Despite this, linear models do very well because sentiment is dominated by strong individual words.

Part 1: four classical models

Logistic regression (F1 0.875)

Learns a signed weight per word and sums the weights of words present. Best grid-search setting: C = 0.01 (strong L2 regularization) with max_features = 20,000; 5-fold CV F1 0.8663. Heavy regularization stops it from over-trusting rare words among tens of thousands of features.

Bernoulli Naive Bayes (F1 0.812)

Models P(word present | class) for each word independently and multiplies them. Bernoulli, not Multinomial, because features are binary and Bernoulli also uses "word absent" as evidence. Lowest classical score: recall only 0.763, missing about 24% of positive reviews. The independence assumption is its main weakness.

Linear SVM (F1 0.875, joint best)

Finds the maximum-margin hyperplane, focusing on the hardest reviews near the boundary (support vectors). Best setting: C = 0.01 with max_features = 50,000; CV F1 0.8615. Needed a larger vocabulary than logistic regression to tie it.

Random forest (F1 0.854)

200 trees, max_depth 50, max_features 50,000; CV F1 0.8519. Each tree sees a random subset of features and rows and votes. Slightly behind linear models on sparse text (trees split on one word at a time), but provides feature importances for debugging.

Analogy

Naive Bayes is a detective who scores each clue alone ("footprints: +5 suspect", "alibi: −8 suspect") and never considers that two clues together mean something different. The SVM is someone drawing the widest possible chalk line between red and blue dots on a whiteboard, caring only about the dots nearest the line. The random forest is a panel of 200 judges who each read a different random excerpt and vote. Mapping back: clues are words, the chalk line is the decision boundary with maximum margin, and the judges are decision trees trained on bootstrapped data and feature subsets.

Part 2: four neural models

The motivation: bag-of-words cannot distinguish "not bad" from "bad not." Neural models add dense embeddings, sequential context, bidirectional context and finally attention between all words.

  1. Vanilla RNN (F1 0.568, 3.9M parameters) Reads left to right with one hidden state and classifies from the final state. Training loss barely moved, 0.7006 → 0.6987 over 10 epochs, which is essentially ln 2: the model guessed throughout. Diagnosis: vanishing gradients over 400-step sequences, no gradient clipping, and a final state computed after padding.
  2. LSTM (F1 0.528, 4.6M parameters) Gates should fix vanishing gradients, and training loss did improve (0.693 → 0.569), yet test F1 was even lower than the RNN. Precision 0.694 but recall 0.426: it learned a skewed decision rule that over-predicted one class. Lesson: a falling training loss does not guarantee a good test metric; check the confusion matrix and the threshold.
  3. BiLSTM + FastText embeddings (F1 0.811, 11.6M parameters) Two LSTMs (forward and backward) concatenated, initialized with 300-dimensional pretrained FastText vectors, so "brilliant" and "masterpiece" start near each other. Jumped about 24 F1 points over the RNN. But training loss crashed to 0.048 while precision was 0.764 and recall 0.865: classic overfitting that over-predicted the positive class.
  4. Cross-attention classifier (F1 0.863, 3.8M parameters, best neural) Every token attends directly to every other token with softmax(QKT/√d)V, so "not" can attend to "bad" with no sequential bottleneck and no vanishing gradient. Fewest parameters of the neural models, best neural result, still 1.2 points behind the linear models.
Analogy

The BiLSTM is a book club where one reader goes front-to-back and another back-to-front, and both were handed an encyclopedia (FastText) before starting. The attention model is a party of 400 people where everyone can call out to anyone and decide whom to listen to, instead of whispering along a telephone chain. Mapping back: the two readers are the forward and backward LSTMs, the encyclopedia is pretrained embeddings, the party is self-attention over all token pairs, and the telephone chain is recurrent state passing.

Review: "the film was not as bad as critics claimed"

Forward LSTM  →  the → film → was → not → as → bad ...   (sees "not" before "bad")
Backward LSTM ←  ... bad ← as ← not ← was ← film ← the   (sees "claimed", "critics" first)
Concatenate [h_fwd ; h_bwd] → classifier → POSITIVE

Attention: row for "bad" puts high weight on "not" and "critics claimed"
           → direct 1-hop connection, regardless of distance

Full results: all eight models on the test set

RankModelTypeAccuracyPrecisionRecallF1ParamsStatus
1Linear SVMClassical0.87460.87190.87820.875–Best
1Logistic regressionClassical0.87340.86590.88370.875–Best
3Cross-attentionNeural0.86380.86530.86160.8633.8MBest neural
4Random forestClassical0.85300.84630.86270.854–Good
5Bernoulli Naive BayesClassical0.82340.86810.76260.812–Low recall
6BiLSTM + FastTextNeural0.79880.76370.86540.81111.6MOverfit
7Vanilla RNNNeural0.50720.50560.64760.5683.9MVanishing gradient
8LSTMNeural0.61900.69400.42550.5284.6MUnbalanced predictions

Note that Bernoulli NB (0.812) and BiLSTM + FastText (0.811) are essentially tied on F1 but fail in opposite directions: NB is precise but misses positives; the BiLSTM finds positives but raises many false alarms.

Key takeaways

  1. Classical baselines are hard to beat Even binary bag-of-words with a linear model reached F1 0.875. Never assume neural means better; the baseline sets the bar every fancier model must clear.
  2. Gradient clipping is not optional for RNNs on long text One line, clip_grad_norm_(model.parameters(), 1.0), plus shorter sequences (or packing), would likely have rescued the RNN and LSTM.
  3. Training loss is not test performance The BiLSTM's training loss of 0.048 looked excellent and was a symptom of memorization. Use a validation set, early stopping, dropout and weight decay.
  4. Pretrained embeddings matter enormously FastText initialization plus bidirectionality moved F1 from 0.568 to 0.811.
  5. Attention removes the sequential bottleneck With the fewest parameters, the attention model beat every recurrent model.
  6. Always inspect precision and recall, not just F1 Two models with the same F1 can have opposite failure modes that matter differently to the business.

How to push beyond 0.875

  • TF-IDF with word 1–2 grams (and optionally character n-grams) instead of binary unigrams: typically around 0.89–0.91 on IMDB with a linear model; NB-SVM (Naive Bayes log-count-ratio features fed to an SVM) is a classic strong baseline around 0.91.
  • Fine-tune a pretrained encoder (DistilBERT, RoBERTa, DeBERTa) for 2–3 epochs with learning rate around 2e-5: roughly 0.93–0.96 accuracy. Handle reviews longer than 512 tokens by head+tail truncation or chunk-and-average.
  • Zero-shot or few-shot LLM prompting: strong without training data, but costlier per prediction and needs output parsing.

Python: the classical half end to end

from datasets import load_dataset
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.svm import LinearSVC
from sklearn.naive_bayes import BernoulliNB
from sklearn.pipeline import Pipeline
from sklearn.model_selection import GridSearchCV
from sklearn.metrics import classification_report
import re

ds = load_dataset("imdb")
clean = lambda s: re.sub(r"<br\s*/?>", " ", s)
Xtr, ytr = [clean(t) for t in ds["train"]["text"]], ds["train"]["label"]
Xte, yte = [clean(t) for t in ds["test"]["text"]], ds["test"]["label"]

onehot = lambda n: CountVectorizer(binary=True, max_features=n)   # binary bag-of-words

pipe = Pipeline([("vec", onehot(20000)), ("clf", LogisticRegression(max_iter=2000))])
grid = GridSearchCV(pipe, {"vec__max_features": [20000, 50000],
                           "clf__C": [0.01, 0.1, 1.0]}, cv=5, scoring="f1", n_jobs=-1)
grid.fit(Xtr, ytr)
print(grid.best_params_, round(grid.best_score_, 4))
print(classification_report(yte, grid.predict(Xte), digits=4))

for name, clf in [("svm", LinearSVC(C=0.01)), ("bnb", BernoulliNB())]:
    p = Pipeline([("vec", onehot(50000)), ("clf", clf)]).fit(Xtr, ytr)
    print(name, classification_report(yte, p.predict(Xte), digits=4))

# Stronger baseline: TF-IDF uni+bigrams
strong = Pipeline([("vec", TfidfVectorizer(ngram_range=(1, 2), min_df=3, sublinear_tf=True)),
                   ("clf", LogisticRegression(C=4, max_iter=2000))]).fit(Xtr, ytr)
print(classification_report(yte, strong.predict(Xte), digits=4))

Python: the attention classifier (neural half, compact)

import torch, torch.nn as nn

class AttnClassifier(nn.Module):
    def __init__(self, vocab, d=128, heads=4, max_len=400, n_classes=2):
        super().__init__()
        self.tok = nn.Embedding(vocab, d, padding_idx=0)
        self.pos = nn.Embedding(max_len, d)                      # learned positions
        self.attn = nn.MultiheadAttention(d, heads, batch_first=True, dropout=0.1)
        self.norm = nn.LayerNorm(d)
        self.ff = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))
        self.out = nn.Linear(d, n_classes)

    def forward(self, ids):
        pad = ids.eq(0)                                          # True where padding
        pos = torch.arange(ids.size(1), device=ids.device)
        x = self.tok(ids) + self.pos(pos)
        a, _ = self.attn(x, x, x, key_padding_mask=pad)          # every token looks at every token
        x = self.norm(x + a)
        x = self.norm(x + self.ff(x))
        x = x.masked_fill(pad.unsqueeze(-1), 0).sum(1) / (~pad).sum(1, keepdim=True)
        return self.out(x)                                       # mean-pooled logits
Common pitfall Declaring "classical beats deep learning" as a general law from this experiment. The neural models here were trained from scratch (or with static embeddings) on 25,000 examples and under-tuned. Pretrained Transformers fine-tuned on the same data clearly beat the linear baseline. The real lesson is to build baselines and debug training properly.
Interview angle Interviewers love this story: "Your LSTM got 0.53 F1 and logistic regression got 0.875. What happened and what do you do?" Walk through: check loss curves (flat at 0.693 means no learning), gradient norms, clipping, sequence length and packing, learning rate, class balance of predictions, threshold, then consider pretrained embeddings or a pretrained Transformer. Close by noting that for many production problems the 0.875 linear model may be the right answer because it is cheap and explainable.

Multilingual and code-mixed text

Most of the world's users do not write in clean English. A production NLP system must cope with many languages, scripts, transliteration and switching between languages within a single sentence.

Analogy

A multilingual model is like an interpreter who grew up in a port city: they learned many languages at once and noticed that "water", "agua" and "paani" are used in the same situations, so they file them together. Code-mixed text is that interpreter overhearing a conversation that switches language mid-sentence; they cope because they never kept the languages in separate boxes. Mapping back: the shared filing system is a shared multilingual embedding space, and not keeping separate boxes is training one model and one tokenizer on all languages together.

Challenges

  • Scripts and segmentation: no spaces in Chinese, Japanese, Thai; complex scripts (Devanagari, Arabic right-to-left, Tamil) with combining characters; Unicode normalization becomes critical.
  • Morphology: agglutinative languages (Turkish, Finnish, Tamil) create huge numbers of word forms; subwords or characters are essential.
  • Low-resource languages: little labelled or even unlabelled data.
  • Tokenizer fertility: English-heavy tokenizers split other languages into many more tokens per word, raising cost, shrinking effective context, and hurting quality. This is sometimes called the token tax.
  • Code-mixing: "Delivery bahut late thi, totally disappointed" mixes Hindi and English in one sentence; "Hinglish" is often written in Roman script (transliterated), with non-standard, inconsistent spelling (bahut / bohot / bhot).
  • Culture and pragmatics: politeness forms, sarcasm conventions and sentiment words differ by culture; translated labels may not carry the same intensity.

Approaches

  1. Language identification first (fastText lid.176, CLD3), at the sentence or even token level for code-mixed text.
  2. Normalization and transliteration Normalize spelling variants; optionally transliterate Romanized Hindi to Devanagari (or the reverse) so it matches the model's training data.
  3. Multilingual pretrained encoders mBERT, XLM-RoBERTa, MuRIL / IndicBERT for Indian languages, LaBSE and multilingual E5/BGE for embeddings. They support zero-shot cross-lingual transfer: fine-tune on English labels, apply to Hindi or Spanish with reasonable accuracy.
  4. Translate-train or translate-test Machine-translate training data into the target language, or translate incoming text into English and use an English model. Simple and effective, but translation loses slang and sentiment nuance.
  5. Collect in-domain code-mixed labels Even a few thousand labelled Hinglish examples usually beat any zero-shot transfer. Use LLMs to pre-label and humans to verify.
  6. Multilingual LLMs Prompt directly in the user's language or ask for output in a fixed schema regardless of input language (for example, always English aspect labels).
from transformers import AutoTokenizer
for name in ["bert-base-uncased", "xlm-roberta-base"]:
    t = AutoTokenizer.from_pretrained(name)
    for s in ["The delivery was very late", "Delivery bahut late thi", "डिलीवरी बहुत देर से आई"]:
        print(name, len(t.tokenize(s)), t.tokenize(s)[:8])
# The English-only tokenizer fragments Devanagari into many pieces or [UNK];
# the multilingual SentencePiece tokenizer keeps it compact.
Common pitfall Evaluating a multilingual system only on clean, monolingual test sets. Real traffic is transliterated, code-mixed and misspelled. Build a test set sampled from actual user messages, stratified by language and script, and report metrics per slice.
Interview angle Asked in design rounds for Indian, Southeast Asian or European markets: "Your sentiment model works in English; now 40% of traffic is Hinglish. What do you do?" Cover language ID, transliteration handling, a multilingual encoder, cross-lingual transfer, targeted labelling, per-language metrics, and the extra token cost for LLM-based approaches.

Applying NLP inside GenAI systems

LLMs did not make classic NLP tasks disappear; they changed how the tasks are implemented. Information extraction, sentiment, NER, entity linking, classification and summarization are now frequently done by prompting an LLM with a schema, and the results feed databases, rules engines, dashboards, RAG indexes and agents. The fundamentals on this page are what let you design, evaluate and debug those systems.

Analogy

An LLM is a brilliant new analyst who can read anything but occasionally invents a convincing detail. You do not fire your filing system and auditors when you hire them; you give the analyst a form to fill in, require them to quote the page they read, and have an auditor spot-check the forms. Mapping back: the form is a JSON schema with constrained decoding, the quote is span grounding or citation, and the auditor is validation code plus evaluation against a labelled gold set.

Information extraction with LLMs

  1. Define the schema Entity types, relation types, allowed enum values, required vs optional fields, and a null/unknown option so the model is not forced to invent values.
  2. Prompt with definitions and examples Clear type definitions, 2–5 few-shot examples including hard negatives ("a city mentioned in a signature is not the incident location").
  3. Constrain the output JSON mode, function/tool calling or grammar-constrained decoding so outputs always parse; validate with a schema library (for example Pydantic) and retry on failure.
  4. Ground every value Ask for the exact supporting quote or character offsets for each field and verify the quote exists in the source. Reject or flag ungrounded values; this is the main defence against hallucinated entities.
  5. Normalize and link Map extracted strings to canonical IDs (dates to ISO format, drug names to a medical ontology code, companies to a registry ID).
  6. Evaluate like NER Field-level precision, recall and F1 against a human-labelled gold set; track per-field error rates and drift over time.
  7. Hybridize for cost and reliability Regex and rules for rigid formats (IBANs, phone numbers, IP addresses), a fine-tuned small encoder for high-volume standard entities, and the LLM for long-tail, ambiguous or relational extraction.
from pydantic import BaseModel, Field
from typing import Literal, Optional
import json

class Entity(BaseModel):
    type: Literal["ACCOUNT", "DEVICE", "LOCATION", "AMOUNT", "DATE"]
    value: str
    evidence: str = Field(description="exact substring copied from the note")

class Extraction(BaseModel):
    entities: list[Entity]
    fraud_pattern: Optional[str] = None      # null when the note does not describe one

PROMPT = """Extract entities from the analyst note. Return JSON matching this schema:
{schema}
Rules: copy 'evidence' verbatim from the note; if unsure, omit the entity; never guess.
Note: {note}"""

def extract(note, llm):
    raw = llm(PROMPT.format(schema=Extraction.model_json_schema(), note=note))
    data = Extraction.model_validate(json.loads(raw))
    data.entities = [e for e in data.entities if e.evidence in note]   # grounding check
    return data

Aspect-based sentiment with LLMs

Document sentiment ("3 stars, negative") cannot drive action. Aspect-based sentiment extracts which aspect is being judged and how, so it can be routed:

"Phone is great but it came 5 days late and the box was crushed"
   │  LLM / fine-tuned ABSA model, fixed aspect taxonomy
   ▼
[{aspect: product_quality, polarity: positive, evidence: "Phone is great"},
 {aspect: delivery_delay,  polarity: negative, evidence: "came 5 days late"},
 {aspect: packaging,       polarity: negative, evidence: "box was crushed"}]
   │  aggregate by aspect, SKU, warehouse, time window
   ▼
delivery_delay spike at one warehouse → trigger logistics audit workflow

Design points: a fixed, versioned aspect taxonomy (with an "other" bucket reviewed periodically for new aspects), evidence spans, calibration against human labels, aggregation with volume thresholds before triggering actions, and multilingual/code-mixed support.

Entity linking and normalization

Entity linking maps a mention to a unique knowledge-base entry. "Paris" may be the city, a person or a hotel; "MI" may be myocardial infarction or Michigan. Pipeline: (1) candidate generation by alias dictionaries, fuzzy string matching and embedding search over entity descriptions; (2) disambiguation by comparing the mention's context with each candidate (cross-encoder or LLM choosing from a short list); (3) NIL detection when no candidate fits. Letting an LLM pick from retrieved candidates, instead of generating an ID from memory, prevents invented codes.

Where classic NLP shows up in GenAI stacks

RAG ingestion

Cleaning, sentence segmentation, chunking, metadata extraction (NER for filters), deduplication, BM25 + embeddings for hybrid search. See RAG.

Routing and guardrails

Fast intent classifiers and toxicity/PII detectors decide which model or tool handles a request, before an expensive LLM call.

Agents

Agents read text, extract structured facts, decide actions and write summaries: every step is an NLP task with its own failure modes and metrics.

Evaluation

BLEU/ROUGE/BERTScore, NLI-based faithfulness, LLM-as-judge and classic F1 on extracted fields are how GenAI output quality is measured.

Cost engineering

Distill LLM labels into a small fine-tuned encoder for high-volume classification: often 100× cheaper and lower latency with similar accuracy.

Knowledge graphs

Extract entities and relations from documents into a graph for multi-hop queries, root-cause analysis and GraphRAG.

LLM vs fine-tuned small model vs classical model

FactorClassical (TF-IDF + linear)Fine-tuned encoder (BERT-size)LLM prompting
Labelled data neededThousandsHundreds to thousandsZero to a few examples
Latency / cost per itemMicroseconds, nearly freeMilliseconds, cheap on CPU/GPUHundreds of ms to seconds, per-token billing
Accuracy on narrow, stable taskGoodVery goodGood to very good
New labels / changing schemaRetrainRetrainEdit the prompt
ExplainabilityWord weightsAttribution methodsEvidence quotes, rationales (not guaranteed faithful)
Hallucination riskNone (fixed labels)None (fixed labels)Present; needs grounding and validation
Common pitfall Trusting LLM "confidence" statements or self-reported probabilities as calibrated scores. For confidence you need either token log-probabilities, agreement across several samples (self-consistency), a verifier model, or calibration against labelled data.
Interview angle GenAI system-design rounds increasingly hide classic NLP inside the prompt: "extract entities from notes", "turn reviews into actions", "moderate content". Interviewers want to hear schema design, grounding, validation, evaluation with a gold set and F1, hybrid rule/model/LLM pipelines, cost and latency, and human-in-the-loop for high-risk decisions. See also GenAI System Design.

Quick revision

  • NLP pipeline: normalize → tokenize → map to ids → represent as vectors → model → post-process and evaluate.
  • Language is hard because of ambiguity, word order and composition, long-range dependencies, noise, world knowledge, and a long tail of rare words (Zipf's law).
  • Aggressive cleaning (lowercase, stop words, stemming) helps bag-of-words models and hurts pretrained Transformers.
  • Keep negations ("not", "no", "never") when removing stop words for sentiment tasks.
  • Stemming chops suffixes by rule ("studies" → "studi"); lemmatization maps to a dictionary form using part of speech ("better" → "good").
  • Distinct one-hot vectors always have cosine similarity 0, so one-hot encodes identity but no similarity.
  • Bag-of-words ignores order: "dog bites man" equals "man bites dog"; n-grams restore some local order.
  • TF-IDF = term frequency × log(N / document frequency); a term in every document gets idf 0; fit it on training data only.
  • Cosine similarity compares direction and ignores length; on unit vectors it ranks identically to Euclidean distance and dot product.
  • BM25 adds term-frequency saturation and length normalization to TF-IDF and is the standard lexical retriever.
  • Distributional hypothesis: words in similar contexts have similar meanings; embeddings are learned purely from co-occurrence.
  • CBOW predicts the center word from context (faster, frequent words); skip-gram predicts context from the center word (better for rare words).
  • Negative sampling replaces the full-vocabulary softmax with k binary real-vs-random classifications, sampled from unigram3/4.
  • GloVe fits word vectors to log global co-occurrence counts; FastText sums character n-gram vectors, so it can embed unseen words.
  • Embedding arithmetic works because consistent relations become roughly constant offset directions (king − man + woman ≈ queen).
  • Static embeddings give one vector per word type; contextual models (ELMo, BERT) give one vector per occurrence and resolve polysemy.
  • Antonyms are often close in embedding space because they share contexts.
  • Raw BERT [CLS] vectors are poor for similarity; SBERT and contrastively trained embedding models fix this.
  • Bi-encoders embed separately and scale to millions of documents; cross-encoders score pairs jointly and are used for re-ranking.
  • BPE repeatedly merges the most frequent adjacent pair (deterministic, GPT default); WordPiece merges the pair that most increases likelihood; Unigram prunes a large vocabulary top-down and can sample segmentations.
  • SentencePiece trains on raw text with the space as a symbol, so it needs no language-specific pre-tokenizer.
  • Byte-level BPE starts from 256 bytes and never produces an unknown token.
  • English averages about 1.3 tokens per word; many non-Latin-script languages need several times more tokens, raising cost.
  • Always use the exact tokenizer, version and chat template the model was trained with.
  • A language model factorizes P(sequence) into next-token probabilities by the chain rule.
  • n-gram LMs use the Markov assumption and need smoothing (Laplace, backoff, Kneser-Ney) for unseen n-grams.
  • Perplexity = exp(average negative log-likelihood per token); lower is better; a uniform model over V tokens has perplexity V; compare only across identical tokenizers.
  • Vanilla RNNs suffer vanishing and exploding gradients through backpropagation through time; clip gradients.
  • The LSTM's additive cell-state update with a forget gate near 1 lets gradients flow over long distances.
  • GRU uses update and reset gates, fewer parameters than LSTM, similar accuracy.
  • Bidirectional RNNs suit tagging and classification, not left-to-right generation.
  • Seq2seq squeezes the source into one vector; attention lets the decoder look back at every encoder state.
  • Scaled dot-product attention: softmax(QKT / √dk) V; the scaling prevents softmax saturation.
  • Transformers replaced RNNs because of parallel training, 1-hop paths between any two tokens, and no fixed-size bottleneck; they need positional encodings.
  • NER is token classification with BIO tags; evaluate with entity-level F1, not token accuracy.
  • Aspect-based sentiment extracts (aspect, opinion, polarity) and is what makes sentiment actionable.
  • LDA models documents as mixtures of topics and topics as distributions over words; BERTopic works better for short text.
  • Extractive summarization is faithful but choppy; abstractive is fluent but can hallucinate.
  • BLEU is clipped n-gram precision with a brevity penalty; ROUGE is n-gram or LCS recall.
  • BERTScore matches tokens by contextual embedding similarity and rewards paraphrases; METEOR uses stems and synonyms.
  • In the IMDB study, linear SVM and logistic regression (F1 0.875) beat all neural models; cross-attention was the best neural model (0.863).
  • IMDB lessons: build baselines, clip gradients, do not trust training loss, use pretrained embeddings, and inspect precision and recall separately.
  • Code-mixed text (for example Hinglish) needs language ID, transliteration handling, multilingual encoders and in-domain labels.
  • LLM extraction needs a schema, constrained output, evidence spans verified against the source, normalization and F1 evaluation.
  • Entity linking maps mentions to knowledge-base IDs via candidate generation plus contextual disambiguation, with a NIL option.

Glossary

Aspect-based sentiment analysis (ABSA)
Extracting which aspect of a product or service is discussed and the sentiment toward that aspect, rather than one label per document.
Attention
A mechanism that computes a weighted average of value vectors, with weights from query–key similarity, letting a model focus on relevant tokens.
Attention mask
A 0/1 array telling the model which positions are real tokens and which are padding (or, in causal models, which future positions are hidden).
Backpropagation through time (BPTT)
Training an RNN by unrolling it over time steps and backpropagating through every step.
Bag-of-words (BoW)
A document vector of word counts or presence flags that ignores word order.
Beam search
A decoding strategy that keeps the B most probable partial outputs at each step instead of only the single best.
BERTScore
A generation metric that matches candidate and reference tokens by cosine similarity of contextual embeddings and reports precision, recall and F1.
Bi-encoder
A model that embeds query and document independently so documents can be pre-indexed for fast similarity search.
BIO tagging
A labelling scheme for spans: B- begins an entity, I- continues it, O is outside any entity.
BLEU
A translation metric based on clipped n-gram precision (n = 1–4) combined geometrically and multiplied by a brevity penalty.
BM25
A lexical ranking function extending TF-IDF with term-frequency saturation and document-length normalization.
Byte-Pair Encoding (BPE)
A subword tokenizer that learns a vocabulary by repeatedly merging the most frequent adjacent symbol pair.
CBOW
Continuous bag of words: the word2vec variant that predicts a center word from its averaged context words.
Code-mixing
Switching between two or more languages within a sentence or conversation, often in a single script (for example Romanized Hindi with English).
Conditional random field (CRF)
A probabilistic sequence model that scores whole tag sequences, enforcing valid transitions; often placed on top of BiLSTMs for NER.
Contextual embedding
A token vector computed from the whole sentence, so the same word gets different vectors in different contexts.
Cosine similarity
The dot product of two vectors divided by the product of their lengths; measures the angle between them.
Cross-encoder
A model that reads a query and a document together and outputs a relevance score; accurate but cannot pre-compute document vectors.
Dependency parsing
Predicting a tree where each word points to its grammatical head with a labelled relation such as subject or object.
Distributional hypothesis
The idea that words occurring in similar contexts have similar meanings.
Document frequency (df)
The number of documents in a corpus that contain a given term.
Embedding
A dense, learned vector representing a token, word, sentence or document so that geometric closeness reflects similarity.
Entity linking
Mapping a text mention to a unique entry in a knowledge base or ontology.
Exposure bias
The mismatch between training with true previous tokens (teacher forcing) and inference with the model's own, possibly wrong, predictions.
F1 score
The harmonic mean of precision and recall.
FastText
A word-embedding method that represents words as sums of character n-gram vectors, enabling vectors for unseen words.
GloVe
Global Vectors: embeddings learned by fitting dot products to the logarithm of global word co-occurrence counts.
GRU
Gated recurrent unit: a recurrent cell with update and reset gates, lighter than an LSTM.
Hallucination
Generated content that is fluent but not supported by the source or by facts.
Inverse document frequency (idf)
log(N / df): a weight that is high for rare terms and zero for terms in every document.
Language model
A model that assigns probabilities to token sequences, usually by predicting the next token.
LDA
Latent Dirichlet Allocation: a generative topic model treating documents as mixtures of topics and topics as distributions over words.
Lemmatization
Reducing a word to its dictionary form using vocabulary and part of speech.
LSTM
Long short-term memory: a recurrent cell with forget, input and output gates and an additive cell state that preserves long-range information.
METEOR
A generation metric that aligns words by exact match, stem and synonym, weights recall over precision and penalizes fragmented order.
n-gram
A sequence of n consecutive tokens or characters.
Named entity recognition (NER)
Finding spans that name people, organizations, locations, dates and domain-specific entities, and labelling their types.
Negative sampling
Training embeddings by distinguishing real context words from a few randomly sampled words instead of computing a full softmax.
Normalization
Making text consistent: Unicode forms, casing, whitespace, markup removal, spelling variants.
One-hot vector
A vector as long as the vocabulary with a single 1 at the word's index.
Out-of-vocabulary (OOV)
A word not present in the model's vocabulary.
Perplexity
The exponential of the average per-token negative log-likelihood; the effective number of choices a language model is uncertain between.
Polysemy
One word having several related or unrelated meanings, such as "bank".
POS tagging
Assigning a part-of-speech category (noun, verb, adjective, and so on) to each token.
Positional encoding
Information about token position added to embeddings so an order-blind attention model can use word order.
Precision
The fraction of predicted positives (or extracted items) that are correct.
Recall
The fraction of true positives (or gold items) that the system found.
RNN
Recurrent neural network: processes a sequence one step at a time, updating a hidden state with shared weights.
ROUGE
A family of summarization metrics measuring n-gram (ROUGE-N) or longest-common-subsequence (ROUGE-L) overlap, emphasizing recall.
Self-attention
Attention where queries, keys and values all come from the same sequence, so every token can use every other token as context.
Sentence-BERT (SBERT)
A BERT model fine-tuned in a siamese setup so that pooled sentence vectors are comparable with cosine similarity.
SentencePiece
A tokenizer library that learns BPE or Unigram vocabularies directly from raw text, treating whitespace as an ordinary symbol.
Sentiment analysis
Classifying the opinion or emotional polarity expressed in text.
Seq2seq
An encoder–decoder architecture that maps an input sequence to an output sequence of possibly different length.
Skip-gram
The word2vec variant that predicts surrounding context words from the center word.
Smoothing
Techniques that reserve probability mass for unseen n-grams in count-based language models.
Stemming
Stripping affixes with heuristic rules to reach a crude root form that may not be a real word.
Stop words
Very frequent function words often removed in classical pipelines.
Subword tokenization
Splitting text into units between characters and words so frequent words stay whole and rare words decompose into known pieces.
Teacher forcing
Feeding the true previous output token to a decoder during training rather than its own prediction.
TF-IDF
Term frequency multiplied by inverse document frequency; weights words that are frequent in a document but rare in the corpus.
Token
The basic unit a model processes: a word, subword, character or byte.
Tokenizer fertility
The average number of tokens a tokenizer produces per word; higher fertility means higher cost and shorter effective context.
Topic coherence
A score of how semantically related the top words of a topic are, used to compare topic models.
Transliteration
Writing a language in another script, such as Hindi in Latin letters.
Unigram tokenization
A subword method that starts with a large vocabulary and prunes tokens whose removal least reduces corpus likelihood.
Vanishing gradient
Gradients shrinking exponentially as they propagate back through many steps or layers, preventing early inputs from being learned.
word2vec
A family of shallow neural methods (CBOW and skip-gram) that learn static word embeddings from local context windows.
WordPiece
BERT's subword tokenizer, which merges pairs that most increase training likelihood and marks continuations with ##.
Zero-shot cross-lingual transfer
Fine-tuning a multilingual model on one language and applying it to others without target-language labels.
Zipf's law
In natural language, the r-th most frequent word occurs about 1/r as often as the most frequent word, so most of the vocabulary is rare.

Interview questions

Fundamentals

What is NLP, and what is the difference between NLU and NLG?

NLP is the field of building systems that process human language: reading, understanding, transforming and producing it. NLU (understanding) maps text to structure or labels: classification, NER, parsing, question answering. NLG (generation) produces text: summarization, translation, dialogue. Most real systems combine both, for example a support bot that understands the intent and then generates a reply.

Why is natural language hard for computers?
  • Ambiguity: lexical ("bank"), syntactic ("I saw the man with the telescope"), referential ("it").
  • Order and composition: "not bad" vs "bad"; "dog bites man" vs "man bites dog".
  • Long-range dependencies between distant words.
  • Variation and noise: slang, typos, code-mixing, speech-recognition errors.
  • Pragmatics and world knowledge: sarcasm, idioms, implied intent.
  • Sparsity: Zipf's law means most words are rare, and new words keep appearing.
Describe a typical NLP pipeline from raw text to prediction.

Normalize the text (Unicode, casing, markup), tokenize it, map tokens to ids with a vocabulary, turn ids into vectors (one-hot, TF-IDF, embeddings or a contextual encoder), run a model (linear classifier, CRF, LSTM, Transformer), post-process the output (thresholds, span decoding, formatting) and evaluate with a task-appropriate metric. The same preprocessing and tokenizer must be used in training and serving.

What is tokenization? Compare word, character and subword tokenization.

Tokenization splits text into units that become vocabulary entries. Word-level: intuitive and short sequences, but huge vocabularies and out-of-vocabulary words. Character-level: tiny vocabulary and no OOV, but very long sequences with little meaning per unit. Subword (BPE, WordPiece, Unigram): frequent words stay whole, rare words split into known pieces; a vocabulary of 30k–200k covers any text. Modern models use subwords.

What are stop words, and when should you not remove them?

Very frequent function words ("the", "is", "of") with little topical content. Removing them shrinks bag-of-words features and helps topic modeling and keyword search. Do not remove them for sentiment (negations like "not" are often in stop lists), for tasks depending on syntax (parsing, QA, translation), for phrase queries ("to be or not to be") or when using pretrained Transformers, which rely on full text.

What is the difference between stemming and lemmatization?

Stemming chops affixes with heuristic rules (Porter, Snowball): fast, no dictionary, but the output may not be a word ("studies" → "studi") and it can over- or under-stem. Lemmatization uses a lexicon and the part of speech to return the dictionary form ("studies" → "study", "better" → "good", "ran" → "run"): slower but accurate and readable. Use stemming for quick search recall, lemmatization for interpretable features, and neither with subword Transformers.

What is text normalization? Give examples.

Making equivalent text identical: Unicode normalization (NFC/NFKC), lowercasing, removing HTML tags and URLs (or replacing them with placeholders), collapsing whitespace, expanding contractions, standardizing numbers and dates, fixing common misspellings, and mapping emoticons to words. How far to go depends on the model: aggressive for bag-of-words, minimal for pretrained Transformers.

What is one-hot encoding of words, and what are its limitations?

Each word gets an index in a vocabulary of size V and is represented by a length-V vector with a single 1. Limitations: no similarity (all pairs are orthogonal), extreme dimensionality and sparsity (100k dimensions, 99.99% zeros), no representation for unseen words, and when summed into documents, no word order.

Prove that the cosine similarity between two distinct one-hot vectors is always zero.

Let ei and ej be one-hot vectors with i ≠ j. Their dot product is Σk ei,kej,k. For every k, at least one factor is 0 because ei is non-zero only at k = i and ej only at k = j, and i ≠ j. So the dot product is 0. Both norms are 1, so the cosine is 0 / (1 × 1) = 0. Hence every word is equally dissimilar to every other word.

What is bag-of-words, and what information does it lose?

A document vector where each dimension is a word's count (or presence) in the document. It loses word order, syntax and negation scope ("not good" vs "good, not"), treats synonyms as unrelated dimensions and cannot represent unseen words. It keeps which words appear, which is often enough for topic and sentiment classification.

What are n-grams, and why are they useful?

Sequences of n consecutive tokens (bigrams, trigrams) or characters. Word n-grams capture local order and phrases ("not bad", "New York", "very good"), improving bag-of-words classifiers. Character n-grams handle typos, morphology and language identification. The cost is a far larger, sparser feature space, so use min_df thresholds and regularization.

Explain TF-IDF and why it works better than raw counts.

TF-IDF multiplies how often a term appears in a document (tf) by how rare it is across the corpus, idf = log(N/df). Raw counts are dominated by common words ("the", "movie"); TF-IDF downweights terms that appear everywhere and highlights terms distinctive to each document. It is cheap, interpretable and a strong baseline for classification and retrieval.

Compute the idf of a word that appears in every document. What does that imply?

idf = log(N/N) = log 1 = 0, so its TF-IDF weight is zero in every document: it carries no information for distinguishing documents. Implementations with smoothing (for example scikit-learn's ln((1+N)/(1+df)) + 1) give it a small positive weight instead of zero, but it is still the minimum.

What is cosine similarity, and why is it preferred for text?

cos(a, b) = a·b / (‖a‖‖b‖), the cosine of the angle between vectors, from −1 to 1 (0 to 1 for non-negative TF-IDF). It ignores vector length, so a short and a long document on the same topic are judged similar, whereas Euclidean distance would call them far apart. For unit-normalized vectors, cosine, dot product and Euclidean distance give the same ranking.

What is a word embedding?

A dense, low-dimensional (for example 300-d) learned vector for a word, in which similar words are close in space and some relationships appear as consistent directions. Unlike one-hot vectors, embeddings encode similarity, are compact, and can be pretrained on huge unlabelled corpora and reused.

What is the distributional hypothesis?

"A word is characterized by the company it keeps": words that appear in similar contexts tend to have similar meanings. It is the foundation for word2vec, GloVe and even masked language modeling: predicting context forces similar-context words toward similar representations, without any dictionary or labels.

What is the difference between CBOW and skip-gram in word2vec?

CBOW averages the context word vectors and predicts the center word: faster, smoother, good for frequent words. Skip-gram takes the center word and predicts each context word: slower but better for rare words and small corpora, because each occurrence of a rare word generates several training pairs and therefore several gradient updates to its vector. They are alternative formulations; you pick one.

How does word2vec learn that "king" and "queen" are related?

Purely from co-occurrence statistics: both appear near "throne", "rules", "kingdom", "royal". Because the training objective predicts context, words with similar context distributions receive similar vectors. No thesaurus, spelling, grammar parser or human labels are involved. The analogy king − man + woman ≈ queen emerges because the male–female contrast shows up as a consistent direction.

What is the out-of-vocabulary (OOV) problem, and how is it handled?

A word at inference time that was not in the training vocabulary gets no representation (a zero vector or a generic [UNK]). Solutions: character n-gram embeddings (FastText), subword tokenization (BPE, WordPiece, Unigram), byte-level tokenization (no unknown symbol at all), character-level models, and for static vocabularies, mapping to [UNK] with frequency thresholds. Note that OOV is a vocabulary limitation, not a sign of overfitting.

What is sentiment analysis, and what makes it difficult?

Classifying the opinion polarity or emotion in text. Difficulties: negation and its scope, contrastive clauses ("I expected to hate it, but"), sarcasm and irony, domain-specific polarity ("unpredictable" plot vs brakes), mixed opinions about different aspects, comparative statements, emojis and slang, and code-mixed text.

What is Named Entity Recognition?

Identifying spans of text that name entities and assigning types such as PERSON, ORG, LOCATION, DATE, MONEY, or domain types like DRUG or ACCOUNT_ID. It is framed as token classification with BIO tags, solved with CRFs, BiLSTM-CRFs, fine-tuned BERT or LLM prompting, and evaluated with entity-level F1.

What is part-of-speech tagging, and where is it used?

Labelling each token with its grammatical category (noun, verb, adjective...). Classic approaches are HMMs with Viterbi decoding and CRFs; modern taggers reach about 97–98% accuracy in English. Uses: accurate lemmatization, rule-based phrase extraction (adjective + noun for aspects), features for parsers and NER, and disambiguation ("book a flight" vs "read a book").

What is a language model?

A model that assigns probabilities to sequences, typically by predicting the next token: P(w1..T) = ∏ P(wt | w<t). Examples range from bigram count models to GPT-style Transformers. Uses: autocomplete, speech recognition and translation rescoring, spelling correction, text generation, and pretraining representations for downstream tasks.

What is perplexity?

The exponential of the average negative log-likelihood per token on held-out text. Lower is better. A perplexity of k means the model is as uncertain as if choosing uniformly among k tokens at each step; a uniform model over vocabulary V has perplexity V. It equals ecross-entropy loss when the loss is in nats.

What are precision, recall and F1, and when do you prefer each?

Precision = TP/(TP+FP), recall = TP/(TP+FN), F1 = their harmonic mean. Favour precision when false positives are costly (auto-blocking content, flagging a customer as fraudulent), recall when misses are costly (detecting symptoms, safety issues), and F1 when you need a single balanced number. On imbalanced data, report per-class and macro-averaged values.

What is BLEU?

A translation metric: the geometric mean of clipped n-gram precisions for n = 1–4, multiplied by a brevity penalty that punishes candidates shorter than the reference. It is precision-oriented (how much of the output matches a reference), cheap and reproducible at corpus level, but ignores synonyms and meaning.

What is ROUGE, and how does it differ from BLEU?

ROUGE measures overlap between a generated summary and a reference, emphasizing recall: how much of the reference content is covered. ROUGE-N counts n-gram overlaps; ROUGE-L uses the longest common subsequence to reward in-order matches. BLEU emphasizes precision and has a brevity penalty; ROUGE is used for summarization, BLEU for translation.

What is topic modeling?

Unsupervised discovery of themes in a collection. LDA treats each document as a mixture of topics and each topic as a distribution over words; NMF factorizes a TF-IDF matrix; BERTopic clusters sentence embeddings and names the clusters with class-based TF-IDF. Outputs are top words per topic and a topic mixture per document, used for exploration, trend tracking and routing.

What is the difference between extractive and abstractive summarization?

Extractive selects existing sentences (TextRank, sentence classifiers): faithful and simple but can be redundant or choppy. Abstractive generates new text (BART, T5, LLMs): concise and fluent but can hallucinate facts, numbers or names not in the source. Many production systems generate abstractively and then verify each claim against the source.

What are the main kinds of question answering systems?

Extractive: predict a start and end span in a given passage (SQuAD style, encoder models). Open-domain: retrieve passages from a large corpus, then read (retriever–reader, the ancestor of RAG). Generative/closed-book: an LLM answers from its parameters. Knowledge-base QA: translate the question into a structured query. Metrics: exact match and token F1 for extractive QA, faithfulness and correctness for generative.

Why do we use pretrained embeddings instead of training from scratch?

They encode general word meaning learned from billions of tokens, so the model starts knowing that "brilliant" and "masterpiece" are similar even with small labelled data. This speeds convergence and improves accuracy, especially for rare words. In the IMDB study, adding pretrained FastText vectors and bidirectionality raised F1 from 0.568 to 0.811.

What does an attention mask do in a Transformer tokenizer output?

It marks which positions are real tokens (1) and which are padding (0), so attention ignores padding and pooled representations are not diluted. In decoder models, a separate causal mask prevents each position from attending to future tokens.

Going deeper

Compute TF-IDF for "cat" and "the" in "the cat sat on the mat" given a 3-document corpus where "the" appears in 2 documents and "cat" in 1.

Document length 6. tf(the) = 2/6 = 0.333, tf(cat) = 1/6 = 0.167. idf(the) = ln(3/2) = 0.405, idf(cat) = ln(3/1) = 1.099. TF-IDF(the) = 0.333 × 0.405 = 0.135; TF-IDF(cat) = 0.167 × 1.099 = 0.183. "cat" wins despite occurring half as often, because it is distinctive to this document.

What is negative sampling in word2vec, and why is it needed?

The full skip-gram objective needs a softmax over the entire vocabulary for every training pair, which is O(V) per update. Negative sampling replaces it with binary logistic classification: the true (center, context) pair should score high, and k randomly sampled "negative" words should score low. Negatives are sampled from the unigram distribution raised to 3/4, which slightly boosts rare words. Cost per update drops to O(k) with k = 5–20, and quality stays high. Hierarchical softmax (a binary tree over the vocabulary, O(log V)) is the alternative.

How do you choose the window size and embedding dimension for word2vec?

Both are hyperparameters tuned on downstream or intrinsic evaluation. Small windows (2–5) emphasize syntactic, functional similarity ("walking" near "running"); large windows (5–10+) emphasize topical similarity ("walking" near "park"). Dimensions of 100–300 became standard empirically: large enough to capture relations, small enough to train efficiently and avoid overfitting on limited data. Larger corpora can support larger dimensions.

Word2vec takes one-hot vectors as input. How can it preserve relationships?

The one-hot vector only selects a row of the input weight matrix; the embedding is that learned row. Training adjusts these rows so that words predicting similar contexts end up with similar rows. The relationships live in the learned weights, not in the one-hot inputs, which remain orthogonal.

Compare word2vec, GloVe and FastText.
word2vecGloVeFastText
SignalLocal context windows, predictiveGlobal co-occurrence counts, regression on log countsLocal windows like skip-gram/CBOW
UnitWordWordWord + character n-grams
OOVNo vectorNo vectorBuilt from n-grams
StrengthSimple, efficient, good analogiesUses corpus-wide statisticsMorphology, typos, rare words, many languages

All three are static: one vector per word type.

Why do antonyms like "hot" and "cold" often have high cosine similarity?

They occur in almost identical contexts ("the water is ___", "___ weather"), and distributional methods only see contexts. Embeddings therefore capture relatedness and substitutability rather than polarity or truth. For sentiment tasks, the downstream classifier must learn polarity; specialized sentiment-aware embeddings or fine-tuning help.

How does word2vec represent "bank" in "withdraw cash at the bank" and "sat on the river bank"? How does ELMo differ?

word2vec (either CBOW or skip-gram) gives exactly the same vector in both sentences, a blend of all senses weighted by frequency; context is used only during training. ELMo runs a deep bidirectional LSTM language model over each sentence and combines its layer outputs, producing a different vector for each occurrence based on the surrounding words. The enabling feature is that the embedding is a function of the whole sentence, computed on the fly by a bidirectional recurrent language model. BERT does the same with self-attention.

How would you build a sentence embedding, and why is averaging word vectors not ideal?

Options in increasing quality: average static word vectors; idf- or SIF-weighted averages with the common component removed; mean pooling of a pretrained encoder; a model trained for similarity (SBERT or a contrastively trained embedding model). Plain averaging ignores order and negation ("not good" ≈ "good"), is dominated by frequent words, and dilutes key words in long texts.

Explain bi-encoders vs cross-encoders and how they are combined in search.

A bi-encoder embeds queries and documents separately, allowing documents to be pre-indexed and searched in milliseconds with approximate nearest neighbours, but query and document tokens never interact. A cross-encoder reads the concatenated pair and scores relevance with full attention: far more accurate but one forward pass per pair. Standard design: bi-encoder (plus BM25) retrieves the top 50–100, then a cross-encoder re-ranks them.

Walk through BPE training on a small corpus.

Split every word into characters plus an end marker, weighted by word frequency. Count all adjacent pairs, merge the most frequent into a new symbol, record the merge, and repeat until the target vocabulary size. Example: with "newest" (6) and "widest" (3), the pair (e, s) occurs 9 times, so "es" is created; then (es, t) gives "est"; then (est, _) gives "est_". With "low" (5) and "lower" (2), (l, o) and then (lo, w) produce "low". The unseen word "lowest" then encodes as ["low", "est_"] by replaying the merges in order.

How do BPE, WordPiece and Unigram differ?

BPE: bottom-up, merges the most frequent adjacent pair. WordPiece: bottom-up, merges the pair that most increases training likelihood (roughly count(ab)/(count(a)·count(b))), uses "##" for continuations and greedy longest-match encoding. Unigram LM: top-down, starts with a large vocabulary and removes tokens whose loss least reduces likelihood; encodes with Viterbi and can sample segmentations for regularization. SentencePiece is a library that runs BPE or Unigram on raw text, treating spaces as symbols.

What is byte-level BPE, and why do GPT models use it?

BPE whose base alphabet is the 256 byte values of UTF-8 rather than Unicode characters. Any string (every script, emoji, code, corrupted text) can be encoded, so there is never an unknown token, and the base vocabulary is tiny. The downside is that non-Latin characters take multiple bytes and often more tokens, making those languages more expensive.

How does vocabulary size affect a model?

Larger vocabularies produce shorter sequences (cheaper attention, more content per context window, better for multilingual coverage) but increase embedding and output-softmax parameters, and more tokens are rarely seen in training, so their vectors are undertrained. Smaller vocabularies produce longer sequences and fragment rare words. Typical: 30k for BERT, 50k for GPT-2, 100k–200k+ for recent multilingual LLMs.

Why can LLMs struggle to count letters or do digit arithmetic?

They see tokens, not characters. "strawberry" may be two or three tokens, so the model never directly observes each letter. Numbers are split inconsistently ("1234" vs "12" + "34"), so digit alignment for arithmetic is not explicit. Mitigations: digit-level tokenization of numbers, asking the model to spell out characters first, or delegating arithmetic to a tool.

Why do non-English languages cost more on LLM APIs?

APIs bill per token, and tokenizers trained mostly on English text learn long merges for English words but fragment other scripts into many short pieces (high tokenizer fertility). The same sentence in Hindi or Tamil can take several times more tokens than in English, costing more, fitting less in the context window, and often producing lower quality.

Compute a bigram probability and explain why smoothing is needed.

With the corpus "<s> I like cats </s>", "<s> I like dogs </s>", "<s> you like cats </s>": P(cats | like) = count(like cats)/count(like) = 2/3, and P(I like cats) = 2/3 × 1 × 2/3 × 1 = 4/9. Any unseen bigram ("cats like") has probability 0, making the whole sentence probability 0 and perplexity infinite. Smoothing (add-k, backoff, interpolation, Kneser-Ney) reserves some probability for unseen events.

What is Kneser-Ney smoothing's key idea?

It uses absolute discounting of observed counts and, crucially, a lower-order distribution based on continuation counts: how many different contexts a word appears after, rather than how often it appears. "Francisco" is frequent but almost always follows "San", so it should get low probability in a new context. This makes backoff estimates much better and made Kneser-Ney the best classical n-gram smoother.

Calculate perplexity for a model that assigns probabilities 0.5, 0.25, 0.5, 0.25 to four test tokens.

Average log-probability = (ln 0.5 + ln 0.25 + ln 0.5 + ln 0.25)/4 = (−0.693 − 1.386 − 0.693 − 1.386)/4 = −1.04. Perplexity = e1.04 ≈ 2.83 (equivalently (0.5 × 0.25 × 0.5 × 0.25)−1/4 = √8). The model is about as uncertain as a choice among 2.8 options per token.

Why does a vanilla RNN suffer from vanishing gradients?

In backpropagation through time, the gradient from step T to step t is a product of T−t Jacobians ∂hk/∂hk−1 = diag(1 − hk2) Wh. The tanh derivative is at most 1, and if the largest singular value of Wh is below 1, the product shrinks exponentially (for example 0.9100 ≈ 3 × 10−5). Early tokens then have essentially no influence on the weight updates, so long-range dependencies cannot be learned. If the singular values exceed 1, gradients explode instead.

How does an LSTM mitigate vanishing gradients?

It keeps a separate cell state updated additively: Ct = ft ⊙ Ct−1 + it ⊙ C̃t. The gradient along the cell path is multiplied by the forget gate ft rather than by a weight matrix and squashing derivative. When the network learns f ≈ 1 for information worth keeping, gradients pass back almost unchanged over many steps. Input and output gates control writing and reading. It mitigates but does not eliminate the problem for very long sequences.

Describe each LSTM gate in plain language with a language example.

Forget gate: what to erase from memory, for example drop the previous subject's gender when a new subject appears ("Alice finished. Bob then..."). Input gate: what new information to store, for example record that "not" was seen so the next adjective is negated. Output gate: what part of memory to expose at this step, for example reveal the stored subject number when a verb must agree with it.

Compare LSTM and GRU.

GRU merges cell and hidden states and uses two gates (update, reset) instead of three; it has about 25% fewer parameters and trains faster. LSTM has a separate cell state and output gate, giving slightly more control, which can help on very long or complex sequences. Empirically they perform similarly on most NLP tasks; choose by validation results and compute budget.

What is teacher forcing, and what problem does it create?

During seq2seq training, the decoder is fed the ground-truth previous token instead of its own prediction, which stabilizes and speeds training. At inference, it must consume its own outputs, so an early mistake leads to inputs it never saw in training and errors compound: exposure bias. Mitigations: scheduled sampling, sequence-level training objectives, beam search, and simply larger, better models.

What was the bottleneck in RNN encoder–decoder models, and how did attention solve it?

The encoder compressed the entire source into one fixed-size vector, so long sentences lost information and translation quality dropped with length. Attention lets the decoder, at each step, compute weights over all encoder hidden states and take their weighted sum as a context vector, so it can focus on the relevant source words ("soft alignment") and no longer depends on a single summary vector.

Why divide by √dk in scaled dot-product attention?

If query and key components have unit variance, their dot product has variance dk, so scores grow with dimension. Large scores push softmax into saturation (nearly one-hot), where gradients are tiny. Dividing by √dk restores unit variance, keeping softmax in a sensitive range and training stable.

List the limitations of RNNs that motivated the Transformer.
  • Sequential computation prevents parallelism across time steps, making training slow.
  • Vanishing and exploding gradients limit effective memory.
  • A fixed-size hidden state must compress the whole history.
  • Long-range dependencies fade because information travels through many steps.

Transformers process all tokens in parallel and connect any two tokens directly through self-attention, at the cost of O(n2) attention and needing positional encodings.

Why do Transformers need positional encodings?

Self-attention is permutation-invariant: shuffling input tokens shuffles outputs identically, so without extra information "dog bites man" and "man bites dog" look the same. Positional encodings (sinusoidal, learned absolute, or relative/rotary schemes like RoPE) add or inject position information so word order influences attention.

Explain BIO tagging and why a CRF layer helps NER.

BIO labels each token as B-TYPE (begins an entity), I-TYPE (inside) or O (outside), turning span extraction into token classification. Independent per-token softmax can produce invalid sequences such as O followed by I-PER, or B-LOC followed by I-ORG. A CRF layer learns transition scores between tags and decodes the best whole sequence with Viterbi, enforcing consistency and usually adding 1–2 F1 points over plain softmax on BiLSTMs (less gain with BERT).

How do you align word-level NER labels with subword tokens?

Use the tokenizer's word ids or offset mapping. Assign the word's label to its first subword and set the label of continuation subwords to −100 (ignored by the cross-entropy loss), or propagate I- labels to them. At inference, take the first subword's prediction per word and use offsets to recover character spans. Special tokens ([CLS], [SEP], padding) also get −100.

Why is accuracy a poor metric for NER, and what should be used?

Most tokens are O, so predicting O everywhere yields high token accuracy while finding no entities. Use entity-level precision, recall and F1 where a prediction counts only if span boundaries and type both match exactly (strict), with per-type breakdowns; lenient (partial overlap) scores are useful for diagnosis. Libraries such as seqeval implement this.

Work through BLEU for candidate "the cat is on mat" vs reference "the cat is on the mat".

p1 = 5/5 = 1.0, p2 = 3/4 (the cat, cat is, is on match; on mat does not), p3 = 2/3, p4 = 1/2. Geometric mean = (1 × 0.75 × 0.667 × 0.5)1/4 = 0.251/4 ≈ 0.707. The candidate is shorter (c = 5 < r = 6), so BP = e1−6/5 = e−0.2 ≈ 0.819. BLEU ≈ 0.819 × 0.707 ≈ 0.58.

Why does BLEU clip n-gram counts and apply a brevity penalty?

Without clipping, "the the the the" would get perfect unigram precision against any reference containing "the"; clipping caps each n-gram's credit at its count in the reference. Without the brevity penalty, a model could output one or two very safe words and achieve high precision. BP multiplies the score by e1−r/c when the candidate is shorter than the reference.

Explain ROUGE-1, ROUGE-2 and ROUGE-L with an example.

Reference "the cat is on the mat", candidate "the cat is on mat". ROUGE-1 recall = matched unigrams / reference unigrams = 5/6 = 0.83. ROUGE-2 recall = matched bigrams / reference bigrams = 3/5 = 0.6. ROUGE-L uses the longest common subsequence: "the cat is on mat" has length 5, so recall = 5/6 = 0.83. ROUGE-L rewards correct order without requiring contiguous matches.

What are METEOR and BERTScore, and when are they better than BLEU/ROUGE?

METEOR aligns words by exact match, stem and synonym, computes a recall-weighted harmonic mean, and penalizes fragmented alignments; it correlates better with human judgement at sentence level. BERTScore embeds tokens with a contextual model and greedily matches each candidate token to its most similar reference token by cosine, yielding precision, recall and F1. Both give credit for paraphrases ("large" vs "big") that n-gram metrics score as zero, so they suit paraphrase-heavy generation.

Macro-F1 vs micro-F1: which should you report for imbalanced intent classification?

Micro-F1 pools all predictions, so it is dominated by frequent classes (and equals accuracy in single-label multi-class). Macro-F1 averages per-class F1 equally, so poor performance on rare intents is visible. For imbalanced problems report macro-F1 (plus per-class metrics); weighted-F1 is a compromise that weights by support.

How does LDA work, and how do you choose the number of topics?

LDA assumes each document has a topic mixture drawn from a Dirichlet, each topic has a word distribution drawn from another Dirichlet, and each word is produced by sampling a topic from the document's mixture and then a word from that topic. Inference (Gibbs sampling or variational Bayes) recovers topics and mixtures from word counts. Choose K by sweeping values and comparing topic coherence (Cv, NPMI), perplexity on held-out data (which often disagrees with human judgement), and manual inspection for interpretability and usefulness.

Advanced

Derive why skip-gram with negative sampling is related to matrix factorization.

It has been shown that, at optimum, skip-gram with negative sampling implicitly factorizes a word–context matrix whose entries are the pointwise mutual information shifted by log k: w·c ≈ PMI(w, c) − log k, where PMI(w, c) = log[P(w,c)/(P(w)P(c))]. So word2vec and count-based methods (PMI matrices + SVD, GloVe) are closely related; differences in practice come largely from hyperparameters such as subsampling, context distribution smoothing (the 3/4 power) and window weighting.

Why does embedding arithmetic (king − man + woman) work, and what are its limits?

If a relation such as gender changes context distributions consistently (man:woman as king:queen), log-linear training makes the difference vectors approximately parallel, so adding the offset lands near the target. Limits: the nearest neighbour must exclude the query words (otherwise "king" often wins); it works for frequent, regular relations and fails for many others; analogy benchmarks overstate quality; and the same geometry encodes social biases (man:programmer as woman:homemaker).

How would you detect and mitigate bias in word or sentence embeddings?

Detect with association tests (WEAT/SEAT: compare cosine associations between target sets like male/female names and attribute sets like career/family), projection onto a bias direction (for example he−she), and downstream audits comparing error rates across demographic slices. Mitigate with counterfactual data augmentation (swap gendered terms), debiasing projections (with the caveat that they often hide rather than remove bias), balanced fine-tuning data, and, most importantly, measuring downstream fairness metrics rather than only embedding geometry.

Why are raw BERT sentence embeddings poor for semantic similarity, and how do SBERT and contrastive training fix this?

BERT was pretrained with masked language modeling (and next-sentence prediction), not to make cosine similarity meaningful. Its representation space is anisotropic: vectors occupy a narrow cone, so almost all pairs have high cosine and frequency effects dominate. SBERT fine-tunes a siamese BERT on sentence pairs (NLI, paraphrase) with classification or cosine losses; modern embedders use InfoNCE contrastive loss with large batches of in-batch negatives and mined hard negatives, spreading the space so that distance reflects meaning. Post-hoc whitening also helps partially.

What are hard negatives in contrastive embedding training, and what can go wrong with them?

Hard negatives are passages that look relevant (high BM25 or embedding score) but are not the correct answer; they teach fine distinctions that random negatives cannot. The risk is false negatives: many "hard negatives" are actually relevant but unlabelled, and pushing them away hurts the model. Mitigations: filter mined negatives with a cross-encoder (drop those it scores as relevant), skip the top few ranks, and use denoised or distilled labels.

Compare subword regularization (Unigram sampling, BPE-dropout) with deterministic tokenization.

Deterministic tokenizers always produce the same segmentation, so the model never sees alternative splits of a word and can be brittle to typos or unusual forms. Subword regularization samples different segmentations during training (Unigram's probabilistic segmentation, or BPE-dropout randomly skipping merges), acting as data augmentation. It improves robustness and low-resource translation quality; at inference the most likely segmentation is used.

How would you extend a pretrained model's tokenizer for a new domain or language?

Train a tokenizer on domain text, identify frequent domain strings that fragment heavily, and add them as new tokens (or merge vocabularies). Resize the embedding and output matrices; initialize new rows sensibly (for example the mean of the embeddings of the old subwords that make up the new token, not random). Then continue pretraining on domain text so new embeddings are learned before task fine-tuning. Measure fertility and downstream accuracy before and after; if gains are small, keep the original tokenizer, since changing it has integration costs.

Explain the exploding gradient problem and the fixes for recurrent networks.

When the recurrent Jacobians have norms above 1, gradients grow exponentially through time, producing huge updates and NaN losses. Fixes: gradient-norm clipping (rescale the gradient if its norm exceeds a threshold such as 1.0), careful initialization (orthogonal recurrent weights), lower learning rates, gated cells (LSTM/GRU), layer normalization, and truncated backpropagation through time over shorter windows.

What is truncated BPTT and when do you use it?

Instead of backpropagating through an entire very long sequence, split it into windows (for example 100–200 steps), carry the hidden state forward across windows but stop gradients at the window boundary. It bounds memory and compute and reduces exploding gradients, at the cost of not learning dependencies longer than the window through gradients. Used for character-level language modeling and streaming data.

Compare Bahdanau (additive) and Luong (multiplicative) attention.

Bahdanau attention scores with a small feed-forward net, score = vT tanh(W1henc + W2sdec), using the previous decoder state; it is flexible when dimensions differ. Luong attention uses dot or bilinear products, score = sTW h, computed with the current decoder state; it is cheaper and maps to matrix multiplication. Transformers adopted scaled dot-product attention for efficiency, with the √dk scaling to fix the large-dimension saturation that makes unscaled dot products worse than additive attention.

When would you still choose a BiLSTM-CRF over a Transformer for sequence labelling?

When latency, memory or hardware is extremely constrained (on-device, CPU-only, very high throughput), when sequences are very long and streaming is required, when you have a small amount of in-domain data and strong domain embeddings, or when interpretability of transition constraints matters. Otherwise a fine-tuned small Transformer (DistilBERT, MiniLM) usually wins and a distilled version can meet the same latency.

How do encoder-only, decoder-only and encoder–decoder models map to NLP tasks?

Encoder-only (BERT, RoBERTa, DeBERTa): bidirectional context, best cost-performance for classification, NER, extractive QA, embeddings and re-ranking. Decoder-only (GPT, Llama): causal generation, instruction following, chat, few-shot tasks; can do any task via prompting but at higher cost per prediction. Encoder–decoder (T5, BART, mT5): input-to-output transformations such as translation and summarization. For NER on legal text with labelled data, an encoder-only model is the natural and cheapest choice because output is aligned one-to-one with input tokens.

Why is BLEU criticized, and what metrics are preferred for machine translation today?

BLEU correlates weakly with human judgements at the sentence level, ignores synonyms and paraphrase, depends heavily on tokenization and the number of references, and does not measure adequacy or fluency directly. Different BLEU implementations are not comparable (hence sacreBLEU). Current practice uses neural metrics trained on human ratings (COMET, BLEURT) plus chrF for morphologically rich languages, with human evaluation (for example MQM error annotation) for final decisions.

How would you evaluate faithfulness of an abstractive summary or RAG answer?

Split the output into atomic claims; for each claim, check whether the source entails it using an NLI model or an LLM judge constrained to the source; report the fraction supported (and flag contradictions). QA-based methods generate questions from the summary and check that answers from the source match. For RAG, also measure citation precision/recall (does each cited passage support the sentence?). Validate the automatic judge against a human-labelled set and track agreement.

What are the pitfalls of using LLM-as-a-judge, and how do you mitigate them?

Pitfalls: position bias (prefers the first or second answer), verbosity bias (prefers longer answers), self-preference (prefers outputs from its own family), sensitivity to rubric wording, inconsistency across runs, and inability to verify facts it does not know. Mitigations: randomize and swap order, use pairwise comparison with explicit rubrics and reference answers, require evidence-based rationales, use a different model family as judge, average multiple samples, and calibrate against human labels with agreement statistics.

Explain zero-shot cross-lingual transfer and why it works.

A multilingual encoder (mBERT, XLM-R) pretrained on many languages develops partially language-agnostic representations because of shared subwords, shared numbers and named entities, and structural similarities across languages. Fine-tuning a classifier on English labels then transfers to other languages with no target-language labels. It works best for related languages and scripts, and degrades for low-resource languages, different scripts, and culturally specific tasks such as sentiment and sarcasm; a small amount of target-language data typically closes much of the gap.

How would you handle Romanized, code-mixed text like Hinglish?
  • Token-level language identification to know which parts are Hindi or English.
  • Spelling normalization for variants (bahut/bohot/bhot) via clustering or learned normalizers.
  • Optional back-transliteration to Devanagari to match the pretraining data of Indic models.
  • Models pretrained on code-mixed or Indic data (MuRIL, IndicBERT, multilingual embedders), or multilingual LLMs with few-shot code-mixed examples.
  • In-domain labelled data (LLM pre-labelling plus human verification).
  • Per-slice evaluation on real traffic, and cost checks because Romanized Hindi fragments into many tokens.
How do you make LLM-based information extraction reliable enough for production?

Define a strict schema with enums and nullable fields; use constrained decoding or function calling; validate and retry on schema errors; require evidence quotes or offsets and verify them against the source; normalize values and link to canonical IDs; route low-agreement or ungrounded cases to humans; evaluate field-level precision/recall/F1 on a gold set and track drift; version prompts and models; and use rules or small fine-tuned models for high-volume, well-defined fields to cut cost and variance.

How would you distill an LLM classifier into a small model?

Run the LLM (with a carefully validated prompt) over a large, representative unlabelled sample to produce labels, ideally with probabilities or several samples for soft labels. Have humans audit a stratified subset to estimate label quality and fix systematic errors. Fine-tune a small encoder (DistilBERT, MiniLM) on the labels, optionally with soft-label (knowledge distillation) loss. Evaluate on a human-labelled test set, not LLM labels. Keep the LLM as a fallback for low-confidence predictions. The result is often orders of magnitude cheaper and faster with similar accuracy.

How do you link extracted entities to an ontology such as ICD codes without hallucinated codes?

Never let a model generate codes freely. Generate candidates from the ontology with synonym dictionaries, fuzzy matching and dense retrieval over code descriptions; then disambiguate with a cross-encoder or an LLM that must choose from the retrieved list (or "none"), using the surrounding clinical context (negation, laterality, acuity). Validate the chosen code exists, attach the evidence span and a confidence score, and route low-confidence or high-impact cases to human coders. Evaluate top-1 and top-k accuracy per code family.

How should negation and uncertainty be handled in clinical or legal extraction?

A mention is not an assertion: "no signs of pneumonia", "rule out MI", "family history of diabetes", "the tenant shall not" all contain entities that must not be recorded as positive facts. Add assertion attributes to each extraction (present, absent, possible, hypothetical, historical, about someone else) using rule-based systems (NegEx-style triggers and scopes), trained assertion classifiers, or explicit schema fields for an LLM. Evaluate assertion accuracy separately from entity detection.

Why can a model with lower training loss have worse test F1, as with the IMDB LSTM and BiLSTM?

Training loss measures fit to the training set; test F1 measures generalization at a specific threshold. Lower training loss can reflect memorization (the BiLSTM reached 0.048 but F1 0.811), a skewed decision rule where probabilities are poorly calibrated around 0.5 (the LSTM improved loss but predicted one class too often, recall 0.43), or a train–test distribution shift. Diagnose with a validation curve, confusion matrix, precision/recall per class, calibration plots and threshold tuning on validation data.

How would you handle documents longer than a model's context window for classification or extraction?

Options: head-and-tail truncation (keep the first and last parts, which often hold the verdict), sliding windows with overlap and aggregation (mean or max of logits, or attention pooling over chunk embeddings), hierarchical models (encode chunks, then a second model over chunk vectors), long-context architectures (Longformer, BigBird, long-context LLMs), or retrieval of relevant chunks first. For extraction, process chunks with overlap and deduplicate entities across chunk boundaries; for contracts, segment by clause structure rather than fixed length.

How do you detect and prevent data leakage in NLP datasets?

Leakage sources: fitting vectorizers or tokenizers on test data, near-duplicate texts across splits (reposted reviews, templated emails), the same user or product in train and test, time leakage (future events in training), label words appearing in the text (a "category: refund" header), and benchmark contamination in LLM pretraining. Prevent with fit-on-train pipelines, near-duplicate detection (MinHash) before splitting, group and time-based splits, and contamination checks for LLM evaluations.

Scenario & debugging

Banking: fraud analysts write free-text investigation notes. Design a system that extracts entities (accounts, devices, locations), identifies fraud patterns and suggests rule updates, with no hallucinated patterns and regulatory explainability.
  1. Ingestion and privacy: pull notes from the case system, mask PII in logs, keep processing inside the secure environment, attach case IDs and outcomes (confirmed fraud or not).
  2. Entity extraction: regex/validators for rigid formats (account numbers, IBANs, IPs, device IDs), a fine-tuned NER model for common types, an LLM with a strict JSON schema for long-tail entities and relations ("device X logged into accounts A and B"). Every entity carries an evidence span verified against the note.
  3. Linking: resolve entities to the bank's master records (account table, device fingerprint store) and build an entity graph joining text-derived links with transaction data.
  4. Pattern discovery: cluster notes by embeddings and extracted attributes; an agent summarizes each emerging cluster into a candidate pattern ("new-device login + change of phone number + large transfer within 2 hours") with citations to the notes that support it.
  5. Rule suggestion without hallucination: the agent outputs a candidate rule only in a formal rule DSL over existing structured fields; the rule is back-tested on historical transactions (precision, recall, alert volume, false-positive cost). Rules with no statistical support are discarded regardless of how plausible the text sounds.
  6. Human approval: fraud analysts review rules with their evidence notes, back-test metrics and rationale; nothing is deployed automatically. Approved rules are versioned with an audit trail.
  7. Explainability: each alert shows the rule, the rule's origin (source notes, approval record) and the contributing features.
  8. Evaluation: extraction F1 on a labelled note set, pattern precision judged by analysts, rule lift in shadow mode, and drift monitoring.
Retail: build a sentiment-to-action engine over reviews, chats and social complaints that performs aspect-based sentiment, maps complaints to operational issues and triggers workflows, including Hinglish and code-mixed text.
  1. Unify sources: normalize reviews, chat transcripts and social posts into one schema (text, channel, order/SKU if known, timestamp, language).
  2. Language handling: language ID at sentence/token level, spelling normalization for Romanized Hindi, multilingual encoder or multilingual LLM.
  3. Aspect-based sentiment: a fixed, versioned taxonomy mapped to owners (delivery_delay → logistics, product_quality → supplier team, refund → payments, packaging, pricing). Start with LLM extraction of (aspect, polarity, severity, evidence), then distill into a fine-tuned multilingual ABSA model for volume.
  4. Map to operations: join with order data (warehouse, carrier, supplier, SKU) so complaints become measurable signals per entity.
  5. Trigger logic: rules plus anomaly detection on aggregated rates ("delivery_delay negative mentions for warehouse W above 3σ of baseline over 24 hours with at least 30 mentions") trigger a logistics audit ticket; product-quality clusters for one SKU alert the supplier team. Individual severe cases (safety, fraud) escalate immediately.
  6. Agent layer: an agent drafts the ticket with summary, example quotes, counts and suspected root cause, and a human confirms before costly actions.
  7. Evaluation: aspect-level F1 per language on a stratified labelled set, precision of triggered workflows, time-to-detection, and business outcomes (return rate, repeat complaints).
Healthcare: design a pipeline that extracts symptoms, diagnoses and medications from clinical notes, maps them to ICD codes, generates patient summaries, and lets an agent suggest missing tests or possible diagnoses, with zero tolerance for hallucination and confidence scores.
  1. Preprocess: section detection (history, assessment, plan), abbreviation expansion, de-identification, sentence segmentation.
  2. Clinical NER: a domain encoder (clinical/biomedical BERT) fine-tuned for PROBLEM, SYMPTOM, MEDICATION, DOSE, TEST, with assertion status (present, negated, possible, historical, family) so "no chest pain" is not recorded as chest pain.
  3. Ontology mapping: candidate retrieval from ICD/SNOMED/RxNorm synonym lists and embedding search; a constrained re-ranker selects from candidates only; unmatched mentions are flagged, never invented.
  4. Summaries: generate from the structured extractions plus retrieved note sentences, with every sentence citing its source span; an NLI verifier rejects unsupported sentences.
  5. Agentic suggestions: the agent compares the patient's structured data against retrieved clinical guidelines ("diabetic patient, no HbA1c in 12 months") and proposes suggestions with the guideline citation and the patient evidence. Suggestions are advisory, shown to clinicians, never auto-ordered.
  6. Confidence: calibrated model probabilities for extraction and linking (temperature scaling on validation data), agreement between methods, and evidence coverage; below-threshold items go to human review.
  7. Evaluation and governance: entity and assertion F1, coding accuracy, clinician review of summaries for faithfulness, audit logs, bias checks across patient groups, and regulatory review before deployment.
Legal: build a contract intelligence system that extracts clauses (termination, liability, indemnity), identifies risky terms, compares against standard templates, and has an agent flag deviations and suggest rewrites, for 100+ page documents in domain-specific language.
  1. Parse structure: convert PDFs (with OCR if needed), detect headings, numbering and clause boundaries, and preserve page and section references. Segment by clause, not fixed-size chunks.
  2. Clause classification: a fine-tuned legal encoder (or LLM with definitions) labels each clause type; long contracts are handled clause by clause, with cross-references ("as defined in Section 2.1") resolved by retrieval.
  3. Field extraction: liability cap amount, notice period, governing law, auto-renewal, with evidence spans.
  4. Template comparison: embed each clause and retrieve the matching clause from the firm's standard playbook; an LLM describes differences in a structured way (missing carve-out, cap raised from 1× to 3× fees, unilateral termination added).
  5. Risk scoring: rules from the firm's playbook (for example uncapped indemnity = high risk) combined with model judgements; each flag cites the clause and the playbook rule.
  6. Rewrite suggestions: the agent proposes fallback language drawn from approved clause libraries rather than free generation, shown as a tracked-changes redline for lawyer approval.
  7. Evaluation: clause-type F1, field accuracy, flag precision/recall versus lawyer review, time saved per contract. Confidentiality: private deployment, no training on client data without consent.
Media: design real-time content moderation that detects hate speech, misinformation and sarcasm, understands conversation context, escalates borderline cases and learns from moderator feedback.
  1. Tiered architecture for latency: tier 1 hash matching of known bad content and fast lexical/regex filters (sub-millisecond); tier 2 a small multilingual fine-tuned classifier with multiple heads (hate, harassment, self-harm, spam) in tens of milliseconds; tier 3 an LLM or larger model for ambiguous cases, asynchronously when possible.
  2. Context: include the parent post, thread history and target of the reply; sarcasm and reclaimed slurs often need context. Encode the conversation window, not just the single message.
  3. Misinformation: claim detection, then retrieval against fact-check databases and trusted sources; label rather than delete unless policy requires.
  4. Thresholds and escalation: two thresholds per category: auto-action above high confidence, human review in the uncertain band, allow below. Prioritize the review queue by severity and reach.
  5. Learning from feedback: moderator decisions become labelled data; active learning samples uncertain and disagreement cases; periodic retraining with evaluation gates to avoid regressions; track adversarial spellings ("h8", leetspeak) and add character-level robustness.
  6. Fairness and policy: per-dialect and per-community false-positive audits (for example, dialects wrongly flagged as toxic), transparent policy definitions, appeal flows.
  7. Metrics: precision and recall per category at operating thresholds, PR-AUC, p95 latency, prevalence of violating content viewed, appeal overturn rate.
Telecom: using call-center transcripts and chat logs, cluster issues dynamically, identify root causes, generate automated responses, and let an agent decide to respond, escalate or trigger a backend fix.
  1. Transcripts: speech recognition with diarization, punctuation restoration, PII redaction; accept that transcripts are noisy.
  2. Per-conversation understanding: intent classification, entity extraction (plan, device, location, error code), sentiment and resolution status; a structured summary per conversation.
  3. Dynamic clustering: embed summaries and cluster with a density method (HDBSCAN or BERTopic) over rolling windows; new clusters that grow quickly are emerging issues. An LLM names each cluster from representative examples.
  4. Root cause: correlate clusters with network telemetry, outages, releases and region (an issue cluster in one city after a firmware push points to the release). Present evidence rather than claiming causation.
  5. Responses: RAG over the knowledge base and resolved tickets to draft answers with citations; automatic sending only for high-confidence, low-risk intents.
  6. Agent decision policy: respond automatically if intent is known, confidence high and a knowledge-base answer exists; trigger a backend action (reset provisioning, re-send SIM activation) only through allow-listed tools with preconditions and audit logs; escalate on low confidence, anger/churn risk, repeated contact or regulated topics.
  7. Metrics: resolution time, first-contact resolution, containment rate with customer satisfaction, escalation precision, false automation rate.
Automotive: design in-car voice command understanding that works on noisy speech, maintains conversation context and tracks preferences, including the ambiguous command "make it cooler."
  1. Speech front end: noise-robust on-device speech recognition with n-best hypotheses (not just the top one) and confidence scores.
  2. NLU: joint intent classification and slot filling (a small on-device encoder for latency and offline use), robust to recognition errors by training on noisy transcripts and using the n-best list.
  3. Dialogue state: track the active domain, last referenced device and slot values ("it" refers to the last device discussed), with a time decay.
  4. Resolving "make it cooler": combine context signals: the active domain (music playing and the last command was about the equalizer vs climate), sensor data (cabin at 29°C vs 19°C), and user history. If the probability of the top interpretation is high, act and confirm briefly ("Lowering temperature to 21"); if two interpretations are close, ask a short clarifying question. Make actions easily reversible ("undo").
  5. Preferences: learn per-driver defaults (preferred temperature, step size) with consent, stored locally.
  6. Safety: minimize distraction, restrict certain commands while driving, never execute safety-critical actions from ambiguous input.
  7. Metrics: intent accuracy and slot F1 on in-car noisy audio, task success rate, clarification rate, latency, correction/undo rate.
Finance: build a RAG-based market intelligence system over news, earnings calls and reports, with an agent that monitors news, detects anomalies and correlates with stock data while distinguishing causality from correlation.
  1. Ingestion: streaming news, call transcripts and filings; deduplicate syndicated stories; extract timestamps, companies (entity-linked to tickers), events (guidance cut, M&A, lawsuit) and sentiment toward each entity.
  2. Retrieval: hybrid BM25 + embeddings with metadata filters (ticker, date range, document type), finance-tuned embeddings, re-ranking, and point-in-time correctness (never retrieve documents published after the query's as-of date).
  3. Anomaly detection: spikes in news volume or sentiment per entity, and unusual price/volume moves; align the two time series.
  4. Causality vs correlation: the agent reports temporal ordering (news before the move?), event studies (abnormal return relative to market and sector in a window around the event), confounders (market-wide moves, same-day earnings), and multiple candidate explanations with evidence. Language is hedged ("coincided with", "likely contributed") unless evidence is strong; no direct claims of causation from co-occurrence alone.
  5. Outputs: analyst briefs with citations, confidence and counter-evidence; no automated trading decisions without separate controls.
  6. Evaluation: retrieval recall on analyst questions, faithfulness of briefs, entity-linking accuracy, and analyst usefulness ratings; backtests for any signals.
Education: design an adaptive learning assistant that analyzes student questions, detects knowledge gaps, adapts the learning path and adjusts difficulty.
  1. Concept map: a curriculum graph of concepts with prerequisites; tag all content and exercises with concepts.
  2. Query analysis: classify each question to concepts (embedding similarity + classifier), detect question type (definition, procedural, misconception) and extract misconceptions ("thinks correlation implies causation").
  3. Knowledge state: model mastery per concept with knowledge tracing (Bayesian or deep), updated by exercise results and question signals; repeated questions on prerequisites signal a gap.
  4. Path adaptation: recommend content for the weakest prerequisite first; choose exercise difficulty to keep success around 70–85%.
  5. Tutoring responses: RAG over vetted course content, Socratic hints before full answers, cited sources; guard against giving away assessment answers.
  6. Privacy and fairness: minimal data retention, parent/institution controls for minors, check that recommendations do not disadvantage groups.
  7. Metrics: learning gains on assessments, concept-classification accuracy, retention and engagement, answer correctness audits.
Manufacturing: factories generate text incident logs. Design a system that extracts causes, actions and outcomes, builds a knowledge graph, and has an agent suggest preventive actions and flag recurring patterns.
  1. Normalize logs: expand plant jargon and abbreviations with a domain dictionary; attach metadata (line, machine ID, shift, timestamp).
  2. Extraction: a schema of EQUIPMENT, COMPONENT, FAILURE_MODE, CAUSE, ACTION_TAKEN, OUTCOME with relations (cause → failure → action → outcome); LLM extraction with evidence spans, then distill to a smaller model; normalize to a failure-mode taxonomy.
  3. Knowledge graph: nodes for machines, components, failure modes, causes and actions; edges with counts and timestamps; entity resolution merges "hyd pump" and "hydraulic pump".
  4. Recurring patterns: frequent subgraph and time-series analysis ("bearing overheating on line 3 recurs every 6 weeks after lubrication was deferred"), plus embedding similarity of new incidents to past ones.
  5. Agent: for a new incident, retrieve similar past incidents and their successful actions via the graph and RAG, propose preventive actions citing precedents, and flag repeats to maintenance leads. Safety-critical recommendations need engineer approval.
  6. Metrics: extraction F1, entity-resolution accuracy, precision of recurring-pattern alerts, reduction in repeat incidents and downtime.
Your LSTM sentiment model on IMDB gets F1 0.53 while logistic regression on bag-of-words gets 0.875. The LSTM's training loss went from 0.693 to 0.569. How do you debug it?
  1. Check predictions: confusion matrix shows recall 0.43 and precision 0.69, so it is heavily favouring one class; check the predicted-class distribution and whether a better threshold on validation data helps.
  2. Check the input: are sequences truncated from the right while the verdict sits at the end of reviews? Is padding placed so the final state is computed after hundreds of pad tokens? Use packed sequences or pool over all outputs with masking.
  3. Check optimization: gradient norms, add clipping (max norm 1.0), tune learning rate, initialize forget-gate bias to 1, use Adam(W).
  4. Check capacity vs data: 4.6M parameters trained from scratch on 25k reviews; use pretrained embeddings (FastText/GloVe), dropout and early stopping on validation F1.
  5. Check data pipeline: vocabulary built on train only, labels aligned after shuffling, same preprocessing at test.
  6. Compare to stronger options: a BiLSTM with pretrained vectors, an attention model, or a fine-tuned DistilBERT (around 0.93).
  7. Decide: if the linear model meets requirements, ship it; it is cheaper and explainable.
Your vanilla RNN's training loss stays at about 0.699 for 10 epochs on binary sentiment. What does that number tell you, and what do you change?

For a balanced binary task, cross-entropy of a model that always predicts 0.5 is ln 2 ≈ 0.693. A loss stuck near 0.69–0.70 means the model has learned nothing beyond chance. Likely causes on 400-token sequences: vanishing gradients, final hidden state dominated by padding, learning rate too high or too low, or a bug (labels shuffled separately from inputs, loss applied to wrong outputs). Fixes: overfit a tiny batch first to prove the pipeline can learn; clip gradients; use packing or masked pooling; shorten sequences; switch to LSTM/GRU or attention; use pretrained embeddings; check learning rate with a range test.

Your BiLSTM with FastText reaches training loss 0.048 but test F1 0.811 with precision 0.76 and recall 0.87. What is happening and how do you fix it?

The model is overfitting (memorizing training reviews) and its decision rule is biased toward the positive class (many false positives). Fixes: monitor validation loss/F1 and stop early; add dropout on embeddings and LSTM outputs; weight decay; freeze the pretrained embeddings for the first epochs, then unfreeze with a smaller learning rate; reduce hidden size; data augmentation; tune the decision threshold on validation data to balance precision and recall; and consider an attention or pretrained Transformer model, which generalized better in the same study.

A sentiment model trained on movie reviews performs poorly on product reviews and tweets. What do you do?

This is domain shift: vocabulary, style and polarity cues differ ("unpredictable" is good for plots, bad for products; tweets have slang, emojis, sarcasm). Steps: measure the drop on a labelled target-domain sample; inspect errors by category; label a few hundred to a few thousand target examples (with active learning and LLM pre-labelling); continue pretraining on unlabelled target text (domain-adaptive pretraining) and fine-tune; or use a strong general-purpose model zero-shot as a baseline. Evaluate per domain and monitor drift after deployment.

Your NER model scores 0.92 F1 on the test set, but users say it misses many entities in production. How do you investigate?

Suspect a test set that does not reflect production. Sample production texts and label them; compare distributions (document length, formatting, entity types, casing, language). Common culprits: all-caps or lowercased input, new entity names (new products, people), OCR noise, longer documents truncated at 512 tokens, different tokenization at serving, or sentence splitting that cuts entities. Check per-type recall; fix with targeted labelling, augmentation (case and noise), sliding windows for long text, gazetteers for new names, and a monitoring pipeline sampling production outputs for review.

A search system using embeddings fails on queries with product codes and rare names. How do you fix it?

Dense embeddings capture meaning but blur exact tokens like "XR-2291" or rare surnames. Add a lexical retriever (BM25) and combine results with hybrid scoring or reciprocal rank fusion; add exact-match or prefix indexes for identifiers; consider sparse learned models (SPLADE); normalize codes consistently at index and query time; and re-rank with a cross-encoder. Evaluate recall@k separately for identifier-style and natural-language queries.

You must classify 50 million support tickets per day into 40 categories. An LLM prompt achieves 91% macro-F1 but is too expensive. What is your plan?

Use the LLM as a teacher: label a few hundred thousand representative tickets, audit a stratified sample with humans, and fine-tune a small encoder (or even TF-IDF + linear for easy categories) on these labels. Evaluate on a human-labelled test set; target near-LLM macro-F1. Deploy the small model on CPUs or small GPUs with batching; route low-confidence predictions (a few percent) to the LLM or humans. Monitor per-category drift and periodically relabel new data with the LLM to retrain. Cost typically drops by orders of magnitude with a small accuracy trade-off.

Your summarizer's ROUGE improved after a model change, but users report more factual errors. What happened, and how do you evaluate properly?

ROUGE rewards lexical overlap with references, not factual consistency; the new model may copy more reference-like phrasing while inventing or distorting details (numbers, names, negations). Add faithfulness evaluation: claim-level NLI or LLM-judge checks against the source, entity and number consistency checks, and human review on a sample. Make faithfulness a release gate alongside ROUGE, and consider constrained or extract-then-abstract generation for high-stakes content.

An LLM extraction pipeline sometimes returns entities that are not in the source document. How do you eliminate these?

Require an evidence quote or character offsets for every extracted value and programmatically verify that the quote exists in the source (exact or normalized match); drop or flag values that fail. Use constrained decoding with a schema that includes null/unknown options, instruct the model to omit uncertain fields, lower temperature, and include negative few-shot examples. For codes and IDs, choose from retrieved candidates only. Track the rate of ungrounded outputs as a metric and add a human review queue for flagged cases.

Your team wants to switch the embedding model in a production RAG system. What must you do?

Vectors from different models live in different spaces, so the entire corpus must be re-embedded; queries must use the new model with its required prefixes and normalization. Build an offline evaluation set of real queries with relevance labels and compare recall@k, MRR and end-to-end answer quality. Check token limits (chunk size may need to change), dimension and index configuration, latency and cost. Run both indexes in parallel (shadow or A/B), then cut over, keeping the old index for rollback.

A chatbot for an Indian bank must understand English, Hindi, Hinglish and Tamil. Token costs are much higher for Hindi and Tamil. How do you design it?

Measure tokenizer fertility per language for candidate LLMs and choose one with an efficient multilingual tokenizer. Use a small multilingual intent/entity model (for example an Indic-pretrained encoder) for routing and for handling frequent intents without an LLM call. For RAG, keep a multilingual embedding model so queries in any language retrieve documents in any language. Keep system prompts compact and in English if quality holds, while answering in the user's language. Evaluate per language and script, including Romanized input, and budget per-language cost.