Edge AI Fundamentals
Edge AI means running a trained model close to where data is produced (on a phone, watch, car, camera, gateway or microcontroller) instead of sending everything to a data centre. The model is rarely the hard part; making it fast, small, cool, private and power-efficient on real, fragmented hardware is a systems-engineering discipline, and this page covers it end to end, from the physics of memory bandwidth to quantization maths, runtimes, on-device LLMs and production trade-offs.
- You move inference to the edge for latency, privacy, offline use, bandwidth and serving cost, and you pay with tight budgets for compute, memory capacity, memory bandwidth, power, heat and storage, plus hardware fragmentation.
- The hardware ladder is CPU (flexible), GPU (parallel, FP16-friendly), NPU/DSP (best performance per watt for low-precision tensor maths, limited operator set) and MCU (milliwatts, kilobytes). Peak TOPS is a ceiling; memory bandwidth and operator coverage decide real speed.
- The roofline model explains almost everything: batch-1 workloads such as LLM decode are memory-bound, so the lever is fewer bytes (quantization), not more TOPS.
- Quantization (FP32 to INT8/INT4, or FP8/block formats) is the single biggest lever: 4-8x smaller, much less bandwidth, access to integer engines, usually a small accuracy drop if calibrated well. GPTQ, AWQ, SmoothQuant and rotation methods make 4-bit LLMs practical.
- Key runtimes: LiteRT (formerly TensorFlow Lite), ONNX Runtime, ExecuTorch, Qualcomm AI Engine Direct (QNN/QAIRT), Core ML, TensorRT on Jetson, llama.cpp/GGUF and MLC-LLM. NNAPI is deprecated from Android 15.
- For on-device LLMs, prefill is compute-bound (sets time to first token) and decode is bandwidth-bound: tokens/s is roughly effective bandwidth divided by bytes read per token, and the KV cache can outgrow the weights at long context.
- Real products are hybrid (small local model plus cloud fallback), are measured on energy per inference and sustained (throttled) performance, and use federated learning and differential privacy when they must learn from user data.
What edge AI is, and why teams do it
Cloud inference is simple to build: send the input to a server with big GPUs, get the answer back. Edge inference runs the same kind of model locally, on the device's own CPU, GPU or dedicated neural accelerator, or on a nearby box rather than a distant data centre. "On-device AI" is the most extreme point of edge AI: the model runs on the very device that owns the sensor and the user.
Cloud AI is like calling an expert on the phone: they know a lot, but you wait for them to pick up, you pay per call, and they hear everything you say. On-device AI is like carrying a pocket reference book: it knows less, but it is instant, private and works in a tunnel. An edge server is like a well-stocked library in your building: bigger than your pocket book, faster to reach than the expert across the country. The pocket book must be small and light (model size and memory), and reading it should not tire you out (power and thermal). Hybrid designs keep the pocket book for everyday questions and call the expert only for the hard ones.
The device-edge-cloud continuum
"Edge" is not one place; it is a spectrum of locations between the sensor and the data centre. Each step toward the sensor lowers latency and data exposure but shrinks the compute and power budget.
Sensor / MCU Device Near edge Far edge / regional Cloud
(TinyML) (phone, watch, (gateway, camera (telco MEC, on-prem (data centre
car ECU, laptop) NVR, Jetson box) server, store server) GPU clusters)
------------------------------------------------------------------------------------------------
mW, KB-MB 1-10 W, 4-24 GB 10-60 W, 8-64 GB 100s of W, GPUs MW, TB of HBM
<1 ms to sensor ~1-50 ms ~1-10 ms LAN ~5-30 ms ~50-300+ ms mobile
keyword spotting camera effects, multi-camera video city-scale video, frontier LLMs,
gesture, anomaly on-device LLM, analytics, robot AR offload, factory training, batch
detection photo search perception quality inspection analytics
------------------------------------------------------------------------------------------------
<-- more private, lower latency, works offline more capable, easier to update, costs per call -->
| Tier | Typical hardware | Typical models | Who owns it |
|---|---|---|---|
| Microcontroller (TinyML) | Arm Cortex-M class MCU, optionally with a micro-NPU; 64 KB-2 MB SRAM, 256 KB-8 MB flash | Keyword spotting, activity recognition, vibration anomaly detection, tiny vision (person detection) | Device maker; firmware updates |
| Personal device | Phone, tablet, laptop, watch, glasses SoC with CPU+GPU+NPU; 2-24 GB shared DRAM | Vision, speech, translation, small LLMs (0.5-4B), recommendations | User owns the hardware; app or OS ships the model |
| Embedded edge box | Jetson-class module, industrial PC, smart camera SoC, car domain controller | Multi-stream detection and tracking, segmentation, driver monitoring, robot perception | Operator or OEM; fleet management |
| Edge server / MEC | Small GPU servers in a store, factory, base station or regional point of presence | Larger models shared by many nearby devices, video analytics, AR rendering offload | Enterprise or telecom operator |
| Cloud | Data-centre GPUs/TPUs with HBM | Frontier LLMs, training, heavy batch jobs | Cloud provider |
Why run AI at the edge?
| Benefit | Why it matters | Typical example |
|---|---|---|
| Latency | No network round trip (often 50-300 ms on mobile networks, much worse on poor coverage). Real-time use cases need single-digit to tens of milliseconds, and predictable tail latency. | Camera effects at 30-60 fps, keyword spotting, live captions, AR tracking, emergency braking |
| Privacy | Raw data (photos, voice, messages, health signals) never leaves the device. This simplifies compliance and builds user trust. | Smart reply, on-device photo search, health anomaly detection |
| Cost | Every cloud inference costs GPU time. At hundreds of millions of users, moving inference to hardware the user already owns saves a lot of serving cost. | Summarization, text suggestions, image enhancement at scale |
| Offline and reliability | Works in airplanes, basements, rural areas and during outages. No dependency on a backend being up. | Translation, navigation assistance, safety features in cars |
| Bandwidth | Sending a summary or event instead of raw video or audio saves uplink bandwidth and data costs. One 1080p camera stream is several Mbit/s, all day. | Smart cameras that upload only detections |
| Data sovereignty and regulation | Some data legally must stay on premises or in a country; edge processing keeps it there. | Hospital imaging, factory data, government deployments |
| Context and personalization | The device sees the richest, freshest personal context (sensors, apps, history) without uploading it. | Personal keyboard models, on-device ranking, context-aware assistants |
The price you pay
Cloud inference
- Large models (hundreds of billions of parameters) are possible
- Easy to update the model any time; one hardware target
- Batching gives high hardware utilization
- Costs money per request, needs network, raises privacy questions, adds latency
Edge / on-device inference
- Models usually under ~4B parameters for phone LLMs, under ~100 MB for vision/audio, under ~1 MB on MCUs
- Must fit in a few GB of shared RAM next to other apps
- Thousands of device models, chipsets and driver versions
- Battery and thermal limits; batch size is usually 1; model updates ship through app, OS or firmware updates
Hybrid is the common answer
Most real products are hybrid. A small on-device model handles the fast, private or frequent cases, and hard cases fall back to the cloud. Examples: on-device wake-word detection that then streams to a cloud assistant; an on-device small LLM that answers simple prompts and escalates long or complex requests; on-device OCR followed by cloud translation for rare languages. A good design states explicitly which requests stay local, what triggers fallback, and what happens offline. The hybrid architectures section below goes deeper.
A decision checklist: should this model run at the edge?
- Latency budget If the end-to-end budget is below the network round trip plus server time (for example under 100 ms at the 99th percentile), the edge is almost mandatory.
- Data sensitivity Would users or regulators object to raw data leaving the device? If yes, process locally and send only results, or nothing.
- Connectivity Must it work offline or on poor networks? Then at least a degraded local path is needed.
- Volume and cost Multiply requests per user per day by users by cloud cost per request; high-frequency features are expensive in the cloud.
- Feasibility Can a model that meets the quality bar fit the device's memory, latency and energy budgets on the lowest tier you must support?
- Update cadence Does the model need to change weekly (cloud-friendly) or is it stable for months (edge-friendly)?
The constraints: compute, memory, bandwidth, power, heat, storage
Every edge design is a negotiation between six budgets. Beginners focus on compute (TOPS); experienced engineers look first at memory capacity and memory bandwidth, then at sustained power and heat, because those are what usually break a feature in the field.
Designing for the edge is like packing for a long hike instead of a road trip. The backpack's volume is memory capacity, how fast you can pull things out of it is memory bandwidth, your leg strength is compute, the food you carry is the battery, how hot you get climbing is thermal, and the pantry at home is storage. A car (the cloud) lets you ignore all of this; on foot, every gram and every minute counts, and the limiting factor on a steep climb is usually overheating, not strength, just as sustained thermal limits usually matter more than peak TOPS.
| Budget | Typical phone numbers | Typical MCU numbers | Data-centre GPU for contrast | What breaks when you exceed it |
|---|---|---|---|---|
| Compute | NPU: tens of INT8 TOPS peak; GPU: 1-4 FP16 TFLOPS; CPU: hundreds of GFLOPS | Tens to hundreds of MOPS (CPU); up to ~0.5 TOPS with a micro-NPU | Hundreds to thousands of dense TFLOPS | Latency too high, frames dropped |
| Memory capacity | 6-24 GB LPDDR shared by OS, apps, GPU and NPU; an app may realistically use 1-4 GB | 64 KB-2 MB SRAM | 80-192 GB HBM per GPU | Out-of-memory crash or the low-memory killer terminating the app |
| Memory bandwidth | ~50-100 GB/s LPDDR5/5X, shared | Hundreds of MB/s to a few GB/s | 2-8 TB/s HBM | Memory-bound layers and LLM decode run slowly regardless of TOPS |
| Power | Sustainable ~2-5 W for the whole SoC; bursts to 10+ W | 1-100 mW, often µW asleep | 700-1200 W per GPU | Battery drain, users disabling the feature |
| Thermal | Passive cooling; skin temperature limits around 40-45 °C | Rarely an issue | Active liquid or air cooling | Throttling: clocks drop and latency grows after minutes |
| Storage and download | App size limits and user patience; models of 50 MB-4 GB | Flash of 256 KB-8 MB for code plus weights | Practically unlimited | Install abandonment, long first-use downloads, update cost |
Memory bandwidth: the constraint people underestimate
Bandwidth is how many bytes per second can move between DRAM and the compute units. It is fixed by the memory technology and bus width: a phone with LPDDR5X at 8533 MT/s on a 64-bit bus delivers at most 8533e6 × 8 bytes ≈ 68 GB/s, and effective bandwidth for one workload is usually 60-80% of that, shared with the display, camera and other units. A laptop may reach 100-250 GB/s; a data-centre GPU 3-8 TB/s. Because every weight must be read at least once per inference at batch size 1, bandwidth alone sets a hard floor on latency:
Energy is dominated by data movement
At modern process nodes, fetching a value from DRAM costs roughly two to three orders of magnitude more energy than an 8-bit multiply-accumulate on it. Rough, widely quoted orders of magnitude: an INT8 add is well under a picojoule, a 32-bit SRAM read from a small local buffer is a few picojoules, and a 32-bit DRAM read is hundreds of picojoules. That is why accelerators have large on-chip SRAM, why operator fusion saves power as well as time, and why smaller data types help energy even when compute is not the bottleneck.
Fragmentation: the hidden seventh budget
A cloud team targets one GPU type. An Android team targets thousands of device models, several chipset vendors, multiple NPU generations, different driver versions and OS versions, and devices with 3 GB to 24 GB of RAM. A model that is fully accelerated on one phone may fall back to the CPU on another. This is why device tiering, capability checks and a reliable CPU fallback path are part of every production design.
Worked example: does this feature fit?
Feature: live background blur in video calls, 30 fps, mid-range phone
Budget per frame : 33 ms total; leave ~10 ms for the model (camera, render, encode need the rest)
Model : segmentation net, 4 GFLOPs per frame (2 GMACs), 3 M parameters
Weights : 3 M x 1 byte (INT8) = 3 MB -> fits easily in memory and cache
Compute : 4 GOPs x 30 fps = 120 GOPS sustained
NPU (say 10 TOPS INT8 at ~20% real utilization = 2 TOPS effective)
Time per frame : 4e9 / 2e12 = 2 ms -> OK
Energy per frame : ~1.5 W while active x 2 ms = 3 mJ -> 90 mW at 30 fps (plus pre/post-processing)
CPU fallback (100 GOPS effective): 40 ms per frame -> misses 30 fps -> needs a smaller model or lower resolution
Conclusion : ship on NPU devices; for CPU-only tiers use a half-resolution variant at 15 fps
The hardware landscape
A modern phone system-on-chip (SoC) has several compute units that can run neural networks, trading flexibility for efficiency. Beyond phones, edge AI runs on embedded GPU modules, USB and M.2 accelerators, PC NPUs and microcontrollers. Picking the right unit, and keeping the model on it without falling back, is a core edge skill.
Think of the SoC as a restaurant kitchen. The CPU is a skilled chef who can cook anything but slowly; the GPU is a line of cooks who are fast at repetitive dishes; the NPU is a specialized machine that makes one kind of dish (matrix maths) incredibly fast and cheaply, but cannot make anything off-menu; the DSP is a small always-on station for simple snacks. DRAM is the storeroom down the hall and on-chip SRAM is the counter next to the stove. If cooks keep walking to the storeroom (memory-bound), a faster stove (more TOPS) does not help, which is exactly why memory bandwidth and on-chip reuse often decide real performance.
| Unit | Strengths | Weaknesses | Good for |
|---|---|---|---|
| CPU (Arm big.LITTLE cores with NEON/SVE SIMD, dot-product and matrix extensions) | Supports every operator; easy to debug; low startup cost; good for small models; excellent optimized libraries (XNNPACK, KleidiAI) | Lowest throughput per watt for large tensor maths; competes with the UI thread | Small models, unsupported ops, fallback, pre/post-processing, LLM decode on some devices |
| GPU (mobile GPU via OpenCL, Vulkan, OpenGL compute or Metal) | Massively parallel; good FP16 throughput; widely available; flexible | Shader compilation/initialization time; less efficient at INT8 on many chips; shared with rendering | Vision models, FP16 models, image effects, LLMs via MLC-LLM or llama.cpp GPU backends |
| NPU (neural processing unit) | Fixed-function or highly specialized matrix/tensor engines; best performance per watt for INT8/INT4/FP16 | Limited operator set; vendor-specific toolchains; prefers static shapes and static quantization; unsupported ops cause fallback | Quantized CNNs, transformers, on-device LLM prefill and decode |
| DSP (digital signal processor) | Very power-efficient, always-on capable, good at fixed-point vector maths | Limited memory and programmability | Audio, keyword spotting, sensor models, always-on tasks |
| ISP (image signal processor) and sensor hub | Processes camera or sensor data at very low power before the main SoC wakes | Fixed pipelines; limited ML capability | Pre-processing, always-on sensing, low-power triggers |
CPU: SIMD and matrix extensions
Arm CPUs process many values per instruction using SIMD (single instruction, multiple data). NEON registers are 128 bits wide, so one instruction can operate on 16 INT8 or 8 FP16 values. The dot-product extension (SDOT/UDOT) multiplies four INT8 pairs and accumulates into an INT32 lane in one instruction; the INT8 matrix-multiply extension (SMMLA, "i8mm") computes a small 2x8 by 8x2 INT8 matrix product per instruction. SVE/SVE2 make vector length hardware-dependent, and the newest cores add SME/SME2 (Scalable Matrix Extension) with outer-product engines. Libraries such as XNNPACK, KleidiAI (Arm's low-bit matmul micro-kernels), and llama.cpp's hand-written kernels exploit these; enabling the right kernels can change CPU LLM prefill speed by 20% or more. x86 edge devices use AVX2, AVX-512 and VNNI/AMX for the same purpose.
Mobile GPU
Mobile GPUs (Adreno, Mali, Apple GPU, Immortalis, Xclipse) are wide SIMD machines designed for graphics. They excel at FP16 convolutions and element-wise work on large tensors. Their costs: kernels must be compiled at first use (shader compilation can take hundreds of milliseconds to seconds unless cached), the GPU is shared with UI rendering and the camera preview, and on many chips INT8 throughput is not much higher than FP16. They are a strong middle ground when an NPU is unavailable or the model uses operators the NPU lacks.
Inside a typical mobile NPU
Vendors design NPUs differently, but many combine three kinds of units. Qualcomm's Hexagon NPU is a well-known example: a scalar unit for control flow, a wide vector unit (HVX, Hexagon Vector eXtensions) for SIMD maths, and a tensor unit (HMX, Hexagon Matrix eXtensions) for dense matrix multiplication, together often called the HTP (Hexagon Tensor Processor). Other vendors (MediaTek APU, Samsung NPU, Google Tensor TPU, Apple Neural Engine) follow similar ideas: a large on-chip SRAM (tightly coupled memory) so data does not keep going out to DRAM, and hardware for low-precision multiply-accumulate (MAC) operations arranged as systolic arrays or dot-product engines.
+---------------------------- SoC ----------------------------+
| |
App / Runtime | CPU cluster GPU NPU (e.g. Hexagon) |
(LiteRT, ORT, | big + LITTLE shader scalar | vector | tensor |
ExecuTorch) | NEON/SVE/SME cores ctrl | HVX | HMX |
| | | | \ | / |
| | +-------+-------+----------------+ on-chip SRAM (TCM) |
| | | |
+-------->| System cache / interconnect |
| | |
+--------------+----------------------------------------------+
|
LPDDR DRAM (shared by everything, ~50-100 GB/s)
Why an NPU is efficient
- Low-precision MAC arrays: an INT8 multiplier is several times smaller and cheaper in energy than an FP32 one, so far more fit in the same area and power.
- Data reuse in SRAM: a systolic array passes each loaded weight and activation through many MACs before discarding it, so each DRAM byte feeds many operations.
- Less control overhead: a CPU spends most of its energy fetching, decoding and scheduling instructions; an NPU runs a fixed dataflow with coarse-grained commands.
- Ahead-of-time compilation: static shapes and static quantization let the compiler plan tiling, memory and scheduling offline, which is also why NPUs dislike dynamic shapes and runtime-computed scales.
Platform-specific accelerators
Apple Neural Engine (ANE)
A dedicated NPU in Apple's A- and M-series chips, reached only through Core ML (and frameworks on top of it). Prefers FP16 and specific tensor layouts; the framework decides per layer whether to use CPU, GPU or ANE. Apple's unified memory lets CPU, GPU and ANE share weights without copies.
Google Tensor TPU
The on-device TPU in Pixel phones runs Google's own models (for example Gemini Nano through the system AICore service, camera and speech features). Third-party access is mainly through platform APIs and LiteRT paths rather than direct programming.
Qualcomm Hexagon / MediaTek APU / Samsung NPU
Android flagship NPUs with vendor SDKs (QNN/QAIRT, NeuroPilot, Samsung's SDK) and delegates for LiteRT, ONNX Runtime and ExecuTorch. Tens of INT8 TOPS, INT4 and FP16 support, and increasingly optimized LLM paths.
PC NPUs
Laptop chips from Qualcomm, Intel and AMD now include NPUs of 40+ TOPS for local AI features, programmed through ONNX Runtime execution providers, OpenVINO, DirectML/Windows ML or vendor SDKs.
Jetson-class edge GPUs
NVIDIA Jetson modules (Orin Nano to AGX Orin, and newer Thor for robotics) combine Arm CPUs, an Ampere-or-newer GPU with tensor cores, and deep-learning accelerators (DLA), sharing LPDDR5 memory (roughly 68-205 GB/s on Orin). Programmed with CUDA, TensorRT and DeepStream; power modes from ~7 W to 60 W. The standard choice for robots, drones and multi-camera analytics.
USB/M.2 accelerators
Small add-on accelerators such as the Coral Edge TPU (about 4 INT8 TOPS at roughly 2 W, INT8-only, needs a compiled LiteRT model) or Hailo-class modules add inference to a Raspberry Pi-class host or an existing camera box.
Microcontrollers and micro-NPUs
Cortex-M4/M7/M33 cores with DSP instructions, Cortex-M55/M85 with Helium (M-profile vector extension), optionally paired with an Ethos-U55/U65/U85 micro-NPU. Kilobytes of SRAM, milliwatts of power, no operating system or a small RTOS. Programmed with LiteRT for Microcontrollers (formerly TFLite Micro), CMSIS-NN, or vendor compilers.
FPGAs and custom ASICs
Used in cameras, industrial and automotive systems where volume or latency justifies custom silicon: deterministic latency and excellent efficiency, at high engineering cost.
TOPS versus reality
Vendors quote peak TOPS (trillions of operations per second), usually for INT8 (sometimes INT4 or with sparsity, which doubles the number) with ideal utilization, counting a multiply-accumulate as two operations. Real models rarely reach that because of:
- Memory bandwidth. If weights or activations must stream from DRAM, the compute units wait. LPDDR5/5X bandwidth on phones is roughly 50-100 GB/s, far below desktop and data-centre GPUs (1-8 TB/s).
- Unsupported operators. An op the NPU cannot run forces a round trip to the CPU, which adds copy and synchronization time.
- Small batch sizes. On device, batch size is usually 1, so there is little data reuse to hide memory latency.
- Poor mapping. Depthwise convolutions, small channel counts, softmax, normalization and reshapes do not fill a large MAC array.
- Thermal throttling. Peak clocks last seconds to minutes; sustained performance is what users feel.
- Marketing arithmetic. Check whether the quoted figure is INT4 or INT8, dense or sparse, and whether it sums CPU, GPU and NPU together.
The roofline model: compute-bound versus memory-bound
The roofline model is the single most useful mental tool in edge AI. It tells you, before writing any code, whether a layer will be limited by arithmetic or by memory traffic, and therefore which optimization will help.
Picture a pizza kitchen with a very fast oven and one delivery van bringing ingredients from a warehouse. If each delivery contains enough dough for many pizzas, the oven is the bottleneck (compute-bound). If each delivery holds dough for just one pizza, the oven sits idle waiting for the van (memory-bound), and buying a bigger oven changes nothing. Arithmetic intensity is "pizzas per delivery"; the roofline tells you whether to buy a bigger oven (more TOPS) or pack the van more densely (quantization, fusion, reuse).
The two numbers you need
Arithmetic intensity (AI) is the number of operations performed per byte moved to or from DRAM. The ridge point of a machine is its peak compute divided by its bandwidth: the intensity at which a workload stops being memory-bound.
attainable
performance | ______________ peak compute (TOPS)
| /
| / compute-bound region
| /
| / memory-bound region
| / (slope = memory bandwidth)
|_____/_______________________________
arithmetic intensity (ops per byte)
LLM decode: far left (memory-bound) big conv / prefill: right (compute-bound)
Worked example 1: a phone NPU
Machine: NPU with 40 TOPS INT8 peak, 60 GB/s effective DRAM bandwidth
Ridge point = 40e12 / 60e9 = ~667 ops per byte
Layer A: matrix-VECTOR multiply, W is 4096 x 4096 INT8 (LLM decode, batch 1)
ops = 2 x 4096 x 4096 = 33.6 M ops
bytes = 4096 x 4096 x 1 byte = 16.8 MB (weights dominate; the vector is tiny)
AI = 33.6 M / 16.8 M = 2 ops/byte (far below 667 -> memory-bound)
attainable = 60 GB/s x 2 = 120 GOPS = 0.3% of peak
time = 16.8 MB / 60 GB/s = 0.28 ms
Layer B: same weights, but N = 512 tokens at once (LLM prefill)
ops = 2 x 4096 x 4096 x 512 = 17.2 G ops
bytes ~ 16.8 MB weights + 2 x (512 x 4096) activations ~ 21 MB
AI ~ 17.2 G / 21 M = ~820 ops/byte (above the ridge -> compute-bound)
time ~ 17.2e9 / (40e12 x 0.5 utilization) = 0.86 ms for 512 tokens (1.7 us/token)
Same weights, same chip: per token, prefill is ~165x cheaper than decode.
Worked example 2: what quantization does on the roofline
Halving bytes per weight doubles arithmetic intensity. For a memory-bound layer, that moves the point up the sloped line and nearly doubles speed. For a compute-bound layer, lower precision helps only if the hardware has faster low-precision units (for example INT8 MACs that run at twice the FP16 rate).
| Weights of a 4096x4096 layer | Bytes | AI at batch 1 | Decode time on 60 GB/s |
|---|---|---|---|
| FP32 | 67 MB | 0.5 ops/byte | 1.12 ms |
| FP16 | 33.6 MB | 1 op/byte | 0.56 ms |
| INT8 | 16.8 MB | 2 ops/byte | 0.28 ms |
| INT4 (+ group scales, ~4.5 bits) | ~9.4 MB | ~3.6 ops/byte | ~0.16 ms |
Arithmetic intensity of common layers
| Layer type | Intensity | Usually | Lever |
|---|---|---|---|
| Large 3x3 convolution with many channels | High (weights reused across every pixel) | Compute-bound | Faster MAC units, lower-precision compute, Winograd |
| Depthwise convolution | Low (one filter per channel, little reuse) | Memory-bound, poorly mapped on MAC arrays | Fusion with neighbours, keep tiles in SRAM |
| Matrix-matrix multiply (batch or many tokens) | Grows with batch/tokens | Compute-bound past the ridge | NPU tensor units |
| Matrix-vector multiply (batch 1 decode) | ~2 ops/byte at INT8 | Memory-bound | Fewer bytes: quantization, weight sharing |
| Element-wise ops (add, activation, normalization) | < 1 op/byte | Memory-bound | Fuse into the previous kernel so the data never leaves SRAM |
| Attention over a long KV cache during decode | Low | Memory-bound | GQA, KV-cache quantization, fused attention kernels |
Model basics for systems engineers
You do not need to train models to deploy them, but you must understand their shapes, costs and sensitive parts. A few families dominate edge work, and a handful of formulas let you estimate size, compute and memory in your head. For a full treatment of the architectures themselves, see Deep Learning and Transformers and LLMs.
A model is like a very large recipe book (the parameters) plus a set of cooking steps (the layers). Parameters decide how heavy the book is to carry (storage and RAM), and FLOPs decide how many chopping and stirring actions each dish needs (compute). A CNN is a recipe that repeats the same small stamp across a tray of cookies; a transformer is a dinner party where every guest must talk to every other guest before deciding what to say, which is why long conversations (long context) get expensive fast.
CNNs (convolutional neural networks)
Slide small filters over an image. Efficient mobile variants (MobileNet, EfficientNet-Lite, MobileNetV4) use depthwise-separable convolutions and inverted residual blocks to cut compute. Used for classification, detection (YOLO-style, SSD), segmentation and image enhancement. Very NPU-friendly.
Transformers
Use attention: each token looks at every other token to decide what matters. Used for text, speech (Whisper-style encoders/decoders) and vision (ViT, efficient hybrids such as MobileViT and EfficientFormer). Attention cost grows with the square of sequence length; layer norms, softmax and GELU can be harder for NPUs.
LLMs (large language models)
Decoder-only transformers that generate one token at a time. On device: roughly 0.5B-4B parameters (for example small Gemma, Llama, Phi, Qwen, SmolLM variants). Size is dominated by weights; runtime memory also includes the KV cache.
Small specialist models
Keyword spotters, activity classifiers, anomaly detectors, recommendation rankers. Often under 1 MB and run on DSPs or microcontrollers (TinyML).
Recurrent and state-space models
LSTMs/GRUs still power many streaming audio and sensor models; newer state-space models (Mamba-style) keep a fixed-size state instead of a growing KV cache, which is attractive for long streams on small devices.
Multimodal models
A vision or audio encoder feeding a small language model (vision-language models). On device, the encoder often runs on the NPU at fixed resolution and its output tokens are prefilled into the LLM.
Parameters, size and FLOPs
- Parameters are the learned weights. Model file size is roughly
parameters × bytes per parameter. A 3B-parameter model is about 12 GB in FP32, 6 GB in FP16, 3 GB in INT8 and about 1.5-1.8 GB in INT4 (group scales add a little overhead). - FLOPs (floating-point operations) measure compute. A dense layer with
Ninputs andMoutputs costs about2 × N × MFLOPs per sample (one multiply plus one add per weight). - A convolution costs about
2 × K × K × C_in × C_out × H_out × W_outFLOPs. Depthwise-separable convolution splits this into a cheap per-channel filter plus a 1x1 convolution, often cutting FLOPs by 8-9x for 3x3 kernels. - A transformer forward pass costs about
2 × parametersFLOPs per token, plus attention cost that grows with context length. - MACs (multiply-accumulates) are often quoted instead of FLOPs; 1 MAC is roughly 2 FLOPs. Always check which one a paper or vendor uses.
- Activations are intermediate tensors. Peak activation memory matters on small devices, sometimes more than weight size; on MCUs it is often the binding constraint.
Worked examples
1) Convolution cost
3x3 conv, 64 -> 128 channels, 56x56 output
MACs = 3 x 3 x 64 x 128 x 56 x 56 = 231 M MACs (~462 M FLOPs)
Depthwise-separable version:
depthwise 3x3 x 64 x 56 x 56 = 1.8 M MACs
pointwise 64 x 128 x 56 x 56 = 25.7 M MACs
total = 27.5 M MACs -> ~8.4x cheaper
2) Small transformer size
d = 2048, 16 layers, vocabulary 128k
layers = 12 x 2048^2 x 16 = ~805 M parameters
embeddings = 128,000 x 2048 = ~262 M parameters (often tied with the LM head)
total = ~1.07 B -> ~2.1 GB FP16, ~0.6 GB INT4
Note: in a 1B model the embedding table is ~25% of all weights, which is why
embedding quantization matters so much for small LLMs.
3) Per-token compute
1B model: ~2 GFLOPs per token. A 500-token prompt = ~1 TFLOP of prefill work.
Training versus inference
Training adjusts weights using large datasets, gradients and a lot of memory (weights, gradients, optimizer state and saved activations, often 16+ bytes per parameter); it happens in data centres. Inference only runs the forward pass with fixed weights. Edge work is almost entirely inference, with occasional light on-device adaptation (for example small adapters or federated learning updates).
Transformer building blocks worth knowing
- Tokens and embeddings: text is split into tokens; each token maps to a vector (embedding). The vocabulary embedding table can be a large part of a small model.
- Attention: each token creates Query, Key and Value vectors. Scores are
softmax(Q × K^T / sqrt(d)), then multiplied by V. During generation only K and V of past tokens are needed again, which is why the cache holds K and V but not Q. - Multi-head, multi-query and grouped-query attention (MHA, MQA, GQA): GQA shares Key/Value heads across several Query heads, shrinking the KV cache, which is why most modern small LLMs use it.
- Positional encoding (RoPE): rotary embeddings rotate Q and K by position-dependent angles; context-extension tricks rescale those angles. RoPE is applied before K is cached.
- Feed-forward (MLP) blocks: hold most of the parameters; they are big matrix multiplies and quantize well.
- Normalization and activation functions (LayerNorm, RMSNorm, GELU, SiLU): cheap in FLOPs but numerically sensitive; often kept at higher precision.
- Mixture of experts (MoE): only a few expert MLPs run per token, so compute per token is small, but all experts must still be stored, which makes MoE memory-heavy for edge devices.
Number formats: FP32 to INT4, FP8 and block formats
Every tensor on an accelerator is stored in some number format, and the choice decides size, bandwidth, energy, which hardware units can be used and how much accuracy is lost. Floating-point formats split their bits between an exponent (range) and a mantissa (precision); integer formats have uniform steps and need an external scale.
A floating-point number is like writing a distance as "3.2 × 104 km": the exponent says the scale (street, city, continent) and the mantissa says the detail within that scale. More exponent bits let you describe both an ant and a galaxy (range); more mantissa bits let you describe the ant precisely (precision). An integer format is a ruler with evenly spaced ticks: great precision within its length, but you must choose the ruler's length (the scale) in advance, which is exactly what quantization calibration does.
Floating-point anatomy
| Format | Bits (sign/exponent/mantissa) | Max value | Precision | Where it is used |
|---|---|---|---|---|
| FP32 | 1 / 8 / 23 | ~3.4 × 1038 | ~7 decimal digits | Training default; reference accuracy; CPU fallback |
| FP16 (half) | 1 / 5 / 10 | 65,504 | ~3 decimal digits | Mobile GPUs, Apple Neural Engine, NPU activations; overflow risk |
| BF16 (bfloat16) | 1 / 8 / 7 | ~3.4 × 1038 | ~2-3 decimal digits | Training and data-centre inference; growing on NPUs and Arm CPUs; same range as FP32, no overflow worries |
| FP8 E4M3 | 1 / 4 / 3 | 448 | ~1 decimal digit | Inference weights and activations on newer accelerators |
| FP8 E5M2 | 1 / 5 / 2 | 57,344 | Coarser | Gradients in training (needs range more than precision) |
| FP4 E2M1 | 1 / 2 / 1 | 6 | 8 magnitudes: 0, 0.5, 1, 1.5, 2, 3, 4, 6 | Only inside block formats (MXFP4, NVFP4) with a shared scale |
| INT8 / UINT8 | 8-bit integer | 127 / 255 steps (times scale) | Uniform: step = scale | The workhorse for NPUs, DSPs and CPU integer kernels |
| INT16 | 16-bit integer | 32,767 steps | Uniform, fine | Activations on NPUs where INT8 is too coarse (for example W8A16, W4A16) |
| INT4 | 4-bit integer (-8..7) | 16 levels | Very coarse; needs small groups | LLM weights, mostly weight-only |
| NF4 (NormalFloat 4) | 4-bit lookup table | 16 non-uniform levels | Levels placed at quantiles of a normal distribution | QLoRA base weights; matches bell-shaped weight distributions |
| Binary / ternary | 1-2 bits | {-1, +1} or {-1, 0, +1} | Extreme | Research and some tiny models; ternary LLMs trained from scratch |
Worked example: the same number in different formats
Store 0.1:
FP32 -> 0.100000001490116... (error ~1.5e-9)
FP16 -> 0.0999755859375 (error ~2.4e-5)
BF16 -> 0.10009765625 (error ~9.8e-5) less precise than FP16...
Store 70,000:
FP16 -> inf (overflow, max is 65,504)
BF16 -> 70,144 (fine, just coarse) ...but BF16 never overflows where FP32 would not.
Store 1e-8:
FP16 -> 0 (underflow; smallest subnormal is ~6e-8)
BF16 -> ~1e-8 (fine)
Lesson: FP16 has better precision, BF16 has better range.
Activations with large values (attention logits, some LLM hidden states) can overflow FP16.
import numpy as np
print(np.float16(70000)) # inf -> overflow
print(np.float16(300) * np.float16(300)) # inf (90,000 > 65,504)
print(np.float16(0.1)) # 0.1 displayed, stored as 0.09998
print(np.float32(np.float16(0.1))) # 0.099975586
Block (microscaling) formats
A small group of values shares one scale, so each element can use very few bits while the group still covers a wide range. This is the same idea as per-group integer quantization, standardized for hardware.
| Format | Block size | Element type | Shared scale | Effective bits per value |
|---|---|---|---|---|
| MXFP8 / MXFP6 / MXFP4 / MXINT8 (OCP microscaling) | 32 | FP8, FP6, FP4 or INT8 | 8-bit power-of-two exponent (E8M0) | MXFP4: 4 + 8/32 = 4.25 |
| NVFP4 | 16 | FP4 E2M1 | FP8 E4M3 scale plus a per-tensor FP32 scale | 4 + 8/16 = 4.5 |
| GGUF Q4_0 (llama.cpp) | 32 | 4-bit integer | One FP16 scale | 4 + 16/32 = 4.5 |
| GGUF Q8_0 | 32 | 8-bit integer | One FP16 scale | 8.5 |
| GGUF k-quants (Q4_K and friends) | Super-block of 256 = 8 blocks of 32 | 4-bit integer | 6-bit per-block scales and mins, FP16 super-block scale | ~4.5 |
Accumulators and mixed formats
Products are accumulated in a wider format than the inputs: INT8 × INT8 products accumulate in INT32 (a sum of thousands of 16-bit products would overflow INT16), FP16 products usually accumulate in FP32 or FP16 depending on hardware, and FP8 products accumulate in FP32 or FP16. Notation such as W4A16 means 4-bit weights and 16-bit activations; W8A8 means both 8-bit; 8da4w means 8-bit dynamically quantized activations with 4-bit weights; KV8 or KV4 describes the KV-cache precision.
Quantization in depth
Quantization stores and computes numbers with fewer bits. It is the single most important optimization in edge AI: it shrinks the model, cuts memory bandwidth (the real bottleneck for many layers), saves energy, and unlocks integer engines on NPUs and DSPs. This section builds it up from the maths; the next section covers the special techniques for LLMs.
Quantization is like rounding prices to the nearest rupee or dollar instead of tracking fractions of a cent: the receipt is shorter and faster to add up, and the total is still almost right as long as you pick a sensible unit (the scale). If one item on the receipt costs a million while the rest cost a few rupees, a single "unit" that covers the million makes the small items round to zero, which is exactly the outlier problem, and why per-channel and per-group scales (a separate unit for each section of the receipt) help.
The affine (asymmetric) mapping
A real value x is mapped to an integer q using a scale (the step size between integer levels) and a zero-point (the integer that represents real zero exactly):
q = clamp(round(x / scale) + zero_point, q_min, q_max)
x_approx = scale * (q - zero_point)
// Example, asymmetric INT8 (q in 0..255) for a tensor ranging -1.0 .. 3.0
scale = (3.0 - (-1.0)) / 255 = 0.01569
zero_point = round(0 - (-1.0) / scale) = 64
x = 1.0 -> q = round(1.0 / 0.01569) + 64 = 128 -> x_approx = 0.01569 * (128 - 64) = 1.004
Symmetric mapping and a worked example
Symmetric quantization forces the zero-point to 0 and uses a range centred on zero: s = max|x| / 127 for INT8 (many toolchains use the restricted range -127..127 so the grid is symmetric). It removes the zero-point arithmetic from the inner loop, which is why it is the default for weights.
Weights: [-0.80, 0.30, 1.20, -0.05] symmetric INT8, range -127..127
scale = max|w| / 127 = 1.20 / 127 = 0.009449
w w / scale q (rounded) dequantized error
-0.80 -84.67 -85 -0.8031 0.0031
0.30 31.75 32 0.3024 0.0024
1.20 127.00 127 1.2000 0.0000
-0.05 -5.29 -5 -0.0472 0.0028
Max possible rounding error = scale / 2 = 0.0047
Now add one outlier weight, 12.0, to the same tensor:
scale = 12.0 / 127 = 0.0945 -> -0.05 becomes q = -1 -> -0.0945 (error 0.045, ~90%)
0.30 becomes q = 3 -> 0.2835
One outlier made every small weight 10x less precise. Per-channel or per-group
scales isolate the outlier to its own channel or group.
How much error does quantization add?
With round-to-nearest, the error for values inside the range is roughly uniform between -s/2 and +s/2, so its variance is s2/12. Values outside the range are clipped, which can create much larger errors. Every extra bit halves the step and adds about 6 dB of signal-to-quantization-noise ratio (SQNR).
Design choices
| Choice | Options | Trade-off |
|---|---|---|
| Symmetric vs asymmetric | Symmetric: zero_point = 0, range centred on zero. Asymmetric: any zero_point. | Symmetric is simpler and faster (common for weights); asymmetric uses the range better for skewed data (common for activations after ReLU). |
| Granularity | Per-tensor, per-channel, per-group (block), per-token | Finer granularity means more scales but better accuracy. Per-channel is standard for conv/linear weights; per-group (for example groups of 32-128 weights) is standard for INT4 LLM weights; per-token scales are used for dynamic activations. |
| What is quantized | Weight-only, or weights and activations, and optionally the KV cache | Weight-only (e.g. W4A16) shrinks memory and helps memory-bound decode; full integer (W8A8) is needed to use integer-only NPU paths. |
| Static vs dynamic | Static: activation ranges fixed from calibration. Dynamic: computed at runtime. | Static is fastest and NPU-friendly; dynamic is easier and more accurate for varying inputs but adds runtime overhead and is mostly CPU/GPU-only. |
| Rounding | Round-to-nearest, or learned rounding (AdaRound), or error-compensating (GPTQ) | Round-to-nearest is free; smarter rounding recovers accuracy at low bits using a little calibration data. |
| Grid | Uniform integer, floating-point (FP8/FP4), non-uniform lookup (NF4, k-means codebooks) | Uniform grids map to integer hardware; non-uniform grids fit bell-shaped distributions better but need a lookup (dequantization) before compute. |
Per-tensor, per-channel, per-group
Weight matrix W (rows = output channels, columns = input features) per-tensor per-channel (per row) per-group (row split into groups of g) +----------+ +----------+ s0 +----+----+----+ s00 s01 s02 | | |----------| s1 |----+----+----| s10 s11 s12 | one s | |----------| s2 |----+----+----| s20 s21 s22 | | |----------| s3 |----+----+----| s30 s31 s32 +----------+ +----------+ +----+----+----+ 1 scale one scale per row one scale per g weights (g = 32..128) cheapest standard for INT8 standard for INT4 LLM weights
Per-channel scales for weights cost nothing at runtime: because each output channel's dot product uses only its own weights, the channel scale can be applied once to the accumulator. Per-channel scales for activations along the reduction dimension are a different story: they cannot be factored out of the dot product, which is why activation outliers are hard and why SmoothQuant moves them into the weights.
Calibration: choosing the range
For static quantization, a calibration set (typically 100-1000 representative samples) is run through the model and statistics are collected per tensor. The range-setting method matters as much as the data.
| Method | How it picks the range | Good for | Risk |
|---|---|---|---|
| Min-max | Observed minimum and maximum | Weights, well-behaved activations | One outlier wastes most of the range |
| Moving-average min-max | Exponential average across batches | Smoother than raw min-max | Still outlier-sensitive |
| Percentile | Clip at, say, the 99.9th or 99.99th percentile | Long-tailed activations | Clips rare but important values |
| MSE-minimizing | Search the clipping threshold that minimizes quantization error | General default in many toolkits | Slower calibration |
| Entropy / KL divergence | Choose the threshold whose quantized histogram best matches the original distribution | Classic TensorRT INT8 calibration for CNNs | Can under-clip or over-clip on unusual distributions |
Static versus dynamic quantization
Static
- Activation scales fixed at conversion time from calibration
- No runtime statistics; integer-only pipelines possible
- Required by most NPUs (compile-time parameters)
- Accuracy depends on calibration data matching production
Dynamic
- Weights pre-quantized; activation scale computed per input (often per token) at runtime
- No calibration needed; adapts to input range
- Extra reduction pass per tensor adds overhead
- Common on CPU (for example 8-bit dynamic activations with 4-bit weights for LLMs)
Integer matrix multiplication with scales
For y = W x with W ≈ sw(qw - zw) and x ≈ sx(qx - zx), the expensive inner loop is pure integer arithmetic, and all the floating-point scale work happens once per output value.
Weight-only schemes (W4A16) work differently: the kernel loads 4-bit weights, dequantizes them to FP16 in registers using the group scale, and does FP16 maths. The win is bandwidth (fewer bytes loaded), not integer compute, which is exactly what memory-bound decode needs.
PTQ versus QAT
Post-training quantization (PTQ)
- Quantize an already-trained model
- Needs a small calibration set (typically 100-1000 representative samples) to measure activation ranges
- Fast, no training pipeline needed
- Usually fine for INT8 on CNNs; may lose too much accuracy at INT4 or on sensitive models
- Advanced PTQ methods (AdaRound, GPTQ, AWQ, SmoothQuant, rotations) reduce the loss
Quantization-aware training (QAT)
- Insert "fake quantization" nodes during (re)training so the model learns to tolerate rounding
- Needs training data, compute and a training pipeline
- Recovers most accuracy at low bit-widths
- Use when PTQ loses more accuracy than the product allows
- Common practice: try PTQ first, move to QAT only if needed; QAT combined with LoRA fine-tuning is a popular cheap variant for LLMs
QAT needs gradients through the rounding function, whose true derivative is zero almost everywhere. The straight-through estimator (STE) pretends rounding is the identity in the backward pass (with zero gradient outside the clipping range). Methods such as LSQ (learned step size quantization) also learn the scales themselves.
Code: quantization from scratch
import numpy as np
def quantize(x, bits=8, symmetric=True, axis=None, group=None):
"""Return integer tensor q, scale s, zero-point z (fake-quant helpers)."""
if group is not None: # per-group along the last axis
x = x.reshape(*x.shape[:-1], -1, group)
axis = -1
reduce = dict(axis=axis, keepdims=True) if axis is not None else {}
if symmetric:
qmax = 2 ** (bits - 1) - 1 # 127 for INT8, 7 for INT4
s = np.abs(x).max(**reduce) / qmax
z = 0
q = np.clip(np.round(x / s), -qmax, qmax)
else:
qmin, qmax = 0, 2 ** bits - 1 # 0..255 for UINT8
lo, hi = x.min(**reduce), x.max(**reduce)
lo, hi = np.minimum(lo, 0), np.maximum(hi, 0) # zero must be representable
s = (hi - lo) / (qmax - qmin)
z = np.round(qmin - lo / s)
q = np.clip(np.round(x / s) + z, qmin, qmax)
return q, s, z
def sqnr_db(x, x_hat):
return 10 * np.log10((x ** 2).sum() / ((x - x_hat) ** 2).sum())
np.random.seed(0)
W = np.random.randn(256, 1024).astype(np.float32)
W[3, 7] = 40.0 # one outlier
for bits in (4, 8):
for name, kw in [("per-tensor", {}), ("per-channel", {"axis": 1}),
("per-group-32", {"group": 32})]:
q, s, z = quantize(W, bits=bits, **kw)
W_hat = (s * (q - z)).reshape(W.shape)
print(f"INT{bits} {name:13s} SQNR = {sqnr_db(W, W_hat):5.1f} dB")
# Output:
# INT4 per-tensor SQNR = 0.1 dB (scale 40/7 = 5.7: almost every weight rounds to 0)
# INT4 per-channel SQNR = 16.2 dB
# INT4 per-group-32 SQNR = 20.2 dB
# INT8 per-tensor SQNR = 20.8 dB
# INT8 per-channel SQNR = 40.3 dB (~6 dB per extra bit: 16.2 + 4 x 6 = ~40)
# INT8 per-group-32 SQNR = 45.3 dB
Code: PyTorch 2 export-based PTQ (sketch)
import torch
from torch.export import export
from torch.ao.quantization.quantize_pt2e import prepare_pt2e, convert_pt2e
# Module paths move between releases; in recent versions the quantizer lives in
# executorch.backends.xnnpack.quantizer.xnnpack_quantizer
from executorch.backends.xnnpack.quantizer.xnnpack_quantizer import (
XNNPACKQuantizer, get_symmetric_quantization_config)
model = MyModel().eval()
example = (torch.randn(1, 3, 224, 224),)
graph = export(model, example).module() # capture the graph
quantizer = XNNPACKQuantizer().set_global(
get_symmetric_quantization_config(is_per_channel=True))
prepared = prepare_pt2e(graph, quantizer) # inserts observers
with torch.no_grad():
for batch in calibration_loader: # 100-1000 production-like samples
prepared(batch)
quantized = convert_pt2e(prepared) # observers -> quantize/dequantize ops
# then lower `quantized` to a backend (see the runtimes section)
Why quantization hurts accuracy, and how to fix it
- Outliers: a few very large activation values stretch the range, so normal values lose resolution. Fixes: per-channel scales, clipping with percentile or MSE calibration, SmoothQuant (moves difficulty from activations to weights), rotations (QuaRot/SpinQuant), keeping outlier layers in higher precision.
- Sensitive layers: first and last layers, attention softmax, layer norms and the embedding/LM head are often more sensitive. Use mixed precision: keep them at INT16/FP16.
- Bad calibration data: if calibration images are all daytime, night images will clip. Calibrate on data that matches production.
- Error accumulation: errors add up through deep networks. Measure per-layer error (for example SQNR, cosine similarity, mean and max absolute error) between an FP32 reference run and the quantized run to find the layer that breaks.
- Uneven channel ranges: cross-layer equalization rescales consecutive layers (using the fact that ReLU is scale-equivariant) so per-channel ranges become similar without changing the FP32 output; bias correction removes the systematic mean shift that quantization error introduces.
- Rounding choices: AdaRound learns, per weight, whether to round up or down to minimize each layer's output error, using a small calibration set.
Layer-wise error analysis: the debugging workflow
- Build a reference Run the FP32 model on the host and dump the output of every layer for a fixed set of inputs.
- Capture the quantized run Run the quantized model (on host simulation and, ideally, on the device) and dump the same tensors, dequantized.
- Diff per layer Compute SQNR, cosine similarity, mean and max absolute error for each tensor.
- Rank and localize Plot error against depth; look for the layer where error jumps rather than grows smoothly. Separate "this layer is sensitive" (large local error) from "error accumulated" (small local error on a damaged input).
- Fix with evidence Keep the worst layers at higher precision, change granularity, or apply an outlier technique; re-measure accuracy and latency, since each higher-precision island may create a new partition on the NPU.
Quantizing LLMs: GPTQ, AWQ, SmoothQuant, rotations and GGUF
Plain round-to-nearest (RTN) INT8 works well for most CNNs. LLMs are harder at low bit-widths for two reasons: we want to go to 4 bits or below to fit memory and bandwidth budgets, and transformer activations contain a few channels with values tens to hundreds of times larger than the rest. A family of post-training methods addresses this with only a small calibration set and no full retraining.
Imagine compressing a choir recording where a few soloists are far louder than everyone else. RTN turns the volume knob to fit the loudest soloist, so the choir becomes a hiss. GPTQ is a sound engineer who, after rounding each singer, nudges the remaining singers to compensate for the error; AWQ turns up the microphones of the few singers who matter most before compressing; SmoothQuant moves loudness from the singers (activations) into the mixing desk (weights), which can take it; rotation methods remix the channels so that no single one is louder than the others. Each maps to a different way of keeping the important signal above the rounding noise.
Why LLM activations are hard, weights are easy
- Weights are roughly bell-shaped per channel with few outliers; with small groups (32-128) even 4 bits keep good accuracy.
- Activations have a handful of fixed channels with huge magnitudes that appear for almost every token. Per-tensor INT8 activations then waste nearly all of their range on those channels. Because the outliers lie along the reduction dimension, per-channel activation scales cannot be factored out of an integer matmul.
- Decode is memory-bound, so weight-only quantization (W4A16, or W4 with 8-bit dynamic activations) captures most of the speed benefit without touching activations. Full W8A8 or W4A8 matters mostly for compute-bound prefill and for integer-only NPUs.
- The embedding table and LM head are large in small models and unusually sensitive; they are often kept at 6-8 bits even when the body is at 4 bits.
GPTQ: error-compensating rounding
GPTQ quantizes a layer's weight matrix one column (input feature) at a time. After rounding a column, it updates the not-yet-quantized columns to cancel the error that rounding introduced in the layer output, using second-order information from calibration activations.
Practical notes: ~128 calibration sequences are typical; it runs layer by layer so memory stays manageable; a 7B model takes minutes to an hour on a single data-centre GPU. It works best combined with group-wise scales (for example INT4, group size 128).
AWQ: protect the salient channels
Activation-aware weight quantization observes that a small fraction (around 1%) of weight input channels matter most, namely those multiplied by large activations. Instead of keeping them in FP16 (which would complicate kernels), AWQ multiplies those weight channels by a factor s > 1 before quantization and divides the matching activations by s (folded into the previous layer or normalization). The product is unchanged, but the important weights now use more of the integer grid, so their relative rounding error shrinks.
SmoothQuant: move the difficulty into the weights
SmoothQuant targets W8A8 by dividing each activation channel by a smoothing factor and multiplying the corresponding weight row by the same factor. This is mathematically exact in full precision and makes both sides easy to quantize.
Rotation methods: QuaRot and SpinQuant
If you multiply activations by an orthogonal matrix R and the next weights by RT, the output is unchanged (R RT = I), but the outlier energy gets spread across all channels, so no single channel dominates the range. QuaRot uses random Hadamard rotations (cheap to apply with a fast transform); SpinQuant learns the rotations on calibration data. Rotations make 4-bit weights, 4- or 8-bit activations and 4-bit KV caches workable, and they were used, together with QAT plus LoRA, for official quantized mobile releases of small open LLMs, which reached roughly 2.5x faster decode and 50-60% smaller models than BF16 on phone CPUs.
GGUF and k-quants (llama.cpp)
GGUF is llama.cpp's single-file format holding tensors, tokenizer, chat template and metadata. Its quantization types are block formats, named by approximate bits and scheme.
| Type | Scheme | Approx. bits per weight | Typical use |
|---|---|---|---|
| F16 / BF16 | Unquantized | 16 | Reference, conversion source |
| Q8_0 | Blocks of 32, one FP16 scale | 8.5 | Near-lossless; baseline for quality checks |
| Q6_K | k-quant, 6-bit | ~6.6 | Very close to FP16; used for sensitive tensors in mixes |
| Q5_K_M | Mix of 5- and 6-bit k-quants | ~5.7 | Quality-leaning balance |
| Q4_K_M | Mostly Q4_K, some tensors (for example attention V and part of the FFN down projections) at Q6_K | ~4.8 | The popular default balance of size and quality |
| Q4_0 | Blocks of 32, one FP16 scale, symmetric | 4.5 | Simple and fast; has special fast paths (for example repacked Arm kernels) |
| Q3_K_M, Q2_K | Lower-bit k-quants | ~3.9, ~2.6-3.4 | Squeezing larger models into small memory; visible quality loss |
| IQ4_XS, IQ3_XXS, IQ2_XXS | "i-quants": lattice/codebook-based, best with an importance matrix | ~4.25, ~3.1, ~2.1 | Best quality per bit at very low bit-widths; slower on some CPUs |
K-quants organize 256 weights into a super-block of 8 sub-blocks of 32, with small quantized per-sub-block scales (and minimums) plus one FP16 super-block scale: two-level scaling keeps metadata overhead low. An importance matrix (imatrix), computed from calibration text, weights the rounding error by how much each weight matters, similar in spirit to AWQ.
KV-cache quantization
At long context the KV cache can outgrow the weights, so quantizing it pays off directly in memory and decode bandwidth. INT8 KV is close to lossless; 4-bit needs care; research shows usable results near 2-3 bits with the right grouping. Keys and values behave differently: keys have outlier channels (so they quantize better per channel), while values are better quantized per token. Recent tokens are often kept in higher precision in a small residual buffer.
Precision by target: what each runtime typically expects
| Target | Weights | Activations | Notes |
|---|---|---|---|
| CPU/GPU (ExecuTorch XNNPACK, ONNX Runtime, LiteRT) | 4-bit group-wise (group 32 or 128) | 8-bit dynamic, scales computed at runtime per token | "8da4w"; embeddings often 4-bit with small groups; LM head higher |
| Mobile NPU (for example Hexagon via QNN/QAIRT) | 4-bit or 8-bit, per-channel or block | 16-bit static, fixed at compile time | NPUs need compile-time parameters, so dynamic activation quantization is not available; W4A16 is a common LLM recipe |
| llama.cpp | GGUF block formats (Q4_0, Q4_K_M, ...) | Backend-dependent, often dynamically quantized to 8-bit for integer dot products | Output/LM head often kept at 6-bit or higher |
| Apple (Core ML) | Palettized (lookup table) or linear 4/8-bit, per-block | FP16 | Weight compression mainly to cut memory and bandwidth |
| Jetson/TensorRT | INT8, FP8 or INT4 (AWQ-style) depending on generation | FP16, INT8 or FP8 | Calibration built into the TensorRT builder or model-optimizer tooling |
INT8 versus INT4: what interviewers expect
| INT8 weights (or FP8) | INT4 weight-only | |
|---|---|---|
| Size vs FP16 | ~2× smaller | ~4× smaller (group scales → ~4.1-4.8 effective bits) |
| CNN PTQ | Usually under 1% top-1 drop with per-channel scales and a decent calibration set | Rarely used; accuracy loss is larger unless you QAT |
| LLM quality | Near-lossless for most tasks after simple PTQ | Small perplexity rise with GPTQ/AWQ/k-quants; visible drops on reasoning, maths, rare languages and long context |
| Decode speed | ~2× vs FP16 if bandwidth-bound | ~4× vs FP16 if bandwidth-bound (the usual on-device case) |
| NPU path | Native integer engines (W8A8 for CNNs) | Often W4A16: dequantize in registers, compute in 16-bit |
| What to keep higher | Usually nothing | Embeddings, LM head, first/last blocks, sometimes attention V and MLP down-proj |
Treat "INT4" as a family of recipes, not one format. Group size, which tensors stayed at 6-8 bits, RTN versus GPTQ/AWQ, and whether activations are dynamic 8-bit or static 16-bit can move quality more than the nominal bit-width. Always state the full recipe and re-measure the product task.
Evaluating a quantized LLM
- Perplexity on held-out text is cheap and sensitive, but small perplexity changes do not always predict task quality.
- Task benchmarks relevant to the feature (summarization quality, extraction accuracy, instruction following) catch behavioural regressions.
- Agreement with the reference: KL divergence of next-token distributions or top-1 agreement versus the FP16 model on the same prompts.
- Long-context and edge cases: quantization damage often shows first on long inputs, rare languages, maths and formatting.
# llama.cpp: build an importance matrix, quantize with it, and compare perplexity
./llama-imatrix -m model-f16.gguf -f calibration.txt -o imatrix.dat
./llama-quantize --imatrix imatrix.dat model-f16.gguf model-IQ4_XS.gguf IQ4_XS
./llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M
./llama-perplexity -m model-Q4_K_M.gguf -f wiki.test.raw # compare against the F16 run
Beyond quantization: pruning, sparsity, distillation, low-rank and NAS
A model trained in FP32 on a data-centre GPU is almost never shipped as-is. Quantization reduces bits per weight; the techniques here reduce the number of weights or operations, or produce a better small model in the first place. They stack: a distilled, pruned, quantized model is common.
Pruning is removing pages nobody reads from a manual; structured pruning removes whole chapters (the book really gets thinner), unstructured pruning blacks out random words (the book is the same size unless the printer knows to skip them). Distillation is an expert writing a short study guide for a student; low-rank factorization is replacing a giant lookup table with two small ones that multiply together; NAS is an architect trying hundreds of floor plans to find the one that fits a small plot. Each maps directly: fewer weights, a smaller model trained from a bigger one, thinner matrices, and an architecture designed for the hardware.
Pruning
Pruning removes weights that contribute little.
- Unstructured pruning zeroes individual weights (usually those with the smallest magnitude). It can remove 50-90% of weights with small accuracy loss, but most mobile hardware cannot skip scattered zeros, so it mainly helps compression (the zeros compress well on disk), not speed.
- Structured pruning removes whole channels, filters, attention heads or layers. The model genuinely becomes smaller and faster on any hardware, but accuracy drops more and usually needs fine-tuning. For LLMs, width pruning (hidden size, MLP size, heads) and depth pruning (whole layers) followed by distillation is a proven recipe for producing small models from larger ones.
- Semi-structured (N:M) sparsity, for example 2:4 (two zeros in every four weights), is supported by some accelerators and gives real speed-ups (up to about 2x for the matrix maths) because the pattern is regular enough for hardware to exploit.
- Iterative pruning (prune a little, fine-tune, repeat) keeps more accuracy than one-shot pruning. One-shot LLM methods such as SparseGPT and Wanda use calibration data to choose which weights to drop without retraining.
- Activation sparsity: after ReLU-like activations many values are zero; some runtimes and research systems skip the corresponding weight rows, which can reduce bytes read during decode.
import torch.nn.utils.prune as prune
# Unstructured: zero the 50% smallest-magnitude weights of a layer
prune.l1_unstructured(model.fc1, name="weight", amount=0.5)
# Structured: remove 25% of output channels (rows) by L2 norm
prune.ln_structured(model.conv2, name="weight", amount=0.25, n=2, dim=0)
prune.remove(model.fc1, "weight") # make the pruning permanent
# Note: masks only zero values; to get real speed from structured pruning you must
# rebuild the layer with fewer channels (and adjust the next layer's inputs).
Knowledge distillation
A large "teacher" model trains a small "student" model to match its outputs (soft probabilities), not just the hard labels. Soft outputs carry extra information ("this looks 70% like a cat, 25% like a fox"), so the student learns more from the same data. Most good small on-device LLMs are distilled from larger models, often after pruning.
Low-rank factorization
A weight matrix W of size m × n is approximated by the product of two thin matrices U (m × r) and V (r × n), usually from a truncated SVD followed by fine-tuning.
Adapters: LoRA and QLoRA
LoRA (low-rank adaptation) fine-tunes a model by learning small low-rank update matrices while freezing the base weights: W' = W + B A, with rank r of 4-64. On device, one base model can serve several features by swapping small LoRA adapters (a few MB to tens of MB each) instead of shipping several full models; OS-level AI services use exactly this to specialize one shared foundation model per feature. QLoRA fine-tunes LoRA adapters on top of a 4-bit (NF4) quantized base, cutting fine-tuning memory. Adapters can be merged into the weights for zero overhead, or kept separate for hot-swapping at a small compute cost. See Fine-Tuning for training details.
Neural architecture search and efficient architecture design
The biggest wins often come before any compression: pick a mobile-first architecture (MobileNet-style, efficient ViTs, small LLMs with GQA), lower input resolution, shorter context, fewer layers, or split into a small always-on model that triggers a larger one only when needed (cascade). Neural architecture search (NAS) automates this: it searches over block types, widths, kernel sizes and depths, scoring candidates by accuracy and by measured or predicted latency on the target hardware ("hardware-aware NAS"). MobileNetV3 and EfficientNet came from NAS; "once-for-all" approaches train one super-network and extract sub-networks sized for each device tier without retraining.
- Scaling knobs: width multiplier (channels), depth, input resolution; compute roughly scales with width squared and resolution squared.
- Early exit: add intermediate classifiers so easy inputs stop early.
- Weight sharing and clustering: store a small codebook of distinct values and an index per weight (palettization), popular on Apple devices.
- Matryoshka / nested models: one set of weights contains smaller sub-models that can be selected at runtime to fit the device or the moment's power budget.
| Technique | Size gain | Speed gain | Accuracy risk | Effort |
|---|---|---|---|---|
| FP32 to FP16 | 2x | Good on GPU | Very low | Low |
| INT8 PTQ | 4x | High on NPU/DSP | Low-medium | Low |
| INT4 weight-only (LLM) | ~7x | High for memory-bound decode | Medium | Medium |
| QAT | same as target bits | same | Lowest at low bits | High |
| Unstructured pruning | Only with sparse storage/compression | Usually none on mobile | Low-medium | Medium |
| 2:4 semi-structured sparsity | ~1.8x (with metadata) | Up to ~2x on supporting hardware | Medium | Medium-high |
| Structured pruning | 1.5-3x | Real | Medium | High (fine-tune) |
| Distillation | Large (new small model) | Large | Depends on student | High (training) |
| Low-rank factorization | 2-4x on factorized layers | Real if rank is small | Medium | Medium (fine-tune) |
| Hardware-aware NAS | Varies | Large | Low (searched for accuracy) | Very high (compute) |
| Operator fusion | None | 10-30% typical | None | Usually automatic |
Graph and kernel optimizations
Between the model and the silicon sits a compiler (inside the converter or the runtime) that rewrites the computation graph and chooses kernels. These optimizations change no weights (or change them in exactly equivalent ways), so they are free accuracy-wise, yet they often deliver 10-50% speed-ups and large memory savings.
Graph optimization is like planning a day of errands. Fusion is combining "drive to the shop, drive home, drive to the shop again" into a single trip; constant folding is buying things once that you would otherwise buy every day; layout transformation is packing groceries in the order you will unpack them; memory planning is reusing the same shopping bags instead of buying new ones each stop. The errands (the maths) are identical, but the travel (memory traffic) and clutter (peak memory) shrink.
The main rewrites
- Operator fusion: merge sequences like Conv + BatchNorm + ReLU into one kernel, or MatMul + bias + activation, or a whole attention block into a fused attention kernel. This avoids writing intermediate tensors to memory and launching several kernels.
- BatchNorm folding: at inference, BatchNorm is a fixed per-channel affine transform, so it can be folded into the preceding convolution's weights and bias. Essentially free accuracy-wise.
- Constant folding: precompute anything that does not depend on input (shape arithmetic, fixed masks, positional tables).
- Dead-code elimination and common-subexpression elimination: remove unused outputs and compute repeated expressions once.
- Layout transforms: choose the memory layout (NHWC vs NCHW, or vendor-specific tiled layouts) that the accelerator prefers, and avoid repeated transposes. Mobile CPUs and many NPUs prefer channels-last (NHWC).
- Static shapes: fixed input shapes let compilers plan memory and pick kernels ahead of time; dynamic shapes often force CPU fallback. For variable inputs, compile a few fixed buckets and pad.
- Memory planning: reuse activation buffers whose lifetimes do not overlap to cut peak RAM (liveness analysis plus a packing algorithm). On MCUs this decides whether a model fits at all.
- Tiling: split big tensors into tiles that fit in on-chip SRAM so each byte fetched from DRAM is reused as much as possible.
- Kernel selection and weight prepacking: pick algorithm variants (direct, im2col+GEMM, Winograd for 3x3 convolutions) per shape, and reorder weights once at load time into the layout the micro-kernel wants.
- Model splitting: NPUs have per-graph memory and size limits; large LLMs are split into several subgraphs (for example several compiled context binaries executed in sequence) that share buffers.
Worked example: folding BatchNorm into a convolution
Fold before quantizing: otherwise the conv weights are quantized with one range and then rescaled by BN, wasting precision.
Worked example: what fusion saves
Activation tensor: 1 x 112 x 112 x 64 in FP16 = 1.6 MB
Unfused Conv -> BN -> ReLU:
conv writes 1.6 MB, BN reads 1.6 + writes 1.6, ReLU reads 1.6 + writes 1.6
= 8.0 MB of DRAM traffic (+3 kernel launches)
Fused ConvBNReLU:
conv computes, applies BN and ReLU in registers, writes 1.6 MB once
= 1.6 MB of DRAM traffic (1 launch)
At 50 GB/s: 0.16 ms vs 0.03 ms for this block's memory traffic, and ~5x less DRAM energy.
Compilers and graph tools
| Tool | Role |
|---|---|
Converters (LiteRT converter, AI Edge Torch, torch.export plus ExecuTorch lowering, ONNX exporters) | Capture the graph, fold constants, fuse common patterns, apply quantization |
| ONNX Runtime graph optimizer and Olive | Basic, extended and layout optimizations; Olive automates conversion, quantization and tuning pipelines per target |
| TensorRT | Layer and tensor fusion, precision selection, kernel auto-tuning per GPU, producing a device-specific engine |
| Apache TVM, MLC, IREE/MLIR, XLA | ML compilers that generate kernels for many targets; MLC-LLM builds on TVM |
| Vendor compilers (QNN/QAIRT converters, NeuroPilot, Ethos-U Vela) | Map graphs to NPU instructions, tile for SRAM, produce compiled binaries |
| Netron | Visualize a graph to spot unfused patterns, odd ops, transposes and dtype changes |
Runtimes and SDKs
A runtime loads a converted model and executes it on the device, deciding which operators run on which hardware. Most runtimes use a plug-in called a delegate, execution provider or backend to hand parts of the graph to a GPU or NPU. Parts the accelerator cannot run fall back to the CPU.
A runtime is like a travel agent booking a multi-city trip. The model graph is the itinerary; the delegates are airlines. The agent books as many legs as possible on the fast direct airline (the NPU), but if that airline does not fly to one city (an unsupported op), you take a slow bus for that leg (CPU fallback), and every change of vehicle costs waiting time at the station (copies and sync). A good runtime choice is one where your whole route is covered by one fast carrier.
| Runtime / SDK | Model format | Accelerator path | Best known for |
|---|---|---|---|
| LiteRT (formerly TensorFlow Lite; renamed in 2024) | .tflite (FlatBuffer) | XNNPACK (CPU), GPU delegate (OpenCL/OpenGL/Metal), vendor NPU delegates and accelerators (for example Qualcomm and MediaTek), newer compiled-model APIs | The long-standing Android default; huge ecosystem; converters from TensorFlow, JAX and PyTorch (via AI Edge Torch) |
| LiteRT for Microcontrollers (formerly TFLite Micro) | .tflite compiled into firmware as a C array | CMSIS-NN kernels, Ethos-U micro-NPU via the Vela compiler | TinyML on MCUs with no OS and no dynamic allocation (a fixed "tensor arena") |
| ONNX Runtime (Mobile) | .onnx / .ort | Execution providers: CPU, XNNPACK, QNN, Core ML, NNAPI (legacy), plus TensorRT, CUDA, OpenVINO and DirectML on larger edge devices | Framework-agnostic; many models export to ONNX; strong quantization tooling (QDQ format); the common path on Windows NPU PCs |
| ExecuTorch | .pte | Backends: XNNPACK (with Arm KleidiAI kernels), Qualcomm, MediaTek, Arm Ethos-U, Vulkan, Core ML, MPS and others | PyTorch's official on-device path via torch.export (1.0 release in late 2025); tiny core runtime (tens of KB); strong LLM support; used in large production apps |
| Qualcomm AI Engine Direct (QNN), now packaged as the Qualcomm AI Runtime (QAIRT) SDK; older SNPE | QNN context binaries, DLC | Hexagon NPU (HTP), Adreno GPU, CPU | Lowest-level control and best performance on Snapdragon; also used underneath LiteRT, ORT and ExecuTorch Qualcomm backends; includes an LLM runtime layer (Genie) for on-device generative models |
| Qualcomm AI Hub (cloud service) | Takes PyTorch/ONNX in | Compiles, quantizes, profiles and runs models on real hosted devices | Fast way to see real latency on many chipsets without owning them |
| MediaTek NeuroPilot, Samsung and other vendor SDKs | Vendor formats | Vendor NPUs | Same role as QNN for other chip families |
| MediaPipe (Google AI Edge) | Task bundles on top of LiteRT | Inherits LiteRT delegates | Ready-made pipelines: face/hand/pose landmarks, object detection, text and audio classification, and an LLM Inference API for small open models |
| llama.cpp | GGUF | CPU (NEON, i8mm, SVE; AVX on x86), GPU (Vulkan, OpenCL, Metal, CUDA), some NPU work | The de-facto open-source quantized LLM engine for laptops and phones; fastest path to a working baseline |
| MLC-LLM | Compiled model libraries | Mobile GPU via Vulkan/OpenCL/Metal | Compiling LLMs for many backends using the TVM compiler stack |
| Core ML (Apple) | .mlmodel / .mlpackage | CPU, GPU, Apple Neural Engine | The Apple equivalent; convert with coremltools; MLX is Apple's array framework popular for LLMs on Macs |
| TensorRT (NVIDIA, including Jetson) | Serialized engine (.plan/.engine), built from ONNX | GPU tensor cores, DLA cores on Jetson | Best performance on Jetson and NVIDIA GPUs; INT8/FP8/FP16 calibration; DeepStream for video pipelines; TensorRT-LLM/edge LLM variants |
| OpenVINO (Intel) | IR (.xml/.bin) or ONNX | Intel CPU, integrated GPU, NPU | Industrial PCs, Intel laptops and edge boxes |
| Platform-provided models (e.g. Android AICore / Gemini Nano, ML Kit, Apple Foundation Models framework) | Managed by the OS | OS decides | Use a system-provided foundation model without shipping your own weights |
What happened to NNAPI
The Android Neural Networks API (NNAPI) was introduced in Android 8.1 as a vendor-neutral interface: apps (or runtimes) described a graph, and each vendor's NNAPI driver ran it on their accelerator. In practice, driver quality and operator coverage varied a lot across devices, updates were tied to OS releases, and performance was unpredictable: a lowest-common-denominator API implemented inconsistently. Google deprecated NNAPI starting with Android 15. Existing apps keep working, but new work should use LiteRT with its GPU delegate or vendor-specific NPU delegates/accelerators, or other runtimes (ONNX Runtime QNN execution provider, ExecuTorch vendor backends) that talk to vendor SDKs directly.
How a delegate partitions a graph
Original graph: [Conv]-[Conv]-[CustomOp]-[Conv]-[Softmax]-[TopK]
Delegate check: NPU NPU CPU only NPU NPU CPU
Partitioned: { NPU subgraph 1 } -> CPU -> { NPU subgraph 2 } -> CPU
^ copy/sync ^ copy/sync
Each boundary costs memory copies, layout conversion and synchronization.
Goal: as few partitions as possible, ideally one fully delegated graph.
Code: choosing execution providers and engines
# ONNX Runtime with the Qualcomm NPU (QNN execution provider) and CPU fallback
import onnxruntime as ort
sess = ort.InferenceSession(
"model.qdq.onnx", # QDQ-quantized ONNX model
providers=["QNNExecutionProvider", "CPUExecutionProvider"],
provider_options=[{"backend_path": "QnnHtp.dll"}, # libQnnHtp.so on Android/Linux
{}])
print(sess.get_providers()) # confirm what actually loaded
out = sess.run(None, {"input": x})
# Jetson: build a TensorRT engine from ONNX, INT8 + FP16, on the DLA with GPU fallback
trtexec --onnx=detector.onnx --int8 --fp16 \
--useDLACore=0 --allowGPUFallback \
--saveEngine=detector.plan
# Engines are specific to the GPU/DLA generation and TensorRT version: build on (or for) the target.
Choosing a runtime
- Start from the source framework PyTorch models fit ExecuTorch or ONNX Runtime naturally; TensorFlow/Keras/JAX models fit LiteRT. PyTorch can also reach LiteRT via AI Edge Torch.
- Check operator coverage on your target accelerator Convert, then inspect the partition report: how many ops are delegated?
- Check device coverage One chipset family (use its vendor SDK for peak performance) or thousands of Android models (use a cross-platform runtime with a reliable CPU/GPU fallback).
- Check binary size and startup time Runtime library size (from tens of KB to several MB) and model initialization/compilation time matter for app install size and cold start.
- For LLMs look at tokenizer support, KV-cache handling, quantization formats and streaming APIs: llama.cpp, ExecuTorch, MediaPipe/LiteRT LLM APIs, MLC-LLM or vendor LLM stacks.
- For embedded Linux boxes TensorRT on Jetson, OpenVINO on Intel, ONNX Runtime as a portable layer; for MCUs, LiteRT for Microcontrollers or a vendor compiler.
The deployment pipeline at a glance
Going from a research checkpoint to a shipped feature follows a repeatable loop. Expect to go around it many times. This section gives the concepts and representative commands; the step-by-step hands-on pipeline, benchmark harness, NPU deployment walkthrough and project ideas live on the companion page Edge AI Deployment.
Deploying a model is like moving a factory production line into a small workshop. You redesign the machines to fit (export and quantize), check they still make parts within tolerance (validate accuracy), install them with the right power supply (delegate and hardware), then run a full shift, not just one part, to see if the workshop overheats (sustained thermal testing). The benchmark loop is the shift supervisor with a stopwatch and a thermometer.
Pick model (e.g. from Hugging Face / model zoo)
-> Export graph (torch.export, ONNX export, TF SavedModel)
-> Optimize (fold BN, fuse ops, fix static shapes)
-> Quantize (PTQ with calibration set, or QAT; INT8 / INT4 / mixed)
-> Convert (LiteRT .tflite | ONNX/.ort | ExecuTorch .pte | QNN context binary | TensorRT engine)
-> Validate on host (compare outputs vs FP32 reference; accuracy on eval set)
-> Deploy to device (bundle in APK/AAB, download on demand, or system model)
-> Run on NPU/GPU/CPU via delegate / backend
-> PROFILE latency p50/p95, init time, peak memory, tokens/s,
power (mW / mJ per inference), thermal over 10+ minutes
-> Iterate (fix fallback ops, change precision, change model)
Export and convert: two representative paths
Post-training full-integer quantization to LiteRT from a TensorFlow SavedModel:
import tensorflow as tf
converter = tf.lite.TFLiteConverter.from_saved_model("saved_model_dir")
converter.optimizations = [tf.lite.Optimize.DEFAULT]
def representative_dataset():
for sample in calibration_samples[:300]: # real, production-like inputs
yield [sample[None, ...].astype("float32")]
converter.representative_dataset = representative_dataset
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type = tf.int8 # integer-only I/O for NPUs
converter.inference_output_type = tf.int8
with open("model_int8.tflite", "wb") as f:
f.write(converter.convert())
Exporting a PyTorch model to ExecuTorch with the XNNPACK CPU backend:
import torch
from torch.export import export
from executorch.exir import to_edge_transform_and_lower
from executorch.backends.xnnpack.partition.xnnpack_partitioner import XnnpackPartitioner
model = MyModel().eval()
example_inputs = (torch.randn(1, 3, 224, 224),)
exported = export(model, example_inputs) # captures a static graph
program = to_edge_transform_and_lower(
exported, partitioner=[XnnpackPartitioner()] # swap for a vendor NPU partitioner
).to_executorch()
with open("model.pte", "wb") as f:
f.write(program.buffer)
Run it in the app with a delegate
// Kotlin, LiteRT interpreter API with GPU delegate and CPU fallback
val options = Interpreter.Options()
val compat = CompatibilityList()
if (compat.isDelegateSupportedOnThisDevice) {
options.addDelegate(GpuDelegate(compat.bestOptionsForThisDevice))
} else {
options.setNumThreads(4) // XNNPACK CPU path
}
val interpreter = Interpreter(loadModelBuffer(context, "model.tflite"), options)
// Create once, reuse for every frame: initialization is expensive
interpreter.run(inputBuffer, outputBuffer)
Benchmark on the device
# LiteRT benchmark tool (prebuilt binary pushed to the device)
adb push benchmark_model /data/local/tmp/
adb push model_int8.tflite /data/local/tmp/
adb shell chmod +x /data/local/tmp/benchmark_model
adb shell /data/local/tmp/benchmark_model \
--graph=/data/local/tmp/model_int8.tflite \
--num_threads=4 --warmup_runs=10 --num_runs=200 \
--use_gpu=true --enable_op_profiling=true
# System-level view while the model runs
adb shell dumpsys thermalservice # thermal status and throttling
adb shell dumpsys meminfo <package> # app memory (PSS, native heap)
adb shell "for z in /sys/class/thermal/thermal_zone*; do echo \$(cat \$z/type) \$(cat \$z/temp); done"
adb shell "cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_cur_freq"
# Record a Perfetto trace with CPU frequency, scheduling and power rails
# (on devices that expose On-Device Power Monitor rails)
Model delivery options
| Option | Pros | Cons |
|---|---|---|
| Bundle in the APK/AAB | Always available, simplest | Increases install size; model updates need an app update |
| Download on demand (app's own server or Play asset/AI pack delivery) | Smaller install; can target model variants per device; update independently | First-use delay; needs integrity checks, versioning and fallback |
| System-provided model (OS service such as AICore) | No weights to ship; shared across apps; OS-managed updates (often as binary deltas rather than full re-downloads) | Only on supported devices; less control over the model and version |
| Firmware / OTA image (embedded, automotive, IoT) | Versioned with the whole system; can be signed and validated | Slow cadence; needs A/B partitions and rollback for safety |
Production concerns
- Device tiering: ship different model variants (or disable the feature) based on chipset, RAM and accelerator support.
- Warm-up and caching: GPU delegates compile shaders and NPU runtimes may compile graphs on first load; cache compiled artifacts where the runtime supports it, or ship precompiled binaries per SoC.
- Threading: never run inference on the UI thread; reuse interpreter instances; watch for contention with rendering.
- Pre/post-processing: image resize, color conversion, tokenization and non-max suppression can take longer than the model; move them to GPU or optimized native code.
- Monitoring: log latency, fallback rate and crashes per device model; keep a kill switch or remote config to disable a model variant.
- Security: models in APKs can be extracted; if the model is valuable, consider encryption at rest, integrity checks, or server-side parts.
- Validation on real silicon: emulators and host simulators are fine for functional checks but useless for performance, and accelerator numerics can differ from host simulation.
LLMs on device
Running a language model locally is the hottest and hardest edge workload. It stresses memory capacity, memory bandwidth and thermals all at once, and it has two phases with opposite performance characteristics.
An LLM generating text is like a translator who, for every single word they write, must flip through the entire dictionary (the weights) once. Reading the question (prefill) can be done in one sweep, but writing the answer happens word by word, so speed depends on how fast pages can be turned (memory bandwidth), not how smart the translator is (compute). The KV cache is their notepad of what has been said so far; a longer conversation fills more pages. Quantization makes the dictionary thinner, which is why it speeds up every word.
The two phases of generation
Prefill (prompt processing)
- Processes all prompt tokens in parallel
- Matrix-matrix multiplies: high arithmetic intensity
- Compute-bound: NPUs shine here
- Determines time to first token (TTFT)
- Fills the KV cache
Decode (generation)
- Produces one token at a time
- Matrix-vector multiplies: every weight is read for every token
- Memory-bandwidth-bound
- Determines tokens/s the user sees streaming
- Reads and appends to the KV cache
Back-of-envelope: decode speed
Example: 3B-parameter model, INT4 weights (~1.7 GB including scales)
Phone LPDDR5X peak ~ 60-75 GB/s, effective maybe ~ 40-50 GB/s
Short context (KV still small): ceiling ~ 45 / 1.7 ~= 26 tokens/s
Same model in FP16 (~6 GB weights): ceiling ~ 7 tokens/s
=> quantization speeds up decode almost in proportion to the bytes saved.
Do not omit the KV term. This 3B GQA model is ~112 KB/token in FP16 KV
(~56 KB INT8, ~28 KB INT4). At 4k context, INT8 KV is ~224 MB:
45 / (1.7 + 0.22) ~= 23 tok/s. At 16k, INT8 KV is ~0.90 GB:
45 / (1.7 + 0.90) ~= 17 tok/s. Long chats get slower even with fixed weights.
Sanity check against published phone numbers: small (1B-class) INT4 models reach
roughly 40-50 decode tokens/s on recent flagship CPUs, and 3B-class models ~20 tokens/s
at short context, consistent with this bandwidth model.
Back-of-envelope: time to first token
3B model, 1,000-token prompt: 2 x 3e9 x 1000 = 6 TFLOP (6e12 ops)
NPU at ~10 effective TOPS -> ~0.6 s TTFT
CPU at ~0.5 effective TFLOPS -> ~12 s TTFT (this is why prefill belongs on the NPU)
1B model, same prompt: 2 TFLOP -> ~0.2 s on the NPU, ~4 s on the CPU
The KV cache
During attention, each new token needs the Keys and Values of all previous tokens. Recomputing them every step would be wasteful, so they are cached. The cache grows linearly with context length (and with batch size, which is usually 1 on device):
Example (a ~3B model with GQA): 28 layers, 8 KV heads, head_dim 128, 4096 tokens, FP16
= 2 x 28 x 8 x 128 x 4096 x 2 bytes ~= 470 MB (~112 KB per token)
With full multi-head attention (24 KV heads instead of 8) it would be ~1.4 GB.
INT8 KV cache halves it; a 2048-token limit halves it again.
A ~1B model with GQA: 16 layers, 8 KV heads, head_dim 64 -> 32 KB per token
4K context = 128 MB; 32K context = 1 GB (more than its ~0.7 GB of INT4 weights).
An older 7B model with full MHA: 32 layers, 32 heads, head_dim 128 -> 512 KB per token
4K context = 2 GB; it passes its ~3.5 GB of INT4 weights at ~7K tokens.
Crossover point (cache = weights) = weight bytes / KV bytes per token.
At long context the KV cache, not the weights, becomes the memory problem.
- GQA/MQA reduce KV heads, directly shrinking the cache.
- KV-cache quantization (INT8, sometimes INT4 or lower) saves memory with small quality impact.
- Sliding-window attention caps how many past tokens each layer keeps; some models alternate local (windowed) and global layers.
- Attention sinks and eviction: keeping the first few tokens plus a recent window (and dropping the middle) keeps generation stable in streaming settings; smarter policies evict tokens that receive little attention.
- Paged / block-allocated caches allocate the cache in fixed blocks (typically 16-64 tokens) mapped by a block table. Contiguous reservation wastes (Tmax − Tactual) × KV per token; a paged cache wastes at most one partially filled block per sequence and avoids reallocating a giant buffer as the conversation grows.
- Context limits: on device, 2K-8K tokens is typical; long documents need summarization, chunking or retrieval.
- Prompt/prefix caching: reuse the KV of a byte-identical prefix (system prompt, tools, a shared document). Prefill savings are about P / (P + U) when prefix P is cached and only suffix U is new. Put static tokens first; a timestamp or user id at the top of the prompt misses the cache every request.
Speculative decoding and friends
Decode leaves compute idle because it is memory-bound. Speculative decoding spends that spare compute: a tiny draft model proposes k tokens; the large target model checks all of them in one parallel forward pass (which costs about the same as one decode step, since the weights are read once) and accepts the longest prefix that matches what it would have produced. With the right acceptance rule, the output distribution is exactly the target model's.
- Self-speculative variants avoid a second model: extra prediction heads (Medusa-style), a light draft head reusing the target's features (EAGLE-style), or skipping layers to draft.
- Prompt-lookup / n-gram drafting copies candidate continuations from the prompt itself, which works very well for summarization, rewriting and code editing where output repeats input.
- On device, the draft model's memory and the verification cost of rejected tokens must fit the budget; gains are largest when the task is predictable.
Other speed-up techniques
- Operator fusion for attention (fused attention kernels, and runtime options that fuse scaled-dot-product attention with KV-cache updates) reduces memory traffic.
- Heterogeneous execution: prefill on the NPU, decode on whichever unit gives the best tokens per watt; tokenization and sampling on CPU.
- Adapter swapping (LoRA): one shared base model with small task-specific adapters instead of several full models.
- Model cascades: answer simple requests with a tiny model; escalate to a larger local or cloud model when needed.
- Constrained decoding: grammars or JSON schemas restrict sampling to valid outputs, so small models produce reliable structured results without retries.
Small model families (landscape at the time of writing)
| Family | Typical on-device sizes | Notes |
|---|---|---|
| Gemma (Google) | ~270M, 1B, 4B; "n" variants with ~2B/4B effective parameters | The "n" variants use per-layer embeddings that can stay in slower memory and nested sub-models; multimodal options |
| Llama 3.2 (Meta) | 1B, 3B (plus official quantized versions) | Quantized releases used SpinQuant and QAT+LoRA; widely supported by ExecuTorch and llama.cpp |
| Phi (Microsoft) | ~3.8B "mini" models | Strong reasoning for size; trained heavily on curated and synthetic data |
| Qwen (Alibaba) | ~0.5B to 4B | Multilingual; many sizes; popular for on-device experiments |
| SmolLM and similar tiny models | ~135M to 1.7B | Fully open training recipes; good for very constrained devices |
| Platform models | Few-billion-parameter models managed by the OS | For example Gemini Nano via AICore on Android and Apple's on-device foundation model, typically quantized aggressively (down to ~2-4 bits) with LoRA adapters per feature |
Memory reality check
| Model size | INT4 weights | Fits comfortably on |
|---|---|---|
| ~0.5-1B | ~0.3-0.7 GB | Most mid-range phones; good for classification, extraction, short replies |
| ~2-4B | ~1.2-2.2 GB | Recent flagships with 8-12 GB RAM; summarization, rewriting, assistants |
| ~7-8B | ~3.5-4.5 GB | Only high-RAM (12-16 GB+) phones and laptops; risky on mobile |
Android also shares RAM between apps; a large resident model can push other apps out of memory. Memory-mapping weights (mmap) lets the OS page them, but paging from storage during decode destroys performance. Watch peak resident memory (RSS): it is the number that decides whether the low-memory killer terminates your process.
Code: running a quantized LLM
# llama.cpp: convert, quantize, benchmark, run
python convert_hf_to_gguf.py ./my-hf-model --outfile model-f16.gguf --outtype f16
./llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M
./llama-bench -m model-Q4_K_M.gguf -p 512 -n 128 -t 4 # prefill and decode tokens/s
./llama-cli -m model-Q4_K_M.gguf -p "Explain KV cache in one line" -n 64
# ExecuTorch LLM export (shape of the command; flags change between releases)
python -m extension.llm.export.export_llm \
base.model_class="llama3_2" base.checkpoint="$CKPT" base.params="$PARAMS" \
model.use_kv_cache=True model.use_sdpa_with_kv_cache=True \
quantization.qmode="torchao:8da4w" quantization.group_size=128 \
quantization.embedding_quantize="torchao:4,32" \
export.output_name="llm_int4.pte"
# 8da4w = 8-bit dynamic activations, 4-bit weights. Leaving the KV-cache options off
# produces a model that recomputes attention for the whole context every token.
Quality and safety on device
- Small models hallucinate more; constrain tasks (summarize this text, extract fields) rather than open-ended knowledge questions, or ground them with on-device retrieval.
- Evaluate on task-specific sets plus perplexity; compare quantized against unquantized outputs.
- Safety filters and prompt-injection defences still apply, even when everything is local.
- Licences of open-weight models vary; check them before shipping.
Vision, audio, sensor models and TinyML
LLMs get the headlines, but most edge AI running today is perception: camera pipelines, speech and sound, and sensor streams from motion, heart-rate or vibration sensors. These workloads are continuous, latency-critical and often always-on, so they are shaped by duty cycle and energy far more than by peak speed. At the extreme end, TinyML runs them on microcontrollers with kilobytes of memory.
An always-on perception system is like a building's security setup. A cheap motion sensor at the door (a tiny always-on model on a DSP or MCU) watches all the time using almost no power; only when it trips does it wake the guard (a larger model on the NPU), who in turn calls the police (the cloud) only for real incidents. Keeping the guard awake all night would be expensive and pointless; the cascade gives the same safety at a fraction of the cost.
On-device vision
- Tasks: classification, detection (YOLO-style, SSD), segmentation (portrait blur, background replacement), pose and landmark estimation, OCR, super-resolution and denoising, depth estimation, image embeddings for search.
- Resolution is the biggest knob: compute scales with pixel count, so going from 640x640 to 320x320 cuts compute about 4x, at the cost of small-object accuracy.
- Pipeline cost: cameras deliver YUV frames; conversion to RGB, resizing, rotation and normalization can cost more than the model if done naively on the CPU. Use the GPU, the ISP or fused native code; fold normalization into the first layer.
- Post-processing: non-max suppression, decoding anchor boxes and upsampling masks are often left on the CPU and can dominate small models.
- Temporal tricks: run the detector every few frames and a cheap tracker in between; skip inference when the scene is static; process regions of interest instead of full frames.
On-device audio and speech
- Keyword spotting (wake words): tiny always-on models on a DSP, typically a few tens of thousands of parameters, operating on log-mel or MFCC features.
- Speech recognition: streaming architectures (RNN-transducer, conformer-transducer) for live captions and dictation; encoder-decoder models such as Whisper-tiny/base-class for offline transcription.
- Real-time factor (RTF) = processing time / audio duration; streaming needs RTF well below 1 with bounded latency per chunk.
- Other audio: noise suppression in calls, sound event detection (alarms, glass breaking), speaker verification, on-device text-to-speech.
Sensor models
- Motion (IMU): activity recognition, fall detection, gesture recognition from accelerometer and gyroscope windows, often on a sensor hub.
- Health (PPG, ECG): heart-rate estimation, arrhythmia screening, sleep staging on wearables; see Wear OS for platform constraints.
- Industrial: vibration anomaly detection with small autoencoders (flag inputs the model reconstructs badly), predictive maintenance on motors and pumps.
- Radar and ultrasound: presence detection and gesture sensing at very low power.
TinyML: machine learning on microcontrollers
A microcontroller has no DRAM, often no OS, and memory measured in kilobytes. Weights and code live in flash (read-only, hundreds of KB to a few MB); activations live in SRAM (tens to hundreds of KB). The binding constraint is usually peak activation memory, not weight size.
| Resource | Typical MCU | What goes there |
|---|---|---|
| Flash | 256 KB-2 MB (up to ~8 MB) | Firmware, runtime, INT8 weights |
| SRAM | 64 KB-1 MB | The "tensor arena": activations, scratch buffers, input/output |
| Compute | Cortex-M4/M7 at 64-600 MHz with DSP/SIMD instructions; M55/M85 with Helium vectors; optional Ethos-U micro-NPU | INT8 kernels (CMSIS-NN) or NPU commands |
| Power | ~1-100 mW active, µW asleep | Battery or energy harvesting, months to years of life |
Worked example: does a keyword spotter fit?
Input: 1 s of 16 kHz audio -> 49 frames x 10 MFCC features (INT8) = 490 bytes
Model: small depthwise-separable CNN, ~25 K parameters (INT8) = ~25 KB flash
~3-6 M MACs per inference
Largest two live activations: 49x10x64 input feature map + same-size output
= 31 KB + 31 KB = ~63 KB peak arena (after in-place / memory planning tricks, less)
Target MCU: 256 KB SRAM, 1 MB flash, Cortex-M4 at 80 MHz
Latency: a few million MACs with SIMD kernels -> tens of ms; run ~every 100-250 ms
Energy: e.g. 10 mW while active x 30 ms = 0.3 mJ per inference; 4 inferences/s = ~1.2 mW average
Verdict: fits comfortably; spend remaining SRAM on the audio ring buffer.
Code: LiteRT for Microcontrollers
#include "tensorflow/lite/micro/micro_interpreter.h"
#include "tensorflow/lite/micro/micro_mutable_op_resolver.h"
#include "model_data.h" // model compiled into flash as a C array
constexpr int kArenaSize = 64 * 1024; // all activations live here: no malloc
alignas(16) static uint8_t tensor_arena[kArenaSize];
void setup_and_run(const int8_t* features) {
const tflite::Model* model = tflite::GetModel(g_model_data);
static tflite::MicroMutableOpResolver<5> resolver; // register only what you use
resolver.AddConv2D();
resolver.AddDepthwiseConv2D();
resolver.AddFullyConnected();
resolver.AddReshape();
resolver.AddSoftmax();
static tflite::MicroInterpreter interpreter(model, resolver, tensor_arena, kArenaSize);
if (interpreter.AllocateTensors() != kTfLiteOk) { /* arena too small */ return; }
TfLiteTensor* input = interpreter.input(0);
memcpy(input->data.int8, features, input->bytes);
interpreter.Invoke();
int8_t* scores = interpreter.output(0)->data.int8; // dequantize with output scale/zero-point
}
Always-on design patterns
- Stage 0: hardware trigger A sensor interrupt, voice activity detector or motion threshold in the sensor hub, costing microwatts.
- Stage 1: tiny model Keyword spotter or activity classifier on a DSP or MCU, tuned for high recall.
- Stage 2: verifier A larger model on the NPU or CPU, woken only when stage 1 fires, tuned for precision.
- Stage 3: heavy lifting Full ASR, a local LLM or a cloud request, only after verification.
The false-trigger rate of each stage directly drives energy: every false wake costs a stage-2 run. Tune thresholds on real-world audio or sensor data, not clean benchmarks.
Power and thermal budgets
Phones, watches and glasses are passively cooled and battery-powered. A sustained workload of a few watts heats the device, and the thermal governor lowers CPU/GPU/NPU clocks. Design for sustained performance and energy per result, not a single warm run. The platform-level details of power management are covered in Power and Thermal; this section focuses on what matters for inference.
Running a model on a phone is like running on a hot day. You can sprint for a short while (burst clocks), but your body heats up and eventually forces you to slow to a jog (throttling). What matters for a marathon (a 30-minute video call) is the pace you can hold, not your sprint time, and the total energy you burn (battery). Finishing an errand quickly and then resting in the shade (race to idle) is often better than walking slowly in the sun the whole time.
Energy, not power, is the metric
Compare two paths for the same model:
CPU : 1.0 W above idle x 60 ms = 60 mJ per inference
NPU : 2.0 W above idle x 10 ms = 20 mJ per inference (2x the power, 1/3 the energy)
A continuous camera feature at 30 fps on the NPU:
20 mJ x 30 = 0.6 W for the model alone (+ camera, ISP, display, pre/post-processing)
0.6 W / 19.25 Wh = ~3.1% battery per hour from inference alone
The same feature at 15 fps with a model half the size (~10 mJ):
10 mJ x 15 = 0.15 W -> ~0.8% per hour. Duty cycle and model size multiply.
Race to idle versus DVFS
Dynamic power scales roughly as P ≈ C × V2 × f. Lower frequency allows lower voltage, so energy per operation falls at lower clocks; that argues for running slowly. But the chip also burns static (leakage) power and keeps memory, rails and other blocks awake while working; that argues for finishing fast and powering down ("race to idle"). The best operating point depends on the chip and workload: accelerators usually win on both counts because they finish faster at lower energy per operation, and for CPUs the most efficient point is typically a mid frequency on the big or mid cores, not the maximum. Measure rather than assume.
Thermal behaviour
- Thermal mass buys time: a phone can absorb heat for seconds to a few minutes before the skin temperature limit triggers throttling, which is why short benchmarks look great.
- Sustained-to-peak ratio: throughput after 10-30 minutes divided by the first-minute throughput. Ratios of 0.5-0.7 are common for heavy workloads on phones; the ratio is what a shipped feature actually gets.
- Environment matters: ambient temperature, phone cases, charging (which adds heat), and simultaneous camera and display use all change the result.
- Shared budget: the SoC's thermal budget is shared; a hot GPU from gaming or a camera pipeline reduces what the NPU can sustain.
throughput
(tok/s) |****
| ****
| ****** <- throttling begins (skin temp limit)
| *********************** sustained plateau
|
+------------------------------------------------ time
0 1 2 3 5 10 20 min
Designing for the budget
- Duty cycle: a camera feature at 30 fps runs continuously; a photo-edit feature runs once. Budgets differ by orders of magnitude. Lower frame rate, skip frames or run on events.
- Race to idle: finishing quickly and letting hardware sleep is often more efficient than running slowly.
- Android APIs:
PowerManager.getThermalHeadroom()and thermal status listeners let you degrade gracefully (lower frame rate, smaller model, lower resolution) before the OS throttles hard. The Android Dynamic Performance Framework (ADPF) helps give the scheduler hints for steady workloads. - Always-on features (wake words, activity detection) belong on a low-power DSP or sensor hub, waking the main processor only on a trigger.
- Batch background work: index photos or fine-tune adapters only when charging and idle.
// Kotlin: degrade gracefully before the OS throttles
val pm = context.getSystemService(PowerManager::class.java)
fun chooseModelTier(): Tier {
val headroom = pm.getThermalHeadroom(10) // forecast 10 s ahead; 1.0 = severe throttling
return when {
headroom.isNaN() -> Tier.DEFAULT // not supported / called too often
headroom > 0.9f -> Tier.SMALL_15FPS
headroom > 0.7f -> Tier.MEDIUM_30FPS
else -> Tier.LARGE_30FPS
}
}
pm.addThermalStatusListener { status ->
if (status >= PowerManager.THERMAL_STATUS_SEVERE) pauseNonEssentialInference()
}
How to measure power
- On-device power rails (On-Device Power Monitor) captured in Perfetto traces on supported devices: per-subsystem energy counters.
- External power monitor in place of the battery: the most precise lab method.
- Fuel gauge and batterystats: coarse, but available on any device and in the field.
- Always subtract an idle baseline measured with the same screen, radio and camera state, and report energy per inference plus battery percent per hour for continuous features.
Privacy, federated learning and security
Keeping inference on the device keeps raw data local, but products also need to improve models from user data and to protect the model itself. Federated learning, differential privacy and secure aggregation let a fleet learn without collecting raw data; device security measures protect the model and its outputs.
Federated learning is like a cooking school where students practise recipes at home and mail back only their suggested tweaks, never their family's secret ingredients or who they cooked for. The school averages everyone's tweaks into a better recipe (the global model). Secure aggregation is a sealed ballot box, so the school only sees the combined tweaks; differential privacy adds a pinch of random noise to each suggestion, so no single family's habits can ever be reverse-engineered from the final recipe.
Federated learning (FL)
- Select clients The server picks a sample of eligible devices (typically charging, idle, on unmetered Wi-Fi, with enough local data).
- Broadcast Devices download the current global model (or just an adapter).
- Train locally Each device runs a few steps of training on its own data.
- Report updates Devices send weight updates (not data), usually clipped, compressed and protected by secure aggregation.
- Aggregate The server averages the updates, weighted by the amount of local data, and produces a new global model. Repeat for many rounds.
- Challenges: non-IID data (each user's data looks different), unreliable and slow clients, communication cost (updates can be larger than the data), stragglers and dropouts, and debugging a model you cannot inspect the data for.
- Good fits: keyboard next-word prediction and emoji suggestions, wake-word and speaker models, on-device ranking and recommendation.
- Lighter alternative: on-device personalization that never leaves the device, such as fine-tuning a small adapter or a last layer locally, or keeping a local embedding index.
Differential privacy (DP)
Differential privacy is a mathematical guarantee that the output of a computation changes very little whether or not any single individual's data is included.
- DP-SGD / DP-FedAvg: clip each example's (or each client's) update to a maximum norm C, then add Gaussian noise proportional to C before aggregation. A privacy accountant tracks the cumulative ε across rounds.
- Central versus local DP: in central DP a trusted aggregator adds noise to the combined result; in local DP each device adds noise before sending, which needs no trust but costs far more accuracy.
- Trade-off: more noise gives stronger privacy but slower convergence and lower accuracy; large client populations make the noise affordable.
Secure aggregation and trusted hardware
Secure aggregation is a cryptographic protocol in which devices mask their updates so that the server can compute only the sum across many devices, never an individual update. Trusted execution environments (TEEs) and hardware-backed keystores protect keys, model decryption and sensitive processing on the device itself.
Threats to on-device models
| Threat | What it is | Mitigation |
|---|---|---|
| Model extraction | Copying the weights out of the app package or memory | Encryption at rest, hardware-backed keys, vendor-compiled binaries, keeping crown-jewel parts server-side; accept that a rooted device can usually extract weights |
| Adversarial inputs | Crafted inputs (stickers, noise, audio perturbations) that fool the model | Robust training, input validation, multi-sensor confirmation, conservative thresholds for safety-critical actions |
| Membership and data inference | Inferring whether someone's data was used in training, or reconstructing it from updates | Differential privacy, secure aggregation, update clipping |
| Model tampering | Replacing a downloaded model with a malicious one | Signed models, integrity checks before load, secure download channels |
| Prompt injection (LLMs) | Untrusted content (a message, a web page) instructing the local model | Same defences as in the cloud: separate instructions from data, restrict tool permissions, output filtering |
| Telemetry leakage | Debug logs or analytics that quietly upload the private data the model was meant to protect | Privacy review of logging, aggregate-only metrics, DP for analytics |
Hybrid edge-cloud architectures
Pure on-device and pure cloud are the two ends; most successful products sit in between, using each tier for what it does best. The design question is not "edge or cloud?" but "which requests, which stages and which data go where, and what happens when the network or the device cannot cope?"
A hybrid system is like a hospital network. The local clinic (the device) handles routine visits instantly and keeps your records private; hard cases are referred to the specialist hospital (the cloud) with only the relevant notes (features or a summary), not your entire history. A triage nurse (the router) decides who gets referred, and if the road to the hospital is closed (offline), the clinic still treats you as well as it can. Mapping back: routing policy, minimal data transfer and a guaranteed local fallback are the three pillars of a hybrid design.
Common patterns
| Pattern | How it works | Example |
|---|---|---|
| Local trigger, cloud heavy lifting | A tiny always-on model detects an event, then streams to the cloud | Wake word then cloud assistant; motion detection then cloud video analysis |
| Cascade / escalation | A small local model answers when confident; low confidence or complex requests escalate | On-device LLM for short rewrites, cloud LLM for long reasoning |
| Routing by policy | A router sends requests by privacy class, capability, cost and connectivity | Personal data stays local; public-knowledge questions go to the cloud |
| Split computing | The first layers run on the device; compressed intermediate features go to a server for the rest | Camera features sent instead of raw video; research-heavy but used in some vision systems |
| Edge server offload | Nearby edge servers (on-prem or telecom MEC) run heavier models with LAN-like latency | Factory quality inspection, AR rendering, multi-camera analytics |
| Cloud trains, edge infers | Training and evaluation in the cloud; models shipped to the fleet over the air | Nearly every production edge model |
| Fleet learning | Edge devices flag hard or novel examples (with consent) or send federated updates to improve the next model | Autonomous driving data engines, keyboard models |
Designing the router
- Classify the request By privacy sensitivity, expected difficulty (length, task type), latency requirement and cost.
- Try local first where allowed Start local immediately for responsiveness; use a confidence signal (model self-score, a small classifier, or task rules) to decide whether to escalate.
- Escalate with consent and minimal data Send only what is needed (a summary, extracted fields, embeddings) and only if the user and policy allow.
- Keep a consistent contract The UI receives the same response format whichever path answered, with a flag for provenance.
- Handle failure Offline or timeout means a local answer with honest messaging, never a hang.
- Log and tune Record routing decisions and outcomes (privacy-safely) to tune thresholds for quality, cost and latency.
Worked example: routing economics
Assistant feature: 50 M daily users x 20 requests/day = 1 B requests/day
Cloud cost: ~0.0005 per request (illustrative) -> 500,000 per day if all cloud
Local model handles 70% of requests acceptably -> 150,000 per day
Local energy per request: ~2 J (short generation on NPU) -> 20 requests = 40 J/day per user,
~0.06% of a 69 kJ battery
Latency: local first token ~0.3 s vs cloud ~1 s on a good network, several s on a poor one
Trade-off: 70% cost saving and better latency, paid for with device engineering effort,
a download of ~1-2 GB of weights (or an OS-provided model), and a quality gap on hard requests.
A platform AI service: what an OS-level design looks like
Operating systems increasingly host one shared foundation model as a system service rather than letting every app ship its own. The design is a good reference architecture for any multi-client edge AI system:
- One resident model shared across apps instead of N copies, with a memory budget and eviction under memory pressure.
- Device gating: enabled only on devices with enough RAM and a capable NPU.
- Per-feature LoRA adapters (tens of MB each) hot-swapped on the shared base, stored per app or in a shared location.
- Model updates as binary deltas rather than full multi-GB downloads.
- Permission isolation and safety filtering on inputs and outputs, enforced centrally.
- Request queueing and fairness across clients, with thermal-aware rate limiting.
- Graceful degradation policy: NPU, then GPU, then CPU, then refuse or defer, documented and observable.
Benchmarking and evaluation metrics
Edge success is multi-dimensional. A model that is accurate but drains 10% battery per hour will be disabled by users; a fast model that drops accuracy by 8% will fail product review; a model that is fast in the lab but throttles in a user's pocket will get one-star reviews. Benchmarking concepts are summarized here; the companion page Edge AI Deployment walks through building a full measurement harness.
Benchmarking a model is like testing a car: 0-100 time (latency), fuel per 100 km (energy per inference), luggage space (memory), and whether the engine overheats climbing a long hill (sustained thermal). A car that wins a single drag race but overheats on the highway is not a good commuter car, just as a model that looks fast in one warm run can fail in a 20-minute video call.
| Metric | Definition | How to measure | Watch out for |
|---|---|---|---|
| Latency (p50, p95, p99) | Time per inference | Benchmark tools, in-app timers around the full pipeline | Report percentiles, not only averages; include pre/post-processing |
| Initialization time | Load + compile time before first inference | Timer from model load to first result, cold and warm | Can be seconds for large GPU/NPU graphs; affects cold start |
| Throughput / FPS | Inferences per second | Sustained run | May drop sharply after thermal throttling |
| Time to first token (TTFT) | LLM: prompt submitted to first output token | Dominated by prefill of the prompt | Grows with prompt length |
| Decode tokens/s | LLM: generation speed after the first token | llama-bench, runtime stats | Bound by memory bandwidth; drops as context grows |
| Prefill tokens/s | LLM: prompt processing speed | Runtime stats | Compute-bound; NPUs help a lot |
| Peak memory | Max RAM used (weights + activations + KV cache + runtime) | dumpsys meminfo, Perfetto heap profiles, peak RSS | Risk of the low-memory killer terminating your app or background apps |
| Model size | File size on disk / download | File size | Affects install size and update cost |
| Accuracy drop | Metric change vs FP32 reference (top-1, mAP, WER, perplexity, task score) | Fixed evaluation set, same pre-processing | Also check tail cases and fairness slices, not just the average |
| Power / energy | mW average; mJ per inference; battery % per hour | Power rails (ODPM), external power monitor, batterystats | Subtract the idle baseline; screen state matters |
| Thermal | Skin/SoC temperature and throttling over time; sustained-to-peak ratio | Thermal service, sustained 10-30 min tests | Room temperature and phone case change results |
| Delegation coverage | Share of ops on the accelerator | Delegate logs, op profiling | A single CPU op in the middle can halve performance |
| Real-time factor | Speech: processing time / audio duration | Timers on streaming chunks | Chunk latency matters as much as average RTF |
Why percentiles matter
Averages hide the stalls users notice. If a camera feature has a mean latency of 12 ms but a p99 of 45 ms, one frame in a hundred misses a 33 ms frame deadline and the preview visibly stutters. Report p50 (typical), p95/p99 (tail) and the maximum, and look at latency over time to separate random jitter (scheduling, garbage collection) from trends (throttling, growing KV cache).
A fair benchmarking protocol
- Fix the environment Same device, OS build, charge level (or on external power supply for power tests), airplane mode, screen brightness, room temperature.
- Warm up Run 10-50 inferences before measuring and report initialization separately.
- Measure distributions Hundreds of runs; report p50/p95/p99 and standard deviation.
- Test sustained load Run continuously for 10-30 minutes and plot latency and temperature over time.
- Compare against a baseline FP32 on CPU, then each optimization step, in one table.
- Test across device tiers Flagship, mid-range and low-end chipsets behave very differently.
- Measure end to end Include pre/post-processing and data copies, and measure inside the real app as well as in the benchmark tool.
Useful tools
- LiteRT
benchmark_modeland its op profiler; ONNX Runtimeonnxruntime_perf_test; ExecuTorch developer tools (ETDump profiling);llama-benchfor GGUF models;trtexecand Nsight Systems on Jetson. - Perfetto / Android System Trace for CPU frequency, scheduling, GPU activity and power rails.
- Vendor profilers (for example Snapdragon Profiler, Qualcomm AI Hub profile jobs) for per-layer NPU timing.
- Netron to visualise a model graph and spot odd ops or layouts.
- Android GPU Inspector for GPU delegate behaviour.
- Industry benchmark suites (MLPerf Mobile, MLPerf Tiny, MLPerf Inference edge categories) for standardized cross-device comparisons.
Industry use cases
Edge AI is already everywhere; what differs by industry is the dominant constraint. Knowing a few concrete deployments per sector, and the constraint that shaped each one, makes design interviews much easier.
Edge AI across industries is like the same engine fitted into different vehicles: a scooter (wearable), a family car (phone), a lorry (car or robot) and a stationary generator (factory camera). The engine principles are identical, but the scooter cares about fuel, the lorry about safety and reliability, and the generator about running non-stop for years. Mapping back: the same quantization, runtime and benchmarking skills apply everywhere, while the binding constraint (battery, safety certification, uptime, cost) changes by sector.
Smartphones
Computational photography (night mode, portrait segmentation, HDR fusion), live captions and translation, on-device dictation, smart reply and text rewriting with small LLMs, photo search by content, call noise suppression, on-device summarization. Constraints: fragmentation, battery, thermals, app size.
Wearables and hearables
Activity and sleep tracking, heart-rhythm screening, fall detection, gesture control, wake words, adaptive noise cancellation. Constraints: milliwatt budgets, tiny batteries, always-on sensing on a sensor hub or MCU, skin-contact thermal limits. See Wear OS.
Automotive
Driver and occupant monitoring (drowsiness, distraction), ADAS perception (lane, object and sign detection), in-cabin voice assistants, surround-view parking. Constraints: functional safety (deterministic latency, redundancy, certification), automotive temperature ranges, long product lifetimes and OTA updates.
Smart cameras and video analytics
Person, vehicle and package detection, licence-plate reading, retail footfall counting, industrial safety monitoring; only events or metadata are uploaded. Constraints: multi-stream throughput per watt, bandwidth, privacy regulation, 24/7 reliability.
Industrial IoT and manufacturing
Visual quality inspection on production lines, predictive maintenance from vibration and acoustic sensors, anomaly detection in process data. Constraints: harsh environments, deterministic latency, air-gapped networks, long hardware lifetimes.
Smart home and consumer electronics
Voice assistants with local wake word and increasingly local command handling, TV upscaling and picture enhancement, robot vacuums with obstacle recognition, presence detection. Constraints: cost per unit (cents matter), privacy perception.
Robotics, drones and XR
Perception, SLAM, grasp planning, obstacle avoidance; hand and eye tracking and scene understanding in headsets and glasses. Constraints: very low motion-to-photon latency, weight and battery, heat on the head.
Healthcare and PCs
Point-of-care imaging assistance and on-premises clinical data processing where data cannot leave the site; on laptops, NPUs run local assistants, meeting effects (background blur, eye contact, noise removal) and local search. Constraints: regulation and validation; battery life on laptops.
Constraint by sector
| Sector | Binding constraint | Typical hardware | Typical model size |
|---|---|---|---|
| Phones | Battery, thermal, fragmentation | SoC NPU/GPU/CPU | 1-100 MB vision/audio; 0.5-2 GB LLMs |
| Wearables | Milliwatt power, memory | Low-power SoC, sensor hub, MCU | KB to a few MB |
| Automotive | Safety, determinism, lifetime | Automotive SoCs with NPUs/GPUs, safety MCUs | Tens to hundreds of MB; multiple concurrent models |
| Cameras / NVR | Throughput per watt, cost, 24/7 uptime | Camera SoC NPU, Jetson-class, add-on accelerators | 5-100 MB detectors |
| Industrial | Reliability, latency, connectivity | Industrial PCs, Jetson, MCUs on sensors | KB (sensors) to hundreds of MB (vision) |
Common failure modes and how to avoid them
Most edge AI projects that fail do not fail because the model is wrong; they fail because of measurement mistakes, silent fallbacks, unplanned memory pressure or scope creep. This section collects the traps, organized by where they appear.
These failure modes are like the classic mistakes of new pilots: trusting a single instrument, skipping the pre-flight checklist, and practising only in perfect weather. Each one is avoidable with a checklist and honest measurement. In edge AI, the checklist is: measure on real hardware, measure sustained, check what actually ran on the accelerator, and budget memory with the rest of the system running.
| Failure mode | Symptom | Prevention |
|---|---|---|
| Cold-run-only benchmarking | Lab numbers far better than field numbers | Report warm-up separately, percentiles, and a 10-30 minute sustained curve |
| Measuring on emulators | Numbers that no device reproduces | Physical silicon for every performance claim |
| Silent CPU fallback | "NPU enabled" but slow; high CPU usage | Check partition logs and per-op profiles in CI |
| Pre/post-processing blind spot | Model is 5 ms, feature is 40 ms | Time each pipeline stage; move processing to GPU/native code |
| Unrepresentative calibration | Accuracy collapses on night images, accents, small objects | Calibrate and evaluate on production-like data and slices |
| Ignoring memory pressure | App killed on 8 GB devices, other apps evicted | Measure peak RSS with realistic background load; gate by RAM tier |
| Chasing model size | Months spent making a too-large model barely fit | Start from the smallest model that meets the quality bar; scale up only with evidence |
| Toolchain rabbit holes | Weeks lost to SDK version mismatches and build errors | Get a simple baseline working first (for example llama.cpp or the CPU path), pin versions, timebox setup |
| Vendor SDK regressions | An upgrade makes the model slower or changes outputs | Performance and accuracy regression suite on a device farm; pin versions in production |
| Mismatched Hexagon/driver libraries | Cryptic load failures on some devices only | Ship the accelerator libraries that match each SoC generation; test per tier |
| Average-only accuracy | Regressions in specific user groups go unnoticed | Per-slice metrics, worst-case examples, agreement with the reference |
| No kill switch | A bad model variant crashes a device family with no quick fix | Remote config per device model, staged rollout, telemetry |
Skills map, roles and further study
Edge AI rewards people who combine two halves: enough ML understanding to know what a model is doing, and deep systems skill to make it run well on silicon. Engineers from mobile, platform, embedded, driver, DSP or power-optimization backgrounds already own the scarce systems half. Many edge roles are explicitly inference-engineering roles (C++ runtimes, SDKs, kernels, deployment), not model-training roles. Detailed learning schedules, benchmark projects and target numbers are on Edge AI Deployment; this section maps the territory.
Moving from platform engineering into edge AI is like a race-car mechanic joining a Formula team: you do not need to design the aerodynamics from scratch (training models), but you must understand why the car is shaped that way so you can tune the engine, gearbox and cooling to win on a real track (hardware, runtime and thermal). The ML half is learning the car's design; the systems half, which you may already have, is what makes it finish the race.
How platform skills transfer
| Existing skill | Becomes, in edge AI |
|---|---|
| SoC, BSP, HAL and driver experience | Deploying and debugging models on NPU/DSP/GPU; understanding delegates and vendor stacks |
| Power and thermal optimization | The number-one edge constraint: performance per watt, sustained thermal behaviour, battery impact |
| Android framework and system services | Integrating runtimes into apps and platform services; model delivery; system-level inference services |
| DSP / signal processing / fixed-point maths | Natural intuition for quantization, scales, saturation and overflow |
| C++ and performance engineering | Runtimes, kernels, delegates and LLM engines are mostly C/C++ |
| Cross-layer debugging | Profiling inference across app, runtime, driver and hardware |
| Integration and delivery leadership | Productizing research models: gates, benchmarks, device tiering, rollout |
Skills map
| Area | Must-have knowledge | Priority |
|---|---|---|
| AI foundations | Neural network basics, CNNs, transformers, attention, tokens, embeddings, training vs inference, fine-tuning vs RAG vs prompting | Foundation |
| Python and PyTorch | Load a model, run inference, export, NumPy, Hugging Face Transformers | Quick add for C/C++/Java engineers |
| Quantization | INT8/INT4, FP16/BF16/FP8, PTQ vs QAT, per-tensor vs per-channel vs per-group, calibration, mixed precision, layer-wise error analysis | High |
| Compression | Pruning, distillation, low-rank, LoRA/QLoRA | High |
| Graphs and formats | ONNX, torch.export, operator sets, fusion, static shapes, Netron | High |
| Edge runtimes | LiteRT, ONNX Runtime, ExecuTorch, delegates and fallback | High |
| Vendor SDKs | At least one: Qualcomm AI Engine Direct / AI Hub, MediaTek NeuroPilot, TensorRT, or another NPU stack | High for silicon/OEM roles |
| Accelerator hardware | NPU/DSP/GPU architecture, memory bandwidth, roofline, numeric formats | High |
| On-device LLMs | llama.cpp, GGUF, KV cache, prefill vs decode, speculative decoding | High demand |
| Profiling | Latency, memory, power, thermal; Perfetto; vendor profilers | Core |
| Deployment | Android packaging, model delivery, device tiering, monitoring | Core |
A generic learning path
- AI literacy Be able to explain transformers, attention, tokens, embeddings, training vs inference and fine-tuning vs RAG in plain words; get comfortable with Python, PyTorch basics and Hugging Face.
- ML through a systems lens Quantization in depth, compression, model formats and graph optimization, and inference maths (compute vs memory bound, KV cache, why NPUs are fast).
- Edge runtimes and toolchains Deploy the same model with two runtimes, use one vendor NPU SDK, run a quantized LLM with llama.cpp or ExecuTorch, and profile latency, memory, tokens/s, power and thermal on a real device.
- Evidence Build two or three benchmarked projects with a clear results table and write them up (see the companion page for project designs).
Roles
Good fit
On-Device / Edge AI Engineer, AI Systems or ML Runtime Engineer, Inference SDK Engineer, Model Optimization / Quantization Engineer, AI Platform Integration Lead, embedded ML on mobile, wearables, automotive, XR and IoT.
Adjacent
GenAI application development, RAG and agent apps, MLOps / AI infrastructure. Useful literacy, but more crowded and less dependent on hardware depth.
Different career
ML research, data science and training frontier models from scratch. These need a different profile; edge work needs understanding of models, not inventing them.
Where demand is
Silicon and platform vendors, phone and PC makers building AI features, wearable and XR programs, automotive (driver monitoring, in-cabin AI), robotics, smart cameras and runtime/tooling vendors.
Portfolio project ideas (summary)
On-device LLM benchmark
A quantized 1-3B model on a phone: TTFT, prefill and decode tokens/s, peak memory, energy and a sustained-load curve, INT4 vs higher precision.
Vision model on the NPU
INT8 detector on NPU vs GPU vs CPU with latency and energy per frame, documenting fallback ops and how they were removed.
Quantization study
Sweep FP32, FP16, INT8 PTQ, QAT and INT4 on one model with per-layer sensitivity analysis and an evidence-derived mixed-precision recipe.
KV-cache study
KV footprint against context length, crossover with weight size, KV quantization at several bit-widths, and maximum usable context under real memory pressure.
Real-time speech or TinyML
Streaming ASR with real-time factor and WER change, or a keyword spotter on an MCU with RAM/flash use and average current.
Custom delegate or kernel (advanced)
Write or extend a delegate, custom operator or optimized kernel for an accelerator: a deep systems signal.
What a good project write-up contains
- Goal and constraints Target device, latency/power budget, accuracy floor.
- Baseline FP32 on CPU numbers with the measurement method.
- Changes Each optimization step with before/after numbers.
- Findings What surprised you (for example an op that fell back, or throttling after 8 minutes).
- Reproducibility Scripts, versions, device build, commands, and raw data.
Curated resources (official sources first)
APIs in this area change quickly, so prefer each project's current official documentation.
Foundations
PyTorch official tutorials; Hugging Face courses and Transformers docs; a solid introductory deep learning course.
Quantization and optimization
PyTorch torchao and PT2E quantization docs; ONNX Runtime quantization docs; TensorFlow Model Optimization Toolkit; Qualcomm AIMET; Hugging Face PEFT; the GPTQ, AWQ, SmoothQuant, QuaRot/SpinQuant and knowledge distillation papers.
Edge runtimes
LiteRT (Google AI Edge) docs; AI Edge Torch; MediaPipe Solutions; ONNX Runtime Mobile docs; ExecuTorch docs and LLM examples; Netron.
Vendor SDKs
Qualcomm AI Hub and AI Engine Direct / AI Runtime docs; MediaTek NeuroPilot; Arm ML developer resources and KleidiAI; Apple Core ML and coremltools; NVIDIA TensorRT and Jetson docs; Intel OpenVINO.
On-device LLMs
llama.cpp repository and docs; MLC-LLM; ExecuTorch LLM examples; official model cards for small open models; survey papers on on-device LLMs.
Android performance and TinyML
Android developer docs on thermal APIs and ADPF; Perfetto docs; Android GPU Inspector; LiteRT for Microcontrollers and CMSIS-NN docs; the tinyML community; MLPerf Mobile and MLPerf Tiny results.
FAQ
Do I need to be good at maths or model training?
Not for most edge roles. You need to understand models (architectures, where compute and memory go, how quantization affects them) well enough to optimize and deploy them. Basic linear algebra and probability intuition are enough to start; the job is inference efficiency, not research.
Is Python a problem if I come from C, C++ or Java?
Python is a quick add for an experienced engineer, usually a few weeks of practice. Much of edge AI (runtimes, kernels, llama.cpp, delegates) is C/C++, where systems engineers are already strong, and many inference-SDK roles require strong C++ specifically.
Is a GenAI application course wasted if I want edge AI?
No. It builds the vocabulary you need to talk to model teams. Just avoid making application-layer GenAI your whole identity if your strength is systems; stack edge-specific skills on top.
How do I prove skill without AI job experience?
Benchmarked projects. A concrete, reproducible result such as "ran model X on device Y at N tokens/s with M% less energy than CPU, sustained over 10 minutes" is verifiable and rare.
Which runtime should I learn first?
LiteRT or ExecuTorch, depending on whether your models come from TensorFlow/JAX or PyTorch, plus llama.cpp for LLMs. Then learn one vendor NPU SDK to understand what happens below the runtime.
Will I lose seniority by switching?
If you position yourself as an AI application engineer, possibly. If you position yourself as an AI systems or on-device engineer, deep platform experience counts fully, because that half is what is scarce. Building skills while in a current role and preferring an internal move reduces the risk.
Quick revision
- Edge AI runs models near the data source, on a continuum from MCUs to devices, edge boxes, edge servers and the cloud; on-device AI is the extreme end.
- Drivers: latency, privacy, cost, offline use, bandwidth, data sovereignty and personalization; costs: compute, memory, bandwidth, power, thermal, storage and fragmentation.
- Most products are hybrid: a small local model for common cases with a clear rule for cloud fallback and a defined offline behaviour.
- Phone DRAM bandwidth is ~50-100 GB/s versus 2-8 TB/s on data-centre GPUs; bandwidth sets a hard latency floor of bytes moved divided by bandwidth.
- Moving data from DRAM costs orders of magnitude more energy than computing on it, so reuse, fusion and smaller data types save power.
- CPU is flexible, GPU is parallel and FP16-friendly, NPU gives the best performance per watt for low-precision tensor maths, DSP suits always-on low-power tasks, MCUs run TinyML at milliwatts.
- NPUs are efficient because of low-precision MAC arrays, on-chip SRAM reuse, fixed dataflow and ahead-of-time compilation, which is also why they want static shapes and static quantization.
- Peak TOPS = MACs x 2 x clock; real utilization is often 20-50% because of bandwidth, operator coverage, batch size 1 and throttling.
- Roofline: attainable performance = min(peak compute, bandwidth x arithmetic intensity); the ridge point is peak / bandwidth.
- Batch-1 matrix-vector work (LLM decode) has ~2 ops/byte at INT8 and is deeply memory-bound; prefill with hundreds of tokens is compute-bound.
- Model size is roughly parameters times bytes per parameter: 3B params is ~6 GB FP16, ~3 GB INT8, ~1.5-1.8 GB INT4.
- A dense layer costs about 2 x inputs x outputs FLOPs; a transformer costs about 2 x parameters FLOPs per token plus attention; depthwise-separable convs cut 3x3 conv cost ~8-9x.
- FP16 has more precision but overflows above 65,504; BF16 has FP32's range with less precision; FP8 comes as E4M3 (precision) and E5M2 (range).
- Block formats (MXFP4, NVFP4, GGUF blocks) share one scale per small group; effective bits = element bits + scale bits / group size.
- Quantization maps reals to integers with a scale and zero-point: x is approximately scale x (q - zero_point); symmetric sets zero_point to 0.
- Rounding error is about scale squared / 12 in variance; each extra bit adds ~6 dB SQNR; calibration balances rounding error against clipping error.
- Per-channel weight scales are standard for INT8; per-group scales (32-128) are standard for INT4 LLM weights; one outlier ruins a per-tensor scale.
- Calibration methods: min-max, percentile, MSE and KL/entropy; calibration data must match production.
- Static quantization fixes activation scales at compile time (NPU-friendly); dynamic computes them at runtime (CPU-friendly, no calibration).
- Integer matmul accumulates in INT32 and requantizes with a fixed-point multiplier M = s_w x s_x / s_y; weight-only W4A16 dequantizes on the fly and wins on bandwidth.
- PTQ needs only a calibration set; QAT uses fake quantization with the straight-through estimator and recovers accuracy at low bits. Try PTQ first.
- Accuracy loss usually comes from outliers, sensitive layers or bad calibration; fix with per-channel scales, clipping, equalization, mixed precision, advanced PTQ or QAT, guided by layer-wise error analysis.
- GPTQ rounds column by column with Hessian-based error compensation; AWQ scales up salient channels; SmoothQuant moves activation outliers into weights; rotations spread outliers.
- GGUF is llama.cpp's single-file format; Q4_K_M (~4.8 bits) is a popular balance; imatrix improves low-bit quants.
- CPU paths commonly use 4-bit group-wise weights with 8-bit dynamic activations; NPUs commonly use 4- or 8-bit weights with 16-bit static activations.
- Unstructured pruning compresses but rarely speeds up mobile hardware; structured and N:M sparsity give real speed-ups.
- Distillation trains a small student on a big teacher's soft outputs (temperature-scaled KL); most good small LLMs are distilled, often after pruning.
- Low-rank factorization turns m x n parameters into r(m + n); LoRA adapters let one base model serve many features.
- Hardware-aware NAS searches architectures using measured latency on the target device.
- Operator fusion and BatchNorm folding remove memory round trips at no accuracy cost; static shapes, layout choice and memory planning matter as much.
- Delegates/execution providers hand subgraphs to accelerators; every CPU fallback boundary adds copies and sync; check partition logs.
- LiteRT is the new name of TensorFlow Lite; ExecuTorch is PyTorch's on-device runtime; ONNX Runtime is framework-agnostic; QNN/QAIRT is Qualcomm's NPU SDK; TensorRT serves Jetson; llama.cpp and MLC-LLM serve LLMs.
- NNAPI is deprecated from Android 15 because of inconsistent drivers and fragmentation; use LiteRT delegates or vendor backends instead.
- The pipeline is: export, optimize, quantize, convert, validate on host, deploy, run on accelerator, profile, iterate.
- LLM prefill is compute-bound and sets time to first token (~2 x params x prompt tokens FLOPs); decode is memory-bound and sets tokens/s.
- Decode tokens/s ceiling is roughly effective memory bandwidth divided by (weight bytes + KV bytes at the current context). Long chats slow down even with fixed weights.
- KV cache = 2 x layers x KV heads x head_dim x context x bytes; at long context it can exceed the weights; GQA, KV quantization, windows and shorter context shrink it.
- Paged KV: contiguous waste is (T_max - T_actual) x KV per token; a block table wastes at most one block per sequence and can share prefix blocks.
- Prefix caching reuses KV for a byte-identical prefix; savings about P / (P + U). Put static tokens first.
- Speculative decoding drafts k tokens and verifies them in one pass: expected tokens per pass = (1 - alpha^(k+1)) / (1 - alpha); speed-up ≈ that / (1 + k c). Unchanged output distribution; helps most at low batch when acceptance is high.
- INT8 weights are ~2x smaller than FP16 and usually near-lossless. INT4 is ~4x smaller and the usual on-device decode lever, but quality loss is larger on reasoning and maths; keep embeddings and the LM head higher and always re-evaluate the task.
- Realistic phone LLMs are about 0.5-4B parameters in INT4; 7-8B needs 12-16 GB+ RAM.
- TinyML's binding constraint is usually peak activation memory in SRAM (the tensor arena), not weight size in flash.
- Always-on features use cascades: hardware trigger, tiny model on DSP/MCU, verifier on NPU, heavy work last; false triggers cost energy.
- Energy per inference (power x time) matters more than peak power; racing to idle often wins; plan for the sustained, throttled clock.
- Use thermal headroom APIs to degrade gracefully (smaller model, lower fps) before the OS throttles hard.
- Federated learning shares updates, not data (FedAvg); secure aggregation hides individual updates; differential privacy bounds what any individual's data can reveal.
- Benchmark with warm-up, hundreds of runs, p50/p95/p99, sustained 10-30 minute runs, baseline rows, device tiers and real silicon only.
- Pre- and post-processing can cost more than the model; profile the whole pipeline inside the real app.
- Ship per-device-tier model variants, monitor latency and fallback in production, and keep a remote kill switch.
Glossary
- ADPF
- Android Dynamic Performance Framework: APIs for performance hints and thermal awareness so apps can sustain steady workloads.
- Arithmetic intensity
- Operations performed per byte of memory traffic; decides compute-bound versus memory-bound behaviour.
- AWQ
- Activation-aware Weight Quantization: a PTQ method that scales up the most important weight channels (found from activation magnitudes) before low-bit quantization.
- BF16
- bfloat16: a 16-bit float with FP32's 8-bit exponent and a 7-bit mantissa; wide range, low precision.
- Block (microscaling) format
- A format in which a small block of values shares one scale, such as MXFP4, NVFP4 or GGUF blocks.
- Calibration set
- A small set of representative inputs used during PTQ to measure activation ranges.
- Cascade
- A chain of models of increasing cost where each stage runs only if the previous one fires or is unsure.
- Core ML
- Apple's on-device inference framework, running on CPU, GPU and the Neural Engine.
- Cross-layer equalization
- Rescaling consecutive layers so per-channel weight ranges are even, improving quantization without changing FP32 outputs.
- Decode
- The token-by-token generation phase of an LLM; memory-bandwidth-bound.
- Delegate
- A runtime plug-in that executes supported parts of a model graph on an accelerator such as a GPU or NPU.
- Depthwise-separable convolution
- A per-channel spatial convolution followed by a 1x1 convolution; far cheaper than a standard convolution.
- Differential privacy
- A mathematical guarantee, parameterized by epsilon and delta, that one individual's data barely changes a computation's output.
- Distillation
- Training a small student model to imitate a larger teacher model's outputs.
- DSP
- Digital signal processor: a power-efficient processor for fixed-point vector maths, often used for audio and always-on models.
- Dynamic quantization
- Quantization where activation scales are computed at runtime for each input.
- Edge server (MEC)
- A server near users, on premises or at a telecom site (multi-access edge computing), offering low-latency compute.
- Energy per inference
- Average power above idle multiplied by latency; the right metric for comparing backends on battery devices.
- Execution provider
- ONNX Runtime's term for a hardware backend (CPU, QNN, Core ML, XNNPACK, TensorRT and so on).
- ExecuTorch
- PyTorch's runtime for on-device inference, using models exported with torch.export into .pte files.
- Fake quantization
- Quantize-then-dequantize operations inserted during QAT to simulate rounding in floating point.
- Fallback
- Running an operator on the CPU because the accelerator does not support it.
- Federated learning
- Training across many devices by sharing model updates instead of raw data, then aggregating them on a server.
- FLOPs
- Floating-point operations; a measure of compute work (one multiply-accumulate is about two FLOPs).
- FP16
- Half-precision float with a 5-bit exponent and 10-bit mantissa; maximum value 65,504.
- FP8
- 8-bit floating point, in E4M3 (more precision, max 448) and E5M2 (more range) variants.
- GGUF
- llama.cpp's single-file model format containing quantized weights, tokenizer and metadata.
- GPTQ
- A post-training quantization algorithm that quantizes weights column by column while compensating the error using second-order information.
- GQA
- Grouped-query attention: several query heads share one key/value head, shrinking the KV cache.
- Hexagon NPU
- Qualcomm's neural processor, combining scalar, vector (HVX) and tensor (HMX) units.
- Importance matrix
- Per-weight importance statistics from calibration text, used by llama.cpp to reduce low-bit quantization error.
- INT4
- 4-bit integer format with 16 levels; used mainly for LLM weights with per-group scales.
- Jetson
- NVIDIA's family of embedded modules combining Arm CPUs, a CUDA GPU with tensor cores and deep-learning accelerators.
- K-quants
- llama.cpp quantization types using super-blocks with quantized sub-block scales, such as Q4_K and Q6_K.
- KV cache
- Stored keys and values of previous tokens so an LLM does not recompute them at each step.
- Paged attention
- Storing the KV cache in fixed-size blocks mapped by a block table, so waste is at most one block per sequence and prefixes can be shared.
- Prefix caching
- Reusing computed KV for a byte-identical prompt prefix (system prompt, tools, shared document) to cut prefill and TTFT.
- LiteRT
- Google's on-device runtime, formerly TensorFlow Lite, using .tflite models.
- LiteRT for Microcontrollers
- The MCU variant of LiteRT (formerly TFLite Micro), running models from a fixed memory arena without an OS.
- llama.cpp
- An open-source C/C++ engine for running quantized LLMs on CPUs and GPUs of laptops and phones.
- LoRA
- Low-rank adaptation: fine-tuning with small low-rank update matrices while freezing the base model.
- MAC
- Multiply-accumulate operation, the basic building block of neural network maths.
- MediaPipe
- Google's framework of ready-made on-device ML pipelines built on LiteRT.
- Memory bandwidth
- Bytes per second that can move between DRAM and compute units; the main limit for memory-bound layers.
- Mixed precision
- Using different numeric formats for different layers, for example INT8 for most and FP16 for sensitive ones.
- MLC-LLM
- A compiler-based LLM deployment stack built on TVM, targeting mobile GPUs and other backends.
- N:M sparsity
- A structured sparsity pattern with N zeros in every M consecutive weights (for example 2:4) that hardware can accelerate.
- NAS
- Neural architecture search: automated search for architectures, often scored by measured latency on the target device.
- NF4
- 4-bit NormalFloat: a non-uniform 4-bit format whose levels follow a normal distribution; used by QLoRA.
- NNAPI
- Android Neural Networks API, a vendor-neutral accelerator interface, deprecated from Android 15.
- NPU
- Neural processing unit: a dedicated accelerator for low-precision tensor maths with high performance per watt.
- ONNX
- Open Neural Network Exchange: an open model format used to move models between frameworks and runtimes.
- ONNX Runtime
- A cross-platform inference engine for ONNX models with pluggable execution providers.
- Operator fusion
- Merging several consecutive operations into one kernel to cut memory traffic and launch overhead.
- Per-channel quantization
- Using a separate scale for each output channel of a weight tensor for better accuracy.
- Per-group quantization
- Using a separate scale for each small group (for example 32-128) of consecutive weights.
- Prefill
- The phase where an LLM processes the whole prompt in parallel; compute-bound; sets time to first token.
- Pruning
- Removing weights, channels, heads or layers that contribute little to the output.
- PTQ
- Post-training quantization: quantizing an already-trained model using calibration data.
- QAT
- Quantization-aware training: training with simulated quantization so the model tolerates low precision.
- QNN / QAIRT
- Qualcomm AI Engine Direct, packaged as the Qualcomm AI Runtime SDK, for running models on the Hexagon NPU, Adreno GPU and CPU.
- Race to idle
- Finishing work quickly so hardware can return to low-power sleep states, often saving energy overall.
- Real-time factor
- Processing time divided by input audio duration; below 1 means faster than real time.
- Ridge point
- Peak compute divided by memory bandwidth; the arithmetic intensity where a workload becomes compute-bound.
- Roofline model
- A chart relating attainable performance to arithmetic intensity, bounded by memory bandwidth and peak compute.
- Rotation methods
- Quantization techniques (QuaRot, SpinQuant) that apply orthogonal transforms to spread outliers across channels.
- Scale and zero-point
- The two parameters that map real values to integers in quantization.
- Secure aggregation
- A cryptographic protocol that lets a server see only the sum of many device updates, never an individual one.
- SmoothQuant
- A technique that divides activation channels and multiplies weight rows by the same factor to move outliers into weights, enabling W8A8.
- Speculative decoding
- A draft model proposes tokens that the main model verifies in parallel, speeding decode without changing outputs.
- SQNR
- Signal-to-quantization-noise ratio in decibels; a standard per-tensor measure of quantization damage.
- Static quantization
- Quantization where activation scales are fixed ahead of time from calibration data.
- Straight-through estimator
- A QAT trick that treats rounding as the identity in the backward pass so gradients can flow.
- Tensor arena
- The fixed, pre-allocated SRAM buffer holding all activations in LiteRT for Microcontrollers.
- TensorRT
- NVIDIA's inference optimizer and runtime that builds device-specific engines with fusion and low precision.
- Thermal headroom
- An OS forecast of how close the device is to thermal throttling, used to degrade workloads gracefully.
- Thermal throttling
- The OS lowering clock speeds to limit heat, reducing sustained performance.
- TinyML
- Machine learning on microcontrollers with kilobytes of memory and milliwatt power budgets.
- TOPS
- Trillions of operations per second; a peak compute figure, usually for INT8.
- TTFT
- Time to first token: delay from submitting a prompt to receiving the first generated token.
- W4A16
- Notation for 4-bit weights with 16-bit activations; similar forms describe other weight/activation precisions.
- XNNPACK
- A highly optimized CPU inference library used as the default CPU backend by LiteRT, ExecuTorch and ONNX Runtime.
Interview questions
Fundamentals
What is on-device AI and why would you use it?
On-device AI runs model inference locally on the phone, watch, car or IoT device instead of on a server. The main reasons are: lower and more predictable latency (no network round trip), privacy (raw data stays on the device), lower serving cost at scale, offline availability, and reduced bandwidth. The costs are tight memory, compute, power and thermal budgets, hardware fragmentation and slower model update cycles.
What is the difference between edge AI, on-device AI and cloud AI?
Cloud AI runs in data centres. Edge AI is any inference close to where data is produced: on the device itself, on a gateway or embedded box on the same site, or on an edge server at a nearby network location (on-premises or telecom MEC). On-device AI is the most local form of edge AI, running on the device that owns the sensor and the user. Moving toward the device lowers latency and data exposure but shrinks compute, memory and power budgets; moving toward the cloud gives capability and easy updates at the cost of latency, per-request cost and privacy exposure.
What are the main constraints of running models on mobile devices?
- Memory capacity: a few GB of RAM shared with the OS and other apps.
- Memory bandwidth: LPDDR around 50-100 GB/s, far below data-centre GPUs, which limits memory-bound layers and LLM decode.
- Compute: far below data-centre GPUs, especially sustained.
- Power and battery: every milliwatt matters.
- Thermal: passive cooling means sustained performance drops after minutes.
- Fragmentation: thousands of devices, chipsets and driver versions.
- App size and update cadence: model weights add to download size and ship with app or system updates.
Compare CPU, GPU, NPU and DSP for inference.
CPU: supports every operator and is easy to debug, but has the lowest performance per watt for large tensor maths. GPU: highly parallel and good at FP16, but needs initialization/shader compilation and competes with rendering. NPU: dedicated low-precision matrix engines with on-chip memory; best performance per watt, but a limited operator set, vendor-specific tools and a preference for static shapes and static quantization. DSP: very power-efficient fixed-point vector processor, suited to always-on audio and sensor models. Real deployments often mix them: pre-processing on CPU/GPU, model on NPU, fallback on CPU.
What is an NPU, in one minute?
A neural processing unit is a dedicated accelerator for the tensor maths in neural networks, mainly matrix multiplications and convolutions at low precision (INT8, INT4, FP16). It packs many small multiply-accumulate units, keeps data in large on-chip SRAM so each value fetched from DRAM is reused many times, and runs a fixed dataflow planned by an ahead-of-time compiler instead of executing general instructions. That gives the best performance per watt on the SoC, at the cost of flexibility: unsupported operators, data types or dynamic shapes fall back to the CPU.
What does TOPS mean and why is it misleading?
TOPS is trillions of operations per second, computed as number of MAC units x 2 x clock frequency, usually quoted for INT8 (sometimes INT4 or with sparsity, which inflates it). It is a theoretical peak. Real throughput depends on memory bandwidth, operator coverage, how well layers map to the MAC array, batch size and thermal limits, so real utilization of 20-50% is common and memory-bound workloads such as LLM decode reach far less. Compare chips by benchmarking your model, not by TOPS.
What is memory bandwidth and why does it matter so much at the edge?
Memory bandwidth is how many bytes per second can move between DRAM and the compute units. At batch size 1, each weight is typically read once per inference, so the minimum latency is bytes moved divided by bandwidth regardless of compute. Phones have roughly 50-100 GB/s shared by CPU, GPU, NPU, display and camera, which is why memory-bound layers and LLM decode are limited by bandwidth, why quantization speeds them up almost in proportion to bytes saved, and why data movement also dominates energy.
What is quantization?
Quantization represents weights and/or activations with fewer bits, for example INT8 instead of FP32. Each tensor (or channel/group) gets a scale and zero-point: q = round(x / scale) + zero_point, and x is approximately scale x (q - zero_point). Benefits: 4x smaller for INT8 (8x for INT4), less memory bandwidth and energy, and access to fast integer units on NPUs/DSPs. The cost is some rounding and clipping error, which may reduce accuracy.
What is the difference between PTQ and QAT?
Post-training quantization (PTQ) quantizes a trained model using only a small calibration dataset to find activation ranges; it is fast and needs no training pipeline. Quantization-aware training (QAT) inserts simulated ("fake") quantization during training or fine-tuning so the model learns to be robust to rounding; it needs data and compute but preserves accuracy better, especially at INT4 or for sensitive models. Standard practice: try PTQ first, move to QAT if accuracy loss is too high.
What is a calibration dataset and why does it matter?
It is a small set (typically hundreds of samples) of representative inputs that the converter runs through the model to record activation ranges for static quantization. If it does not match production data (for example only daytime images), ranges will be wrong, values will clip or lose resolution, and accuracy drops. It should cover the real distribution, including edge cases.
What is the difference between FP32, FP16, BF16 and INT8?
FP32 (1 sign, 8 exponent, 23 mantissa bits) is the training default with about 7 decimal digits of precision. FP16 (1/5/10) halves size and has about 3 digits of precision but a maximum of 65,504, so it can overflow. BF16 (1/8/7) also halves size and keeps FP32's range, trading precision; it converts from FP32 by truncation and rarely overflows. INT8 is an integer format with 256 uniform levels that needs an external scale and zero-point; it is 4x smaller than FP32 and runs on integer engines.
What is LiteRT?
LiteRT is Google's on-device inference runtime, the new name (since 2024) for TensorFlow Lite. It runs .tflite FlatBuffer models, uses XNNPACK for optimized CPU execution, and offers a GPU delegate and vendor NPU delegates/accelerators. Models can come from TensorFlow, Keras, JAX, or PyTorch via AI Edge Torch. A microcontroller variant runs models without an OS.
What is a delegate in LiteRT?
A delegate is a plug-in that takes over execution of the parts of a model graph it supports, running them on a GPU, NPU or DSP. The runtime partitions the graph: supported subgraphs go to the delegate, the rest runs on the CPU. Fewer partitions means fewer copies and synchronizations, so the ideal is full delegation.
What is NNAPI and what is its current status?
The Android Neural Networks API (introduced in Android 8.1) let runtimes describe a model graph that vendor drivers executed on accelerators. It suffered from inconsistent driver quality, limited operator coverage and updates tied to OS releases. Google deprecated it starting with Android 15. New work should use LiteRT with GPU or vendor NPU delegates, or runtimes that call vendor SDKs directly (ONNX Runtime QNN execution provider, ExecuTorch vendor backends).
What is ONNX and ONNX Runtime?
ONNX is an open model format with a standard operator set that many frameworks can export to. ONNX Runtime executes ONNX models on many platforms, including mobile, using execution providers such as CPU, XNNPACK, QNN (Qualcomm NPU), Core ML, TensorRT and OpenVINO. It is a good choice when you want one framework-agnostic format across platforms.
What is ExecuTorch?
ExecuTorch is PyTorch's official on-device runtime. You capture the model with torch.export, lower it (optionally partitioning to backends such as XNNPACK, Qualcomm, MediaTek, Vulkan, Core ML or Arm), and save a .pte program that a small C++ runtime executes. It keeps you inside the PyTorch ecosystem and is widely used for on-device LLMs.
What is llama.cpp and GGUF?
llama.cpp is an open-source C/C++ engine for running LLMs efficiently on CPUs and GPUs of laptops and phones, with heavy use of quantization and hand-optimized kernels. GGUF is its single-file format that holds quantized weights, tokenizer and metadata. Common quant types include Q8_0, Q5_K_M and Q4_K_M.
What is TensorRT and where is it used at the edge?
TensorRT is NVIDIA's inference optimizer and runtime. It takes a model (usually ONNX), fuses layers, selects precisions (FP16, INT8, FP8, INT4 depending on hardware) with calibration, auto-tunes kernels for the exact GPU, and serializes a device-specific engine. At the edge it is the standard runtime on Jetson modules, where it can also target the deep-learning accelerator (DLA) cores; DeepStream builds multi-camera video pipelines on top of it.
How do you estimate a model's size from its parameter count?
Size is roughly parameters x bytes per parameter. 1B parameters is about 4 GB in FP32, 2 GB in FP16, 1 GB in INT8 and about 0.5-0.6 GB in INT4 (group scales add some overhead). Runtime memory also needs activations, KV cache for LLMs and runtime buffers.
What are FLOPs and MACs?
FLOPs count floating-point operations and measure compute cost. A multiply-accumulate (MAC) is one multiply plus one add, so 1 MAC is about 2 FLOPs. A dense layer with N inputs and M outputs costs about N x M MACs (2 x N x M FLOPs). Always check which unit a vendor or paper uses, since they differ by 2x.
What is pruning?
Pruning removes parameters that contribute little. Unstructured pruning zeroes individual weights (good compression, little speed-up on typical mobile hardware). Structured pruning removes whole channels, filters, heads or layers (real speed-up everywhere, more accuracy loss, usually needs fine-tuning). N:M semi-structured sparsity like 2:4 gets hardware acceleration on some chips.
What is knowledge distillation?
A large teacher model's output probabilities (soft labels, usually softened with a temperature) are used to train a smaller student model. Soft labels carry information about how similar classes are, so the student learns more than from hard labels alone. It is a key technique behind strong small models for on-device use.
What is operator fusion?
Combining several consecutive ops into one kernel, for example Conv + BatchNorm + ReLU or a fused attention kernel. It avoids writing intermediate tensors to memory and reduces kernel launches, giving speed-ups and energy savings with no accuracy change. BatchNorm can also be folded directly into convolution weights.
What metrics do you track for an on-device model?
Latency (p50/p95/p99) including pre/post-processing, initialization time, throughput or FPS, peak memory, model size, accuracy drop versus the FP32 reference, power/energy per inference, thermal behaviour over sustained use, and delegation coverage. For LLMs add time to first token, prefill tokens/s and decode tokens/s.
Why report p95/p99 latency instead of the average?
Users notice the slow outliers, not the mean. A camera feature with a 12 ms average but a 45 ms p99 misses a 33 ms frame deadline once every hundred frames and visibly stutters. Tail latency exposes scheduling contention, garbage collection, cache misses and throttling that averages hide. Stable tail estimates need hundreds to thousands of runs.
What is energy per inference and why is it the right metric?
Energy per inference is the average power above the idle baseline multiplied by the latency, usually in millijoules. It captures both how hard and how long the hardware works, which is what drains the battery. A backend with higher peak power can still use less energy if it finishes much faster, so comparing peak power or latency alone can pick the wrong backend.
What is thermal throttling and why does it matter for AI?
When the device heats up, the thermal governor lowers CPU/GPU/NPU clock frequencies to protect hardware and keep the skin temperature comfortable. AI workloads that run continuously (camera effects, long LLM sessions, video calls) can see latency grow or FPS drop after a few minutes. You must design and test for sustained performance, not a single fast run.
What is TinyML?
TinyML is machine learning on microcontrollers: devices with tens to hundreds of KB of SRAM, up to a few MB of flash, no DRAM and often no OS, running at milliwatts. Typical tasks are keyword spotting, activity recognition, simple vision and vibration anomaly detection. Models are INT8, often under 100 KB, run with LiteRT for Microcontrollers, CMSIS-NN or a micro-NPU, and the main constraint is peak activation memory in SRAM.
What is a Jetson-class device and when would you choose one?
Jetson modules combine Arm CPUs, a CUDA GPU with tensor cores, deep-learning accelerators and shared LPDDR5 memory in a 7-60 W envelope. They run full Linux and the CUDA/TensorRT stack, so they suit robots, drones, industrial vision and multi-camera analytics where you need tens to hundreds of TOPS, flexibility for many models and a mains or large battery power supply. For a phone-sized power budget or a cost-sensitive consumer device, an SoC NPU or small accelerator is a better fit.
What is MediaPipe?
MediaPipe (part of Google AI Edge) provides ready-made, cross-platform on-device ML pipelines, such as face, hand and pose landmarks, object detection, image segmentation, text and audio classification and an LLM inference API. It runs on top of LiteRT and handles pre/post-processing, so it is the fastest way to ship common tasks.
What is Core ML?
Core ML is Apple's on-device inference framework. Models are converted with coremltools into .mlmodel or .mlpackage and run on the CPU, GPU or Apple Neural Engine, with the framework choosing the compute unit. It plays the same role on Apple devices that LiteRT and vendor SDKs play on Android.
What is federated learning, in simple terms?
Federated learning trains a shared model across many devices without collecting their raw data. Each selected device downloads the current model, trains briefly on local data, and sends back only a model update; the server averages updates (weighted by data size) into a new global model and repeats. It is used for keyboards, wake words and ranking, usually combined with secure aggregation and differential privacy because updates can still leak information.
What is a small language model and what sizes run on phones today?
A small language model is a decoder-only LLM of roughly 0.1-4B parameters, usually distilled or pruned from larger models, designed to fit device memory and bandwidth. On phones, 0.5-1B models run on most mid-range devices, 2-4B models on recent flagships with 8-12 GB RAM, both in INT4. They are best at constrained tasks: summarizing, rewriting, extraction, classification and short replies, rather than open-ended knowledge questions.
Going deeper
Symmetric versus asymmetric quantization: when do you use each?
Symmetric quantization fixes zero-point at 0 and uses a range centred on zero; the maths is simpler and faster (no zero-point cross terms in the integer dot product), so it is the default for weights, which are roughly zero-centred. Asymmetric quantization allows any zero-point, using the integer range efficiently for skewed data, such as activations after ReLU that are all non-negative. Many toolchains use symmetric per-channel weights with asymmetric per-tensor activations.
Quantize the weights [-0.8, 0.3, 1.2, -0.05] to symmetric INT8 by hand.
Scale = max|w| / 127 = 1.2 / 127 ≈ 0.009449, zero-point 0. Divide and round: -0.8 / 0.009449 = -84.7 → -85; 0.3 → 31.75 → 32; 1.2 → 127; -0.05 → -5.3 → -5. Dequantize: -0.8031, 0.3024, 1.2000, -0.0472. The maximum rounding error is scale / 2 ≈ 0.0047. If one weight were 12.0 instead, the scale would grow tenfold and -0.05 would dequantize to -0.094, a 90% error, which is why per-channel or per-group scales are used.
Per-tensor, per-channel and per-group quantization: what is the difference?
Per-tensor uses one scale for the whole tensor: cheapest, least accurate. Per-channel uses one scale per output channel: standard for convolution and linear weights because channel ranges differ a lot, and free at runtime because the scale applies to each output's accumulator. Per-group (block-wise) uses one scale per small block of weights (for example 32-128): standard for INT4 LLM weights, where per-channel is not fine enough. Finer granularity improves accuracy at the cost of storing more scales (effective bits = bits + scale bits / group size) and slightly more complex kernels.
INT8 versus INT4 for on-device models: how do you choose?
INT8 (or FP8) weights are about 2× smaller than FP16 and usually near-lossless after PTQ; for CNNs the typical top-1 drop is under 1% with per-channel scales. INT4 weight-only is about 4× smaller and is the usual decode-speed lever on phones, but quality loss is larger and shows first on reasoning, maths, rare languages and long context. Use GPTQ/AWQ or k-quants, keep embeddings and the LM head at 6-8 bits, and re-measure the product task. NPUs often want W8A8 for CNNs and W4A16 for LLMs (static 16-bit activations). Two "INT4" checkpoints are not interchangeable: group size, leftover high-precision tensors and the algorithm matter more than the label.
What is the difference between weight-only and full integer quantization?
Weight-only quantization (for example W4A16 or W8A16) stores weights in low precision and dequantizes them on the fly, keeping activations in FP16. It cuts memory and bandwidth, ideal for memory-bound LLM decode, but compute still happens in floating point. Full integer quantization (W8A8) also quantizes activations, allowing integer-only execution on NPUs and DSPs for maximum speed and efficiency, but it is more sensitive to activation outliers and needs good calibration.
Static versus dynamic quantization?
Static quantization fixes activation scales ahead of time using calibration data; no range calculation happens at runtime, so it is fastest and suits NPUs, which need compile-time parameters. Dynamic quantization stores weights as integers but computes activation scales at runtime per input (often per token); it needs no calibration and adapts to input, but adds overhead and is mainly a CPU/GPU technique (common for LSTMs and transformer linear layers, and for 8-bit dynamic activations with 4-bit weights in LLMs).
What calibration methods exist for choosing activation ranges?
Min-max uses the observed extremes (simple, outlier-sensitive); moving-average min-max smooths across batches; percentile clipping (for example 99.99th) ignores rare extremes; MSE-based search picks the clipping threshold that minimizes quantization error; entropy/KL calibration picks the threshold whose quantized histogram best matches the original distribution (the classic TensorRT INT8 method). The trade-off is always rounding error (wide range) against clipping error (narrow range); MSE or percentile usually beat raw min-max for activations.
Why do some layers not quantize well?
Layers with wide or outlier-heavy value ranges (for example certain transformer activations), layers where small errors amplify (softmax inputs, layer norms), and the first and last layers (raw input and final logits) are commonly sensitive. Embedding tables and LM heads in LLMs can also be sensitive. The fix is mixed precision (keep them in INT16/FP16), per-channel scales, outlier handling like SmoothQuant or rotations, or QAT.
Why are activations harder to quantize than weights?
Weights are fixed and known in advance, roughly bell-shaped per channel, and can use fine-grained scales that factor out of the dot product. Activations depend on the input, so their range must be estimated (calibration) or computed at runtime; transformers also produce a few channels with magnitudes 10-100x larger than the rest for almost every token. Because those outliers lie along the reduction dimension, per-channel activation scales cannot be factored out of an integer matmul, so a single per-tensor scale either clips outliers or crushes normal values.
What are GPTQ, AWQ and SmoothQuant?
All are advanced PTQ techniques for transformers. GPTQ quantizes weights column by column and adjusts remaining weights to compensate for the error, using second-order (Hessian) information from calibration data. AWQ identifies the small fraction of weight channels that matter most (based on activation magnitudes) and scales them up before quantization, folding the inverse scale into the activations, so their relative error shrinks. SmoothQuant migrates quantization difficulty from activations (which have outliers) to weights by a mathematically equivalent per-channel scaling, making W8A8 practical.
What is the straight-through estimator and why does QAT need it?
Rounding has zero gradient almost everywhere, so backpropagation through fake-quantization nodes would stop learning. The straight-through estimator treats the quantize-dequantize step as the identity in the backward pass (passing the gradient through unchanged) while keeping true rounding in the forward pass, and zeroes the gradient for values outside the clipping range. The model therefore learns weights that sit well on the quantization grid. Variants such as LSQ also learn the scale parameters.
FP8 E4M3 versus E5M2: what is the difference and where is each used?
Both are 8-bit floats. E4M3 has 4 exponent and 3 mantissa bits, maximum 448, better precision; it is used for weights and activations in inference and forward passes. E5M2 has 5 exponent and 2 mantissa bits, maximum 57,344, more range but coarser; it is used mainly for gradients in training. Both usually still use a per-tensor or per-block scale to place values in the best part of the range. Unlike INT8, FP8's spacing is non-uniform, which fits bell-shaped distributions with outliers better.
What are block (microscaling) formats such as MXFP4?
Block formats let a small block of values share one scale so each element can use very few bits. MXFP4 (from the OCP microscaling specification) uses blocks of 32 FP4 (E2M1) elements sharing an 8-bit power-of-two scale, about 4.25 bits per value; NVFP4 uses blocks of 16 with an FP8 scale plus a per-tensor scale, about 4.5 bits. They standardize in hardware what per-group integer quantization and GGUF blocks do in software: fine-grained scaling that handles varying ranges at very low bit-widths.
What is NF4 and how is it different from INT4?
INT4 has 16 evenly spaced levels. NF4 (4-bit NormalFloat, from QLoRA) has 16 levels placed at quantiles of a normal distribution, so more levels sit near zero where most weights are, giving lower error for bell-shaped weights at the same bit count. Values are stored as 4-bit indices into that table, with per-block scales (and optionally quantized scales, "double quantization"). Because it needs a lookup before maths, it is mainly a storage format, used for QLoRA fine-tuning rather than for integer NPU execution.
What are llama.cpp k-quants and what does Q4_K_M mean?
K-quants group 256 weights into a super-block of 8 sub-blocks of 32, with quantized per-sub-block scales (and minimums) plus an FP16 super-block scale: two-level scaling keeps metadata overhead low while keeping fine granularity. Q4_K is the 4-bit variant. The suffix S/M/L is a mix level: Q4_K_M uses Q4_K for most tensors but keeps some sensitive ones (such as attention value projections and part of the FFN down projections) at Q6_K, for about 4.8 bits per weight overall. An importance matrix computed from calibration text can further reduce error.
How does depthwise-separable convolution reduce compute?
A standard KxK convolution mixes spatial and channel information at once: about K x K x C_in x C_out MACs per output pixel. Depthwise-separable convolution splits it into a depthwise KxK filter per channel (K x K x C_in) and a 1x1 pointwise convolution (C_in x C_out). For 3x3 kernels and many channels this cuts compute by roughly 8-9x (for example 231 M to 27.5 M MACs for a 64-to-128-channel layer at 56x56), which is why MobileNet-style networks use it. The catch: depthwise layers have low arithmetic intensity and map poorly onto large MAC arrays.
What is the roofline model and how do you use it?
It plots attainable performance against arithmetic intensity (ops per byte). Below a ridge point (peak compute divided by bandwidth), performance is limited by memory bandwidth (the sloped line); above it, by peak compute (the flat line). You compute a layer's intensity and see which bound applies. If memory-bound, reduce bytes (quantization, fusion, better cache reuse); if compute-bound, use faster units, lower precision maths or fewer FLOPs.
Compute the ridge point for a 40 TOPS NPU with 60 GB/s and classify a batch-1 INT8 matrix-vector layer.
Ridge point = 40e12 / 60e9 ≈ 667 ops per byte. A 4096x4096 INT8 matrix-vector multiply does 2 x 40962 ≈ 33.6 M ops and reads 16.8 MB of weights, so its intensity is 2 ops per byte, far below the ridge. Attainable performance is 60 GB/s x 2 = 120 GOPS, 0.3% of peak, and the layer takes about 0.28 ms. The same weights applied to 512 tokens at once reach roughly 800 ops per byte and become compute-bound.
Why does a mobile NPU rarely reach its advertised TOPS?
Peak TOPS assumes perfect utilization of all MAC units on INT8 with data always ready. Real models have memory-bound layers, batch size 1 (little data reuse), ops that do not map to the tensor unit (softmax, normalization, reshapes, depthwise convolutions), CPU fallback boundaries, layout conversions, synchronization and thermal clock limits. Real utilization of 20-50% is common; always benchmark the actual model.
How do you check whether a model is fully delegated to the accelerator?
Read the runtime's delegation or partition log (LiteRT prints how many nodes and partitions were delegated; ONNX Runtime logs node assignment per execution provider; ExecuTorch reports partitioner results). Use per-op profiling in benchmark tools to see where time goes. Vendor profilers show per-layer NPU timing. Visualizing the graph in Netron helps identify unsupported ops.
What causes CPU fallback and how do you fix it?
Causes: an operator the accelerator does not support, an unsupported data type (for example FP32 on an INT8-only NPU), dynamic or unusual shapes, unsupported attributes (such as a strange padding mode) or ops above a size limit. Fixes: replace the op with an equivalent supported pattern, fix shapes to static, quantize consistently, move pre/post-processing out of the graph, update to a newer runtime/SDK, or write a custom op for the accelerator as a last resort.
Why are static shapes preferred on NPUs?
Accelerator compilers plan memory, tiling and kernel selection ahead of time. With known shapes they can allocate buffers once and pick the best kernels. Dynamic shapes force either recompilation, generic slower kernels or CPU fallback. For variable-length inputs, a common trick is to compile a few fixed sizes (buckets) and pad inputs to the nearest one; LLM stacks often compile separate prefill (many tokens) and decode (one token) graphs.
How do you fold BatchNorm into a convolution?
At inference BN is a fixed per-channel affine map: y = γ(z - μ)/√(σ2 + ε) + β, where z = Wx + b. Substituting gives new weights W' = W · γ/√(σ2 + ε) and bias b' = (b - μ)·γ/√(σ2 + ε) + β, per output channel. The folded conv produces identical outputs with one op fewer. Fold before quantization so the quantizer sees the final weight ranges.
NHWC versus NCHW: why does layout matter?
Layout determines which elements are adjacent in memory. NHWC (channels last) keeps all channels of a pixel together, which suits SIMD and many mobile CPU/NPU kernels that vectorize across channels; NCHW (channels first) is the traditional GPU/PyTorch training layout. If the model's layout differs from what the accelerator wants, the runtime inserts transposes, which cost memory traffic and may not be delegated. Export in the deployment layout and check the graph for stray transposes.
What is hardware-aware neural architecture search?
NAS automatically searches over architecture choices (block types, kernel sizes, widths, depths, resolutions). Hardware-aware NAS adds measured or predicted latency (or energy) on the target device to the objective, so the result is Pareto-optimal for that hardware rather than for FLOPs, which correlate poorly with real latency. Efficient approaches train one weight-sharing super-network once and extract sub-networks sized for each device tier without retraining.
What are the prefill and decode phases of LLM inference?
Prefill processes the whole prompt in parallel, building the KV cache; it uses matrix-matrix multiplies, is compute-bound and determines time to first token. Decode generates one token at a time, reading all weights per token with matrix-vector multiplies; it is memory-bandwidth-bound and determines streaming tokens/s. Optimizations differ: NPUs and more compute help prefill; quantization and bandwidth help decode.
What is the KV cache and how do you calculate its size?
It stores each layer's keys and values for all previous tokens so they are not recomputed during decode. Size = 2 x layers x KV heads x head dimension x context length x bytes per value (x batch). For a model with 28 layers, 8 KV heads, head dimension 128 and 4096 tokens in FP16, that is about 470 MB. It grows linearly with context length and batch size. A contiguous reservation sized for Tmax wastes (Tmax − Tactual) × KV per token; a paged (block-allocated) cache wastes at most one block and can share prefix blocks across requests.
How do GQA and MQA help on-device LLMs?
In multi-head attention every query head has its own key and value head. Multi-query attention (MQA) shares a single K/V head across all query heads; grouped-query attention (GQA) shares one K/V head per group of query heads. Both shrink the KV cache (for example 3x-8x) and reduce memory traffic during decode, with little quality loss for GQA. That is why most modern small LLMs use GQA.
How would you benchmark a model fairly on Android?
- Fix the environment: same device and build, airplane mode, stable brightness, known temperature, consistent charge state.
- Separate initialization time from inference time; warm up 10-50 runs.
- Run hundreds of iterations and report p50/p95/p99 and variance.
- Run sustained tests for 10-30 minutes and plot latency and temperature.
- Measure the whole pipeline, including pre/post-processing.
- Include a baseline (FP32 on CPU) and each optimization step.
- Repeat across device tiers, on physical devices only.
How do you measure power consumption of inference?
Options: on-device power rails (On-Device Power Monitor) captured in Perfetto traces on supported devices; an external power monitor connected in place of the battery for precise lab measurements; battery statistics (batterystats, fuel gauge) for coarse field data. Always subtract an idle baseline with the same screen state, and report energy per inference (average power x latency) plus sustained battery drain per hour for continuous features.
Why can a faster accelerator use less battery even at higher peak power?
Energy is power multiplied by time. If an NPU draws 2 W for 10 ms (20 mJ) and the CPU draws 1 W for 60 ms (60 mJ), the NPU uses a third of the energy despite doubling peak power. Finishing quickly also lets the system return to low-power idle states ("race to idle"). Energy per inference is therefore the right comparison metric.
Bundle the model in the app or download it later?
Bundling guarantees availability and is simple, but increases install size and ties model updates to app updates. Downloading on demand keeps the install small, allows per-device variants and independent updates, but adds first-use delay, needs versioning, integrity checks, storage management and an offline fallback. A system-provided model (such as an OS AI service) avoids shipping weights but only works on supported devices. Large models often use downloads; small critical models are bundled.
What is LoRA and why is it useful on device?
LoRA fine-tunes a model by learning two small low-rank matrices per adapted layer (W' = W + BA) while the base weights stay frozen. The adapter is tiny (a few MB to tens of MB) compared with the base model (GBs). On device, one shared base model can serve many features by loading different adapters, saving storage and memory; OS AI services use this for per-feature specialization. QLoRA trains LoRA adapters on top of a 4-bit quantized base, reducing fine-tuning memory.
What is speculative decoding?
A small, fast draft model proposes several next tokens; the large target model checks all of them in one parallel forward pass and accepts the longest correct prefix. Because verification is parallel (compute-bound, like prefill) and decode is memory-bound, accepted tokens cost little extra. Expected tokens per target pass = (1 − αk+1) / (1 − α); speed-up ≈ that / (1 + k · c). The output distribution matches the target model exactly, and speed-ups of 1.5-3x are common when acceptance is high and batch size is small.
How do pre- and post-processing affect end-to-end latency?
Resizing, color conversion (for example YUV to RGB), normalization, tokenization, non-max suppression and decoding can take as long as, or longer than, the model itself, especially on the CPU in Java/Kotlin. Optimizations: use GPU or native code (for example vectorized C++), fuse normalization into the model, request the right camera format and resolution, avoid copies with direct buffers, and pipeline stages across frames.
What is the difference between central and local differential privacy?
In central DP, devices send (securely aggregated) data or updates to a trusted aggregator, which adds calibrated noise to the combined result; accuracy is good because noise is added once. In local DP, each device randomizes its own data before sending, so no one needs to be trusted, but the total noise is much larger and needs huge populations to be useful. Production federated systems typically combine secure aggregation with central (often user-level) DP.
What is split computing and when is it useful?
Split computing runs the first part of a network on the device and sends the intermediate features (often compressed or quantized) to a server that runs the rest. It can reduce uplink bandwidth versus raw data, keep some privacy (features are less interpretable than raw images, though not private by guarantee), and use server compute for the heavy tail. It suits camera analytics or AR with a good network; it adds complexity, depends on connectivity and requires co-versioning the two halves.
What does peak activation memory mean on a microcontroller, and how do you reduce it?
It is the maximum total size of tensors that are alive at the same moment during inference, which must fit in the SRAM tensor arena alongside input and scratch buffers. It is usually set by early layers with high resolution and many channels. Reduce it with lower input resolution, fewer early channels, in-place operations and better memory planning, operator reordering, or patch-based inference that computes early layers tile by tile. Weights live in flash and do not count against SRAM.
Advanced
Estimate decode tokens/s for a 3B INT4 model on a phone.
Decode reads essentially all weights per token. 3B parameters at 4 bits is about 1.5 GB, plus scales, say ~1.7 GB. Phone LPDDR5X peak bandwidth is around 60-75 GB/s; effective sustained bandwidth for one workload might be 40-50 GB/s. Ceiling: 45 / 1.7, roughly 25 tokens/s. KV-cache reads add more bytes as context grows, and kernel efficiency and thermal limits lower it further, so 10-20 tokens/s is a realistic expectation. The same model in FP16 (~6 GB) would cap at about 7 tokens/s, which shows why quantization is essential.
Estimate TTFT for a 3B model with a 1,000-token prompt on NPU versus CPU.
Prefill costs about 2 x parameters x tokens = 2 x 3e9 x 1000 = 6e12 operations, plus attention (small at this length). At an effective 10 TOPS on the NPU that is about 0.6 s; at an effective 0.5 TFLOPS on the CPU about 12 s. Add tokenization and, if cold, model load time. This is why prefill is the phase where NPUs matter most, and why prefix caching of fixed system prompts and shorter prompts are the main software levers.
Why does TTFT grow with prompt length, and how do you reduce it?
Prefill must process every prompt token through every layer; compute grows linearly with prompt length for the dense parts and quadratically for attention. Reductions: run prefill on the NPU (compute-bound work suits it), cache the KV state for fixed system prompts (prefix caching), shorten prompts (summarize retrieved context, trim instructions), chunk prefill to keep the UI responsive, and use efficient attention kernels.
At what context length does the KV cache exceed the weights, and why does it matter?
Crossover = weight bytes / KV bytes per token. A ~1B GQA model (16 layers, 8 KV heads, head dim 64, FP16 KV) uses 32 KB per token, so its ~0.7 GB of INT4 weights are matched at about 22K tokens. An older 7B model with full multi-head attention uses 512 KB per token and crosses its ~3.5 GB of INT4 weights near 7K tokens. Beyond the crossover the cache dominates memory and decode bandwidth, so KV quantization, GQA and windowing matter more than further weight compression, and on an 8 GB phone the cache often decides the usable context before the low-memory killer intervenes.
How does the Hexagon-style NPU architecture map to neural network workloads?
The scalar unit handles control flow and orchestration. The vector unit (HVX) executes wide SIMD operations: element-wise ops, activations, some convolutions and data rearrangement. The tensor unit (HMX) performs dense matrix multiply-accumulate at low precision, where convolutions and linear layers spend most time. A tightly coupled on-chip memory keeps tiles of weights and activations close to the compute. Compilers tile the graph to maximize reuse in that memory and minimize DRAM traffic; ops that fit neither unit well become bottlenecks or fall back.
Why do mobile NPUs typically require static quantization, and what does it cost?
The NPU compiler bakes quantization parameters into the compiled graph: requantization multipliers and shifts, fused activation ranges, tiling and buffer sizes all depend on fixed scales. Computing scales at runtime would need extra reduction passes, floating-point logic and recompilation-like flexibility the hardware lacks. The cost is accuracy: scales must come from calibration, so inputs outside the calibrated range clip. That is why NPU LLM recipes often use 16-bit static activations (W4A16 or W8A16) instead of the 8-bit dynamic activations common on CPUs.
Explain how quantized integer matrix multiplication works with scales.
For y = W x with W approx s_w (q_w - z_w) and x approx s_x (q_x - z_x), the core is an integer dot product of (q_w - z_w) and (q_x - z_x) accumulated in INT32 to avoid overflow. The result is multiplied by the combined scale s_w x s_x, then requantized to the output scale and zero-point: q_y = round(acc x (s_w x s_x / s_y)) + z_y. Hardware implements the rescale as an integer multiplier plus shift. Bias is stored as INT32 with scale s_w x s_x. Symmetric weights (z_w = 0) remove cross terms and simplify the maths.
How does a W4A16 kernel work, and why is it fast even though maths is in FP16?
The kernel loads packed 4-bit weights (two per byte) plus the group's scale (and zero-point), unpacks and dequantizes them to FP16 in registers, then performs FP16 multiply-accumulates with FP16 activations. The arithmetic is no cheaper than FP16, but decode is memory-bound, so reading 4x fewer weight bytes than FP16 nearly quadruples the achievable speed. The dequantization adds a few instructions per weight, which is affordable because compute units are idle waiting for memory anyway; in compute-bound prefill, that overhead matters more, so some stacks use integer activation paths there.
Walk through the GPTQ algorithm in more detail.
For each linear layer, GPTQ collects calibration inputs X and forms H = 2XXT (plus damping on the diagonal). It aims to minimize ||WX - ŝX||2. Processing input columns in order, it quantizes column q, computes the error scaled by the inverse Hessian diagonal, e = (wq - quant(wq)) / [H-1]qq, and subtracts e x [H-1]q,: from the not-yet-quantized columns so the layer output is preserved. It uses a Cholesky decomposition of H-1 for stability, lazy block updates for speed, optional act-order (largest Hessian diagonal first) and group-wise scales. Layers are processed sequentially, feeding each the outputs of the already-quantized previous layers.
Explain AWQ's scaling trick and why it helps.
For a linear layer y = Wx, multiplying a weight input channel by s and dividing the matching activation channel by s leaves y unchanged. If a channel's activations are large, its weights are salient: their errors are amplified. Scaling those weights up by s > 1 before group-wise quantization makes them use more of the grid, so their relative rounding error shrinks roughly by s, while the group's scale grows only slightly if few channels are scaled. AWQ picks s = (mean activation magnitude)α per channel, grid-searching α to minimize output error on calibration data, and folds 1/s into the preceding normalization or linear layer. No backprop or mixed-precision kernels are needed.
Derive SmoothQuant's smoothing factor and give a numeric example.
Y = XW = (X diag(s)-1)(diag(s) W). Choose sj per input channel to balance the maximum magnitude of activations and weights in that channel: sj = max|Xj|α / max|Wj|1-α. With α = 0.5, a channel with activation max 100 and weight max 0.5 gets s = 10 / 0.707 ≈ 14.1, giving both new maxima ≈ 7.1. Larger α pushes more difficulty into weights (useful when activation outliers are extreme). The activation division is folded into the preceding LayerNorm, so runtime cost is zero and W8A8 kernels can be used.
How do rotation-based methods (QuaRot, SpinQuant) enable 4-bit activations?
For an orthogonal matrix R, (XR)(RTW) = XW. Multiplying activations by a rotation spreads the energy of a few outlier channels across all channels, making the distribution close to Gaussian with no dominant channel, so a single scale fits well even at 4 bits. QuaRot uses randomized Hadamard matrices (applied with a fast O(n log n) transform, and many rotations can be folded into adjacent weights); SpinQuant learns the rotations on calibration data for lower error. Combined, they allow W4A4 or W4A8 with a 4-bit KV cache at modest quality loss, and they were part of official quantized mobile releases of small open LLMs.
How would you quantize a KV cache, and why treat keys and values differently?
Start with INT8 per-token or per-head scales, which is close to lossless. Going to 4 bits or below, observe that keys have outlier channels (consistent across tokens, partly due to RoPE), so quantizing keys per channel (grouped along the channel dimension) works better, while values have no such structure and quantize well per token. Keep the most recent tokens in higher precision in a small buffer (they are both most attended and still being written), and quantize in groups as the buffer fills. Validate on long-context tasks and needle-in-a-haystack style retrieval, not only perplexity.
What are attention sinks and how do they relate to KV-cache eviction?
Models learn to dump excess attention weight on the first few tokens, which act as "sinks" regardless of their content. If a sliding window evicts them, generation quality collapses. Streaming approaches therefore keep a handful of initial tokens plus a recent window, bounding KV memory for arbitrarily long streams at the cost of forgetting the middle. Smarter eviction policies keep tokens that historically received high attention. On device, this enables long-running assistants within a fixed memory budget, trading some long-range recall.
Derive the expected speed-up of speculative decoding.
If each drafted token is accepted independently with probability α and k tokens are drafted, the number of tokens produced per target pass (including the target's own correction or bonus token) is 1 + α + α2 + ... + αk = (1 - αk+1)/(1 - α). Each cycle costs one target pass plus k draft steps of relative cost c, so speed-up ≈ [(1 - αk+1)/(1 - α)] / (1 + kc). For α = 0.8, k = 4, c = 0.1: 3.36 / 1.4 ≈ 2.4x. Too large a k wastes draft work on tokens likely to be rejected; the optimum depends on α and c. On device, also account for the draft model's memory and the verification pass reading the KV cache.
How would you debug a large accuracy drop after INT8 quantization?
- Verify the FP32 converted model matches the original (rules out conversion bugs, pre-processing mismatch, wrong normalization or channel order).
- Check calibration data is representative and big enough; try percentile/MSE-based range selection instead of min/max.
- Switch weights to per-channel.
- Run per-layer comparison (for example SQNR or cosine similarity between FP32 and quantized activations) to find the layer where error jumps.
- Keep sensitive layers in INT16/FP16 (mixed precision) or apply outlier techniques such as SmoothQuant or cross-layer equalization.
- If still insufficient, use QAT.
- Re-measure latency, since mixed precision can create new fallback partitions.
What is cross-layer equalization and bias correction?
Cross-layer equalization rescales the weights of consecutive layers (for example two convolutions separated by ReLU, which is scale-equivariant) so their per-channel ranges are more even, without changing the network's output in full precision. This makes per-tensor quantization much more accurate. Bias correction estimates the systematic shift in outputs introduced by quantization error and compensates by adjusting the bias. Both are data-free or low-data PTQ improvements.
How do you run an LLM that barely fits in memory?
Reduce weights (INT4 group-wise, or a smaller distilled model), reduce KV cache (GQA model, INT8 KV cache, shorter context, sliding window), memory-map weights so clean pages can be dropped instead of counted as dirty heap, share the embedding and LM head if the architecture ties them, free intermediate buffers through memory planning, and avoid duplicating weights between runtime and accelerator memory. Test with other apps in the background, because the low-memory killer may terminate your process under pressure.
What are the trade-offs of splitting prefill and decode across NPU and CPU/GPU?
Prefill is compute-bound, so the NPU gives large speed-ups. Decode is bandwidth-bound, and all units share the same DRAM, so the NPU's advantage is smaller, though it may still win on energy. Splitting phases across units means sharing or copying the KV cache and possibly keeping two weight layouts, which costs memory. Many stacks run both phases on the NPU with different compiled graphs (for example a batch-of-tokens graph for prefill and a single-token graph for decode) to avoid copies.
Why and how are large models split into multiple graphs for an NPU?
NPU toolchains impose limits on a single compiled graph: maximum size, addressable memory per context, on-chip buffer planning and compile time. A multi-GB LLM is therefore split into several sub-graphs (for example groups of transformer layers), each compiled into its own context binary, executed in sequence with shared input/output buffers and a shared KV cache. Splitting well means balancing sizes, minimizing tensors crossing boundaries, and loading or memory-mapping parts efficiently; poor splits add copies and synchronization between parts.
How would you design a hybrid on-device and cloud assistant?
Route by capability, privacy and cost: a small local model handles intent detection, short rewrites, summarization of on-device content and anything privacy-sensitive; complex reasoning, long context or fresh knowledge goes to the cloud with user consent. Use a router (rules or a small classifier with a confidence threshold). Keep a consistent response format so the UI does not care which path answered. Handle offline mode (local only, with graceful messaging), latency budgets (start local immediately, stream cloud if needed), and log routing decisions to tune thresholds. Enforce the same safety policies on both paths.
What is the effect of context length on on-device LLM performance?
Longer context increases prefill compute (attention grows quadratically), KV-cache memory (linear), and per-token decode bandwidth because the cache must be read each step. As a result TTFT rises, tokens/s declines over a long conversation, and memory pressure increases. Mitigations: cap context, summarize history, use sliding-window or GQA models, quantize the KV cache and reuse prefix caches.
How do you choose between INT8 and INT4 for an LLM on device?
INT4 halves memory and roughly doubles the decode speed ceiling compared with INT8, enabling larger models on the same device, but loses more quality, especially for small models and reasoning tasks. INT8 is closer to FP16 quality. Evaluate on task-specific benchmarks and perplexity, check what the NPU supports efficiently (many prefer W4A16 or W8A16), and consider mixed schemes: INT4 for large MLP weights, INT8 for attention projections or sensitive layers. Often a larger model at INT4 beats a smaller model at INT8 for the same memory.
How does memory bandwidth contention affect on-device AI?
CPU, GPU, NPU, display, camera ISP and modem share the same DRAM. A camera pipeline at high resolution plus rendering plus model inference can saturate bandwidth, raising latency for all. Symptoms include jank during inference and unstable model latency. Mitigations: reduce resolution or frame rate, keep data on-chip (fusion, tiling), avoid redundant copies between units, quantize to cut bytes, and schedule heavy work when the UI is idle.
Race to idle versus running at lower frequency: which saves more energy?
Dynamic power scales roughly as C V2 f, and lower frequency allows lower voltage, so energy per operation falls at lower clocks, favouring slow and steady. But static leakage and the fixed cost of keeping rails, memory and other blocks awake accrue for as long as work continues, favouring finishing fast and sleeping. The optimum is chip- and workload-specific: accelerators usually win both ways (fast and low energy per op); for CPUs the most efficient point is often a middle frequency rather than the maximum or minimum. The answer is to measure energy per inference across operating points, and to consider thermal effects on sustained performance.
How do you ship one feature across thousands of Android device models?
Define device tiers by chipset, RAM, accelerator and OS version. Build a model family (for example large INT8 for NPU flagships, medium FP16 for GPU devices, small INT8 for CPU-only devices) and choose at runtime from an allow-list plus runtime capability checks. Keep a reliable CPU fallback path. Download variants on demand. Collect telemetry (latency, fallback rate, crashes, thermal events) per device model and use remote configuration to change the variant or disable the feature. Run a device-lab or cloud-device benchmark suite in CI for each model release.
How would you add a custom operator for an NPU?
First try to express the op with supported primitives or rewrite the model. If a custom op is necessary, implement it in the vendor's custom-op framework (for example a vector-unit kernel), register it with the vendor SDK and the runtime's delegate/partitioner, define its quantization behaviour and shapes, and write reference tests against a CPU implementation. Benchmark carefully: a slow custom op that avoids two CPU round trips may still be a net win. Maintenance cost across SDK versions is the main downside.
How do you protect model IP on the device?
Models in an APK can be extracted. Options: encrypt the weights at rest and decrypt in native code (raises the bar but keys can still be found), use platform-backed key storage, load decrypted weights only into memory, compile to vendor binary formats that are harder to reverse, keep the most valuable parts server-side, and add integrity checks. Accept that a determined attacker with a rooted device can usually extract weights; decide based on the model's value.
What role does federated learning play on device?
Federated learning trains or fine-tunes a shared model across many devices without collecting raw data: each device computes an update on local data, and a server aggregates updates (often with secure aggregation and differential privacy). It suits keyboards, personalization and ranking. Challenges: device availability (train only when charging and on Wi-Fi), non-uniform data, communication cost, and privacy guarantees. On-device fine-tuning of small adapters is a lighter-weight alternative for personalization.
How does DP-SGD (or DP-FedAvg) provide differential privacy, and what does it cost?
Each example's gradient (or each client's update, for user-level DP) is clipped to a maximum L2 norm C, bounding any individual's influence; Gaussian noise with standard deviation proportional to C (the noise multiplier times C) is added to the sum before the model update. A privacy accountant composes the per-step guarantees across all rounds, accounting for sampling, into a final (ε, δ). Costs: clipping biases updates and noise slows convergence and lowers accuracy, compensated by larger cohorts, more rounds or smaller models; per-example clipping also adds compute and memory. Stating ε and the unit of privacy (example versus user) is essential.
Why does data movement dominate energy, and what does that imply for accelerator design and model choice?
At modern process nodes, reading a word from off-chip DRAM costs roughly two to three orders of magnitude more energy than an 8-bit MAC, and even SRAM reads cost more than the arithmetic. Implications: accelerators maximize on-chip reuse (systolic arrays, large SRAM, tiling), fuse operations to avoid round trips, and use narrow data types; models should favour architectures with high reuse, fewer bytes per inference (quantization, smaller activations), and avoid memory-bound patterns where possible. Energy per inference therefore tracks bytes moved at least as much as FLOPs.
Scenario & debugging
Your model runs at 15 ms on the NPU in the benchmark tool but 60 ms in the app. What do you check?
- Is the app using the same model file, delegate options and precision? Check delegate logs in the app.
- Is the interpreter recreated per frame (paying initialization each time)? Create once and reuse.
- Pre/post-processing time: image conversion and resizing in Kotlin can dominate. Measure each stage separately.
- Threading: inference on the UI thread, or contention with rendering and camera threads.
- Data copies between Java and native buffers; use direct buffers or zero-copy paths.
- Thermal state and clock frequencies differ between a cold benchmark and a warm app.
- Process priority: a background process gets lower CPU and scheduling priority.
Latency is fine for the first two minutes, then doubles. Diagnose.
This is the signature of thermal throttling. Confirm with a Perfetto trace showing CPU/GPU/NPU frequency drops and thermal status changes, and log thermal headroom. Fixes: lower energy per inference (more quantization, smaller model, full NPU delegation), lower duty cycle (skip frames, run at 15 fps instead of 30), reduce input resolution, use thermal headroom APIs to degrade gracefully before hard throttling, and test sustained runs routinely. Also check for memory growth causing garbage collection or swapping.
After switching to the NPU delegate, results are different from the CPU. What happened?
Possible causes: the NPU runs at lower precision (FP16 or INT8) where the CPU ran FP32; different rounding modes or accumulation order; saturation in INT8 due to poor calibration; a vendor kernel bug for a specific op or shape; or a layout issue (NHWC/NCHW mismatch) in custom pre-processing. Compare layer by layer against the CPU reference, test with the delegate limited to parts of the graph to bisect, and check vendor release notes. Decide acceptable tolerance based on task metrics, not bit-exactness.
The delegate log says only 60% of ops are delegated with 5 partitions. What do you do?
List the non-delegated ops and why (unsupported type, shape, attribute or op). Replace unsupported ops with supported equivalents (for example swap a custom activation for a supported one, rewrite reshape/transpose chains), make shapes static, quantize the whole graph consistently, move non-neural pre/post-processing out of the graph, and try a newer runtime or vendor SDK. Re-benchmark: sometimes delegating fewer, larger partitions is faster than more, smaller ones, and occasionally CPU-only is fastest for small models.
The app gets killed while the LLM feature is active on 8 GB devices. What do you investigate?
Measure peak memory (weights, KV cache, runtime buffers, activations) with dumpsys meminfo and heap profiles; check logs for low-memory killer events. Likely fixes: smaller or lower-bit model, shorter max context or INT8 KV cache, memory-map weights instead of reading into the heap, release the model when the feature is not in use, avoid duplicate copies of weights, and gate the feature by RAM tier. Test with realistic background app load.
Product wants an on-device summarizer that works on all phones sold in the last four years. How do you approach it?
Clarify requirements: input length, quality bar, latency budget, languages and offline needs. Profile the device population by RAM, chipset and accelerator. Propose tiers: a 1-3B INT4 model with NPU acceleration on capable devices, a smaller model or extractive summarizer on low-end devices, and optional cloud fallback with consent. Define metrics (quality scores on a test set, TTFT, tokens/s, battery per summary). Build a benchmark matrix across representative devices, download models on demand per tier, and roll out gradually with telemetry and a kill switch.
Your INT8 detector misses small objects that the FP32 model finds. Why and what do you do?
Small objects produce weak activations that can be rounded away, especially if calibration ranges are dominated by large, confident detections, or if the final box/score layers are quantized aggressively. Fixes: calibrate with images containing small objects, use percentile clipping rather than min/max, keep the detection head at higher precision, use per-channel weights, evaluate mAP per object size, and use QAT if needed. Also check that input resolution was not reduced as part of the optimization.
A keyword-spotting feature drains 5% battery per hour. How do you fix it?
An always-on feature must not keep the application processor awake. Move the first-stage detector to a low-power DSP or sensor hub with a tiny model, and wake the main processor only for a second-stage verification model when the first stage fires. Reduce the audio frame rate and feature computation cost, quantize to INT8, batch audio frames, and verify with power rail traces that the CPU enters deep idle between triggers. Tune the false-trigger rate, because each false wake costs energy.
Tokens/s is good initially but drops during a long chat session. Why?
Two effects: the KV cache grows with every token, so each decode step reads more bytes; and sustained load heats the device, triggering throttling. Also check for memory pressure causing paging. Mitigate with a context cap and conversation summarization, KV-cache quantization, sliding-window attention, efficient attention kernels, and thermal-aware pacing. Measure tokens/s against context length and time separately to tell the two effects apart.
A new vendor SDK version makes your model 30% slower. How do you handle it?
Reproduce with the benchmark tool on both versions on the same device and build to confirm. Compare per-layer profiles to locate the regressed ops, check release notes for changed defaults (precision, graph optimizations, memory mode), try toggling options, and report a minimal reproduction to the vendor. Meanwhile pin the previous SDK version for production. Add the model to a performance regression suite in CI so future upgrades are caught before release.
You must choose between a 1B model at INT8 and a 3B model at INT4 with similar memory. How do you decide?
Memory is similar (~1 GB versus ~1.5-1.7 GB), but the 3B INT4 model often gives better quality because capacity matters more than precision at these sizes, while the 1B INT8 model decodes faster (fewer bytes per token) and has shorter TTFT. Evaluate both on the actual task set, measure TTFT, tokens/s, energy per response and peak memory on target devices, and check NPU support for each format. Choose based on the product's quality bar and latency budget; for simple extraction tasks the smaller model may win.
The GPU delegate makes the camera preview stutter. What is going on?
The GPU is shared between rendering (UI, camera preview composition) and inference. Long-running compute work can delay frame rendering, causing jank. Options: move inference to the NPU, split the model into smaller GPU workloads, lower model resolution or frequency, use the delegate's options to reduce GPU priority where supported, or run inference on alternate frames. Profile with a system trace to see GPU queue contention and frame deadlines.
First inference after app start takes 3 seconds. How do you reduce it?
That is initialization: model loading, graph compilation for GPU shaders or NPU binaries, and memory allocation. Fixes: enable serialization/caching of compiled artifacts where the runtime supports it, precompile to the vendor's context binary format ahead of time, load the model asynchronously at a suitable moment (for example when the user opens the relevant screen), memory-map the model file, and consider a smaller model for the first interaction while the large one loads.
An interviewer asks: design an on-device photo search ("find photos of my dog at the beach").
Use a compact image-text embedding model (CLIP-style) quantized to INT8. Index offline: when the device is charging and idle, embed each photo on the NPU and store vectors in a local vector index (with incremental updates for new photos). At query time, embed the text on device and do a nearest-neighbour search, combining with metadata filters (date, location). Budget: embedding thousands of photos must not drain the battery, so batch work under charging constraints. Privacy: nothing leaves the device. Evaluate recall on a labelled set and latency per query; handle languages and model updates (re-indexing cost).
Your team says "the model is accurate enough" after quantization, based on average accuracy only. What else do you check?
Average accuracy can hide regressions in specific slices: low light, accents, skin tones, languages, small objects or rare classes. Check per-slice metrics, worst-case examples, calibration of confidence scores, and agreement rate with the FP32 model. For LLMs, check task-specific outputs, refusal and safety behaviour, and formatting. Also verify behaviour on real devices, since some accelerators produce slightly different numerics than the host simulator.
How would you set up CI for on-device models?
On each model or runtime change: convert and quantize reproducibly, run host-side accuracy tests against a golden evaluation set with thresholds, run on a device farm (or a cloud device service) across representative tiers to collect latency, memory, delegation coverage and, where possible, power; compare against the last release with regression thresholds; store artifacts with version metadata; and block the release on failures. Track results over time on a dashboard so slow drifts are visible.
After INT4 quantization, a small LLM starts repeating itself or producing gibberish. How do you debug it?
- Confirm the unquantized model works in the same runtime with the same tokenizer, chat template and sampling settings (template or special-token mistakes cause similar symptoms).
- Compare next-token distributions against FP16 on a few prompts (KL divergence, top-1 agreement) to measure damage.
- Check the recipe: group size too large, embeddings or LM head quantized too aggressively, or round-to-nearest where GPTQ/AWQ or an importance matrix is needed.
- Keep the LM head and embeddings at 6-8 bits, reduce group size, or switch to a higher-quality quant type; try Q8_0 to bracket the problem.
- Check the KV-cache precision and context handling (overflowing the context or a broken cache update causes degeneration at a fixed length).
- Tune sampling (repetition penalty, temperature) only after the numerics are right.
A tiny model (under 1 MB) runs slower on the NPU than on the CPU. Is something broken?
Probably not. For tiny models the fixed costs dominate: dispatching work to the NPU, synchronizing, copying and converting input/output buffers, and possibly waking the accelerator from a low-power state can take longer than the few microseconds of maths. The CPU with XNNPACK has near-zero dispatch overhead and data already in cache. Measure end to end and energy per inference; keep tiny models on CPU or DSP unless they run continuously and the NPU path is zero-copy, or batch several inferences per NPU call.
Design a smart camera box that analyses 8 video streams on a Jetson-class module.
Budget first: 8 streams x 15 fps = 120 frames/s; decode with the hardware video decoder, not the CPU. Use a batched detector (for example INT8 TensorRT engine at 640x640, batch 8) on the GPU and, if supported, a second model on the DLA; run a lightweight tracker so the detector can run every second frame while tracking fills gaps. Keep frames in GPU memory end to end (zero-copy pipeline, as in DeepStream) to avoid bandwidth waste. Check the power mode and sustained thermal behaviour in the enclosure at the maximum ambient temperature. Upload only events and metadata. Plan OTA updates with A/B partitions and a health check with rollback, and monitor per-stream FPS and dropped frames.
Your keyword-spotting model does not fit in the MCU's 128 KB of SRAM. What are your options?
Check what actually consumes SRAM: the interpreter reports the arena size needed. Weights should sit in flash, not SRAM. Reduce peak activations: lower the feature resolution (fewer frames or coefficients), shrink early-layer channels, use depthwise-separable blocks, reorder operators and enable in-place ops, or process the input in patches. Make sure everything is INT8 (a stray float tensor quadruples its size). Consider a streaming model that processes one frame at a time with a small state instead of a full 1 s window. Shrink the audio ring buffer, and, if still stuck, distil into a smaller student or move to an MCU with more SRAM or a micro-NPU.
Design heart-rhythm anomaly detection for a smartwatch with a one-day battery target.
Split the pipeline by power: the sensor hub samples PPG at low rate and runs signal-quality checks and a tiny INT8 model continuously (microwatts to a milliwatt); only suspicious windows wake the application processor for a larger model, and confirmed events prompt the user for an ECG reading or notify them. Budget energy: for example 1 mW average for sensing and inference over 24 h is 86 J, a small fraction of a ~1-2 Wh watch battery. Handle motion artefacts (use the accelerometer to gate), personalize thresholds on device, evaluate sensitivity and specificity per population slice, and consider regulatory requirements for medical claims. Keep raw health data on the device.
You are building in-car driver drowsiness detection. What is different from a phone feature?
It is safety-related, so determinism and reliability dominate: guaranteed worst-case latency (not p50), behaviour under all lighting (IR cameras at night), sunglasses and occlusions, and fail-safe handling when the model or camera fails. Hardware must work across automotive temperature ranges for many years, often with redundancy and a safety island, and software follows functional-safety processes with traceable validation datasets. Models are updated rarely and carefully via OTA with rollback. Privacy matters (in-cabin cameras), so processing stays in the vehicle. Evaluate false-alarm and miss rates on diverse drivers, since both annoy or endanger users.
How would you roll out a new model to 100,000 deployed IoT cameras safely?
Sign and version the model artifact; validate it against the target runtime version and hardware revision in a device lab first. Use staged rollout (internal devices, then 1%, 10%, 100%) with automatic health checks: model loads, latency and FPS within bounds, detection rate sanity, memory and temperature. Deliver as a delta when possible to save bandwidth, install to an inactive slot (A/B), and roll back automatically on failed health checks or on command. Keep the old model available, log the running version per device, and monitor aggregate metrics for drift after rollout.
A vendor claims their chip runs your model 2x faster than your measurements show. How do you reconcile?
Align the conditions: model version and input size, precision (their INT8 or INT4 versus your FP16), sparsity assumptions, batch size, whether pre/post-processing is included, SDK version and flags, which compute units were used, and warm versus sustained measurements (and at what temperature or power mode). Ask for their exact command and artifacts and reproduce on the same device. Often the gap is precision, batch size, excluded processing or peak clocks. Report both numbers with conditions, and base decisions on your sustained, end-to-end measurement.
The privacy team asks you to "prove" that federated keyboard training does not leak what users type. How do you respond?
Explain the layered protections and their guarantees: raw text never leaves the device; updates are clipped and combined with secure aggregation so the server sees only sums over many users; differential privacy with a stated user-level (ε, δ) bounds what any single user's data can change in the model; minimum cohort sizes prevent small-group inference. Add empirical checks: canary or "secret sharer" tests measuring whether planted rare sequences can be extracted from the trained model, and memorization audits. Be honest that DP gives a bounded, not zero, risk and that telemetry and logging need separate review.
An offline translation feature must add at most 150 MB per language pair. How do you get there?
Start with a compact encoder-decoder transformer trained or distilled for the language pair (sequence-level distillation from a large teacher works well), with a shared vocabulary and tied embeddings. Quantize weights to INT8 (or 4-bit for the largest matrices) with per-channel or per-group scales, check BLEU/COMET-style quality and human spot checks per domain. Use a deeper encoder and shallower decoder, since the decoder runs per token and dominates latency. Share one multilingual encoder across pairs if several languages are needed, download packs on demand, and keep the tokenizer small. Validate latency per sentence and energy on low-end devices.
Design on-device question answering over a user's personal notes and messages.
Use on-device retrieval-augmented generation: chunk and embed content with a small quantized embedding model in the background while charging; store vectors in a local index with incremental updates and deletion when content is removed. At query time, embed the question, retrieve top chunks with metadata filters, and prompt a small local LLM with a strict context budget (for example 2-4K tokens), asking it to cite sources and refuse when evidence is missing. Keep prefill small (short, relevant chunks), cache the system prompt, and quantize the KV cache. Everything stays local; escalation to the cloud, if any, requires explicit consent. Evaluate retrieval recall and answer faithfulness on a labelled set.
AR glasses must run hand tracking continuously, but the frame gets hot after five minutes. What do you do?
Head-worn devices have very tight skin-temperature limits and tiny batteries, so the budget is perhaps a few hundred milliwatts for the whole perception stack. Reduce work: lower camera resolution and frame rate when hands are absent, run a cheap hand-presence detector and the full landmark model only when needed, track between detections, crop to a region of interest, and fully quantize to run on the NPU or DSP. Offload heavy work to a paired phone or compute puck when connected. Use thermal headroom signals to scale fidelity, and measure sustained power at realistic ambient temperatures, not on a bench fan.
The quantized model matches the host simulator exactly but differs on the device. Why?
Simulators often emulate quantized maths with floating point and may not reproduce the device's exact rounding modes, accumulator widths, saturation behaviour, fused-kernel ordering, lookup-table approximations for activations (such as sigmoid or softmax) or FP16 intermediate precision. Driver or SDK versions may also differ. Dump per-layer outputs on the device and diff against the simulator to find the first diverging op, check the SDK release notes, and evaluate whether the difference matters on task metrics. Always validate final accuracy on the device, not only in simulation.