Edge AI

Edge AI Deployment & Projects

This is the hands-on companion to the concepts page: how to actually take a model from a checkpoint to a fast, stable, shippable feature on a phone-class device, with every command, code path and measurement you need. It is organised as a build-along course of five projects (benchmark harness, quantization error analysis, NPU deployment, KV-cache engineering and a full system) plus the profiling, debugging and shipping knowledge interviewers probe.

~165 min read 0 interview questions
In 30 seconds
  • Deployment is a loop: choose model, export a static graph, convert, quantize, compile for the target, integrate into the app, benchmark on real silicon, monitor in the field, and repeat.
  • Pick one primary path per job: ExecuTorch (.pte) or LiteRT (.tflite) for general apps, llama.cpp (GGUF) for fast LLM baselines, ONNX Runtime for cross-platform, vendor SDKs (QNN/QAIRT context binaries) for NPU efficiency, Core ML on Apple.
  • Measure like a product engineer: warm-up, many runs, p50/p90/p99, TTFT, prefill and decode tokens/s, peak RSS, load time, energy per request and a 10+ minute thermal soak. Never trust emulator numbers.
  • Quantization problems are found layer by layer: compare every intermediate tensor against an FP32 reference (cosine similarity, SQNR), then keep only the sensitive layers at higher precision.
  • NPUs want static shapes, static quantization and full graph coverage; one unsupported op in the middle can make an NPU slower than the CPU.
  • For LLMs, decode speed is set by memory bandwidth and long-context feasibility by the KV cache, which can outgrow the weights; manage it with GQA, quantized KV, sliding windows, paging and prefix caching.
  • Shipping means tiering by device, downloading models safely, versioning, A/B testing, remote kill switches and a fallback chain (NPU, GPU, CPU, cloud, or refuse).

The end-to-end deployment pipeline

Every edge AI deployment, whether it is a 5 MB keyword spotter or a 1-billion-parameter language model, walks the same eight stages. The names of the tools change; the stages do not. If you can explain each stage, its artifact, and what goes wrong there, you can reason about any stack an interviewer throws at you. The theory behind each stage (what an NPU is, how quantization maths works, how runtimes differ) lives on the concepts page, Edge AI; this page is about doing it.

Analogy

Think of deploying a model like adapting a restaurant recipe for an airline kitchen. You choose a dish that can survive the trip (choose the model), write it down as an exact, fixed recipe card with no improvisation (export a static graph), translate it into the airline's format (convert), swap ingredients for lighter ones that still taste right (quantize), pre-cook for the specific oven on that aircraft type (compile for the target), load it onto the plane (integrate into the app), test-serve it on a real flight at altitude, not in the test kitchen (benchmark on device), and read passenger feedback after every route (monitor).

Each step maps one-to-one: the recipe card is the exported graph, the airline format is the runtime file (.pte, .tflite, .gguf, context binary), the lighter ingredients are INT8/INT4 numbers, the aircraft oven is the specific NPU generation, and passenger feedback is field telemetry on latency, crashes and quality.

  +-------------+   +-----------+   +-----------+   +------------+
  | 1. CHOOSE / |--▶| 2. EXPORT |--▶| 3. CONVERT|--▶| 4. QUANTIZE|
  |    TRAIN    |   | static    |   | to runtime|   | PTQ / QAT  |
  +-------------+   | graph     |   | format    |   | mixed prec.|
                    +-----------+   +-----------+   +------------+
                                                          |
  +-------------+   +-----------+   +-----------+   +------------+
  | 8. MONITOR  |◀--| 7. BENCH- |◀--| 6. INTE-  |◀--| 5. COMPILE |
  | field data, |   |  MARK on  |   |  GRATE in |   | for target |
  | A/B, rollbk |   |  device   |   |  the app  |   | CPU/GPU/NPU|
  +-------------+   +-----------+   +-----------+   +------------+
         |                                               ▲
         +------------- iterate: fix ops, precision, model ----+

The eight stages and their artifacts

StageInputOutput artifactTypical toolsWhat usually breaks
1. Choose or trainTask, latency and memory budgetFP32/BF16 checkpointHugging Face, model zoos, vendor model hubs, your own trainingModel too large for the RAM tier; licence not suitable for shipping
2. ExportPyTorch/TF module + example inputsStatic graph (ExportedProgram, ONNX, SavedModel)torch.export, torch.onnx.export, TF SavedModelPython control flow, data-dependent shapes, unsupported custom ops
3. ConvertStatic graphRuntime format: .pte, .tflite, .onnx/.ort, .gguf, .mlpackage, QNN modelExecuTorch, ai-edge-torch, ONNX tools, llama.cpp converters, coremltools, QAIRT convertersOp not in the target op set; layout (NCHW vs NHWC) inserted transposes
4. QuantizeConverted or exported graph + calibration dataINT8/INT4/W4A16 model with scalesPT2E quantizers, torchao, AIMET, ORT quantization, LiteRT converter, llama-quantize, AI HubAccuracy drop from outliers, poor calibration data, sensitive layers
5. Compile for targetQuantized graphDelegated program, context binary, compiled model cacheBackend partitioners, QNN context binary generator, AI Hub compile jobs, GPU shader cachesGraph partitioning, CPU fallback, dynamic shapes rejected, SoC version mismatch
6. IntegrateModel + runtime libraryAPK/AAB or system component with native libsKotlin/Java, JNI/C++, Gradle, CMake, NDKUI-thread inference, asset compression breaking mmap, ABI mismatch, missing libs
7. BenchmarkApp or CLI on real hardwareLatency/memory/power/thermal reportCustom harness, benchmark_model, llama-bench, Perfetto, vendor profilersCold-run-only numbers, emulator numbers, uncontrolled thermal state
8. MonitorShipped featureField telemetry, rollback decisionsRemote config, analytics, crash reporting, staged rolloutsNo per-device breakdown; no kill switch; silent quality regressions

How to read this page

  1. Set up once Read the stack decisions, buy or borrow the right hardware, and complete the environment checklist before writing any model code.
  2. Learn the paths Walk through the export paths section and run at least two of them end to end (ExecuTorch and llama.cpp are the easiest starting pair).
  3. Integrate Put one model into a real Android app so you understand packaging, threading and memory.
  4. Build the five projects in order Each project reuses the previous project's harness and understanding; do not reorder them.
  5. Harden for production Profiling, accuracy debugging, shipping and the failure-mode checklist turn a demo into a product.
  6. Revise Use the quick revision list, glossary and interview questions at the end.
Interview angle "Walk me through deploying a model to a phone" is the most common opener. A strong answer names every stage above, the artifact at each stage, one concrete tool per stage, the two biggest risks (operator coverage on the accelerator and accuracy after quantization), and closes with how you would measure success on-device (sustained latency, peak memory, energy) and in the field (telemetry, staged rollout, rollback).
Common pitfall Treating the pipeline as linear. In practice stages 3 to 7 are a tight loop: an unsupported operator discovered in stage 5 sends you back to stage 2 to rewrite a module, and a latency miss in stage 7 sends you back to stage 1 to choose a smaller model. Plan time for several passes.

Stack decisions: what to learn and what to skip

The on-device stack churned heavily in 2024 and 2025 and has now largely consolidated. Choosing the right tools up front saves months of learning APIs that are being retired. The rule of thumb: learn one general-purpose runtime deeply, one fast LLM prototyping tool, and one vendor NPU toolchain. Everything else you should recognise and be able to discuss.

Analogy

Picking a stack is like choosing which languages to learn before moving abroad. You learn the national language fluently (your primary runtime), a handful of phrases for the neighbouring regions (secondary runtimes), and the local dialect of the city you will actually work in (the vendor NPU toolchain). You do not spend a year on a language that the government just stopped using (a deprecated API).

The national language is ExecuTorch or LiteRT, the neighbouring regions are ONNX Runtime and Core ML, the city dialect is QNN/QAIRT or NeuroPilot, and the retired language is NNAPI.

Do not invest in these

  • NNAPI (Android Neural Networks API). Introduced in Android 8.1 and marked deprecated / not recommended for new work from Android 15. Vendors implemented it inconsistently, so behaviour and performance varied wildly between devices. The replacement direction is vendor delegates and backends that plug directly into frameworks (LiteRT accelerators, ExecuTorch backends, ONNX Runtime execution providers). If a tutorial is built on NNAPI, treat it as historical.
  • Training from scratch, distributed training, cloud MLOps pipelines, data-centre serving. Useful elsewhere, but not what edge deployment roles screen for. Fine-tuning small adapters is the one training skill worth having (see Fine-tuning).
  • Chasing 7B-plus models on phones early. Quantizing them needs 80 GB-class GPUs, iteration is slow, and a 1B-class model teaches every lesson.

The stack that matters

ToolWhat it isWhen to use itPriority
ExecuTorch (PyTorch Edge)PyTorch's on-device runtime. Exports via torch.export to a .pte program. Small core runtime (tens of KB), many backends: XNNPACK (CPU), Vulkan (GPU), Qualcomm QNN, MediaTek, Arm Ethos-U, Core ML, MPS.Primary runtime for PyTorch models and LLMs on Android and iOS.Core
Qualcomm QAIRT / QNN SDKQualcomm's AI runtime SDK (the QNN SDK was folded into QAIRT). Converters, quantizers, HTP (Hexagon Tensor Processor) backend, context binaries, profiling, and Genie for generative models.Maximum efficiency on Snapdragon NPUs.Core (if targeting Snapdragon)
Qualcomm AI HubHosted service that compiles, quantizes and profiles models on real devices in the cloud; model zoo with ready export scripts.Fastest path to an NPU context binary and per-layer profile without owning every device.Core
llama.cpp / GGUFC/C++ LLM inference engine with hand-tuned CPU kernels (Arm NEON, i8mm, SVE), GPU backends (Vulkan, OpenCL, Metal) and the GGUF single-file format.First working LLM baseline in an hour; CPU reference numbers; format for distributing quantized models.Core
LiteRT (formerly TensorFlow Lite)Google's runtime for .tflite FlatBuffers with XNNPACK CPU, GPU delegate and NPU accelerators; PyTorch models arrive via ai-edge-torch.Vision, audio and classic models in Android apps; Google ecosystem (MediaPipe).Useful
MediaPipe LLM Inference / LiteRT-LMHigh-level on-device LLM API on top of LiteRT, loading bundled .task or .litertlm models.Shipping a supported LLM (Gemma family and others) with minimal code.Useful
ONNX Runtime MobileCross-platform runtime; execution providers for CPU, NNAPI (legacy), QNN, Core ML, XNNPACK.When the model pipeline is ONNX-based or you need one runtime across Windows, Android and iOS.Useful
Core ML / coremltoolsApple's runtime and converter; dispatches across CPU, GPU and Neural Engine.Any Apple target.Know it
AIMETQualcomm's quantization and compression toolkit (AdaRound, cross-layer equalisation, QuantAnalyzer, QAT).Recovering accuracy for NPU INT8/INT4 deployments.Useful
MediaTek NeuroPilot / Neuron runtimeMediaTek's APU toolchain (also reachable via ExecuTorch and LiteRT).Second vendor, after you are comfortable with one.Later
System AI services (for example AICore with Gemini Nano)OS-level service hosting a shared foundation model; apps call it through an SDK and cannot load their own weights.Study the architecture: it is the reference design for a platform AI service.Study

Why a closed system AI service is worth studying

A platform AI service such as Android's AICore shows what a production-grade on-device design looks like. It is updated independently of the OS (as an updatable system module), it only enables itself on devices with enough RAM and an NPU, it ships model updates as binary deltas instead of re-downloading gigabytes, it runs safety filtering around the model, it isolates the model from callers behind a permission boundary, and it supports small LoRA adapters (tens of MB) that specialise one shared base model per feature. Every one of those decisions is an answer to a real constraint: storage, bandwidth, memory, privacy, safety and multi-tenancy. It is the blueprint for the shared-inference-service option in Project 5.

Choosing a runtime for a given job

SituationGood first choiceWhy
PyTorch LLM on Android and iOSExecuTorch (XNNPACK, then QNN/Core ML backends)Same export flow, many backends, strong LLM tooling
Validate an LLM use case this afternoonllama.cpp with a GGUFBuilds in minutes, runs anything on CPU
Vision/audio model in an Android appLiteRT with GPU/NPU acceleratorMature Android tooling, MediaPipe tasks, easy benchmarks
Supported LLM shipped fastMediaPipe LLM Inference / LiteRT-LMA few lines of Kotlin; GPU path included
Best perf/W on SnapdragonQNN context binary via AI Hub or QAIRT (directly, through ExecuTorch QNN backend, or ORT QNN EP)Direct HTP access, precompiled graphs
Same model on Windows on Arm, Android and iOSONNX Runtime with per-platform EPsOne API and model format
Apple devicesCore ML (or ExecuTorch Core ML backend)Only supported path to the Neural Engine
Interview angle Expect "Which runtime would you pick and why?" Give a decision, not a survey: state the constraints (source framework, target SoCs, model type, need for NPU, team skills, app size), pick one, and name the fallback. Mentioning that NNAPI is deprecated and why (inconsistent vendor drivers, lowest-common-denominator op set) signals current knowledge.
Common pitfall API names in this space move fast (TensorFlow Lite became LiteRT, the QNN SDK became part of QAIRT, MediaPipe LLM Inference is converging with LiteRT-LM, torch.ao quantization moved into torchao). Re-verify SDK versions and flag names against current documentation before quoting them, and say so in interviews: it reads as experience, not uncertainty.

Hardware, software and budget

Performance work on edge AI is only meaningful on physical silicon. Emulators and simulators do not model memory bandwidth, cache sizes, DVFS, thermal throttling, or the NPU at all; a number measured on an emulator is not wrong by a small factor, it is meaningless.

Analogy

Benchmarking on an emulator is like testing running shoes on a treadmill in an air-conditioned showroom and then promising they will perform in a mountain marathon in summer. The terrain (memory system), the heat (thermal limits) and the fatigue (sustained throttling) are exactly what you did not test.

The showroom treadmill is your laptop or emulator; the mountain marathon is a mid-range phone in a user's pocket at 35 degrees, running your model for ten minutes.

Devices: match the NPU generation

Vendor NPU libraries are versioned per hardware generation. For Snapdragon, the Hexagon library folder must match the SoC; using the wrong one fails to load or silently falls back to CPU.

SoC (flagship tier)Hexagon architecture / library folderNotes
Snapdragon 8 Gen 1v69 (confirm the hexagon-vXX folder in your QAIRT SDK)Older; usable for INT8 CNNs, limited for LLMs
Snapdragon 8 Gen 2v73 (hexagon-v73)Good minimum for LLM NPU work; INT4 weight support
Snapdragon 8 Gen 3v75 (hexagon-v75)Common reference device generation
Snapdragon 8 Elitev79 (hexagon-v79)Current-generation class; strongest LLM NPU numbers

For other vendors the same principle applies: MediaTek Dimensity APU generations map to specific NeuroPilot versions; Google Tensor chips expose their TPU through LiteRT accelerators; Apple's Neural Engine is reached only through Core ML.

Minimum viable kit

  • One Android phone with a recent flagship SoC (for example Snapdragon 8 Gen 2 or newer) and 12 GB+ RAM. A second-hand device is perfectly fine.
  • A Linux workstation (Ubuntu 22.04+ native, or WSL2 on Windows) with 16-32 GB RAM and 100 GB+ free disk. Most vendor toolchains are Linux-first.
  • Android SDK platform-tools (adb) and NDK r26 or newer.
  • A reliable USB-C data cable. Flaky cables cause intermittent adb disconnects that look like software bugs.

Ideal kit

  • Two phones from different SoC generations or vendors: cross-silicon comparisons are the most informative measurements you can produce.
  • A Pixel device to explore the platform AI service path and Tensor TPU accelerators.
  • A MediaTek device for second-vendor work later.
  • A USB power meter or, better, a bench power supply / power monitor for repeatable energy numbers; otherwise use on-device power rails and the fuel gauge.
  • Optional: a single-board computer or microcontroller dev kit (Arm Cortex-M with an Ethos-U NPU, or a Linux SBC with an NPU) for the wearable/sensor project.

GPU access for quantization

Running a model is cheap; quantizing a large one is not. Algorithms such as GPTQ, SpinQuant, AdaRound or QAT need the full-precision model in GPU memory plus calibration activations. Rough requirements for LLM quantization workflows:

Model classTypical GPU memory needed to quantizePractical advice
~0.5-1B16-24 GB (often fits a consumer GPU)Do almost all experiments here
~3B~40 GB (32 GB is borderline and may OOM)Rent by the hour only when needed
~7-8B~80 GBOnly once, at the end, to prove you can scale

Strategy: use 1B-class models with post-training quantization for most learning; rent a cloud GPU by the hour for the one or two runs that genuinely need it (a QAT comparison, a vendor LLM export). Shut instances down immediately after; idle GPUs are the main cost leak.

Budget (approximate, varies by region)

ItemApproximate costNotes
Used flagship Android phoneUSD 200-450The one thing worth spending on
Second device (optional)USD 150-400Enables cross-silicon comparison
Cloud GPU hours for quantizationUSD 150-300 over six monthsHourly rental, only for heavy runs
USB power meter (optional)USD 20-60Rough external power numbers
Vendor model hub accountsFree tierRegistration usually required
ExecuTorch, llama.cpp, LiteRT, ONNX RuntimeFreeOpen source
Realistic totalUSD 400-800Spread over roughly six months
Tip If you work somewhere that has device labs, cloud device farms, or vendor-hosted real-device profiling (such as AI Hub profile jobs), use them for breadth across SoCs, but keep one device on your desk for deep, interactive thermal and power work.
Common pitfall Buying a device with too little RAM. A 1B INT4 LLM needs roughly 1-2 GB resident, a 3B model 2-3.5 GB, plus the app, the OS, and other apps. On an 8 GB phone the low-memory killer becomes part of your experiment; on 12 GB+ you can study it deliberately rather than fight it.
Interview angle "Why not benchmark on the emulator or your laptop?" The answer: different memory bandwidth and cache hierarchy, no DVFS or thermal model, no NPU, different kernels (x86 vs Arm), and no realistic contention with the rest of the OS. Host runs are for numerical reference only, never for performance claims.

Environment setup

Toolchain setup is where most people lose several weekends and give up. Do it once, carefully, in a reproducible way, and timebox it. If the heavy toolchain fights you for more than a day, build llama.cpp first to get a model running on the phone, then come back.

Analogy

Setting up the environment is like a chef's mise en place: every knife sharpened, every ingredient measured into its bowl, before the stove is lit. Cooking goes fast and calmly once everything is in place; cooking while hunting for ingredients burns the dish.

The knives are your NDK, CMake and Python environment; the measured bowls are the downloaded model weights and SDKs; lighting the stove is your first export and on-device run.

Checklist

  1. Operating system Ubuntu 22.04+ (native or WSL2) with 100 GB+ free disk. Model artifacts, SDKs and build trees add up quickly.
  2. Python Python 3.10-3.12 in a dedicated virtual environment or conda environment per toolchain. Never install into system Python; vendor SDKs often pin conflicting versions.
  3. Android tools SDK platform-tools with adb devices showing your phone as device (not unauthorized).
  4. NDK Android NDK r26+ installed; ANDROID_NDK exported.
  5. Build tools CMake 3.24+, Ninja, a recent Clang, git with LFS.
  6. Model access A model hub account with accepted licences for gated models (Llama-family weights require accepting the licence).
  7. Vendor SDK A vendor developer account, AI Hub token, and the QAIRT SDK downloaded and unpacked; QNN_SDK_ROOT exported.
  8. Phone settings Developer options on, USB debugging on, "stay awake while charging" on, battery optimisation disabled for your test app, adaptive brightness off.
  9. Version control One git repository for the whole effort. Commit measurement CSVs and configs, not just code, so every number is reproducible.

Host environment

# Base tools (Ubuntu)
sudo apt update && sudo apt install -y build-essential cmake ninja-build git git-lfs \
     python3-venv python3-dev clang unzip wget

# Android NDK (adjust version/path to what you installed)
export ANDROID_HOME=$HOME/Android/Sdk
export ANDROID_NDK=$ANDROID_HOME/ndk/26.3.11579264
export PATH=$ANDROID_HOME/platform-tools:$PATH

# Sanity checks
adb devices                     # should list your phone as "device"
adb shell getprop ro.soc.model  # e.g. SM8650 for Snapdragon 8 Gen 3
adb shell getprop ro.product.cpu.abi   # must be arm64-v8a
adb shell cat /proc/meminfo | head -3  # total RAM

ExecuTorch host build

git clone --recursive https://github.com/pytorch/executorch.git
cd executorch
python -m venv .venv && source .venv/bin/activate
./install_executorch.sh          # installs torch, torchao and the executorch pip package

# Verify before touching Android
python -c "import executorch, torch; print(executorch.__file__, torch.__version__)"

ExecuTorch Android runner build (the flags that matter)

cmake -DCMAKE_TOOLCHAIN_FILE=$ANDROID_NDK/build/cmake/android.toolchain.cmake \
      -DANDROID_ABI=arm64-v8a \
      -DANDROID_PLATFORM=android-26 \
      -DCMAKE_BUILD_TYPE=Release \
      -DEXECUTORCH_BUILD_XNNPACK=ON \
      -DEXECUTORCH_BUILD_KERNELS_QUANTIZED=ON \
      -DEXECUTORCH_BUILD_KERNELS_OPTIMIZED=ON \
      -DEXECUTORCH_BUILD_EXTENSION_MODULE=ON \
      -DEXECUTORCH_BUILD_EXTENSION_DATA_LOADER=ON \
      -DEXECUTORCH_BUILD_EXTENSION_TENSOR=ON \
      -DEXECUTORCH_BUILD_EXTENSION_LLM=ON \
      -DEXECUTORCH_XNNPACK_ENABLE_KLEIDI=ON \
      -Bcmake-out-android .
cmake --build cmake-out-android -j$(nproc) --target install

# Then build the LLM example runner against that install
# (see examples/models/llama in the repo for the exact current target)
Tip EXECUTORCH_XNNPACK_ENABLE_KLEIDI enables Arm's KleidiAI low-bit matrix-multiply micro-kernels inside XNNPACK. They give a large prefill improvement (on the order of 20% or more) on Arm CPUs at identical accuracy. Recent versions enable it by default on Arm, but if your prefill numbers look about 20% low against reference tables, check this flag first. Always build Release: a debug build can be several times slower.

llama.cpp for Android (fast baseline)

git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build-android \
  -DCMAKE_TOOLCHAIN_FILE=$ANDROID_NDK/build/cmake/android.toolchain.cmake \
  -DANDROID_ABI=arm64-v8a -DANDROID_PLATFORM=android-28 \
  -DCMAKE_BUILD_TYPE=Release -DBUILD_SHARED_LIBS=OFF -DGGML_OPENMP=OFF
cmake --build build-android -j$(nproc) --target llama-cli llama-bench

adb shell mkdir -p /data/local/tmp/lcpp
adb push build-android/bin/llama-cli build-android/bin/llama-bench /data/local/tmp/lcpp/

Vendor SDK environment (Qualcomm example)

# After unpacking QAIRT
export QNN_SDK_ROOT=/opt/qcom/aistack/qairt/<version>
source $QNN_SDK_ROOT/bin/envsetup.sh        # sets PATH, PYTHONPATH, LD_LIBRARY_PATH
python $QNN_SDK_ROOT/bin/check-python-dependency   # installs pinned Python deps

# AI Hub client (separate venv recommended)
python -m venv ~/venvs/aihub && source ~/venvs/aihub/bin/activate
pip install qai-hub qai-hub-models
qai-hub configure --api_token <YOUR_TOKEN>
qai-hub list-devices | head

Phone preparation for stable measurements

# Keep the screen on while plugged in, and set a fixed brightness
adb shell svc power stayon usb
adb shell settings put system screen_brightness_mode 0
adb shell settings put system screen_brightness 100

# Record the software build you measured on
adb shell getprop ro.build.fingerprint

# Create a scratch area on the device
adb shell mkdir -p /data/local/tmp/bench
Common pitfall Mixing Python environments. The ExecuTorch, QAIRT and AI Hub stacks each pin specific versions of torch, onnx and numpy. Installing them into one environment produces import errors that look like model bugs. One environment per toolchain, and record pip freeze next to every result.
Interview angle Interviewers for SDK roles sometimes ask how you would make a model build reproducible. Mention pinned toolchain versions, containerised or locked environments, recording the device build fingerprint, SoC and SDK version alongside results, deterministic export scripts, and checking artifacts (hashes) into a registry.

Export and conversion paths

"Export" captures a model as a static, framework-independent graph; "conversion" turns that graph into the format a specific runtime executes. Most modern paths start from PyTorch. Below are the six paths you will meet most often, with working command and code patterns. Treat exact flag names as version-dependent and check the current documentation.

Analogy

Exporting is like recording a live jazz performance to sheet music. The live performance (eager PyTorch) can improvise: it branches, loops and changes tempo depending on the audience. Sheet music (the exported graph) must fix every note in advance so any orchestra can play it. Conversion is then transcribing that sheet music for a particular band: a string quartet (CPU), a brass band (GPU) or a single virtuoso synthesiser (NPU) that only plays certain notes.

Improvisation is Python control flow and dynamic shapes; the notes the synthesiser cannot play are unsupported operators that fall back to the CPU.

                         PyTorch nn.Module (FP32/BF16)
                                    |
      +-------------+-------------+-+-----------+-------------+--------------+
      |             |             |             |             |              |
 torch.onnx    torch.export  ai-edge-torch  HF safetensors  coremltools   AI Hub /
  .export           |        (torch.export  convert_hf_to_   ct.convert    QAIRT
      |             |          inside)       gguf.py            |        converters
   .onnx        ExecuTorch        |             |            .mlpackage      |
      |          to_edge +     .tflite       .gguf (f16)        |       QNN model
 ORT quantize    partitioner   (LiteRT)    llama-quantize    Core ML     + context
  / QNN EP         |              |             |            runtime     binary .bin
      |          .pte         LiteRT +      Q4_K_M .gguf       (ANE)          |
 ONNX Runtime   ExecuTorch    delegates     llama.cpp                   QNN HTP /
   Mobile        runtime                                                  Genie
PathArtifactBest forMain gotcha
PyTorch to ONNX.onnx (optionally .ort)Cross-platform, ORT execution providers, many vendor converters accept ONNXOpset versions, dynamic axes, exporter differences (TorchScript vs dynamo)
PyTorch to ExecuTorch.ptePyTorch-native on-device, LLMs, multi-backendGraph breaks in torch.export; backend partition coverage
PyTorch to LiteRT.tfliteAndroid apps, MediaPipe, Google acceleratorsLayout transposes (NCHW to NHWC), op coverage in the converter
Hugging Face to GGUF.ggufLLM prototyping and distribution, CPU inferenceArchitecture must be supported by llama.cpp; tokenizer metadata
To QNN context binary.bin context + libsSnapdragon HTP at full efficiencyStatic shapes, static quantization, SoC-specific binaries
To Core ML.mlpackageApple Neural EngineANE op/shape constraints silently push work to GPU/CPU

Path 1: PyTorch to ONNX (and ONNX Runtime quantization)

import torch, torchvision

model = torchvision.models.mobilenet_v3_small(weights="DEFAULT").eval()
dummy = torch.randn(1, 3, 224, 224)

torch.onnx.export(
    model, (dummy,), "mobilenet_v3.onnx",
    input_names=["image"], output_names=["logits"],
    opset_version=17,
    dynamic_axes=None,          # keep static for NPUs; use {"image": {0: "batch"}} only if needed
    do_constant_folding=True,
)

# Validate numerically against PyTorch before doing anything else
import onnxruntime as ort, numpy as np
sess = ort.InferenceSession("mobilenet_v3.onnx", providers=["CPUExecutionProvider"])
ref = model(dummy).detach().numpy()
out = sess.run(None, {"image": dummy.numpy()})[0]
print("max abs diff:", np.abs(ref - out).max())   # expect ~1e-5 for FP32

Static (QDQ) INT8 quantization with a calibration reader:

from onnxruntime.quantization import (quantize_static, CalibrationDataReader,
                                      QuantFormat, QuantType, CalibrationMethod)
from onnxruntime.quantization.shape_inference import quant_pre_process

quant_pre_process("mobilenet_v3.onnx", "mobilenet_v3.pre.onnx")   # shape inference + folding

class Reader(CalibrationDataReader):
    def __init__(self, samples):          # samples: list of float32 arrays, shape (1,3,224,224)
        self.it = iter([{"image": s} for s in samples])
    def get_next(self):
        return next(self.it, None)

quantize_static(
    "mobilenet_v3.pre.onnx", "mobilenet_v3.int8.onnx",
    calibration_data_reader=Reader(calibration_samples[:300]),
    quant_format=QuantFormat.QDQ,             # QuantizeLinear/DequantizeLinear pairs; EPs fuse them
    activation_type=QuantType.QUInt8, weight_type=QuantType.QInt8,
    per_channel=True, calibrate_method=CalibrationMethod.MinMax,
)
# For the QNN EP, ORT provides helpers (qnn_preprocess_model, get_qnn_qdq_config)
# that produce a QDQ model with the uint8/uint16 activation types the HTP expects.

Optionally convert to the ORT format and build a reduced-operator runtime to shrink the mobile binary: python -m onnxruntime.tools.convert_onnx_models_to_ort model.int8.onnx produces a .ort file and a list of required operators you can feed into a custom minimal build.

Path 2: PyTorch to ExecuTorch (.pte)

Generic model with the XNNPACK CPU backend and PT2E INT8 quantization:

import torch
from torch.export import export
from executorch.exir import to_edge_transform_and_lower
from executorch.backends.xnnpack.partition.xnnpack_partitioner import XnnpackPartitioner
from executorch.backends.xnnpack.quantizer.xnnpack_quantizer import (
    XNNPACKQuantizer, get_symmetric_quantization_config)
from torchao.quantization.pt2e.quantize_pt2e import prepare_pt2e, convert_pt2e

model = MyModel().eval()
example = (torch.randn(1, 3, 224, 224),)

# 1. Capture a graph suitable for quantization
graph = export(model, example).module()

# 2. Insert observers, calibrate, convert (post-training static quantization)
quantizer = XNNPACKQuantizer().set_global(get_symmetric_quantization_config(is_per_channel=True))
prepared = prepare_pt2e(graph, quantizer)
for x in calibration_batches:          # ~100-500 representative inputs
    prepared(x)
quantized = convert_pt2e(prepared)

# 3. Re-export, lower supported subgraphs to XNNPACK, serialize
program = to_edge_transform_and_lower(
    export(quantized, example), partitioner=[XnnpackPartitioner()]
).to_executorch()
with open("model_xnnpack_int8.pte", "wb") as f:
    f.write(program.buffer)

LLM export using the built-in LLM exporter (running example: Llama-3.2-1B-Instruct; you need consolidated.00.pth, params.json and tokenizer.model from the model repository):

python -m extension.llm.export.export_llm \
  base.model_class="llama3_2" \
  base.checkpoint="${LLAMA_DIR}/consolidated.00.pth" \
  base.params="${LLAMA_DIR}/params.json" \
  model.use_kv_cache=True \
  model.use_sdpa_with_kv_cache=True \
  model.dtype_override="fp32" \
  base.metadata='"{\"get_bos_id\":128000, \"get_eos_ids\":[128009, 128001]}"' \
  quantization.qmode="torchao:8da4w" \
  quantization.group_size=128 \
  quantization.embedding_quantize="torchao:4,32" \
  backend.xnnpack.enabled=True \
  export.max_seq_length=2048 \
  export.output_name="llama3_2_1b_8da4w.pte"
  • 8da4w means 8-bit dynamic activations and 4-bit weights; group_size=128 gives one scale per 128 weights.
  • embedding_quantize="4,32" quantizes the large embedding table (128256 x 2048 for this model, about a fifth of all parameters) to 4-bit with groups of 32.
  • use_kv_cache and use_sdpa_with_kv_cache are essential for speed. Without them every decode step recomputes attention over the whole sequence and numbers are badly wrong.
  • The BOS/EOS ids in metadata must match the tokenizer, or generation will not stop (or will stop immediately).
  • Older ExecuTorch versions used python -m examples.models.llama.export_llama with flags like -kv --use_sdpa_with_kv_cache -X -qmode 8da4w; recognise both.

Swap XnnpackPartitioner for a vendor partitioner (Qualcomm QNN, MediaTek, Core ML, Vulkan) to target accelerators; the rest of the flow stays the same. That uniformity is ExecuTorch's main selling point.

Path 3: PyTorch to LiteRT via ai-edge-torch

import torch, torchvision
import ai_edge_torch        # package name may be litert-torch in newer releases

model = torchvision.models.resnet18(weights="DEFAULT").eval()
sample = (torch.randn(1, 3, 224, 224),)

edge_model = ai_edge_torch.convert(model, sample)      # uses torch.export under the hood
edge_model.export("resnet18_fp32.tflite")

# Check parity on host
import numpy as np
ref = model(*sample).detach().numpy()
out = edge_model(*sample)
print("max abs diff:", np.abs(ref - out).max())

# Optional: NHWC inputs to avoid transposes on GPU/NPU delegates
nhwc_model = ai_edge_torch.to_channel_last_io(model, args=[0])
edge_nhwc = ai_edge_torch.convert(nhwc_model, (torch.randn(1, 224, 224, 3),))

Quantization uses the PT2E flow with a PT2E quantizer for LiteRT (ai_edge_torch.quantize) before conversion, or post-conversion quantization for weight-only compression. For LLMs, ai-edge-torch's generative API includes re-authored model definitions and conversion scripts that emit prefill and decode signatures with an external KV cache, then bundle them into a .task or .litertlm file for MediaPipe / LiteRT-LM.

From TensorFlow/Keras, the classic converter still applies: tf.lite.TFLiteConverter.from_saved_model() with optimizations=[tf.lite.Optimize.DEFAULT], a representative_dataset generator for full-integer quantization, and TFLITE_BUILTINS_INT8 ops for integer-only NPU targets.

Path 4: Hugging Face checkpoint to GGUF (llama.cpp)

cd llama.cpp
pip install -r requirements.txt

# 1. Convert HF safetensors to a 16-bit GGUF (weights + tokenizer + metadata in one file)
python convert_hf_to_gguf.py /models/Llama-3.2-1B-Instruct \
       --outfile llama-3.2-1b-f16.gguf --outtype f16

# 2. Optional: importance matrix from calibration text improves low-bit quality
./build/bin/llama-imatrix -m llama-3.2-1b-f16.gguf -f calib.txt -o imatrix.dat

# 3. Quantize
./build/bin/llama-quantize --imatrix imatrix.dat llama-3.2-1b-f16.gguf llama-3.2-1b-Q4_K_M.gguf Q4_K_M
./build/bin/llama-quantize llama-3.2-1b-f16.gguf llama-3.2-1b-Q4_0.gguf Q4_0   # Arm repack-friendly

# 4. Quality check on host: perplexity on held-out text
./build/bin/llama-perplexity -m llama-3.2-1b-Q4_K_M.gguf -f wiki.test.raw

# 5. On device
adb push llama-3.2-1b-Q4_0.gguf /data/local/tmp/lcpp/
adb shell "cd /data/local/tmp/lcpp && ./llama-bench -m llama-3.2-1b-Q4_0.gguf -p 128,512 -n 128 -t 4"
adb shell "cd /data/local/tmp/lcpp && ./llama-cli -m llama-3.2-1b-Q4_0.gguf -p 'Explain KV caching in one paragraph.' -n 128 -t 4 -no-cnv"
  • Q4_0 is a simple 4-bit format with a scale per 32 weights; on recent Arm CPUs llama.cpp repacks it at load time into layouts that use i8mm/dotprod instructions, which is often the fastest CPU option on phones.
  • Q4_K_M uses super-blocks with mixed 4/6-bit precision for sensitive tensors; better quality per byte, a popular default.
  • Thread count matters: on big.LITTLE phones, use the number of performance cores (often 4-6); using all cores can be slower because little cores become stragglers.

Path 5: To a QNN context binary (Snapdragon NPU)

Option A, via AI Hub (hosted compile, quantize and profile on real devices):

import torch, qai_hub as hub

model = MyModel().eval()
example = torch.randn(1, 3, 224, 224)
traced = torch.jit.trace(model, example)

device = hub.Device("Samsung Galaxy S24 (Family)")

# Compile straight to a QNN context binary for this SoC
compile_job = hub.submit_compile_job(
    model=traced, device=device,
    input_specs={"image": (1, 3, 224, 224)},
    options="--target_runtime qnn_context_binary",
)
target_model = compile_job.get_target_model()

# Profile on a real device in the cloud: per-layer timing, compute unit per op, memory
profile_job = hub.submit_profile_job(model=target_model, device=device)

# Run inference on-device to check numerics against the PyTorch reference
infer_job = hub.submit_inference_job(model=target_model, device=device,
                                     inputs={"image": [example.numpy()]})
out = infer_job.download_output_data()
target_model.download("model_ctx.bin")

For quantized models, submit a quantize job (INT8 or W4A16/W8A16 style activations) with calibration data before compiling, or export a pre-quantized ONNX (QDQ) model. For supported LLMs, the model zoo has ready export scripts that quantize and split the model into several context binaries:

python -m qai_hub_models.models.llama_v3_2_3b_instruct.export \
    --device "Snapdragon 8 Elite QRD" \
    --skip-inferencing --skip-profiling \
    --output-dir ./genie_bundle_src

Option B, locally with the QAIRT tools. Binary names and flags are release-specific and are the most common source of copy-paste failures. Older QNN SDK releases ship qnn-onnx-converter, qnn-model-lib-generator and qnn-context-binary-generator. Newer QAIRT releases document qairt-converter / qairt-quantizer (exact names: check the SDK bin/ of the version you installed). The sketch below is the older converter shape so you can recognise it; copy flags from that release's HTML docs or --help, do not invent them.

# 1. Convert (and quantize with a calibration input list of raw tensor files)
# Flag names (--act_bitwidth, --weights_bitwidth, --input_list) are from older
# qnn-onnx-converter help text. Confirm on your SDK before running.
qnn-onnx-converter --input_network model.onnx \
    --input_list calib_list.txt \
    --act_bitwidth 16 --weights_bitwidth 8 \
    --output_path model_qnn.cpp

# 2. Build a model library for the target
qnn-model-lib-generator -c model_qnn.cpp -b model_qnn.bin \
    -t x86_64-linux-clang -o model_libs

# 3. Generate an HTP context binary (graph finalized for a specific SoC arch)
qnn-context-binary-generator \
    --backend $QNN_SDK_ROOT/lib/x86_64-linux-clang/libQnnHtp.so \
    --model model_libs/x86_64-linux-clang/libmodel_qnn.so \
    --config_file htp_config.json \
    --binary_file model_ctx

# 4. Run on device from the context binary (no on-device graph compile)
adb push model_ctx.bin $QNN_SDK_ROOT/lib/aarch64-android/libQnnHtp*.so \
         $QNN_SDK_ROOT/lib/hexagon-v75/unsigned/libQnnHtpV75Skel.so \
         $QNN_SDK_ROOT/bin/aarch64-android/qnn-net-run /data/local/tmp/qnn/
adb shell "cd /data/local/tmp/qnn && export LD_LIBRARY_PATH=. ADSP_LIBRARY_PATH=. && \
  ./qnn-net-run --backend libQnnHtp.so --retrieve_context model_ctx.bin \
                --input_list inputs.txt --profiling_level basic"

Option C, via frameworks: the ExecuTorch Qualcomm backend (a QnnPartitioner plus QnnQuantizer, producing a .pte with embedded context binaries) or the ONNX Runtime QNN execution provider with EP context caching (covered in the Android integration section).

Path 6: To Core ML (Apple)

import torch, coremltools as ct

model = MyModel().eval()
example = torch.randn(1, 3, 224, 224)
traced = torch.jit.trace(model, example)          # or torch.export program in newer coremltools

mlmodel = ct.convert(
    traced,
    inputs=[ct.TensorType(name="image", shape=example.shape)],
    convert_to="mlprogram",
    compute_units=ct.ComputeUnit.ALL,              # CPU + GPU + Neural Engine
    minimum_deployment_target=ct.target.iOS17,
    compute_precision=ct.precision.FLOAT16,
)

# Weight compression: 8-bit linear or 4-bit palettization / per-block quantization
from coremltools.optimize.coreml import (OpLinearQuantizerConfig, OptimizationConfig,
                                         linear_quantize_weights)
cfg = OptimizationConfig(global_config=OpLinearQuantizerConfig(mode="linear_symmetric"))
mlmodel = linear_quantize_weights(mlmodel, config=cfg)
mlmodel.save("MyModel.mlpackage")

Use Xcode's Core ML performance report to see which operations actually ran on the Neural Engine; ops that the ANE cannot run (certain shapes, dynamic dimensions, some activation types) fall back to GPU or CPU without any error.

Making a model exportable

  1. Remove Python-level dynamism Replace data-dependent if branches with torch.cond or tensor operations; make loops fixed-length.
  2. Fix shapes Pick concrete input sizes (and for LLMs, fixed prefill chunk sizes such as 128 plus a decode length of 1). Use torch.export.Dim only where the target runtime truly supports dynamic dimensions.
  3. Replace unsupported ops Rewrite custom or exotic operators with standard ones (for example GELU-tanh approximation, explicit RMSNorm, attention written as matmul + softmax).
  4. Keep the KV cache explicit For LLMs, make the cache an input/output or registered buffer with a fixed maximum length rather than a growing Python list.
  5. Validate parity After every conversion, run the same inputs through the original and converted model; FP32 differences should be around 1e-5 to 1e-4.
Common pitfall Exporting with dynamic axes "for flexibility" and then targeting an NPU. Most NPU compilers require fully static shapes; a dynamic batch or sequence dimension either fails compilation or forces the whole graph to CPU. Decide shapes early and export separate graphs (for example a prefill graph and a decode graph) instead.
Interview angle "How would you get this PyTorch model onto a Snapdragon NPU?" A complete answer: make it exportable (static shapes, standard ops), export (ONNX or torch.export), quantize with representative calibration data to the precision the HTP wants (INT8 or W4/W8 weights with 16-bit activations), compile to a context binary for the exact SoC, check per-op placement in a profile to confirm nothing falls back to CPU, validate numerics against FP32, then integrate via QNN directly, the ExecuTorch QNN backend, or ORT's QNN EP.

Integrating into an Android app

A model file is not a feature. Integration means choosing where inference runs (Kotlin via a runtime's Java API, or native C++ via JNI), how the model gets onto the device and into memory, which accelerator executes it, and how the app behaves when things go wrong. For Android framework background (services, lifecycles, threading) see Android frameworks.

Analogy

Integrating a model into an app is like installing a commercial espresso machine in a small cafe. You must get it through the door (packaging and download), connect it to the right power and water supply (delegates and native libraries), put it where it does not block the queue at the counter (off the UI thread), keep it warm between customers instead of cold-starting it every time (reuse the interpreter), and have a backup kettle if it breaks (CPU or cloud fallback).

The door is APK size limits, the power supply is the NPU/GPU delegate, the counter queue is the main thread, keeping it warm is session reuse and compiled-graph caching, and the backup kettle is the fallback chain.

Architecture of an on-device inference feature

+---------------------------- App process ------------------------------+
|  UI (Compose/View)  --▶  ViewModel  --▶  InferenceRepository          |
|                                   |  (single background dispatcher)   |
|                                   ▼                                    |
|              Kotlin runtime API  or  JNI bridge ──▶ C++ engine        |
|              (LiteRT / ORT /        (ExecuTorch, llama.cpp, QNN,      |
|               ExecuTorch / MediaPipe)  custom pre/post-processing)    |
|                                   |                                    |
|        pre-processing ──▶ model ──▶ post-processing ──▶ result         |
+-----------------------------------|------------------------------------+
                                    ▼
      Delegate / backend:  NPU (HTP, APU, TPU)  |  GPU (OpenCL/Vulkan)  |  CPU (XNNPACK/KleidiAI)
                                    ▼
      Model file: mmap from app storage (downloaded) or uncompressed asset

Option 1: LiteRT Interpreter with GPU delegate and CPU fallback (Kotlin)

// build.gradle.kts
// Artifact IDs moved from org.tensorflow:tensorflow-lite* to com.google.ai.edge.litert:*.
// Confirm the current Maven coordinates and the Java package they export in the
// official LiteRT Android docs for the version you pin.
dependencies {
    implementation("com.google.ai.edge.litert:litert:<version>")
    implementation("com.google.ai.edge.litert:litert-gpu:<version>")
}
android {
    androidResources { noCompress += listOf("tflite") }   // required to mmap from assets
}
// Interpreter is the long-lived, widely documented path. Some LiteRT artifacts
// still export org.tensorflow.lite.* for compatibility; newer packages may not.
import org.tensorflow.lite.Interpreter
import org.tensorflow.lite.gpu.CompatibilityList
import org.tensorflow.lite.gpu.GpuDelegate
import java.io.FileInputStream
import java.nio.MappedByteBuffer
import java.nio.channels.FileChannel

class Classifier(context: Context) : AutoCloseable {
    private var gpu: GpuDelegate? = null
    private val interpreter: Interpreter

    init {
        val options = Interpreter.Options()
        val compat = CompatibilityList()
        if (compat.isDelegateSupportedOnThisDevice) {
            gpu = GpuDelegate(compat.bestOptionsForThisDevice)
            options.addDelegate(gpu)
        } else {
            options.setNumThreads(4)          // XNNPACK CPU path
        }
        interpreter = Interpreter(mapAsset(context, "model_int8.tflite"), options)
    }

    private fun mapAsset(ctx: Context, name: String): MappedByteBuffer {
        ctx.assets.openFd(name).use { fd ->
            FileInputStream(fd.fileDescriptor).channel.use { ch ->
                return ch.map(FileChannel.MapMode.READ_ONLY, fd.startOffset, fd.declaredLength)
            }
        }
    }

    // Reuse buffers: allocate once, fill per call
    private val input = ByteBuffer.allocateDirect(1 * 224 * 224 * 3).order(ByteOrder.nativeOrder())
    private val output = Array(1) { ByteArray(1000) }

    fun classify(pixels: ByteArray): ByteArray {
        input.rewind(); input.put(pixels)
        interpreter.run(input, output)       // call from a background thread only
        return output[0]
    }

    override fun close() { interpreter.close(); gpu?.close() }
}

LiteRT Next also documents a CompiledModel-style API that binds a model to an accelerator (CPU, GPU, or an NPU compiler plugin) and can consume SoC-specific compiled artifacts. Class names, packages, accelerator enums and compile options change between releases. Check the current official LiteRT docs before writing production code; do not assume Interpreter.Options maps one-to-one onto CompiledModel, and do not copy guessed method names from older blog posts. The concepts stay the same: create once, choose an accelerator, reuse buffers, fall back gracefully.

Option 2: ONNX Runtime Mobile with the QNN execution provider

// build.gradle.kts: use the onnxruntime-android-qnn package for the QNN EP
dependencies { implementation("com.microsoft.onnxruntime:onnxruntime-android-qnn:<version>") }
import ai.onnxruntime.*

val env = OrtEnvironment.getEnvironment()
val opts = OrtSession.SessionOptions().apply {
    // Java helper name (addQnn vs addConfigEntry) and option keys are
    // version-specific. Confirm against the current ONNX Runtime QNN EP docs.
    // Common documented keys include backend_path and an HTP performance mode.
    addQnn(mapOf(
        "backend_path" to "libQnnHtp.so",
        "htp_performance_mode" to "burst"
        // Other documented modes include sustained_high_performance; do not
        // invent extra flags. Graph-finalization and context-cache keys also
        // move between ORT releases — look them up for the AAR you ship.
    ))
    addConfigEntry("ep.context_enable", "1")
    addConfigEntry("ep.context_file_path", File(context.filesDir, "model_ctx.onnx").path)
    addConfigEntry("session.disable_cpu_ep_fallback", "1")
}
val session = env.createSession(modelFile.path, opts)

val input = OnnxTensor.createTensor(env, floatBuffer, longArrayOf(1, 3, 224, 224))
session.run(mapOf("image" to input)).use { result ->
    val logits = (result[0].value as Array<FloatArray>)[0]
}
  • The QNN HTP backend requires a quantized (QDQ) model for NPU execution; FP32 models either use the HTP's FP16 path on newer chips or fall back.
  • Turn off CPU fallback in development so partitioning problems surface as errors; turn it back on in production for robustness.
  • EP context caching avoids re-finalizing the graph on every launch, which can take seconds for large models.

Option 3: MediaPipe LLM Inference API (fastest LLM integration)

// dependencies { implementation("com.google.mediapipe:tasks-genai:<version>") }
// Engine options vs session options (temperature / top-k / LoRA) were split
// across releases. Check the current MediaPipe Tasks GenAI docs and the
// LiteRT-LM notes; method names below are the long-lived engine-level ones.
import com.google.mediapipe.tasks.genai.llminference.LlmInference

val options = LlmInference.LlmInferenceOptions.builder()
    .setModelPath(File(context.filesDir, "models/gemma.task").path)
    .setMaxTokens(1024)                 // prompt + output budget: sizes the KV cache
    .build()
val llm = LlmInference.createFromOptions(context, options)   // expensive: do once, off main thread

val answer = llm.generateResponse("Summarise this note in two lines: ...")

llm.generateResponseAsync(prompt) { partial, done ->
    uiFlow.tryEmit(partial)
    if (done) onFinished()
}

This path only loads models converted into its bundle format (with tokenizer and metadata). It is ideal when a supported model family fits the product; for arbitrary architectures, use ExecuTorch or llama.cpp.

Option 4: ExecuTorch from Kotlin

// Packages and constructors move between ExecuTorch Android releases
// (Module vs ExecutorchModule, LlmModule constructor args). Confirm against
// the current ExecuTorch Android / LLM runner docs for the AAR you ship.
import org.pytorch.executorch.EValue
import org.pytorch.executorch.Module
import org.pytorch.executorch.Tensor

val module = Module.load(File(filesDir, "model_xnnpack_int8.pte").path)   // mmap-based loader
val input = Tensor.fromBlob(floatArray, longArrayOf(1, 3, 224, 224))
val out = module.forward(EValue.from(input))[0].toTensor().dataAsFloatArray

import org.pytorch.executorch.extension.llm.LlmCallback
import org.pytorch.executorch.extension.llm.LlmModule

val llm = LlmModule(ptePath, tokenizerPath, /* temperature = */ 0.7f)
llm.load()
llm.generate(prompt, /* seqLen = */ 512, object : LlmCallback {
    override fun onResult(token: String) { appendToUi(token) }
    override fun onStats(stats: String) { Log.i("LLM", stats) }
})

Option 5: your own C++ engine through JNI

Native integration gives you full control (custom pre/post-processing, direct QNN or Genie calls, llama.cpp, zero-copy buffers). Keep the JNI surface small and coarse-grained: one call per request, not one per tensor.

// Kotlin side
object NativeEngine {
    init { System.loadLibrary("edgeengine") }
    external fun create(modelPath: String, backend: Int): Long      // returns a handle
    external fun generate(handle: Long, prompt: String, maxTokens: Int, cb: TokenCallback): String
    external fun destroy(handle: Long)
}
fun interface TokenCallback { fun onToken(piece: String) }
// C++ side: engine_jni.cpp
#include <jni.h>
#include <memory>
#include <string>
#include "engine.h"          // your wrapper around ExecuTorch / llama.cpp / QNN

extern "C" JNIEXPORT jlong JNICALL
Java_com_example_edge_NativeEngine_create(JNIEnv* env, jobject, jstring jpath, jint backend) {
    const char* path = env->GetStringUTFChars(jpath, nullptr);
    auto* engine = new Engine(path, static_cast<Backend>(backend));   // loads + mmaps weights
    env->ReleaseStringUTFChars(jpath, path);
    return reinterpret_cast<jlong>(engine);
}

extern "C" JNIEXPORT jstring JNICALL
Java_com_example_edge_NativeEngine_generate(JNIEnv* env, jobject, jlong h, jstring jprompt,
                                            jint maxTokens, jobject cb) {
    auto* engine = reinterpret_cast<Engine*>(h);
    const char* p = env->GetStringUTFChars(jprompt, nullptr);
    std::string prompt(p);
    env->ReleaseStringUTFChars(jprompt, p);

    jclass cls = env->GetObjectClass(cb);
    jmethodID onToken = env->GetMethodID(cls, "onToken", "(Ljava/lang/String;)V");

    std::string full;
    engine->generate(prompt, maxTokens, [&](const std::string& piece) {
        full += piece;
        jstring js = env->NewStringUTF(piece.c_str());
        env->CallVoidMethod(cb, onToken, js);
        env->DeleteLocalRef(js);            // avoid local-ref table overflow in long loops
    });
    return env->NewStringUTF(full.c_str());
}

extern "C" JNIEXPORT void JNICALL
Java_com_example_edge_NativeEngine_destroy(JNIEnv*, jobject, jlong h) {
    delete reinterpret_cast<Engine*>(h);
}
# CMakeLists.txt (app/src/main/cpp)
cmake_minimum_required(VERSION 3.22)
project(edgeengine CXX)
set(CMAKE_CXX_STANDARD 17)
add_library(edgeengine SHARED engine_jni.cpp engine.cpp)
target_compile_options(edgeengine PRIVATE -O3 -march=armv8.2-a+dotprod+fp16)
# Link prebuilt runtime libs (ExecuTorch, llama, or QNN) imported as IMPORTED targets
target_link_libraries(edgeengine executorch_llm_runner log android)
  • Token-by-token callbacks into Java cost microseconds each; fine for tens of tokens per second, but batch pieces if you stream very fast.
  • Call JNI from a thread that was attached to the JVM; callbacks from a native worker thread need AttachCurrentThread.
  • Vendor NPU libraries must be packaged in jniLibs/arm64-v8a (or extracted to app storage) and the DSP library search path (ADSP_LIBRARY_PATH) must point to where the Hexagon skeleton libraries live; set it before initializing the backend.
  • Set useLegacyPackaging or extractNativeLibs according to whether your runtime can load libraries directly from the APK.

Threading and lifecycle rules

  1. Never on the main thread Use a single dedicated background dispatcher or thread for inference; most interpreters are not thread-safe for concurrent calls.
  2. Create once, reuse Loading, delegate initialization and graph compilation can take hundreds of milliseconds to seconds. Keep the session in an application-scoped object.
  3. Warm up Run one dummy inference after load, while the user is not waiting, to trigger shader compilation and cache fills.
  4. Release on memory pressure Respond to onTrimMemory by freeing caches and, for large models, unloading entirely when backgrounded.
  5. Cancel LLM generation must be cancellable (a flag checked every token) so leaving the screen stops work and power draw.

Model packaging and delivery

MethodSize limit and behaviourGood forWatch out for
Bundled asset in the APK/AABCounts towards install size; update needs app updateSmall models (under ~50-100 MB)Must be stored uncompressed to mmap
Play Asset Delivery (install-time, fast-follow, on-demand packs)Large packs served by the storeMedium-large models tied to app versionsPack availability timing; handle "not yet downloaded"
Store-provided on-device AI packs (device-targeted delivery)Different model variants per device classShipping NPU-specific binaries per SoCNewer mechanism; check availability
Own download server / CDNAny size; full control of versioningFrequent model updates, A/B testsResume, integrity checks, storage checks, metered networks
System-provided model (OS AI service)Nothing to shipSupported devices and model familiesNo control of model or version; availability varies

A robust self-managed download flow:

// Pseudocode for a WorkManager download worker
1. Fetch manifest: {modelId, version, url, sha256, sizeBytes, minRamMb, socAllowList, runtimeVersion}
2. Check eligibility: RAM tier, SoC, free storage (size x 2 for safety), unmetered network, charging
3. Download to  files/models/.tmp/<id>-<version>.part  with HTTP Range resume
4. Verify SHA-256 (and signature if the model is sensitive)
5. Atomically rename to files/models/<id>/<version>/model.pte
6. Update "active version" pointer only after a smoke-test inference succeeds
7. Keep the previous version until the new one has run successfully N times; then delete

Memory-mapped weights

Memory mapping (mmap) makes the model file appear as memory without copying it into the heap. Pages are read from storage on first touch and can be dropped by the kernel under pressure because they are clean and file-backed.

mmap (recommended default)

  • Near-instant "load"; pages fault in lazily
  • Clean file pages are reclaimable, so lower kill risk
  • Shared between processes that map the same file
  • Counts in RSS only when touched
  • Requires uncompressed, page-aligned storage

Read into heap

  • Full load cost up front; double memory during load if copying
  • Anonymous memory: not reclaimable, counts fully against you
  • Needed when the runtime must repack or transform weights
  • Predictable latency once loaded
  • Works with compressed or encrypted files

Two subtleties: first, if the runtime repacks weights into a kernel-friendly layout at load time (common for CPU kernels and GPU delegates), the repacked copy is anonymous memory and the mmap benefit shrinks; second, if pages are evicted during generation, decode stalls on storage reads, which appears as sudden latency spikes. For latency-critical paths, touch all pages after load (prefault) or use mlock where permitted, and budget memory honestly.

Common pitfall Forgetting noCompress for model assets. A compressed asset cannot be memory-mapped, so the runtime silently copies the whole model into the Java or native heap, doubling peak memory at load and sometimes causing an out-of-memory crash only on low-RAM devices.
Interview angle "How do you integrate a 1 GB model into an Android app?" Strong answers cover: do not bundle it (download with resume, checksum and atomic install), mmap from app storage, create the session once on a background thread with warm-up, pick accelerator with a fallback chain, stream results, respond to onTrimMemory, gate the feature by device tier, and keep a remote kill switch.

Project 1: baseline and benchmark harness

Duration: 2-3 weeks. Goal: an LLM running on your phone plus a measurement harness that every later project reuses. Skills: latency, memory and power profiling; token-rate measurement; PyTorch export tooling; accelerator basics. Output: a repository containing the harness and a sustained-load thermal analysis write-up.

Analogy

A benchmark harness is like the timing system at an athletics track. A stopwatch in someone's hand gives one noisy number; a proper system has the same start line every time (fixed environment), a warm-up lap that is not timed (warm-up runs), dozens of heats (repeated runs), photo-finish percentiles rather than one lucky time (p50/p90/p99), and a long-distance event, not just the sprint (the thermal soak).

The start line is device state (charge, temperature, brightness), the warm-up lap is discarded initial runs, heats are repetitions, and the long-distance event is ten or more minutes of continuous generation.

What you build

Run Llama-3.2-1B-Instruct (or an even smaller model such as a 0.5-0.6B one for a quicker first win) on Android through ExecuTorch with the XNNPACK backend, and in parallel through llama.cpp as a cross-check. Then build a harness that records, on every run:

MetricDefinitionWhy it matters
Model load time (cold and warm)From process start / file open to ready-for-first-inference; cold = after dropping caches or rebootDominates perceived start-up; almost nobody reports it
TTFT (time to first token)Prompt submitted to first output token visiblePerceived responsiveness; grows with prompt length
Prefill throughputPrompt tokens / prefill timeCompute-bound phase; accelerators help most here
Decode throughputGenerated tokens / decode time (excluding first token)Memory-bandwidth-bound; the streaming speed users see
Per-token latency percentilesp50/p90/p99 of inter-token timeStutters are noticed even if the average is fine
Peak RSS / PSSMax resident memory of the processDecides whether the low-memory killer (lmkd) terminates you or other apps
Model file size.pte/.gguf bytes on diskDownload and storage cost
Energy per requestJoules per 100 generated tokens or per inferenceBattery impact; the fair way to compare CPU, GPU and NPU
Sustained / peak ratioThroughput after 10 min divided by throughput in the first minuteWhat a shipped product actually experiences

Step 1: export, push and run

# Export (see the ExecuTorch path above), then:
adb shell mkdir -p /data/local/tmp/llama
adb push llama3_2_1b_8da4w.pte  /data/local/tmp/llama/
adb push tokenizer.model        /data/local/tmp/llama/
adb push cmake-out-android/examples/models/llama/llama_main /data/local/tmp/llama/

adb shell "cd /data/local/tmp/llama && ./llama_main \
    --model_path=llama3_2_1b_8da4w.pte \
    --tokenizer_path=tokenizer.model \
    --prompt='Explain KV caching in one paragraph.' \
    --seq_len=256 --cpu_threads=4 --warmup=1"
# The runner prints prompt/generated token counts and timing stats
# (model load, first token latency, prefill and generation rates).

For an instruct model, wrap the prompt in the model's chat template (for Llama 3.x: <|begin_of_text|><|start_header_id|>user<|end_header_id|> ... <|eot_id|><|start_header_id|>assistant<|end_header_id|>). A missing template is the most common cause of rambling or empty output.

Step 2: define a benchmark protocol

  1. Fix the environment Same device, OS build, airplane mode, fixed screen brightness (or screen off via a CLI run), battery 50-90% or external supply, device at room temperature with the case removed, no other apps, same prompt set.
  2. Cool down between configurations Wait until the SoC thermal zone returns near its idle temperature (for example within 2 degrees of baseline) before the next configuration.
  3. Warm-up Discard 1-3 full generations (or 10-50 runs for small models) so caches, page faults and kernel selection settle. Record load time and first-run latency separately.
  4. Repeat At least 10 generations for LLMs, 100-1000 inferences for small models; report median and p90/p99, plus standard deviation.
  5. Vary prompt length Measure prefill at several lengths (for example 64, 256, 1024 tokens) because attention cost and cache behaviour change with length.
  6. Record everything Device, SoC, build fingerprint, runtime version, model hash, quantization config, thread count, ambient temperature, start/end temperatures.

Step 3: the harness (host-side driver)

#!/usr/bin/env python3
"""bench.py - drive an on-device LLM runner over adb and collect metrics into CSV."""
import csv, json, re, statistics, subprocess, time, datetime

DEV_DIR = "/data/local/tmp/llama"
RUNNER = f"cd {DEV_DIR} && ./llama_main --model_path={{pte}} --tokenizer_path=tokenizer.model " \
         "--prompt=\"{prompt}\" --seq_len={seq} --cpu_threads={threads}"

def adb(cmd: str) -> str:
    return subprocess.run(["adb", "shell", cmd], capture_output=True, text=True, timeout=900).stdout

def soc_temp_c() -> float:
    # pick the zone(s) that represent CPU/SoC on your device; names vary by vendor
    out = adb("for z in /sys/class/thermal/thermal_zone*; do echo $(cat $z/type) $(cat $z/temp); done")
    temps = [int(t) / 1000 for name, t in (l.split() for l in out.splitlines() if l.strip())
             if re.search(r"cpu|soc|tsens", name, re.I) and t.lstrip("-").isdigit()]
    return max(temps) if temps else float("nan")

def wait_cool(target_c: float, timeout_s=600):
    t0 = time.time()
    while soc_temp_c() > target_c and time.time() - t0 < timeout_s:
        time.sleep(10)

def parse_stats(text: str) -> dict:
    # Adapt these regexes to your runner's output format
    pats = {
        "load_ms":    r"Model load time:\s*([\d.]+)",
        "ttft_ms":    r"(?:Time to first generated token|first token).*?([\d.]+)",
        "prefill_tps": r"Prompt evaluation:.*?([\d.]+)\s*tokens/s",
        "decode_tps":  r"Generated \d+ tokens:.*?([\d.]+)\s*tokens/s",
    }
    return {k: float(m.group(1)) if (m := re.search(p, text, re.S)) else None for k, p in pats.items()}

def peak_rss_kb(proc_name="llama_main") -> int:
    out = adb(f"pid=$(pidof {proc_name}); [ -n \"$pid\" ] && grep VmHWM /proc/$pid/status")
    m = re.search(r"(\d+)", out)
    return int(m.group(1)) if m else -1

def run_config(pte, prompt, seq=256, threads=4, warmup=2, reps=10, idle_c=None):
    if idle_c: wait_cool(idle_c + 2)
    cmd = RUNNER.format(pte=pte, prompt=prompt, seq=seq, threads=threads)
    for _ in range(warmup):
        adb(cmd)
    rows = []
    for i in range(reps):
        t_start = soc_temp_c()
        text = adb(cmd)
        s = parse_stats(text)
        s.update(rep=i, temp_start=t_start, temp_end=soc_temp_c(), pte=pte, threads=threads)
        rows.append(s)
    return rows

def summarize(rows, key):
    xs = sorted(r[key] for r in rows if r.get(key) is not None)
    if not xs: return {}
    pct = lambda p: xs[min(len(xs) - 1, int(round(p / 100 * (len(xs) - 1))))]
    return {"p50": pct(50), "p90": pct(90), "p99": pct(99),
            "mean": statistics.mean(xs), "stdev": statistics.pstdev(xs)}

if __name__ == "__main__":
    meta = {"fingerprint": adb("getprop ro.build.fingerprint").strip(),
            "soc": adb("getprop ro.soc.model").strip(),
            "time": datetime.datetime.now().isoformat()}
    idle = soc_temp_c()
    rows = run_config("llama3_2_1b_8da4w.pte", "Explain KV caching in one paragraph.", idle_c=idle)
    with open("results.csv", "w", newline="") as f:
        w = csv.DictWriter(f, fieldnames=rows[0].keys()); w.writeheader(); w.writerows(rows)
    print(json.dumps({"meta": meta, **{k: summarize(rows, k)
                      for k in ("ttft_ms", "prefill_tps", "decode_tps")}}, indent=2))

Note: VmHWM in /proc/<pid>/status is the peak resident set ("high water mark"); poll it while the process runs or print it from the runner itself at exit, because it disappears with the process. For app-based runs, dumpsys meminfo <package> gives PSS broken down into native heap, graphics, and file mappings.

Step 4: in-app timing (C++ side)

#include <chrono>
#include <vector>
#include <algorithm>

using Clock = std::chrono::steady_clock;            // monotonic; never use wall clock

struct GenStats { double load_ms, ttft_ms, prefill_tps, decode_tps; std::vector<double> itl_ms; };

GenStats timed_generate(Engine& e, const std::vector<int>& prompt, int max_new) {
    GenStats s{};
    auto t0 = Clock::now();
    e.prefill(prompt);                                // fills KV cache, returns logits
    int tok = e.sample();
    auto t1 = Clock::now();
    s.ttft_ms = std::chrono::duration<double, std::milli>(t1 - t0).count();
    s.prefill_tps = prompt.size() / (s.ttft_ms / 1000.0);

    auto prev = t1;
    for (int i = 1; i < max_new && tok != e.eos(); ++i) {
        e.decode(tok);
        tok = e.sample();
        auto now = Clock::now();
        s.itl_ms.push_back(std::chrono::duration<double, std::milli>(now - prev).count());
        prev = now;
    }
    double decode_s = std::chrono::duration<double>(prev - t1).count();
    s.decode_tps = s.itl_ms.size() / decode_s;
    return s;
}

double percentile(std::vector<double> v, double p) {
    std::sort(v.begin(), v.end());
    size_t idx = static_cast<size_t>(p / 100.0 * (v.size() - 1) + 0.5);
    return v[idx];
}
TTFT ≈ ttokenize + tprefill(Nprompt) + tfirst decode + sample prefill tok/s = Nprompt / tprefill; decode tok/s = (Ngenerated − 1) / tdecode. Always state the prompt length and generated length with the number.

Step 5: power and energy

Energy is what users feel as battery drain. Three levels of rigour:

MethodHowAccuracy
Fuel gauge samplingSample current_now and voltage_now every 100-500 ms during the run; integrate V x I over time; subtract an idle baseline measured the same wayCoarse (fuel gauges average and update slowly), but fine for relative comparisons over long runs
batterystatsdumpsys batterystats --reset, run the workload unplugged, then dump and inspect estimated drain per UIDModel-based estimates; good for app-level attribution
On-device power rails (ODPM) via PerfettoRecord the android.power data source; many recent devices expose per-rail energy counters (CPU big/mid/little, GPU, memory, NPU/DSP rails vary)Best on-device option; per-subsystem breakdown
External power monitorPower the phone from a monitor (battery bypass on dev boards, or USB meter for rough numbers)Most accurate total power; needs hardware
# Fuel gauge sampling (units are usually microamps / microvolts; sign convention varies by device)
adb shell "while true; do echo $(date +%s%N) $(cat /sys/class/power_supply/battery/current_now) \
  $(cat /sys/class/power_supply/battery/voltage_now); sleep 0.2; done" > power.log

# batterystats
adb shell dumpsys batterystats --reset
adb shell dumpsys battery unplug          # simulate unplugged so stats accumulate over USB
# ... run workload ...
adb shell dumpsys batterystats > batterystats.txt
adb shell dumpsys battery reset
Erequest = ∑i (Vi · Ii − Pidle) · Δti    J per 100 tokens = E / (Ntokens / 100) Pidle is the baseline power in the same screen and radio state. A faster run at higher power can still use less energy (race to idle).

Step 6: the thermal soak (the differentiating measurement)

Most published numbers are a single cold run. Instead, generate continuously for 10-20 minutes and plot throughput, temperature, CPU frequency and power over time. This reveals when throttling begins, how far throughput falls, and the sustained-to-peak ratio, which is the number that actually matters for a shipped feature. For the thermal framework itself see Power and thermal.

# thermal_soak.sh - run generations back-to-back and sample system state every 2 s
DUR=${1:-600}
adb shell "cd /data/local/tmp/llama; end=\$((\$(date +%s)+$DUR)); \
  while [ \$(date +%s) -lt \$end ]; do ./llama_main --model_path=llama3_2_1b_8da4w.pte \
  --tokenizer_path=tokenizer.model --prompt='Write a long story about a lighthouse.' \
  --seq_len=512 --cpu_threads=4 2>&1 | grep -i 'tokens/s'; done" > soak_tps.log &

while kill -0 $! 2>/dev/null; do
  ts=$(date +%s)
  temps=$(adb shell "for z in /sys/class/thermal/thermal_zone*; do echo -n \$(cat \$z/type)=\$(cat \$z/temp),; done")
  freqs=$(adb shell "cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_cur_freq | tr '\n' ,")
  status=$(adb shell "dumpsys thermalservice | grep -m1 'Thermal Status'")
  echo "$ts;$temps;$freqs;$status" >> soak_sys.log
  sleep 2
done
decode tok/s
  45 |■■■■■■■■■■■
  40 |           ■■■■■                          peak  (first 60-120 s)
  35 |                ■■■■■
  30 |                     ■■■■■■■■■■■■■■■■■■■■■  sustained plateau
  25 |
     +-----+-----+-----+-----+-----+-----+-----+--▶ time (min)
     0     1     2     3     4     5     6     10
  SoC temp rises ~15-25 °C; big-core frequency steps down; ratio sustained/peak ≈ 0.6-0.8 (typical, device-dependent)

Step 7: compare BF16 and INT4

Export the same model twice (BF16/FP32 and 8da4w INT4) and fill a comparison table on all metrics: load time, TTFT, prefill tok/s, decode tok/s, peak RSS, file size, energy per 100 tokens and sustained ratio. Expect roughly 2-3x faster decode, 3-4x faster prefill (with KleidiAI), about 50-60% smaller file and 30-40% lower peak memory for the quantized model. Deviations beyond that mean something is misconfigured (see the target numbers section).

Deliverables checklist

  • Model exported and running on device with coherent output.
  • Harness scripted so a full measurement run is one command, producing a CSV plus a summary.
  • BF16 vs INT4 comparison across all metrics.
  • Ten-minute sustained-load curve with temperature and frequency overlay.
  • Numbers sanity-checked against reference values.
  • Short write-up of what surprised you, with the raw data committed.
Common pitfall Measuring TTFT with a different prompt length each time, or reporting "tokens/s" without saying whether it is prefill or decode. The two differ by 5-10x on the same device; a single mixed number is useless.
Interview angle "How would you benchmark an on-device LLM fairly?" Cover environment control, warm-up, repeated runs with percentiles, separating load/TTFT/prefill/decode, peak memory from VmHWM or PSS, energy per token with idle baseline subtraction, a 10+ minute thermal soak with the sustained/peak ratio, and cross-checking against a second runtime.

Project 2: quantization sweep and layer-wise error analysis

Duration: 3-4 weeks. Goal: understand quantization numerically and build the tool that finds where it breaks. Skills: PTQ vs QAT, weight and activation quantization, mixed precision, layer-wise mismatch, host reference vs on-device validation, fixed-point representation. Output: a layer-wise activation diff tool plus a sweep results table. For the theory of scales, zero-points and granularity, see Edge AI.

Analogy

Finding quantization error layer by layer is like finding a leak in a long water pipeline. Measuring only at the tap (final accuracy) tells you water is missing, not where. Installing a pressure gauge at every junction (per-layer comparison against the FP32 reference) shows exactly which segment loses pressure, whether the loss compounds downstream, and which few segments need thicker pipe (higher precision) instead of replacing the entire line.

Junctions are layer boundaries, gauges are cosine similarity and SQNR, the leaky segment is a quantization-sensitive layer, and thicker pipe is keeping that layer in INT8 or FP16 while the rest stays INT4.

Part A: the configuration sweep

Take one model (Llama-3.2-1B is ideal) and walk the configuration space, measuring both quality and cost every time. Record every result in one CSV; the sweep table is itself a useful artifact.

VariableValues to tryWhat you learn
Weight bit-width4-bit, 8-bitThe primary size and decode-speed lever
Group size32, 64, 128, 256, per-channelGranularity vs scale-metadata overhead (a 16-bit scale per 32 weights adds 0.5 bits per weight)
Activation quantizationnone (weight-only), dynamic 8-bit, static 8-bit, static 16-bitWhy NPUs demand static quantization, and its accuracy cost
Embedding quantizationoff, 8-bit, 4-bit group 32Embeddings are a large share of small models (about 21% for Llama-3.2-1B)
LM head (output projection) precision4, 6, 8 bitsThe output layer is unusually sensitive: errors land directly on logits
Algorithmround-to-nearest, GPTQ, AWQ, SpinQuant (rotations), QAT (+LoRA)How much accuracy smarter PTQ or training recovers
Calibration setsize (32-512 samples), domain matchSensitivity of static ranges to data

Measure quality two ways: perplexity on held-out text (statistical degradation) and at least one task-level score (for example a small multiple-choice benchmark, or exact-match on your product's prompts) for behavioural degradation. Perplexity can move very little while a task collapses, and vice versa.

INT8 versus INT4, what to expect before you sweep. INT8 (or FP8) weights are about 2× smaller than FP16 and usually near-lossless after PTQ; that is the control row. INT4 weight-only is about 4× smaller and is the decode-speed lever, but quality loss is larger and shows first on reasoning, maths, rare languages and long context. Keep embeddings and the LM head at 6-8 bits in the first INT4 recipe. Two "4-bit" rows that differ in group size or algorithm are not comparable; write the full recipe on every CSV line.

# Sweep driver sketch (ExecuTorch exporter; one row per config)
for bits in 4 8; do for gs in 32 128 256; do for emb in "none" "torchao:8,0" "torchao:4,32"; do
  name="l32_1b_w${bits}_g${gs}_emb${emb//[:,]/_}"
  python -m extension.llm.export.export_llm base.model_class=llama3_2 \
    base.checkpoint=$CKPT base.params=$PARAMS model.use_kv_cache=True \
    model.use_sdpa_with_kv_cache=True quantization.qmode="torchao:8da${bits}w" \
    quantization.group_size=$gs quantization.embedding_quantize="$emb" \
    backend.xnnpack.enabled=True export.output_name="$name.pte"
  python eval_ppl.py --pte $name.pte --data heldout.txt >> sweep.csv   # host-side quality
  python bench.py --pte $name.pte >> sweep.csv                         # device-side cost
done; done; done

Part B: layer-wise error analysis (the core skill)

Analysing quantization error and layer-wise mismatches is a named responsibility in many edge AI roles, and very few candidates can demonstrate it. The method:

  1. Reference run Run the FP32 model on the host (x86 or Arm workstation) and capture the output of every layer boundary (attention projections, attention output, MLP gate/up/down, norms, residual stream after each block).
  2. Quantized run Run the quantized model on the same inputs and capture the same tensors: first as a simulated (fake-quantized) model on the host, then from the device.
  3. Diff per tensor Compute cosine similarity, mean and max absolute error, and SQNR for each captured tensor.
  4. Rank and plot Rank layers by divergence; plot SQNR across depth to see whether error compounds or recovers.
  5. Isolate Quantize one layer at a time (everything else FP32) to measure each layer's individual sensitivity; this separates "this layer is fragile" from "this layer receives bad inputs".
  6. Derive a recipe Keep the most sensitive layers at higher precision, re-run the sweep and the Project 1 harness, and quantify quality recovered versus latency and memory cost.
SQNRdB = 10 · log10( Σ xi2 / Σ (xi − x̂i)2 )     cos(x, x̂) = (x · x̂) / (‖x‖ ‖x̂‖) x is the FP32 reference tensor and x̂ the quantized one. Each extra bit of uniform quantization adds about 6 dB of SQNR in theory. As rough rules of thumb: above 30 dB is usually safe, 20-30 dB is worth watching, below ~15-20 dB or cosine below ~0.99 on the residual stream usually shows up as quality loss.

Host-side tool: capture and diff with forward hooks

import torch, math, csv
from collections import OrderedDict

def capture(model, inputs, names_filter=lambda n, m: isinstance(m, torch.nn.Linear)
                                              or "norm" in n.lower() or n.endswith("layers")):
    acts, hooks = OrderedDict(), []
    for name, mod in model.named_modules():
        if names_filter(name, mod):
            hooks.append(mod.register_forward_hook(
                lambda m, i, o, name=name: acts.__setitem__(name, (o[0] if isinstance(o, tuple) else o)
                                                             .detach().float().cpu())))
    with torch.no_grad():
        model(*inputs)
    for h in hooks: h.remove()
    return acts

def metrics(ref, q):
    ref, q = ref.flatten(), q.flatten()
    err = ref - q
    sqnr = 10 * math.log10((ref.pow(2).sum() / err.pow(2).sum().clamp_min(1e-20)).item())
    cos = torch.nn.functional.cosine_similarity(ref, q, dim=0).item()
    return {"sqnr_db": sqnr, "cosine": cos,
            "mae": err.abs().mean().item(), "max_abs": err.abs().max().item(),
            "ref_absmax": ref.abs().max().item()}        # large absmax hints at outliers

ref_acts = capture(fp32_model, (tokens,))
q_acts   = capture(fake_quant_model, (tokens,))          # e.g. torchao-quantized copy

rows = [{"layer": n, **metrics(ref_acts[n], q_acts[n])} for n in ref_acts if n in q_acts]
rows.sort(key=lambda r: r["sqnr_db"])                    # worst first
with open("layerwise.csv", "w", newline="") as f:
    w = csv.DictWriter(f, fieldnames=rows[0].keys()); w.writeheader(); w.writerows(rows)
for r in rows[:10]:
    print(f'{r["layer"]:50s} SQNR {r["sqnr_db"]:6.1f} dB  cos {r["cosine"]:.5f}')

Single-layer sensitivity sweep (the "isolate" step):

def per_layer_sensitivity(build_model, layer_names, quantize_only, eval_fn):
    base = eval_fn(build_model())                      # FP32 score (e.g. perplexity)
    out = []
    for name in layer_names:
        m = quantize_only(build_model(), {name})       # quantize a single layer
        out.append((name, eval_fn(m) - base))          # degradation attributable to this layer
    return sorted(out, key=lambda t: -t[1])

Capturing intermediate tensors on the device

StackHow to get per-layer outputs
ExecuTorchGenerate an ETRecord at export time and run with ETDump plus a debug buffer enabled; the devtools Inspector aligns on-device intermediate outputs with the original graph nodes so you can compute numeric gaps per operator.
QNN / QAIRTqnn-net-run --debug writes every intermediate tensor; the SDK's accuracy debugger tooling compares them against a framework reference layer by layer.
ONNX Runtimeonnxruntime.quantization.qdq_loss_debug augments a model to output intermediate tensors and matches FP32 and QDQ activations (create_activation_matching, compute_activation_error).
AIMETQuantAnalyzer reports per-layer sensitivity, activation/weight ranges, and the effect of enabling quantizers one at a time.
LiteRTThe quantization debugger reports per-layer statistics between float and quantized models; or add intermediate outputs to the model signature.
llama.cppEvaluation callbacks (the eval-callback example) print per-tensor statistics during a forward pass; compare against the f16 GGUF.

Comparing host-simulated quantization with the device run is itself informative. If fake-quant on the host matches FP32 well but the device diverges, the problem is not the quantization scheme but the device implementation: different rounding, accumulator width, requantization, fused operator behaviour, or an FP16 overflow in a GPU/NPU kernel.

Typical findings to look for

Outlier channels

A few hidden dimensions in transformer activations carry values 10-100x larger than the rest. Per-tensor activation scales then waste most of the integer range. Signs: huge ref_absmax, low SQNR at inputs to the next projection. Fixes: per-channel or per-group, SmoothQuant-style scale migration, rotations (SpinQuant/QuaRot), or 16-bit activations.

Down projection and LM head

The MLP down projection sees the widest activation ranges; the LM head writes directly into logits. Both are common candidates to keep at 8-bit when everything else is 4-bit.

First and last blocks

Early layers shape the residual stream; late layers are close to the output. Sensitivity is often U-shaped across depth.

Norms and softmax

RMSNorm, softmax and residual additions are precision-sensitive; most stacks keep them at FP16/FP32 or 16-bit integer even in "INT8" models.

Error accumulation

SQNR of the residual stream often degrades gradually with depth, while individual layer outputs may recover; residual connections both carry and dilute error.

Attention vs MLP

Q/K projections affect attention scores exponentially through softmax; V and O projections behave more linearly. Different sensitivities justify different precisions.

Mixed-precision fallback recipe

  1. Start uniform Everything at the target precision (for example W4 group 32/128, A8 dynamic).
  2. Rank Use the single-layer sensitivity list, not just the full-model diff.
  3. Promote greedily Move the top-k layers to W8 (or FP16) one at a time until the quality target is met.
  4. Price it Each promoted layer costs size and decode time in proportion to its parameter share; report the Pareto curve of quality vs size vs latency.
  5. Validate on device Re-run the Project 1 harness and the device-side diff; mixed precision can split graphs on NPUs if the backend cannot run both precisions in one partition.

Also explore

  • SpinQuant and QAT + LoRA: the two quantization-aware recipes with published Llama 3.2 numbers. Compare their quality against your PTQ results at equal bit-width.
  • GPTQ / AWQ: weight-only PTQ that uses calibration data to minimise output error or protect salient channels.
  • AIMET techniques: cross-layer equalisation, bias correction and AdaRound for CNN-style models on NPUs.
  • Operator fusion: inspect the exported graph before and after lowering (for example with a graph visualiser or the printed edge program) and understand which ops fused and why fused quantized kernels change numerics slightly.
  • Fixed-point representation: verify by hand how an INT8 x INT8 multiply accumulates into INT32 and requantizes with a multiplier and shift; this is exactly what NPU kernels do.
q = clamp( round(x / s) + z, qmin, qmax )    x̂ = s · (q − z)    yint32 = Σ (qw − zw)(qa − za),   qy = round(yint32 · swsa/sy) + zy s is scale, z zero-point. The combined rescale factor swsa/sy is implemented as an integer multiplier and right shift on integer hardware.

Deliverables checklist

  • Full sweep table: quality, latency, size and memory for every configuration.
  • Layer-wise diff tool working for host reference vs host fake-quant vs device.
  • Ranked list of the most sensitive layers with a hypothesis for each.
  • Evidence-derived mixed-precision recipe validated on the Project 1 harness.
  • PTQ vs rotation-based vs QAT comparison at equal bit-width.
Common pitfall Judging quantization by one metric on a handful of prompts. A 0.2 perplexity increase can hide a collapse on arithmetic or non-English inputs, and a few "looks fine" chats prove nothing. Use a fixed evaluation set that includes your product's hard cases and compare token-level outputs against the FP32 model.
Interview angle "Accuracy dropped after INT8 or INT4 quantization; how do you debug it?" Walk through: confirm the FP32 converted model matches the original (rule out conversion bugs), check calibration data, run a layer-wise diff (cosine/SQNR) to find where divergence starts, run single-layer sensitivity, inspect ranges for outliers, then fix with finer granularity, better algorithms (GPTQ/AWQ/rotations/AdaRound), 16-bit activations or mixed precision for the few sensitive layers, and QAT as the last resort.

Project 3: NPU deployment on vendor silicon

Duration: 4-5 weeks. Goal: a model running on the NPU (Snapdragon Hexagon HTP as the worked example) integrated into an Android app through the native C API. Skills: vendor SDK (QAIRT/QNN), HTP deployment, static quantization, model splitting, operator support gaps. Output: working NPU inference plus an honest CPU vs GPU vs NPU comparison including power. This is the hardest project and the most valuable, because few engineers outside chipset vendors have done a full NPU deployment. Expect SDK version mismatches, unsupported operators and cryptic errors; debugging them is the curriculum.

Analogy

An NPU is like a high-speed bottling plant. It fills thousands of identical bottles per minute with astonishing efficiency, but only if every bottle is the same size (static shapes), the recipe is fixed before the shift starts (static quantization, compiled graph), and every step is one the machines know (supported operators). Hand it an odd-shaped bottle and it stops the line, sends the bottle to a person at the side table (CPU fallback), waits, and restarts. A few odd bottles per batch and the plant is slower than just having people do it all by hand.

Bottle size is tensor shape, the fixed recipe is the context binary with baked-in scales, the side table is the CPU, and the stop-start cost is the data transfer and synchronisation at every partition boundary.

The constraint that changes everything: static quantization

The Hexagon NPU is built around fixed-point arithmetic with quantization parameters known at compile time. For LLMs the typical scheme is 4-bit (or 8-bit) weights with 16-bit static activations (W4A16 / W8A16); for CNNs, full INT8. On the CPU path in Projects 1-2, activations were quantized dynamically: scales computed at runtime from each tensor's actual range. The NPU cannot afford a data-dependent range computation per tensor per step, so ranges must come from calibration data. If production inputs exceed the calibrated range, values clip. Understanding why the NPU needs this, and measuring what it costs in accuracy, is the core lesson of the project.

TargetWeightsActivationsGranularity
CPU / GPU (ExecuTorch XNNPACK, ONNX Runtime, LiteRT)4-bit group-wise (or 8-bit)8-bit dynamic, scale computed at runtimeGroup size 32 or 128
Qualcomm NPU (QNN / QAIRT)4-bit or 8-bit16-bit static, parameters fixed at compile time (8-bit for many CNNs)Per-channel in QAIRT; ExecuTorch QNN delegate commonly uses block/group 32 for 4-bit
llama.cpp (Q4_0 on CPU)4-bit, group 32 (output layer often 6-bit)Quantized on the fly to 8-bit for dot products, backend-dependentBlocks of 32

Step 1: prepare the model (AI Hub path)

pip install qai-hub-models
qai-hub configure --api_token <YOUR_TOKEN>

# Export a supported LLM to QNN context binaries, quantized and split for the NPU
python -m qai_hub_models.models.llama_v3_2_3b_instruct.export \
    --device "Snapdragon 8 Elite QRD" \
    --skip-inferencing --skip-profiling \
    --output-dir ./export_8_elite
# If you produced a custom quantized checkpoint in Project 2, pass it through
# with the export script's checkpoint option instead of the default weights.

For the 1B model, use the corresponding 1B export if available in the zoo, or go through the ExecuTorch Qualcomm backend, which has its own Llama export scripts that produce a .pte with embedded HTP context binaries.

Why the model is split into several binaries

  • The NPU process has limits on how much memory a single graph and its mapped weights can use; large models are therefore partitioned by layers into several context binaries executed in sequence.
  • Prefill and decode are compiled as separate graphs with different static shapes (for example 128-token chunks for prefill, 1 token for decode) that share weights ("weight sharing") so memory is not doubled.
  • The KV cache is managed outside the graphs as input/output tensors of fixed maximum context length; the runtime updates it between calls.

Step 2: assemble the Genie bundle

mkdir -p genie_bundle

# 1. Model config for Genie (start from the sample config for your model in the vendor tutorials)
cp <tutorial-configs>/llama_v3_2_3b.json genie_bundle/genie_config.json
cp <tutorial-configs>/htp_backend_ext_config.json genie_bundle/

# 2. Hexagon DSP-side libraries: MUST match the SoC architecture
cp "$QNN_SDK_ROOT"/lib/hexagon-v73/unsigned/* genie_bundle/   # Snapdragon 8 Gen 2
# cp "$QNN_SDK_ROOT"/lib/hexagon-v75/unsigned/* genie_bundle/ # 8 Gen 3
# cp "$QNN_SDK_ROOT"/lib/hexagon-v79/unsigned/* genie_bundle/ # 8 Elite

# 3. Android (ARM) side libraries and the CLI runner
cp "$QNN_SDK_ROOT"/lib/aarch64-android/* genie_bundle/
cp "$QNN_SDK_ROOT"/bin/aarch64-android/genie-t2t-run genie_bundle/

# 4. Context binaries and tokenizer from the export step
cp export_8_elite/*.bin export_8_elite/tokenizer.json genie_bundle/

# 5. Verify: every file listed in ctx-bins[] of genie_config.json exists in the bundle

An abridged Genie dialog config. Field names (ctx-bins, kv-dim, sampler keys) vary between SDK versions. Always start from the sample JSON shipped for your model in that QAIRT / AI Hub export, then change only paths and context length.

{
  "dialog": {
    "type": "basic",
    "context":  { "size": 4096, "n-vocab": 128256, "bos-token": 128000, "eos-token": [128001, 128009] },
    "sampler":  { "seed": 42, "temp": 0.7, "top-k": 40, "top-p": 0.95 },
    "tokenizer": { "path": "tokenizer.json" },
    "engine": {
      "n-threads": 3,
      "backend": {
        "type": "QnnHtp",
        "QnnHtp": { "use-mmap": true, "poll": true, "kv-dim": 128 },
        "extensions": "htp_backend_ext_config.json"
      },
      "model": {
        "type": "binary",
        "binary": { "ctx-bins": ["model_part_1_of_3.bin", "model_part_2_of_3.bin", "model_part_3_of_3.bin"] }
      }
    }
  }
}

Step 3: run on device

adb push genie_bundle /data/local/tmp/
adb shell
  cd /data/local/tmp/genie_bundle
  export LD_LIBRARY_PATH=$PWD          # ARM-side libraries
  export ADSP_LIBRARY_PATH=$PWD        # DSP-side (Hexagon skeleton) libraries
  ./genie-t2t-run -c genie_config.json \
    -p "<|begin_of_text|><|start_header_id|>user<|end_header_id|>

Explain NPU static quantization.<|eot_id|><|start_header_id|>assistant<|end_header_id|>

"

Step 4: integrate the C API into an app

Genie wraps the tokenizer, the QNN backend, KV-cache management, decoding and sampling behind a dialog API; the engine executes forward passes on the HTP (with CPU for the rest). Wrapping it in JNI gives you the complete path: PyTorch checkpoint to quantized model to context binaries to NPU to Android UI. The C symbols below match the public dialog-API shape (JSON config, create, query with a callback, free). Header names, status codes, callback signatures and JSON keys are release-specific — copy them from the QAIRT samples for the version you link. Do not invent extra flags or config fields.

#include "GenieDialog.h"
#include <fstream>
#include <sstream>
#include <string>

struct Ctx { std::string out; void (*on_piece)(const char*, void*); void* user; };

static void on_response(const char* piece, const GenieDialog_SentenceCode_t code, const void* ud) {
    auto* c = static_cast<Ctx*>(const_cast<void*>(ud));
    if (piece) { c->out += piece; if (c->on_piece) c->on_piece(piece, c->user); }
    (void)code;   // BEGIN / CONTINUE / END / COMPLETE / ABORT
}

class NpuLlm {
    GenieDialogConfig_Handle_t cfg_ = nullptr;
    GenieDialog_Handle_t dlg_ = nullptr;
public:
    bool init(const std::string& config_path) {
        std::ifstream f(config_path); std::stringstream ss; ss << f.rdbuf();
        if (GenieDialogConfig_createFromJson(ss.str().c_str(), &cfg_) != GENIE_STATUS_SUCCESS) return false;
        return GenieDialog_create(cfg_, &dlg_) == GENIE_STATUS_SUCCESS;   // loads context binaries
    }
    std::string ask(const std::string& prompt, void (*cb)(const char*, void*), void* user) {
        Ctx c{{}, cb, user};
        GenieDialog_query(dlg_, prompt.c_str(), GENIE_DIALOG_SENTENCE_COMPLETE, on_response, &c);
        return c.out;
    }
    ~NpuLlm() { if (dlg_) GenieDialog_free(dlg_); if (cfg_) GenieDialogConfig_free(cfg_); }
};

Before creating the dialog inside an app, set ADSP_LIBRARY_PATH (with setenv) to the directory containing the Hexagon skeleton libraries (for example the app's native library directory or an extracted folder), and ship the ARM-side QNN libraries in jniLibs/arm64-v8a.

Step 5: the honest three-way comparison

BackendMeasureTypical pattern (same 1-3B model)
CPU (XNNPACK + KleidiAI or llama.cpp)TTFT, prefill, decode, RSS, power, sustained curveUniversal fallback; decent decode, weakest prefill, highest power, throttles first
GPU (OpenCL/Vulkan, e.g. Adreno)SameGood prefill; decode similar to CPU (both bandwidth-bound); competes with UI rendering
NPU (Hexagon HTP)Same plus accuracy delta from static quantizationMuch faster prefill (often several times), comparable or better decode, lowest energy per token, best sustained ratio

Report energy per 100 tokens and the ten-minute curve, not just peak tok/s. Vendors publish peak numbers; independent, power-inclusive, sustained comparisons are rare and far more useful.

NPU deployment gotchas (and fixes)

SymptomLikely causeFix
NPU slower than CPUGraph split into many partitions; unsupported op in the middle forces NPU-to-CPU round tripsProfile per-op placement; rewrite or replace the op; fuse; move pre/post-processing out of the graph
"Op not supported" or silent CPU fallbackOp, data type, rank, or attribute not in the backend op set (for example 5-D tensors, certain gather/scatter, dynamic slicing, int64)Check the backend's supported-ops table; decompose into supported ops; cast int64 indices to int32
Compilation fails with dynamic shapesNPU compilers require static shapesExport fixed shapes; separate prefill and decode graphs; pad inputs to buckets
Multi-second first loadGraph compiled/finalized on device at startupUse precompiled context binaries or runtime context caching (ORT EP context, LiteRT compilation cache)
Library load or "unable to open session" errorsHexagon library version does not match SoC; ADSP_LIBRARY_PATH wrong; SDK version of runtime libs differs from the one that built the context binaryMatch hexagon-vXX to the SoC; keep build and runtime SDK versions identical; check logcat for FastRPC errors
Accuracy worse than CPU INT8Static activation ranges clip; 8-bit activations too coarse for transformersBetter calibration data; 16-bit activations; per-channel weights; keep sensitive ops in FP16 where supported
Out-of-memory on the NPU sideGraph or weights exceed the NPU session memory limitsSplit the model into multiple context binaries; weight sharing between graphs; smaller context length
Latency jitterNPU power/clock mode set to power-saver; DSP shared with other clients (camera, audio)Set performance mode (burst or sustained high performance) appropriately; avoid contention; measure with other workloads active
Great benchmark, poor appPre/post-processing on CPU dominates; copies between Java and native buffersZero-copy shared buffers (ION/dmabuf, rpcmem), native pre-processing, batching

Graph partitioning, visualised

Ideal (1 partition):     [ NPU: conv..attn..mlp..conv ................ ]      1 dispatch

Bad (5 partitions):      [NPU][CPU:op X][NPU][CPU:op Y][NPU]
                             ▲copy+sync  ▲copy+sync ▲copy+sync ▲copy+sync
Each boundary costs data transfer, format conversion (quantize/dequantize, layout)
and a synchronisation round trip; with many small partitions the NPU sits idle.

HTP context caching and performance settings

  • Context binary: the serialized, finalized graph for one SoC architecture. Loading it skips graph optimisation and is typically 10-100x faster than compiling from the model at startup. It is not portable across Hexagon versions, and often not across SDK versions.
  • On-device caching: if you must compile on device (for example via ORT QNN EP), write the compiled context to app storage on first run and reuse it; invalidate the cache when the model, SDK or OS/driver version changes.
  • Performance modes: burst gives maximum clocks for short tasks; sustained high performance gives a stable profile for long workloads; power saver trades latency for energy. Match to the duty cycle.
  • VTCM and spill-fill: the HTP has fast on-chip memory (VTCM); graphs that exceed it spill to DDR. Tiling and model splitting keep working sets on-chip.

Deliverables checklist

  • Model exported to context binaries (via AI Hub, QAIRT or ExecuTorch QNN backend).
  • CLI runner generating coherent output on the NPU.
  • C API integrated into an Android app via JNI.
  • Three-way CPU/GPU/NPU comparison with power and sustained curves.
  • Documented accuracy cost of static vs dynamic activation quantization.
  • A written log of every operator gap and SDK issue encountered, with the workaround. This log is gold in interviews.
Common pitfall Mixing SDK versions: building context binaries with one QAIRT release and running them with runtime libraries from another. Symptoms range from load failures to subtly wrong outputs. Pin one SDK version per artifact and ship the matching libraries with it.
Interview angle "Why can't the NPU use dynamic quantization, and what does that cost?" Fixed-function integer pipelines need scales at compile time to pre-compute requantization multipliers and to avoid a data-dependent reduction per tensor; static ranges from calibration can clip outliers or waste resolution. Mitigations: 16-bit activations for LLMs, good calibration data, per-channel weights, mixed precision. Then mention partitions, CPU fallback, static shapes and context caching unprompted.

Project 4: KV-cache engineering

Duration: 3-4 weeks. Goal: own the memory bottleneck that limits on-device long context. Skills: KV caching, KV-cache quantization, footprint profiling, attention internals. Output: a KV-cache profiler plus a quantization study across bit-widths and context lengths. The KV cache grows linearly with sequence length and at long contexts it exceeds the model weights; for long-context edge deployment, compressing it is often more impactful than squeezing weights further. Yet most tutorials focus only on weight quantization. For attention mechanics, see Transformers and LLMs.

Analogy

The KV cache is like the minutes of a long meeting. Every new speaker (token) must be able to glance back at everything said so far, so the secretary writes down a summary card (key) and the content (value) for each contribution. Rereading the whole recording each time would be absurd, so the cards are kept. But the pile of cards grows with every sentence, and in a long meeting it can outweigh the rulebook on the table (the model weights). You can write cards in shorthand (quantized KV), keep only the last hour plus the opening remarks (sliding window with attention sinks), file cards in fixed-size folders instead of one giant stack (paging), and reuse the standard agenda cards every meeting (prefix caching).

Cards are K and V vectors per token per layer, the rulebook is the weights, shorthand is INT8/INT4 KV, the opening remarks are attention-sink tokens, folders are fixed-size blocks, and the agenda is the shared system prompt.

The memory math

KV bytes = 2 × nlayers × nkv_heads × dhead × T × B × bytesper element 2 for keys and values; T context length in tokens; B batch (usually 1 on device). With grouped-query attention, nkv_heads is smaller than the number of query heads, shrinking the cache proportionally.

Worked example with Llama-3.2-1B (from its params.json: dim 2048, 16 layers, 32 query heads, 8 KV heads, so head dim = 2048 / 32 = 64):

per token (FP16) = 2 x 16 layers x 8 kv_heads x 64 head_dim x 2 bytes = 32,768 B = 32 KiB
per token (INT8)  = 16 KiB (+ small scale overhead)
per token (INT4)  =  8 KiB (+ scales)

Llama-3.2-3B: 28 layers, 24 query heads, 8 KV heads, head_dim 128
per token (FP16) = 2 x 28 x 8 x 128 x 2 = 114,688 B = 112 KiB

A 7B model WITHOUT GQA (32 layers, 32 KV heads, head_dim 128)
per token (FP16) = 2 x 32 x 32 x 128 x 2 = 524,288 B = 512 KiB   (16x the 1B model)
Context1B FP16 KV1B INT4 KV3B FP16 KV3B INT4 KV7B MHA FP16 KV
51216 MiB4 MiB56 MiB14 MiB256 MiB
2k64 MiB16 MiB224 MiB56 MiB1 GiB
4k128 MiB32 MiB448 MiB112 MiB2 GiB
8k256 MiB64 MiB896 MiB224 MiB4 GiB
16k512 MiB128 MiB1.75 GiB448 MiB8 GiB
32k1 GiB256 MiB3.5 GiB896 MiB16 GiB
128k4 GiB1 GiB14 GiB3.5 GiB64 GiB

INT4 weights are about 0.7-1.1 GiB for the 1B model and roughly 1.7-2.3 GiB for the 3B (depending on how embeddings and the output layer are quantized). Crossover (FP16 cache = INT4 weights) is weight bytes / 32 KiB per token for the 1B: about 22k tokens at 0.7 GiB and about 35k at 1.1 GiB. For the 3B (112 KiB/token) it is about 16-20k tokens. A 7B without GQA (512 KiB/token, ~3.5 GiB INT4 weights) crosses below 8k. Those are exactly the context lengths long-document and chat-history features want. Also remember that NPU and many CPU runtimes allocate the cache statically for the maximum context length, so you pay the full cost even for short prompts.

Part A: profile it

  1. Footprint table Measure actual RSS (not just the formula) at 512, 1k, 2k, 4k, 8k and 16k context for the 1B and 3B models; the difference from the formula reveals allocator overhead and duplicated buffers.
  2. Crossover point Find the context length where cache bytes exceed weight bytes.
  3. Speed impact Plot decode tok/s against current context length: each decode step reads the whole cache, so decode slows as the conversation grows.
  4. Allocation pattern Trace allocations over a long generation (heapprofd in Perfetto, or malloc hooks): growth, fragmentation and reallocation stalls when a contiguous cache is resized.
  5. Kill threshold Find the context length at which the low-memory killer starts terminating your process on an 8 GB device with a real foreground workload.
def kv_bytes(layers, kv_heads, head_dim, tokens, bytes_per=2, batch=1, group=None, scale_bytes=2):
    base = 2 * layers * kv_heads * head_dim * tokens * batch * bytes_per
    if group:                                   # per-group scales for quantized KV
        base += 2 * layers * kv_heads * (head_dim // group) * tokens * batch * scale_bytes
    return base

for T in (512, 1024, 2048, 4096, 8192, 16384):
    fp16 = kv_bytes(16, 8, 64, T) / 2**20
    int4 = kv_bytes(16, 8, 64, T, bytes_per=0.5, group=32) / 2**20
    print(f"{T:6d}  fp16 {fp16:7.1f} MiB   int4(g32) {int4:7.1f} MiB")

Why decode slows as context grows

bytes read per decode step ≈ Wbytes + KVbytes(t)     decode tok/s ≤ BWeffective / (Wbytes + KVbytes(t)) W is weight bytes, t the current context length, BW effective memory bandwidth (typically 40-60% of the LPDDR peak on phones). When KV(t) approaches W, decode speed roughly halves compared with a short context.

Part B: compress and manage it

TechniqueHow it worksWhat to measure
KV quantization (8 / 4 / 3 bits)Store K and V as low-bit integers with per-group scales; dequantize inside attentionMemory saved vs perplexity and task quality; where it falls apart. INT8 is usually near-lossless; well-designed 4-bit and even ~3-bit schemes can have small loss
Per-channel keys, per-token valuesKeys have outlier channels (a few dimensions consistently large), so quantize keys along channels; values are better quantized per tokenVerify on your model by plotting K and V magnitude per channel; compare both groupings
Sliding-window attentionKeep only the last W tokens per layer (ring buffer); memory bounded at WQuality on inputs longer than W; retrieval of facts from early in the context
Attention sinksAlways keep the first few tokens plus the window; models place large attention mass on initial tokens, and dropping them destabilises generationPerplexity on very long streams with and without sinks
Paged / block-allocated cacheFixed-size blocks (for example 16-64 tokens) mapped by a block table; no large contiguous reallocations; blocks shareable between sequences. Contiguous waste ≈ (Tmax − Tactual) × KV per token; paged waste ≤ one blockFragmentation, allocation stalls, memory overhead of the last partially filled block
Eviction policiesDrop tokens that receive little attention (heavy-hitter style), keep recent and sink tokensWhat can be dropped with least damage; stability across tasks
Prefix (prompt) cachingCompute KV for a byte-identical system prompt or shared document once, persist it, reuse for every request. Prefill savings ≈ P / (P + U)TTFT saved; storage cost; invalidation when model or prompt changes; a timestamp at the top of the prompt misses every time
Speculative decodingA tiny draft (or extra heads / prompt n-grams) proposes k tokens; the target verifies them in one pass. E[tokens per pass] = (1 − αk+1) / (1 − α); speed-up ≈ E / (1 + k · c)Draft memory, rejection rollback, extra power, weaker gains on high-entropy text and when the NPU needs a static k-token graph
Architecture choicesGQA/MQA, cross-layer KV sharing, smaller head dimsOnly when you can choose or fine-tune the model

Code: per-token INT8 value cache and ring-buffer window

import torch

def quant_per_token(x, bits=8):                     # x: [heads, T, head_dim]
    qmax = 2 ** (bits - 1) - 1
    scale = x.abs().amax(dim=-1, keepdim=True).clamp_min(1e-8) / qmax
    q = torch.clamp(torch.round(x / scale), -qmax - 1, qmax).to(torch.int8)
    return q, scale.to(torch.float16)

def quant_per_channel(x, bits=8):                   # keys: scale per head_dim channel
    qmax = 2 ** (bits - 1) - 1
    scale = x.abs().amax(dim=-2, keepdim=True).clamp_min(1e-8) / qmax
    q = torch.clamp(torch.round(x / scale), -qmax - 1, qmax).to(torch.int8)
    return q, scale.to(torch.float16)

class RingKV:
    """Sliding window + attention sinks, fixed memory."""
    def __init__(self, layers, heads, head_dim, window=2048, sinks=4, dtype=torch.float16):
        self.W, self.S = window, sinks
        shape = (layers, heads, sinks + window, head_dim)
        self.k = torch.zeros(shape, dtype=dtype); self.v = torch.zeros(shape, dtype=dtype)
        self.n = 0                                   # total tokens seen
    def slot(self):
        if self.n < self.S: return self.n            # sinks are never overwritten
        return self.S + (self.n - self.S) % self.W   # ring over the window
    def append(self, layer, k_t, v_t):               # k_t, v_t: [heads, head_dim]
        s = self.slot()
        self.k[layer, :, s] = k_t; self.v[layer, :, s] = v_t
    def advance(self): self.n += 1
    def valid(self): return min(self.n, self.S + self.W)
# Note: with RoPE, positions of cached keys are already baked in; implementations that
# roll a window either keep original positions or re-assign positions within the cache.

Code sketch: paged KV allocator in C++

struct BlockPool {
    size_t block_tokens, bytes_per_block;
    std::vector<void*> free_list;
    explicit BlockPool(size_t tokens, size_t bpt, size_t n_blocks)
        : block_tokens(tokens), bytes_per_block(tokens * bpt) {
        for (size_t i = 0; i < n_blocks; ++i) free_list.push_back(aligned_alloc(64, bytes_per_block));
    }
    void* get() { if (free_list.empty()) return nullptr; auto* b = free_list.back(); free_list.pop_back(); return b; }
    void put(void* b) { free_list.push_back(b); }
};

struct SequenceKV {
    std::vector<void*> blocks;      // block table: logical block i -> physical block
    size_t tokens = 0;
    bool append_token(BlockPool& pool) {
        if (tokens % pool.block_tokens == 0) {
            void* b = pool.get();
            if (!b) return false;       // caller evicts, compresses, or stops generation
            blocks.push_back(b);
        }
        ++tokens; return true;
    }
};
// Pre-allocating the pool once at startup avoids reallocation stalls and makes the
// memory budget explicit; the attention kernel walks the block table.

Part C: the Android systems finding

Tie the work to device reality. On a real 8 GB phone with a foreground app and background services, what is the maximum usable context length before the system kills your process, and how does each compression technique move that number? This requires understanding Android's memory management (see Linux kernel and BSP and Power and thermal for related platform detail):

  • lmkd (the userspace low-memory killer daemon) monitors memory pressure (PSI, pressure stall information) and kills processes in order of oom_score_adj: cached apps first, then services, then perceptible and finally foreground.
  • A foreground app is killed last, but a large model can push the system into killing everything else (including the launcher or music playback), which users experience as "the phone got weird".
  • Anonymous memory (heap-allocated KV cache, repacked weights) cannot be reclaimed without swap/zRAM compression; file-backed mmapped weights can be dropped and re-read.
  • zRAM compresses anonymous pages; low-entropy tensors compress well, quantized ones do not.
# Observe memory pressure and kills while increasing context length
adb shell cat /proc/pressure/memory             # PSI: some/full stall percentages
adb shell dumpsys meminfo -s <package>         # PSS summary for the app
adb logcat -b events | grep -i am_kill           # framework kill events
adb logcat | grep -i lowmemorykiller             # lmkd decisions and reasons
adb shell cat /proc/<pid>/oom_score_adj

Deliverables checklist

  • KV footprint curve vs context length for two model sizes, formula vs measured.
  • Documented crossover point where the cache exceeds the weights.
  • Quantization study at 8/4/3 bits with quality cost quantified, including key vs value grouping.
  • At least one alternative strategy implemented (sliding window with sinks, or paged).
  • Maximum usable context under real Android memory pressure, with lmkd behaviour.
Common pitfall Quoting KV size with the number of query heads instead of KV heads. For a GQA model like Llama-3.2-1B, using 32 heads instead of 8 overestimates the cache by 4x. Always read n_kv_heads from the config.
Interview angle "How much memory does the KV cache need for model X at 8k tokens, and how would you reduce it?" Write the formula, plug in layers, KV heads, head dim and bytes, give the number, then list GQA (architecture), KV quantization (with per-channel keys and per-token values), sliding window with sinks, paging (fragmentation), prefix caching (TTFT), limiting max context, and note the decode-speed effect of large caches.

Project 5: the system project

Duration: 4-6 weeks. Goal: a complete system where the model is one component among many, built to survive memory pressure, heat and bad inputs. Skills: LoRA adapters, system integration, memory- and thermal-aware inference, platform architecture. Output: a working system, a short demo, and an architecture write-up of the trade-offs. Pick one of the four designs below; each reuses the harness (P1), quantization recipe (P2), NPU path (P3) and cache management (P4). For wider design framing, see GenAI system design.

Analogy

The earlier projects built and tuned an engine; the system project builds the whole car. A great engine in a car with no brakes, no fuel gauge and no cooling system is not drivable. The system project adds the cooling (thermal-aware throttling), the fuel gauge (memory budgeting), the brakes (cancellation and refusal), the gearbox (NPU, GPU, CPU fallback) and the dashboard (telemetry).

The engine is the model runtime; the car is the service or app around it that users actually operate.

Option A: a shared on-device inference service

Build a small version of what a platform AI service is: one Android service hosting one model, shared by several client apps.

  Client app A ─┐                                      ┌─ LoRA: summarise (≈20-40 MB)
  Client app B ─┼─ AIDL (bound service, permission) ─▶ │  Scheduler / queue (fairness, priorities)
  Client app C ─┘                                      │  Policy: memory + thermal governor
                                                       │  Engine: one resident base model (mmap)
                                                       │     NPU ──▶ GPU ──▶ CPU ──▶ refuse
                                                       └─ LoRA: rewrite / classify (hot-swap)
  • Stable IPC surface: an AIDL interface with request id, prompt, options and a streaming callback; enforce a custom permission; validate inputs; cap sizes (see Binder transaction limits: stream results rather than returning large payloads).
  • One model resident, not one copy per client: memory is the scarcest resource.
  • Memory-pressure awareness: listen to onTrimMemory and PSI; shrink the KV budget, unload adapters, or unload the model when idle and under pressure.
  • Thermal-aware throttling: poll PowerManager.getThermalHeadroom() and thermal status listeners; lower the generation rate or max tokens proactively instead of being throttled arbitrarily by the OS.
  • Hot-swappable LoRA adapters: fine-tune two small adapters for different tasks (see Fine-tuning), swap at runtime, and measure the switch cost in latency and memory. Merged weights are fastest but not swappable; unmerged adapters add a small matmul per adapted layer.
  • Queueing and fairness: per-client quotas, priority for foreground callers, cancellation when the client dies (linkToDeath).
  • Graceful degradation: NPU, then GPU, then CPU, then refuse with a clear error, following a documented policy.
// IInferenceService.aidl
package com.example.edge;
import com.example.edge.IInferenceCallback;

interface IInferenceService {
    int  getApiVersion();
    long submit(String task, String prompt, in Bundle options, IInferenceCallback cb); // returns request id
    void cancel(long requestId);
    Bundle getStatus();     // backend in use, thermal state, loaded adapters, queue depth
}

// IInferenceCallback.aidl
oneway interface IInferenceCallback {
    void onToken(long requestId, String piece);
    void onComplete(long requestId, in Bundle stats);
    void onError(long requestId, int code, String message);
}
// Thermal- and memory-aware policy (Kotlin, inside the service)
class Governor(private val pm: PowerManager) {
    fun decide(req: Request): Plan {
        val headroom = pm.getThermalHeadroom(10)          // forecast 10 s ahead; 1.0 = throttling
        val status = pm.currentThermalStatus
        return when {
            status >= PowerManager.THERMAL_STATUS_SEVERE -> Plan.Refuse("device too hot")
            headroom > 0.85f -> Plan.Run(backend = Backend.NPU, maxTokens = req.maxTokens / 2,
                                          tokenDelayMs = 20)   // pace generation
            memoryTight()     -> Plan.Run(backend = Backend.NPU, maxContext = 1024)
            else              -> Plan.Run(backend = Backend.NPU, maxTokens = req.maxTokens)
        }
    }
}

Option B: an on-device assistant with local RAG (documents or device logs)

A fully local assistant that ingests documents (or device logs such as a bug report or logcat capture) and answers questions with citations. Nothing leaves the device.

Ingest ──▶ Parse/clean ──▶ Chunk (200-500 tokens, overlap) ──▶ Embed (small encoder, INT8)
                                                                   │
                                                          Vector index (flat / HNSW) + keyword index (BM25)
Query ──▶ Embed ──▶ Hybrid retrieve top-k ──▶ Rerank (optional) ──▶ Pack context within budget
                                                                   │
                         Local LLM (NPU, KV budget from P4) ──▶ Answer with cited chunk ids ──▶ UI
  • Chunking: for logs, chunk on natural boundaries (timestamps, process ids, crash blocks) rather than fixed characters, and keep metadata (time, tag, pid) for filtering.
  • Embeddings: a small sentence encoder (tens of millions of parameters) quantized to INT8 runs in milliseconds per chunk; batch ingestion in the background while charging.
  • Index: a flat index is fine up to tens of thousands of chunks; beyond that use an approximate index (HNSW) in SQLite or a small native library. Combine with keyword search, because identifiers and error codes are poorly captured by embeddings.
  • Context budget: a 40 MB log does not fit a 4-8k context; retrieval plus filtering is mandatory. Budget tokens explicitly: system prompt (prefix-cached) + retrieved chunks + question + answer.
  • Grounding and refusal: instruct the model to answer only from retrieved text with chunk citations, and to say it cannot find evidence otherwise; verify citations exist before displaying.
  • Measure the whole pipeline: ingestion time per MB, retrieval recall@k on a labelled question set, TTFT and total answer latency, peak memory, energy per question.

Option C: a wearable sensor model

An always-on activity, gesture or anomaly classifier on accelerometer/gyroscope/PPG data for a watch or band, where the budget is milliwatts, not watts.

  • Model: a tiny 1-D CNN or small recurrent model (tens of KB to a few hundred KB), INT8, with fixed windows (for example 2-second windows at 50 Hz, 50% overlap).
  • Placement: run on the sensor hub / low-power DSP or microcontroller (for example a Cortex-M with a micro NPU via a micro runtime), and wake the application processor only on detected events. Batch sensor data with hardware FIFOs to avoid waking the CPU every sample.
  • Budget: energy per inference in microjoules, average power in low milliwatts, RAM in tens to hundreds of KB, flash for weights.
  • Evaluation: per-user variation matters; evaluate with leave-one-subject-out splits, and measure false wake-ups per hour, which dominate battery cost.
  • Updates: over-the-air model updates via the companion phone, with version checks and rollback.

Option D: a real-time camera pipeline

Detection, segmentation or super-resolution on live camera frames at 30 fps, where the budget is a 33 ms frame period and the model shares the device with the camera stack and display.

Camera (YUV_420_888) ──▶ GPU/ISP resize + colour convert ──▶ NPU model (INT8, static 256-640 px)
          │                                                        │
          └──── preview surface (no copy) ◀── overlay render ◀── post-process (NMS, masks) on GPU/CPU
  • Zero-copy: use hardware buffers (AHardwareBuffer) shared between camera, GPU and NPU where the runtime supports it; each CPU copy of a 1080p frame costs milliseconds.
  • Pipelining: overlap capture, pre-processing, inference and rendering across frames; drop frames rather than queueing them (latency matters more than throughput).
  • Pre/post-processing: often slower than the model itself; move resize/colour conversion to GPU or ISP and NMS to optimized native code.
  • Thermal: camera plus display plus NPU is a heavy sustained load; measure 20-30 minutes, and adapt resolution or frame rate to thermal headroom.
  • Metrics: end-to-end frame latency (sensor to overlay), p99 frame time, dropped-frame rate, accuracy on device-captured test footage (not only curated datasets).

Deliverables for any option

  • System runs end to end on device, reproducibly.
  • Behaviour under memory pressure documented (not just the happy path).
  • Thermal behaviour under sustained use documented.
  • A short demo recording.
  • Architecture write-up explaining the design trade-offs and why.
Tip Document and share your work: a repository with raw measurement data, scripts that reproduce every chart, and a short write-up per project of "what I measured and what surprised me" is far more convincing than a demo alone. Write findings, not tutorials.
Common pitfall Designing only the happy path. Real devices are hot, low on memory, on an old SDK version, or missing an NPU. A system project that documents what happens when the model cannot be loaded, the device is throttling, or the process is killed mid-generation demonstrates product engineering.
Interview angle "Design an on-device assistant for a phone" is a common system-design prompt. Cover: model choice by RAM tier, quantization and runtime, a shared service vs in-app model, memory budget (weights + KV + runtime), local RAG for grounding, LoRA adapters per feature, thermal and memory governors, fallback chain including cloud with consent, safety filtering, model delivery and updates, and metrics to monitor.

Profiling tools and how to use them

Benchmark numbers tell you what is slow; profilers tell you why. Profile at three levels: the system (CPU scheduling, frequencies, thermal, power), the runtime (per-operator timings and placement), and the native code (hot functions and cache behaviour).

Analogy

Profiling is like diagnosing a slow delivery company. The fleet dashboard (system trace) shows which trucks were moving, idle or stuck in traffic; the per-depot logs (operator profile) show which warehouse took longest to load; and riding along with one driver (sampling profiler) shows that he stops at every red light because the route planner is wrong.

Trucks are CPU cores and accelerators, traffic is thermal throttling and contention, depots are operators or partitions, and the ride-along is simpleperf on the native engine.

ToolLevelUse it to answer
Perfetto / Android System TraceSystemWhich cores ran my threads, at what frequency; did the thermal governor cap clocks; where did the UI thread block; power rails over time
heapprofd (Perfetto)MemoryWhich native call stacks allocated the peak memory; KV cache growth; leaks
simpleperfNative codeHot functions in the engine; whether optimized kernels (dotprod, i8mm, SVE) are used; cache misses via hardware counters
Android GPU Inspector (AGI)GPUGPU utilisation, shader occupancy, memory bandwidth counters for GPU delegate workloads
Snapdragon ProfilerSoC (Qualcomm)Real-time CPU, GPU, DSP/NPU and memory bandwidth metrics on Snapdragon devices
QNN / QAIRT profiling (--profiling_level, qnn-profile-viewer)NPU runtimePer-op/per-layer HTP execution time, graph init time, DDR spill
AI Hub profile jobsNPU runtime (hosted)Per-layer timing and compute unit assignment (NPU/GPU/CPU) for each op on real devices
ExecuTorch devtools (ETDump + ETRecord + Inspector)RuntimePer-operator and per-delegate timings mapped back to the source graph; intermediate outputs
LiteRT benchmark_model --enable_op_profilingRuntimePer-op latency and which ops ran on the delegate
onnxruntime_perf_test and ORT profilingRuntimePer-node timings and EP assignment
llama-benchRuntime (LLM)Prefill and decode tok/s across thread counts, quant types, prompt lengths
Arm Performance Studio (Streamline)CPU/GPU (Arm)Hardware counters on Arm cores and Mali GPUs
Xcode Instruments / Core ML reportAppleWhich ops ran on the Neural Engine; load and prediction times

Recording a Perfetto trace with CPU frequency, scheduling and power rails

# config.pbtxt
buffers { size_kb: 131072 fill_policy: RING_BUFFER }
data_sources { config { name: "linux.ftrace" ftrace_config {
    ftrace_events: "sched/sched_switch"
    ftrace_events: "power/cpu_frequency"
    ftrace_events: "power/cpu_idle"
    ftrace_events: "thermal/thermal_temperature"
    atrace_categories: "gfx"
    atrace_categories: "view"
    atrace_apps: "com.example.edge"
} } }
data_sources { config { name: "android.power" android_power_config {
    battery_poll_ms: 250
    collect_power_rails: true
    battery_counters: BATTERY_COUNTER_CURRENT
    battery_counters: BATTERY_COUNTER_VOLTAGE
} } }
data_sources { config { name: "linux.process_stats" target_buffer: 0 } }
duration_ms: 60000

# record (the config is streamed through stdin)
adb push config.pbtxt /data/local/tmp/
adb shell "cat /data/local/tmp/config.pbtxt | perfetto --txt -c - -o /data/misc/perfetto-traces/edge.pftrace"
adb pull /data/misc/perfetto-traces/edge.pftrace
# open the file in the Perfetto UI (runs locally in a browser; the trace stays on your machine)

Add custom slices so the trace shows model phases alongside system activity:

// Kotlin
Trace.beginSection("llm_prefill"); engine.prefill(tokens); Trace.endSection()

// C++ (NDK)
#include <android/trace.h>
ATrace_beginSection("llm_decode_step");
engine.decode(tok);
ATrace_endSection();

simpleperf on the native engine

adb shell simpleperf record -p $(adb shell pidof llama_main) -g --duration 15 \
    -o /data/local/tmp/perf.data
adb shell simpleperf report -i /data/local/tmp/perf.data --sort symbol | head -40

# Hardware counters (availability varies by device)
adb shell simpleperf stat -e cpu-cycles,instructions,cache-misses,raw-l2d-cache-refill \
    -p $(adb shell pidof llama_main) --duration 10
# For app processes, the NDK's simpleperf scripts (app_profiler.py) handle symbols and flame graphs.

QNN per-layer profiling

adb shell "cd /data/local/tmp/qnn && ./qnn-net-run --backend libQnnHtp.so \
   --retrieve_context model_ctx.bin --input_list inputs.txt \
   --profiling_level detailed --output_dir out"
adb pull /data/local/tmp/qnn/out/qnn-profiling-data_0.log
qnn-profile-viewer --input_log qnn-profiling-data_0.log   # per-op cycles/time, init vs execute

A profiling workflow that finds real problems

  1. Start at the system trace Is the work actually on the expected unit? Are threads on big cores? Are frequencies capped (thermal) or low (governor idle)? Is anything else competing?
  2. Check placement In the runtime profile, confirm what fraction of ops or time runs on the accelerator and how many partitions exist.
  3. Find the top ops Typically matmuls and attention dominate; if a reshape, transpose, gather or quantize/dequantize is near the top, it is a conversion artifact to fix.
  4. Roofline sanity check For decode, compare achieved bytes/s (weights + KV per token x tok/s) with the device's memory bandwidth; if you are far below, the kernels or threading are the problem, not the hardware.
  5. Drill into native code Only then use simpleperf to see whether the expected SIMD kernels are running.
Common pitfall Profiling with high-overhead settings and then quoting the profiled latency. Detailed op profiling, debug buffers and tracing with many categories slow execution. Use profiling runs to find where time goes; use clean runs for the headline numbers.
Interview angle "Your model hits 12 tok/s but the bandwidth math says 30 should be possible; how do you find out why?" Start with a system trace (frequencies, thermal caps, core placement, other load), then runtime per-op profile (unexpected ops, partitions, fallback), then simpleperf (kernel selection, threads spinning or waiting), and validate the bandwidth assumption with an effective-bandwidth microbenchmark.

Debugging accuracy drift

Accuracy drift is when the deployed model's outputs differ from the reference by more than expected. It can come from conversion, quantization, the device implementation, or the pipeline around the model. The key is to bisect systematically instead of guessing.

Analogy

Debugging accuracy drift is like tracing why a photocopy of a photocopy looks bad. You compare each generation against the original: was the first copy already blurred (conversion), did one machine use the wrong paper (pre-processing), was the toner low (quantization), or was the last machine miscalibrated (device kernel)? Comparing adjacent generations tells you exactly which step introduced the damage.

Each generation is one stage in the pipeline, and comparing adjacent generations is the bisection between FP32 original, converted FP32, fake-quant host, and real device.

The bisection ladder

  Stage                               Compare against          Expect
  ─────────────────────────────────────────────────────────────────────────
  1 Original framework model (FP32)    ground truth labels      baseline accuracy
  2 Exported / converted FP32 model    stage 1 outputs          max abs diff ~1e-5..1e-4
  3 Converted FP16 model               stage 2                  small diff, no inf/NaN
  4 Quantized model, run on host       stage 2                  small accuracy delta
  5 Quantized model, run on device     stage 4                  near-identical outputs
  6 Full app pipeline on device        stage 5 with same inputs identical if pre/post match
  The first rung where the diff jumps is where the bug lives.

Common causes by rung

RungCauseHow to confirm
Conversion (2)Op semantics differ (for example GELU approximation, align_corners in resize, epsilon in norms), wrong opset, fused op bugLayer-wise diff FP32 vs FP32 finds the first divergent op
FP16 (3)Overflow in large activations (FP16 max is 65504), underflow in small ones, accumulation in FP16Look for inf/NaN or saturated values; keep those ops in FP32; check accumulator precision setting
Quantization (4)Bad calibration data, outliers, per-tensor scales, sensitive layersProject 2 tooling: SQNR per layer, single-layer sensitivity
Device (5)Different rounding or saturation in kernels, requantization details, NPU-specific precision, driver bugsSame quantized model host vs device, layer by layer; test on another device or backend
Pipeline (6)Wrong colour order (RGB vs BGR), normalisation, resize method, NCHW vs NHWC, image rotation from camera, wrong tokenizer or chat template, different sampling settingsDump the exact input tensor the model receives in the app and feed it to the host model

LLM-specific drift checks

  • Greedy decode comparison: run with temperature 0 on host and device with identical token ids; the first differing token position shows how quickly outputs diverge. Small logit differences can flip an argmax late in a sequence, which is expected; early divergence is not.
  • Logit metrics: compare top-1 agreement and KL divergence of next-token distributions over a fixed prompt set, not just generated text.
  • Tokenizer parity: the device tokenizer must produce the same ids as the reference (special tokens, BOS handling, whitespace normalisation).
  • Chat template: a missing or different template changes behaviour more than any quantization.
  • KV-cache bugs: outputs fine for short prompts but degrade after a certain length point to cache indexing, position (RoPE) offsets, or window wrap-around errors.
import torch.nn.functional as F
def logit_agreement(ref_logits, dev_logits):          # [T, vocab] for the same token sequence
    top1 = (ref_logits.argmax(-1) == dev_logits.argmax(-1)).float().mean().item()
    kl = F.kl_div(F.log_softmax(dev_logits, -1), F.log_softmax(ref_logits, -1),
                  log_target=True, reduction="batchmean").item()
    return {"top1_agreement": top1, "kl": kl}
Common pitfall Blaming quantization for a pre-processing bug. A large fraction of "the model is less accurate on device" reports turn out to be different image resizing, colour channel order, normalisation constants or camera orientation. Always dump and compare the exact input tensor before touching the quantization recipe.
Interview angle Interviewers like "The model is accurate on the laptop but wrong on the phone." Describe the ladder: validate conversion in FP32, then FP16, then quantized-on-host, then on-device, then the full pipeline; use layer-wise comparison at the first rung that diverges; check pre-processing and tokenizer parity; and test on a second device/backend to separate model issues from driver issues.

Shipping and updating models

Shipping turns a working build into a feature that runs on thousands of device models, improves over time, and can be switched off safely. Treat models like code: versioned, tested, staged, observable and reversible.

Analogy

Shipping a model is like a restaurant chain launching a new dish. It is trialled in a few branches first (staged rollout), compared against the old dish on sales and complaints (A/B testing), each kitchen gets a version suited to its equipment (device tiering), head office can pull it from the menu overnight (remote kill switch), and if a branch runs out of an ingredient it serves the old dish rather than nothing (fallback).

Branches are device cohorts, sales and complaints are quality and performance telemetry, kitchen equipment is RAM and NPU tier, and the old dish is the previous model version or a cloud path.

Device tiering

TierTypical signalsWhat to ship
High12 GB+ RAM, recent NPU with INT4 support, allow-listed SoCLargest model (for example 3B W4A16 on NPU), longer context
Mid8 GB RAM, capable GPU or older NPU1B INT4 on GPU/CPU, shorter context, smaller adapters
Low6 GB or less, no usable acceleratorTiny task model, or cloud-only with consent, or feature disabled

Decide the tier with a combination of static facts (RAM from ActivityManager.MemoryInfo.totalMem, SoC model from build properties) and a quick on-device capability probe at first run (load a tiny test graph on the NPU and check it runs), then cache the result.

Versioning and compatibility

  • Model manifest: model id, semantic version, runtime and SDK version required, SoC/NPU architecture, input/output schema, tokenizer version, quantization recipe, checksum, minimum app version.
  • Schema contracts: pre/post-processing code in the app must match the model's expected inputs; version them together or embed metadata in the model file and assert on load.
  • Runtime compatibility: context binaries and compiled caches are tied to NPU generation and SDK version; ship per-SoC variants and invalidate caches on OS or driver updates.
  • Adapters: LoRA adapters are tied to a specific base model version; a base model update invalidates all adapters.

Rollout, A/B testing and remote config

  1. Offline gates Quality evaluation on a fixed set, on-device benchmark on a device matrix (at least one device per tier), and memory/thermal soak all pass.
  2. Internal and beta cohort Enable via remote config for internal users; monitor crashes and ANRs per device.
  3. Staged rollout 1%, 5%, 20%, 50%, 100% with automatic halt thresholds (crash rate, p90 latency, fallback rate, user-visible errors).
  4. A/B test Compare new vs old model on product metrics (task success, acceptance of suggestions, retention) and device metrics (latency, battery), segmented by tier.
  5. Kill switch A remote flag that disables the model or reverts to the previous version without an app update.

Fallback chain

request ──▶ NPU path ok? ──yes──▶ run
               │ no (unsupported / init failed / too hot)
               ▼
            GPU path ok? ──yes──▶ run (maybe smaller model or shorter context)
               │ no
               ▼
            CPU path within budget? ──yes──▶ run (reduced max tokens)
               │ no
               ▼
            cloud allowed (network + user consent + privacy policy)? ──yes──▶ cloud model
               │ no
               ▼
            refuse gracefully with a clear message

What to monitor in the field

  • Load success rate and load time per device model and OS version.
  • Backend actually used (NPU/GPU/CPU/cloud) and fallback rate.
  • Latency percentiles (TTFT, tok/s, end-to-end) by tier.
  • Crashes, ANRs, native crashes inside runtime libraries, low-memory kills during inference.
  • Thermal status at request start, and share of requests degraded by the governor.
  • Quality proxies: user edits or rejections, thumbs up/down, retries; never upload raw private inputs without explicit consent.
  • Model version distribution, to know when old versions can be retired.

Security and integrity

  • Verify checksums and signatures of downloaded models before loading; a malicious model file can be a crash or code-path attack surface on native parsers.
  • Store models in app-private storage; encrypt at rest if the weights are valuable, knowing that a determined attacker with root can still extract them at runtime.
  • Apply safety filters and prompt-injection defences for LLM features even though inference is local.
  • Check model licences before shipping open weights.
Common pitfall Rolling out a new model without a per-device breakdown of metrics. Averages hide the fact that one SoC family falls back to CPU and doubles latency, or that one OS version crashes on load. Always segment by SoC, RAM tier and OS build.
Interview angle "How do you update a 1.5 GB on-device model safely for millions of users?" Answer with: manifest and versioning, eligibility checks, download on unmetered network with resume and checksum (or binary deltas to save bandwidth), atomic swap after a smoke test, keep the previous version, staged rollout with automatic halt metrics, A/B test on product metrics, per-device monitoring and a kill switch.

Target numbers to validate against

Use these as a smoke test. If your measurements are wildly off (for example half the expected decode speed), something is misconfigured: wrong build type, missing KV-cache flags, kernels not enabled, wrong thread count, or the device is already hot. All numbers are approximate, depend heavily on SoC, runtime version, prompt length and thermal state, and will improve with newer releases; re-verify before quoting.

Analogy

Reference numbers are like the typical fuel economy printed for a car model. Your own figure will never match exactly, because driving style and roads differ, but if you get half the rated mileage, you check the tyre pressure and the handbrake before blaming the engine.

The rated mileage is the published benchmark, driving conditions are prompt length and thermal state, and the handbrake is a debug build or a disabled KV cache.

Llama 3.2 on Android with ExecuTorch (CPU, flagship 2024 phone, prompt length 64, KleidiAI enabled)

ConfigDecode tok/sTTFT (s)Prefill tok/sModel size changeMemory change
1B BF16 (baseline)~19~1.1~60--
1B SpinQuant (INT4)~50 (2.6x)~0.3 (-77%)~260 (4.3x)-54%-40%
1B QAT + LoRA (INT4)~46 (2.4x)~0.3 (-76%)~252 (4.2x)-52%-29%
3B BF16 (baseline)~7.6~3.0~21--
3B SpinQuant (INT4)~20 (2.6x)~0.7 (-76%)~90 (4.2x)-60%-50%
3B QAT + LoRA (INT4)~18.5 (2.4x)~0.7 (-76%)~89 (4.2x)-59%-45%

Another reference point: quantized Llama 3.2 1B (INT4 weights, 8-bit dynamic activations) on a flagship 2024 phone

  • Prefill above ~350 tok/s and decode above ~40 tok/s on CPU with KleidiAI kernels.
  • About 2 seconds to respond to a ~600-token input.
  • .pte size about 1.1 GiB quantized vs about 2.3 GiB in BF16 (not a clean 4x because embeddings and some layers stay at 8-bit and scales add overhead).
  • Peak RSS about 1.9 GiB vs about 3.1 GiB for BF16 at a 2048-token maximum sequence length (roughly 40% lower).
  • KleidiAI contributes more than 20% of prefill speed at identical accuracy.

Rule-of-thumb ranges for phone-class flagships (approximate)

ModelWeights (INT4)CPU prefill tok/sCPU decode tok/sNPU prefill tok/sNPU decode tok/sPeak RSS (2k ctx)
~0.5B~0.3-0.5 GB300-70060-1001000-250060-1200.6-1 GB
~1B~0.7-1.1 GB150-40030-55800-200030-701.2-2 GB
~3B~1.7-2.3 GB60-15012-22300-100015-302.5-3.5 GB
~7-8B~3.8-4.8 GB20-605-10150-6008-185-6.5 GB

Classic models (INT8, flagship phone, approximate per-inference latency)

ModelCPU (4 threads)GPU delegateNPU
MobileNet-class classifier (224 px)3-8 ms2-5 msunder 1 ms
ResNet-50 (224 px)20-50 ms8-15 ms1-3 ms
Small YOLO-style detector (640 px)40-120 ms15-30 ms3-10 ms
Small speech encoder (Whisper tiny/base class), 30 s audio0.5-2 s0.3-1 s0.1-0.5 s

Sanity rules

  • Decode ceiling: tok/s cannot exceed effective bandwidth / bytes read per token. With ~50 GB/s effective and ~1 GB of INT4 weights, about 50 tok/s is the ceiling for a 1B model.
  • Prefill vs decode: prefill tok/s should be several times (5-20x) decode tok/s; if they are similar, prefill is not batching tokens (for example the KV-cache or SDPA path is disabled).
  • Quantization gain: moving BF16 to INT4 should give about 2.4-3x decode (bandwidth ratio minus overheads); much less suggests dequantization overhead or fallback kernels.
  • Threads: decode usually peaks at the number of performance cores; more threads can reduce speed.
  • Sustained: expect 60-85% of peak after 10 minutes on a phone in free air; much lower means aggressive throttling or a hot environment.
Common pitfall Comparing your numbers with published ones measured at a different prompt length, context size, thread count or device temperature. Always match the conditions before concluding something is wrong, and publish your own conditions with every number.
Interview angle Interviewers often ask for an estimate: "What tok/s would you expect for a 3B INT4 model on a phone?" Do the bandwidth calculation out loud (bytes per token, effective bandwidth, overheads), give a range, and mention that prefill is compute-bound and much faster, especially on the NPU.

A 24-week learning plan

At 10-15 hours per week, this sequence takes a systems or application engineer to the point of having shipped-quality, measured deployments across CPU, GPU and NPU. Each project depends on the previous project's harness and understanding, so keep the order.

Analogy

The plan is like a marathon training block. Early weeks build a base (setup and the first run), middle weeks add hard specific sessions (quantization and NPU work), a taper consolidates, and every week has a logbook. Skipping the long runs because they hurt is how people fail on race day.

The base is the environment and Project 1, the hard sessions are Projects 2-4, the race is the system project, and the logbook is your committed measurement data.

WeeksFocusMilestones
0Device and toolchainComplete the environment checklist; adb is solid; ExecuTorch builds for host; llama.cpp running on the phone as a fallback baseline. Read a platform AI service architecture overview.
1-3Project 1: benchmark harnessWeek 1 model on device; week 2 harness scripted; week 3 sustained-load thermal study and first write-up.
4-7Project 2: quantization and error analysisWeeks 4-5 sweep; weeks 5-6 layer-wise diff tool; week 7 mixed-precision recipe validated. Rent a GPU only for the QAT comparison.
8-12Project 3: NPU deploymentWeeks 8-9 toolchain friction (normal); week 10 NPU inference; week 11 JNI integration; week 12 three-way comparison with power.
13-16Project 4: KV cacheWeek 13 profiling; weeks 14-15 compression study; week 16 memory-pressure ceiling on a real device.
17-22Project 5: system projectWeek 17 choose a design and commit; weeks 17-21 build; week 22 demo and architecture write-up.
23-24ConsolidateTurn each project's data into concise stories; practise the interview questions using your own measurements as evidence; optionally start a second vendor toolchain.

Theory to study alongside (about 20-25 hours total)

Transformer internals (~8 h)

Self-attention and why only K and V are cached; cross-attention; multi-head vs grouped-query vs multi-query and their KV footprints; prefill (compute-bound) vs decode (bandwidth-bound); RoPE and context extension.

Numerics and quantization (~8 h)

FP32/FP16/BF16 layouts and range; fixed-point, scale and zero-point; symmetric vs asymmetric; per-tensor, per-channel, per-group; PTQ vs QAT; outliers; rotation methods at the intuition level.

Efficiency techniques (~5 h)

LoRA/QLoRA; structured vs unstructured pruning (and why unstructured rarely helps on real hardware); distillation; speculative decoding; operator fusion and graph optimisation.

Hardware mapping (~4 h)

NPU block architecture and why it wants static shapes and static quantization; memory bandwidth as the decode ceiling; when GPU beats NPU; heterogeneous scheduling; model splitting for NPU memory limits.

Slip discipline

  • If you fall behind, cut scope inside a project, never skip a project. A shallow NPU project still teaches the toolchain; skipping it leaves the biggest gap open.
  • If the NPU project stalls completely on hardware or SDK access, switch vendor or move to Project 4 and return later.
  • Timebox setup and toolchain debugging; when stuck for more than a day, get a baseline working with a simpler tool and come back.
Tip Keep a running engineering log: every error message, SDK version, workaround and measurement, dated. It becomes your write-ups, your interview stories and your debugging memory.
Interview angle "Tell me about a deployment you did end to end" is best answered with one project from this plan: the constraint, the measurement, the surprise, the fix and the final numbers. Concrete numbers you measured yourself (tok/s, peak RSS, sustained ratio, SQNR of the worst layer) are far more persuasive than general knowledge.

Failure modes checklist

These are the ways edge AI work most often goes wrong, both while learning and in production. Read them before you start and again before you ship.

Analogy

This checklist is like a pilot's pre-flight checklist. Experienced pilots still use it, not because they do not know how to fly, but because the failures it prevents are simple, common and catastrophic when missed.

Each line is a known way a deployment "crashes": a wrong number, a misleading benchmark, an app killed in the field.

While learning

Failure modeSymptomPrevention
Toolchain rabbit holeWeekends lost to build errors before any model runsTimebox setup; get llama.cpp running first; one environment per toolchain
Measuring on an emulator or laptopNumbers that do not survive contact with a phoneEvery performance claim from physical silicon
Cold-run-only benchmarkingOne impressive number no user experiencesWarm-up, percentiles, sustained curve; report sustained next to peak
Chasing model size80 GB GPU bills and slow iteration on 7B modelsLearn on 1B-class models; scale up once at the end
Drifting into researchReading papers instead of deployingCap theory time; stay on deployment, numerics and systems
Not documentingFinished projects with no evidenceCommit raw data and write a short findings note per project

Before shipping

  • Converted FP32 model matches the original within tolerance.
  • Quantized model evaluated on a fixed set including hard and edge cases, not only a few prompts.
  • Per-op placement confirmed: no unexpected CPU fallback partitions on the target accelerator.
  • Static shapes and max context chosen deliberately; memory budget written down (weights + KV + activations + runtime + app).
  • Release build, correct flags (optimized kernels, KV cache, SDPA), correct thread count.
  • Load time measured cold and warm; compiled graph or context cached.
  • Inference off the main thread; cancellable; session reused; warm-up done off the critical path.
  • Model assets uncompressed and memory-mapped, or downloaded with checksum and atomic install.
  • Behaviour tested under memory pressure (other apps open, onTrimMemory), with no low-memory kills of the foreground app.
  • Ten- to thirty-minute thermal soak done; thermal headroom policy in place.
  • Energy per request measured against an idle baseline.
  • Tokenizer and chat template parity verified; greedy outputs compared host vs device.
  • Hexagon/NPU libraries match the SoC; SDK versions of build and runtime match.
  • Fallback chain implemented and tested by forcing each failure.
  • Telemetry segmented by SoC, RAM tier and OS build; staged rollout thresholds and kill switch ready.
  • Model licence checked; safety filtering in place for generative features; downloaded models verified.
Common pitfall Testing only on the one flagship on your desk. The long tail of mid-range devices, older OS versions and vendor driver differences is where most field failures come from. Test at least one device per tier before rollout.
Interview angle "What are the top risks when shipping an on-device model?" Give a prioritised list: memory (kills, OOM), thermal (sustained performance), accelerator coverage (fallback and fragmentation), accuracy drift after quantization or on specific devices, model delivery and versioning, and observability. For each, one prevention and one detection mechanism.

Quick revision

  • The deployment loop is choose, export, convert, quantize, compile for target, integrate, benchmark on device, monitor; stages 3-7 iterate many times.
  • Export captures a static graph (torch.export, ONNX); conversion produces a runtime format (.pte, .tflite, .gguf, .onnx/.ort, .mlpackage, QNN context binary).
  • NNAPI is deprecated from Android 15; new work uses LiteRT delegates/accelerators, ExecuTorch backends or ONNX Runtime execution providers. Confirm current LiteRT CompiledModel / Interpreter package names in official docs; do not invent flags.
  • ExecuTorch swaps backends by changing the partitioner (XNNPACK, Vulkan, QNN, MediaTek, Core ML) while keeping one export flow.
  • llama.cpp with GGUF is the fastest way to a working on-device LLM baseline; Q4_0 is repacked for fast Arm kernels, Q4_K_M is a quality-per-byte default.
  • For ExecuTorch LLM export, use_kv_cache and use_sdpa_with_kv_cache are essential; 8da4w means 8-bit dynamic activations, 4-bit weights.
  • KleidiAI kernels in XNNPACK add more than 20% prefill speed on Arm CPUs; always build Release.
  • Validate numerics after every conversion: FP32 differences should be around 1e-5 to 1e-4.
  • Never benchmark on an emulator: no realistic memory system, DVFS, thermal model or NPU.
  • Report load time (cold and warm), TTFT, prefill tok/s, decode tok/s, per-token p50/p90/p99, peak RSS, file size, energy per request and sustained/peak ratio.
  • Always state prompt length and generated length with LLM numbers; prefill and decode differ by 5-20x.
  • A fair protocol fixes the environment, cools down between runs, discards warm-up runs, repeats, and records device, build and runtime versions.
  • Peak RSS comes from VmHWM in /proc/<pid>/status; app memory breakdown from dumpsys meminfo.
  • Energy = integral of (power minus idle power) over time; use power rails via Perfetto, fuel gauge sampling, batterystats, or an external monitor.
  • A 10-30 minute thermal soak reveals throttling; phones typically sustain 60-85% of peak.
  • Quantization quality needs both perplexity and task-level evaluation on a fixed set including hard cases.
  • Layer-wise error analysis compares each intermediate tensor to an FP32 reference using cosine similarity, MAE, max error and SQNR.
  • SQNR = 10 log10(signal power / error power); each bit adds about 6 dB; low SQNR on the residual stream predicts quality loss.
  • Single-layer sensitivity (quantize one layer at a time) separates fragile layers from layers receiving bad inputs.
  • Typical sensitive parts: outlier activation channels, MLP down projection, LM head, first/last blocks, norms and softmax.
  • Mixed precision promotes only the most sensitive layers to 8-bit or FP16 and prices each promotion in size and latency.
  • Host fake-quant matching FP32 but device diverging means a device implementation issue (rounding, accumulators, FP16 overflow), not the scheme.
  • NPUs need static quantization: ranges fixed at compile time from calibration; LLMs on Hexagon commonly use W4A16 or W8A16.
  • NPUs need static shapes; export separate prefill (chunked) and decode (one token) graphs that share weights.
  • One unsupported op mid-graph creates partitions and CPU round trips that can make the NPU slower than the CPU.
  • Context binaries are finalized graphs for one Hexagon architecture and SDK version; loading them avoids on-device compilation.
  • Hexagon library folders (v73, v75, v79) must match the SoC; set ADSP_LIBRARY_PATH for DSP-side libraries.
  • Large models are split into several context binaries because of NPU session memory limits.
  • KV bytes = 2 x layers x KV heads x head dim x tokens x bytes; Llama-3.2-1B uses 32 KiB per token in FP16 (verified: 2 × 16 × 8 × 64 × 2).
  • At long context the KV cache exceeds the weights (about 22-35k tokens for 1B INT4, about 16-20k for 3B, under 8k for models without GQA).
  • Decode speed falls as context grows because each step reads the weights plus the whole cache: tok/s ≤ BW / (W + KV(t)).
  • INT8 weights are usually near-lossless; INT4 is the decode lever but needs a task eval, plus higher precision on embeddings and the LM head.
  • Quantize keys per channel (outlier channels) and values per token; INT8 KV is near-lossless.
  • Sliding-window attention bounds memory; keep attention-sink tokens to avoid collapse on long streams.
  • Paged KV: contiguous waste is (T_max - T_actual) x KV per token; paged waste is at most one block. Prefix caching saves about P / (P + U) of prefill if the prefix is byte-identical.
  • Speculative decoding: E[tokens per pass] = (1 - alpha^(k+1)) / (1 - alpha); speed-up ≈ E / (1 + k c). Helps on-device when decode is bandwidth-bound and the draft is accurate.
  • lmkd kills by oom_score_adj under memory pressure (PSI); anonymous memory is not reclaimable, mmapped clean file pages are.
  • Store model assets uncompressed (noCompress) so they can be memory-mapped; downloaded models need resume, checksum, atomic install and a smoke test.
  • Create sessions once, off the main thread, warm them up, reuse buffers, make generation cancellable and respond to onTrimMemory.
  • ONNX Runtime's QNN EP needs a QDQ model; disable CPU fallback during development and enable EP context caching. Provider option keys (backend_path, HTP performance mode) are version-specific: check the current ORT QNN EP docs.
  • QAIRT converter and Genie binary/API names move between SDK releases. Copy flags, JSON keys and C symbols from the samples for the version you link; do not invent them.
  • Profile at three levels: system (Perfetto), runtime (per-op profiles, AI Hub, QNN profiler, ETDump), native (simpleperf).
  • Accuracy drift is bisected: original, converted FP32, FP16, quantized host, quantized device, full app pipeline; pre-processing bugs are the most common cause.
  • Shipping needs device tiering, versioned manifests, staged rollout with halt thresholds, A/B tests, a kill switch and a fallback chain (NPU, GPU, CPU, cloud, refuse).
  • Monitor load success, backend used, fallback rate, latency by tier, low-memory kills and thermal state, segmented by SoC, RAM and OS build.

Glossary

Activation quantization
Representing the intermediate outputs of layers with low-bit integers; harder than weight quantization because ranges depend on inputs and contain outliers.
AI Hub
Qualcomm's hosted service to compile, quantize, profile and run models on real Snapdragon devices, with a model zoo of export scripts.
AIMET
Qualcomm's open-source toolkit for quantization and compression, including cross-layer equalisation, AdaRound, QuantAnalyzer and QAT.
Attention sink
The first few tokens of a sequence, which attract a large share of attention; keeping them in a sliding-window cache keeps long generations stable.
Block table
The mapping from a sequence's logical KV-token ranges to physical KV blocks in a paged cache, analogous to a page table.
CompiledModel
LiteRT Next's compiled-model API that binds a .tflite (or a vendor-compiled artifact) to an accelerator. Class names and options are release-specific; confirm in current LiteRT docs.
AWQ
Activation-aware weight quantization: a PTQ method that protects the weight channels most important to activations by scaling before quantizing.
batterystats
Android's battery accounting service; dumpsys batterystats reports estimated energy use per app over a measurement window.
BF16
Brain float 16: 16-bit floating point with the same exponent range as FP32 but fewer mantissa bits; robust to overflow.
Calibration dataset
A small, representative set of inputs run through a model to record activation ranges for static quantization.
Context binary
A serialized, fully compiled QNN graph for a specific Hexagon architecture that loads without on-device graph compilation.
CPU fallback
Execution of operators the accelerator cannot run on the CPU instead, creating partitions and data transfers.
Decode phase
The token-by-token generation phase of an LLM; memory-bandwidth-bound because all weights and the KV cache are read for each token.
Draft acceptance rate
The per-token probability that a speculative draft token is accepted by the target model; it sets expected tokens per verification pass.
Delegate
A LiteRT plug-in that takes over supported parts of a graph and runs them on a GPU, NPU or DSP.
Dynamic quantization
Quantizing activations at runtime with scales computed from each tensor's actual values; common on CPUs, not possible on most NPUs.
EP context cache
An ONNX Runtime feature that saves a compiled execution-provider graph (for example a QNN context) so later sessions skip compilation.
ETDump and ETRecord
ExecuTorch developer-tool artifacts: runtime profiling and debug data (ETDump) and export-time graph metadata (ETRecord) used together by the Inspector.
Execution provider
An ONNX Runtime backend (CPU, QNN, Core ML, XNNPACK and others) that executes the subgraphs assigned to it.
ExecuTorch
PyTorch's on-device runtime that executes .pte programs produced from torch.export, with pluggable hardware backends.
Fake quantization
Simulating quantization in floating point (quantize then dequantize) to measure or train for its effect on a host.
Genie
Qualcomm's generative AI runtime in QAIRT that wraps tokenizer, QNN execution, KV-cache management and sampling behind a dialog API. Confirm headers, JSON keys and CLI names against the SDK samples for your release.
GGUF
llama.cpp's single-file model format containing tensors, quantization types, tokenizer and metadata.
GPTQ
A post-training weight quantization method that quantizes weights column by column while compensating the error using second-order information from calibration data.
GQA (grouped-query attention)
Attention where several query heads share each key/value head, shrinking the KV cache in proportion.
Graph partitioning
Splitting a model graph into subgraphs for different processors; each boundary adds transfer and synchronisation cost.
Group-wise quantization
Using one scale (and zero-point) per small block of weights, for example 32 or 128, to follow local ranges more closely.
heapprofd
Perfetto's native heap profiler that attributes allocations to call stacks in a running process.
Hexagon HTP
The Hexagon Tensor Processor, the matrix/tensor accelerator in Snapdragon NPUs, reached through QNN/QAIRT.
KleidiAI
Arm's library of optimized low-bit matrix-multiply micro-kernels, integrated into XNNPACK and other runtimes.
KV cache
Stored key and value vectors for all previous tokens in every layer, so attention does not recompute them each step.
LiteRT
Google's on-device runtime, formerly TensorFlow Lite, executing .tflite models with CPU, GPU and NPU accelerators.
llama.cpp
A C/C++ LLM inference engine with optimized CPU kernels and GPU backends, using the GGUF format.
lmkd
Android's userspace low-memory killer daemon, which kills processes by priority when memory pressure is high.
LoRA adapter
A small set of low-rank weight updates that specialise a shared base model for a task and can be swapped at runtime.
Memory mapping (mmap)
Mapping a file into the address space so pages load lazily and can be reclaimed by the kernel, avoiding a full copy into the heap.
Mixed precision
Using different numeric precisions for different layers, typically keeping sensitive layers at higher precision.
NNAPI
Android Neural Networks API, the older accelerator abstraction, deprecated from Android 15 in favour of vendor delegates and backends.
ODPM
On-Device Power Monitor: hardware energy counters for power rails, readable via Perfetto on supported devices.
Paged KV cache
A KV cache stored in fixed-size blocks mapped by a block table, avoiding contiguous reallocations and fragmentation.
Perfetto
Android's system-wide tracing tool for scheduling, frequencies, thermal, power rails, memory and custom trace slices.
Prefill phase
Processing all prompt tokens in parallel to fill the KV cache; compute-bound and the main contributor to TTFT.
Prefix caching
Reusing a precomputed KV cache for a shared prompt prefix, such as a system prompt, to reduce TTFT.
PSI
Pressure stall information: kernel metrics showing how much time tasks stall waiting for memory, CPU or IO.
PTQ
Post-training quantization: quantizing a trained model using calibration data, without retraining.
QAIRT
Qualcomm AI Runtime SDK, which includes the QNN libraries, converters, quantizers, HTP backend, profiling and Genie.
QAT
Quantization-aware training: fine-tuning with simulated quantization so the model learns to tolerate rounding.
QDQ format
An ONNX representation of quantization using explicit QuantizeLinear/DequantizeLinear nodes that backends fuse into integer kernels.
RSS and PSS
Resident set size (all resident pages of a process) and proportional set size (shared pages divided among sharers).
simpleperf
Android's sampling CPU profiler for native code, with call graphs and hardware counters.
Sliding-window attention
Attention restricted to the most recent W tokens, bounding KV-cache memory.
SpinQuant
A quantization method that applies learned rotations to weights and activations to spread outliers before low-bit quantization.
Speculative decoding
A draft model (or extra heads / n-grams) proposes tokens that the target verifies in one pass. Expected tokens per pass = (1 − αk+1) / (1 − α). Output distribution is unchanged.
SQNR
Signal-to-quantization-noise ratio in decibels, comparing reference tensor energy with quantization error energy.
Static quantization
Quantization with activation scales fixed ahead of time from calibration; required by most NPUs.
Static shapes
Tensor dimensions fixed at export/compile time, allowing NPU compilers to plan memory and tiling.
Sustained-to-peak ratio
Throughput after a long run divided by initial throughput; captures the effect of thermal throttling.
Thermal headroom
Android's forecast of how close the device is to severe throttling, from PowerManager.getThermalHeadroom().
Thermal soak
A long continuous workload (10-30 minutes) used to measure throttling and sustained performance.
torch.export
PyTorch's ahead-of-time graph capture producing an ExportedProgram with no Python control flow, used by ExecuTorch and ai-edge-torch.
TTFT
Time to first token: from prompt submission to the first generated token appearing.
VTCM
Vector tightly coupled memory: fast on-chip memory in the Hexagon NPU; working sets that exceed it spill to DRAM.
W4A16
A scheme with 4-bit weights and 16-bit activations, common for LLMs on NPUs.
XNNPACK
A highly optimized CPU inference library for Arm and x86 used as the CPU backend by LiteRT, ExecuTorch and ONNX Runtime.

Interview questions

Fundamentals

What are the stages of deploying a model to an edge device?

Eight stages: (1) choose or train a model that fits the latency and memory budget; (2) export it as a static graph (torch.export, ONNX); (3) convert to a runtime format (.pte, .tflite, .gguf, .onnx, Core ML, QNN); (4) quantize with calibration data or QAT; (5) compile for the target accelerator (partitioning, context binaries, caches); (6) integrate into the app (threading, packaging, fallback); (7) benchmark on real devices (latency percentiles, memory, energy, thermal); (8) monitor in the field with staged rollout and rollback. Stages 3-7 iterate: an unsupported op or a latency miss sends you back.

What is the difference between exporting and converting a model?

Exporting captures the model's computation as a static, framework-level graph with fixed operators and (usually) fixed shapes, removing Python control flow: for example a torch.export ExportedProgram or an ONNX file. Converting transforms that graph into the format and operator set of a specific runtime: a .pte for ExecuTorch, a .tflite for LiteRT, a QNN model or context binary, a Core ML package. Export problems are about capturability (dynamic control flow, data-dependent shapes); conversion problems are about operator coverage and layout.

Why is NNAPI no longer recommended, and what replaced it?

NNAPI offered a common Android interface to accelerators, but vendors implemented its drivers inconsistently, the op set was a lowest common denominator, and behaviour and performance varied across devices, causing fragmentation. From Android 15 it is deprecated for new work. The replacement is vendor-specific delegates and backends integrated directly into frameworks: LiteRT GPU and NPU accelerators, ExecuTorch backends (QNN, MediaTek, Vulkan), and ONNX Runtime execution providers such as QNN.

What is ExecuTorch and what is a .pte file?

ExecuTorch is PyTorch's on-device inference runtime. You capture a model with torch.export, lower it to an "edge" dialect, hand supported subgraphs to backends through partitioners (XNNPACK for CPU, Vulkan, QNN, MediaTek, Core ML, etc.), and serialize the result into a .pte program. The .pte contains the execution plan, constant weights (or references to external weight files) and delegate blobs. The runtime core is small and portable C++, with Java/Kotlin and Swift bindings and an LLM runner.

What is LiteRT, and how do you get a PyTorch model into it?

LiteRT is the new name for TensorFlow Lite: Google's runtime for .tflite FlatBuffer models with XNNPACK on CPU, a GPU delegate and NPU accelerators. NNAPI is deprecated from Android 15; new Android work should use LiteRT delegates or vendor backends, not NNAPI. For PyTorch models you use ai-edge-torch, which runs torch.export internally and emits a .tflite. The long-lived app API is still an Interpreter-style session; LiteRT Next also documents a CompiledModel API whose class names and accelerator options you must confirm in the current official docs. TensorFlow/Keras models use the TFLiteConverter with a representative dataset for integer quantization.

What is GGUF, and why start LLM work with llama.cpp?

GGUF is llama.cpp's single-file format holding quantized tensors, tokenizer and metadata. llama.cpp builds in minutes with the NDK, runs most popular architectures on CPU with highly tuned Arm kernels, and has llama-bench for prefill/decode numbers and llama-perplexity for quality. That makes it the fastest way to validate that a use case works on a phone and to obtain a CPU reference baseline before investing in a heavier export path such as ExecuTorch or a vendor NPU flow.

What is a QNN context binary?

It is a serialized, fully prepared QNN graph for a specific Hexagon NPU architecture (and SDK version): the graph has been optimised, tiled and finalized, with weights and quantization parameters embedded. Loading a context binary skips on-device graph compilation, turning multi-second initialization into a fast load. It is not portable across Hexagon generations, so you produce one per target SoC family, via AI Hub, the QAIRT context-binary generator, or a framework backend.

Why can't you trust performance numbers from an emulator or a laptop?

They do not reproduce the device's memory bandwidth and cache hierarchy, DVFS governors, thermal limits and throttling, big.LITTLE scheduling, or the NPU/GPU drivers at all, and they run different (x86) kernels. The error is not a constant factor; it can change which option is faster. Host runs are useful for numerical reference and functional tests only; all performance claims must come from physical silicon.

What are TTFT, prefill throughput and decode throughput?

TTFT (time to first token) is the time from submitting a prompt to the first generated token; it is dominated by prefill. Prefill throughput is prompt tokens divided by prefill time: all prompt tokens are processed in parallel with matrix-matrix operations, so it is compute-bound. Decode throughput is generated tokens per second after the first: one token at a time, reading all weights and the cache each step, so it is memory-bandwidth-bound. Always report them separately with prompt and output lengths.

Why is LLM decode memory-bandwidth-bound?

Each decode step multiplies a single token's activation vector by every weight matrix: a matrix-vector product with about 2 FLOPs per weight read. The arithmetic intensity is so low that the processor waits on memory, not compute. So decode speed is roughly effective bandwidth divided by bytes read per token (weights plus KV cache). This is why weight quantization speeds up decode almost in proportion to the bytes saved, and why NPUs help less for decode than for prefill.

Why do benchmarks need warm-up runs?

The first runs include one-time costs: page faults as mmapped weights are touched, cache and TLB warming, GPU shader compilation, NPU graph finalization, kernel auto-selection, memory pool allocation and CPU frequency ramp-up. Including them mixes initialization with steady-state latency. Measure and report load and first-run time separately, discard a few warm-up runs, then measure steady state.

Why report p50, p90 and p99 rather than the mean?

Latency distributions on phones are skewed: scheduler preemption, frequency changes, garbage collection, thermal events and background work create long tails. The mean hides them and is distorted by outliers. p50 describes the typical experience; p90/p99 describe what users notice as stutter or lag. For LLMs, per-token inter-token latency percentiles reveal stutters that an average tok/s hides.

What does "8da4w" mean in ExecuTorch LLM export?

8-bit dynamic activations, 4-bit weights. Linear layer weights are stored as 4-bit integers with group-wise scales (for example one per 128 weights); at runtime, each activation tensor is quantized to 8-bit with a scale computed from its actual values, and integer matmul kernels run the product. It gives most of the INT4 size and bandwidth benefit while keeping activation error low, and it suits CPU backends like XNNPACK with KleidiAI kernels.

What is a calibration dataset?

A small representative set of inputs (typically 100-500 samples, or a few hundred text sequences) passed through the model during static quantization so observers can record activation ranges and choose scales. It must resemble production data, including edge cases; otherwise ranges are wrong, values clip or lose resolution and accuracy drops. For LLMs, the calibration text should match the target domain and languages.

What is the difference between static and dynamic quantization?

Dynamic quantization computes activation scales at runtime from the current tensor's range: accurate and needs no calibration, but costs a reduction per tensor per step and requires flexible hardware (CPUs). Static quantization fixes activation scales ahead of time using calibration data: no runtime overhead and compatible with fixed-function integer pipelines (NPUs, DSPs), but values outside the calibrated range clip. Weights are always effectively static.

What is a delegate or execution provider?

It is a plug-in backend that claims the parts of a model graph it supports and runs them on specific hardware (GPU, NPU, DSP, optimized CPU library). LiteRT calls them delegates or accelerators, ONNX Runtime calls them execution providers, ExecuTorch calls them backends selected via partitioners. Unsupported parts remain on the default CPU path, which creates partitions.

What is CPU fallback and why does it hurt performance?

When an accelerator cannot run an operator (unsupported type, shape, data type or attribute), the runtime executes that operator on the CPU. The graph is split into partitions and at each boundary data must be transferred, possibly re-laid out and requantized, and the processors synchronise. With several partitions the accelerator idles while waiting, and total time can exceed a CPU-only run. Aim for full delegation and check placement in profiles.

What does memory-mapping the model weights give you?

With mmap, the model file is mapped into the address space instead of copied into the heap. Loading becomes almost instant (pages are read on first access), clean file-backed pages can be reclaimed by the kernel under pressure instead of forcing a kill, and multiple processes mapping the same file share physical pages. The downsides: evicted pages must be re-read, causing latency spikes, and runtimes that repack weights create anonymous copies that lose these benefits.

Why must model assets be stored uncompressed in an APK?

Compressed APK entries cannot be memory-mapped directly; the runtime must decompress the whole file into memory, doubling peak memory at load and slowing startup. Marking model extensions as noCompress in Gradle (androidResources { noCompress += "tflite" }) keeps them stored and page-aligned so AssetFileDescriptor plus FileChannel.map can mmap them.

Why should inference never run on the UI thread, and why reuse the session?

Inference can take tens of milliseconds to seconds; on the main thread it blocks rendering and input, causing jank and ANRs. Session creation is also expensive (loading, delegate init, graph compilation), so creating it per request multiplies latency and memory churn. Create the session once on a background thread, warm it up, keep it in an application-scoped owner, reuse input/output buffers, and serialise calls on a dedicated executor.

What is the KV cache and why is it needed?

In self-attention, each new token attends to the keys and values of all previous tokens. Without caching, each step would recompute keys and values for the entire sequence, making generation quadratic. The KV cache stores K and V for every past token in every layer, so each decode step only computes K and V for the new token and appends them. It trades memory (growing linearly with context) for compute.

What is thermal throttling, and what is a thermal soak test?

Phones are passively cooled; sustained power heats the SoC and skin, and the thermal framework lowers CPU/GPU/NPU frequencies to stay within temperature limits, reducing throughput. A thermal soak test runs the workload continuously for 10-30 minutes while logging throughput, temperatures, frequencies and power, showing when throttling starts and the sustained-to-peak ratio. It is the measurement that predicts real product behaviour.

How do you measure peak memory of an on-device inference process?

For a native process, read VmHWM (peak resident set) and VmRSS from /proc/<pid>/status, polling during the run or printing at exit. For an app, use dumpsys meminfo <package> for PSS broken down by native heap, graphics and file mappings, and Perfetto's heapprofd to attribute native allocations to call stacks. Distinguish file-backed (mmapped weights) from anonymous memory, because only the latter is non-reclaimable.

How can you measure energy consumption of inference on Android?

Options in increasing accuracy: batterystats (model-based per-UID estimates), sampling the fuel gauge (current_now and voltage_now) and integrating power over time, on-device power rails (ODPM) recorded via Perfetto's android.power data source for per-subsystem energy, and an external power monitor. Always subtract an idle baseline with the same screen and radio state, and report energy per inference or per 100 tokens.

What is device tiering?

Grouping devices by capability (RAM, SoC and NPU generation, GPU, OS version) and shipping different model variants or settings per tier: for example a 3B NPU model on high-end phones, a 1B CPU/GPU model on mid-range, and cloud or no feature on low-end. Tier decisions combine static facts with a first-run capability probe, and are controlled remotely so they can be adjusted after launch.

Should you bundle a model in the APK or download it?

Bundle small models (tens of MB) that must work immediately and offline: simplest, always available, but updates require an app update and increase install size. Download larger models (asset packs, device-targeted AI packs or your own CDN): smaller install, per-device variants and independent updates, at the cost of first-use delay and the need for resume, integrity checks, storage checks, versioning and fallback while not yet downloaded.

What is a chat template, and why does it matter on device?

Instruction-tuned LLMs were trained with a specific prompt format of special tokens marking system, user and assistant turns (for Llama 3.x, header and end-of-turn tokens). Sending raw text without the template or with the wrong special token ids produces rambling, off-task or empty outputs and missing stop conditions. On-device runners often do not apply templates automatically, so it is a frequent cause of "the model is worse on the phone".

What is mixed-precision quantization?

Using different bit-widths in different parts of the model: most layers at the target low precision (for example INT4 weights) and a few sensitive layers (often the LM head, some down projections, first or last blocks, norms) at INT8 or FP16. It recovers most of the quality lost by uniform low-bit quantization at a small size and latency cost, and should be derived from measured layer sensitivity.

What is lmkd and why does it matter for on-device LLMs?

lmkd is Android's low-memory killer daemon. It watches memory pressure (PSI) and kills processes in order of oom_score_adj: cached apps first, then services, then perceptible and finally foreground apps. A large model plus KV cache can push the system into killing other apps (music, launcher) or, under extreme pressure, the foreground app itself. Memory budgets for LLM features must leave room for the rest of the system.

What are SQNR and cosine similarity used for in quantization work?

They measure how close a quantized tensor is to its FP32 reference. SQNR (in dB) is the ratio of signal energy to error energy; each extra bit of uniform quantization adds about 6 dB. Cosine similarity measures directional agreement, insensitive to uniform scaling. Computed per layer, they locate where quantization error is introduced and how it accumulates, guiding mixed-precision decisions.

Going deeper

Walk through exporting Llama-3.2-1B to ExecuTorch for an Android CPU. Which flags matter most?

Get consolidated.00.pth, params.json and tokenizer.model. Run the LLM exporter with the Llama 3.2 model class, enabling the KV cache (use_kv_cache) and the fused SDPA-with-KV-cache op (use_sdpa_with_kv_cache), XNNPACK backend, 8da4w quantization with group size 128 (or 32 for better quality), 4-bit embedding quantization with group 32, a chosen max_seq_length, and metadata with the correct BOS/EOS ids (128000; 128001/128009). Push the .pte, tokenizer and a Release-built runner with KleidiAI enabled. The KV-cache and SDPA flags are the most important; without them decode recomputes attention over the full sequence and speed collapses. Wrong EOS ids make generation never stop.

What is KleidiAI and what would you check if prefill is 20% below reference numbers?

KleidiAI is Arm's set of optimized low-bit matmul micro-kernels using dotprod/i8mm (and SME where available), integrated into XNNPACK. It improves prefill by over 20% at identical accuracy. If prefill is low: confirm the build has EXECUTORCH_XNNPACK_ENABLE_KLEIDI on and is Release; confirm the quantization scheme is one the kernels support (for example 4-bit group-wise weights with 8-bit dynamic activations); check thread count and that threads run on performance cores; check the device is not already thermally throttled; and compare prompt lengths with the reference.

How do you choose between GGUF Q4_0, Q4_K_M and Q8_0, and how many threads to use?

Q8_0 is near-lossless but twice the bytes of 4-bit, so decode is roughly half as fast. Q4_K_M uses super-blocks with some 6-bit tensors, giving better quality per byte and a common default. Q4_0 is simpler; on Arm CPUs with i8mm/dotprod, llama.cpp repacks it into interleaved layouts for very fast kernels, so it is often the fastest phone option with slightly lower quality (mitigated with an importance matrix). Measure perplexity and your task. For threads, start with the number of performance cores; using all cores often slows decode because little cores become stragglers, and prefill may benefit from a different count than decode.

Explain the ONNX static quantization flow and the QDQ format.

Pre-process the model (shape inference, constant folding), implement a CalibrationDataReader yielding representative inputs, and call quantize_static with a calibration method (MinMax, Entropy, Percentile), per-channel weights, activation and weight types, and QuantFormat.QDQ. QDQ inserts QuantizeLinear/DequantizeLinear pairs around tensors; the graph remains valid in float, and execution providers pattern-match DQ-op-Q sequences into integer kernels. For the QNN EP, use ORT's QNN helpers to produce the activation types the HTP expects (8- or 16-bit) and ensure all ops are supported.

How do you use ONNX Runtime with the QNN execution provider on Android?

Use the QNN-enabled Android package, create SessionOptions and add the QNN EP. The Java helper name and option keys are version-specific: confirm them in the current ONNX Runtime QNN EP docs. Common documented keys include backend_path (HTP library) and an HTP performance mode. Feed a QDQ-quantized model. During development disable CPU EP fallback so unsupported nodes error instead of silently falling back. Enable EP context caching so the compiled QNN graph is saved and reused on later launches. Profile to confirm all nodes are on QNN, and re-enable CPU fallback in production for robustness.

When would you use MediaPipe LLM Inference vs ExecuTorch vs llama.cpp in an app?

MediaPipe LLM Inference (and LiteRT-LM) when a supported model family fits the product and you want minimal code, GPU support and Google-maintained bundles. ExecuTorch when you need arbitrary PyTorch models, multiple backends including vendor NPUs, custom quantization, and one flow for Android and iOS. llama.cpp for quick prototyping, broad architecture support on CPU, GGUF distribution and fine control in C++, accepting weaker NPU support. Many teams prototype with llama.cpp and ship with ExecuTorch or a vendor path.

Design a benchmark harness for on-device LLMs. What must it record?

A host-side driver (adb) plus on-device timing in the runner. Per run: model load time cold and warm, TTFT, prefill tok/s at several prompt lengths, decode tok/s, inter-token latency percentiles, peak RSS (VmHWM) and PSS, file size, energy per 100 tokens with idle baseline, temperatures and CPU frequencies at start and end. Protocol: fixed environment, cool-down to a temperature threshold between configs, warm-up runs discarded, repetitions with p50/p90/p99 and stdev, a sustained soak mode. Metadata: device, SoC, build fingerprint, runtime and SDK versions, model hash, quantization config, threads, ambient temperature. Output a CSV and a summary, one command per full run.

How do you correctly time TTFT and per-token latency in code?

Use a monotonic clock (std::chrono::steady_clock, SystemClock.elapsedRealtimeNanos), never wall-clock time. TTFT spans from prompt submission (including tokenization if user-visible) to the first token being sampled; prefill rate is prompt tokens over prefill time. For decode, record a timestamp after each sampled token and compute inter-token intervals; decode rate is (generated minus 1) over decode time. Exclude detokenization or UI rendering only if you report them separately, and do not print per token to the console during timing, which adds I/O overhead.

Why is energy per inference a better metric than power, and what is race to idle?

Energy (power times time) is what drains the battery. A backend drawing more power but finishing much faster can use less energy per request than a slow, low-power one. Race to idle is the strategy of completing work quickly at high performance and letting the hardware return to deep idle, often more efficient than running slowly, as long as the high-power bursts do not trigger throttling. For continuous workloads (30 fps camera, long generation) the average power and thermal steady state matter more.

How do you run and interpret a thermal soak test?

Run back-to-back generations (or inferences) for 10-30 minutes at a fixed ambient temperature, with the device in its realistic state (case on or off noted, screen state fixed). Log throughput per generation, thermal zones, CPU/GPU frequencies, thermal status from thermalservice, and power. Plot throughput and temperature against time. Interpret: time until first throttle step, depth of each drop, the plateau level, the sustained-to-peak ratio, and whether throughput oscillates (governor hunting). Compare backends: NPUs usually sustain better due to lower power.

Describe how you would do layer-wise quantization error analysis.

Run the FP32 model on the host with hooks capturing every layer's output on a fixed input set; run the quantized model (host fake-quant, then device with intermediate-output dumping) on the same inputs. For each tensor, compute cosine similarity, MAE, max absolute error and SQNR; rank layers and plot SQNR across depth. Then quantize one layer at a time to measure individual sensitivity, check activation ranges for outliers, and derive a mixed-precision recipe that promotes only the most sensitive layers. Validate the recipe with quality metrics and the performance harness.

How does quantization group size affect quality and size?

Smaller groups adapt scales to local ranges, reducing error, especially with outliers; larger groups use fewer scales. The overhead is scale bits divided by group size: a 16-bit scale per 32 weights adds 0.5 bits per weight (4.5 effective bits), per 128 adds 0.125. Small groups also add dequantization work and can reduce kernel efficiency. Common choices: 32 for quality (and some NPU backends), 128 for CPU LLM exports; per-channel for INT8.

Why do the embedding table and LM head deserve special treatment?

In small LLMs with large vocabularies they are a big share of parameters: Llama-3.2-1B's 128256 x 2048 embedding is about 263M of 1.24B parameters (about 21%). Quantizing the embedding saves a lot of storage with little quality loss because it is a lookup. The LM head, if not tied, is also large, but its errors land directly on logits, so it is unusually sensitive; it is often kept at 6-8 bits even when the rest is 4-bit. With tied weights, the choice affects both.

Perplexity vs task metrics: why do you need both?

Perplexity measures how well the model predicts held-out text on average: sensitive, cheap and good for comparing configurations. But small perplexity changes can hide collapses on specific skills (arithmetic, code, structured output, non-English) and large ones may not matter for a narrow task. Task-level metrics (accuracy on multiple-choice, exact match, format validity on your product's prompts) measure behaviour. Use perplexity to sweep, task metrics to decide, plus top-1 agreement or KL against the FP32 model.

Compare round-to-nearest, GPTQ, AWQ, SpinQuant and QAT.

Round-to-nearest quantizes each weight independently: fastest, worst at 4-bit. GPTQ quantizes weights column by column and updates remaining weights to compensate using second-order (Hessian) information from calibration data. AWQ identifies weight channels important to large activations and scales them to protect them before quantization. SpinQuant learns rotations applied to weights and activations that spread outliers, enabling low-bit weights and activations with small loss. QAT fine-tunes with simulated quantization (optionally with LoRA to keep it cheap) and recovers the most quality at the highest cost. PTQ methods need minutes to hours; QAT needs a training setup and GPUs.

How do you capture intermediate tensors on the device for comparison?

ExecuTorch: generate an ETRecord at export time, run with ETDump and a debug buffer enabled, and use the devtools Inspector to map on-device outputs back to graph nodes. QNN/QAIRT: qnn-net-run --debug dumps all intermediate outputs; the SDK accuracy debugger compares against a framework reference. ONNX Runtime: the QDQ loss debug utilities add intermediate outputs and match FP32 and quantized activations. LiteRT: the quantization debugger or extra model outputs. llama.cpp: evaluation callbacks print per-tensor stats. Compare on identical inputs and align names or node ids.

Why do NPUs need static shapes, and how do you handle variable-length LLM input?

NPU compilers plan tiling, on-chip memory allocation, DMA schedules and instruction streams for exact tensor sizes at compile time; dynamic shapes break that planning. For LLMs, export two graphs: a prefill graph processing fixed-size chunks (for example 128 tokens, padding the last chunk and masking) and a decode graph processing one token, both with a KV cache of fixed maximum length passed as inputs/outputs and an attention mask marking valid positions. Weights are shared between the graphs. For other models, pad or resize inputs into a few fixed buckets.

Why are LLMs split into several context binaries for the NPU?

NPU sessions have limits on graph size and on the memory that can be mapped into the NPU's address space, and very large graphs compile slowly and exceed on-chip planning limits. Splitting the model by layers into a few binaries (for example 3-5 parts for a 3B model) keeps each within limits; the runtime executes them in sequence, passing hidden states between them. Weight sharing between the prefill and decode variants of each part avoids storing weights twice.

What are Hexagon library versions and ADSP_LIBRARY_PATH about?

The QNN HTP backend has an ARM-side library (loaded by your process) and DSP-side "skeleton" libraries that run on the Hexagon processor, built per architecture: v73 for Snapdragon 8 Gen 2, v75 for 8 Gen 3, v79 for 8 Elite. The skeleton must match the SoC. ADSP_LIBRARY_PATH tells the DSP loader (via FastRPC) where to find them; LD_LIBRARY_PATH or the app's native library directory covers the ARM side. A mismatch causes load failures or fallback. Build-time and runtime SDK versions must also match.

Compute the KV cache size for Llama-3.2-1B at 8k tokens, and for the 3B.

1B: 16 layers, 8 KV heads, head dim 64 (2048 / 32 heads). Per token in FP16: 2 x 16 x 8 x 64 x 2 bytes = 32 KiB. At 8192 tokens: 256 MiB; INT8 about 128 MiB; INT4 about 64 MiB plus scales. 3B: 28 layers, 8 KV heads, head dim 128, so 112 KiB per token and about 896 MiB at 8k in FP16. Both use GQA; with full multi-head attention the numbers would be several times larger.

Why quantize keys per channel and values per token?

Key vectors have a few channels with consistently large magnitudes across tokens (outlier channels, partly due to RoPE and learned structure). Per-token scales would be dominated by those channels and crush the others, so per-channel scales (computed across tokens) fit keys better. Values do not show such fixed-channel outliers but vary per token, so per-token scales fit them. Verify by plotting magnitude per channel for K and V on your model; practical implementations may group tokens for per-channel key scales since the cache grows.

Explain sliding-window attention and attention sinks.

Sliding-window attention limits each token to attend to the last W tokens, so the KV cache is a fixed-size ring buffer and memory is bounded regardless of conversation length. The cost is losing direct access to older context. Models trained with full attention tend to dump a lot of attention mass on the first few tokens ("sinks"); evicting them destabilises generation and perplexity explodes. Keeping a handful of initial tokens plus the recent window restores stability for long streams. It does not restore retrieval of facts that fell out of the window.

What does a paged KV cache buy you on a device?

Instead of one contiguous buffer per sequence (which must be reallocated and copied as it grows, or preallocated for the maximum), the cache is split into fixed-size blocks from a preallocated pool, mapped by a block table. Benefits: no large reallocation stalls, no fragmentation of big contiguous regions, memory proportional to actual length, easy sharing of prefix blocks between sequences, and an explicit, enforceable memory budget. Costs: indirection in the attention kernel and partly filled last blocks. It matters most for multi-session services.

What is prefix caching and when does it help?

If many requests start with the same tokens (system prompt, tool instructions, a document being questioned), compute their KV cache once and reuse it, so each request only prefills the new suffix. It cuts TTFT and energy, often dramatically when the shared prefix is long. On device you can persist the prefix cache to storage and mmap it. It must be invalidated when the model, quantization, prompt text or position handling changes, and costs storage equal to the KV size of the prefix.

What are good practices for a JNI bridge to a native inference engine?

Keep the interface coarse: create, run/generate, cancel, destroy with an opaque handle, rather than per-tensor calls. Pass large buffers as direct ByteBuffers to avoid copies. Delete local references in long loops, cache method ids, attach native worker threads to the JVM before calling back, and never hold JNI references across threads without global refs. Handle errors by returning status codes or throwing Java exceptions, not crashing. Make generation cancellable via an atomic flag, and ensure destroy is idempotent and thread-safe.

Describe a robust model download and install flow.

Fetch a signed manifest (id, version, URL, size, SHA-256, required runtime, SoC and RAM requirements). Check eligibility and free storage. Download in a background worker with constraints (unmetered, optionally charging) and HTTP range resume into a temporary file. Verify checksum and signature. Atomically move into a versioned directory. Run a smoke-test inference, then switch the active version pointer. Keep the previous version until the new one is proven, then garbage-collect. Handle "not yet downloaded" in the UI and expose a remote kill switch.

What would you put in a Perfetto trace for an inference feature?

Scheduling (sched_switch) to see which cores threads run on, CPU frequency and idle events, thermal events, the app's atrace categories and custom slices (Trace.beginSection / ATrace_beginSection) around load, pre-processing, prefill, decode steps and post-processing, the android.power data source with power rails and battery counters, and process stats for memory. Optionally heapprofd for native allocations. This lets you correlate model phases with frequency drops, thermal events and power.

What are the drawbacks of memory-mapped weights?

Page faults on first access add latency to the first inference unless you prefault. Under memory pressure the kernel may evict clean pages, and re-reading them from storage during decode causes large latency spikes. If the runtime repacks weights into a different layout at load, it allocates anonymous memory anyway, losing reclaimability and adding load time. Encrypted or compressed models cannot be mapped directly. Storage speed and file system also affect cold-load behaviour.

How do LoRA adapters work at inference time, and what is the trade-off between merged and unmerged?

A LoRA adapter adds a low-rank update B·A to selected weight matrices (for example attention projections), so the effective weight is W + (alpha/r)·B·A. Merged: fold the update into W once; no runtime overhead, but switching tasks means re-merging or keeping multiple full copies, and it complicates quantized weights. Unmerged: keep the small A and B matrices separate and compute the extra low-rank product each forward pass; a few percent overhead but instant hot-swap between adapters on one shared base model. Adapters are tied to one base model version and quantization.

Advanced

Why can't most NPUs do dynamic activation quantization, and what does static quantization cost?

Integer NPU pipelines precompute requantization parameters (multipliers and shifts combining input, weight and output scales) when the graph is compiled, and schedule data movement assuming fixed formats. Dynamic quantization needs a data-dependent reduction (min/max) over each activation tensor before the matmul, a synchronisation point that breaks streaming dataflow and requires flexible scalar logic. The cost of static ranges: outliers in production data clip, or ranges set wide to include outliers waste resolution for normal values. For transformers this is why NPUs use 16-bit activations (W4A16/W8A16), and why calibration data quality and outlier-handling methods (rotations, smoothing) matter so much.

Explain how an INT8 matmul is computed and requantized on integer hardware.

Weights and activations are stored as int8 with scales s_w, s_a and zero-points. The kernel multiplies int8 values and accumulates into int32 (subtracting zero-point terms, often precomputed into a bias correction). The real result equals s_w·s_a times the int32 accumulator. To produce int8 output with scale s_y, multiply by M = s_w·s_a/s_y, which is represented as an integer multiplier and a right shift (fixed-point), add the output zero-point, round and saturate. Bias is pre-quantized to int32 with scale s_w·s_a. Per-channel weights mean one M per output channel. Differences in rounding mode and saturation between implementations explain small host vs device mismatches.

What are activation outliers in LLMs and how do different methods handle them?

A few hidden dimensions carry values much larger than the rest, consistently across tokens. With per-tensor activation scales, they force a large scale and most values quantize to a few levels. Remedies: keep activations at 16-bit (NPUs) or quantize dynamically per token (CPUs); per-channel handling where hardware allows; SmoothQuant-style migration that divides activations by per-channel factors and multiplies weights by the same factors, moving difficulty into weights; rotation methods (QuaRot, SpinQuant) that multiply by orthogonal matrices to spread outlier energy across all channels, making both weights and activations easier to quantize; and mixed precision for the affected layers.

Derive an upper bound on decode speed and use it to diagnose a slow deployment.

Bytes read per token is roughly the weight bytes actually used per token plus the KV bytes at the current context. Decode tok/s is at most effective bandwidth divided by that. Example: 1B model at about 1 GB INT4 including scales and 8-bit embeddings (embeddings are only gathered, so the actual read is somewhat less), effective bandwidth about 45 GB/s, so the ceiling is about 45 tok/s. If you measure 15, you are far from the bound: suspect dequantization-heavy or scalar kernels, wrong thread placement, disabled KV cache or SDPA, fallback ops, frequency caps, or reading FP32 copies. If you measure 40, you are near the roofline and only fewer bytes (lower precision, smaller model) or speculative decoding will help.

Would you run prefill and decode on different processors? What are the trade-offs?

Prefill is compute-bound, so the NPU's high integer throughput cuts TTFT dramatically. Decode is bandwidth-bound, and all processors share the same DRAM, so NPU, GPU and CPU decode speeds are closer; the choice then comes down to energy per token, thermal behaviour and contention. Splitting phases across processors requires the KV cache in a format and memory both can access (shared buffers, same quantization and layout) or conversion costs at the handover, two sets of compiled kernels and more memory. Many shipping stacks keep both phases on the NPU for simplicity and energy, and use the CPU as fallback.

What is weight sharing between prefill and decode graphs, and why is it needed?

Static shapes force separate graphs for prefill (chunk of N tokens) and decode (1 token). If each graph embedded its own copy of the weights, memory and storage would double. Weight sharing compiles both graphs against a single set of weight buffers in the same context, so switching graphs costs nothing in memory. It requires both graphs to use identical quantization encodings for the shared weights, which constrains per-graph quantization choices.

What role does on-chip memory (VTCM) play in NPU performance?

VTCM is fast scratch memory next to the Hexagon vector and tensor units. The compiler tiles operators so that working sets (weight tiles, activation tiles) fit in VTCM, streaming data from DRAM via DMA while computing on previous tiles. Operators or graphs whose tiles do not fit spill to DRAM, adding bandwidth and latency. Large activation tensors (high-resolution images, long prefill chunks) and wide layers are typical spill sources. Mitigations: smaller prefill chunk sizes, model splitting, layout choices, and compiler options that control VTCM usage.

Why do FP16 overflows happen on GPU/NPU paths, and how do you fix them?

FP16 has a maximum of 65504 and limited precision. Transformer activations with outliers, attention logits before softmax, sums of squares in norms, and large accumulations can overflow or lose precision, producing inf/NaN or degraded outputs, even though the FP32 model is fine. Fixes: keep sensitive ops (norms, softmax, final layers) in FP32; use FP32 accumulation where supported; rescale (for example compute norms with a pre-scaling factor); use BF16 on hardware that supports it; or quantize with 16-bit integer activations which have well-defined ranges. Layer-wise diffing locates the first overflowing op quickly.

Does speculative decoding help on device, and what does it cost?

It helps because decode is bandwidth-bound: verifying k draft tokens in one forward pass of the target model reads the weights once for several tokens, so accepted tokens are nearly free. Expected tokens per pass = (1 − αk+1) / (1 − α); speed-up ≈ that / (1 + k · c). Speedups of 1.5-2.5x are common when the draft is accurate. Costs: a draft model's memory and its own KV cache (or extra heads for self-speculative methods), extra compute that raises power, complexity in cache rollback when drafts are rejected, and poorer gains on creative, high-entropy text. On NPUs, verification needs a static k-token graph. It pays off most on long, predictable outputs such as summaries or code.

Why does static KV allocation for the maximum context matter for memory planning?

NPU graphs and many optimized CPU runtimes allocate KV tensors at their maximum length because shapes are static. A 3B model compiled for 4k context reserves about 448 MiB of FP16 KV even for a 20-token question; at 16k it would be about 1.75 GiB. So choosing max context is a memory decision, not just a capability one. Options: compile several context variants and choose per request or device tier; use quantized KV; use paging on runtimes that support it; or cap context and use summarisation or retrieval to stay within it.

How does GQA change KV memory and NPU efficiency?

Grouped-query attention shares each K/V head among several query heads (for example 32 query heads and 8 KV heads in Llama-3.2-1B), cutting KV memory and KV bandwidth by the group factor (4x there) with small quality impact. For decode, less KV to read means higher tok/s at long contexts. For NPUs, implementations may broadcast K/V to match query heads (costing memory traffic) or reshape queries to batch the heads in a group; the exported attention layout affects whether the compiler maps it efficiently.

How do you design mixed precision under NPU constraints?

Start from layer sensitivity, but check what the backend supports in one partition: some NPU stacks support per-op precision (INT4 and INT8 weights, 8- and 16-bit activations) within a graph, others force a partition break or CPU fallback when precision changes, which can erase the benefit. Prefer promoting whole blocks or op types consistently, keep promoted ops on the NPU (for example W8A16 instead of FP32), keep encodings consistent across graphs that share weights, and re-profile placement after every change. Measure the cost in both latency and partition count, not just size.

How would you estimate the cost of graph partitioning?

For each boundary: data transfer time (tensor bytes over effective bandwidth, possibly two copies), format conversion (quantize/dequantize, layout transpose), synchronisation latency (a round trip to the NPU driver, often tens to hundreds of microseconds), plus lost pipelining. Multiply by the number of boundaries per inference (or per token for LLMs). If a model has 20 boundaries at 200 microseconds each, that is 4 ms of overhead, which can exceed the NPU compute time of a small model. The fix is removing boundaries (op rewrites, moving pre/post-processing out of the graph), not faster kernels.

How do you achieve zero-copy data flow between camera, GPU and NPU?

Use hardware buffers that all components can import: AHardwareBuffer (backed by dmabuf) from the camera or ImageReader, import them into the GPU (EGL/Vulkan) for resize and colour conversion into another hardware buffer, and pass that buffer to the NPU runtime through its shared-memory API (for Qualcomm, rpcmem/ION-dmabuf registered with QNN, or runtime-specific buffer interop). Avoid round-trips through Java arrays or CPU memcpy. Watch for format requirements (NHWC, alignment, quantized input types) that force a conversion, and ensure cache coherency and synchronisation fences between producers and consumers.

Design a shared on-device inference service used by several apps.

A bound system or privileged service exposing a versioned AIDL interface guarded by a permission. One base model resident (mmapped), with per-task LoRA adapters loaded on demand. A scheduler with a request queue, per-client quotas, priority for foreground callers, streaming callbacks (avoid large Binder transactions), cancellation and linkToDeath cleanup. A memory governor using PSI and onTrimMemory to shrink KV budgets, evict adapters or unload the model. A thermal governor using thermal headroom to pace or refuse. A backend policy NPU, GPU, CPU, refuse. Safety filtering on inputs and outputs, input size limits, and telemetry. Updates delivered as deltas and swapped atomically.

Explain Android memory accounting relevant to model deployment.

RSS counts all resident pages of a process, including shared ones; PSS divides shared pages among sharers and is what dumpsys meminfo reports as the app's footprint; USS is private-only. Anonymous memory (heap, repacked weights, KV cache) can only be reclaimed by compressing into zRAM (swap), which costs CPU and compresses quantized data poorly. File-backed clean pages (mmapped weights) can be dropped and re-read. lmkd decisions depend on overall pressure and oom_score_adj, not your RSS directly. GPU and NPU memory may be accounted under graphics or dmabuf and missed if you only watch heap.

How would you measure the maximum usable context on an 8 GB phone under realistic pressure?

Create a realistic background: a foreground workload (for example a memory-heavy app or a synthetic allocator holding a typical footprint) plus normal services. Run generation while increasing context in steps (1k, 2k, 4k...), recording PSS, PSI, zRAM usage, kills (am_kill events, lmkd logs) of your process and others, and decode speed. The usable limit is the largest context without killing the foreground app and without killing important background apps or severe PSI stalls. Repeat for FP16, INT8 and INT4 KV and sliding-window configurations to show how each technique moves the limit.

How do big.LITTLE scheduling and thread affinity affect CPU inference?

Phones combine prime, performance and efficiency cores with very different speeds. Parallel matmuls split work evenly across threads, so a thread on an efficiency core becomes the straggler that everyone waits for. Use as many threads as performance-class cores, pin or hint them to those cores where the runtime allows, and avoid oversubscription with UI and rendering threads. Performance hints (ADPF performance hint sessions) tell the scheduler your target work duration so it can choose appropriate frequencies. Spin-waiting thread pools can burn power; tune spin times for decode.

How do you keep model builds reproducible across SDK and runtime versions?

Pin every tool version (framework, exporter, quantizer, vendor SDK) in a locked environment or container; make export scripts deterministic (fixed seeds, fixed calibration set with a hash); record a manifest per artifact (source checkpoint hash, recipe, tool versions, target SoC, runtime version required); store artifacts in a registry keyed by hash; ship the matching runtime libraries with the artifact; run a numeric regression test (outputs on fixed inputs within tolerance) and a performance regression test on a device farm for every build. Invalidate on-device compiled caches when any version changes.

How do you evaluate a quantized LLM reliably against its reference?

Use several complementary signals: perplexity on held-out, domain-relevant text; next-token top-1 agreement and KL divergence against the FP32 model on a fixed prompt set (teacher-forced, so errors do not compound); greedy decoding comparison measuring the first divergence position; task benchmarks relevant to the product (including structured output validity, arithmetic, multilingual); and, for product features, a rubric-based or human evaluation on real prompts. Report confidence intervals, and evaluate on the device output, not only host simulation.

Why does unstructured pruning rarely speed up edge inference, while structured pruning can?

Unstructured pruning zeroes individual weights; unless sparsity is very high and hardware or kernels support sparse formats, dense kernels still read and multiply the zeros, and index overheads can make sparse kernels slower. Structured pruning removes whole channels, heads or layers, producing a smaller dense model that every runtime accelerates directly. Semi-structured patterns (like 2:4) help only on hardware with dedicated support. On phones, structured pruning plus distillation or simply choosing a smaller model is usually the practical route.

How do runtimes decide which ops go to an accelerator, and how can you influence it?

A partitioner walks the graph, asks the backend whether each node (with its data types, shapes and attributes) is supported, groups contiguous supported nodes into subgraphs (respecting dependencies and sometimes minimum partition sizes), and replaces each with a delegate call. You influence it by rewriting unsupported ops into supported equivalents before export, fixing data types (int64 to int32, FP32 to quantized), making shapes static, using the backend's quantizer so encodings are compatible, configuring partitioner options (skip lists, precision), and moving pre/post-processing out of the graph.

How do you protect a valuable on-device model?

Accept that anything that runs on a user's device can ultimately be extracted by a determined attacker with root. Raise the cost: download at runtime instead of bundling, store in app-private storage, encrypt at rest with keys from the Android Keystore and decrypt into memory (losing mmap benefits), verify integrity and signatures before loading, use runtime integrity checks for the app, split the most valuable part to the server, and use licensing and legal measures. For compiled NPU binaries, the format itself offers some obfuscation but not security.

How do binary delta updates for models work, and when do they make sense?

Compute a binary diff (bsdiff-style or chunk-based) between the installed and new model files on the server; the device downloads the patch and reconstructs the new file, verifying its hash. It saves bandwidth when changes are localized, such as a fine-tuned adapter or partially changed layers. Quantized weights after re-quantization often change almost everywhere, reducing delta effectiveness; chunk-aligned formats and stable layouts help. Patching needs temporary storage for both versions and CPU time, so run it while charging and idle.

When does a mobile GPU beat the NPU?

When the model uses ops, shapes or precisions the NPU does not support well (dynamic shapes, unusual attention variants, FP16-only accuracy requirements), when a model changes often and NPU compilation friction is too high, for moderate-size models where the GPU's FP16 throughput is enough, and for workloads already on the GPU (image processing, rendering) where staying on the GPU avoids transfers. The NPU generally wins on energy efficiency and sustained performance for large quantized models, especially prefill.

How would you design a fair CPU vs GPU vs NPU comparison?

Same model, same inputs, same prompt and output lengths, each backend with its best realistic quantization (documented, with the accuracy delta measured against FP32), same thermal starting point, warm-up and repetitions. Report load/compile time, TTFT, prefill and decode tok/s, peak memory, energy per 100 tokens with idle subtraction, and a sustained 10-minute curve for each. Note what runs where (partitions, fallback ops) and include pre/post-processing time in an end-to-end number. Publish configuration, versions and raw data so others can reproduce it.

How do RoPE positions interact with sliding windows and cache eviction?

With rotary position embeddings, positions are applied to keys before they enter the cache, so cached keys already encode their absolute positions. If you evict middle tokens or wrap a window, the relative distances between the new query and remaining keys stay consistent as long as you keep using the true absolute positions, but positions can grow beyond the trained range in long streams. Some streaming methods instead assign positions within the cache (re-indexing) and store keys before rotation, applying RoPE at attention time, which keeps positions within range but costs extra compute. Getting this wrong shows up as degradation after the window first fills.

Scenario & debugging

The model is 3x slower on the NPU than the vendor's published numbers. How do you investigate?
  1. Match conditions: same model variant, precision, input size or prompt length, SoC, SDK version, and whether they reported compute-only time.
  2. Check placement: per-op profile (QNN profiler, AI Hub profile, runtime op profiling) to find CPU fallback ops and count partitions.
  3. Check initialization: is graph compilation or context generation counted in each run? Use a precompiled context binary or cache.
  4. Check performance mode: burst or sustained high performance rather than default or power saver; check that the NPU is not shared with camera or other clients.
  5. Check data movement: copies between Java and native, quantize/dequantize of inputs and outputs on CPU, layout transposes (NCHW vs NHWC) outside the graph; use shared buffers.
  6. Check precision: an FP32 model running as FP16 or partially on CPU; ensure the model is quantized in the format the HTP expects.
  7. Check thermal state and spill: DDR spill from VTCM for large tensors; try smaller tiles or chunk sizes.

Fix the largest contributor, re-profile, and repeat.

Accuracy dropped by 6 points after INT8 quantization. What do you do?

First rule out non-quantization causes: compare the converted FP32 model with the original (conversion bug?) and confirm the evaluation pipeline and pre-processing are identical. Then inspect the calibration set (size, representativeness, pre-processing identical to inference). Run layer-wise diffs (SQNR, cosine) and single-layer sensitivity to find where error originates. Typical fixes in order: per-channel weights, better calibration method (percentile or entropy instead of min-max), cross-layer equalisation and bias correction or AdaRound for CNNs, 16-bit activations for sensitive layers, mixed precision for the worst few layers, and QAT if PTQ cannot close the gap. Re-validate on device, because device kernels can differ from host simulation.

The phone throttles after two minutes of LLM generation and tok/s drops 40%. What can you do?

Measure first: thermal soak with temperatures, frequencies, power per rail and throughput to confirm it is thermal and find the dominant power consumer. Then reduce energy per token: move to the NPU (lower power per token than CPU), use lower precision weights and KV, reduce thread count (fewer cores at lower power sometimes sustain better), avoid spin-waiting thread pools, cut wasted work (stop tokens, shorter outputs, prefix caching to avoid repeated prefill). Manage the budget: use thermal headroom APIs to pace generation proactively, pick a sustained performance mode, and degrade gracefully (smaller model or shorter answers) when hot. Test with the case on at realistic ambient temperature, and set product expectations on sustained, not peak, numbers.

The app gets killed when a conversation grows past about 3k tokens on 8 GB phones. Why, and how do you fix it?

The KV cache and activation buffers grow with context (or were allocated for a large max context), pushing total anonymous memory high enough that lmkd kills processes; logcat lmkd messages and PSI confirm it. Fixes: calculate and enforce a memory budget per tier; quantize the KV cache (INT8 halves it); cap context for this tier and use sliding window with sinks, summarisation of older turns, or retrieval; use a paged cache to avoid reallocation peaks; mmap weights so they are reclaimable; free caches on onTrimMemory; make sure there is no duplicate copy of weights (compressed asset, repacking). Verify with the pressure test at several contexts.

The first inference after launch takes 8 seconds. How do you reduce it?

Break it down with trace slices: file read or decompression, weight repacking, delegate or NPU graph compilation, GPU shader compilation, tokenizer load, first-run page faults. Fixes: store uncompressed and mmap; ship precompiled context binaries or enable runtime compiled-model caching (EP context, GPU serialization caches) so compilation happens once; move weight repacking offline into the exported format; initialize asynchronously at app start or on a trigger that predicts use; warm up off the critical path; split the model so a small part is ready first. Measure cold (after reboot) and warm separately.

Decode tok/s is half of the reference number for the same model and phone. What do you check?

Release vs debug build; KV-cache and SDPA flags enabled in export; the right quantization (INT4 weights, not an FP32 fallback); optimized kernels enabled (KleidiAI, dotprod/i8mm); thread count equal to performance cores and threads not landing on efficiency cores; device temperature at start; background load; prompt and output lengths matching the reference; power-saving mode off; and whether the reference measured with a different context length. Profile with simpleperf to see which kernels are hot. Compare with llama.cpp as a second opinion on the same device.

On device the LLM outputs gibberish or never stops, but on the host it is fine.

Most likely a tokenizer or prompt issue: different tokenizer file, missing BOS, wrong special token ids, chat template not applied, or EOS ids missing from the runner's metadata (so it never stops). Then check the export: KV-cache positions, max sequence length exceeded, RoPE scaling parameters. Then numerics: run greedy decoding with identical token ids on host and device and compare logits at the first step; if logits differ strongly, do a layer-wise diff to find an FP16 overflow or a miscompiled op. Also check sampling parameters (temperature, top-k) are what you expect.

The feature works on your flagship but crashes on a mid-range phone.

Collect the native crash (tombstone) and logcat: common causes are out-of-memory during load (a model too large for the RAM tier, or compressed assets doubling memory), missing CPU instructions (a library built with i8mm or other extensions on a core without them), an unsupported accelerator path that is not guarded (NPU libraries for another Hexagon version), or GPU driver bugs. Fix with device-tier gating, runtime CPU feature detection and multiple kernel variants, capability probes before enabling a backend, graceful fallback, and a test matrix including low and mid-tier devices.

Outputs are fine for short prompts but quality degrades after 1-2k tokens.

Suspect the cache and positions: KV cache index or wrap-around bugs, sliding window evicting sink tokens, RoPE position handling when the window rolls, exceeding the compiled max context, or KV quantization error accumulating with length (especially keys quantized per token instead of per channel). Also check whether the model itself was trained for that context (and whether RoPE scaling is applied in the exported model). Reproduce with FP16 KV and full attention to isolate, then compare logits at increasing positions against the host reference.

The INT4 model runs at the same speed as the FP16 model. Why?

The kernels may be dequantizing weights to FP16/FP32 into a temporary buffer and then running a dense float matmul, so bandwidth is not reduced; or the quantized ops fell back to reference (scalar) kernels because the scheme or group size is not supported by the optimized path; or the time is dominated by something else (unquantized embeddings or LM head, attention with a large KV cache, pre/post-processing, CPU fallback partitions); or you are compute-bound in prefill where INT4 weights help less without integer compute. Profile to see which kernels run and whether bytes read per token actually dropped.

The NPU session fails to load on one SoC generation but works on others.

The context binary was built for a different Hexagon architecture, the DSP skeleton libraries for that architecture are missing or not on ADSP_LIBRARY_PATH, the runtime library version differs from the SDK that built the binary, or the device's firmware/driver is older than required. Check logcat for FastRPC or QNN errors. Fix by shipping per-architecture binaries and libraries selected at runtime from the SoC model, pinning SDK versions, and falling back to GPU/CPU when the probe fails.

The camera feature drops frames even though the model runs in 5 ms.

The model is not the bottleneck; the pipeline is. Trace the full frame path: YUV to RGB conversion and resize on the CPU, copying frames into Java arrays, allocating buffers per frame, running post-processing (NMS, mask upsampling) on the CPU, synchronously waiting on the GPU, or rendering overlays on the main thread. Fix with GPU or ISP-based pre-processing, hardware buffers and zero-copy, buffer reuse, pipelining stages across frames, dropping stale frames rather than queueing, and native post-processing. Measure end-to-end latency and p99 frame time, not model time.

After shipping, users complain about battery drain. How do you find and fix the cause?

Segment telemetry by device and usage: how often the model runs, on which backend, for how long, and whether it runs in the background. Reproduce with batterystats and power rails. Common causes: inference running more often than needed (every frame when every fifth would do), CPU fallback instead of NPU, spin-waiting thread pools, repeated model loading, generation not cancelled when the user leaves, or background work not constrained. Fix with duty cycling, event triggers, batching, lower frame rates, NPU placement, cancellation, WorkManager constraints, and track energy per request as a release gate.

p99 latency spikes periodically, even though p50 is stable.

Correlate spikes in a Perfetto trace with system events: thread migration to efficiency cores, CPU frequency drops, thermal mitigation steps, garbage collection pauses in the Java layer, page faults from evicted mmapped weights, memory allocation in the hot path, contention with rendering or other apps, or periodic background jobs. Fixes: preallocate buffers, prefault or lock weights, use performance hints, avoid allocations per inference, move work off threads that compete with the UI, and use sustained performance modes for continuous workloads.

The same model gives noticeably different results on two phone models.

Different backends or kernels run on each: one may use the NPU with static INT8 and another the GPU with FP16 or the CPU; vendor drivers implement ops with different rounding, accumulation precision or approximations (for example exp and softmax); FP16 overflow may occur on one GPU and not another. Log the backend used per device, compare intermediate outputs on a fixed input, and test the same backend on both. Mitigate by constraining precision for sensitive ops, pinning backends for critical features, and setting accuracy tolerances per device family.

A new model rollout increased crash rate, but only on one OS version.

Halt the staged rollout (or flip the kill switch for that cohort) immediately. Symbolize the native crashes and check whether they come from the runtime or driver libraries; an OS update may have changed the GPU/NPU driver or a system library the runtime depends on, or invalidated a compiled cache format. Reproduce on that build, add a device/OS deny-list or a different backend for it, report to the vendor with a minimal repro, and add that OS version to the pre-release test matrix. Longer term, validate caches against driver versions.

The quantized LLM has fine perplexity but fails at arithmetic and producing valid JSON.

Perplexity averages over typical text and hides skill-specific damage; digits, brackets and rare formatting tokens can be disproportionately affected, often via the LM head or embedding quantization and outlier-heavy layers. Build targeted evaluations (arithmetic set, JSON schema validity, your product prompts), run layer sensitivity with those tasks as the metric, keep the LM head and sensitive layers at higher precision, try better PTQ (GPTQ, AWQ, rotations) or QAT with LoRA on task-relevant data, and use constrained decoding (grammar-guided sampling) for JSON so the format is guaranteed.

The GPU delegate is slower than the CPU for your model.

Likely causes: the model is small, so GPU dispatch and synchronisation overhead dominate; some ops are unsupported and fall back to CPU, creating copies; data transfers of inputs and outputs between CPU and GPU memory each call; shader compilation included in timing; quantized INT8 models running on a GPU path that dequantizes to FP16; or the GPU is busy rendering. Check delegation coverage and partition count, exclude initialization, use GPU buffers directly, enable serialization caches, and consider that the CPU with XNNPACK is often the right choice for small models.

Peak memory at load is twice the model size.

Typical causes: the asset is compressed so it is decompressed into memory in addition to being read; the runtime reads the file into a buffer and then repacks weights into another layout, keeping both until load finishes; a Java byte array copy plus a native copy; a GPU delegate uploading weights while CPU copies are still resident. Fix with uncompressed mmapped files, offline pre-packing into the runtime's preferred layout, releasing source buffers after upload, streaming loads, and measuring with heapprofd and dmabuf accounting to confirm.

Product wants 8k context with a 3B model on 8 GB phones. What is your plan?

Do the budget: 3B INT4 weights about 2 GB, FP16 KV at 8k about 900 MB, plus activations and runtime overhead, roughly 3.5 GB total, which is risky on 8 GB phones with a foreground app. Options: INT8 or INT4 KV (450 or about 250 MB), a sliding window with sinks for chat history plus retrieval for documents, summarising older turns, prefix caching of the system prompt to save TTFT, or a smaller model for this tier with 8k context. Validate with the memory-pressure test and thermal soak, and propose a tiered plan: full 8k on 12 GB+, 4k plus retrieval on 8 GB.

A product manager asks to run a 7B model on mid-range phones. How do you respond?

Quantify it: 7B at INT4 is about 4 GB of weights, plus KV and runtime, over 5 GB resident; decode on a mid-range phone with maybe 25-35 GB/s effective bandwidth is at most around 6-8 tok/s, falling with context and heat; energy per response and load time are high; and many mid-range devices lack an NPU path for it. Offer alternatives tied to the product goal: a 1-3B model fine-tuned or distilled for the specific task, LoRA adapters, retrieval to supply knowledge, a cascade where hard queries go to the cloud with consent, or 7B only on high-end tiers. Back the recommendation with a quick prototype measurement.

The streaming chat UI janks while tokens are being generated.

Check that inference is not on the main thread and that token callbacks do not trigger heavy work per token (full text re-layout, markdown parsing of the whole message, list diffing). Batch UI updates (for example every 50 ms or every few tokens), append incrementally, and do text processing off the main thread. Also check CPU contention: inference threads saturating all performance cores can starve the render thread; leave a core free, lower inference thread priority, or move inference to the NPU. Verify in a trace that frames meet their deadlines.

ONNX Runtime with the QNN EP is silently running most of the model on the CPU.

Enable verbose logging and set session.disable_cpu_ep_fallback to make unsupported nodes fail loudly and name them. Typical reasons: the model is not QDQ-quantized in the format the HTP expects; unsupported ops or data types (int64, dynamic shapes, certain reductions); shapes not static; the QNN backend library not found so the EP was not registered. Fix by running the QNN preprocessing and quantization helpers, making shapes static, rewriting unsupported ops, and verifying placement through the profile before re-enabling fallback.

A vision model is accurate in lab tests but poor with the real camera.

Dump the exact tensors the app feeds the model and compare with the lab pipeline: colour order, normalisation, resize and crop method, rotation from sensor orientation, YUV conversion ranges (full vs limited), and aspect-ratio handling. Then consider domain shift: lighting, noise, motion blur and lens differences not present in the dataset, and a calibration set that did not include real camera frames. Fix the pipeline mismatch first; then recalibrate with device-captured data, augment or fine-tune with such data, and build a device-captured evaluation set.

Three apps from your company each want their own LLM. What do you propose?

Three separate copies would triple memory and storage and fight for the NPU. Propose one shared service hosting a single base model, with per-app LoRA adapters swapped at request time, a stable AIDL API with permissions, a scheduler with quotas and priorities, and memory and thermal governors. If the OS provides a platform AI service with a suitable model, evaluate using it instead. Measure adapter swap cost and service overhead against in-process inference, and define versioning so adapters are retrained when the base model updates.

A wearable activity model drains the battery through false wake-ups.

Measure false wake-ups per hour and energy per wake-up, since waking the application processor costs far more than the model. Move the first-stage detector to the sensor hub or low-power core with a tiny model and a stricter threshold, add a second-stage confirmation model on the main processor, batch sensor data in hardware FIFOs, tune thresholds on realistic day-long recordings (not only labelled activity clips), add hysteresis or temporal smoothing, and evaluate per user. Track battery impact per day as the release metric.

Swapping LoRA adapters takes 1.5 seconds, which is too slow.

Break down the time: reading the adapter file, converting or quantizing it at load, merging into base weights, or recompiling an NPU graph. Fixes: keep adapters unmerged and apply them as separate low-rank ops so swapping is a pointer change; preload likely adapters into memory; store adapters in the runtime's final format and precision; on NPUs, compile graphs with adapter weights as updatable inputs rather than constants so swapping does not require recompilation; and cache per-adapter prefix KV if system prompts differ per task.

Your benchmark numbers vary by 20% between runs on the same device.

Control the sources of variance: start temperature (cool down to a threshold before each configuration), battery level and charging state (charging adds heat and may change governor behaviour), screen state and brightness, background activity (disable sync, use airplane mode), thread placement (pin or use performance hints), and prompt differences. Increase repetitions and report median and spread. Check for thermal throttling within the run and for memory pressure causing page eviction. If variance remains, report it honestly with confidence intervals.