Edge AI Deployment & Projects
This is the hands-on companion to the concepts page: how to actually take a model from a checkpoint to a fast, stable, shippable feature on a phone-class device, with every command, code path and measurement you need. It is organised as a build-along course of five projects (benchmark harness, quantization error analysis, NPU deployment, KV-cache engineering and a full system) plus the profiling, debugging and shipping knowledge interviewers probe.
- Deployment is a loop: choose model, export a static graph, convert, quantize, compile for the target, integrate into the app, benchmark on real silicon, monitor in the field, and repeat.
- Pick one primary path per job: ExecuTorch (.pte) or LiteRT (.tflite) for general apps, llama.cpp (GGUF) for fast LLM baselines, ONNX Runtime for cross-platform, vendor SDKs (QNN/QAIRT context binaries) for NPU efficiency, Core ML on Apple.
- Measure like a product engineer: warm-up, many runs, p50/p90/p99, TTFT, prefill and decode tokens/s, peak RSS, load time, energy per request and a 10+ minute thermal soak. Never trust emulator numbers.
- Quantization problems are found layer by layer: compare every intermediate tensor against an FP32 reference (cosine similarity, SQNR), then keep only the sensitive layers at higher precision.
- NPUs want static shapes, static quantization and full graph coverage; one unsupported op in the middle can make an NPU slower than the CPU.
- For LLMs, decode speed is set by memory bandwidth and long-context feasibility by the KV cache, which can outgrow the weights; manage it with GQA, quantized KV, sliding windows, paging and prefix caching.
- Shipping means tiering by device, downloading models safely, versioning, A/B testing, remote kill switches and a fallback chain (NPU, GPU, CPU, cloud, or refuse).
The end-to-end deployment pipeline
Every edge AI deployment, whether it is a 5 MB keyword spotter or a 1-billion-parameter language model, walks the same eight stages. The names of the tools change; the stages do not. If you can explain each stage, its artifact, and what goes wrong there, you can reason about any stack an interviewer throws at you. The theory behind each stage (what an NPU is, how quantization maths works, how runtimes differ) lives on the concepts page, Edge AI; this page is about doing it.
Think of deploying a model like adapting a restaurant recipe for an airline kitchen. You choose a dish that can survive the trip (choose the model), write it down as an exact, fixed recipe card with no improvisation (export a static graph), translate it into the airline's format (convert), swap ingredients for lighter ones that still taste right (quantize), pre-cook for the specific oven on that aircraft type (compile for the target), load it onto the plane (integrate into the app), test-serve it on a real flight at altitude, not in the test kitchen (benchmark on device), and read passenger feedback after every route (monitor).
Each step maps one-to-one: the recipe card is the exported graph, the airline format is the runtime file (.pte, .tflite, .gguf, context binary), the lighter ingredients are INT8/INT4 numbers, the aircraft oven is the specific NPU generation, and passenger feedback is field telemetry on latency, crashes and quality.
+-------------+ +-----------+ +-----------+ +------------+
| 1. CHOOSE / |--▶| 2. EXPORT |--▶| 3. CONVERT|--▶| 4. QUANTIZE|
| TRAIN | | static | | to runtime| | PTQ / QAT |
+-------------+ | graph | | format | | mixed prec.|
+-----------+ +-----------+ +------------+
|
+-------------+ +-----------+ +-----------+ +------------+
| 8. MONITOR |◀--| 7. BENCH- |◀--| 6. INTE- |◀--| 5. COMPILE |
| field data, | | MARK on | | GRATE in | | for target |
| A/B, rollbk | | device | | the app | | CPU/GPU/NPU|
+-------------+ +-----------+ +-----------+ +------------+
| ▲
+------------- iterate: fix ops, precision, model ----+
The eight stages and their artifacts
| Stage | Input | Output artifact | Typical tools | What usually breaks |
|---|---|---|---|---|
| 1. Choose or train | Task, latency and memory budget | FP32/BF16 checkpoint | Hugging Face, model zoos, vendor model hubs, your own training | Model too large for the RAM tier; licence not suitable for shipping |
| 2. Export | PyTorch/TF module + example inputs | Static graph (ExportedProgram, ONNX, SavedModel) | torch.export, torch.onnx.export, TF SavedModel | Python control flow, data-dependent shapes, unsupported custom ops |
| 3. Convert | Static graph | Runtime format: .pte, .tflite, .onnx/.ort, .gguf, .mlpackage, QNN model | ExecuTorch, ai-edge-torch, ONNX tools, llama.cpp converters, coremltools, QAIRT converters | Op not in the target op set; layout (NCHW vs NHWC) inserted transposes |
| 4. Quantize | Converted or exported graph + calibration data | INT8/INT4/W4A16 model with scales | PT2E quantizers, torchao, AIMET, ORT quantization, LiteRT converter, llama-quantize, AI Hub | Accuracy drop from outliers, poor calibration data, sensitive layers |
| 5. Compile for target | Quantized graph | Delegated program, context binary, compiled model cache | Backend partitioners, QNN context binary generator, AI Hub compile jobs, GPU shader caches | Graph partitioning, CPU fallback, dynamic shapes rejected, SoC version mismatch |
| 6. Integrate | Model + runtime library | APK/AAB or system component with native libs | Kotlin/Java, JNI/C++, Gradle, CMake, NDK | UI-thread inference, asset compression breaking mmap, ABI mismatch, missing libs |
| 7. Benchmark | App or CLI on real hardware | Latency/memory/power/thermal report | Custom harness, benchmark_model, llama-bench, Perfetto, vendor profilers | Cold-run-only numbers, emulator numbers, uncontrolled thermal state |
| 8. Monitor | Shipped feature | Field telemetry, rollback decisions | Remote config, analytics, crash reporting, staged rollouts | No per-device breakdown; no kill switch; silent quality regressions |
How to read this page
- Set up once Read the stack decisions, buy or borrow the right hardware, and complete the environment checklist before writing any model code.
- Learn the paths Walk through the export paths section and run at least two of them end to end (ExecuTorch and llama.cpp are the easiest starting pair).
- Integrate Put one model into a real Android app so you understand packaging, threading and memory.
- Build the five projects in order Each project reuses the previous project's harness and understanding; do not reorder them.
- Harden for production Profiling, accuracy debugging, shipping and the failure-mode checklist turn a demo into a product.
- Revise Use the quick revision list, glossary and interview questions at the end.
Stack decisions: what to learn and what to skip
The on-device stack churned heavily in 2024 and 2025 and has now largely consolidated. Choosing the right tools up front saves months of learning APIs that are being retired. The rule of thumb: learn one general-purpose runtime deeply, one fast LLM prototyping tool, and one vendor NPU toolchain. Everything else you should recognise and be able to discuss.
Picking a stack is like choosing which languages to learn before moving abroad. You learn the national language fluently (your primary runtime), a handful of phrases for the neighbouring regions (secondary runtimes), and the local dialect of the city you will actually work in (the vendor NPU toolchain). You do not spend a year on a language that the government just stopped using (a deprecated API).
The national language is ExecuTorch or LiteRT, the neighbouring regions are ONNX Runtime and Core ML, the city dialect is QNN/QAIRT or NeuroPilot, and the retired language is NNAPI.
Do not invest in these
- NNAPI (Android Neural Networks API). Introduced in Android 8.1 and marked deprecated / not recommended for new work from Android 15. Vendors implemented it inconsistently, so behaviour and performance varied wildly between devices. The replacement direction is vendor delegates and backends that plug directly into frameworks (LiteRT accelerators, ExecuTorch backends, ONNX Runtime execution providers). If a tutorial is built on NNAPI, treat it as historical.
- Training from scratch, distributed training, cloud MLOps pipelines, data-centre serving. Useful elsewhere, but not what edge deployment roles screen for. Fine-tuning small adapters is the one training skill worth having (see Fine-tuning).
- Chasing 7B-plus models on phones early. Quantizing them needs 80 GB-class GPUs, iteration is slow, and a 1B-class model teaches every lesson.
The stack that matters
| Tool | What it is | When to use it | Priority |
|---|---|---|---|
| ExecuTorch (PyTorch Edge) | PyTorch's on-device runtime. Exports via torch.export to a .pte program. Small core runtime (tens of KB), many backends: XNNPACK (CPU), Vulkan (GPU), Qualcomm QNN, MediaTek, Arm Ethos-U, Core ML, MPS. | Primary runtime for PyTorch models and LLMs on Android and iOS. | Core |
| Qualcomm QAIRT / QNN SDK | Qualcomm's AI runtime SDK (the QNN SDK was folded into QAIRT). Converters, quantizers, HTP (Hexagon Tensor Processor) backend, context binaries, profiling, and Genie for generative models. | Maximum efficiency on Snapdragon NPUs. | Core (if targeting Snapdragon) |
| Qualcomm AI Hub | Hosted service that compiles, quantizes and profiles models on real devices in the cloud; model zoo with ready export scripts. | Fastest path to an NPU context binary and per-layer profile without owning every device. | Core |
| llama.cpp / GGUF | C/C++ LLM inference engine with hand-tuned CPU kernels (Arm NEON, i8mm, SVE), GPU backends (Vulkan, OpenCL, Metal) and the GGUF single-file format. | First working LLM baseline in an hour; CPU reference numbers; format for distributing quantized models. | Core |
| LiteRT (formerly TensorFlow Lite) | Google's runtime for .tflite FlatBuffers with XNNPACK CPU, GPU delegate and NPU accelerators; PyTorch models arrive via ai-edge-torch. | Vision, audio and classic models in Android apps; Google ecosystem (MediaPipe). | Useful |
| MediaPipe LLM Inference / LiteRT-LM | High-level on-device LLM API on top of LiteRT, loading bundled .task or .litertlm models. | Shipping a supported LLM (Gemma family and others) with minimal code. | Useful |
| ONNX Runtime Mobile | Cross-platform runtime; execution providers for CPU, NNAPI (legacy), QNN, Core ML, XNNPACK. | When the model pipeline is ONNX-based or you need one runtime across Windows, Android and iOS. | Useful |
| Core ML / coremltools | Apple's runtime and converter; dispatches across CPU, GPU and Neural Engine. | Any Apple target. | Know it |
| AIMET | Qualcomm's quantization and compression toolkit (AdaRound, cross-layer equalisation, QuantAnalyzer, QAT). | Recovering accuracy for NPU INT8/INT4 deployments. | Useful |
| MediaTek NeuroPilot / Neuron runtime | MediaTek's APU toolchain (also reachable via ExecuTorch and LiteRT). | Second vendor, after you are comfortable with one. | Later |
| System AI services (for example AICore with Gemini Nano) | OS-level service hosting a shared foundation model; apps call it through an SDK and cannot load their own weights. | Study the architecture: it is the reference design for a platform AI service. | Study |
Why a closed system AI service is worth studying
A platform AI service such as Android's AICore shows what a production-grade on-device design looks like. It is updated independently of the OS (as an updatable system module), it only enables itself on devices with enough RAM and an NPU, it ships model updates as binary deltas instead of re-downloading gigabytes, it runs safety filtering around the model, it isolates the model from callers behind a permission boundary, and it supports small LoRA adapters (tens of MB) that specialise one shared base model per feature. Every one of those decisions is an answer to a real constraint: storage, bandwidth, memory, privacy, safety and multi-tenancy. It is the blueprint for the shared-inference-service option in Project 5.
Choosing a runtime for a given job
| Situation | Good first choice | Why |
|---|---|---|
| PyTorch LLM on Android and iOS | ExecuTorch (XNNPACK, then QNN/Core ML backends) | Same export flow, many backends, strong LLM tooling |
| Validate an LLM use case this afternoon | llama.cpp with a GGUF | Builds in minutes, runs anything on CPU |
| Vision/audio model in an Android app | LiteRT with GPU/NPU accelerator | Mature Android tooling, MediaPipe tasks, easy benchmarks |
| Supported LLM shipped fast | MediaPipe LLM Inference / LiteRT-LM | A few lines of Kotlin; GPU path included |
| Best perf/W on Snapdragon | QNN context binary via AI Hub or QAIRT (directly, through ExecuTorch QNN backend, or ORT QNN EP) | Direct HTP access, precompiled graphs |
| Same model on Windows on Arm, Android and iOS | ONNX Runtime with per-platform EPs | One API and model format |
| Apple devices | Core ML (or ExecuTorch Core ML backend) | Only supported path to the Neural Engine |
Hardware, software and budget
Performance work on edge AI is only meaningful on physical silicon. Emulators and simulators do not model memory bandwidth, cache sizes, DVFS, thermal throttling, or the NPU at all; a number measured on an emulator is not wrong by a small factor, it is meaningless.
Benchmarking on an emulator is like testing running shoes on a treadmill in an air-conditioned showroom and then promising they will perform in a mountain marathon in summer. The terrain (memory system), the heat (thermal limits) and the fatigue (sustained throttling) are exactly what you did not test.
The showroom treadmill is your laptop or emulator; the mountain marathon is a mid-range phone in a user's pocket at 35 degrees, running your model for ten minutes.
Devices: match the NPU generation
Vendor NPU libraries are versioned per hardware generation. For Snapdragon, the Hexagon library folder must match the SoC; using the wrong one fails to load or silently falls back to CPU.
| SoC (flagship tier) | Hexagon architecture / library folder | Notes |
|---|---|---|
| Snapdragon 8 Gen 1 | v69 (confirm the hexagon-vXX folder in your QAIRT SDK) | Older; usable for INT8 CNNs, limited for LLMs |
| Snapdragon 8 Gen 2 | v73 (hexagon-v73) | Good minimum for LLM NPU work; INT4 weight support |
| Snapdragon 8 Gen 3 | v75 (hexagon-v75) | Common reference device generation |
| Snapdragon 8 Elite | v79 (hexagon-v79) | Current-generation class; strongest LLM NPU numbers |
For other vendors the same principle applies: MediaTek Dimensity APU generations map to specific NeuroPilot versions; Google Tensor chips expose their TPU through LiteRT accelerators; Apple's Neural Engine is reached only through Core ML.
Minimum viable kit
- One Android phone with a recent flagship SoC (for example Snapdragon 8 Gen 2 or newer) and 12 GB+ RAM. A second-hand device is perfectly fine.
- A Linux workstation (Ubuntu 22.04+ native, or WSL2 on Windows) with 16-32 GB RAM and 100 GB+ free disk. Most vendor toolchains are Linux-first.
- Android SDK platform-tools (
adb) and NDK r26 or newer. - A reliable USB-C data cable. Flaky cables cause intermittent
adbdisconnects that look like software bugs.
Ideal kit
- Two phones from different SoC generations or vendors: cross-silicon comparisons are the most informative measurements you can produce.
- A Pixel device to explore the platform AI service path and Tensor TPU accelerators.
- A MediaTek device for second-vendor work later.
- A USB power meter or, better, a bench power supply / power monitor for repeatable energy numbers; otherwise use on-device power rails and the fuel gauge.
- Optional: a single-board computer or microcontroller dev kit (Arm Cortex-M with an Ethos-U NPU, or a Linux SBC with an NPU) for the wearable/sensor project.
GPU access for quantization
Running a model is cheap; quantizing a large one is not. Algorithms such as GPTQ, SpinQuant, AdaRound or QAT need the full-precision model in GPU memory plus calibration activations. Rough requirements for LLM quantization workflows:
| Model class | Typical GPU memory needed to quantize | Practical advice |
|---|---|---|
| ~0.5-1B | 16-24 GB (often fits a consumer GPU) | Do almost all experiments here |
| ~3B | ~40 GB (32 GB is borderline and may OOM) | Rent by the hour only when needed |
| ~7-8B | ~80 GB | Only once, at the end, to prove you can scale |
Strategy: use 1B-class models with post-training quantization for most learning; rent a cloud GPU by the hour for the one or two runs that genuinely need it (a QAT comparison, a vendor LLM export). Shut instances down immediately after; idle GPUs are the main cost leak.
Budget (approximate, varies by region)
| Item | Approximate cost | Notes |
|---|---|---|
| Used flagship Android phone | USD 200-450 | The one thing worth spending on |
| Second device (optional) | USD 150-400 | Enables cross-silicon comparison |
| Cloud GPU hours for quantization | USD 150-300 over six months | Hourly rental, only for heavy runs |
| USB power meter (optional) | USD 20-60 | Rough external power numbers |
| Vendor model hub accounts | Free tier | Registration usually required |
| ExecuTorch, llama.cpp, LiteRT, ONNX Runtime | Free | Open source |
| Realistic total | USD 400-800 | Spread over roughly six months |
Environment setup
Toolchain setup is where most people lose several weekends and give up. Do it once, carefully, in a reproducible way, and timebox it. If the heavy toolchain fights you for more than a day, build llama.cpp first to get a model running on the phone, then come back.
Setting up the environment is like a chef's mise en place: every knife sharpened, every ingredient measured into its bowl, before the stove is lit. Cooking goes fast and calmly once everything is in place; cooking while hunting for ingredients burns the dish.
The knives are your NDK, CMake and Python environment; the measured bowls are the downloaded model weights and SDKs; lighting the stove is your first export and on-device run.
Checklist
- Operating system Ubuntu 22.04+ (native or WSL2) with 100 GB+ free disk. Model artifacts, SDKs and build trees add up quickly.
- Python Python 3.10-3.12 in a dedicated virtual environment or conda environment per toolchain. Never install into system Python; vendor SDKs often pin conflicting versions.
- Android tools SDK platform-tools with
adb devicesshowing your phone asdevice(notunauthorized). - NDK Android NDK r26+ installed;
ANDROID_NDKexported. - Build tools CMake 3.24+, Ninja, a recent Clang, git with LFS.
- Model access A model hub account with accepted licences for gated models (Llama-family weights require accepting the licence).
- Vendor SDK A vendor developer account, AI Hub token, and the QAIRT SDK downloaded and unpacked;
QNN_SDK_ROOTexported. - Phone settings Developer options on, USB debugging on, "stay awake while charging" on, battery optimisation disabled for your test app, adaptive brightness off.
- Version control One git repository for the whole effort. Commit measurement CSVs and configs, not just code, so every number is reproducible.
Host environment
# Base tools (Ubuntu)
sudo apt update && sudo apt install -y build-essential cmake ninja-build git git-lfs \
python3-venv python3-dev clang unzip wget
# Android NDK (adjust version/path to what you installed)
export ANDROID_HOME=$HOME/Android/Sdk
export ANDROID_NDK=$ANDROID_HOME/ndk/26.3.11579264
export PATH=$ANDROID_HOME/platform-tools:$PATH
# Sanity checks
adb devices # should list your phone as "device"
adb shell getprop ro.soc.model # e.g. SM8650 for Snapdragon 8 Gen 3
adb shell getprop ro.product.cpu.abi # must be arm64-v8a
adb shell cat /proc/meminfo | head -3 # total RAM
ExecuTorch host build
git clone --recursive https://github.com/pytorch/executorch.git
cd executorch
python -m venv .venv && source .venv/bin/activate
./install_executorch.sh # installs torch, torchao and the executorch pip package
# Verify before touching Android
python -c "import executorch, torch; print(executorch.__file__, torch.__version__)"
ExecuTorch Android runner build (the flags that matter)
cmake -DCMAKE_TOOLCHAIN_FILE=$ANDROID_NDK/build/cmake/android.toolchain.cmake \
-DANDROID_ABI=arm64-v8a \
-DANDROID_PLATFORM=android-26 \
-DCMAKE_BUILD_TYPE=Release \
-DEXECUTORCH_BUILD_XNNPACK=ON \
-DEXECUTORCH_BUILD_KERNELS_QUANTIZED=ON \
-DEXECUTORCH_BUILD_KERNELS_OPTIMIZED=ON \
-DEXECUTORCH_BUILD_EXTENSION_MODULE=ON \
-DEXECUTORCH_BUILD_EXTENSION_DATA_LOADER=ON \
-DEXECUTORCH_BUILD_EXTENSION_TENSOR=ON \
-DEXECUTORCH_BUILD_EXTENSION_LLM=ON \
-DEXECUTORCH_XNNPACK_ENABLE_KLEIDI=ON \
-Bcmake-out-android .
cmake --build cmake-out-android -j$(nproc) --target install
# Then build the LLM example runner against that install
# (see examples/models/llama in the repo for the exact current target)
EXECUTORCH_XNNPACK_ENABLE_KLEIDI enables Arm's KleidiAI low-bit matrix-multiply micro-kernels inside XNNPACK. They give a large prefill improvement (on the order of 20% or more) on Arm CPUs at identical accuracy. Recent versions enable it by default on Arm, but if your prefill numbers look about 20% low against reference tables, check this flag first. Always build Release: a debug build can be several times slower.llama.cpp for Android (fast baseline)
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build-android \
-DCMAKE_TOOLCHAIN_FILE=$ANDROID_NDK/build/cmake/android.toolchain.cmake \
-DANDROID_ABI=arm64-v8a -DANDROID_PLATFORM=android-28 \
-DCMAKE_BUILD_TYPE=Release -DBUILD_SHARED_LIBS=OFF -DGGML_OPENMP=OFF
cmake --build build-android -j$(nproc) --target llama-cli llama-bench
adb shell mkdir -p /data/local/tmp/lcpp
adb push build-android/bin/llama-cli build-android/bin/llama-bench /data/local/tmp/lcpp/
Vendor SDK environment (Qualcomm example)
# After unpacking QAIRT
export QNN_SDK_ROOT=/opt/qcom/aistack/qairt/<version>
source $QNN_SDK_ROOT/bin/envsetup.sh # sets PATH, PYTHONPATH, LD_LIBRARY_PATH
python $QNN_SDK_ROOT/bin/check-python-dependency # installs pinned Python deps
# AI Hub client (separate venv recommended)
python -m venv ~/venvs/aihub && source ~/venvs/aihub/bin/activate
pip install qai-hub qai-hub-models
qai-hub configure --api_token <YOUR_TOKEN>
qai-hub list-devices | head
Phone preparation for stable measurements
# Keep the screen on while plugged in, and set a fixed brightness
adb shell svc power stayon usb
adb shell settings put system screen_brightness_mode 0
adb shell settings put system screen_brightness 100
# Record the software build you measured on
adb shell getprop ro.build.fingerprint
# Create a scratch area on the device
adb shell mkdir -p /data/local/tmp/bench
pip freeze next to every result.Export and conversion paths
"Export" captures a model as a static, framework-independent graph; "conversion" turns that graph into the format a specific runtime executes. Most modern paths start from PyTorch. Below are the six paths you will meet most often, with working command and code patterns. Treat exact flag names as version-dependent and check the current documentation.
Exporting is like recording a live jazz performance to sheet music. The live performance (eager PyTorch) can improvise: it branches, loops and changes tempo depending on the audience. Sheet music (the exported graph) must fix every note in advance so any orchestra can play it. Conversion is then transcribing that sheet music for a particular band: a string quartet (CPU), a brass band (GPU) or a single virtuoso synthesiser (NPU) that only plays certain notes.
Improvisation is Python control flow and dynamic shapes; the notes the synthesiser cannot play are unsupported operators that fall back to the CPU.
PyTorch nn.Module (FP32/BF16)
|
+-------------+-------------+-+-----------+-------------+--------------+
| | | | | |
torch.onnx torch.export ai-edge-torch HF safetensors coremltools AI Hub /
.export | (torch.export convert_hf_to_ ct.convert QAIRT
| | inside) gguf.py | converters
.onnx ExecuTorch | | .mlpackage |
| to_edge + .tflite .gguf (f16) | QNN model
ORT quantize partitioner (LiteRT) llama-quantize Core ML + context
/ QNN EP | | | runtime binary .bin
| .pte LiteRT + Q4_K_M .gguf (ANE) |
ONNX Runtime ExecuTorch delegates llama.cpp QNN HTP /
Mobile runtime Genie
| Path | Artifact | Best for | Main gotcha |
|---|---|---|---|
| PyTorch to ONNX | .onnx (optionally .ort) | Cross-platform, ORT execution providers, many vendor converters accept ONNX | Opset versions, dynamic axes, exporter differences (TorchScript vs dynamo) |
| PyTorch to ExecuTorch | .pte | PyTorch-native on-device, LLMs, multi-backend | Graph breaks in torch.export; backend partition coverage |
| PyTorch to LiteRT | .tflite | Android apps, MediaPipe, Google accelerators | Layout transposes (NCHW to NHWC), op coverage in the converter |
| Hugging Face to GGUF | .gguf | LLM prototyping and distribution, CPU inference | Architecture must be supported by llama.cpp; tokenizer metadata |
| To QNN context binary | .bin context + libs | Snapdragon HTP at full efficiency | Static shapes, static quantization, SoC-specific binaries |
| To Core ML | .mlpackage | Apple Neural Engine | ANE op/shape constraints silently push work to GPU/CPU |
Path 1: PyTorch to ONNX (and ONNX Runtime quantization)
import torch, torchvision
model = torchvision.models.mobilenet_v3_small(weights="DEFAULT").eval()
dummy = torch.randn(1, 3, 224, 224)
torch.onnx.export(
model, (dummy,), "mobilenet_v3.onnx",
input_names=["image"], output_names=["logits"],
opset_version=17,
dynamic_axes=None, # keep static for NPUs; use {"image": {0: "batch"}} only if needed
do_constant_folding=True,
)
# Validate numerically against PyTorch before doing anything else
import onnxruntime as ort, numpy as np
sess = ort.InferenceSession("mobilenet_v3.onnx", providers=["CPUExecutionProvider"])
ref = model(dummy).detach().numpy()
out = sess.run(None, {"image": dummy.numpy()})[0]
print("max abs diff:", np.abs(ref - out).max()) # expect ~1e-5 for FP32
Static (QDQ) INT8 quantization with a calibration reader:
from onnxruntime.quantization import (quantize_static, CalibrationDataReader,
QuantFormat, QuantType, CalibrationMethod)
from onnxruntime.quantization.shape_inference import quant_pre_process
quant_pre_process("mobilenet_v3.onnx", "mobilenet_v3.pre.onnx") # shape inference + folding
class Reader(CalibrationDataReader):
def __init__(self, samples): # samples: list of float32 arrays, shape (1,3,224,224)
self.it = iter([{"image": s} for s in samples])
def get_next(self):
return next(self.it, None)
quantize_static(
"mobilenet_v3.pre.onnx", "mobilenet_v3.int8.onnx",
calibration_data_reader=Reader(calibration_samples[:300]),
quant_format=QuantFormat.QDQ, # QuantizeLinear/DequantizeLinear pairs; EPs fuse them
activation_type=QuantType.QUInt8, weight_type=QuantType.QInt8,
per_channel=True, calibrate_method=CalibrationMethod.MinMax,
)
# For the QNN EP, ORT provides helpers (qnn_preprocess_model, get_qnn_qdq_config)
# that produce a QDQ model with the uint8/uint16 activation types the HTP expects.
Optionally convert to the ORT format and build a reduced-operator runtime to shrink the mobile binary: python -m onnxruntime.tools.convert_onnx_models_to_ort model.int8.onnx produces a .ort file and a list of required operators you can feed into a custom minimal build.
Path 2: PyTorch to ExecuTorch (.pte)
Generic model with the XNNPACK CPU backend and PT2E INT8 quantization:
import torch
from torch.export import export
from executorch.exir import to_edge_transform_and_lower
from executorch.backends.xnnpack.partition.xnnpack_partitioner import XnnpackPartitioner
from executorch.backends.xnnpack.quantizer.xnnpack_quantizer import (
XNNPACKQuantizer, get_symmetric_quantization_config)
from torchao.quantization.pt2e.quantize_pt2e import prepare_pt2e, convert_pt2e
model = MyModel().eval()
example = (torch.randn(1, 3, 224, 224),)
# 1. Capture a graph suitable for quantization
graph = export(model, example).module()
# 2. Insert observers, calibrate, convert (post-training static quantization)
quantizer = XNNPACKQuantizer().set_global(get_symmetric_quantization_config(is_per_channel=True))
prepared = prepare_pt2e(graph, quantizer)
for x in calibration_batches: # ~100-500 representative inputs
prepared(x)
quantized = convert_pt2e(prepared)
# 3. Re-export, lower supported subgraphs to XNNPACK, serialize
program = to_edge_transform_and_lower(
export(quantized, example), partitioner=[XnnpackPartitioner()]
).to_executorch()
with open("model_xnnpack_int8.pte", "wb") as f:
f.write(program.buffer)
LLM export using the built-in LLM exporter (running example: Llama-3.2-1B-Instruct; you need consolidated.00.pth, params.json and tokenizer.model from the model repository):
python -m extension.llm.export.export_llm \
base.model_class="llama3_2" \
base.checkpoint="${LLAMA_DIR}/consolidated.00.pth" \
base.params="${LLAMA_DIR}/params.json" \
model.use_kv_cache=True \
model.use_sdpa_with_kv_cache=True \
model.dtype_override="fp32" \
base.metadata='"{\"get_bos_id\":128000, \"get_eos_ids\":[128009, 128001]}"' \
quantization.qmode="torchao:8da4w" \
quantization.group_size=128 \
quantization.embedding_quantize="torchao:4,32" \
backend.xnnpack.enabled=True \
export.max_seq_length=2048 \
export.output_name="llama3_2_1b_8da4w.pte"
8da4wmeans 8-bit dynamic activations and 4-bit weights;group_size=128gives one scale per 128 weights.embedding_quantize="4,32"quantizes the large embedding table (128256 x 2048 for this model, about a fifth of all parameters) to 4-bit with groups of 32.use_kv_cacheanduse_sdpa_with_kv_cacheare essential for speed. Without them every decode step recomputes attention over the whole sequence and numbers are badly wrong.- The BOS/EOS ids in metadata must match the tokenizer, or generation will not stop (or will stop immediately).
- Older ExecuTorch versions used
python -m examples.models.llama.export_llamawith flags like-kv --use_sdpa_with_kv_cache -X -qmode 8da4w; recognise both.
Swap XnnpackPartitioner for a vendor partitioner (Qualcomm QNN, MediaTek, Core ML, Vulkan) to target accelerators; the rest of the flow stays the same. That uniformity is ExecuTorch's main selling point.
Path 3: PyTorch to LiteRT via ai-edge-torch
import torch, torchvision
import ai_edge_torch # package name may be litert-torch in newer releases
model = torchvision.models.resnet18(weights="DEFAULT").eval()
sample = (torch.randn(1, 3, 224, 224),)
edge_model = ai_edge_torch.convert(model, sample) # uses torch.export under the hood
edge_model.export("resnet18_fp32.tflite")
# Check parity on host
import numpy as np
ref = model(*sample).detach().numpy()
out = edge_model(*sample)
print("max abs diff:", np.abs(ref - out).max())
# Optional: NHWC inputs to avoid transposes on GPU/NPU delegates
nhwc_model = ai_edge_torch.to_channel_last_io(model, args=[0])
edge_nhwc = ai_edge_torch.convert(nhwc_model, (torch.randn(1, 224, 224, 3),))
Quantization uses the PT2E flow with a PT2E quantizer for LiteRT (ai_edge_torch.quantize) before conversion, or post-conversion quantization for weight-only compression. For LLMs, ai-edge-torch's generative API includes re-authored model definitions and conversion scripts that emit prefill and decode signatures with an external KV cache, then bundle them into a .task or .litertlm file for MediaPipe / LiteRT-LM.
From TensorFlow/Keras, the classic converter still applies: tf.lite.TFLiteConverter.from_saved_model() with optimizations=[tf.lite.Optimize.DEFAULT], a representative_dataset generator for full-integer quantization, and TFLITE_BUILTINS_INT8 ops for integer-only NPU targets.
Path 4: Hugging Face checkpoint to GGUF (llama.cpp)
cd llama.cpp
pip install -r requirements.txt
# 1. Convert HF safetensors to a 16-bit GGUF (weights + tokenizer + metadata in one file)
python convert_hf_to_gguf.py /models/Llama-3.2-1B-Instruct \
--outfile llama-3.2-1b-f16.gguf --outtype f16
# 2. Optional: importance matrix from calibration text improves low-bit quality
./build/bin/llama-imatrix -m llama-3.2-1b-f16.gguf -f calib.txt -o imatrix.dat
# 3. Quantize
./build/bin/llama-quantize --imatrix imatrix.dat llama-3.2-1b-f16.gguf llama-3.2-1b-Q4_K_M.gguf Q4_K_M
./build/bin/llama-quantize llama-3.2-1b-f16.gguf llama-3.2-1b-Q4_0.gguf Q4_0 # Arm repack-friendly
# 4. Quality check on host: perplexity on held-out text
./build/bin/llama-perplexity -m llama-3.2-1b-Q4_K_M.gguf -f wiki.test.raw
# 5. On device
adb push llama-3.2-1b-Q4_0.gguf /data/local/tmp/lcpp/
adb shell "cd /data/local/tmp/lcpp && ./llama-bench -m llama-3.2-1b-Q4_0.gguf -p 128,512 -n 128 -t 4"
adb shell "cd /data/local/tmp/lcpp && ./llama-cli -m llama-3.2-1b-Q4_0.gguf -p 'Explain KV caching in one paragraph.' -n 128 -t 4 -no-cnv"
Q4_0is a simple 4-bit format with a scale per 32 weights; on recent Arm CPUs llama.cpp repacks it at load time into layouts that use i8mm/dotprod instructions, which is often the fastest CPU option on phones.Q4_K_Muses super-blocks with mixed 4/6-bit precision for sensitive tensors; better quality per byte, a popular default.- Thread count matters: on big.LITTLE phones, use the number of performance cores (often 4-6); using all cores can be slower because little cores become stragglers.
Path 5: To a QNN context binary (Snapdragon NPU)
Option A, via AI Hub (hosted compile, quantize and profile on real devices):
import torch, qai_hub as hub
model = MyModel().eval()
example = torch.randn(1, 3, 224, 224)
traced = torch.jit.trace(model, example)
device = hub.Device("Samsung Galaxy S24 (Family)")
# Compile straight to a QNN context binary for this SoC
compile_job = hub.submit_compile_job(
model=traced, device=device,
input_specs={"image": (1, 3, 224, 224)},
options="--target_runtime qnn_context_binary",
)
target_model = compile_job.get_target_model()
# Profile on a real device in the cloud: per-layer timing, compute unit per op, memory
profile_job = hub.submit_profile_job(model=target_model, device=device)
# Run inference on-device to check numerics against the PyTorch reference
infer_job = hub.submit_inference_job(model=target_model, device=device,
inputs={"image": [example.numpy()]})
out = infer_job.download_output_data()
target_model.download("model_ctx.bin")
For quantized models, submit a quantize job (INT8 or W4A16/W8A16 style activations) with calibration data before compiling, or export a pre-quantized ONNX (QDQ) model. For supported LLMs, the model zoo has ready export scripts that quantize and split the model into several context binaries:
python -m qai_hub_models.models.llama_v3_2_3b_instruct.export \
--device "Snapdragon 8 Elite QRD" \
--skip-inferencing --skip-profiling \
--output-dir ./genie_bundle_src
Option B, locally with the QAIRT tools. Binary names and flags are release-specific and are the most common source of copy-paste failures. Older QNN SDK releases ship qnn-onnx-converter, qnn-model-lib-generator and qnn-context-binary-generator. Newer QAIRT releases document qairt-converter / qairt-quantizer (exact names: check the SDK bin/ of the version you installed). The sketch below is the older converter shape so you can recognise it; copy flags from that release's HTML docs or --help, do not invent them.
# 1. Convert (and quantize with a calibration input list of raw tensor files)
# Flag names (--act_bitwidth, --weights_bitwidth, --input_list) are from older
# qnn-onnx-converter help text. Confirm on your SDK before running.
qnn-onnx-converter --input_network model.onnx \
--input_list calib_list.txt \
--act_bitwidth 16 --weights_bitwidth 8 \
--output_path model_qnn.cpp
# 2. Build a model library for the target
qnn-model-lib-generator -c model_qnn.cpp -b model_qnn.bin \
-t x86_64-linux-clang -o model_libs
# 3. Generate an HTP context binary (graph finalized for a specific SoC arch)
qnn-context-binary-generator \
--backend $QNN_SDK_ROOT/lib/x86_64-linux-clang/libQnnHtp.so \
--model model_libs/x86_64-linux-clang/libmodel_qnn.so \
--config_file htp_config.json \
--binary_file model_ctx
# 4. Run on device from the context binary (no on-device graph compile)
adb push model_ctx.bin $QNN_SDK_ROOT/lib/aarch64-android/libQnnHtp*.so \
$QNN_SDK_ROOT/lib/hexagon-v75/unsigned/libQnnHtpV75Skel.so \
$QNN_SDK_ROOT/bin/aarch64-android/qnn-net-run /data/local/tmp/qnn/
adb shell "cd /data/local/tmp/qnn && export LD_LIBRARY_PATH=. ADSP_LIBRARY_PATH=. && \
./qnn-net-run --backend libQnnHtp.so --retrieve_context model_ctx.bin \
--input_list inputs.txt --profiling_level basic"
Option C, via frameworks: the ExecuTorch Qualcomm backend (a QnnPartitioner plus QnnQuantizer, producing a .pte with embedded context binaries) or the ONNX Runtime QNN execution provider with EP context caching (covered in the Android integration section).
Path 6: To Core ML (Apple)
import torch, coremltools as ct
model = MyModel().eval()
example = torch.randn(1, 3, 224, 224)
traced = torch.jit.trace(model, example) # or torch.export program in newer coremltools
mlmodel = ct.convert(
traced,
inputs=[ct.TensorType(name="image", shape=example.shape)],
convert_to="mlprogram",
compute_units=ct.ComputeUnit.ALL, # CPU + GPU + Neural Engine
minimum_deployment_target=ct.target.iOS17,
compute_precision=ct.precision.FLOAT16,
)
# Weight compression: 8-bit linear or 4-bit palettization / per-block quantization
from coremltools.optimize.coreml import (OpLinearQuantizerConfig, OptimizationConfig,
linear_quantize_weights)
cfg = OptimizationConfig(global_config=OpLinearQuantizerConfig(mode="linear_symmetric"))
mlmodel = linear_quantize_weights(mlmodel, config=cfg)
mlmodel.save("MyModel.mlpackage")
Use Xcode's Core ML performance report to see which operations actually ran on the Neural Engine; ops that the ANE cannot run (certain shapes, dynamic dimensions, some activation types) fall back to GPU or CPU without any error.
Making a model exportable
- Remove Python-level dynamism Replace data-dependent
ifbranches withtorch.condor tensor operations; make loops fixed-length. - Fix shapes Pick concrete input sizes (and for LLMs, fixed prefill chunk sizes such as 128 plus a decode length of 1). Use
torch.export.Dimonly where the target runtime truly supports dynamic dimensions. - Replace unsupported ops Rewrite custom or exotic operators with standard ones (for example GELU-tanh approximation, explicit RMSNorm, attention written as matmul + softmax).
- Keep the KV cache explicit For LLMs, make the cache an input/output or registered buffer with a fixed maximum length rather than a growing Python list.
- Validate parity After every conversion, run the same inputs through the original and converted model; FP32 differences should be around 1e-5 to 1e-4.
torch.export), quantize with representative calibration data to the precision the HTP wants (INT8 or W4/W8 weights with 16-bit activations), compile to a context binary for the exact SoC, check per-op placement in a profile to confirm nothing falls back to CPU, validate numerics against FP32, then integrate via QNN directly, the ExecuTorch QNN backend, or ORT's QNN EP.Integrating into an Android app
A model file is not a feature. Integration means choosing where inference runs (Kotlin via a runtime's Java API, or native C++ via JNI), how the model gets onto the device and into memory, which accelerator executes it, and how the app behaves when things go wrong. For Android framework background (services, lifecycles, threading) see Android frameworks.
Integrating a model into an app is like installing a commercial espresso machine in a small cafe. You must get it through the door (packaging and download), connect it to the right power and water supply (delegates and native libraries), put it where it does not block the queue at the counter (off the UI thread), keep it warm between customers instead of cold-starting it every time (reuse the interpreter), and have a backup kettle if it breaks (CPU or cloud fallback).
The door is APK size limits, the power supply is the NPU/GPU delegate, the counter queue is the main thread, keeping it warm is session reuse and compiled-graph caching, and the backup kettle is the fallback chain.
Architecture of an on-device inference feature
+---------------------------- App process ------------------------------+
| UI (Compose/View) --▶ ViewModel --▶ InferenceRepository |
| | (single background dispatcher) |
| ▼ |
| Kotlin runtime API or JNI bridge ──▶ C++ engine |
| (LiteRT / ORT / (ExecuTorch, llama.cpp, QNN, |
| ExecuTorch / MediaPipe) custom pre/post-processing) |
| | |
| pre-processing ──▶ model ──▶ post-processing ──▶ result |
+-----------------------------------|------------------------------------+
▼
Delegate / backend: NPU (HTP, APU, TPU) | GPU (OpenCL/Vulkan) | CPU (XNNPACK/KleidiAI)
▼
Model file: mmap from app storage (downloaded) or uncompressed asset
Option 1: LiteRT Interpreter with GPU delegate and CPU fallback (Kotlin)
// build.gradle.kts
// Artifact IDs moved from org.tensorflow:tensorflow-lite* to com.google.ai.edge.litert:*.
// Confirm the current Maven coordinates and the Java package they export in the
// official LiteRT Android docs for the version you pin.
dependencies {
implementation("com.google.ai.edge.litert:litert:<version>")
implementation("com.google.ai.edge.litert:litert-gpu:<version>")
}
android {
androidResources { noCompress += listOf("tflite") } // required to mmap from assets
}
// Interpreter is the long-lived, widely documented path. Some LiteRT artifacts
// still export org.tensorflow.lite.* for compatibility; newer packages may not.
import org.tensorflow.lite.Interpreter
import org.tensorflow.lite.gpu.CompatibilityList
import org.tensorflow.lite.gpu.GpuDelegate
import java.io.FileInputStream
import java.nio.MappedByteBuffer
import java.nio.channels.FileChannel
class Classifier(context: Context) : AutoCloseable {
private var gpu: GpuDelegate? = null
private val interpreter: Interpreter
init {
val options = Interpreter.Options()
val compat = CompatibilityList()
if (compat.isDelegateSupportedOnThisDevice) {
gpu = GpuDelegate(compat.bestOptionsForThisDevice)
options.addDelegate(gpu)
} else {
options.setNumThreads(4) // XNNPACK CPU path
}
interpreter = Interpreter(mapAsset(context, "model_int8.tflite"), options)
}
private fun mapAsset(ctx: Context, name: String): MappedByteBuffer {
ctx.assets.openFd(name).use { fd ->
FileInputStream(fd.fileDescriptor).channel.use { ch ->
return ch.map(FileChannel.MapMode.READ_ONLY, fd.startOffset, fd.declaredLength)
}
}
}
// Reuse buffers: allocate once, fill per call
private val input = ByteBuffer.allocateDirect(1 * 224 * 224 * 3).order(ByteOrder.nativeOrder())
private val output = Array(1) { ByteArray(1000) }
fun classify(pixels: ByteArray): ByteArray {
input.rewind(); input.put(pixels)
interpreter.run(input, output) // call from a background thread only
return output[0]
}
override fun close() { interpreter.close(); gpu?.close() }
}
LiteRT Next also documents a CompiledModel-style API that binds a model to an accelerator (CPU, GPU, or an NPU compiler plugin) and can consume SoC-specific compiled artifacts. Class names, packages, accelerator enums and compile options change between releases. Check the current official LiteRT docs before writing production code; do not assume Interpreter.Options maps one-to-one onto CompiledModel, and do not copy guessed method names from older blog posts. The concepts stay the same: create once, choose an accelerator, reuse buffers, fall back gracefully.
Option 2: ONNX Runtime Mobile with the QNN execution provider
// build.gradle.kts: use the onnxruntime-android-qnn package for the QNN EP
dependencies { implementation("com.microsoft.onnxruntime:onnxruntime-android-qnn:<version>") }
import ai.onnxruntime.*
val env = OrtEnvironment.getEnvironment()
val opts = OrtSession.SessionOptions().apply {
// Java helper name (addQnn vs addConfigEntry) and option keys are
// version-specific. Confirm against the current ONNX Runtime QNN EP docs.
// Common documented keys include backend_path and an HTP performance mode.
addQnn(mapOf(
"backend_path" to "libQnnHtp.so",
"htp_performance_mode" to "burst"
// Other documented modes include sustained_high_performance; do not
// invent extra flags. Graph-finalization and context-cache keys also
// move between ORT releases — look them up for the AAR you ship.
))
addConfigEntry("ep.context_enable", "1")
addConfigEntry("ep.context_file_path", File(context.filesDir, "model_ctx.onnx").path)
addConfigEntry("session.disable_cpu_ep_fallback", "1")
}
val session = env.createSession(modelFile.path, opts)
val input = OnnxTensor.createTensor(env, floatBuffer, longArrayOf(1, 3, 224, 224))
session.run(mapOf("image" to input)).use { result ->
val logits = (result[0].value as Array<FloatArray>)[0]
}
- The QNN HTP backend requires a quantized (QDQ) model for NPU execution; FP32 models either use the HTP's FP16 path on newer chips or fall back.
- Turn off CPU fallback in development so partitioning problems surface as errors; turn it back on in production for robustness.
- EP context caching avoids re-finalizing the graph on every launch, which can take seconds for large models.
Option 3: MediaPipe LLM Inference API (fastest LLM integration)
// dependencies { implementation("com.google.mediapipe:tasks-genai:<version>") }
// Engine options vs session options (temperature / top-k / LoRA) were split
// across releases. Check the current MediaPipe Tasks GenAI docs and the
// LiteRT-LM notes; method names below are the long-lived engine-level ones.
import com.google.mediapipe.tasks.genai.llminference.LlmInference
val options = LlmInference.LlmInferenceOptions.builder()
.setModelPath(File(context.filesDir, "models/gemma.task").path)
.setMaxTokens(1024) // prompt + output budget: sizes the KV cache
.build()
val llm = LlmInference.createFromOptions(context, options) // expensive: do once, off main thread
val answer = llm.generateResponse("Summarise this note in two lines: ...")
llm.generateResponseAsync(prompt) { partial, done ->
uiFlow.tryEmit(partial)
if (done) onFinished()
}
This path only loads models converted into its bundle format (with tokenizer and metadata). It is ideal when a supported model family fits the product; for arbitrary architectures, use ExecuTorch or llama.cpp.
Option 4: ExecuTorch from Kotlin
// Packages and constructors move between ExecuTorch Android releases
// (Module vs ExecutorchModule, LlmModule constructor args). Confirm against
// the current ExecuTorch Android / LLM runner docs for the AAR you ship.
import org.pytorch.executorch.EValue
import org.pytorch.executorch.Module
import org.pytorch.executorch.Tensor
val module = Module.load(File(filesDir, "model_xnnpack_int8.pte").path) // mmap-based loader
val input = Tensor.fromBlob(floatArray, longArrayOf(1, 3, 224, 224))
val out = module.forward(EValue.from(input))[0].toTensor().dataAsFloatArray
import org.pytorch.executorch.extension.llm.LlmCallback
import org.pytorch.executorch.extension.llm.LlmModule
val llm = LlmModule(ptePath, tokenizerPath, /* temperature = */ 0.7f)
llm.load()
llm.generate(prompt, /* seqLen = */ 512, object : LlmCallback {
override fun onResult(token: String) { appendToUi(token) }
override fun onStats(stats: String) { Log.i("LLM", stats) }
})
Option 5: your own C++ engine through JNI
Native integration gives you full control (custom pre/post-processing, direct QNN or Genie calls, llama.cpp, zero-copy buffers). Keep the JNI surface small and coarse-grained: one call per request, not one per tensor.
// Kotlin side
object NativeEngine {
init { System.loadLibrary("edgeengine") }
external fun create(modelPath: String, backend: Int): Long // returns a handle
external fun generate(handle: Long, prompt: String, maxTokens: Int, cb: TokenCallback): String
external fun destroy(handle: Long)
}
fun interface TokenCallback { fun onToken(piece: String) }
// C++ side: engine_jni.cpp
#include <jni.h>
#include <memory>
#include <string>
#include "engine.h" // your wrapper around ExecuTorch / llama.cpp / QNN
extern "C" JNIEXPORT jlong JNICALL
Java_com_example_edge_NativeEngine_create(JNIEnv* env, jobject, jstring jpath, jint backend) {
const char* path = env->GetStringUTFChars(jpath, nullptr);
auto* engine = new Engine(path, static_cast<Backend>(backend)); // loads + mmaps weights
env->ReleaseStringUTFChars(jpath, path);
return reinterpret_cast<jlong>(engine);
}
extern "C" JNIEXPORT jstring JNICALL
Java_com_example_edge_NativeEngine_generate(JNIEnv* env, jobject, jlong h, jstring jprompt,
jint maxTokens, jobject cb) {
auto* engine = reinterpret_cast<Engine*>(h);
const char* p = env->GetStringUTFChars(jprompt, nullptr);
std::string prompt(p);
env->ReleaseStringUTFChars(jprompt, p);
jclass cls = env->GetObjectClass(cb);
jmethodID onToken = env->GetMethodID(cls, "onToken", "(Ljava/lang/String;)V");
std::string full;
engine->generate(prompt, maxTokens, [&](const std::string& piece) {
full += piece;
jstring js = env->NewStringUTF(piece.c_str());
env->CallVoidMethod(cb, onToken, js);
env->DeleteLocalRef(js); // avoid local-ref table overflow in long loops
});
return env->NewStringUTF(full.c_str());
}
extern "C" JNIEXPORT void JNICALL
Java_com_example_edge_NativeEngine_destroy(JNIEnv*, jobject, jlong h) {
delete reinterpret_cast<Engine*>(h);
}
# CMakeLists.txt (app/src/main/cpp)
cmake_minimum_required(VERSION 3.22)
project(edgeengine CXX)
set(CMAKE_CXX_STANDARD 17)
add_library(edgeengine SHARED engine_jni.cpp engine.cpp)
target_compile_options(edgeengine PRIVATE -O3 -march=armv8.2-a+dotprod+fp16)
# Link prebuilt runtime libs (ExecuTorch, llama, or QNN) imported as IMPORTED targets
target_link_libraries(edgeengine executorch_llm_runner log android)
- Token-by-token callbacks into Java cost microseconds each; fine for tens of tokens per second, but batch pieces if you stream very fast.
- Call JNI from a thread that was attached to the JVM; callbacks from a native worker thread need
AttachCurrentThread. - Vendor NPU libraries must be packaged in
jniLibs/arm64-v8a(or extracted to app storage) and the DSP library search path (ADSP_LIBRARY_PATH) must point to where the Hexagon skeleton libraries live; set it before initializing the backend. - Set
useLegacyPackagingorextractNativeLibsaccording to whether your runtime can load libraries directly from the APK.
Threading and lifecycle rules
- Never on the main thread Use a single dedicated background dispatcher or thread for inference; most interpreters are not thread-safe for concurrent calls.
- Create once, reuse Loading, delegate initialization and graph compilation can take hundreds of milliseconds to seconds. Keep the session in an application-scoped object.
- Warm up Run one dummy inference after load, while the user is not waiting, to trigger shader compilation and cache fills.
- Release on memory pressure Respond to
onTrimMemoryby freeing caches and, for large models, unloading entirely when backgrounded. - Cancel LLM generation must be cancellable (a flag checked every token) so leaving the screen stops work and power draw.
Model packaging and delivery
| Method | Size limit and behaviour | Good for | Watch out for |
|---|---|---|---|
| Bundled asset in the APK/AAB | Counts towards install size; update needs app update | Small models (under ~50-100 MB) | Must be stored uncompressed to mmap |
| Play Asset Delivery (install-time, fast-follow, on-demand packs) | Large packs served by the store | Medium-large models tied to app versions | Pack availability timing; handle "not yet downloaded" |
| Store-provided on-device AI packs (device-targeted delivery) | Different model variants per device class | Shipping NPU-specific binaries per SoC | Newer mechanism; check availability |
| Own download server / CDN | Any size; full control of versioning | Frequent model updates, A/B tests | Resume, integrity checks, storage checks, metered networks |
| System-provided model (OS AI service) | Nothing to ship | Supported devices and model families | No control of model or version; availability varies |
A robust self-managed download flow:
// Pseudocode for a WorkManager download worker
1. Fetch manifest: {modelId, version, url, sha256, sizeBytes, minRamMb, socAllowList, runtimeVersion}
2. Check eligibility: RAM tier, SoC, free storage (size x 2 for safety), unmetered network, charging
3. Download to files/models/.tmp/<id>-<version>.part with HTTP Range resume
4. Verify SHA-256 (and signature if the model is sensitive)
5. Atomically rename to files/models/<id>/<version>/model.pte
6. Update "active version" pointer only after a smoke-test inference succeeds
7. Keep the previous version until the new one has run successfully N times; then delete
Memory-mapped weights
Memory mapping (mmap) makes the model file appear as memory without copying it into the heap. Pages are read from storage on first touch and can be dropped by the kernel under pressure because they are clean and file-backed.
mmap (recommended default)
- Near-instant "load"; pages fault in lazily
- Clean file pages are reclaimable, so lower kill risk
- Shared between processes that map the same file
- Counts in RSS only when touched
- Requires uncompressed, page-aligned storage
Read into heap
- Full load cost up front; double memory during load if copying
- Anonymous memory: not reclaimable, counts fully against you
- Needed when the runtime must repack or transform weights
- Predictable latency once loaded
- Works with compressed or encrypted files
Two subtleties: first, if the runtime repacks weights into a kernel-friendly layout at load time (common for CPU kernels and GPU delegates), the repacked copy is anonymous memory and the mmap benefit shrinks; second, if pages are evicted during generation, decode stalls on storage reads, which appears as sudden latency spikes. For latency-critical paths, touch all pages after load (prefault) or use mlock where permitted, and budget memory honestly.
noCompress for model assets. A compressed asset cannot be memory-mapped, so the runtime silently copies the whole model into the Java or native heap, doubling peak memory at load and sometimes causing an out-of-memory crash only on low-RAM devices.onTrimMemory, gate the feature by device tier, and keep a remote kill switch.Project 1: baseline and benchmark harness
Duration: 2-3 weeks. Goal: an LLM running on your phone plus a measurement harness that every later project reuses. Skills: latency, memory and power profiling; token-rate measurement; PyTorch export tooling; accelerator basics. Output: a repository containing the harness and a sustained-load thermal analysis write-up.
A benchmark harness is like the timing system at an athletics track. A stopwatch in someone's hand gives one noisy number; a proper system has the same start line every time (fixed environment), a warm-up lap that is not timed (warm-up runs), dozens of heats (repeated runs), photo-finish percentiles rather than one lucky time (p50/p90/p99), and a long-distance event, not just the sprint (the thermal soak).
The start line is device state (charge, temperature, brightness), the warm-up lap is discarded initial runs, heats are repetitions, and the long-distance event is ten or more minutes of continuous generation.
What you build
Run Llama-3.2-1B-Instruct (or an even smaller model such as a 0.5-0.6B one for a quicker first win) on Android through ExecuTorch with the XNNPACK backend, and in parallel through llama.cpp as a cross-check. Then build a harness that records, on every run:
| Metric | Definition | Why it matters |
|---|---|---|
| Model load time (cold and warm) | From process start / file open to ready-for-first-inference; cold = after dropping caches or reboot | Dominates perceived start-up; almost nobody reports it |
| TTFT (time to first token) | Prompt submitted to first output token visible | Perceived responsiveness; grows with prompt length |
| Prefill throughput | Prompt tokens / prefill time | Compute-bound phase; accelerators help most here |
| Decode throughput | Generated tokens / decode time (excluding first token) | Memory-bandwidth-bound; the streaming speed users see |
| Per-token latency percentiles | p50/p90/p99 of inter-token time | Stutters are noticed even if the average is fine |
| Peak RSS / PSS | Max resident memory of the process | Decides whether the low-memory killer (lmkd) terminates you or other apps |
| Model file size | .pte/.gguf bytes on disk | Download and storage cost |
| Energy per request | Joules per 100 generated tokens or per inference | Battery impact; the fair way to compare CPU, GPU and NPU |
| Sustained / peak ratio | Throughput after 10 min divided by throughput in the first minute | What a shipped product actually experiences |
Step 1: export, push and run
# Export (see the ExecuTorch path above), then:
adb shell mkdir -p /data/local/tmp/llama
adb push llama3_2_1b_8da4w.pte /data/local/tmp/llama/
adb push tokenizer.model /data/local/tmp/llama/
adb push cmake-out-android/examples/models/llama/llama_main /data/local/tmp/llama/
adb shell "cd /data/local/tmp/llama && ./llama_main \
--model_path=llama3_2_1b_8da4w.pte \
--tokenizer_path=tokenizer.model \
--prompt='Explain KV caching in one paragraph.' \
--seq_len=256 --cpu_threads=4 --warmup=1"
# The runner prints prompt/generated token counts and timing stats
# (model load, first token latency, prefill and generation rates).
For an instruct model, wrap the prompt in the model's chat template (for Llama 3.x: <|begin_of_text|><|start_header_id|>user<|end_header_id|> ... <|eot_id|><|start_header_id|>assistant<|end_header_id|>). A missing template is the most common cause of rambling or empty output.
Step 2: define a benchmark protocol
- Fix the environment Same device, OS build, airplane mode, fixed screen brightness (or screen off via a CLI run), battery 50-90% or external supply, device at room temperature with the case removed, no other apps, same prompt set.
- Cool down between configurations Wait until the SoC thermal zone returns near its idle temperature (for example within 2 degrees of baseline) before the next configuration.
- Warm-up Discard 1-3 full generations (or 10-50 runs for small models) so caches, page faults and kernel selection settle. Record load time and first-run latency separately.
- Repeat At least 10 generations for LLMs, 100-1000 inferences for small models; report median and p90/p99, plus standard deviation.
- Vary prompt length Measure prefill at several lengths (for example 64, 256, 1024 tokens) because attention cost and cache behaviour change with length.
- Record everything Device, SoC, build fingerprint, runtime version, model hash, quantization config, thread count, ambient temperature, start/end temperatures.
Step 3: the harness (host-side driver)
#!/usr/bin/env python3
"""bench.py - drive an on-device LLM runner over adb and collect metrics into CSV."""
import csv, json, re, statistics, subprocess, time, datetime
DEV_DIR = "/data/local/tmp/llama"
RUNNER = f"cd {DEV_DIR} && ./llama_main --model_path={{pte}} --tokenizer_path=tokenizer.model " \
"--prompt=\"{prompt}\" --seq_len={seq} --cpu_threads={threads}"
def adb(cmd: str) -> str:
return subprocess.run(["adb", "shell", cmd], capture_output=True, text=True, timeout=900).stdout
def soc_temp_c() -> float:
# pick the zone(s) that represent CPU/SoC on your device; names vary by vendor
out = adb("for z in /sys/class/thermal/thermal_zone*; do echo $(cat $z/type) $(cat $z/temp); done")
temps = [int(t) / 1000 for name, t in (l.split() for l in out.splitlines() if l.strip())
if re.search(r"cpu|soc|tsens", name, re.I) and t.lstrip("-").isdigit()]
return max(temps) if temps else float("nan")
def wait_cool(target_c: float, timeout_s=600):
t0 = time.time()
while soc_temp_c() > target_c and time.time() - t0 < timeout_s:
time.sleep(10)
def parse_stats(text: str) -> dict:
# Adapt these regexes to your runner's output format
pats = {
"load_ms": r"Model load time:\s*([\d.]+)",
"ttft_ms": r"(?:Time to first generated token|first token).*?([\d.]+)",
"prefill_tps": r"Prompt evaluation:.*?([\d.]+)\s*tokens/s",
"decode_tps": r"Generated \d+ tokens:.*?([\d.]+)\s*tokens/s",
}
return {k: float(m.group(1)) if (m := re.search(p, text, re.S)) else None for k, p in pats.items()}
def peak_rss_kb(proc_name="llama_main") -> int:
out = adb(f"pid=$(pidof {proc_name}); [ -n \"$pid\" ] && grep VmHWM /proc/$pid/status")
m = re.search(r"(\d+)", out)
return int(m.group(1)) if m else -1
def run_config(pte, prompt, seq=256, threads=4, warmup=2, reps=10, idle_c=None):
if idle_c: wait_cool(idle_c + 2)
cmd = RUNNER.format(pte=pte, prompt=prompt, seq=seq, threads=threads)
for _ in range(warmup):
adb(cmd)
rows = []
for i in range(reps):
t_start = soc_temp_c()
text = adb(cmd)
s = parse_stats(text)
s.update(rep=i, temp_start=t_start, temp_end=soc_temp_c(), pte=pte, threads=threads)
rows.append(s)
return rows
def summarize(rows, key):
xs = sorted(r[key] for r in rows if r.get(key) is not None)
if not xs: return {}
pct = lambda p: xs[min(len(xs) - 1, int(round(p / 100 * (len(xs) - 1))))]
return {"p50": pct(50), "p90": pct(90), "p99": pct(99),
"mean": statistics.mean(xs), "stdev": statistics.pstdev(xs)}
if __name__ == "__main__":
meta = {"fingerprint": adb("getprop ro.build.fingerprint").strip(),
"soc": adb("getprop ro.soc.model").strip(),
"time": datetime.datetime.now().isoformat()}
idle = soc_temp_c()
rows = run_config("llama3_2_1b_8da4w.pte", "Explain KV caching in one paragraph.", idle_c=idle)
with open("results.csv", "w", newline="") as f:
w = csv.DictWriter(f, fieldnames=rows[0].keys()); w.writeheader(); w.writerows(rows)
print(json.dumps({"meta": meta, **{k: summarize(rows, k)
for k in ("ttft_ms", "prefill_tps", "decode_tps")}}, indent=2))
Note: VmHWM in /proc/<pid>/status is the peak resident set ("high water mark"); poll it while the process runs or print it from the runner itself at exit, because it disappears with the process. For app-based runs, dumpsys meminfo <package> gives PSS broken down into native heap, graphics, and file mappings.
Step 4: in-app timing (C++ side)
#include <chrono>
#include <vector>
#include <algorithm>
using Clock = std::chrono::steady_clock; // monotonic; never use wall clock
struct GenStats { double load_ms, ttft_ms, prefill_tps, decode_tps; std::vector<double> itl_ms; };
GenStats timed_generate(Engine& e, const std::vector<int>& prompt, int max_new) {
GenStats s{};
auto t0 = Clock::now();
e.prefill(prompt); // fills KV cache, returns logits
int tok = e.sample();
auto t1 = Clock::now();
s.ttft_ms = std::chrono::duration<double, std::milli>(t1 - t0).count();
s.prefill_tps = prompt.size() / (s.ttft_ms / 1000.0);
auto prev = t1;
for (int i = 1; i < max_new && tok != e.eos(); ++i) {
e.decode(tok);
tok = e.sample();
auto now = Clock::now();
s.itl_ms.push_back(std::chrono::duration<double, std::milli>(now - prev).count());
prev = now;
}
double decode_s = std::chrono::duration<double>(prev - t1).count();
s.decode_tps = s.itl_ms.size() / decode_s;
return s;
}
double percentile(std::vector<double> v, double p) {
std::sort(v.begin(), v.end());
size_t idx = static_cast<size_t>(p / 100.0 * (v.size() - 1) + 0.5);
return v[idx];
}
Step 5: power and energy
Energy is what users feel as battery drain. Three levels of rigour:
| Method | How | Accuracy |
|---|---|---|
| Fuel gauge sampling | Sample current_now and voltage_now every 100-500 ms during the run; integrate V x I over time; subtract an idle baseline measured the same way | Coarse (fuel gauges average and update slowly), but fine for relative comparisons over long runs |
| batterystats | dumpsys batterystats --reset, run the workload unplugged, then dump and inspect estimated drain per UID | Model-based estimates; good for app-level attribution |
| On-device power rails (ODPM) via Perfetto | Record the android.power data source; many recent devices expose per-rail energy counters (CPU big/mid/little, GPU, memory, NPU/DSP rails vary) | Best on-device option; per-subsystem breakdown |
| External power monitor | Power the phone from a monitor (battery bypass on dev boards, or USB meter for rough numbers) | Most accurate total power; needs hardware |
# Fuel gauge sampling (units are usually microamps / microvolts; sign convention varies by device)
adb shell "while true; do echo $(date +%s%N) $(cat /sys/class/power_supply/battery/current_now) \
$(cat /sys/class/power_supply/battery/voltage_now); sleep 0.2; done" > power.log
# batterystats
adb shell dumpsys batterystats --reset
adb shell dumpsys battery unplug # simulate unplugged so stats accumulate over USB
# ... run workload ...
adb shell dumpsys batterystats > batterystats.txt
adb shell dumpsys battery reset
Step 6: the thermal soak (the differentiating measurement)
Most published numbers are a single cold run. Instead, generate continuously for 10-20 minutes and plot throughput, temperature, CPU frequency and power over time. This reveals when throttling begins, how far throughput falls, and the sustained-to-peak ratio, which is the number that actually matters for a shipped feature. For the thermal framework itself see Power and thermal.
# thermal_soak.sh - run generations back-to-back and sample system state every 2 s
DUR=${1:-600}
adb shell "cd /data/local/tmp/llama; end=\$((\$(date +%s)+$DUR)); \
while [ \$(date +%s) -lt \$end ]; do ./llama_main --model_path=llama3_2_1b_8da4w.pte \
--tokenizer_path=tokenizer.model --prompt='Write a long story about a lighthouse.' \
--seq_len=512 --cpu_threads=4 2>&1 | grep -i 'tokens/s'; done" > soak_tps.log &
while kill -0 $! 2>/dev/null; do
ts=$(date +%s)
temps=$(adb shell "for z in /sys/class/thermal/thermal_zone*; do echo -n \$(cat \$z/type)=\$(cat \$z/temp),; done")
freqs=$(adb shell "cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_cur_freq | tr '\n' ,")
status=$(adb shell "dumpsys thermalservice | grep -m1 'Thermal Status'")
echo "$ts;$temps;$freqs;$status" >> soak_sys.log
sleep 2
done
decode tok/s
45 |■■■■■■■■■■■
40 | ■■■■■ peak (first 60-120 s)
35 | ■■■■■
30 | ■■■■■■■■■■■■■■■■■■■■■ sustained plateau
25 |
+-----+-----+-----+-----+-----+-----+-----+--▶ time (min)
0 1 2 3 4 5 6 10
SoC temp rises ~15-25 °C; big-core frequency steps down; ratio sustained/peak ≈ 0.6-0.8 (typical, device-dependent)
Step 7: compare BF16 and INT4
Export the same model twice (BF16/FP32 and 8da4w INT4) and fill a comparison table on all metrics: load time, TTFT, prefill tok/s, decode tok/s, peak RSS, file size, energy per 100 tokens and sustained ratio. Expect roughly 2-3x faster decode, 3-4x faster prefill (with KleidiAI), about 50-60% smaller file and 30-40% lower peak memory for the quantized model. Deviations beyond that mean something is misconfigured (see the target numbers section).
Deliverables checklist
- Model exported and running on device with coherent output.
- Harness scripted so a full measurement run is one command, producing a CSV plus a summary.
- BF16 vs INT4 comparison across all metrics.
- Ten-minute sustained-load curve with temperature and frequency overlay.
- Numbers sanity-checked against reference values.
- Short write-up of what surprised you, with the raw data committed.
Project 2: quantization sweep and layer-wise error analysis
Duration: 3-4 weeks. Goal: understand quantization numerically and build the tool that finds where it breaks. Skills: PTQ vs QAT, weight and activation quantization, mixed precision, layer-wise mismatch, host reference vs on-device validation, fixed-point representation. Output: a layer-wise activation diff tool plus a sweep results table. For the theory of scales, zero-points and granularity, see Edge AI.
Finding quantization error layer by layer is like finding a leak in a long water pipeline. Measuring only at the tap (final accuracy) tells you water is missing, not where. Installing a pressure gauge at every junction (per-layer comparison against the FP32 reference) shows exactly which segment loses pressure, whether the loss compounds downstream, and which few segments need thicker pipe (higher precision) instead of replacing the entire line.
Junctions are layer boundaries, gauges are cosine similarity and SQNR, the leaky segment is a quantization-sensitive layer, and thicker pipe is keeping that layer in INT8 or FP16 while the rest stays INT4.
Part A: the configuration sweep
Take one model (Llama-3.2-1B is ideal) and walk the configuration space, measuring both quality and cost every time. Record every result in one CSV; the sweep table is itself a useful artifact.
| Variable | Values to try | What you learn |
|---|---|---|
| Weight bit-width | 4-bit, 8-bit | The primary size and decode-speed lever |
| Group size | 32, 64, 128, 256, per-channel | Granularity vs scale-metadata overhead (a 16-bit scale per 32 weights adds 0.5 bits per weight) |
| Activation quantization | none (weight-only), dynamic 8-bit, static 8-bit, static 16-bit | Why NPUs demand static quantization, and its accuracy cost |
| Embedding quantization | off, 8-bit, 4-bit group 32 | Embeddings are a large share of small models (about 21% for Llama-3.2-1B) |
| LM head (output projection) precision | 4, 6, 8 bits | The output layer is unusually sensitive: errors land directly on logits |
| Algorithm | round-to-nearest, GPTQ, AWQ, SpinQuant (rotations), QAT (+LoRA) | How much accuracy smarter PTQ or training recovers |
| Calibration set | size (32-512 samples), domain match | Sensitivity of static ranges to data |
Measure quality two ways: perplexity on held-out text (statistical degradation) and at least one task-level score (for example a small multiple-choice benchmark, or exact-match on your product's prompts) for behavioural degradation. Perplexity can move very little while a task collapses, and vice versa.
INT8 versus INT4, what to expect before you sweep. INT8 (or FP8) weights are about 2× smaller than FP16 and usually near-lossless after PTQ; that is the control row. INT4 weight-only is about 4× smaller and is the decode-speed lever, but quality loss is larger and shows first on reasoning, maths, rare languages and long context. Keep embeddings and the LM head at 6-8 bits in the first INT4 recipe. Two "4-bit" rows that differ in group size or algorithm are not comparable; write the full recipe on every CSV line.
# Sweep driver sketch (ExecuTorch exporter; one row per config)
for bits in 4 8; do for gs in 32 128 256; do for emb in "none" "torchao:8,0" "torchao:4,32"; do
name="l32_1b_w${bits}_g${gs}_emb${emb//[:,]/_}"
python -m extension.llm.export.export_llm base.model_class=llama3_2 \
base.checkpoint=$CKPT base.params=$PARAMS model.use_kv_cache=True \
model.use_sdpa_with_kv_cache=True quantization.qmode="torchao:8da${bits}w" \
quantization.group_size=$gs quantization.embedding_quantize="$emb" \
backend.xnnpack.enabled=True export.output_name="$name.pte"
python eval_ppl.py --pte $name.pte --data heldout.txt >> sweep.csv # host-side quality
python bench.py --pte $name.pte >> sweep.csv # device-side cost
done; done; done
Part B: layer-wise error analysis (the core skill)
Analysing quantization error and layer-wise mismatches is a named responsibility in many edge AI roles, and very few candidates can demonstrate it. The method:
- Reference run Run the FP32 model on the host (x86 or Arm workstation) and capture the output of every layer boundary (attention projections, attention output, MLP gate/up/down, norms, residual stream after each block).
- Quantized run Run the quantized model on the same inputs and capture the same tensors: first as a simulated (fake-quantized) model on the host, then from the device.
- Diff per tensor Compute cosine similarity, mean and max absolute error, and SQNR for each captured tensor.
- Rank and plot Rank layers by divergence; plot SQNR across depth to see whether error compounds or recovers.
- Isolate Quantize one layer at a time (everything else FP32) to measure each layer's individual sensitivity; this separates "this layer is fragile" from "this layer receives bad inputs".
- Derive a recipe Keep the most sensitive layers at higher precision, re-run the sweep and the Project 1 harness, and quantify quality recovered versus latency and memory cost.
Host-side tool: capture and diff with forward hooks
import torch, math, csv
from collections import OrderedDict
def capture(model, inputs, names_filter=lambda n, m: isinstance(m, torch.nn.Linear)
or "norm" in n.lower() or n.endswith("layers")):
acts, hooks = OrderedDict(), []
for name, mod in model.named_modules():
if names_filter(name, mod):
hooks.append(mod.register_forward_hook(
lambda m, i, o, name=name: acts.__setitem__(name, (o[0] if isinstance(o, tuple) else o)
.detach().float().cpu())))
with torch.no_grad():
model(*inputs)
for h in hooks: h.remove()
return acts
def metrics(ref, q):
ref, q = ref.flatten(), q.flatten()
err = ref - q
sqnr = 10 * math.log10((ref.pow(2).sum() / err.pow(2).sum().clamp_min(1e-20)).item())
cos = torch.nn.functional.cosine_similarity(ref, q, dim=0).item()
return {"sqnr_db": sqnr, "cosine": cos,
"mae": err.abs().mean().item(), "max_abs": err.abs().max().item(),
"ref_absmax": ref.abs().max().item()} # large absmax hints at outliers
ref_acts = capture(fp32_model, (tokens,))
q_acts = capture(fake_quant_model, (tokens,)) # e.g. torchao-quantized copy
rows = [{"layer": n, **metrics(ref_acts[n], q_acts[n])} for n in ref_acts if n in q_acts]
rows.sort(key=lambda r: r["sqnr_db"]) # worst first
with open("layerwise.csv", "w", newline="") as f:
w = csv.DictWriter(f, fieldnames=rows[0].keys()); w.writeheader(); w.writerows(rows)
for r in rows[:10]:
print(f'{r["layer"]:50s} SQNR {r["sqnr_db"]:6.1f} dB cos {r["cosine"]:.5f}')
Single-layer sensitivity sweep (the "isolate" step):
def per_layer_sensitivity(build_model, layer_names, quantize_only, eval_fn):
base = eval_fn(build_model()) # FP32 score (e.g. perplexity)
out = []
for name in layer_names:
m = quantize_only(build_model(), {name}) # quantize a single layer
out.append((name, eval_fn(m) - base)) # degradation attributable to this layer
return sorted(out, key=lambda t: -t[1])
Capturing intermediate tensors on the device
| Stack | How to get per-layer outputs |
|---|---|
| ExecuTorch | Generate an ETRecord at export time and run with ETDump plus a debug buffer enabled; the devtools Inspector aligns on-device intermediate outputs with the original graph nodes so you can compute numeric gaps per operator. |
| QNN / QAIRT | qnn-net-run --debug writes every intermediate tensor; the SDK's accuracy debugger tooling compares them against a framework reference layer by layer. |
| ONNX Runtime | onnxruntime.quantization.qdq_loss_debug augments a model to output intermediate tensors and matches FP32 and QDQ activations (create_activation_matching, compute_activation_error). |
| AIMET | QuantAnalyzer reports per-layer sensitivity, activation/weight ranges, and the effect of enabling quantizers one at a time. |
| LiteRT | The quantization debugger reports per-layer statistics between float and quantized models; or add intermediate outputs to the model signature. |
| llama.cpp | Evaluation callbacks (the eval-callback example) print per-tensor statistics during a forward pass; compare against the f16 GGUF. |
Comparing host-simulated quantization with the device run is itself informative. If fake-quant on the host matches FP32 well but the device diverges, the problem is not the quantization scheme but the device implementation: different rounding, accumulator width, requantization, fused operator behaviour, or an FP16 overflow in a GPU/NPU kernel.
Typical findings to look for
Outlier channels
A few hidden dimensions in transformer activations carry values 10-100x larger than the rest. Per-tensor activation scales then waste most of the integer range. Signs: huge ref_absmax, low SQNR at inputs to the next projection. Fixes: per-channel or per-group, SmoothQuant-style scale migration, rotations (SpinQuant/QuaRot), or 16-bit activations.
Down projection and LM head
The MLP down projection sees the widest activation ranges; the LM head writes directly into logits. Both are common candidates to keep at 8-bit when everything else is 4-bit.
First and last blocks
Early layers shape the residual stream; late layers are close to the output. Sensitivity is often U-shaped across depth.
Norms and softmax
RMSNorm, softmax and residual additions are precision-sensitive; most stacks keep them at FP16/FP32 or 16-bit integer even in "INT8" models.
Error accumulation
SQNR of the residual stream often degrades gradually with depth, while individual layer outputs may recover; residual connections both carry and dilute error.
Attention vs MLP
Q/K projections affect attention scores exponentially through softmax; V and O projections behave more linearly. Different sensitivities justify different precisions.
Mixed-precision fallback recipe
- Start uniform Everything at the target precision (for example W4 group 32/128, A8 dynamic).
- Rank Use the single-layer sensitivity list, not just the full-model diff.
- Promote greedily Move the top-k layers to W8 (or FP16) one at a time until the quality target is met.
- Price it Each promoted layer costs size and decode time in proportion to its parameter share; report the Pareto curve of quality vs size vs latency.
- Validate on device Re-run the Project 1 harness and the device-side diff; mixed precision can split graphs on NPUs if the backend cannot run both precisions in one partition.
Also explore
- SpinQuant and QAT + LoRA: the two quantization-aware recipes with published Llama 3.2 numbers. Compare their quality against your PTQ results at equal bit-width.
- GPTQ / AWQ: weight-only PTQ that uses calibration data to minimise output error or protect salient channels.
- AIMET techniques: cross-layer equalisation, bias correction and AdaRound for CNN-style models on NPUs.
- Operator fusion: inspect the exported graph before and after lowering (for example with a graph visualiser or the printed edge program) and understand which ops fused and why fused quantized kernels change numerics slightly.
- Fixed-point representation: verify by hand how an INT8 x INT8 multiply accumulates into INT32 and requantizes with a multiplier and shift; this is exactly what NPU kernels do.
Deliverables checklist
- Full sweep table: quality, latency, size and memory for every configuration.
- Layer-wise diff tool working for host reference vs host fake-quant vs device.
- Ranked list of the most sensitive layers with a hypothesis for each.
- Evidence-derived mixed-precision recipe validated on the Project 1 harness.
- PTQ vs rotation-based vs QAT comparison at equal bit-width.
Project 3: NPU deployment on vendor silicon
Duration: 4-5 weeks. Goal: a model running on the NPU (Snapdragon Hexagon HTP as the worked example) integrated into an Android app through the native C API. Skills: vendor SDK (QAIRT/QNN), HTP deployment, static quantization, model splitting, operator support gaps. Output: working NPU inference plus an honest CPU vs GPU vs NPU comparison including power. This is the hardest project and the most valuable, because few engineers outside chipset vendors have done a full NPU deployment. Expect SDK version mismatches, unsupported operators and cryptic errors; debugging them is the curriculum.
An NPU is like a high-speed bottling plant. It fills thousands of identical bottles per minute with astonishing efficiency, but only if every bottle is the same size (static shapes), the recipe is fixed before the shift starts (static quantization, compiled graph), and every step is one the machines know (supported operators). Hand it an odd-shaped bottle and it stops the line, sends the bottle to a person at the side table (CPU fallback), waits, and restarts. A few odd bottles per batch and the plant is slower than just having people do it all by hand.
Bottle size is tensor shape, the fixed recipe is the context binary with baked-in scales, the side table is the CPU, and the stop-start cost is the data transfer and synchronisation at every partition boundary.
The constraint that changes everything: static quantization
The Hexagon NPU is built around fixed-point arithmetic with quantization parameters known at compile time. For LLMs the typical scheme is 4-bit (or 8-bit) weights with 16-bit static activations (W4A16 / W8A16); for CNNs, full INT8. On the CPU path in Projects 1-2, activations were quantized dynamically: scales computed at runtime from each tensor's actual range. The NPU cannot afford a data-dependent range computation per tensor per step, so ranges must come from calibration data. If production inputs exceed the calibrated range, values clip. Understanding why the NPU needs this, and measuring what it costs in accuracy, is the core lesson of the project.
| Target | Weights | Activations | Granularity |
|---|---|---|---|
| CPU / GPU (ExecuTorch XNNPACK, ONNX Runtime, LiteRT) | 4-bit group-wise (or 8-bit) | 8-bit dynamic, scale computed at runtime | Group size 32 or 128 |
| Qualcomm NPU (QNN / QAIRT) | 4-bit or 8-bit | 16-bit static, parameters fixed at compile time (8-bit for many CNNs) | Per-channel in QAIRT; ExecuTorch QNN delegate commonly uses block/group 32 for 4-bit |
| llama.cpp (Q4_0 on CPU) | 4-bit, group 32 (output layer often 6-bit) | Quantized on the fly to 8-bit for dot products, backend-dependent | Blocks of 32 |
Step 1: prepare the model (AI Hub path)
pip install qai-hub-models
qai-hub configure --api_token <YOUR_TOKEN>
# Export a supported LLM to QNN context binaries, quantized and split for the NPU
python -m qai_hub_models.models.llama_v3_2_3b_instruct.export \
--device "Snapdragon 8 Elite QRD" \
--skip-inferencing --skip-profiling \
--output-dir ./export_8_elite
# If you produced a custom quantized checkpoint in Project 2, pass it through
# with the export script's checkpoint option instead of the default weights.
For the 1B model, use the corresponding 1B export if available in the zoo, or go through the ExecuTorch Qualcomm backend, which has its own Llama export scripts that produce a .pte with embedded HTP context binaries.
Why the model is split into several binaries
- The NPU process has limits on how much memory a single graph and its mapped weights can use; large models are therefore partitioned by layers into several context binaries executed in sequence.
- Prefill and decode are compiled as separate graphs with different static shapes (for example 128-token chunks for prefill, 1 token for decode) that share weights ("weight sharing") so memory is not doubled.
- The KV cache is managed outside the graphs as input/output tensors of fixed maximum context length; the runtime updates it between calls.
Step 2: assemble the Genie bundle
mkdir -p genie_bundle
# 1. Model config for Genie (start from the sample config for your model in the vendor tutorials)
cp <tutorial-configs>/llama_v3_2_3b.json genie_bundle/genie_config.json
cp <tutorial-configs>/htp_backend_ext_config.json genie_bundle/
# 2. Hexagon DSP-side libraries: MUST match the SoC architecture
cp "$QNN_SDK_ROOT"/lib/hexagon-v73/unsigned/* genie_bundle/ # Snapdragon 8 Gen 2
# cp "$QNN_SDK_ROOT"/lib/hexagon-v75/unsigned/* genie_bundle/ # 8 Gen 3
# cp "$QNN_SDK_ROOT"/lib/hexagon-v79/unsigned/* genie_bundle/ # 8 Elite
# 3. Android (ARM) side libraries and the CLI runner
cp "$QNN_SDK_ROOT"/lib/aarch64-android/* genie_bundle/
cp "$QNN_SDK_ROOT"/bin/aarch64-android/genie-t2t-run genie_bundle/
# 4. Context binaries and tokenizer from the export step
cp export_8_elite/*.bin export_8_elite/tokenizer.json genie_bundle/
# 5. Verify: every file listed in ctx-bins[] of genie_config.json exists in the bundle
An abridged Genie dialog config. Field names (ctx-bins, kv-dim, sampler keys) vary between SDK versions. Always start from the sample JSON shipped for your model in that QAIRT / AI Hub export, then change only paths and context length.
{
"dialog": {
"type": "basic",
"context": { "size": 4096, "n-vocab": 128256, "bos-token": 128000, "eos-token": [128001, 128009] },
"sampler": { "seed": 42, "temp": 0.7, "top-k": 40, "top-p": 0.95 },
"tokenizer": { "path": "tokenizer.json" },
"engine": {
"n-threads": 3,
"backend": {
"type": "QnnHtp",
"QnnHtp": { "use-mmap": true, "poll": true, "kv-dim": 128 },
"extensions": "htp_backend_ext_config.json"
},
"model": {
"type": "binary",
"binary": { "ctx-bins": ["model_part_1_of_3.bin", "model_part_2_of_3.bin", "model_part_3_of_3.bin"] }
}
}
}
}
Step 3: run on device
adb push genie_bundle /data/local/tmp/
adb shell
cd /data/local/tmp/genie_bundle
export LD_LIBRARY_PATH=$PWD # ARM-side libraries
export ADSP_LIBRARY_PATH=$PWD # DSP-side (Hexagon skeleton) libraries
./genie-t2t-run -c genie_config.json \
-p "<|begin_of_text|><|start_header_id|>user<|end_header_id|>
Explain NPU static quantization.<|eot_id|><|start_header_id|>assistant<|end_header_id|>
"
Step 4: integrate the C API into an app
Genie wraps the tokenizer, the QNN backend, KV-cache management, decoding and sampling behind a dialog API; the engine executes forward passes on the HTP (with CPU for the rest). Wrapping it in JNI gives you the complete path: PyTorch checkpoint to quantized model to context binaries to NPU to Android UI. The C symbols below match the public dialog-API shape (JSON config, create, query with a callback, free). Header names, status codes, callback signatures and JSON keys are release-specific — copy them from the QAIRT samples for the version you link. Do not invent extra flags or config fields.
#include "GenieDialog.h"
#include <fstream>
#include <sstream>
#include <string>
struct Ctx { std::string out; void (*on_piece)(const char*, void*); void* user; };
static void on_response(const char* piece, const GenieDialog_SentenceCode_t code, const void* ud) {
auto* c = static_cast<Ctx*>(const_cast<void*>(ud));
if (piece) { c->out += piece; if (c->on_piece) c->on_piece(piece, c->user); }
(void)code; // BEGIN / CONTINUE / END / COMPLETE / ABORT
}
class NpuLlm {
GenieDialogConfig_Handle_t cfg_ = nullptr;
GenieDialog_Handle_t dlg_ = nullptr;
public:
bool init(const std::string& config_path) {
std::ifstream f(config_path); std::stringstream ss; ss << f.rdbuf();
if (GenieDialogConfig_createFromJson(ss.str().c_str(), &cfg_) != GENIE_STATUS_SUCCESS) return false;
return GenieDialog_create(cfg_, &dlg_) == GENIE_STATUS_SUCCESS; // loads context binaries
}
std::string ask(const std::string& prompt, void (*cb)(const char*, void*), void* user) {
Ctx c{{}, cb, user};
GenieDialog_query(dlg_, prompt.c_str(), GENIE_DIALOG_SENTENCE_COMPLETE, on_response, &c);
return c.out;
}
~NpuLlm() { if (dlg_) GenieDialog_free(dlg_); if (cfg_) GenieDialogConfig_free(cfg_); }
};
Before creating the dialog inside an app, set ADSP_LIBRARY_PATH (with setenv) to the directory containing the Hexagon skeleton libraries (for example the app's native library directory or an extracted folder), and ship the ARM-side QNN libraries in jniLibs/arm64-v8a.
Step 5: the honest three-way comparison
| Backend | Measure | Typical pattern (same 1-3B model) |
|---|---|---|
| CPU (XNNPACK + KleidiAI or llama.cpp) | TTFT, prefill, decode, RSS, power, sustained curve | Universal fallback; decent decode, weakest prefill, highest power, throttles first |
| GPU (OpenCL/Vulkan, e.g. Adreno) | Same | Good prefill; decode similar to CPU (both bandwidth-bound); competes with UI rendering |
| NPU (Hexagon HTP) | Same plus accuracy delta from static quantization | Much faster prefill (often several times), comparable or better decode, lowest energy per token, best sustained ratio |
Report energy per 100 tokens and the ten-minute curve, not just peak tok/s. Vendors publish peak numbers; independent, power-inclusive, sustained comparisons are rare and far more useful.
NPU deployment gotchas (and fixes)
| Symptom | Likely cause | Fix |
|---|---|---|
| NPU slower than CPU | Graph split into many partitions; unsupported op in the middle forces NPU-to-CPU round trips | Profile per-op placement; rewrite or replace the op; fuse; move pre/post-processing out of the graph |
| "Op not supported" or silent CPU fallback | Op, data type, rank, or attribute not in the backend op set (for example 5-D tensors, certain gather/scatter, dynamic slicing, int64) | Check the backend's supported-ops table; decompose into supported ops; cast int64 indices to int32 |
| Compilation fails with dynamic shapes | NPU compilers require static shapes | Export fixed shapes; separate prefill and decode graphs; pad inputs to buckets |
| Multi-second first load | Graph compiled/finalized on device at startup | Use precompiled context binaries or runtime context caching (ORT EP context, LiteRT compilation cache) |
| Library load or "unable to open session" errors | Hexagon library version does not match SoC; ADSP_LIBRARY_PATH wrong; SDK version of runtime libs differs from the one that built the context binary | Match hexagon-vXX to the SoC; keep build and runtime SDK versions identical; check logcat for FastRPC errors |
| Accuracy worse than CPU INT8 | Static activation ranges clip; 8-bit activations too coarse for transformers | Better calibration data; 16-bit activations; per-channel weights; keep sensitive ops in FP16 where supported |
| Out-of-memory on the NPU side | Graph or weights exceed the NPU session memory limits | Split the model into multiple context binaries; weight sharing between graphs; smaller context length |
| Latency jitter | NPU power/clock mode set to power-saver; DSP shared with other clients (camera, audio) | Set performance mode (burst or sustained high performance) appropriately; avoid contention; measure with other workloads active |
| Great benchmark, poor app | Pre/post-processing on CPU dominates; copies between Java and native buffers | Zero-copy shared buffers (ION/dmabuf, rpcmem), native pre-processing, batching |
Graph partitioning, visualised
Ideal (1 partition): [ NPU: conv..attn..mlp..conv ................ ] 1 dispatch
Bad (5 partitions): [NPU][CPU:op X][NPU][CPU:op Y][NPU]
▲copy+sync ▲copy+sync ▲copy+sync ▲copy+sync
Each boundary costs data transfer, format conversion (quantize/dequantize, layout)
and a synchronisation round trip; with many small partitions the NPU sits idle.
HTP context caching and performance settings
- Context binary: the serialized, finalized graph for one SoC architecture. Loading it skips graph optimisation and is typically 10-100x faster than compiling from the model at startup. It is not portable across Hexagon versions, and often not across SDK versions.
- On-device caching: if you must compile on device (for example via ORT QNN EP), write the compiled context to app storage on first run and reuse it; invalidate the cache when the model, SDK or OS/driver version changes.
- Performance modes: burst gives maximum clocks for short tasks; sustained high performance gives a stable profile for long workloads; power saver trades latency for energy. Match to the duty cycle.
- VTCM and spill-fill: the HTP has fast on-chip memory (VTCM); graphs that exceed it spill to DDR. Tiling and model splitting keep working sets on-chip.
Deliverables checklist
- Model exported to context binaries (via AI Hub, QAIRT or ExecuTorch QNN backend).
- CLI runner generating coherent output on the NPU.
- C API integrated into an Android app via JNI.
- Three-way CPU/GPU/NPU comparison with power and sustained curves.
- Documented accuracy cost of static vs dynamic activation quantization.
- A written log of every operator gap and SDK issue encountered, with the workaround. This log is gold in interviews.
Project 4: KV-cache engineering
Duration: 3-4 weeks. Goal: own the memory bottleneck that limits on-device long context. Skills: KV caching, KV-cache quantization, footprint profiling, attention internals. Output: a KV-cache profiler plus a quantization study across bit-widths and context lengths. The KV cache grows linearly with sequence length and at long contexts it exceeds the model weights; for long-context edge deployment, compressing it is often more impactful than squeezing weights further. Yet most tutorials focus only on weight quantization. For attention mechanics, see Transformers and LLMs.
The KV cache is like the minutes of a long meeting. Every new speaker (token) must be able to glance back at everything said so far, so the secretary writes down a summary card (key) and the content (value) for each contribution. Rereading the whole recording each time would be absurd, so the cards are kept. But the pile of cards grows with every sentence, and in a long meeting it can outweigh the rulebook on the table (the model weights). You can write cards in shorthand (quantized KV), keep only the last hour plus the opening remarks (sliding window with attention sinks), file cards in fixed-size folders instead of one giant stack (paging), and reuse the standard agenda cards every meeting (prefix caching).
Cards are K and V vectors per token per layer, the rulebook is the weights, shorthand is INT8/INT4 KV, the opening remarks are attention-sink tokens, folders are fixed-size blocks, and the agenda is the shared system prompt.
The memory math
Worked example with Llama-3.2-1B (from its params.json: dim 2048, 16 layers, 32 query heads, 8 KV heads, so head dim = 2048 / 32 = 64):
per token (FP16) = 2 x 16 layers x 8 kv_heads x 64 head_dim x 2 bytes = 32,768 B = 32 KiB
per token (INT8) = 16 KiB (+ small scale overhead)
per token (INT4) = 8 KiB (+ scales)
Llama-3.2-3B: 28 layers, 24 query heads, 8 KV heads, head_dim 128
per token (FP16) = 2 x 28 x 8 x 128 x 2 = 114,688 B = 112 KiB
A 7B model WITHOUT GQA (32 layers, 32 KV heads, head_dim 128)
per token (FP16) = 2 x 32 x 32 x 128 x 2 = 524,288 B = 512 KiB (16x the 1B model)
| Context | 1B FP16 KV | 1B INT4 KV | 3B FP16 KV | 3B INT4 KV | 7B MHA FP16 KV |
|---|---|---|---|---|---|
| 512 | 16 MiB | 4 MiB | 56 MiB | 14 MiB | 256 MiB |
| 2k | 64 MiB | 16 MiB | 224 MiB | 56 MiB | 1 GiB |
| 4k | 128 MiB | 32 MiB | 448 MiB | 112 MiB | 2 GiB |
| 8k | 256 MiB | 64 MiB | 896 MiB | 224 MiB | 4 GiB |
| 16k | 512 MiB | 128 MiB | 1.75 GiB | 448 MiB | 8 GiB |
| 32k | 1 GiB | 256 MiB | 3.5 GiB | 896 MiB | 16 GiB |
| 128k | 4 GiB | 1 GiB | 14 GiB | 3.5 GiB | 64 GiB |
INT4 weights are about 0.7-1.1 GiB for the 1B model and roughly 1.7-2.3 GiB for the 3B (depending on how embeddings and the output layer are quantized). Crossover (FP16 cache = INT4 weights) is weight bytes / 32 KiB per token for the 1B: about 22k tokens at 0.7 GiB and about 35k at 1.1 GiB. For the 3B (112 KiB/token) it is about 16-20k tokens. A 7B without GQA (512 KiB/token, ~3.5 GiB INT4 weights) crosses below 8k. Those are exactly the context lengths long-document and chat-history features want. Also remember that NPU and many CPU runtimes allocate the cache statically for the maximum context length, so you pay the full cost even for short prompts.
Part A: profile it
- Footprint table Measure actual RSS (not just the formula) at 512, 1k, 2k, 4k, 8k and 16k context for the 1B and 3B models; the difference from the formula reveals allocator overhead and duplicated buffers.
- Crossover point Find the context length where cache bytes exceed weight bytes.
- Speed impact Plot decode tok/s against current context length: each decode step reads the whole cache, so decode slows as the conversation grows.
- Allocation pattern Trace allocations over a long generation (heapprofd in Perfetto, or malloc hooks): growth, fragmentation and reallocation stalls when a contiguous cache is resized.
- Kill threshold Find the context length at which the low-memory killer starts terminating your process on an 8 GB device with a real foreground workload.
def kv_bytes(layers, kv_heads, head_dim, tokens, bytes_per=2, batch=1, group=None, scale_bytes=2):
base = 2 * layers * kv_heads * head_dim * tokens * batch * bytes_per
if group: # per-group scales for quantized KV
base += 2 * layers * kv_heads * (head_dim // group) * tokens * batch * scale_bytes
return base
for T in (512, 1024, 2048, 4096, 8192, 16384):
fp16 = kv_bytes(16, 8, 64, T) / 2**20
int4 = kv_bytes(16, 8, 64, T, bytes_per=0.5, group=32) / 2**20
print(f"{T:6d} fp16 {fp16:7.1f} MiB int4(g32) {int4:7.1f} MiB")
Why decode slows as context grows
Part B: compress and manage it
| Technique | How it works | What to measure |
|---|---|---|
| KV quantization (8 / 4 / 3 bits) | Store K and V as low-bit integers with per-group scales; dequantize inside attention | Memory saved vs perplexity and task quality; where it falls apart. INT8 is usually near-lossless; well-designed 4-bit and even ~3-bit schemes can have small loss |
| Per-channel keys, per-token values | Keys have outlier channels (a few dimensions consistently large), so quantize keys along channels; values are better quantized per token | Verify on your model by plotting K and V magnitude per channel; compare both groupings |
| Sliding-window attention | Keep only the last W tokens per layer (ring buffer); memory bounded at W | Quality on inputs longer than W; retrieval of facts from early in the context |
| Attention sinks | Always keep the first few tokens plus the window; models place large attention mass on initial tokens, and dropping them destabilises generation | Perplexity on very long streams with and without sinks |
| Paged / block-allocated cache | Fixed-size blocks (for example 16-64 tokens) mapped by a block table; no large contiguous reallocations; blocks shareable between sequences. Contiguous waste ≈ (Tmax − Tactual) × KV per token; paged waste ≤ one block | Fragmentation, allocation stalls, memory overhead of the last partially filled block |
| Eviction policies | Drop tokens that receive little attention (heavy-hitter style), keep recent and sink tokens | What can be dropped with least damage; stability across tasks |
| Prefix (prompt) caching | Compute KV for a byte-identical system prompt or shared document once, persist it, reuse for every request. Prefill savings ≈ P / (P + U) | TTFT saved; storage cost; invalidation when model or prompt changes; a timestamp at the top of the prompt misses every time |
| Speculative decoding | A tiny draft (or extra heads / prompt n-grams) proposes k tokens; the target verifies them in one pass. E[tokens per pass] = (1 − αk+1) / (1 − α); speed-up ≈ E / (1 + k · c) | Draft memory, rejection rollback, extra power, weaker gains on high-entropy text and when the NPU needs a static k-token graph |
| Architecture choices | GQA/MQA, cross-layer KV sharing, smaller head dims | Only when you can choose or fine-tune the model |
Code: per-token INT8 value cache and ring-buffer window
import torch
def quant_per_token(x, bits=8): # x: [heads, T, head_dim]
qmax = 2 ** (bits - 1) - 1
scale = x.abs().amax(dim=-1, keepdim=True).clamp_min(1e-8) / qmax
q = torch.clamp(torch.round(x / scale), -qmax - 1, qmax).to(torch.int8)
return q, scale.to(torch.float16)
def quant_per_channel(x, bits=8): # keys: scale per head_dim channel
qmax = 2 ** (bits - 1) - 1
scale = x.abs().amax(dim=-2, keepdim=True).clamp_min(1e-8) / qmax
q = torch.clamp(torch.round(x / scale), -qmax - 1, qmax).to(torch.int8)
return q, scale.to(torch.float16)
class RingKV:
"""Sliding window + attention sinks, fixed memory."""
def __init__(self, layers, heads, head_dim, window=2048, sinks=4, dtype=torch.float16):
self.W, self.S = window, sinks
shape = (layers, heads, sinks + window, head_dim)
self.k = torch.zeros(shape, dtype=dtype); self.v = torch.zeros(shape, dtype=dtype)
self.n = 0 # total tokens seen
def slot(self):
if self.n < self.S: return self.n # sinks are never overwritten
return self.S + (self.n - self.S) % self.W # ring over the window
def append(self, layer, k_t, v_t): # k_t, v_t: [heads, head_dim]
s = self.slot()
self.k[layer, :, s] = k_t; self.v[layer, :, s] = v_t
def advance(self): self.n += 1
def valid(self): return min(self.n, self.S + self.W)
# Note: with RoPE, positions of cached keys are already baked in; implementations that
# roll a window either keep original positions or re-assign positions within the cache.
Code sketch: paged KV allocator in C++
struct BlockPool {
size_t block_tokens, bytes_per_block;
std::vector<void*> free_list;
explicit BlockPool(size_t tokens, size_t bpt, size_t n_blocks)
: block_tokens(tokens), bytes_per_block(tokens * bpt) {
for (size_t i = 0; i < n_blocks; ++i) free_list.push_back(aligned_alloc(64, bytes_per_block));
}
void* get() { if (free_list.empty()) return nullptr; auto* b = free_list.back(); free_list.pop_back(); return b; }
void put(void* b) { free_list.push_back(b); }
};
struct SequenceKV {
std::vector<void*> blocks; // block table: logical block i -> physical block
size_t tokens = 0;
bool append_token(BlockPool& pool) {
if (tokens % pool.block_tokens == 0) {
void* b = pool.get();
if (!b) return false; // caller evicts, compresses, or stops generation
blocks.push_back(b);
}
++tokens; return true;
}
};
// Pre-allocating the pool once at startup avoids reallocation stalls and makes the
// memory budget explicit; the attention kernel walks the block table.
Part C: the Android systems finding
Tie the work to device reality. On a real 8 GB phone with a foreground app and background services, what is the maximum usable context length before the system kills your process, and how does each compression technique move that number? This requires understanding Android's memory management (see Linux kernel and BSP and Power and thermal for related platform detail):
- lmkd (the userspace low-memory killer daemon) monitors memory pressure (PSI, pressure stall information) and kills processes in order of
oom_score_adj: cached apps first, then services, then perceptible and finally foreground. - A foreground app is killed last, but a large model can push the system into killing everything else (including the launcher or music playback), which users experience as "the phone got weird".
- Anonymous memory (heap-allocated KV cache, repacked weights) cannot be reclaimed without swap/zRAM compression; file-backed mmapped weights can be dropped and re-read.
- zRAM compresses anonymous pages; low-entropy tensors compress well, quantized ones do not.
# Observe memory pressure and kills while increasing context length
adb shell cat /proc/pressure/memory # PSI: some/full stall percentages
adb shell dumpsys meminfo -s <package> # PSS summary for the app
adb logcat -b events | grep -i am_kill # framework kill events
adb logcat | grep -i lowmemorykiller # lmkd decisions and reasons
adb shell cat /proc/<pid>/oom_score_adj
Deliverables checklist
- KV footprint curve vs context length for two model sizes, formula vs measured.
- Documented crossover point where the cache exceeds the weights.
- Quantization study at 8/4/3 bits with quality cost quantified, including key vs value grouping.
- At least one alternative strategy implemented (sliding window with sinks, or paged).
- Maximum usable context under real Android memory pressure, with lmkd behaviour.
n_kv_heads from the config.Project 5: the system project
Duration: 4-6 weeks. Goal: a complete system where the model is one component among many, built to survive memory pressure, heat and bad inputs. Skills: LoRA adapters, system integration, memory- and thermal-aware inference, platform architecture. Output: a working system, a short demo, and an architecture write-up of the trade-offs. Pick one of the four designs below; each reuses the harness (P1), quantization recipe (P2), NPU path (P3) and cache management (P4). For wider design framing, see GenAI system design.
The earlier projects built and tuned an engine; the system project builds the whole car. A great engine in a car with no brakes, no fuel gauge and no cooling system is not drivable. The system project adds the cooling (thermal-aware throttling), the fuel gauge (memory budgeting), the brakes (cancellation and refusal), the gearbox (NPU, GPU, CPU fallback) and the dashboard (telemetry).
The engine is the model runtime; the car is the service or app around it that users actually operate.
Option A: a shared on-device inference service
Build a small version of what a platform AI service is: one Android service hosting one model, shared by several client apps.
Client app A ─┐ ┌─ LoRA: summarise (≈20-40 MB)
Client app B ─┼─ AIDL (bound service, permission) ─▶ │ Scheduler / queue (fairness, priorities)
Client app C ─┘ │ Policy: memory + thermal governor
│ Engine: one resident base model (mmap)
│ NPU ──▶ GPU ──▶ CPU ──▶ refuse
└─ LoRA: rewrite / classify (hot-swap)
- Stable IPC surface: an AIDL interface with request id, prompt, options and a streaming callback; enforce a custom permission; validate inputs; cap sizes (see Binder transaction limits: stream results rather than returning large payloads).
- One model resident, not one copy per client: memory is the scarcest resource.
- Memory-pressure awareness: listen to
onTrimMemoryand PSI; shrink the KV budget, unload adapters, or unload the model when idle and under pressure. - Thermal-aware throttling: poll
PowerManager.getThermalHeadroom()and thermal status listeners; lower the generation rate or max tokens proactively instead of being throttled arbitrarily by the OS. - Hot-swappable LoRA adapters: fine-tune two small adapters for different tasks (see Fine-tuning), swap at runtime, and measure the switch cost in latency and memory. Merged weights are fastest but not swappable; unmerged adapters add a small matmul per adapted layer.
- Queueing and fairness: per-client quotas, priority for foreground callers, cancellation when the client dies (
linkToDeath). - Graceful degradation: NPU, then GPU, then CPU, then refuse with a clear error, following a documented policy.
// IInferenceService.aidl
package com.example.edge;
import com.example.edge.IInferenceCallback;
interface IInferenceService {
int getApiVersion();
long submit(String task, String prompt, in Bundle options, IInferenceCallback cb); // returns request id
void cancel(long requestId);
Bundle getStatus(); // backend in use, thermal state, loaded adapters, queue depth
}
// IInferenceCallback.aidl
oneway interface IInferenceCallback {
void onToken(long requestId, String piece);
void onComplete(long requestId, in Bundle stats);
void onError(long requestId, int code, String message);
}
// Thermal- and memory-aware policy (Kotlin, inside the service)
class Governor(private val pm: PowerManager) {
fun decide(req: Request): Plan {
val headroom = pm.getThermalHeadroom(10) // forecast 10 s ahead; 1.0 = throttling
val status = pm.currentThermalStatus
return when {
status >= PowerManager.THERMAL_STATUS_SEVERE -> Plan.Refuse("device too hot")
headroom > 0.85f -> Plan.Run(backend = Backend.NPU, maxTokens = req.maxTokens / 2,
tokenDelayMs = 20) // pace generation
memoryTight() -> Plan.Run(backend = Backend.NPU, maxContext = 1024)
else -> Plan.Run(backend = Backend.NPU, maxTokens = req.maxTokens)
}
}
}
Option B: an on-device assistant with local RAG (documents or device logs)
A fully local assistant that ingests documents (or device logs such as a bug report or logcat capture) and answers questions with citations. Nothing leaves the device.
Ingest ──▶ Parse/clean ──▶ Chunk (200-500 tokens, overlap) ──▶ Embed (small encoder, INT8)
│
Vector index (flat / HNSW) + keyword index (BM25)
Query ──▶ Embed ──▶ Hybrid retrieve top-k ──▶ Rerank (optional) ──▶ Pack context within budget
│
Local LLM (NPU, KV budget from P4) ──▶ Answer with cited chunk ids ──▶ UI
- Chunking: for logs, chunk on natural boundaries (timestamps, process ids, crash blocks) rather than fixed characters, and keep metadata (time, tag, pid) for filtering.
- Embeddings: a small sentence encoder (tens of millions of parameters) quantized to INT8 runs in milliseconds per chunk; batch ingestion in the background while charging.
- Index: a flat index is fine up to tens of thousands of chunks; beyond that use an approximate index (HNSW) in SQLite or a small native library. Combine with keyword search, because identifiers and error codes are poorly captured by embeddings.
- Context budget: a 40 MB log does not fit a 4-8k context; retrieval plus filtering is mandatory. Budget tokens explicitly: system prompt (prefix-cached) + retrieved chunks + question + answer.
- Grounding and refusal: instruct the model to answer only from retrieved text with chunk citations, and to say it cannot find evidence otherwise; verify citations exist before displaying.
- Measure the whole pipeline: ingestion time per MB, retrieval recall@k on a labelled question set, TTFT and total answer latency, peak memory, energy per question.
Option C: a wearable sensor model
An always-on activity, gesture or anomaly classifier on accelerometer/gyroscope/PPG data for a watch or band, where the budget is milliwatts, not watts.
- Model: a tiny 1-D CNN or small recurrent model (tens of KB to a few hundred KB), INT8, with fixed windows (for example 2-second windows at 50 Hz, 50% overlap).
- Placement: run on the sensor hub / low-power DSP or microcontroller (for example a Cortex-M with a micro NPU via a micro runtime), and wake the application processor only on detected events. Batch sensor data with hardware FIFOs to avoid waking the CPU every sample.
- Budget: energy per inference in microjoules, average power in low milliwatts, RAM in tens to hundreds of KB, flash for weights.
- Evaluation: per-user variation matters; evaluate with leave-one-subject-out splits, and measure false wake-ups per hour, which dominate battery cost.
- Updates: over-the-air model updates via the companion phone, with version checks and rollback.
Option D: a real-time camera pipeline
Detection, segmentation or super-resolution on live camera frames at 30 fps, where the budget is a 33 ms frame period and the model shares the device with the camera stack and display.
Camera (YUV_420_888) ──▶ GPU/ISP resize + colour convert ──▶ NPU model (INT8, static 256-640 px)
│ │
└──── preview surface (no copy) ◀── overlay render ◀── post-process (NMS, masks) on GPU/CPU
- Zero-copy: use hardware buffers (
AHardwareBuffer) shared between camera, GPU and NPU where the runtime supports it; each CPU copy of a 1080p frame costs milliseconds. - Pipelining: overlap capture, pre-processing, inference and rendering across frames; drop frames rather than queueing them (latency matters more than throughput).
- Pre/post-processing: often slower than the model itself; move resize/colour conversion to GPU or ISP and NMS to optimized native code.
- Thermal: camera plus display plus NPU is a heavy sustained load; measure 20-30 minutes, and adapt resolution or frame rate to thermal headroom.
- Metrics: end-to-end frame latency (sensor to overlay), p99 frame time, dropped-frame rate, accuracy on device-captured test footage (not only curated datasets).
Deliverables for any option
- System runs end to end on device, reproducibly.
- Behaviour under memory pressure documented (not just the happy path).
- Thermal behaviour under sustained use documented.
- A short demo recording.
- Architecture write-up explaining the design trade-offs and why.
Profiling tools and how to use them
Benchmark numbers tell you what is slow; profilers tell you why. Profile at three levels: the system (CPU scheduling, frequencies, thermal, power), the runtime (per-operator timings and placement), and the native code (hot functions and cache behaviour).
Profiling is like diagnosing a slow delivery company. The fleet dashboard (system trace) shows which trucks were moving, idle or stuck in traffic; the per-depot logs (operator profile) show which warehouse took longest to load; and riding along with one driver (sampling profiler) shows that he stops at every red light because the route planner is wrong.
Trucks are CPU cores and accelerators, traffic is thermal throttling and contention, depots are operators or partitions, and the ride-along is simpleperf on the native engine.
| Tool | Level | Use it to answer |
|---|---|---|
| Perfetto / Android System Trace | System | Which cores ran my threads, at what frequency; did the thermal governor cap clocks; where did the UI thread block; power rails over time |
| heapprofd (Perfetto) | Memory | Which native call stacks allocated the peak memory; KV cache growth; leaks |
| simpleperf | Native code | Hot functions in the engine; whether optimized kernels (dotprod, i8mm, SVE) are used; cache misses via hardware counters |
| Android GPU Inspector (AGI) | GPU | GPU utilisation, shader occupancy, memory bandwidth counters for GPU delegate workloads |
| Snapdragon Profiler | SoC (Qualcomm) | Real-time CPU, GPU, DSP/NPU and memory bandwidth metrics on Snapdragon devices |
QNN / QAIRT profiling (--profiling_level, qnn-profile-viewer) | NPU runtime | Per-op/per-layer HTP execution time, graph init time, DDR spill |
| AI Hub profile jobs | NPU runtime (hosted) | Per-layer timing and compute unit assignment (NPU/GPU/CPU) for each op on real devices |
| ExecuTorch devtools (ETDump + ETRecord + Inspector) | Runtime | Per-operator and per-delegate timings mapped back to the source graph; intermediate outputs |
LiteRT benchmark_model --enable_op_profiling | Runtime | Per-op latency and which ops ran on the delegate |
onnxruntime_perf_test and ORT profiling | Runtime | Per-node timings and EP assignment |
llama-bench | Runtime (LLM) | Prefill and decode tok/s across thread counts, quant types, prompt lengths |
| Arm Performance Studio (Streamline) | CPU/GPU (Arm) | Hardware counters on Arm cores and Mali GPUs |
| Xcode Instruments / Core ML report | Apple | Which ops ran on the Neural Engine; load and prediction times |
Recording a Perfetto trace with CPU frequency, scheduling and power rails
# config.pbtxt
buffers { size_kb: 131072 fill_policy: RING_BUFFER }
data_sources { config { name: "linux.ftrace" ftrace_config {
ftrace_events: "sched/sched_switch"
ftrace_events: "power/cpu_frequency"
ftrace_events: "power/cpu_idle"
ftrace_events: "thermal/thermal_temperature"
atrace_categories: "gfx"
atrace_categories: "view"
atrace_apps: "com.example.edge"
} } }
data_sources { config { name: "android.power" android_power_config {
battery_poll_ms: 250
collect_power_rails: true
battery_counters: BATTERY_COUNTER_CURRENT
battery_counters: BATTERY_COUNTER_VOLTAGE
} } }
data_sources { config { name: "linux.process_stats" target_buffer: 0 } }
duration_ms: 60000
# record (the config is streamed through stdin)
adb push config.pbtxt /data/local/tmp/
adb shell "cat /data/local/tmp/config.pbtxt | perfetto --txt -c - -o /data/misc/perfetto-traces/edge.pftrace"
adb pull /data/misc/perfetto-traces/edge.pftrace
# open the file in the Perfetto UI (runs locally in a browser; the trace stays on your machine)
Add custom slices so the trace shows model phases alongside system activity:
// Kotlin
Trace.beginSection("llm_prefill"); engine.prefill(tokens); Trace.endSection()
// C++ (NDK)
#include <android/trace.h>
ATrace_beginSection("llm_decode_step");
engine.decode(tok);
ATrace_endSection();
simpleperf on the native engine
adb shell simpleperf record -p $(adb shell pidof llama_main) -g --duration 15 \
-o /data/local/tmp/perf.data
adb shell simpleperf report -i /data/local/tmp/perf.data --sort symbol | head -40
# Hardware counters (availability varies by device)
adb shell simpleperf stat -e cpu-cycles,instructions,cache-misses,raw-l2d-cache-refill \
-p $(adb shell pidof llama_main) --duration 10
# For app processes, the NDK's simpleperf scripts (app_profiler.py) handle symbols and flame graphs.
QNN per-layer profiling
adb shell "cd /data/local/tmp/qnn && ./qnn-net-run --backend libQnnHtp.so \
--retrieve_context model_ctx.bin --input_list inputs.txt \
--profiling_level detailed --output_dir out"
adb pull /data/local/tmp/qnn/out/qnn-profiling-data_0.log
qnn-profile-viewer --input_log qnn-profiling-data_0.log # per-op cycles/time, init vs execute
A profiling workflow that finds real problems
- Start at the system trace Is the work actually on the expected unit? Are threads on big cores? Are frequencies capped (thermal) or low (governor idle)? Is anything else competing?
- Check placement In the runtime profile, confirm what fraction of ops or time runs on the accelerator and how many partitions exist.
- Find the top ops Typically matmuls and attention dominate; if a reshape, transpose, gather or quantize/dequantize is near the top, it is a conversion artifact to fix.
- Roofline sanity check For decode, compare achieved bytes/s (weights + KV per token x tok/s) with the device's memory bandwidth; if you are far below, the kernels or threading are the problem, not the hardware.
- Drill into native code Only then use simpleperf to see whether the expected SIMD kernels are running.
Debugging accuracy drift
Accuracy drift is when the deployed model's outputs differ from the reference by more than expected. It can come from conversion, quantization, the device implementation, or the pipeline around the model. The key is to bisect systematically instead of guessing.
Debugging accuracy drift is like tracing why a photocopy of a photocopy looks bad. You compare each generation against the original: was the first copy already blurred (conversion), did one machine use the wrong paper (pre-processing), was the toner low (quantization), or was the last machine miscalibrated (device kernel)? Comparing adjacent generations tells you exactly which step introduced the damage.
Each generation is one stage in the pipeline, and comparing adjacent generations is the bisection between FP32 original, converted FP32, fake-quant host, and real device.
The bisection ladder
Stage Compare against Expect ───────────────────────────────────────────────────────────────────────── 1 Original framework model (FP32) ground truth labels baseline accuracy 2 Exported / converted FP32 model stage 1 outputs max abs diff ~1e-5..1e-4 3 Converted FP16 model stage 2 small diff, no inf/NaN 4 Quantized model, run on host stage 2 small accuracy delta 5 Quantized model, run on device stage 4 near-identical outputs 6 Full app pipeline on device stage 5 with same inputs identical if pre/post match The first rung where the diff jumps is where the bug lives.
Common causes by rung
| Rung | Cause | How to confirm |
|---|---|---|
| Conversion (2) | Op semantics differ (for example GELU approximation, align_corners in resize, epsilon in norms), wrong opset, fused op bug | Layer-wise diff FP32 vs FP32 finds the first divergent op |
| FP16 (3) | Overflow in large activations (FP16 max is 65504), underflow in small ones, accumulation in FP16 | Look for inf/NaN or saturated values; keep those ops in FP32; check accumulator precision setting |
| Quantization (4) | Bad calibration data, outliers, per-tensor scales, sensitive layers | Project 2 tooling: SQNR per layer, single-layer sensitivity |
| Device (5) | Different rounding or saturation in kernels, requantization details, NPU-specific precision, driver bugs | Same quantized model host vs device, layer by layer; test on another device or backend |
| Pipeline (6) | Wrong colour order (RGB vs BGR), normalisation, resize method, NCHW vs NHWC, image rotation from camera, wrong tokenizer or chat template, different sampling settings | Dump the exact input tensor the model receives in the app and feed it to the host model |
LLM-specific drift checks
- Greedy decode comparison: run with temperature 0 on host and device with identical token ids; the first differing token position shows how quickly outputs diverge. Small logit differences can flip an argmax late in a sequence, which is expected; early divergence is not.
- Logit metrics: compare top-1 agreement and KL divergence of next-token distributions over a fixed prompt set, not just generated text.
- Tokenizer parity: the device tokenizer must produce the same ids as the reference (special tokens, BOS handling, whitespace normalisation).
- Chat template: a missing or different template changes behaviour more than any quantization.
- KV-cache bugs: outputs fine for short prompts but degrade after a certain length point to cache indexing, position (RoPE) offsets, or window wrap-around errors.
import torch.nn.functional as F
def logit_agreement(ref_logits, dev_logits): # [T, vocab] for the same token sequence
top1 = (ref_logits.argmax(-1) == dev_logits.argmax(-1)).float().mean().item()
kl = F.kl_div(F.log_softmax(dev_logits, -1), F.log_softmax(ref_logits, -1),
log_target=True, reduction="batchmean").item()
return {"top1_agreement": top1, "kl": kl}
Shipping and updating models
Shipping turns a working build into a feature that runs on thousands of device models, improves over time, and can be switched off safely. Treat models like code: versioned, tested, staged, observable and reversible.
Shipping a model is like a restaurant chain launching a new dish. It is trialled in a few branches first (staged rollout), compared against the old dish on sales and complaints (A/B testing), each kitchen gets a version suited to its equipment (device tiering), head office can pull it from the menu overnight (remote kill switch), and if a branch runs out of an ingredient it serves the old dish rather than nothing (fallback).
Branches are device cohorts, sales and complaints are quality and performance telemetry, kitchen equipment is RAM and NPU tier, and the old dish is the previous model version or a cloud path.
Device tiering
| Tier | Typical signals | What to ship |
|---|---|---|
| High | 12 GB+ RAM, recent NPU with INT4 support, allow-listed SoC | Largest model (for example 3B W4A16 on NPU), longer context |
| Mid | 8 GB RAM, capable GPU or older NPU | 1B INT4 on GPU/CPU, shorter context, smaller adapters |
| Low | 6 GB or less, no usable accelerator | Tiny task model, or cloud-only with consent, or feature disabled |
Decide the tier with a combination of static facts (RAM from ActivityManager.MemoryInfo.totalMem, SoC model from build properties) and a quick on-device capability probe at first run (load a tiny test graph on the NPU and check it runs), then cache the result.
Versioning and compatibility
- Model manifest: model id, semantic version, runtime and SDK version required, SoC/NPU architecture, input/output schema, tokenizer version, quantization recipe, checksum, minimum app version.
- Schema contracts: pre/post-processing code in the app must match the model's expected inputs; version them together or embed metadata in the model file and assert on load.
- Runtime compatibility: context binaries and compiled caches are tied to NPU generation and SDK version; ship per-SoC variants and invalidate caches on OS or driver updates.
- Adapters: LoRA adapters are tied to a specific base model version; a base model update invalidates all adapters.
Rollout, A/B testing and remote config
- Offline gates Quality evaluation on a fixed set, on-device benchmark on a device matrix (at least one device per tier), and memory/thermal soak all pass.
- Internal and beta cohort Enable via remote config for internal users; monitor crashes and ANRs per device.
- Staged rollout 1%, 5%, 20%, 50%, 100% with automatic halt thresholds (crash rate, p90 latency, fallback rate, user-visible errors).
- A/B test Compare new vs old model on product metrics (task success, acceptance of suggestions, retention) and device metrics (latency, battery), segmented by tier.
- Kill switch A remote flag that disables the model or reverts to the previous version without an app update.
Fallback chain
request ──▶ NPU path ok? ──yes──▶ run
│ no (unsupported / init failed / too hot)
▼
GPU path ok? ──yes──▶ run (maybe smaller model or shorter context)
│ no
▼
CPU path within budget? ──yes──▶ run (reduced max tokens)
│ no
▼
cloud allowed (network + user consent + privacy policy)? ──yes──▶ cloud model
│ no
▼
refuse gracefully with a clear message
What to monitor in the field
- Load success rate and load time per device model and OS version.
- Backend actually used (NPU/GPU/CPU/cloud) and fallback rate.
- Latency percentiles (TTFT, tok/s, end-to-end) by tier.
- Crashes, ANRs, native crashes inside runtime libraries, low-memory kills during inference.
- Thermal status at request start, and share of requests degraded by the governor.
- Quality proxies: user edits or rejections, thumbs up/down, retries; never upload raw private inputs without explicit consent.
- Model version distribution, to know when old versions can be retired.
Security and integrity
- Verify checksums and signatures of downloaded models before loading; a malicious model file can be a crash or code-path attack surface on native parsers.
- Store models in app-private storage; encrypt at rest if the weights are valuable, knowing that a determined attacker with root can still extract them at runtime.
- Apply safety filters and prompt-injection defences for LLM features even though inference is local.
- Check model licences before shipping open weights.
Target numbers to validate against
Use these as a smoke test. If your measurements are wildly off (for example half the expected decode speed), something is misconfigured: wrong build type, missing KV-cache flags, kernels not enabled, wrong thread count, or the device is already hot. All numbers are approximate, depend heavily on SoC, runtime version, prompt length and thermal state, and will improve with newer releases; re-verify before quoting.
Reference numbers are like the typical fuel economy printed for a car model. Your own figure will never match exactly, because driving style and roads differ, but if you get half the rated mileage, you check the tyre pressure and the handbrake before blaming the engine.
The rated mileage is the published benchmark, driving conditions are prompt length and thermal state, and the handbrake is a debug build or a disabled KV cache.
Llama 3.2 on Android with ExecuTorch (CPU, flagship 2024 phone, prompt length 64, KleidiAI enabled)
| Config | Decode tok/s | TTFT (s) | Prefill tok/s | Model size change | Memory change |
|---|---|---|---|---|---|
| 1B BF16 (baseline) | ~19 | ~1.1 | ~60 | - | - |
| 1B SpinQuant (INT4) | ~50 (2.6x) | ~0.3 (-77%) | ~260 (4.3x) | -54% | -40% |
| 1B QAT + LoRA (INT4) | ~46 (2.4x) | ~0.3 (-76%) | ~252 (4.2x) | -52% | -29% |
| 3B BF16 (baseline) | ~7.6 | ~3.0 | ~21 | - | - |
| 3B SpinQuant (INT4) | ~20 (2.6x) | ~0.7 (-76%) | ~90 (4.2x) | -60% | -50% |
| 3B QAT + LoRA (INT4) | ~18.5 (2.4x) | ~0.7 (-76%) | ~89 (4.2x) | -59% | -45% |
Another reference point: quantized Llama 3.2 1B (INT4 weights, 8-bit dynamic activations) on a flagship 2024 phone
- Prefill above ~350 tok/s and decode above ~40 tok/s on CPU with KleidiAI kernels.
- About 2 seconds to respond to a ~600-token input.
- .pte size about 1.1 GiB quantized vs about 2.3 GiB in BF16 (not a clean 4x because embeddings and some layers stay at 8-bit and scales add overhead).
- Peak RSS about 1.9 GiB vs about 3.1 GiB for BF16 at a 2048-token maximum sequence length (roughly 40% lower).
- KleidiAI contributes more than 20% of prefill speed at identical accuracy.
Rule-of-thumb ranges for phone-class flagships (approximate)
| Model | Weights (INT4) | CPU prefill tok/s | CPU decode tok/s | NPU prefill tok/s | NPU decode tok/s | Peak RSS (2k ctx) |
|---|---|---|---|---|---|---|
| ~0.5B | ~0.3-0.5 GB | 300-700 | 60-100 | 1000-2500 | 60-120 | 0.6-1 GB |
| ~1B | ~0.7-1.1 GB | 150-400 | 30-55 | 800-2000 | 30-70 | 1.2-2 GB |
| ~3B | ~1.7-2.3 GB | 60-150 | 12-22 | 300-1000 | 15-30 | 2.5-3.5 GB |
| ~7-8B | ~3.8-4.8 GB | 20-60 | 5-10 | 150-600 | 8-18 | 5-6.5 GB |
Classic models (INT8, flagship phone, approximate per-inference latency)
| Model | CPU (4 threads) | GPU delegate | NPU |
|---|---|---|---|
| MobileNet-class classifier (224 px) | 3-8 ms | 2-5 ms | under 1 ms |
| ResNet-50 (224 px) | 20-50 ms | 8-15 ms | 1-3 ms |
| Small YOLO-style detector (640 px) | 40-120 ms | 15-30 ms | 3-10 ms |
| Small speech encoder (Whisper tiny/base class), 30 s audio | 0.5-2 s | 0.3-1 s | 0.1-0.5 s |
Sanity rules
- Decode ceiling: tok/s cannot exceed effective bandwidth / bytes read per token. With ~50 GB/s effective and ~1 GB of INT4 weights, about 50 tok/s is the ceiling for a 1B model.
- Prefill vs decode: prefill tok/s should be several times (5-20x) decode tok/s; if they are similar, prefill is not batching tokens (for example the KV-cache or SDPA path is disabled).
- Quantization gain: moving BF16 to INT4 should give about 2.4-3x decode (bandwidth ratio minus overheads); much less suggests dequantization overhead or fallback kernels.
- Threads: decode usually peaks at the number of performance cores; more threads can reduce speed.
- Sustained: expect 60-85% of peak after 10 minutes on a phone in free air; much lower means aggressive throttling or a hot environment.
A 24-week learning plan
At 10-15 hours per week, this sequence takes a systems or application engineer to the point of having shipped-quality, measured deployments across CPU, GPU and NPU. Each project depends on the previous project's harness and understanding, so keep the order.
The plan is like a marathon training block. Early weeks build a base (setup and the first run), middle weeks add hard specific sessions (quantization and NPU work), a taper consolidates, and every week has a logbook. Skipping the long runs because they hurt is how people fail on race day.
The base is the environment and Project 1, the hard sessions are Projects 2-4, the race is the system project, and the logbook is your committed measurement data.
| Weeks | Focus | Milestones |
|---|---|---|
| 0 | Device and toolchain | Complete the environment checklist; adb is solid; ExecuTorch builds for host; llama.cpp running on the phone as a fallback baseline. Read a platform AI service architecture overview. |
| 1-3 | Project 1: benchmark harness | Week 1 model on device; week 2 harness scripted; week 3 sustained-load thermal study and first write-up. |
| 4-7 | Project 2: quantization and error analysis | Weeks 4-5 sweep; weeks 5-6 layer-wise diff tool; week 7 mixed-precision recipe validated. Rent a GPU only for the QAT comparison. |
| 8-12 | Project 3: NPU deployment | Weeks 8-9 toolchain friction (normal); week 10 NPU inference; week 11 JNI integration; week 12 three-way comparison with power. |
| 13-16 | Project 4: KV cache | Week 13 profiling; weeks 14-15 compression study; week 16 memory-pressure ceiling on a real device. |
| 17-22 | Project 5: system project | Week 17 choose a design and commit; weeks 17-21 build; week 22 demo and architecture write-up. |
| 23-24 | Consolidate | Turn each project's data into concise stories; practise the interview questions using your own measurements as evidence; optionally start a second vendor toolchain. |
Theory to study alongside (about 20-25 hours total)
Transformer internals (~8 h)
Self-attention and why only K and V are cached; cross-attention; multi-head vs grouped-query vs multi-query and their KV footprints; prefill (compute-bound) vs decode (bandwidth-bound); RoPE and context extension.
Numerics and quantization (~8 h)
FP32/FP16/BF16 layouts and range; fixed-point, scale and zero-point; symmetric vs asymmetric; per-tensor, per-channel, per-group; PTQ vs QAT; outliers; rotation methods at the intuition level.
Efficiency techniques (~5 h)
LoRA/QLoRA; structured vs unstructured pruning (and why unstructured rarely helps on real hardware); distillation; speculative decoding; operator fusion and graph optimisation.
Hardware mapping (~4 h)
NPU block architecture and why it wants static shapes and static quantization; memory bandwidth as the decode ceiling; when GPU beats NPU; heterogeneous scheduling; model splitting for NPU memory limits.
Slip discipline
- If you fall behind, cut scope inside a project, never skip a project. A shallow NPU project still teaches the toolchain; skipping it leaves the biggest gap open.
- If the NPU project stalls completely on hardware or SDK access, switch vendor or move to Project 4 and return later.
- Timebox setup and toolchain debugging; when stuck for more than a day, get a baseline working with a simpler tool and come back.
Failure modes checklist
These are the ways edge AI work most often goes wrong, both while learning and in production. Read them before you start and again before you ship.
This checklist is like a pilot's pre-flight checklist. Experienced pilots still use it, not because they do not know how to fly, but because the failures it prevents are simple, common and catastrophic when missed.
Each line is a known way a deployment "crashes": a wrong number, a misleading benchmark, an app killed in the field.
While learning
| Failure mode | Symptom | Prevention |
|---|---|---|
| Toolchain rabbit hole | Weekends lost to build errors before any model runs | Timebox setup; get llama.cpp running first; one environment per toolchain |
| Measuring on an emulator or laptop | Numbers that do not survive contact with a phone | Every performance claim from physical silicon |
| Cold-run-only benchmarking | One impressive number no user experiences | Warm-up, percentiles, sustained curve; report sustained next to peak |
| Chasing model size | 80 GB GPU bills and slow iteration on 7B models | Learn on 1B-class models; scale up once at the end |
| Drifting into research | Reading papers instead of deploying | Cap theory time; stay on deployment, numerics and systems |
| Not documenting | Finished projects with no evidence | Commit raw data and write a short findings note per project |
Before shipping
- Converted FP32 model matches the original within tolerance.
- Quantized model evaluated on a fixed set including hard and edge cases, not only a few prompts.
- Per-op placement confirmed: no unexpected CPU fallback partitions on the target accelerator.
- Static shapes and max context chosen deliberately; memory budget written down (weights + KV + activations + runtime + app).
- Release build, correct flags (optimized kernels, KV cache, SDPA), correct thread count.
- Load time measured cold and warm; compiled graph or context cached.
- Inference off the main thread; cancellable; session reused; warm-up done off the critical path.
- Model assets uncompressed and memory-mapped, or downloaded with checksum and atomic install.
- Behaviour tested under memory pressure (other apps open,
onTrimMemory), with no low-memory kills of the foreground app. - Ten- to thirty-minute thermal soak done; thermal headroom policy in place.
- Energy per request measured against an idle baseline.
- Tokenizer and chat template parity verified; greedy outputs compared host vs device.
- Hexagon/NPU libraries match the SoC; SDK versions of build and runtime match.
- Fallback chain implemented and tested by forcing each failure.
- Telemetry segmented by SoC, RAM tier and OS build; staged rollout thresholds and kill switch ready.
- Model licence checked; safety filtering in place for generative features; downloaded models verified.
Quick revision
- The deployment loop is choose, export, convert, quantize, compile for target, integrate, benchmark on device, monitor; stages 3-7 iterate many times.
- Export captures a static graph (
torch.export, ONNX); conversion produces a runtime format (.pte, .tflite, .gguf, .onnx/.ort, .mlpackage, QNN context binary). - NNAPI is deprecated from Android 15; new work uses LiteRT delegates/accelerators, ExecuTorch backends or ONNX Runtime execution providers. Confirm current LiteRT
CompiledModel/ Interpreter package names in official docs; do not invent flags. - ExecuTorch swaps backends by changing the partitioner (XNNPACK, Vulkan, QNN, MediaTek, Core ML) while keeping one export flow.
- llama.cpp with GGUF is the fastest way to a working on-device LLM baseline; Q4_0 is repacked for fast Arm kernels, Q4_K_M is a quality-per-byte default.
- For ExecuTorch LLM export,
use_kv_cacheanduse_sdpa_with_kv_cacheare essential;8da4wmeans 8-bit dynamic activations, 4-bit weights. - KleidiAI kernels in XNNPACK add more than 20% prefill speed on Arm CPUs; always build Release.
- Validate numerics after every conversion: FP32 differences should be around 1e-5 to 1e-4.
- Never benchmark on an emulator: no realistic memory system, DVFS, thermal model or NPU.
- Report load time (cold and warm), TTFT, prefill tok/s, decode tok/s, per-token p50/p90/p99, peak RSS, file size, energy per request and sustained/peak ratio.
- Always state prompt length and generated length with LLM numbers; prefill and decode differ by 5-20x.
- A fair protocol fixes the environment, cools down between runs, discards warm-up runs, repeats, and records device, build and runtime versions.
- Peak RSS comes from
VmHWMin/proc/<pid>/status; app memory breakdown fromdumpsys meminfo. - Energy = integral of (power minus idle power) over time; use power rails via Perfetto, fuel gauge sampling, batterystats, or an external monitor.
- A 10-30 minute thermal soak reveals throttling; phones typically sustain 60-85% of peak.
- Quantization quality needs both perplexity and task-level evaluation on a fixed set including hard cases.
- Layer-wise error analysis compares each intermediate tensor to an FP32 reference using cosine similarity, MAE, max error and SQNR.
- SQNR = 10 log10(signal power / error power); each bit adds about 6 dB; low SQNR on the residual stream predicts quality loss.
- Single-layer sensitivity (quantize one layer at a time) separates fragile layers from layers receiving bad inputs.
- Typical sensitive parts: outlier activation channels, MLP down projection, LM head, first/last blocks, norms and softmax.
- Mixed precision promotes only the most sensitive layers to 8-bit or FP16 and prices each promotion in size and latency.
- Host fake-quant matching FP32 but device diverging means a device implementation issue (rounding, accumulators, FP16 overflow), not the scheme.
- NPUs need static quantization: ranges fixed at compile time from calibration; LLMs on Hexagon commonly use W4A16 or W8A16.
- NPUs need static shapes; export separate prefill (chunked) and decode (one token) graphs that share weights.
- One unsupported op mid-graph creates partitions and CPU round trips that can make the NPU slower than the CPU.
- Context binaries are finalized graphs for one Hexagon architecture and SDK version; loading them avoids on-device compilation.
- Hexagon library folders (v73, v75, v79) must match the SoC; set
ADSP_LIBRARY_PATHfor DSP-side libraries. - Large models are split into several context binaries because of NPU session memory limits.
- KV bytes = 2 x layers x KV heads x head dim x tokens x bytes; Llama-3.2-1B uses 32 KiB per token in FP16 (verified: 2 × 16 × 8 × 64 × 2).
- At long context the KV cache exceeds the weights (about 22-35k tokens for 1B INT4, about 16-20k for 3B, under 8k for models without GQA).
- Decode speed falls as context grows because each step reads the weights plus the whole cache: tok/s ≤ BW / (W + KV(t)).
- INT8 weights are usually near-lossless; INT4 is the decode lever but needs a task eval, plus higher precision on embeddings and the LM head.
- Quantize keys per channel (outlier channels) and values per token; INT8 KV is near-lossless.
- Sliding-window attention bounds memory; keep attention-sink tokens to avoid collapse on long streams.
- Paged KV: contiguous waste is (T_max - T_actual) x KV per token; paged waste is at most one block. Prefix caching saves about P / (P + U) of prefill if the prefix is byte-identical.
- Speculative decoding: E[tokens per pass] = (1 - alpha^(k+1)) / (1 - alpha); speed-up ≈ E / (1 + k c). Helps on-device when decode is bandwidth-bound and the draft is accurate.
- lmkd kills by
oom_score_adjunder memory pressure (PSI); anonymous memory is not reclaimable, mmapped clean file pages are. - Store model assets uncompressed (
noCompress) so they can be memory-mapped; downloaded models need resume, checksum, atomic install and a smoke test. - Create sessions once, off the main thread, warm them up, reuse buffers, make generation cancellable and respond to
onTrimMemory. - ONNX Runtime's QNN EP needs a QDQ model; disable CPU fallback during development and enable EP context caching. Provider option keys (
backend_path, HTP performance mode) are version-specific: check the current ORT QNN EP docs. - QAIRT converter and Genie binary/API names move between SDK releases. Copy flags, JSON keys and C symbols from the samples for the version you link; do not invent them.
- Profile at three levels: system (Perfetto), runtime (per-op profiles, AI Hub, QNN profiler, ETDump), native (simpleperf).
- Accuracy drift is bisected: original, converted FP32, FP16, quantized host, quantized device, full app pipeline; pre-processing bugs are the most common cause.
- Shipping needs device tiering, versioned manifests, staged rollout with halt thresholds, A/B tests, a kill switch and a fallback chain (NPU, GPU, CPU, cloud, refuse).
- Monitor load success, backend used, fallback rate, latency by tier, low-memory kills and thermal state, segmented by SoC, RAM and OS build.
Glossary
- Activation quantization
- Representing the intermediate outputs of layers with low-bit integers; harder than weight quantization because ranges depend on inputs and contain outliers.
- AI Hub
- Qualcomm's hosted service to compile, quantize, profile and run models on real Snapdragon devices, with a model zoo of export scripts.
- AIMET
- Qualcomm's open-source toolkit for quantization and compression, including cross-layer equalisation, AdaRound, QuantAnalyzer and QAT.
- Attention sink
- The first few tokens of a sequence, which attract a large share of attention; keeping them in a sliding-window cache keeps long generations stable.
- Block table
- The mapping from a sequence's logical KV-token ranges to physical KV blocks in a paged cache, analogous to a page table.
- CompiledModel
- LiteRT Next's compiled-model API that binds a
.tflite(or a vendor-compiled artifact) to an accelerator. Class names and options are release-specific; confirm in current LiteRT docs. - AWQ
- Activation-aware weight quantization: a PTQ method that protects the weight channels most important to activations by scaling before quantizing.
- batterystats
- Android's battery accounting service;
dumpsys batterystatsreports estimated energy use per app over a measurement window. - BF16
- Brain float 16: 16-bit floating point with the same exponent range as FP32 but fewer mantissa bits; robust to overflow.
- Calibration dataset
- A small, representative set of inputs run through a model to record activation ranges for static quantization.
- Context binary
- A serialized, fully compiled QNN graph for a specific Hexagon architecture that loads without on-device graph compilation.
- CPU fallback
- Execution of operators the accelerator cannot run on the CPU instead, creating partitions and data transfers.
- Decode phase
- The token-by-token generation phase of an LLM; memory-bandwidth-bound because all weights and the KV cache are read for each token.
- Draft acceptance rate
- The per-token probability that a speculative draft token is accepted by the target model; it sets expected tokens per verification pass.
- Delegate
- A LiteRT plug-in that takes over supported parts of a graph and runs them on a GPU, NPU or DSP.
- Dynamic quantization
- Quantizing activations at runtime with scales computed from each tensor's actual values; common on CPUs, not possible on most NPUs.
- EP context cache
- An ONNX Runtime feature that saves a compiled execution-provider graph (for example a QNN context) so later sessions skip compilation.
- ETDump and ETRecord
- ExecuTorch developer-tool artifacts: runtime profiling and debug data (ETDump) and export-time graph metadata (ETRecord) used together by the Inspector.
- Execution provider
- An ONNX Runtime backend (CPU, QNN, Core ML, XNNPACK and others) that executes the subgraphs assigned to it.
- ExecuTorch
- PyTorch's on-device runtime that executes
.pteprograms produced fromtorch.export, with pluggable hardware backends. - Fake quantization
- Simulating quantization in floating point (quantize then dequantize) to measure or train for its effect on a host.
- Genie
- Qualcomm's generative AI runtime in QAIRT that wraps tokenizer, QNN execution, KV-cache management and sampling behind a dialog API. Confirm headers, JSON keys and CLI names against the SDK samples for your release.
- GGUF
- llama.cpp's single-file model format containing tensors, quantization types, tokenizer and metadata.
- GPTQ
- A post-training weight quantization method that quantizes weights column by column while compensating the error using second-order information from calibration data.
- GQA (grouped-query attention)
- Attention where several query heads share each key/value head, shrinking the KV cache in proportion.
- Graph partitioning
- Splitting a model graph into subgraphs for different processors; each boundary adds transfer and synchronisation cost.
- Group-wise quantization
- Using one scale (and zero-point) per small block of weights, for example 32 or 128, to follow local ranges more closely.
- heapprofd
- Perfetto's native heap profiler that attributes allocations to call stacks in a running process.
- Hexagon HTP
- The Hexagon Tensor Processor, the matrix/tensor accelerator in Snapdragon NPUs, reached through QNN/QAIRT.
- KleidiAI
- Arm's library of optimized low-bit matrix-multiply micro-kernels, integrated into XNNPACK and other runtimes.
- KV cache
- Stored key and value vectors for all previous tokens in every layer, so attention does not recompute them each step.
- LiteRT
- Google's on-device runtime, formerly TensorFlow Lite, executing
.tflitemodels with CPU, GPU and NPU accelerators. - llama.cpp
- A C/C++ LLM inference engine with optimized CPU kernels and GPU backends, using the GGUF format.
- lmkd
- Android's userspace low-memory killer daemon, which kills processes by priority when memory pressure is high.
- LoRA adapter
- A small set of low-rank weight updates that specialise a shared base model for a task and can be swapped at runtime.
- Memory mapping (mmap)
- Mapping a file into the address space so pages load lazily and can be reclaimed by the kernel, avoiding a full copy into the heap.
- Mixed precision
- Using different numeric precisions for different layers, typically keeping sensitive layers at higher precision.
- NNAPI
- Android Neural Networks API, the older accelerator abstraction, deprecated from Android 15 in favour of vendor delegates and backends.
- ODPM
- On-Device Power Monitor: hardware energy counters for power rails, readable via Perfetto on supported devices.
- Paged KV cache
- A KV cache stored in fixed-size blocks mapped by a block table, avoiding contiguous reallocations and fragmentation.
- Perfetto
- Android's system-wide tracing tool for scheduling, frequencies, thermal, power rails, memory and custom trace slices.
- Prefill phase
- Processing all prompt tokens in parallel to fill the KV cache; compute-bound and the main contributor to TTFT.
- Prefix caching
- Reusing a precomputed KV cache for a shared prompt prefix, such as a system prompt, to reduce TTFT.
- PSI
- Pressure stall information: kernel metrics showing how much time tasks stall waiting for memory, CPU or IO.
- PTQ
- Post-training quantization: quantizing a trained model using calibration data, without retraining.
- QAIRT
- Qualcomm AI Runtime SDK, which includes the QNN libraries, converters, quantizers, HTP backend, profiling and Genie.
- QAT
- Quantization-aware training: fine-tuning with simulated quantization so the model learns to tolerate rounding.
- QDQ format
- An ONNX representation of quantization using explicit QuantizeLinear/DequantizeLinear nodes that backends fuse into integer kernels.
- RSS and PSS
- Resident set size (all resident pages of a process) and proportional set size (shared pages divided among sharers).
- simpleperf
- Android's sampling CPU profiler for native code, with call graphs and hardware counters.
- Sliding-window attention
- Attention restricted to the most recent W tokens, bounding KV-cache memory.
- SpinQuant
- A quantization method that applies learned rotations to weights and activations to spread outliers before low-bit quantization.
- Speculative decoding
- A draft model (or extra heads / n-grams) proposes tokens that the target verifies in one pass. Expected tokens per pass = (1 − αk+1) / (1 − α). Output distribution is unchanged.
- SQNR
- Signal-to-quantization-noise ratio in decibels, comparing reference tensor energy with quantization error energy.
- Static quantization
- Quantization with activation scales fixed ahead of time from calibration; required by most NPUs.
- Static shapes
- Tensor dimensions fixed at export/compile time, allowing NPU compilers to plan memory and tiling.
- Sustained-to-peak ratio
- Throughput after a long run divided by initial throughput; captures the effect of thermal throttling.
- Thermal headroom
- Android's forecast of how close the device is to severe throttling, from
PowerManager.getThermalHeadroom(). - Thermal soak
- A long continuous workload (10-30 minutes) used to measure throttling and sustained performance.
- torch.export
- PyTorch's ahead-of-time graph capture producing an ExportedProgram with no Python control flow, used by ExecuTorch and ai-edge-torch.
- TTFT
- Time to first token: from prompt submission to the first generated token appearing.
- VTCM
- Vector tightly coupled memory: fast on-chip memory in the Hexagon NPU; working sets that exceed it spill to DRAM.
- W4A16
- A scheme with 4-bit weights and 16-bit activations, common for LLMs on NPUs.
- XNNPACK
- A highly optimized CPU inference library for Arm and x86 used as the CPU backend by LiteRT, ExecuTorch and ONNX Runtime.
Interview questions
Fundamentals
What are the stages of deploying a model to an edge device?
Eight stages: (1) choose or train a model that fits the latency and memory budget; (2) export it as a static graph (torch.export, ONNX); (3) convert to a runtime format (.pte, .tflite, .gguf, .onnx, Core ML, QNN); (4) quantize with calibration data or QAT; (5) compile for the target accelerator (partitioning, context binaries, caches); (6) integrate into the app (threading, packaging, fallback); (7) benchmark on real devices (latency percentiles, memory, energy, thermal); (8) monitor in the field with staged rollout and rollback. Stages 3-7 iterate: an unsupported op or a latency miss sends you back.
What is the difference between exporting and converting a model?
Exporting captures the model's computation as a static, framework-level graph with fixed operators and (usually) fixed shapes, removing Python control flow: for example a torch.export ExportedProgram or an ONNX file. Converting transforms that graph into the format and operator set of a specific runtime: a .pte for ExecuTorch, a .tflite for LiteRT, a QNN model or context binary, a Core ML package. Export problems are about capturability (dynamic control flow, data-dependent shapes); conversion problems are about operator coverage and layout.
Why is NNAPI no longer recommended, and what replaced it?
NNAPI offered a common Android interface to accelerators, but vendors implemented its drivers inconsistently, the op set was a lowest common denominator, and behaviour and performance varied across devices, causing fragmentation. From Android 15 it is deprecated for new work. The replacement is vendor-specific delegates and backends integrated directly into frameworks: LiteRT GPU and NPU accelerators, ExecuTorch backends (QNN, MediaTek, Vulkan), and ONNX Runtime execution providers such as QNN.
What is ExecuTorch and what is a .pte file?
ExecuTorch is PyTorch's on-device inference runtime. You capture a model with torch.export, lower it to an "edge" dialect, hand supported subgraphs to backends through partitioners (XNNPACK for CPU, Vulkan, QNN, MediaTek, Core ML, etc.), and serialize the result into a .pte program. The .pte contains the execution plan, constant weights (or references to external weight files) and delegate blobs. The runtime core is small and portable C++, with Java/Kotlin and Swift bindings and an LLM runner.
What is LiteRT, and how do you get a PyTorch model into it?
LiteRT is the new name for TensorFlow Lite: Google's runtime for .tflite FlatBuffer models with XNNPACK on CPU, a GPU delegate and NPU accelerators. NNAPI is deprecated from Android 15; new Android work should use LiteRT delegates or vendor backends, not NNAPI. For PyTorch models you use ai-edge-torch, which runs torch.export internally and emits a .tflite. The long-lived app API is still an Interpreter-style session; LiteRT Next also documents a CompiledModel API whose class names and accelerator options you must confirm in the current official docs. TensorFlow/Keras models use the TFLiteConverter with a representative dataset for integer quantization.
What is GGUF, and why start LLM work with llama.cpp?
GGUF is llama.cpp's single-file format holding quantized tensors, tokenizer and metadata. llama.cpp builds in minutes with the NDK, runs most popular architectures on CPU with highly tuned Arm kernels, and has llama-bench for prefill/decode numbers and llama-perplexity for quality. That makes it the fastest way to validate that a use case works on a phone and to obtain a CPU reference baseline before investing in a heavier export path such as ExecuTorch or a vendor NPU flow.
What is a QNN context binary?
It is a serialized, fully prepared QNN graph for a specific Hexagon NPU architecture (and SDK version): the graph has been optimised, tiled and finalized, with weights and quantization parameters embedded. Loading a context binary skips on-device graph compilation, turning multi-second initialization into a fast load. It is not portable across Hexagon generations, so you produce one per target SoC family, via AI Hub, the QAIRT context-binary generator, or a framework backend.
Why can't you trust performance numbers from an emulator or a laptop?
They do not reproduce the device's memory bandwidth and cache hierarchy, DVFS governors, thermal limits and throttling, big.LITTLE scheduling, or the NPU/GPU drivers at all, and they run different (x86) kernels. The error is not a constant factor; it can change which option is faster. Host runs are useful for numerical reference and functional tests only; all performance claims must come from physical silicon.
What are TTFT, prefill throughput and decode throughput?
TTFT (time to first token) is the time from submitting a prompt to the first generated token; it is dominated by prefill. Prefill throughput is prompt tokens divided by prefill time: all prompt tokens are processed in parallel with matrix-matrix operations, so it is compute-bound. Decode throughput is generated tokens per second after the first: one token at a time, reading all weights and the cache each step, so it is memory-bandwidth-bound. Always report them separately with prompt and output lengths.
Why is LLM decode memory-bandwidth-bound?
Each decode step multiplies a single token's activation vector by every weight matrix: a matrix-vector product with about 2 FLOPs per weight read. The arithmetic intensity is so low that the processor waits on memory, not compute. So decode speed is roughly effective bandwidth divided by bytes read per token (weights plus KV cache). This is why weight quantization speeds up decode almost in proportion to the bytes saved, and why NPUs help less for decode than for prefill.
Why do benchmarks need warm-up runs?
The first runs include one-time costs: page faults as mmapped weights are touched, cache and TLB warming, GPU shader compilation, NPU graph finalization, kernel auto-selection, memory pool allocation and CPU frequency ramp-up. Including them mixes initialization with steady-state latency. Measure and report load and first-run time separately, discard a few warm-up runs, then measure steady state.
Why report p50, p90 and p99 rather than the mean?
Latency distributions on phones are skewed: scheduler preemption, frequency changes, garbage collection, thermal events and background work create long tails. The mean hides them and is distorted by outliers. p50 describes the typical experience; p90/p99 describe what users notice as stutter or lag. For LLMs, per-token inter-token latency percentiles reveal stutters that an average tok/s hides.
What does "8da4w" mean in ExecuTorch LLM export?
8-bit dynamic activations, 4-bit weights. Linear layer weights are stored as 4-bit integers with group-wise scales (for example one per 128 weights); at runtime, each activation tensor is quantized to 8-bit with a scale computed from its actual values, and integer matmul kernels run the product. It gives most of the INT4 size and bandwidth benefit while keeping activation error low, and it suits CPU backends like XNNPACK with KleidiAI kernels.
What is a calibration dataset?
A small representative set of inputs (typically 100-500 samples, or a few hundred text sequences) passed through the model during static quantization so observers can record activation ranges and choose scales. It must resemble production data, including edge cases; otherwise ranges are wrong, values clip or lose resolution and accuracy drops. For LLMs, the calibration text should match the target domain and languages.
What is the difference between static and dynamic quantization?
Dynamic quantization computes activation scales at runtime from the current tensor's range: accurate and needs no calibration, but costs a reduction per tensor per step and requires flexible hardware (CPUs). Static quantization fixes activation scales ahead of time using calibration data: no runtime overhead and compatible with fixed-function integer pipelines (NPUs, DSPs), but values outside the calibrated range clip. Weights are always effectively static.
What is a delegate or execution provider?
It is a plug-in backend that claims the parts of a model graph it supports and runs them on specific hardware (GPU, NPU, DSP, optimized CPU library). LiteRT calls them delegates or accelerators, ONNX Runtime calls them execution providers, ExecuTorch calls them backends selected via partitioners. Unsupported parts remain on the default CPU path, which creates partitions.
What is CPU fallback and why does it hurt performance?
When an accelerator cannot run an operator (unsupported type, shape, data type or attribute), the runtime executes that operator on the CPU. The graph is split into partitions and at each boundary data must be transferred, possibly re-laid out and requantized, and the processors synchronise. With several partitions the accelerator idles while waiting, and total time can exceed a CPU-only run. Aim for full delegation and check placement in profiles.
What does memory-mapping the model weights give you?
With mmap, the model file is mapped into the address space instead of copied into the heap. Loading becomes almost instant (pages are read on first access), clean file-backed pages can be reclaimed by the kernel under pressure instead of forcing a kill, and multiple processes mapping the same file share physical pages. The downsides: evicted pages must be re-read, causing latency spikes, and runtimes that repack weights create anonymous copies that lose these benefits.
Why must model assets be stored uncompressed in an APK?
Compressed APK entries cannot be memory-mapped directly; the runtime must decompress the whole file into memory, doubling peak memory at load and slowing startup. Marking model extensions as noCompress in Gradle (androidResources { noCompress += "tflite" }) keeps them stored and page-aligned so AssetFileDescriptor plus FileChannel.map can mmap them.
Why should inference never run on the UI thread, and why reuse the session?
Inference can take tens of milliseconds to seconds; on the main thread it blocks rendering and input, causing jank and ANRs. Session creation is also expensive (loading, delegate init, graph compilation), so creating it per request multiplies latency and memory churn. Create the session once on a background thread, warm it up, keep it in an application-scoped owner, reuse input/output buffers, and serialise calls on a dedicated executor.
What is the KV cache and why is it needed?
In self-attention, each new token attends to the keys and values of all previous tokens. Without caching, each step would recompute keys and values for the entire sequence, making generation quadratic. The KV cache stores K and V for every past token in every layer, so each decode step only computes K and V for the new token and appends them. It trades memory (growing linearly with context) for compute.
What is thermal throttling, and what is a thermal soak test?
Phones are passively cooled; sustained power heats the SoC and skin, and the thermal framework lowers CPU/GPU/NPU frequencies to stay within temperature limits, reducing throughput. A thermal soak test runs the workload continuously for 10-30 minutes while logging throughput, temperatures, frequencies and power, showing when throttling starts and the sustained-to-peak ratio. It is the measurement that predicts real product behaviour.
How do you measure peak memory of an on-device inference process?
For a native process, read VmHWM (peak resident set) and VmRSS from /proc/<pid>/status, polling during the run or printing at exit. For an app, use dumpsys meminfo <package> for PSS broken down by native heap, graphics and file mappings, and Perfetto's heapprofd to attribute native allocations to call stacks. Distinguish file-backed (mmapped weights) from anonymous memory, because only the latter is non-reclaimable.
How can you measure energy consumption of inference on Android?
Options in increasing accuracy: batterystats (model-based per-UID estimates), sampling the fuel gauge (current_now and voltage_now) and integrating power over time, on-device power rails (ODPM) recorded via Perfetto's android.power data source for per-subsystem energy, and an external power monitor. Always subtract an idle baseline with the same screen and radio state, and report energy per inference or per 100 tokens.
What is device tiering?
Grouping devices by capability (RAM, SoC and NPU generation, GPU, OS version) and shipping different model variants or settings per tier: for example a 3B NPU model on high-end phones, a 1B CPU/GPU model on mid-range, and cloud or no feature on low-end. Tier decisions combine static facts with a first-run capability probe, and are controlled remotely so they can be adjusted after launch.
Should you bundle a model in the APK or download it?
Bundle small models (tens of MB) that must work immediately and offline: simplest, always available, but updates require an app update and increase install size. Download larger models (asset packs, device-targeted AI packs or your own CDN): smaller install, per-device variants and independent updates, at the cost of first-use delay and the need for resume, integrity checks, storage checks, versioning and fallback while not yet downloaded.
What is a chat template, and why does it matter on device?
Instruction-tuned LLMs were trained with a specific prompt format of special tokens marking system, user and assistant turns (for Llama 3.x, header and end-of-turn tokens). Sending raw text without the template or with the wrong special token ids produces rambling, off-task or empty outputs and missing stop conditions. On-device runners often do not apply templates automatically, so it is a frequent cause of "the model is worse on the phone".
What is mixed-precision quantization?
Using different bit-widths in different parts of the model: most layers at the target low precision (for example INT4 weights) and a few sensitive layers (often the LM head, some down projections, first or last blocks, norms) at INT8 or FP16. It recovers most of the quality lost by uniform low-bit quantization at a small size and latency cost, and should be derived from measured layer sensitivity.
What is lmkd and why does it matter for on-device LLMs?
lmkd is Android's low-memory killer daemon. It watches memory pressure (PSI) and kills processes in order of oom_score_adj: cached apps first, then services, then perceptible and finally foreground apps. A large model plus KV cache can push the system into killing other apps (music, launcher) or, under extreme pressure, the foreground app itself. Memory budgets for LLM features must leave room for the rest of the system.
What are SQNR and cosine similarity used for in quantization work?
They measure how close a quantized tensor is to its FP32 reference. SQNR (in dB) is the ratio of signal energy to error energy; each extra bit of uniform quantization adds about 6 dB. Cosine similarity measures directional agreement, insensitive to uniform scaling. Computed per layer, they locate where quantization error is introduced and how it accumulates, guiding mixed-precision decisions.
Going deeper
Walk through exporting Llama-3.2-1B to ExecuTorch for an Android CPU. Which flags matter most?
Get consolidated.00.pth, params.json and tokenizer.model. Run the LLM exporter with the Llama 3.2 model class, enabling the KV cache (use_kv_cache) and the fused SDPA-with-KV-cache op (use_sdpa_with_kv_cache), XNNPACK backend, 8da4w quantization with group size 128 (or 32 for better quality), 4-bit embedding quantization with group 32, a chosen max_seq_length, and metadata with the correct BOS/EOS ids (128000; 128001/128009). Push the .pte, tokenizer and a Release-built runner with KleidiAI enabled. The KV-cache and SDPA flags are the most important; without them decode recomputes attention over the full sequence and speed collapses. Wrong EOS ids make generation never stop.
What is KleidiAI and what would you check if prefill is 20% below reference numbers?
KleidiAI is Arm's set of optimized low-bit matmul micro-kernels using dotprod/i8mm (and SME where available), integrated into XNNPACK. It improves prefill by over 20% at identical accuracy. If prefill is low: confirm the build has EXECUTORCH_XNNPACK_ENABLE_KLEIDI on and is Release; confirm the quantization scheme is one the kernels support (for example 4-bit group-wise weights with 8-bit dynamic activations); check thread count and that threads run on performance cores; check the device is not already thermally throttled; and compare prompt lengths with the reference.
How do you choose between GGUF Q4_0, Q4_K_M and Q8_0, and how many threads to use?
Q8_0 is near-lossless but twice the bytes of 4-bit, so decode is roughly half as fast. Q4_K_M uses super-blocks with some 6-bit tensors, giving better quality per byte and a common default. Q4_0 is simpler; on Arm CPUs with i8mm/dotprod, llama.cpp repacks it into interleaved layouts for very fast kernels, so it is often the fastest phone option with slightly lower quality (mitigated with an importance matrix). Measure perplexity and your task. For threads, start with the number of performance cores; using all cores often slows decode because little cores become stragglers, and prefill may benefit from a different count than decode.
Explain the ONNX static quantization flow and the QDQ format.
Pre-process the model (shape inference, constant folding), implement a CalibrationDataReader yielding representative inputs, and call quantize_static with a calibration method (MinMax, Entropy, Percentile), per-channel weights, activation and weight types, and QuantFormat.QDQ. QDQ inserts QuantizeLinear/DequantizeLinear pairs around tensors; the graph remains valid in float, and execution providers pattern-match DQ-op-Q sequences into integer kernels. For the QNN EP, use ORT's QNN helpers to produce the activation types the HTP expects (8- or 16-bit) and ensure all ops are supported.
How do you use ONNX Runtime with the QNN execution provider on Android?
Use the QNN-enabled Android package, create SessionOptions and add the QNN EP. The Java helper name and option keys are version-specific: confirm them in the current ONNX Runtime QNN EP docs. Common documented keys include backend_path (HTP library) and an HTP performance mode. Feed a QDQ-quantized model. During development disable CPU EP fallback so unsupported nodes error instead of silently falling back. Enable EP context caching so the compiled QNN graph is saved and reused on later launches. Profile to confirm all nodes are on QNN, and re-enable CPU fallback in production for robustness.
When would you use MediaPipe LLM Inference vs ExecuTorch vs llama.cpp in an app?
MediaPipe LLM Inference (and LiteRT-LM) when a supported model family fits the product and you want minimal code, GPU support and Google-maintained bundles. ExecuTorch when you need arbitrary PyTorch models, multiple backends including vendor NPUs, custom quantization, and one flow for Android and iOS. llama.cpp for quick prototyping, broad architecture support on CPU, GGUF distribution and fine control in C++, accepting weaker NPU support. Many teams prototype with llama.cpp and ship with ExecuTorch or a vendor path.
Design a benchmark harness for on-device LLMs. What must it record?
A host-side driver (adb) plus on-device timing in the runner. Per run: model load time cold and warm, TTFT, prefill tok/s at several prompt lengths, decode tok/s, inter-token latency percentiles, peak RSS (VmHWM) and PSS, file size, energy per 100 tokens with idle baseline, temperatures and CPU frequencies at start and end. Protocol: fixed environment, cool-down to a temperature threshold between configs, warm-up runs discarded, repetitions with p50/p90/p99 and stdev, a sustained soak mode. Metadata: device, SoC, build fingerprint, runtime and SDK versions, model hash, quantization config, threads, ambient temperature. Output a CSV and a summary, one command per full run.
How do you correctly time TTFT and per-token latency in code?
Use a monotonic clock (std::chrono::steady_clock, SystemClock.elapsedRealtimeNanos), never wall-clock time. TTFT spans from prompt submission (including tokenization if user-visible) to the first token being sampled; prefill rate is prompt tokens over prefill time. For decode, record a timestamp after each sampled token and compute inter-token intervals; decode rate is (generated minus 1) over decode time. Exclude detokenization or UI rendering only if you report them separately, and do not print per token to the console during timing, which adds I/O overhead.
Why is energy per inference a better metric than power, and what is race to idle?
Energy (power times time) is what drains the battery. A backend drawing more power but finishing much faster can use less energy per request than a slow, low-power one. Race to idle is the strategy of completing work quickly at high performance and letting the hardware return to deep idle, often more efficient than running slowly, as long as the high-power bursts do not trigger throttling. For continuous workloads (30 fps camera, long generation) the average power and thermal steady state matter more.
How do you run and interpret a thermal soak test?
Run back-to-back generations (or inferences) for 10-30 minutes at a fixed ambient temperature, with the device in its realistic state (case on or off noted, screen state fixed). Log throughput per generation, thermal zones, CPU/GPU frequencies, thermal status from thermalservice, and power. Plot throughput and temperature against time. Interpret: time until first throttle step, depth of each drop, the plateau level, the sustained-to-peak ratio, and whether throughput oscillates (governor hunting). Compare backends: NPUs usually sustain better due to lower power.
Describe how you would do layer-wise quantization error analysis.
Run the FP32 model on the host with hooks capturing every layer's output on a fixed input set; run the quantized model (host fake-quant, then device with intermediate-output dumping) on the same inputs. For each tensor, compute cosine similarity, MAE, max absolute error and SQNR; rank layers and plot SQNR across depth. Then quantize one layer at a time to measure individual sensitivity, check activation ranges for outliers, and derive a mixed-precision recipe that promotes only the most sensitive layers. Validate the recipe with quality metrics and the performance harness.
How does quantization group size affect quality and size?
Smaller groups adapt scales to local ranges, reducing error, especially with outliers; larger groups use fewer scales. The overhead is scale bits divided by group size: a 16-bit scale per 32 weights adds 0.5 bits per weight (4.5 effective bits), per 128 adds 0.125. Small groups also add dequantization work and can reduce kernel efficiency. Common choices: 32 for quality (and some NPU backends), 128 for CPU LLM exports; per-channel for INT8.
Why do the embedding table and LM head deserve special treatment?
In small LLMs with large vocabularies they are a big share of parameters: Llama-3.2-1B's 128256 x 2048 embedding is about 263M of 1.24B parameters (about 21%). Quantizing the embedding saves a lot of storage with little quality loss because it is a lookup. The LM head, if not tied, is also large, but its errors land directly on logits, so it is unusually sensitive; it is often kept at 6-8 bits even when the rest is 4-bit. With tied weights, the choice affects both.
Perplexity vs task metrics: why do you need both?
Perplexity measures how well the model predicts held-out text on average: sensitive, cheap and good for comparing configurations. But small perplexity changes can hide collapses on specific skills (arithmetic, code, structured output, non-English) and large ones may not matter for a narrow task. Task-level metrics (accuracy on multiple-choice, exact match, format validity on your product's prompts) measure behaviour. Use perplexity to sweep, task metrics to decide, plus top-1 agreement or KL against the FP32 model.
Compare round-to-nearest, GPTQ, AWQ, SpinQuant and QAT.
Round-to-nearest quantizes each weight independently: fastest, worst at 4-bit. GPTQ quantizes weights column by column and updates remaining weights to compensate using second-order (Hessian) information from calibration data. AWQ identifies weight channels important to large activations and scales them to protect them before quantization. SpinQuant learns rotations applied to weights and activations that spread outliers, enabling low-bit weights and activations with small loss. QAT fine-tunes with simulated quantization (optionally with LoRA to keep it cheap) and recovers the most quality at the highest cost. PTQ methods need minutes to hours; QAT needs a training setup and GPUs.
How do you capture intermediate tensors on the device for comparison?
ExecuTorch: generate an ETRecord at export time, run with ETDump and a debug buffer enabled, and use the devtools Inspector to map on-device outputs back to graph nodes. QNN/QAIRT: qnn-net-run --debug dumps all intermediate outputs; the SDK accuracy debugger compares against a framework reference. ONNX Runtime: the QDQ loss debug utilities add intermediate outputs and match FP32 and quantized activations. LiteRT: the quantization debugger or extra model outputs. llama.cpp: evaluation callbacks print per-tensor stats. Compare on identical inputs and align names or node ids.
Why do NPUs need static shapes, and how do you handle variable-length LLM input?
NPU compilers plan tiling, on-chip memory allocation, DMA schedules and instruction streams for exact tensor sizes at compile time; dynamic shapes break that planning. For LLMs, export two graphs: a prefill graph processing fixed-size chunks (for example 128 tokens, padding the last chunk and masking) and a decode graph processing one token, both with a KV cache of fixed maximum length passed as inputs/outputs and an attention mask marking valid positions. Weights are shared between the graphs. For other models, pad or resize inputs into a few fixed buckets.
Why are LLMs split into several context binaries for the NPU?
NPU sessions have limits on graph size and on the memory that can be mapped into the NPU's address space, and very large graphs compile slowly and exceed on-chip planning limits. Splitting the model by layers into a few binaries (for example 3-5 parts for a 3B model) keeps each within limits; the runtime executes them in sequence, passing hidden states between them. Weight sharing between the prefill and decode variants of each part avoids storing weights twice.
What are Hexagon library versions and ADSP_LIBRARY_PATH about?
The QNN HTP backend has an ARM-side library (loaded by your process) and DSP-side "skeleton" libraries that run on the Hexagon processor, built per architecture: v73 for Snapdragon 8 Gen 2, v75 for 8 Gen 3, v79 for 8 Elite. The skeleton must match the SoC. ADSP_LIBRARY_PATH tells the DSP loader (via FastRPC) where to find them; LD_LIBRARY_PATH or the app's native library directory covers the ARM side. A mismatch causes load failures or fallback. Build-time and runtime SDK versions must also match.
Compute the KV cache size for Llama-3.2-1B at 8k tokens, and for the 3B.
1B: 16 layers, 8 KV heads, head dim 64 (2048 / 32 heads). Per token in FP16: 2 x 16 x 8 x 64 x 2 bytes = 32 KiB. At 8192 tokens: 256 MiB; INT8 about 128 MiB; INT4 about 64 MiB plus scales. 3B: 28 layers, 8 KV heads, head dim 128, so 112 KiB per token and about 896 MiB at 8k in FP16. Both use GQA; with full multi-head attention the numbers would be several times larger.
Why quantize keys per channel and values per token?
Key vectors have a few channels with consistently large magnitudes across tokens (outlier channels, partly due to RoPE and learned structure). Per-token scales would be dominated by those channels and crush the others, so per-channel scales (computed across tokens) fit keys better. Values do not show such fixed-channel outliers but vary per token, so per-token scales fit them. Verify by plotting magnitude per channel for K and V on your model; practical implementations may group tokens for per-channel key scales since the cache grows.
Explain sliding-window attention and attention sinks.
Sliding-window attention limits each token to attend to the last W tokens, so the KV cache is a fixed-size ring buffer and memory is bounded regardless of conversation length. The cost is losing direct access to older context. Models trained with full attention tend to dump a lot of attention mass on the first few tokens ("sinks"); evicting them destabilises generation and perplexity explodes. Keeping a handful of initial tokens plus the recent window restores stability for long streams. It does not restore retrieval of facts that fell out of the window.
What does a paged KV cache buy you on a device?
Instead of one contiguous buffer per sequence (which must be reallocated and copied as it grows, or preallocated for the maximum), the cache is split into fixed-size blocks from a preallocated pool, mapped by a block table. Benefits: no large reallocation stalls, no fragmentation of big contiguous regions, memory proportional to actual length, easy sharing of prefix blocks between sequences, and an explicit, enforceable memory budget. Costs: indirection in the attention kernel and partly filled last blocks. It matters most for multi-session services.
What is prefix caching and when does it help?
If many requests start with the same tokens (system prompt, tool instructions, a document being questioned), compute their KV cache once and reuse it, so each request only prefills the new suffix. It cuts TTFT and energy, often dramatically when the shared prefix is long. On device you can persist the prefix cache to storage and mmap it. It must be invalidated when the model, quantization, prompt text or position handling changes, and costs storage equal to the KV size of the prefix.
What are good practices for a JNI bridge to a native inference engine?
Keep the interface coarse: create, run/generate, cancel, destroy with an opaque handle, rather than per-tensor calls. Pass large buffers as direct ByteBuffers to avoid copies. Delete local references in long loops, cache method ids, attach native worker threads to the JVM before calling back, and never hold JNI references across threads without global refs. Handle errors by returning status codes or throwing Java exceptions, not crashing. Make generation cancellable via an atomic flag, and ensure destroy is idempotent and thread-safe.
Describe a robust model download and install flow.
Fetch a signed manifest (id, version, URL, size, SHA-256, required runtime, SoC and RAM requirements). Check eligibility and free storage. Download in a background worker with constraints (unmetered, optionally charging) and HTTP range resume into a temporary file. Verify checksum and signature. Atomically move into a versioned directory. Run a smoke-test inference, then switch the active version pointer. Keep the previous version until the new one is proven, then garbage-collect. Handle "not yet downloaded" in the UI and expose a remote kill switch.
What would you put in a Perfetto trace for an inference feature?
Scheduling (sched_switch) to see which cores threads run on, CPU frequency and idle events, thermal events, the app's atrace categories and custom slices (Trace.beginSection / ATrace_beginSection) around load, pre-processing, prefill, decode steps and post-processing, the android.power data source with power rails and battery counters, and process stats for memory. Optionally heapprofd for native allocations. This lets you correlate model phases with frequency drops, thermal events and power.
What are the drawbacks of memory-mapped weights?
Page faults on first access add latency to the first inference unless you prefault. Under memory pressure the kernel may evict clean pages, and re-reading them from storage during decode causes large latency spikes. If the runtime repacks weights into a different layout at load, it allocates anonymous memory anyway, losing reclaimability and adding load time. Encrypted or compressed models cannot be mapped directly. Storage speed and file system also affect cold-load behaviour.
How do LoRA adapters work at inference time, and what is the trade-off between merged and unmerged?
A LoRA adapter adds a low-rank update B·A to selected weight matrices (for example attention projections), so the effective weight is W + (alpha/r)·B·A. Merged: fold the update into W once; no runtime overhead, but switching tasks means re-merging or keeping multiple full copies, and it complicates quantized weights. Unmerged: keep the small A and B matrices separate and compute the extra low-rank product each forward pass; a few percent overhead but instant hot-swap between adapters on one shared base model. Adapters are tied to one base model version and quantization.
Advanced
Why can't most NPUs do dynamic activation quantization, and what does static quantization cost?
Integer NPU pipelines precompute requantization parameters (multipliers and shifts combining input, weight and output scales) when the graph is compiled, and schedule data movement assuming fixed formats. Dynamic quantization needs a data-dependent reduction (min/max) over each activation tensor before the matmul, a synchronisation point that breaks streaming dataflow and requires flexible scalar logic. The cost of static ranges: outliers in production data clip, or ranges set wide to include outliers waste resolution for normal values. For transformers this is why NPUs use 16-bit activations (W4A16/W8A16), and why calibration data quality and outlier-handling methods (rotations, smoothing) matter so much.
Explain how an INT8 matmul is computed and requantized on integer hardware.
Weights and activations are stored as int8 with scales s_w, s_a and zero-points. The kernel multiplies int8 values and accumulates into int32 (subtracting zero-point terms, often precomputed into a bias correction). The real result equals s_w·s_a times the int32 accumulator. To produce int8 output with scale s_y, multiply by M = s_w·s_a/s_y, which is represented as an integer multiplier and a right shift (fixed-point), add the output zero-point, round and saturate. Bias is pre-quantized to int32 with scale s_w·s_a. Per-channel weights mean one M per output channel. Differences in rounding mode and saturation between implementations explain small host vs device mismatches.
What are activation outliers in LLMs and how do different methods handle them?
A few hidden dimensions carry values much larger than the rest, consistently across tokens. With per-tensor activation scales, they force a large scale and most values quantize to a few levels. Remedies: keep activations at 16-bit (NPUs) or quantize dynamically per token (CPUs); per-channel handling where hardware allows; SmoothQuant-style migration that divides activations by per-channel factors and multiplies weights by the same factors, moving difficulty into weights; rotation methods (QuaRot, SpinQuant) that multiply by orthogonal matrices to spread outlier energy across all channels, making both weights and activations easier to quantize; and mixed precision for the affected layers.
Derive an upper bound on decode speed and use it to diagnose a slow deployment.
Bytes read per token is roughly the weight bytes actually used per token plus the KV bytes at the current context. Decode tok/s is at most effective bandwidth divided by that. Example: 1B model at about 1 GB INT4 including scales and 8-bit embeddings (embeddings are only gathered, so the actual read is somewhat less), effective bandwidth about 45 GB/s, so the ceiling is about 45 tok/s. If you measure 15, you are far from the bound: suspect dequantization-heavy or scalar kernels, wrong thread placement, disabled KV cache or SDPA, fallback ops, frequency caps, or reading FP32 copies. If you measure 40, you are near the roofline and only fewer bytes (lower precision, smaller model) or speculative decoding will help.
Would you run prefill and decode on different processors? What are the trade-offs?
Prefill is compute-bound, so the NPU's high integer throughput cuts TTFT dramatically. Decode is bandwidth-bound, and all processors share the same DRAM, so NPU, GPU and CPU decode speeds are closer; the choice then comes down to energy per token, thermal behaviour and contention. Splitting phases across processors requires the KV cache in a format and memory both can access (shared buffers, same quantization and layout) or conversion costs at the handover, two sets of compiled kernels and more memory. Many shipping stacks keep both phases on the NPU for simplicity and energy, and use the CPU as fallback.
What is weight sharing between prefill and decode graphs, and why is it needed?
Static shapes force separate graphs for prefill (chunk of N tokens) and decode (1 token). If each graph embedded its own copy of the weights, memory and storage would double. Weight sharing compiles both graphs against a single set of weight buffers in the same context, so switching graphs costs nothing in memory. It requires both graphs to use identical quantization encodings for the shared weights, which constrains per-graph quantization choices.
What role does on-chip memory (VTCM) play in NPU performance?
VTCM is fast scratch memory next to the Hexagon vector and tensor units. The compiler tiles operators so that working sets (weight tiles, activation tiles) fit in VTCM, streaming data from DRAM via DMA while computing on previous tiles. Operators or graphs whose tiles do not fit spill to DRAM, adding bandwidth and latency. Large activation tensors (high-resolution images, long prefill chunks) and wide layers are typical spill sources. Mitigations: smaller prefill chunk sizes, model splitting, layout choices, and compiler options that control VTCM usage.
Why do FP16 overflows happen on GPU/NPU paths, and how do you fix them?
FP16 has a maximum of 65504 and limited precision. Transformer activations with outliers, attention logits before softmax, sums of squares in norms, and large accumulations can overflow or lose precision, producing inf/NaN or degraded outputs, even though the FP32 model is fine. Fixes: keep sensitive ops (norms, softmax, final layers) in FP32; use FP32 accumulation where supported; rescale (for example compute norms with a pre-scaling factor); use BF16 on hardware that supports it; or quantize with 16-bit integer activations which have well-defined ranges. Layer-wise diffing locates the first overflowing op quickly.
Does speculative decoding help on device, and what does it cost?
It helps because decode is bandwidth-bound: verifying k draft tokens in one forward pass of the target model reads the weights once for several tokens, so accepted tokens are nearly free. Expected tokens per pass = (1 − αk+1) / (1 − α); speed-up ≈ that / (1 + k · c). Speedups of 1.5-2.5x are common when the draft is accurate. Costs: a draft model's memory and its own KV cache (or extra heads for self-speculative methods), extra compute that raises power, complexity in cache rollback when drafts are rejected, and poorer gains on creative, high-entropy text. On NPUs, verification needs a static k-token graph. It pays off most on long, predictable outputs such as summaries or code.
Why does static KV allocation for the maximum context matter for memory planning?
NPU graphs and many optimized CPU runtimes allocate KV tensors at their maximum length because shapes are static. A 3B model compiled for 4k context reserves about 448 MiB of FP16 KV even for a 20-token question; at 16k it would be about 1.75 GiB. So choosing max context is a memory decision, not just a capability one. Options: compile several context variants and choose per request or device tier; use quantized KV; use paging on runtimes that support it; or cap context and use summarisation or retrieval to stay within it.
How does GQA change KV memory and NPU efficiency?
Grouped-query attention shares each K/V head among several query heads (for example 32 query heads and 8 KV heads in Llama-3.2-1B), cutting KV memory and KV bandwidth by the group factor (4x there) with small quality impact. For decode, less KV to read means higher tok/s at long contexts. For NPUs, implementations may broadcast K/V to match query heads (costing memory traffic) or reshape queries to batch the heads in a group; the exported attention layout affects whether the compiler maps it efficiently.
How do you design mixed precision under NPU constraints?
Start from layer sensitivity, but check what the backend supports in one partition: some NPU stacks support per-op precision (INT4 and INT8 weights, 8- and 16-bit activations) within a graph, others force a partition break or CPU fallback when precision changes, which can erase the benefit. Prefer promoting whole blocks or op types consistently, keep promoted ops on the NPU (for example W8A16 instead of FP32), keep encodings consistent across graphs that share weights, and re-profile placement after every change. Measure the cost in both latency and partition count, not just size.
How would you estimate the cost of graph partitioning?
For each boundary: data transfer time (tensor bytes over effective bandwidth, possibly two copies), format conversion (quantize/dequantize, layout transpose), synchronisation latency (a round trip to the NPU driver, often tens to hundreds of microseconds), plus lost pipelining. Multiply by the number of boundaries per inference (or per token for LLMs). If a model has 20 boundaries at 200 microseconds each, that is 4 ms of overhead, which can exceed the NPU compute time of a small model. The fix is removing boundaries (op rewrites, moving pre/post-processing out of the graph), not faster kernels.
How do you achieve zero-copy data flow between camera, GPU and NPU?
Use hardware buffers that all components can import: AHardwareBuffer (backed by dmabuf) from the camera or ImageReader, import them into the GPU (EGL/Vulkan) for resize and colour conversion into another hardware buffer, and pass that buffer to the NPU runtime through its shared-memory API (for Qualcomm, rpcmem/ION-dmabuf registered with QNN, or runtime-specific buffer interop). Avoid round-trips through Java arrays or CPU memcpy. Watch for format requirements (NHWC, alignment, quantized input types) that force a conversion, and ensure cache coherency and synchronisation fences between producers and consumers.
Design a shared on-device inference service used by several apps.
A bound system or privileged service exposing a versioned AIDL interface guarded by a permission. One base model resident (mmapped), with per-task LoRA adapters loaded on demand. A scheduler with a request queue, per-client quotas, priority for foreground callers, streaming callbacks (avoid large Binder transactions), cancellation and linkToDeath cleanup. A memory governor using PSI and onTrimMemory to shrink KV budgets, evict adapters or unload the model. A thermal governor using thermal headroom to pace or refuse. A backend policy NPU, GPU, CPU, refuse. Safety filtering on inputs and outputs, input size limits, and telemetry. Updates delivered as deltas and swapped atomically.
Explain Android memory accounting relevant to model deployment.
RSS counts all resident pages of a process, including shared ones; PSS divides shared pages among sharers and is what dumpsys meminfo reports as the app's footprint; USS is private-only. Anonymous memory (heap, repacked weights, KV cache) can only be reclaimed by compressing into zRAM (swap), which costs CPU and compresses quantized data poorly. File-backed clean pages (mmapped weights) can be dropped and re-read. lmkd decisions depend on overall pressure and oom_score_adj, not your RSS directly. GPU and NPU memory may be accounted under graphics or dmabuf and missed if you only watch heap.
How would you measure the maximum usable context on an 8 GB phone under realistic pressure?
Create a realistic background: a foreground workload (for example a memory-heavy app or a synthetic allocator holding a typical footprint) plus normal services. Run generation while increasing context in steps (1k, 2k, 4k...), recording PSS, PSI, zRAM usage, kills (am_kill events, lmkd logs) of your process and others, and decode speed. The usable limit is the largest context without killing the foreground app and without killing important background apps or severe PSI stalls. Repeat for FP16, INT8 and INT4 KV and sliding-window configurations to show how each technique moves the limit.
How do big.LITTLE scheduling and thread affinity affect CPU inference?
Phones combine prime, performance and efficiency cores with very different speeds. Parallel matmuls split work evenly across threads, so a thread on an efficiency core becomes the straggler that everyone waits for. Use as many threads as performance-class cores, pin or hint them to those cores where the runtime allows, and avoid oversubscription with UI and rendering threads. Performance hints (ADPF performance hint sessions) tell the scheduler your target work duration so it can choose appropriate frequencies. Spin-waiting thread pools can burn power; tune spin times for decode.
How do you keep model builds reproducible across SDK and runtime versions?
Pin every tool version (framework, exporter, quantizer, vendor SDK) in a locked environment or container; make export scripts deterministic (fixed seeds, fixed calibration set with a hash); record a manifest per artifact (source checkpoint hash, recipe, tool versions, target SoC, runtime version required); store artifacts in a registry keyed by hash; ship the matching runtime libraries with the artifact; run a numeric regression test (outputs on fixed inputs within tolerance) and a performance regression test on a device farm for every build. Invalidate on-device compiled caches when any version changes.
How do you evaluate a quantized LLM reliably against its reference?
Use several complementary signals: perplexity on held-out, domain-relevant text; next-token top-1 agreement and KL divergence against the FP32 model on a fixed prompt set (teacher-forced, so errors do not compound); greedy decoding comparison measuring the first divergence position; task benchmarks relevant to the product (including structured output validity, arithmetic, multilingual); and, for product features, a rubric-based or human evaluation on real prompts. Report confidence intervals, and evaluate on the device output, not only host simulation.
Why does unstructured pruning rarely speed up edge inference, while structured pruning can?
Unstructured pruning zeroes individual weights; unless sparsity is very high and hardware or kernels support sparse formats, dense kernels still read and multiply the zeros, and index overheads can make sparse kernels slower. Structured pruning removes whole channels, heads or layers, producing a smaller dense model that every runtime accelerates directly. Semi-structured patterns (like 2:4) help only on hardware with dedicated support. On phones, structured pruning plus distillation or simply choosing a smaller model is usually the practical route.
How do runtimes decide which ops go to an accelerator, and how can you influence it?
A partitioner walks the graph, asks the backend whether each node (with its data types, shapes and attributes) is supported, groups contiguous supported nodes into subgraphs (respecting dependencies and sometimes minimum partition sizes), and replaces each with a delegate call. You influence it by rewriting unsupported ops into supported equivalents before export, fixing data types (int64 to int32, FP32 to quantized), making shapes static, using the backend's quantizer so encodings are compatible, configuring partitioner options (skip lists, precision), and moving pre/post-processing out of the graph.
How do you protect a valuable on-device model?
Accept that anything that runs on a user's device can ultimately be extracted by a determined attacker with root. Raise the cost: download at runtime instead of bundling, store in app-private storage, encrypt at rest with keys from the Android Keystore and decrypt into memory (losing mmap benefits), verify integrity and signatures before loading, use runtime integrity checks for the app, split the most valuable part to the server, and use licensing and legal measures. For compiled NPU binaries, the format itself offers some obfuscation but not security.
How do binary delta updates for models work, and when do they make sense?
Compute a binary diff (bsdiff-style or chunk-based) between the installed and new model files on the server; the device downloads the patch and reconstructs the new file, verifying its hash. It saves bandwidth when changes are localized, such as a fine-tuned adapter or partially changed layers. Quantized weights after re-quantization often change almost everywhere, reducing delta effectiveness; chunk-aligned formats and stable layouts help. Patching needs temporary storage for both versions and CPU time, so run it while charging and idle.
When does a mobile GPU beat the NPU?
When the model uses ops, shapes or precisions the NPU does not support well (dynamic shapes, unusual attention variants, FP16-only accuracy requirements), when a model changes often and NPU compilation friction is too high, for moderate-size models where the GPU's FP16 throughput is enough, and for workloads already on the GPU (image processing, rendering) where staying on the GPU avoids transfers. The NPU generally wins on energy efficiency and sustained performance for large quantized models, especially prefill.
How would you design a fair CPU vs GPU vs NPU comparison?
Same model, same inputs, same prompt and output lengths, each backend with its best realistic quantization (documented, with the accuracy delta measured against FP32), same thermal starting point, warm-up and repetitions. Report load/compile time, TTFT, prefill and decode tok/s, peak memory, energy per 100 tokens with idle subtraction, and a sustained 10-minute curve for each. Note what runs where (partitions, fallback ops) and include pre/post-processing time in an end-to-end number. Publish configuration, versions and raw data so others can reproduce it.
How do RoPE positions interact with sliding windows and cache eviction?
With rotary position embeddings, positions are applied to keys before they enter the cache, so cached keys already encode their absolute positions. If you evict middle tokens or wrap a window, the relative distances between the new query and remaining keys stay consistent as long as you keep using the true absolute positions, but positions can grow beyond the trained range in long streams. Some streaming methods instead assign positions within the cache (re-indexing) and store keys before rotation, applying RoPE at attention time, which keeps positions within range but costs extra compute. Getting this wrong shows up as degradation after the window first fills.
Scenario & debugging
The model is 3x slower on the NPU than the vendor's published numbers. How do you investigate?
- Match conditions: same model variant, precision, input size or prompt length, SoC, SDK version, and whether they reported compute-only time.
- Check placement: per-op profile (QNN profiler, AI Hub profile, runtime op profiling) to find CPU fallback ops and count partitions.
- Check initialization: is graph compilation or context generation counted in each run? Use a precompiled context binary or cache.
- Check performance mode: burst or sustained high performance rather than default or power saver; check that the NPU is not shared with camera or other clients.
- Check data movement: copies between Java and native, quantize/dequantize of inputs and outputs on CPU, layout transposes (NCHW vs NHWC) outside the graph; use shared buffers.
- Check precision: an FP32 model running as FP16 or partially on CPU; ensure the model is quantized in the format the HTP expects.
- Check thermal state and spill: DDR spill from VTCM for large tensors; try smaller tiles or chunk sizes.
Fix the largest contributor, re-profile, and repeat.
Accuracy dropped by 6 points after INT8 quantization. What do you do?
First rule out non-quantization causes: compare the converted FP32 model with the original (conversion bug?) and confirm the evaluation pipeline and pre-processing are identical. Then inspect the calibration set (size, representativeness, pre-processing identical to inference). Run layer-wise diffs (SQNR, cosine) and single-layer sensitivity to find where error originates. Typical fixes in order: per-channel weights, better calibration method (percentile or entropy instead of min-max), cross-layer equalisation and bias correction or AdaRound for CNNs, 16-bit activations for sensitive layers, mixed precision for the worst few layers, and QAT if PTQ cannot close the gap. Re-validate on device, because device kernels can differ from host simulation.
The phone throttles after two minutes of LLM generation and tok/s drops 40%. What can you do?
Measure first: thermal soak with temperatures, frequencies, power per rail and throughput to confirm it is thermal and find the dominant power consumer. Then reduce energy per token: move to the NPU (lower power per token than CPU), use lower precision weights and KV, reduce thread count (fewer cores at lower power sometimes sustain better), avoid spin-waiting thread pools, cut wasted work (stop tokens, shorter outputs, prefix caching to avoid repeated prefill). Manage the budget: use thermal headroom APIs to pace generation proactively, pick a sustained performance mode, and degrade gracefully (smaller model or shorter answers) when hot. Test with the case on at realistic ambient temperature, and set product expectations on sustained, not peak, numbers.
The app gets killed when a conversation grows past about 3k tokens on 8 GB phones. Why, and how do you fix it?
The KV cache and activation buffers grow with context (or were allocated for a large max context), pushing total anonymous memory high enough that lmkd kills processes; logcat lmkd messages and PSI confirm it. Fixes: calculate and enforce a memory budget per tier; quantize the KV cache (INT8 halves it); cap context for this tier and use sliding window with sinks, summarisation of older turns, or retrieval; use a paged cache to avoid reallocation peaks; mmap weights so they are reclaimable; free caches on onTrimMemory; make sure there is no duplicate copy of weights (compressed asset, repacking). Verify with the pressure test at several contexts.
The first inference after launch takes 8 seconds. How do you reduce it?
Break it down with trace slices: file read or decompression, weight repacking, delegate or NPU graph compilation, GPU shader compilation, tokenizer load, first-run page faults. Fixes: store uncompressed and mmap; ship precompiled context binaries or enable runtime compiled-model caching (EP context, GPU serialization caches) so compilation happens once; move weight repacking offline into the exported format; initialize asynchronously at app start or on a trigger that predicts use; warm up off the critical path; split the model so a small part is ready first. Measure cold (after reboot) and warm separately.
Decode tok/s is half of the reference number for the same model and phone. What do you check?
Release vs debug build; KV-cache and SDPA flags enabled in export; the right quantization (INT4 weights, not an FP32 fallback); optimized kernels enabled (KleidiAI, dotprod/i8mm); thread count equal to performance cores and threads not landing on efficiency cores; device temperature at start; background load; prompt and output lengths matching the reference; power-saving mode off; and whether the reference measured with a different context length. Profile with simpleperf to see which kernels are hot. Compare with llama.cpp as a second opinion on the same device.
On device the LLM outputs gibberish or never stops, but on the host it is fine.
Most likely a tokenizer or prompt issue: different tokenizer file, missing BOS, wrong special token ids, chat template not applied, or EOS ids missing from the runner's metadata (so it never stops). Then check the export: KV-cache positions, max sequence length exceeded, RoPE scaling parameters. Then numerics: run greedy decoding with identical token ids on host and device and compare logits at the first step; if logits differ strongly, do a layer-wise diff to find an FP16 overflow or a miscompiled op. Also check sampling parameters (temperature, top-k) are what you expect.
The feature works on your flagship but crashes on a mid-range phone.
Collect the native crash (tombstone) and logcat: common causes are out-of-memory during load (a model too large for the RAM tier, or compressed assets doubling memory), missing CPU instructions (a library built with i8mm or other extensions on a core without them), an unsupported accelerator path that is not guarded (NPU libraries for another Hexagon version), or GPU driver bugs. Fix with device-tier gating, runtime CPU feature detection and multiple kernel variants, capability probes before enabling a backend, graceful fallback, and a test matrix including low and mid-tier devices.
Outputs are fine for short prompts but quality degrades after 1-2k tokens.
Suspect the cache and positions: KV cache index or wrap-around bugs, sliding window evicting sink tokens, RoPE position handling when the window rolls, exceeding the compiled max context, or KV quantization error accumulating with length (especially keys quantized per token instead of per channel). Also check whether the model itself was trained for that context (and whether RoPE scaling is applied in the exported model). Reproduce with FP16 KV and full attention to isolate, then compare logits at increasing positions against the host reference.
The INT4 model runs at the same speed as the FP16 model. Why?
The kernels may be dequantizing weights to FP16/FP32 into a temporary buffer and then running a dense float matmul, so bandwidth is not reduced; or the quantized ops fell back to reference (scalar) kernels because the scheme or group size is not supported by the optimized path; or the time is dominated by something else (unquantized embeddings or LM head, attention with a large KV cache, pre/post-processing, CPU fallback partitions); or you are compute-bound in prefill where INT4 weights help less without integer compute. Profile to see which kernels run and whether bytes read per token actually dropped.
The NPU session fails to load on one SoC generation but works on others.
The context binary was built for a different Hexagon architecture, the DSP skeleton libraries for that architecture are missing or not on ADSP_LIBRARY_PATH, the runtime library version differs from the SDK that built the binary, or the device's firmware/driver is older than required. Check logcat for FastRPC or QNN errors. Fix by shipping per-architecture binaries and libraries selected at runtime from the SoC model, pinning SDK versions, and falling back to GPU/CPU when the probe fails.
The camera feature drops frames even though the model runs in 5 ms.
The model is not the bottleneck; the pipeline is. Trace the full frame path: YUV to RGB conversion and resize on the CPU, copying frames into Java arrays, allocating buffers per frame, running post-processing (NMS, mask upsampling) on the CPU, synchronously waiting on the GPU, or rendering overlays on the main thread. Fix with GPU or ISP-based pre-processing, hardware buffers and zero-copy, buffer reuse, pipelining stages across frames, dropping stale frames rather than queueing, and native post-processing. Measure end-to-end latency and p99 frame time, not model time.
After shipping, users complain about battery drain. How do you find and fix the cause?
Segment telemetry by device and usage: how often the model runs, on which backend, for how long, and whether it runs in the background. Reproduce with batterystats and power rails. Common causes: inference running more often than needed (every frame when every fifth would do), CPU fallback instead of NPU, spin-waiting thread pools, repeated model loading, generation not cancelled when the user leaves, or background work not constrained. Fix with duty cycling, event triggers, batching, lower frame rates, NPU placement, cancellation, WorkManager constraints, and track energy per request as a release gate.
p99 latency spikes periodically, even though p50 is stable.
Correlate spikes in a Perfetto trace with system events: thread migration to efficiency cores, CPU frequency drops, thermal mitigation steps, garbage collection pauses in the Java layer, page faults from evicted mmapped weights, memory allocation in the hot path, contention with rendering or other apps, or periodic background jobs. Fixes: preallocate buffers, prefault or lock weights, use performance hints, avoid allocations per inference, move work off threads that compete with the UI, and use sustained performance modes for continuous workloads.
The same model gives noticeably different results on two phone models.
Different backends or kernels run on each: one may use the NPU with static INT8 and another the GPU with FP16 or the CPU; vendor drivers implement ops with different rounding, accumulation precision or approximations (for example exp and softmax); FP16 overflow may occur on one GPU and not another. Log the backend used per device, compare intermediate outputs on a fixed input, and test the same backend on both. Mitigate by constraining precision for sensitive ops, pinning backends for critical features, and setting accuracy tolerances per device family.
A new model rollout increased crash rate, but only on one OS version.
Halt the staged rollout (or flip the kill switch for that cohort) immediately. Symbolize the native crashes and check whether they come from the runtime or driver libraries; an OS update may have changed the GPU/NPU driver or a system library the runtime depends on, or invalidated a compiled cache format. Reproduce on that build, add a device/OS deny-list or a different backend for it, report to the vendor with a minimal repro, and add that OS version to the pre-release test matrix. Longer term, validate caches against driver versions.
The quantized LLM has fine perplexity but fails at arithmetic and producing valid JSON.
Perplexity averages over typical text and hides skill-specific damage; digits, brackets and rare formatting tokens can be disproportionately affected, often via the LM head or embedding quantization and outlier-heavy layers. Build targeted evaluations (arithmetic set, JSON schema validity, your product prompts), run layer sensitivity with those tasks as the metric, keep the LM head and sensitive layers at higher precision, try better PTQ (GPTQ, AWQ, rotations) or QAT with LoRA on task-relevant data, and use constrained decoding (grammar-guided sampling) for JSON so the format is guaranteed.
The GPU delegate is slower than the CPU for your model.
Likely causes: the model is small, so GPU dispatch and synchronisation overhead dominate; some ops are unsupported and fall back to CPU, creating copies; data transfers of inputs and outputs between CPU and GPU memory each call; shader compilation included in timing; quantized INT8 models running on a GPU path that dequantizes to FP16; or the GPU is busy rendering. Check delegation coverage and partition count, exclude initialization, use GPU buffers directly, enable serialization caches, and consider that the CPU with XNNPACK is often the right choice for small models.
Peak memory at load is twice the model size.
Typical causes: the asset is compressed so it is decompressed into memory in addition to being read; the runtime reads the file into a buffer and then repacks weights into another layout, keeping both until load finishes; a Java byte array copy plus a native copy; a GPU delegate uploading weights while CPU copies are still resident. Fix with uncompressed mmapped files, offline pre-packing into the runtime's preferred layout, releasing source buffers after upload, streaming loads, and measuring with heapprofd and dmabuf accounting to confirm.
Product wants 8k context with a 3B model on 8 GB phones. What is your plan?
Do the budget: 3B INT4 weights about 2 GB, FP16 KV at 8k about 900 MB, plus activations and runtime overhead, roughly 3.5 GB total, which is risky on 8 GB phones with a foreground app. Options: INT8 or INT4 KV (450 or about 250 MB), a sliding window with sinks for chat history plus retrieval for documents, summarising older turns, prefix caching of the system prompt to save TTFT, or a smaller model for this tier with 8k context. Validate with the memory-pressure test and thermal soak, and propose a tiered plan: full 8k on 12 GB+, 4k plus retrieval on 8 GB.
A product manager asks to run a 7B model on mid-range phones. How do you respond?
Quantify it: 7B at INT4 is about 4 GB of weights, plus KV and runtime, over 5 GB resident; decode on a mid-range phone with maybe 25-35 GB/s effective bandwidth is at most around 6-8 tok/s, falling with context and heat; energy per response and load time are high; and many mid-range devices lack an NPU path for it. Offer alternatives tied to the product goal: a 1-3B model fine-tuned or distilled for the specific task, LoRA adapters, retrieval to supply knowledge, a cascade where hard queries go to the cloud with consent, or 7B only on high-end tiers. Back the recommendation with a quick prototype measurement.
The streaming chat UI janks while tokens are being generated.
Check that inference is not on the main thread and that token callbacks do not trigger heavy work per token (full text re-layout, markdown parsing of the whole message, list diffing). Batch UI updates (for example every 50 ms or every few tokens), append incrementally, and do text processing off the main thread. Also check CPU contention: inference threads saturating all performance cores can starve the render thread; leave a core free, lower inference thread priority, or move inference to the NPU. Verify in a trace that frames meet their deadlines.
ONNX Runtime with the QNN EP is silently running most of the model on the CPU.
Enable verbose logging and set session.disable_cpu_ep_fallback to make unsupported nodes fail loudly and name them. Typical reasons: the model is not QDQ-quantized in the format the HTP expects; unsupported ops or data types (int64, dynamic shapes, certain reductions); shapes not static; the QNN backend library not found so the EP was not registered. Fix by running the QNN preprocessing and quantization helpers, making shapes static, rewriting unsupported ops, and verifying placement through the profile before re-enabling fallback.
A vision model is accurate in lab tests but poor with the real camera.
Dump the exact tensors the app feeds the model and compare with the lab pipeline: colour order, normalisation, resize and crop method, rotation from sensor orientation, YUV conversion ranges (full vs limited), and aspect-ratio handling. Then consider domain shift: lighting, noise, motion blur and lens differences not present in the dataset, and a calibration set that did not include real camera frames. Fix the pipeline mismatch first; then recalibrate with device-captured data, augment or fine-tune with such data, and build a device-captured evaluation set.
Three apps from your company each want their own LLM. What do you propose?
Three separate copies would triple memory and storage and fight for the NPU. Propose one shared service hosting a single base model, with per-app LoRA adapters swapped at request time, a stable AIDL API with permissions, a scheduler with quotas and priorities, and memory and thermal governors. If the OS provides a platform AI service with a suitable model, evaluate using it instead. Measure adapter swap cost and service overhead against in-process inference, and define versioning so adapters are retrained when the base model updates.
A wearable activity model drains the battery through false wake-ups.
Measure false wake-ups per hour and energy per wake-up, since waking the application processor costs far more than the model. Move the first-stage detector to the sensor hub or low-power core with a tiny model and a stricter threshold, add a second-stage confirmation model on the main processor, batch sensor data in hardware FIFOs, tune thresholds on realistic day-long recordings (not only labelled activity clips), add hysteresis or temporal smoothing, and evaluate per user. Track battery impact per day as the release metric.
Swapping LoRA adapters takes 1.5 seconds, which is too slow.
Break down the time: reading the adapter file, converting or quantizing it at load, merging into base weights, or recompiling an NPU graph. Fixes: keep adapters unmerged and apply them as separate low-rank ops so swapping is a pointer change; preload likely adapters into memory; store adapters in the runtime's final format and precision; on NPUs, compile graphs with adapter weights as updatable inputs rather than constants so swapping does not require recompilation; and cache per-adapter prefix KV if system prompts differ per task.
Your benchmark numbers vary by 20% between runs on the same device.
Control the sources of variance: start temperature (cool down to a threshold before each configuration), battery level and charging state (charging adds heat and may change governor behaviour), screen state and brightness, background activity (disable sync, use airplane mode), thread placement (pin or use performance hints), and prompt differences. Increase repetitions and report median and spread. Check for thermal throttling within the run and for memory pressure causing page eviction. If variance remains, report it honestly with confidence intervals.