Android Platform

Trace a Path Through the Android Stack

Interviewers love "trace X from top to bottom." This page is that skill: the same five layers, then one walk each for voice, input, display, Binder, launch, OTA, audio, camera and location. The cellular packet path is only a map here; the full control-plane and netd/eBPF walk is Android Data Call. Two watch examples (sensor hub and Data Layer) show the same cake on Wear OS; the product page is Wear OS Platform.

~72 min read 0 interview questions
In 30 seconds
  • Almost every flow crosses the same five layers: app, framework (system_server), native daemons and HALs, kernel, and hardware or firmware.
  • Between app and framework the glue is Binder; between framework and vendor code it is a stable HAL interface (AIDL, or HIDL on older devices); between user space and kernel it is syscalls, ioctl and interrupts.
  • The high-frequency traces are voice, input, frames, Binder, launch, OTA, audio, camera and location. Cellular packets have their own page. Sensor-hub and Data Layer are Wear examples of the same layers.
  • Performance and power problems usually come from one hop doing too much work: a blocked main thread, a slow Binder call, a sensor read that bypasses batching, or a data path that keeps the CPU awake.
  • For any "it's broken" question: name the flow, walk the layers, say which tool you would read at each layer, then bisect to the failing layer.

The layer cake every flow crosses

Almost every data journey on an Android device moves up or down the same vertical stack. Learn the stack once, and every specific flow becomes a question of "which boxes, in which order, over which interface."

┌─────────────────────────────────────────────────┐
│ APP            (your process, ART runtime)      │  Java/Kotlin, SDK APIs
├─────────────────────────────────────────────────┤
│ FRAMEWORK      (system_server services)         │  AMS, WMS, Telephony, SensorService
│                ↕ Binder IPC (AIDL)              │
├─────────────────────────────────────────────────┤
│ NATIVE / HAL   (daemons + vendor HALs)          │  SurfaceFlinger, rild, Sensors HAL
│                ↕ AIDL/HIDL HALs, sockets, FMQ   │
├─────────────────────────────────────────────────┤
│ KERNEL         (drivers, subsystems)            │  TCP/IP, input, DRM/KMS, IPA, binder
│                ↕ syscalls / ioctl / IRQ         │
├─────────────────────────────────────────────────┤
│ HARDWARE / FW  (SoC, modem, sensors, radio)     │  modem (MPSS), sensor hub, panel, GPU
└─────────────────────────────────────────────────┘
Analogy

Think of a large office building. The customer (app) talks to the front desk (framework), which passes the request to a specialist department (native daemon or HAL), which uses the building's plumbing and wiring (kernel) to reach the machines in the basement (hardware and firmware). In Android, the front desk is system_server, the specialist departments are daemons and HALs, the internal mail system between floors is Binder, and the plumbing is the kernel's drivers and syscalls.

BoundaryMechanismTypical example
App ↔ frameworkBinder IPC with AIDL-generated proxies and stubsActivityManager.startActivity() calls into AMS
Framework ↔ vendor HALStable AIDL HAL (Android 11+) or HIDL over hwbinder (Treble era); checked by VINTFTelephony to IRadioVoice, SensorService to ISensors
HAL / daemon ↔ kernelsyscalls, ioctl, sysfs, netlink, /dev nodesSurfaceFlinger to DRM/KMS via the HWC HAL; Radio HAL to modem via QMI (QRTR on modern Qualcomm SoCs)
Kernel ↔ hardwareMMIO registers, DMA, interrupts, shared memory with co-processorstouch controller IRQ, modem shared memory, IPA descriptors
Tip Whenever you are asked to trace something, say the layers out loud in order and name one concrete component per layer. That structure alone sounds senior. Deeper dives: Android frameworks, Binder and AIDL, Telephony, RIL and modem.
Interview angle "Trace X end to end" questions test whether you have a mental model or just memorised API names. A strong answer names the layers, the component at each layer, the interface between them, and at least one thing that commonly goes wrong at a boundary (a HAL version mismatch, a blocked Binder thread, a missing SELinux permission).

A voice call: dial to RF

Here is a mobile-originated VoLTE call, from the moment the user taps "call" to radio waves leaving the antenna. The same path applies on a standalone LTE watch, tuned for power.

Dialer app ──▶ TelecomManager.placeCall()
  ──▶ Telecom (system service)            picks the PhoneAccount, owns call state/audio routing
  ──▶ TelephonyConnectionService          creates a Connection
  ──▶ Telephony (GsmCdmaPhone / ImsPhone) decides CS vs IMS; VoLTE goes to ImsPhone
  ──▶ ImsService (vendor IMS stack)       builds SIP INVITE with SDP offer
  ──▶ RIL (RIL.java) ──Binder──▶ Radio HAL (IRadioVoice / IRadioIms)
  ──▶ vendor Radio HAL process (classically rild) ──QMI over QRTR or shared memory──▶ modem (MPSS)
  ──▶ modem: SIP INVITE over the IMS bearer ──▶ eNB/gNB ──▶ IMS core (P/S-CSCF) ──▶ callee
  ◀── 100 Trying / 183 / 180 Ringing / 200 OK ◀── ACK
  ◀══ media: RTP/RTCP voice over a dedicated QoS bearer (QCI 1 on LTE, 5QI 1 on 5G)
Analogy

Placing a call is like booking a table and then eating the meal. First you phone the restaurant to reserve (signalling); only once the booking is confirmed do the courses arrive on a separate, reserved lane (media). In VoLTE, the reservation is the SIP exchange through the IMS core, and the reserved lane is the dedicated QoS bearer that carries RTP audio with guaranteed bit rate and low latency.

Signalling plane

SIP (INVITE, 180 Ringing, 200 OK, ACK, BYE) carries call setup and teardown. SDP inside SIP negotiates the codec (AMR-WB, EVS), ports and precondition QoS.

Media plane

RTP carries voice frames, RTCP carries quality reports. The network creates a dedicated bearer for them after SDP negotiation so voice is protected from best-effort data.

Framework ownership

Telecom owns calls from the user's point of view (UI, audio routing, multiple call apps). Telephony owns the radio side (service state, IMS registration, RIL).

Fallbacks

If IMS is not registered, the call can go circuit-switched: CSFB from LTE to 3G/2G, or EPS fallback from 5G NR to LTE for VoNR-less networks. Emergency calls have their own routing rules.

Note Older devices use the HIDL IRadio interface (android.hardware.radio@1.x). From Android 13 the Radio HAL is AIDL and split into service-specific interfaces such as IRadioVoice, IRadioData, IRadioNetwork, IRadioSim, IRadioMessaging, IRadioModem and IRadioIms. More in Telephony, RIL and modem.
Interview angle Interviewers check that you separate signalling from media, know where Telecom ends and Telephony begins, and can place the vendor boundary (Radio HAL). Bonus points for mentioning IMS registration as a precondition, CSFB or EPS fallback, and which logs you would pull (logcat for Telecom/Telephony/IMS, modem logs for the air interface, a SIP trace).

A mobile-data packet: socket to PDN and back

How an app's HTTPS request leaves the phone over cellular and how the response comes back. Split it into two planes or you will mix signalling with payload. The control plane leases a data call (PDN connection on LTE, PDU session on 5G). The data plane is every packet after that: app socket, kernel, netd policy, eBPF, the cellular driver, the modem, the radio, the core, the server, and the reverse. This section is the map. The hop-by-hop walk, including netd, eBPF, fwmark, 464XLAT, and both directions, is on Android Data Call: Control, netd, eBPF and Packets.

Control plane: leasing the data call

  1. Need a network An app opens a socket on the default network, or calls ConnectivityManager.requestNetwork(). ConnectivityService matches NetworkRequests against network offers (cellular, Wi-Fi, VPN).
  2. Telephony policy The data stack (DcTracker on older releases; DataNetworkController / DataNetwork from Android 13) picks an APN or DNN, checks data-enabled, roaming, SIM state and carrier policy.
  3. RIL setupDataCall() Telephony talks to IRadioData. The modem runs NAS + RRC signalling: default EPS bearer (LTE) or PDU session establishment (5G). The network returns IP, DNS, MTU, and a CID.
  4. Interface and netd A kernel interface appears (often rmnet_dataX on Qualcomm, ccmni on MediaTek). netd programs addresses, policy routing, DNS, and firewall / eBPF rules. ConnectivityService registers a NetworkAgent, scores the network, and runs validation (HTTPS probe). Only then is the network the default for apps.
App / ConnectivityService
  -- Telephony DataNetworkController (APN, policy)
  -- IRadioData.setupDataCall() -- modem NAS/RRC
  -- PDN / PDU session -- P-GW or UPF assigns IP
  -- kernel iface (rmnet_dataX / ccmni)
  -- netd: IP, ip rule, routes, DNS, eBPF/firewall
  -- ConnectivityService validates -- default Network

Data plane: uplink (app to server)

App  HttpURLConnection / OkHttp / Socket
  -- ART / libcore -- bionic connect/send/write
  -- kernel socket + TCP/IP (sk_buff)
  -- eBPF cgroup/skb + netfilter (UID, Data Saver, Doze)
  -- fwmark + ip rule -- routing table for this Network
  -- 464XLAT (clat) if IPv6-only cellular
  -- rmnet / CCMNI -- QMAP aggregation
  -- IPA / vendor offload (optional)
  -- modem PDCP/RLC/MAC/PHY -- eNB/gNB
  -- S-GW+P-GW or UPF (GTP-U) -- Internet -- server

Data plane: downlink (server to app)

Server -- Internet -- P-GW / UPF -- GTP-U -- RAN -- modem
  -- IPA filters/aggregates, DMA to AP (one IRQ per batch)
  -- rmnet de-QMAP -- NAPI / GRO / softirq
  -- eBPF / netfilter (UID stats, firewall)
  -- TCP stack -- socket wait queue / epoll
  -- bionic recv -- app
Analogy

The data call is a leased private lane: APN is the destination, the IP is the badge, rmnet_dataX is the lane number. netd is the traffic office that paints the signs (routes) and installs cameras (eBPF). IPA is an automatic toll plaza that sorts trucks without waking the office. Packets are the trucks; NAS signalling that leased the lane is a different conversation and must not be mixed into this walk.

PieceRole
ConnectivityServiceScores networks, matches requests, validation, Wi-Fi vs cellular fallback, VPN overlay.
netdNative daemon: interfaces, policy routing, DNS, firewall, eBPF load, per-UID fwmark.
eBPFModern replacement for much of xt_qtaguid: UID accounting, cgroup skb filters, tethering, CLAT offload. Loaded by netd / bpfloader.
fwmark / ip ruleHow Android sends one UID or one Network onto a specific routing table instead of the main table.
rmnet / CCMNIVendor net driver: one physical modem link, several logical data calls as separate interfaces.
IPA (and peers)Hardware offload for routing, NAT, filter, aggregation so the AP can sleep during bulk transfer.
Tip "Data connected but no internet" is almost never the modem alone. Bisect: CID and IP on the iface, routes and ip rule, validation in dumpsys connectivity, DNS, eBPF/firewall UID rules, then tcpdump vs modem logs. Full playbook: Android data call.
Interview angle Walk control plane and data plane separately, name netd and eBPF (not only rmnet and IPA), and say downlink is not "the reverse with the same interrupts": IPA batches so the AP sees aggregated IRQs. Link radio signalling detail to Telephony, RIL and modem and 5G NR.

A sensor sample: PPG to app (wearable)

A heart-rate reading on a watch starts as light reflected through the skin (photoplethysmography, PPG) and ends as a number in a fitness app. The whole design goal is to do as much as possible on the low-power sensor hub and wake the main application processor (AP) as rarely as possible.

PPG / accelerometer / gyroscope silicon
  ──▶ sensor hub firmware on the always-on (AON) low-power core
        (Qualcomm: SLPI / ADSP running the Sensors Execution Environment)
        sampling, filtering, sensor fusion, FIFO batching  ─ AP stays asleep
  ──▶ AP woken only when the batch is full or max report latency expires
  ──▶ Sensors HAL (AIDL ISensors, events via Fast Message Queue)
  ──▶ SensorService (framework, in system_server)
  ──▶ SensorManager in the app  /  Health Services (Wear OS)
  ──▶ app: SensorEventListener, or ExerciseClient / PassiveMonitoringClient (batched)
Analogy

The sensor hub is like a night-shift security guard who writes down every event in a logbook and only wakes the manager when the page is full or something important happens. The manager (the AP) reads the whole page in one go and goes back to sleep. On the watch, the logbook is the hardware FIFO, the page size is the batch size and max report latency, and "something important" is a wake-up sensor event such as a detected gesture.

ConceptWhat it means
Sampling periodHow often the sensor produces a sample (for example 25 Hz for PPG).
Max report latencyHow long events may wait in the FIFO before delivery. A non-zero value enables batching.
Wake-up vs non-wake-up sensorWake-up sensors can wake the AP from suspend; non-wake-up sensors wait until the AP is awake for another reason (or drop old data if the FIFO overflows).
Fast Message Queue (FMQ)Shared-memory queue used by Sensors HAL 2.x and AIDL to push events to SensorService without a Binder call per event.
Health ServicesWear OS API that wraps sensors into exercise, passive monitoring and measure clients, and handles batching and power for the app.
Common pitfall An app that registers a high-rate listener with zero report latency, or holds a wake lock to poll a sensor, bypasses batching. The AP wakes dozens of times per second and battery life collapses. This is a classic wearable systemic bug, and it shows up in wakeup counts and suspend residency, not in crash reports.
Interview angle Interviewers want the power argument, not only the component list: where batching happens, what wakes the AP, and how you would prove a sensor is draining battery (dumpsys sensorservice for active clients and rates, batterystats and wakeup sources for wakeups). See Wear OS and Power and thermal.

A touch or input event: panel to view

A finger touches the screen, and a few milliseconds later a button's onClick() runs. The input pipeline is short but latency-critical.

Touch panel ──▶ touch controller raises an IRQ
  ──▶ kernel input driver ──▶ /dev/input/eventX  (evdev)
  ──▶ InputFlinger in system_server:
        EventHub (reads evdev) ──▶ InputReader (cooks raw events into MotionEvents)
        ──▶ InputDispatcher (finds target window using window info from WMS/SurfaceFlinger)
  ──▶ InputChannel (a Unix socket pair) ──▶ app's UI thread Looper
  ──▶ ViewRootImpl ──▶ DecorView ──▶ View.dispatchTouchEvent() ──▶ onTouchEvent / onClick
  ◀── app sends a "finished" signal back on the InputChannel
Analogy

The input pipeline is like a hotel reception taking a phone call and transferring it to the right room. The switchboard (InputDispatcher) knows which guest is in which room (window focus and position) and rings that room; if the guest never picks up, reception eventually reports a problem. In Android, the room phone is the InputChannel socket, the guest is the app's main thread, and "never picks up" becomes an Input Dispatch Timeout ANR after about 5 seconds.

Why latency lives here

The dispatcher-to-app hop crosses a socket and then waits for the app's main Looper. Any slow work queued ahead on the main thread delays input handling.

ANR detection

InputDispatcher tracks events that the app has not acknowledged. If the oldest one waits longer than the timeout (5 s by default), the system raises an ANR and dumps stack traces.

Batching and resampling

Motion events are batched and delivered in step with vsync via Choreographer, with resampling so drawing stays smooth.

Wearable inputs

Rotary crown or bezel (MotionEvent with AXIS_SCROLL), wrist-raise gestures and hardware buttons flow through the same pipeline.

# raw kernel events
adb shell getevent -lt
# dispatcher state, focused window, pending events
adb shell dumpsys input
Interview angle Expect "walk a touch from the panel to onClick" and a follow-up on ANRs. A strong answer names EventHub, InputReader and InputDispatcher, explains that focus comes from the window manager, and states clearly that the app side runs on the main thread, which is why blocking it delays input and causes ANRs.

A rendered frame: draw to pixels

Every frame the user sees goes through a producer-consumer pipeline paced by vsync, the display's refresh signal.

App UI thread: View.invalidate() / requestLayout()
  ──▶ Choreographer waits for vsync-app, then runs input ▸ animation ▸ traversal
        (measure, layout, draw records a display list of RenderNodes)
  ──▶ sync to RenderThread ──▶ RenderThread issues GPU commands (OpenGL ES / Vulkan via HWUI/Skia)
  ──▶ GPU renders into a buffer dequeued from the app's BufferQueue
  ──▶ queueBuffer(): app is PRODUCER, SurfaceFlinger is CONSUMER
  ──▶ SurfaceFlinger wakes on vsync-sf, latches the newest buffer of every layer
  ──▶ HWC (Hardware Composer HAL) decides: device composition (display overlay planes)
        or client composition (SurfaceFlinger draws layers with the GPU first)
  ──▶ Display HAL / DRM-KMS ──▶ display processor (Qualcomm DPU) ──▶ DSI ──▶ panel at vsync
Analogy

Think of a newspaper printing press with a fixed deadline every night. Reporters (apps) write articles on reusable sheets; the editor (SurfaceFlinger) collects the latest sheet from each reporter and lays out the page; the press (display) prints at exactly the same time every night (vsync). If a reporter misses the deadline, yesterday's article runs again, which is a dropped frame. The reusable sheets are BufferQueue buffers, the layout step is composition via HWC, and the deadline is about 16.6 ms at 60 Hz, 11.1 ms at 90 Hz and 8.3 ms at 120 Hz.

ComponentJob
ChoreographerPer-thread scheduler that runs input, animation and draw callbacks aligned to vsync.
RenderThreadPer-app thread that turns the recorded display list into GPU commands, so the UI thread can start the next frame.
BufferQueue / BLASTPool of graphic buffers (usually 2 or 3) shared between a producer and a consumer; dequeue, fill, queue, acquire, release. Buffers are passed by handle, not copied. Since Android 10 most app surfaces use BLASTBufferQueue: same producer-consumer buffers, a cheaper SurfaceFlinger transaction path than the older BufferQueue-only setup.
FencesSync objects that say "this buffer is ready to read" or "ready to reuse," so CPU, GPU and display can work in parallel.
SurfaceFlingerSystem compositor. Owns layers, latches buffers, drives composition with HWC.
HWCVendor HAL that maps layers onto display hardware planes; cheaper in power than GPU composition.
FrameTimelineSurfaceFlinger tracking (Android 12+) of expected vs actual present time for every frame; shows up in Perfetto as jank types.
Common pitfall Blaming the GPU for all jank. Most jank comes from the UI thread: heavy work in onDraw() or onBindViewHolder(), layout thrash, disk or network I/O, or a synchronous Binder call during a frame. Check the UI thread and RenderThread slices against vsync before touching GPU settings.
Tip On a watch, ambient mode and always-on display use a low refresh rate and simplified rendering to save power, and composition is tuned aggressively because the display is one of the biggest power consumers.
Interview angle Interviewers probe whether you understand the producer-consumer split and vsync budget, and whether you can go from "janky scroll" to a concrete trace analysis. Mention dumpsys gfxinfo, Perfetto with the UI thread and RenderThread, FrameTimeline jank types, and the difference between app-side jank and SurfaceFlinger-side jank.

A Binder transaction: crossing a process boundary

Almost every arrow between an app and the framework is a Binder call. Here is the data crossing the boundary once.

Client thread: proxy.method(args)            (AIDL-generated Proxy)
  ──▶ args marshalled into a Parcel
  ──▶ ioctl(BINDER_WRITE_READ) on /dev/binder with BC_TRANSACTION
  ──▶ binder driver: checks the target handle, records sender PID/UID,
        copies the Parcel ONCE into the target process's mmap'd receive buffer
  ──▶ wakes a thread in the target's Binder thread pool (BR_TRANSACTION)
  ──▶ Stub.onTransact(code, data, reply, flags) unmarshals and runs the real method
  ◀── reply Parcel travels back the same way (BC_REPLY / BR_REPLY)
  ◀── client thread unblocks (two-way call); oneway calls never block the caller
Analogy

Binder is like a pneumatic tube system in a bank. The teller writes a slip (Parcel), drops it into a tube (ioctl), and the system delivers it directly into the back office's inbox (the receiver's mapped buffer) with a stamp showing who sent it (UID and PID). A clerk from the back office pool (a Binder thread) processes it and sends a reply slip back. The single tube delivery is Binder's single copy, the stamp is what the receiver uses for permission checks, and running out of clerks is thread-pool exhaustion.

Single copy

The driver copies data once from the sender into a buffer the receiver has mapped, instead of copying twice through the kernel like pipes or sockets.

Identity

The driver attaches the caller's UID and PID, so services can call Binder.getCallingUid() and enforce permissions safely.

Limits

Each process has about 1 MB of transaction buffer shared by all in-flight calls; large payloads cause TransactionTooLargeException. Use shared memory or file descriptors for big data.

Thread pool

Each process has a bounded Binder thread pool (15 extra threads by default; system_server raises its limit to 31). If they are all stuck in slow calls, every caller stalls.

Common pitfall Making a synchronous Binder call from the main thread to a service that might be slow (or that calls back into you). This causes jank and ANRs and can deadlock. Use oneway calls or move the call off the main thread.
Interview angle Expect "explain a Binder call at the data level." Say Parcel, ioctl, single copy into mmap'd memory, thread pool, onTransact, reply, and oneway. Follow-ups: death notifications (linkToDeath), the 1 MB limit, and how you debug a stuck call. Full detail in Binder and AIDL.

App cold start: tap to first frame

A cold start ties together Binder, the process model and the frame pipeline. It is "cold" because the app has no process yet.

Launcher tap ──▶ startActivity() ──Binder──▶ ActivityTaskManagerService / AMS
  ──▶ resolve intent, check permissions, create task and ActivityRecord
  ──▶ starting window (splash screen) shown immediately by the system
  ──▶ no process for the app? AMS asks Zygote (over a socket) to fork
  ──▶ Zygote fork(): child inherits warm ART + preloaded classes/resources (copy-on-write)
  ──▶ ActivityThread.main() ──▶ main Looper ──▶ attachApplication() Binder call to AMS
  ──▶ bindApplication: ContentProviders installed, Application.onCreate()
  ──▶ Activity onCreate ▸ onStart ▸ onResume
  ──▶ ViewRootImpl first traversal ──▶ RenderThread ──▶ BufferQueue ──▶ SurfaceFlinger
  ──▶ first frame on screen = TTID (time to initial display)
  ──▶ app calls reportFullyDrawn() when content is ready = TTFD (time to full display)
Analogy

Zygote is like a restaurant kitchen that keeps a pot of stock always simmering. When an order comes in, the chef ladles out a portion and adds the specific ingredients instead of boiling water from scratch. In Android, the simmering stock is the pre-initialised ART runtime with preloaded classes, the ladle is fork() with copy-on-write memory, and the specific ingredients are the app's own code loaded in bindApplication.

Start typeWhat already existsRelative cost
ColdNothing; process must be forked and the app initialisedHighest
WarmProcess exists, but the activity must be recreatedMedium
HotProcess and activity in memory; activity just comes to the frontLowest
# measure launch time (reports TotalTime / WaitTime)
adb shell am start -W -n com.example/.MainActivity
# look for "Displayed" and "Fully drawn" lines
adb logcat -b events,main | grep -E "Displayed|Fully drawn"
Tip Common cold-start wins: lazy-initialise SDKs instead of doing it in Application.onCreate(), avoid disk I/O on the main thread, reduce ContentProviders that run at startup, and ship Baseline Profiles so hot code is compiled ahead of time.
Interview angle Interviewers want the order of events and ownership: what AMS does, why Zygote exists, what runs on the app's main thread, and where the first frame comes from. Follow-ups: how to measure (TTID, TTFD, Perfetto app-startup track) and what you would optimise. The boot-time version of Zygote and system_server is in Android boot.

Watch and phone sync: the Data Layer

A Wear OS watch and its paired phone share state through the Wearable Data Layer API, which is provided by Google Play services on both devices.

Watch app: DataClient.putDataItem()   or   MessageClient.sendMessage()
  ──▶ Wearable Data Layer (Google Play services on the watch)
  ──▶ transport chosen automatically:  Bluetooth (BLE / classic)  ▸  Wi-Fi  ▸  cloud / LTE
  ──▶ phone node receives ──▶ Data Layer on the phone
  ──▶ companion app's WearableListenerService / OnDataChangedListener fires
  ──▶ DataItems are synchronised to all nodes and versioned (last write wins per path)
Large blobs ──▶ Asset (attached to a DataItem) or ChannelClient (streams / files)
Feature discovery ──▶ CapabilityClient ("which node can do X?")
Analogy

DataItems are like a shared family whiteboard that is copied to every room: anyone can write on it, everyone eventually sees the latest version, even if they were out when it changed. Messages are like shouting to someone in another room: fast, but lost if nobody is listening. On the watch, the whiteboard is the synchronised DataItem store keyed by URI path, the shout is MessageClient, and the "rooms" are the connected nodes.

APIUse it forDelivery
DataClient (DataItem)Small shared state (settings, last workout summary); about 100 KB per itemPersistent, synced to all nodes, survives disconnection
AssetImages or binary blobs attached to a DataItemTransferred and cached; de-duplicated
MessageClientOne-shot RPC-style commands ("start music")Fire-and-forget; fails if the node is not connected
ChannelClientStreams and large files (audio, logs)Reliable stream while connected
CapabilityClientFinding which node supports a featureCapability advertisements per node
Common pitfall Failure modes include a transport switch mid-sync, stale capability information, duplicated notifications between phone and watch, and Doze or battery saver deferring delivery. Design sync to be idempotent and transport-agnostic, and never assume a message was received without an acknowledgement.
Interview angle Interviewers ask when to use DataItems vs messages vs channels, and what happens when the phone is out of range. A strong answer covers persistence, eventual consistency, size limits, transport fallback, and how to debug (Data Layer logs, Bluetooth HCI snoop log, connectivity dumpsys). More in Wear OS.

An OTA update: server to slot B

Modern Android devices use A/B (seamless) updates: the new build is written to the inactive slot while the user keeps using the device, and the switch happens on reboot.

OTA server ──▶ client (GmsCore / OEM updater) downloads or streams the payload (full or delta)
  ──▶ update_engine verifies the payload signature and metadata
  ──▶ writes the INACTIVE slot: boot_b, vendor_boot_b, system_b, vendor_b, product_b ...
        (virtual A/B: dynamic partitions are written as copy-on-write snapshots)
  ──▶ post-install step (for example dexopt), verify hashes
  ──▶ boot control HAL marks slot B active ("unbootable" cleared, retry count set)
  ──▶ reboot ──▶ bootloader picks slot B ──▶ AVB verifies vbmeta ──▶ dm-verity on partitions
  ──▶ boot completes ──▶ update_verifier + markBootSuccessful()
  ──▶ virtual A/B: snapshot merge runs in the background (snapuserd)
  ──▶ if slot B fails to boot N times ──▶ bootloader falls back to slot A (rollback)
Analogy

An A/B update is like renovating a second bedroom while you keep sleeping in the first. Only when the new room is finished and inspected do you move in; if the new room has a problem on the first night, you simply go back to the old room. The two rooms are slot A and slot B, the inspection is signature checking plus verified boot, and moving back is the bootloader's automatic rollback after repeated boot failures.

Legacy (non-A/B)

  • Reboot into recovery to apply the update
  • Device unusable during install
  • A failed install can leave the device unbootable
  • Uses less storage

A/B and virtual A/B

  • Install in the background while the device is in use
  • Only a normal reboot is needed
  • Automatic rollback if the new slot does not boot
  • Virtual A/B keeps one copy of dynamic partitions and stores only a COW snapshot, saving space
# slot state
adb shell getprop ro.boot.slot_suffix
adb shell bootctl get-current-slot
# update_engine progress and errors
adb logcat -s update_engine
# snapshot state on virtual A/B devices
adb shell snapshotctl dump
Tip On a watch, OTA is gated by battery level, charging state and connectivity (for example "install overnight while charging on Wi-Fi"), but the A/B mechanism is the same. Payloads may be relayed from the phone when the watch has no direct internet.
Interview angle Expect "how is an OTA applied without bricking the device?" Cover the inactive slot, signature verification, AVB and dm-verity, the boot control HAL, marking boot successful, and rollback. Senior follow-ups: virtual A/B snapshot merge, what happens if power is lost during merge, rollback protection (anti-rollback index), and how OTA success rate becomes a release KPI. Boot chain detail in Android boot.

An audio path: app to speaker (and back)

"Trace audio from the app to the speaker" is asked as often as frames or Binder. Playback and capture share audioserver, and both have an offload story: the AP should not mix or decode every sample if a DSP can do it.

Playback:
App AudioTrack / MediaPlayer / ExoPlayer / AAudio / Oboe
  ──▶ AudioFlinger (audioserver) mixes or hands off a track
  ──▶ AudioPolicyService picks the device (speaker, earpiece, A2DP, USB, telephony)
  ──▶ Audio HAL (AIDL IAudio; HIDL on older devices)
  ──▶ kernel ALSA or vendor DSP / codec  ──▶ speaker, headset, or Bluetooth

Capture (the reverse): mic / BT SCO ──▶ codec / DSP ──▶ Audio HAL ──▶ AudioFlinger ──▶ AudioRecord

Offload paths (AP can sleep):
  compressed offload  ──▶ DSP decodes MP3/AAC during music
  AAudio MMAP         ──▶ shared-memory, low-latency, less wakeups
  hotword / ACD       ──▶ always-on DSP; AP wakes only on a trigger
Analogy

AudioFlinger is a mixing desk. Each app plugs in a channel (a track); the desk (policy) decides whether the output is the house speakers, a headset or a Bluetooth radio, and whether a junior engineer (the AP) must sit there mixing every sample or a hardware rack (the DSP) can play a finished tape unattended. In the real system a deep-buffer or compressed-offload track lets the AP suspend between callbacks; a fast or MMAP track keeps latency low for games and VoIP at a higher power cost.

PieceRole
AudioFlingerMixer and track server in audioserver. Owns playback and record threads, timing and the hand-off to the HAL.
AudioPolicyServiceRouting and volume: which device, which strategy (media, voice, alarm, accessibility), ducking and focus.
Audio HALVendor implementation of output and input streams, devices and, on newer HALs, AIDL IAudio.
Audio focusApps request focus so two media players do not fight; telephony and alarms can preempt.
Offload / MMAPCompressed offload and AAudio MMAP move work off the AP. Fast mixer tracks are the opposite: low latency, more CPU.
adb shell dumpsys media.audio_flinger     # tracks, threads, HAL streams
adb shell dumpsys media.audio_policy      # devices, volumes, focus
adb shell dumpsys audio                   # combined view on recent releases
Tip A music-playback battery bug is often "offload never engaged, so a FastMixer thread and the AP stay awake." Confirm with dumpsys media.audio_flinger (thread type, standby) and power rails, not just the media app's process CPU.
Interview angle Walk app → AudioFlinger → policy → HAL → DSP/codec, then name the offload fork. Follow-ups: audio focus, why VoIP uses a fast or MMAP path, and how a watch plays media through the phone over Bluetooth.

A camera frame: shutter to JPEG (and preview)

Camera is another producer-consumer pipeline, but the producer is the ISP, not the app. Interviewers want HAL3's request/result model and the split between preview (to SurfaceFlinger) and stills (to an ImageReader or encoder).

App CameraX / Camera2  ──Binder──▶ CameraService (cameraserver)
  ──▶ Camera HAL3 (AIDL ICameraDevice; HIDL on older devices)
        repeating request  ──▶ sensor + ISP produce frames
  ──▶ preview Surface   ──▶ BufferQueue / BLAST ──▶ SurfaceFlinger ──▶ display
  ──▶ still ImageReader ──▶ JPEG / YUV buffer ──▶ app saves or uploads
  ──▶ video Surface     ──▶ MediaCodec / encoder HAL ──▶ muxer ──▶ file
Analogy

HAL3 is a photo lab that takes an order form (a capture request: exposure, AF, stream list) and returns a finished packet (a result plus buffers) for every frame. You do not drive the shutter timing yourself; you keep a repeating request on the preview stream and submit a still request when the user taps. The lab (ISP) writes onto shared sheets (buffers) that the shop window (SurfaceFlinger) or the archive (ImageReader) already owns.

ConceptWhat it means
HAL3 request / resultEach frame is one request with settings and output streams, and one result with metadata plus filled buffers. HAL1 (deprecated) was a simpler "start preview / take picture" API.
StreamsConfigured outputs: preview, still, video, analysis. The ISP may produce several at once.
CameraX versus Camera2CameraX is Jetpack on top of Camera2; CameraService and the HAL stay the same.
PowerThe ISP, sensor and AF are large rails. Close the session when the app is not visible; do not leave a repeating request running in the background.
adb shell dumpsys media.camera            # cameras, sessions, HAL
adb shell dumpsys media.camera.provider   # provider / HAL process
Interview angle "Trace taking a photo" should name Camera2/CameraX, CameraService, HAL3 requests, the ISP, and the two buffer destinations (preview versus still). Compose it with the data path if the next question is "and upload it."

A location fix: GNSS and the fused provider

Location is the other always-on-or-else-it-drains flow, especially on a watch workout. The app should almost never talk to the GNSS chip itself.

App FusedLocationProviderClient / LocationManager
  ──▶ LocationManagerService (system_server)
  ──▶ Fused provider: GNSS + network (Wi-Fi / cell) + sensors
  ──▶ GNSS HAL (AIDL) ──▶ GNSS chip (duty-cycled, batched)
  ──▶ Network location (NLP) ──▶ Wi-Fi / cell scans when GNSS is weak
  ──▶ batch or interval expires ──▶ wake AP ──▶ callback to app
Analogy

The fused provider is a receptionist who looks out the window (GNSS), checks the building directory (Wi-Fi and cell), and glances at the step counter before paging you. You asked for "a location every 30 seconds, 20 metres is fine"; you did not ask to keep a satellite radio powered continuously. In the real system interval, displacement and batching decide how often the GNSS rail and the AP wake.

  • Prefer the fused provider with a coarse interval and a displacement threshold; request PRIORITY_HIGH_ACCURACY only while a workout or navigation is visible.
  • GNSS chips can batch fixes (same idea as sensor FIFOs). A watch ExerciseClient uses this; a raw requestLocationUpdates at 1 Hz does not.
  • Permissions: fine versus coarse, and background location if the app is not in the foreground. On Wear, pair this with Health Services rather than rolling your own GPS loop.
adb shell dumpsys location                # providers, clients, last fix
adb shell dumpsys gnss                    # GNSS HAL, batching, measurement
Common pitfall Leaving a high-accuracy GNSS request running after the workout activity is destroyed. The chip and the AP stay in a high-power state. Always remove updates in the matching lifecycle callback and confirm with dumpsys location.
Interview angle Walk fused provider → LMS → GNSS HAL / NLP, then the power knobs (interval, displacement, batching, priority). On a watch, say why Health Services ExerciseClient beats a raw GPS loop. See Wear OS and Power and thermal.

Where to tap each flow when it breaks

For each flow, know at least one tool per layer. The meta-skill is to name the flow, walk the layers, read the tap at each layer, then bisect to the failing layer.

FlowPrimary tools and taps
Voice calllogcat (Telecom, Telephony, ImsService, RIL), dumpsys telecom, dumpsys telephony.registry, vendor modem logs, SIP and RTP capture in Wireshark
Data packetip addr / ip route / ip rule on the cellular interface (often rmnet_data), dumpsys connectivity, dumpsys netd, tcpdump, modem logs, vendor offload stats (IPA on Qualcomm)
Audiodumpsys media.audio_flinger, dumpsys media.audio_policy, Perfetto audio tracks, HAL / DSP logs
Cameradumpsys media.camera, Perfetto camera / GPU, HAL logs
Locationdumpsys location, dumpsys gnss, Perfetto, GNSS chip logs
Sensordumpsys sensorservice (active connections, rates, batching), sensor hub logs, Health Services logs, Perfetto
Inputgetevent, dumpsys input, Perfetto input tracks, ANR traces in /data/anr
Frame and jankdumpsys gfxinfo (framestats), dumpsys SurfaceFlinger, Perfetto or systrace with FrameTimeline, GPU counters
BinderPerfetto binder transactions and thread states, /sys/kernel/debug/binder (debug builds), dumpsys activity stack dumps, ANR traces
App startam start -W (TTID), reportFullyDrawn (TTFD), Perfetto app-startup track, logcat "Displayed"
Watch syncData Layer and companion logs, Bluetooth HCI snoop log, dumpsys bluetooth_manager, dumpsys connectivity
OTAupdate_engine logs, bootctl slot state, snapshotctl, dmesg for verified boot and dm-verity, last kernel log
Power (cross-cutting)batterystats with Battery Historian, /sys/kernel/debug/wakeup_sources, Perfetto power rails, suspend residency. See Power and thermal
Analogy

Debugging a flow is like finding a leak in a long pipeline: you do not dig up the whole thing, you open inspection valves at a few points and check which side still has water. The valves are the per-layer taps (logcat, dumpsys, Perfetto, modem logs, tcpdump), and the "last valve with water" tells you which layer to hand to its owning team.

  1. Name the flow Which journey is failing: voice, data, audio, camera, location, sensor, input, frame, Binder, start, sync or OTA?
  2. Reproduce and quantify Get a reliable repro and a number (drop rate, latency, drain per hour).
  3. Tap each layer Capture the tool for every layer in one timeline, ideally a single Perfetto trace plus logs.
  4. Bisect Find the last layer where data looks correct and the first where it does not; also bisect builds if it is a regression.
  5. Hand off with evidence Give the owning team the trace, the timestamp and the exact hop that fails.
Interview angle Interviewers often end with "and how would you debug it?" A structured, layer-by-layer answer with named tools scores much higher than "I would check the logs." This is the same systemic triage skill described in Platform integration and release.

Quick revision

  • The layer cake: app, framework (system_server), native daemons and HALs, kernel, hardware and firmware.
  • App to framework uses Binder; framework to vendor uses stable AIDL or HIDL HALs checked by VINTF; user space to kernel uses syscalls, ioctl and interrupts.
  • Voice call: Dialer, Telecom, TelephonyConnectionService, ImsPhone and ImsService, RIL, Radio HAL, vendor RIL process, modem, IMS core.
  • On modern Qualcomm SoCs, QMI rides QRTR rather than only shared-memory SMD.
  • VoLTE has two planes: SIP signalling through the IMS core, and RTP media on a dedicated QoS bearer (QCI 1 or 5QI 1).
  • If IMS is unavailable, calls fall back to circuit-switched (CSFB) or, on 5G, EPS fallback to LTE.
  • Mobile data has a setup phase (setupDataCall through RIL, PDN or PDU session, IP assigned) and a transfer phase.
  • Uplink packet: app socket, bionic, kernel TCP/IP, eBPF/netfilter, fwmark routing, optional CLAT, rmnet/QMAP, IPA, modem, RAN, P-GW/UPF, Internet.
  • Downlink is not a mirror of interrupts: IPA/NAPI batch packets so the AP sees aggregated IRQs, then eBPF, TCP, socket, app recv.
  • netd programs routes, eBPF/firewall, DNS and fwmark; ConnectivityService picks and validates networks.
  • Vendor data-path offload (IPA on Qualcomm) does routing, filtering, NAT and aggregation in hardware so the AP can sleep during transfers.
  • Sensor data is sampled and batched on the always-on sensor hub; the AP wakes only when the FIFO fills or report latency expires.
  • Sensors HAL 2.x and AIDL deliver events through a Fast Message Queue, not a Binder call per event.
  • High-rate, unbatched sensor use is a classic wearable battery bug.
  • Input: IRQ, kernel driver, evdev, EventHub, InputReader, InputDispatcher, InputChannel socket, app main thread, View.
  • An unacknowledged input event for about 5 seconds triggers an input dispatch ANR.
  • Audio: app track, AudioFlinger in audioserver, AudioPolicyService, Audio HAL, DSP or codec; prefer compressed offload or AAudio MMAP when the AP should sleep.
  • Camera: CameraX/Camera2, CameraService, HAL3 request/result, ISP; preview Surface to SurfaceFlinger, still to ImageReader.
  • Location: fused provider, LocationManagerService, GNSS HAL (batched) and network location; interval and displacement decide power.
  • Frame: invalidate, Choreographer on vsync, RenderThread and GPU, BufferQueue / BLAST, SurfaceFlinger, HWC, display.
  • BufferQueue: app is the producer, SurfaceFlinger the consumer; buffers are passed by handle with fences. BLASTBufferQueue is the usual Android 10+ path.
  • Frame budget is about 16.6 ms at 60 Hz, 11.1 ms at 90 Hz, 8.3 ms at 120 Hz; missing it drops a frame.
  • HWC chooses device composition (overlay planes) or client composition (GPU), which affects power.
  • Binder: Parcel, ioctl, single copy into the receiver's mmap'd buffer, thread pool, onTransact, reply.
  • Binder transaction buffer is about 1 MB per process; oneway calls do not block the caller.
  • Cold start: AMS, Zygote fork (copy-on-write), ActivityThread, bindApplication, Activity lifecycle, first frame.
  • TTID is time to initial display; TTFD is time to full display, marked by reportFullyDrawn().
  • Data Layer: DataItems are persistent and synced; messages are fire-and-forget; channels stream large data.
  • Data Layer transport falls back automatically between Bluetooth, Wi-Fi and cloud or LTE.
  • A/B OTA: update_engine writes the inactive slot, boot control HAL switches slots, rollback if the new slot fails.
  • Virtual A/B stores only a copy-on-write snapshot for dynamic partitions and merges it after a successful boot.
  • AVB verifies vbmeta at boot; dm-verity verifies partition blocks at read time.
  • For any broken flow: name it, reproduce, tap each layer, bisect, hand off with evidence.

Glossary

A/B (seamless) update
Two copies (slots) of key partitions; the update is written to the inactive slot and activated on reboot, with rollback on failure.
AMS / ATMS
ActivityManagerService and ActivityTaskManagerService; system services that manage processes, activities and tasks.
ANR
Application Not Responding; raised when an app's main thread fails to handle input or a broadcast or service call in time.
AudioFlinger
The mixer and track server in audioserver that sits between app tracks and the Audio HAL.
AVB
Android Verified Boot; checks signed vbmeta so that only trusted images boot.
BLASTBufferQueue
Android 10+ buffer path for most app surfaces: still producer-consumer buffers, with a cheaper SurfaceFlinger transaction than classic BufferQueue-only setup.
Camera HAL3
Request/result camera HAL: each frame is a request with streams and settings, plus a result with metadata and buffers.
Binder
Android's kernel-assisted IPC mechanism with single-copy transfer, caller identity and reference-counted objects.
BufferQueue
A shared pool of graphic buffers connecting a producer (app) to a consumer (SurfaceFlinger).
Choreographer
Framework class that schedules input, animation and drawing work on vsync.
Cold start
Launching an app when no process exists, so it must be forked from Zygote and fully initialised.
CLAT / 464XLAT
Customer-side translator that lets IPv4-only apps work on IPv6-only cellular: the phone NATs IPv4 sockets onto IPv6 toward the network.
CSFB
Circuit-Switched Fallback; moving a call from LTE to 3G or 2G when VoLTE is not available.
Data Layer
Wear OS API (DataClient, MessageClient, ChannelClient, CapabilityClient) for syncing data between watch and phone.
Dedicated bearer
A cellular connection with guaranteed QoS, used for VoLTE voice media.
dm-verity
Kernel feature that checks each block of a read-only partition against a hash tree when it is read.
eBPF
Extended Berkeley Packet Filter; programs that the kernel runs on packets, sockets and cgroups. Android uses it for UID traffic stats, firewall, tethering and CLAT instead of much of the old iptables/qtaguid path.
EventHub
Component of the input system that reads raw events from /dev/input devices.
Fused location provider
Framework provider that combines GNSS, network location and sensors so apps do not talk to the chip themselves.
Fast Message Queue (FMQ)
Shared-memory queue used by HALs (such as Sensors) to pass high-rate data without Binder calls.
Fence
A synchronisation object that signals when a GPU or display operation on a buffer has finished.
fwmark
Firewall mark on a socket; Android uses it with ip rule so traffic for a UID or a specific Network takes the matching routing table.
HAL
Hardware Abstraction Layer; a versioned interface between the Android framework and vendor code.
Health Services
Wear OS API that provides batched, power-efficient access to health and fitness sensor data.
HWC
Hardware Composer HAL; decides how layers are combined by the display hardware.
IMS
IP Multimedia Subsystem; the network core that handles SIP-based voice (VoLTE, VoNR), video and SMS over IP.
InputDispatcher
Part of the input system that delivers events to the correct window and detects input ANRs.
IPA
IP Accelerator; Qualcomm hardware block that offloads the data path (routing, filtering, NAT, aggregation) from the CPU. Other vendors have equivalent offload engines.
netd
Android's native network daemon. It applies interface, routing, DNS, firewall and eBPF configuration on behalf of ConnectivityService.
Parcel
The container used to marshal data for a Binder transaction.
PDN connection / PDU session
A data connection to a packet network with its own IP address, on LTE and 5G respectively.
PPG
Photoplethysmography; optical measurement of blood volume changes used for heart rate.
QMAP
Qualcomm Multiplexing and Aggregation Protocol; header used by rmnet to carry several data calls over one link.
QMI
Qualcomm MSM Interface; the message protocol between AP software and the modem and other subsystems. On modern SoCs it rides QRTR.
QRTR
Qualcomm IPC Router; socket-style transport used for QMI on recent Qualcomm platforms, replacing older shared-memory SMD channels.
RenderThread
Per-app thread that turns display lists into GPU commands.
RIL
Radio Interface Layer; the path from Telephony through the Radio HAL to the vendor modem software.
rmnet
Qualcomm network driver exposing modem data calls as rmnet_data interfaces.
RTP
Real-time Transport Protocol; carries voice and video media packets.
Sensor hub
A low-power processor that samples and batches sensors while the main CPU sleeps.
SIP
Session Initiation Protocol; the signalling protocol that sets up and tears down IMS calls.
SurfaceFlinger
Android's system compositor, which combines layers and sends frames to the display.
TTID / TTFD
Time to initial display and time to full display; standard app start-up metrics.
update_engine
The daemon that downloads, verifies and applies A/B OTA payloads.
Virtual A/B
A/B scheme where dynamic partitions use copy-on-write snapshots instead of full second copies.
VINTF
Vendor Interface object; manifests and compatibility matrices that check framework and vendor HAL versions match.
Vsync
The display's refresh signal, used to pace rendering and composition.
Zygote
A pre-initialised process that forks to create every app process quickly.

Interview questions

Fundamentals

What are the main layers of the Android stack that data flows through?

From top to bottom: the app (Java/Kotlin on ART), the framework (system services in system_server such as AMS, WMS, Telephony and SensorService), native daemons and vendor HALs (SurfaceFlinger, rild, the Sensors HAL), the Linux kernel (drivers, networking, binder driver) and the hardware and firmware (SoC, modem, sensor hub, display). App and framework talk over Binder, framework and vendor over stable HAL interfaces, and user space and kernel over syscalls, ioctl and interrupts.

What is system_server?

system_server is the core system process, forked from Zygote during boot. It hosts most framework services as threads: ActivityManagerService, WindowManagerService, PackageManagerService, PowerManagerService, ConnectivityService, SensorService, InputManagerService and many more. Apps reach these services through Binder. If system_server crashes, the whole framework restarts (a "soft reboot").

What is a HAL and why does Android need one?

A Hardware Abstraction Layer is a versioned interface between the Android framework and vendor-specific code for a piece of hardware (radio, sensors, camera, display, audio). It lets Google update the framework without rewriting vendor drivers, and lets vendors implement hardware support without changing the framework. Since Project Treble, HALs are defined in HIDL or stable AIDL and run in separate vendor processes; VINTF manifests check that versions match.

What is the difference between signalling and media in a VoLTE call?

Signalling is the control conversation that sets up, modifies and ends the call; in VoLTE this is SIP (INVITE, 180 Ringing, 200 OK, ACK, BYE) through the IMS core, with SDP negotiating codecs and ports. Media is the actual voice, carried as RTP packets with RTCP reports. Media runs on a dedicated bearer with guaranteed QoS (QCI 1 on LTE), while signalling uses the IMS default bearer.

What is the difference between Telecom and Telephony in Android?

Telecom is the call-management layer: it tracks calls from any source (SIM calls, VoIP apps via ConnectionService), routes audio, and talks to the in-call UI. Telephony is the cellular stack: SIM, service state, IMS registration, data connections and the RIL to the modem. A cellular call goes from Telecom through TelephonyConnectionService into Telephony.

What is the RIL?

The Radio Interface Layer connects Android Telephony to the modem. On the framework side, RIL.java sends requests and receives responses and unsolicited indications. It calls the Radio HAL (HIDL IRadio or, from Android 13, AIDL interfaces such as IRadioVoice and IRadioData), implemented by a vendor process (classically rild). That process talks to the modem using a vendor protocol; on Qualcomm this is QMI, carried on modern SoCs over QRTR (Qualcomm IPC Router) rather than only shared-memory SMD.

What is a PDN connection or PDU session?

It is a cellular data connection to a specific packet network, identified by an APN or DNN, with its own IP address. LTE calls it a PDN connection (anchored at the P-GW); 5G calls it a PDU session (anchored at the UPF). A device often has several at once, for example internet, IMS and emergency, each appearing as a separate network interface.

What is rmnet?

rmnet is Qualcomm's network driver for modem data. It exposes each data call as a Linux network interface (rmnet_data0, rmnet_data1 and so on) and multiplexes them over one physical link to the modem using QMAP headers, including packet aggregation. The kernel TCP/IP stack sees normal network interfaces.

What is IPA and why does it matter?

IPA (IP Accelerator) is Qualcomm's hardware block in the data path between the modem and the AP (and peripherals like Wi-Fi or USB for tethering). It does routing, filtering, NAT, header processing and aggregation in hardware. This reduces CPU load and interrupts, so the AP can stay asleep during transfers, which is a large power win. Other SoCs have the same idea under different names.

Walk audio from an app to the speaker.

The app writes to AudioTrack, MediaPlayer, ExoPlayer, AAudio or Oboe. AudioFlinger in audioserver mixes or hands off the track. AudioPolicyService picks the output device and volume strategy. The Audio HAL (AIDL IAudio, or HIDL on older devices) talks to ALSA or a vendor DSP and codec, which drive the speaker, headset or Bluetooth. Compressed offload and AAudio MMAP let the AP sleep; a FastMixer or VoIP path does not.

Walk a camera capture from Camera2 to a JPEG.

The app (CameraX or Camera2) talks over Binder to CameraService in cameraserver. CameraService configures streams and sends HAL3 requests to the Camera HAL. The sensor and ISP fill buffers: the preview Surface goes to SurfaceFlinger, the still goes to an ImageReader as JPEG or YUV, and a video Surface can go to MediaCodec. HAL1's "take picture" API is deprecated; HAL3 is request/result per frame.

How does a location fix reach an app?

The app uses the fused location provider (or LocationManager). LocationManagerService combines GNSS (via the GNSS HAL and a duty-cycled, often batched chip), network location (Wi-Fi and cell) and sensors. Fixes are delivered on the requested interval and displacement. High-accuracy 1 Hz GNSS without batching is a classic battery bug; on a watch, Health Services ExerciseClient is the intended API.

What does the sensor hub do on a wearable?

The sensor hub is a low-power processor (on Qualcomm, the SLPI or ADSP) that samples sensors such as the accelerometer, gyroscope and PPG while the main application processor sleeps. It runs filtering, sensor fusion and algorithms like step counting, and stores samples in a FIFO. It wakes the AP only when a batch is ready or an important event occurs.

What is sensor batching?

Batching lets sensor events be stored in a hardware FIFO and delivered together instead of one at a time. An app enables it by passing a non-zero maxReportLatencyUs when registering a listener. The AP then wakes once per batch rather than for every sample, which saves a lot of power for continuous sensing.

Walk a touch event from the panel to onClick.

The touch controller raises an interrupt; the kernel input driver reports events on /dev/input/eventX. In system_server, EventHub reads them, InputReader converts them to MotionEvents, and InputDispatcher finds the target window and sends the event over an InputChannel socket. The app's main thread receives it, ViewRootImpl dispatches it through the view hierarchy via dispatchTouchEvent(), and the view's onTouchEvent() eventually triggers onClick().

What is vsync?

Vsync is the periodic signal from the display that marks the start of a new refresh cycle, for example every 16.6 ms at 60 Hz. Android uses it to pace app rendering (through Choreographer) and composition (in SurfaceFlinger), so that frames are produced at a steady rate and never shown half-drawn.

What does SurfaceFlinger do?

SurfaceFlinger is the system compositor. It receives buffers from every visible surface (apps, status bar, navigation bar, wallpaper), latches the newest buffer of each layer on vsync, and combines them into the final frame, using the Hardware Composer HAL to decide whether display hardware or the GPU does the composition. It then presents the frame to the display.

What is Zygote and why does Android use it?

Zygote is a process started by init at boot that pre-loads the ART runtime and commonly used framework classes and resources. When a new app process is needed, Zygote forks itself. The child starts already warmed up, and unchanged memory pages are shared with other apps through copy-on-write. This makes app start faster and saves RAM compared with starting a new VM for every app.

What is the difference between cold, warm and hot app start?

A cold start has no existing process, so the system must fork from Zygote, bind the application and create the activity. A warm start reuses the process but recreates the activity (for example after it was destroyed). A hot start just brings an existing activity back to the foreground. Cold start is the most expensive and the one usually measured.

What is an A/B OTA update?

An A/B update keeps two copies (slots A and B) of the key partitions. The update is written to the inactive slot while the device keeps running; after reboot the bootloader starts from the updated slot. If the new slot fails to boot several times, the bootloader switches back to the old slot, so a bad update does not brick the device.

What is the Wear OS Data Layer?

It is the API provided by Google Play services for communication between a watch and its paired phone (and other nodes). It offers DataClient for synced, persistent DataItems; MessageClient for one-way messages; ChannelClient for streams and large files; and CapabilityClient for discovering which node offers a feature. It picks the transport (Bluetooth, Wi-Fi or cloud) automatically. Builds without Play services do not have it.

Name one debugging tool for each of the main flows.
  • Voice call: logcat for Telecom, Telephony and IMS, plus modem logs.
  • Data: dumpsys connectivity and tcpdump.
  • Audio: dumpsys media.audio_flinger.
  • Camera: dumpsys media.camera.
  • Location: dumpsys location.
  • Sensor: dumpsys sensorservice.
  • Input: getevent and dumpsys input.
  • Frames: dumpsys gfxinfo and Perfetto.
  • App start: am start -W.
  • OTA: update_engine logs and bootctl.
  • Power: batterystats with Battery Historian.

Going deeper

Trace a VoLTE call from the dialer to the RF.

The Dialer calls TelecomManager.placeCall(). Telecom picks the PhoneAccount and asks TelephonyConnectionService to create a Connection. Telephony chooses ImsPhone because IMS is registered, and the ImsService builds a SIP INVITE with an SDP offer. Commands go through RIL and the Radio HAL (IRadioVoice or IRadioIms, depending on where the IMS stack lives) to the vendor RIL and, over QMI (QRTR on modern Qualcomm SoCs), to the modem. The modem sends the INVITE over the IMS bearer to the P-CSCF and IMS core, which routes it to the callee; after 180 Ringing and 200 OK, the network sets up a dedicated QCI 1 bearer and RTP voice flows over it.

What happens to a VoLTE call if IMS is not registered?

Telephony cannot place the call over IMS, so it uses the circuit-switched path. On LTE-only networks this means CSFB: the modem moves to 3G or 2G to place the call and returns to LTE afterwards. On 5G standalone without VoNR, the network performs EPS fallback to LTE and uses VoLTE there. If no CS or IMS path exists, the call fails; emergency calls have special rules to try every available domain.

How does an app's packet get out over cellular, and what accelerates it?

First a data call must exist: ConnectivityService and the Telephony data stack ask the modem, through RIL setupDataCall(), to attach to a PDN or PDU session. The network assigns an IP, and an rmnet_data interface is configured by netd with routes and DNS. Then an app write() goes through the kernel TCP/IP stack to the cellular driver (rmnet plus QMAP on Qualcomm), then through a hardware offload engine (IPA on Qualcomm), which handles routing and aggregation, to the modem, the radio, the base station and the core network. That offload keeps the bulk path off the CPU, which saves power.

How does Android choose which network an app's traffic uses?

ConnectivityService ranks available networks (Wi-Fi, cellular, Ethernet, VPN) by score and validation status and picks a default. netd then uses Linux policy routing: sockets are tagged with a firewall mark (fwmark) based on the app's UID or explicit network binding, and ip rule entries send each mark to the right routing table. Apps can bind to a specific network with Network.bindSocket() or bindProcessToNetwork().

Why can a device have several rmnet interfaces at once?

Each data call (PDN or PDU session) gets its own interface. Typical examples are the internet APN, the IMS APN used for VoLTE signalling, an emergency APN, and sometimes an MMS or tethering APN. Keeping them separate lets each have its own IP, QoS and routing, and lets IMS stay up even if the user turns mobile data off.

Trace a heart-rate sample from sensor to app and explain the power angle.

The PPG sensor is sampled by firmware on the always-on sensor hub, which filters it, computes heart rate and stores results in a FIFO while the AP sleeps. When the batch is full or the report latency expires, the hub wakes the AP. The Sensors HAL pushes events through a Fast Message Queue to SensorService, which delivers them to the app's listener, or Health Services delivers them to a PassiveMonitoringClient or ExerciseClient. Because the AP wakes once per batch instead of per sample, it can stay in suspend for long periods; an app forcing unbatched high-rate reads destroys that.

What is the difference between wake-up and non-wake-up sensors?

A wake-up sensor can wake the AP from suspend when it has data (for example when its batch is full or a significant-motion event occurs), and the HAL holds a wake lock until the event is delivered. A non-wake-up sensor never wakes the AP; its events wait in the FIFO until the AP wakes for another reason, and old events may be overwritten if the FIFO overflows. Choosing the right type is a power versus data-loss trade-off.

Why does a blocked main thread cause input lag and ANRs?

Input events are delivered to the app's main thread through its Looper, the same thread that runs lifecycle callbacks and drawing. If the main thread is busy (disk I/O, heavy computation, a slow Binder call), the event waits in the queue, so the user sees lag. InputDispatcher waits for the app to acknowledge each event; if it waits longer than about 5 seconds, the system declares an ANR.

How does InputDispatcher know which window gets a touch?

The window manager and SurfaceFlinger provide InputDispatcher with a list of windows (InputWindowInfo) with their positions, z-order, flags and focus. For touches, the dispatcher hit-tests the coordinates against visible, touchable windows from top to bottom. For key events it uses the focused window. It then sends the event on that window's InputChannel.

Describe the frame pipeline and where jank comes from.

A view is invalidated; on the next vsync Choreographer runs input, animation and traversal (measure, layout, draw into a display list) on the UI thread. RenderThread converts the display list to GPU commands, and the GPU draws into a buffer from the app's BufferQueue. The buffer is queued to SurfaceFlinger, which latches it on its vsync, composes all layers with HWC and presents the frame. Jank happens when any stage misses its deadline (about 16.6 ms at 60 Hz): usually main-thread work, a slow Binder call, heavy layout, overdraw or GPU load, or occasionally SurfaceFlinger or composition delays.

What is the role of RenderThread?

RenderThread is a per-app thread created by the hardware-accelerated renderer (HWUI). The UI thread records drawing operations into display lists and hands them off; RenderThread then issues the GPU commands, uploads textures and manages the buffer. This lets the UI thread start the next frame sooner and allows some animations to run on RenderThread even if the UI thread is briefly busy.

What is a BufferQueue and why is it a producer-consumer design?

A BufferQueue is a small pool of graphic buffers shared between a producer (an app's Surface, a camera, a video decoder) and a consumer (SurfaceFlinger, a video encoder). The producer dequeues a free buffer, fills it and queues it; the consumer acquires it, uses it and releases it. Buffers are shared by handle, not copied, and fences say when each side has finished. This decouples rendering from display timing, so the app can prepare the next frame while the current one is shown. Since Android 10 most app surfaces use BLASTBufferQueue, which keeps this model but uses a cheaper SurfaceFlinger transaction path.

What is BLASTBufferQueue?

The usual Android 10+ path for app surfaces. The app still produces graphic buffers and SurfaceFlinger still consumes them; the change is how buffer state is sent to SurfaceFlinger (a BLAST transaction rather than the older BufferQueue-only setup). Interviewers want you to know the producer-consumer and fence story did not go away.

What is QRTR and how does it relate to QMI?

QMI is the message protocol between AP software and Qualcomm subsystems (modem and others). QRTR (Qualcomm IPC Router) is the modern socket-style transport those messages ride on, replacing older shared-memory SMD channels. Saying "QMI over shared memory" is the historic picture; on current SoCs, say QMI over QRTR.

What does AudioPolicyService decide that AudioFlinger does not?

Policy picks the route and the strategy: speaker versus earpiece versus Bluetooth versus USB, media versus voice versus alarm volumes, ducking and focus. AudioFlinger runs the tracks and talks to the HAL streams that policy selected. A "no sound" bug can be either: a mixing/HAL problem (Flinger) or a routing/focus problem (policy).

What is the difference between device composition and client composition?

With device composition, the Hardware Composer assigns layers to hardware overlay planes in the display processor, which blends them while scanning out, using little power. With client composition, SurfaceFlinger first uses the GPU to combine some or all layers into one buffer, then hands that to the display. HWC falls back to client composition when there are too many layers or unsupported features, which costs power and can cause jank.

Explain a Binder call at the data level.

The client calls a method on an AIDL-generated proxy, which marshals the arguments into a Parcel. The proxy calls transact(), which issues ioctl(BINDER_WRITE_READ) on /dev/binder. The driver looks up the target, records the caller's UID and PID, and copies the Parcel once into a buffer that the target process has mapped. It wakes a thread in the target's Binder thread pool, which runs Stub.onTransact(), unmarshals the data and calls the real method. The reply comes back the same way, and the client thread unblocks; oneway calls return immediately.

What is a oneway Binder call and when would you use it?

A oneway call is asynchronous: the caller returns as soon as the driver accepts the transaction, without waiting for the method to run or return a result. Calls to the same object are queued and delivered in order. Use them for notifications and callbacks where the caller must not block, such as a system service calling back into an app. The trade-off is no return value and no exceptions reported to the caller.

What happens between tapping an icon and seeing the first frame?

The launcher calls startActivity() over Binder. ActivityTaskManagerService resolves the intent, creates the activity record and shows a starting window. If the app has no process, AMS asks Zygote to fork one. The new process runs ActivityThread.main(), starts the main Looper and calls attachApplication(); AMS replies with bindApplication, which installs ContentProviders and runs Application.onCreate(). The activity goes through onCreate, onStart and onResume, ViewRootImpl performs the first traversal, RenderThread draws, the buffer goes to SurfaceFlinger, and the first frame appears. That moment is TTID.

How do you measure app start-up time?

Use adb shell am start -W, which reports TotalTime and WaitTime; logcat's "Displayed" line gives TTID. Call reportFullyDrawn() when content is ready to get TTFD. For detail, capture a Perfetto trace with the app-startup track to see bindApplication, activity start, inflation and the first frame. Use Macrobenchmark for repeatable measurements in CI.

When would you use a DataItem, a message, or a channel in the Data Layer?

Use a DataItem for small state that both devices should eventually agree on, such as settings or the latest workout summary; it is persisted and synced even if the devices are disconnected now. Use a message for a quick command that only matters if the other side is connected right now, such as "start playback." Use a channel for large or streaming data such as audio or log files. Attach Assets to DataItems for images and binary blobs.

How is an OTA applied without bricking the device?

With A/B updates, update_engine verifies the payload signature and writes the new images to the inactive slot while the device runs normally. It verifies the written data, then the boot control HAL marks the new slot active with a limited retry count. On reboot the bootloader starts the new slot, AVB verifies it, and once boot completes the system marks the slot successful. If the new slot fails to boot within its retry count, the bootloader falls back to the old slot.

What is the difference between AVB and dm-verity?

AVB (Android Verified Boot) is the boot-time chain of trust: the bootloader checks the signed vbmeta structure, which contains hashes or hash-tree descriptors for each partition, and refuses to boot or warns the user if they do not match. dm-verity is the kernel feature that enforces those hash trees at runtime, checking each block of a read-only partition such as system or vendor as it is read. AVB decides whether to boot; dm-verity catches tampering or corruption afterwards.

Advanced

Why does Binder copy data only once, and what are the limits of that design?

Each process that receives Binder transactions maps a region of kernel-managed memory into its address space. The driver copies the sender's data directly from user space into pages that are mapped in the receiver, so the receiver reads it without a second copy. The limit is that this buffer is small (about 1 MB per process, shared by all in-flight transactions, and less for oneway calls), so large payloads throw TransactionTooLargeException. For big data, pass a file descriptor, SharedMemory, or a HardwareBuffer through Binder instead.

How does Binder support security and permission checks?

The driver stamps each transaction with the caller's real UID and PID, which cannot be forged from user space. Services call Binder.getCallingUid() or checkCallingPermission() to enforce permissions. SELinux adds another layer: policy controls which domains may call, transfer to, or find which service managers and services. Services should use clearCallingIdentity() carefully when acting on their own behalf.

What is Binder thread-pool exhaustion and how do you detect it?

Every process has a fixed maximum number of Binder threads. If all of them are blocked (for example waiting on a lock, slow I/O, or a nested Binder call to another stuck process), new incoming transactions have to wait, so every caller of that process stalls. In system_server this looks like a system-wide freeze and can trigger the watchdog. Detect it with ANR or watchdog stack dumps showing many binder: threads blocked in the same place, Perfetto traces showing binder transactions with long durations, or debug entries under /sys/kernel/debug/binder.

How do HIDL and stable AIDL HALs differ in the data path?

HIDL HALs use the hwbinder driver and a separate hwservicemanager, with interfaces versioned like android.hardware.radio@1.6. Stable AIDL HALs (Android 11 onwards, and the default for new HALs) use the normal binder driver or vndbinder, are registered with servicemanager, and are versioned by integer with frozen API snapshots. Both marshal data into parcels and support FMQ for high-rate data. Google has been migrating HALs from HIDL to AIDL, so a platform upgrade often includes a HAL interface migration.

How does a Fast Message Queue work and why use it for sensors?

An FMQ is a ring buffer in shared memory set up once between two processes over a HAL call. Afterwards the writer and reader exchange data directly through the shared memory, using atomic read and write pointers and optionally an event flag to wake the reader. There is no Binder transaction per item. The Sensors HAL 2.x and AIDL use one FMQ for events and another for wake-lock acknowledgements, which cuts the overhead of delivering thousands of sensor events per second.

Explain vsync-app and vsync-sf offsets.

Android derives two software vsync signals from the hardware vsync. Vsync-app wakes Choreographer in apps to start a frame; vsync-sf wakes SurfaceFlinger to compose. They are offset so that apps have time to render before SurfaceFlinger latches their buffers, and SurfaceFlinger has time to compose before the display scans out. Tuning these offsets trades latency against the chance of missing a frame; high refresh rates make the windows tighter.

What is triple buffering and what problem does it solve?

With two buffers, if the app misses a frame deadline, both buffers may be in use (one displayed, one still being drawn), so the app must wait and may miss the next deadline too. A third buffer lets the app start the next frame while one buffer is on screen and another is waiting to be composed. This smooths occasional long frames at the cost of more memory and up to one extra frame of latency.

How do fences allow CPU, GPU and display to work in parallel?

When the app queues a buffer, the GPU may still be drawing it; an acquire fence signals when the drawing finishes, so SurfaceFlinger or HWC waits only on that fence rather than the CPU waiting for the GPU. When the display releases a buffer, a release fence signals when scan-out has finished, so the producer can safely overwrite it. Present fences report when a frame actually reached the screen. This removes the need for blocking handshakes at every stage.

How does IPA interact with rmnet and the kernel data path?

IPA sits between the modem and the AP memory. On downlink, the modem hands packets to IPA, which applies filter and routing rules, aggregates packets with QMAP headers and writes them to AP memory via DMA, raising one interrupt per aggregated batch. The rmnet driver de-aggregates and hands packets to the kernel network stack. On uplink the reverse happens. For tethering, IPA can forward between USB or Wi-Fi and the modem entirely in hardware, so packets never reach the AP's network stack.

How does the IMS stack stay reachable when the user disables mobile data?

IMS uses its own APN (usually "ims"), which is a separate PDN connection that the "mobile data" toggle does not control. Telephony keeps the IMS data call up while the device is registered on LTE or NR, so SIP registration, incoming call INVITEs and SMS over IMS continue to work. Only the internet APN is brought down when data is disabled.

How does the Sensors HAL avoid dropping wake-up events when the AP suspends?

When the HAL writes a wake-up event into the event FMQ, it holds a wake lock so the AP cannot suspend before the framework has read it. SensorService reads the event, delivers it to clients, and writes an acknowledgement to a separate wake-lock FMQ; the HAL then releases its wake lock. If this handshake is broken (for example a client that never acknowledges), the device can be kept awake, which shows up as a sensor wake lock in batterystats.

What is copy-on-write in the context of Zygote, and how can it be undermined?

After fork(), the child shares all of Zygote's memory pages with the parent, marked read-only. Only when a page is written does the kernel copy it for the writing process. This lets every app share preloaded classes and resources. It is undermined when writes touch many shared pages, for example garbage collection moving objects in the preloaded heap or apps modifying preloaded objects; ART mitigates this by placing preloaded objects in a separate, rarely collected space.

How does virtual A/B differ from classic A/B, and what is the snapshot merge?

Classic A/B keeps two full copies of every updatable partition, doubling storage. Virtual A/B keeps one copy of dynamic partitions (system, vendor, product) inside the super partition and writes the update as a copy-on-write snapshot that stores only changed blocks (compressed from Android 12). After reboot, the new slot is assembled from the base partition plus the snapshot using device-mapper and, with compression, snapuserd. Once boot is marked successful, a background merge writes the snapshot into the base partition; until then rollback is possible. The merge is designed to be resumable if power is lost.

What is anti-rollback protection in OTA?

Each image in vbmeta has a rollback index. The bootloader stores the highest index it has booted in tamper-resistant storage (for example RPMB or fuses) and refuses to boot images with a lower index. This prevents attackers from flashing an older, vulnerable but validly signed build. The stored index is only increased after the new build has booted successfully, so A/B rollback to the previous slot still works during the trial period.

How does the Data Layer handle conflicts and ordering?

DataItems are identified by a URI path including the node that created them; the latest write to a path replaces the previous value and is propagated to all connected nodes, so it behaves like last-writer-wins eventual consistency. Messages have no persistence or conflict resolution. For state edited on both sides, design the data so each device writes its own paths, include timestamps or version numbers in the payload, and make handlers idempotent so re-delivered updates do no harm.

What does a senior engineer look for in a Perfetto trace of a janky frame?
  • The FrameTimeline track, to see whether the frame was late on the app side or on SurfaceFlinger side, and the jank type.
  • The app's UI thread: long Choreographer#doFrame, inflation, layout, or binder transaction slices.
  • RenderThread: long draw, texture upload, or waiting on a GPU fence (dequeueBuffer blocked means no free buffer).
  • CPU scheduling: was the thread runnable but not running (CPU contention, wrong core, frequency too low)?
  • SurfaceFlinger: long composition or client composition fallback.
How does the app's first frame get prioritised during a cold start?

The system shows a starting window (splash screen) immediately, so the user sees feedback before the app draws. The top app is placed in the top-app cgroup with higher CPU priority and access to big cores, and the performance HAL may boost CPU and memory frequencies during launch. Features like Baseline Profiles and cloud profiles make sure start-up code is compiled ahead of time, avoiding interpretation and JIT during the first frame.

How would you design a cross-layer trace that covers an entire data flow?

Use Perfetto as the backbone: enable scheduling, binder, frame timeline, input, power rails and relevant atrace categories, plus app-level trace sections you add with Trace.beginSection(). Capture logcat in the same trace, and align modem or sensor-hub logs by timestamp with a known event (for example a SIP INVITE or a sensor batch). Record a clear marker when the user action happens. The goal is one timeline where every hop in the flow can be seen with its duration.

Scenario & debugging

A frame is janking during scroll. Which flow and which tools do you use?

This is the render flow. Start with dumpsys gfxinfo <package> framestats to confirm and quantify missed frames. Then capture a Perfetto trace while scrolling and look at the UI thread and RenderThread against vsync and FrameTimeline. Typical root causes are heavy onBindViewHolder() work, layout inflation during scroll, image decoding on the main thread, a synchronous Binder call, or GPU overdraw; check dumpsys SurfaceFlinger for client composition. Fix by moving work off the main thread, pre-computing, caching, and flattening layouts, then re-measure.

Users report that VoLTE calls fail but CS calls work. How do you debug?
  1. Check IMS registration state (dumpsys telephony.registry, IMS logs); if not registered, calls fall back to CS.
  2. If registered, capture logcat for Telecom, Telephony and ImsService plus modem logs for the failing call.
  3. Look at the SIP ladder: does the INVITE go out, and what response comes back (403, 488 codec mismatch, 503, timeout)?
  4. Check whether the dedicated bearer is set up and whether preconditions succeed.
  5. Compare with a known-good build or SIM and check carrier configuration (CarrierConfig VoLTE flags, IMS APN settings).

Then hand off to the owning layer (framework, IMS stack, modem or network) with the exact failing step.

Mobile data shows as connected, but apps have no internet. Walk through your debugging.
  1. Confirm the data call: dumpsys telephony.registry and the interface in ip addr (does rmnet_dataX have an IP?).
  2. Check routes, rules and DNS: ip route show table all, ip rule, dumpsys connectivity (is the network validated? is it the default?).
  3. Test from the shell: ping an IP address, then a hostname, to separate routing from DNS problems.
  4. Check firewall and per-app restrictions (Data Saver, background restrictions, VPN).
  5. If packets leave the AP (tcpdump shows them) but nothing returns, collect modem logs and IPA stats to see whether the modem or network drops them; check MTU if small requests work but large ones hang.
After an update, watch battery life dropped sharply overnight. Suspect: sensors. How do you prove it?

Reproduce with a fixed overnight profile and compare drain per hour against the previous build. Take a bug report and load batterystats into Battery Historian: look for frequent wakeups, sensor wake locks and low suspend residency. Check dumpsys sensorservice for active connections, their sampling rate and whether batching (max report latency) is used; look for a client that registered a wake-up sensor with zero latency. Check /sys/kernel/debug/wakeup_sources for sensor-related wake sources. Once the client and change are identified, fix the batching or wake-up usage, then add standby drain as a regression gate.

An app gets "Input dispatching timed out" ANRs. How do you find the cause?

Pull the ANR trace from /data/anr or the bug report and look at the main thread stack at the time of the ANR. Common patterns: blocked on a lock held by another thread, doing disk or network I/O, waiting on a synchronous Binder call (check what the remote side was doing), or a long loop. Also check dumpsys input for the pending event queue and the CPU load section in the ANR log; a system-wide overload can make any app ANR. Fix the specific blocking work and use StrictMode to catch main-thread I/O in testing.

Touch feels laggy on a new board bring-up, but apps are not busy. Where do you look?

Measure where the latency is. Use getevent -lt to see the timestamps of raw events and whether the touch controller reports at the expected rate; a low report rate or firmware filtering points to the touch driver or firmware. Use Perfetto input tracks to see time from kernel event to InputDispatcher to app delivery. Check whether the touch IRQ is threaded and pinned to a slow or sleeping core, whether CPU frequency is low during interaction (touch boost missing in the power HAL), and whether display refresh or vsync configuration is wrong.

A cold start regressed by 300 ms in the latest platform build. How do you find the cause?

Measure with am start -W or Macrobenchmark on both builds under the same conditions (cache dropped, same thermal state). Capture Perfetto traces with the app-startup track on both and compare phases: process fork, bindApplication, activity start, inflation, first frame. If the gap is before bindApplication, look at AMS, Zygote or system load; if inside, look at class loading, dex compilation state (dumpsys package for compiler filter), or I/O. Also check CPU frequency and scheduling differences. Bisect the build change list if the phase is not obvious.

The whole UI freezes for several seconds, then the watchdog restarts system_server. What is your approach?

Collect the watchdog dump and ANR traces from the bug report. The watchdog reports which monitor or handler thread was blocked; look at its stack and at the lock it is waiting for, then find the thread holding that lock. Very often it is a Binder call out of system_server to a slow HAL or app while holding a service lock, or Binder thread-pool exhaustion. Check kernel logs for I/O stalls or memory pressure too. Fix by not holding locks across outgoing Binder calls, adding timeouts, or making the call oneway.

Phone-to-watch notifications sometimes arrive late or twice. How do you debug?

Check whether the delay correlates with transport changes (Bluetooth to Wi-Fi to cloud) or with the watch being in Doze. Collect Bluetooth HCI snoop logs and Data Layer and companion logs on both devices with synced timestamps. Duplicates often come from both the phone bridge and a standalone watch app posting the same notification, or from retries after a transport switch without de-duplication. The fix is usually to use bridging rules or dismissal IDs correctly and make the handling idempotent.

An OTA installs successfully, but after reboot the device returns to the old build. What happened?

The new slot most likely failed to boot and the bootloader rolled back after exhausting its retry count, or the boot was never marked successful. Check bootctl for slot states (unbootable, successful), bootloader and kernel logs from the failed attempts (last kernel log, pstore), and whether update_verifier found dm-verity errors. Other causes include a vbmeta or AVB failure, a rollback index problem, or a service that crashes during boot so boot never completes. Reproduce by flashing the new slot directly and capturing a UART or serial log.

A virtual A/B device lost power during the snapshot merge. Is it bricked?

No, if the implementation is correct. The merge is designed to be resumable: progress is tracked in metadata, and on the next boot the device assembles the partitions from the base plus remaining snapshot and continues merging. During the merge, rollback to the old slot is no longer possible because the base has been partly overwritten, which is why the merge only starts after the new slot is marked successful. Verify with snapshotctl dump and update_engine logs.

Downloads over cellular drain much more battery than over Wi-Fi on a new build. What do you check?

Check whether the IPA offload and aggregation are working: IPA statistics, interrupt counts on the data path (/proc/interrupts), and CPU usage in network softirq processing. If aggregation is off or misconfigured, the AP handles every packet and cannot sleep. Also compare modem power and radio state (for example the device staying in a high-power connected state because of frequent small transfers) and check for apps keeping sockets busy. Compare the IPA and rmnet configuration against the last good build.

A system service gets TransactionTooLargeException in production. How do you fix it?

Find which transaction is too large, usually from the stack trace and by logging parcel sizes. Common culprits are large Bundles in onSaveInstanceState(), big bitmaps in intents, or long lists returned from a service. Fix by sending less data (IDs instead of objects), paging results with ParceledListSlice, or passing a file descriptor or SharedMemory for bulk data. Remember the 1 MB buffer is shared by all in-flight transactions, so many concurrent medium calls can also trigger it.

A health app shows gaps in heart-rate data overnight. What could cause it?

Likely causes are FIFO overflow with non-wake-up sensors (the AP slept too long and old samples were overwritten), the app being killed or restricted in the background, Doze deferring delivery, or the sensor hub pausing the sensor (for example off-wrist detection). Check dumpsys sensorservice for the registration and FIFO settings, the app's standby bucket and process state, and sensor-hub logs. Using Health Services passive monitoring, which buffers data on the hub and handles delivery, usually fixes this.

An interviewer asks you to trace "taking a photo and uploading it over cellular." How do you structure the answer?

Break it into flows you know. Capture: the app calls CameraX or Camera2, CameraService talks to Camera HAL3 with a request/result, the ISP produces frames into buffers (preview goes to SurfaceFlinger, the still goes to an ImageReader), and the JPEG is written to storage. Upload: the app opens a socket on the default network, the packet path goes through the kernel, the cellular interface (rmnet on Qualcomm), the vendor offload engine and the modem to the internet, after the data call is set up. Name one tool per stage. This shows you can compose known flows into a new one.

Music playback drains the watch even with the screen off. How do you debug?

Decide whether decode and mix are on the AP or offloaded. dumpsys media.audio_flinger shows track types (compressed offload, deep buffer, fast) and whether the mixer thread is in standby. If offload never engaged, the AP stays awake mixing PCM. Also check Bluetooth A2DP versus a speaker, a wakelock held by the media session, and whether a visualization or MediaStyle notification is forcing a fast path. On a watch, playing through the phone over Bluetooth is often cheaper than a local speaker plus LTE.

A workout app's GPS is jagged and the battery drops fast. Which flow do you tap?

The location flow. dumpsys location and dumpsys gnss show the client, interval, displacement, batching and whether high accuracy is on. Compare against Health Services ExerciseClient, which batches GNSS on the hub or chip. Also check whether the screen stayed interactive and whether LTE was up. Fix: longer interval, batching, displacement, and stop updates when the exercise ends.

A Binder call from an app into a system service takes 200 ms at random times. How do you investigate?

Capture a Perfetto trace with binder tracing enabled and find the slow transactions. Check the server side: was a Binder thread available (thread-pool exhaustion shows as waiting to start), and what was the server thread doing (waiting on a lock, I/O, or its own outgoing Binder call)? Check CPU scheduling for both threads: runnable but not running points to CPU contention or priority problems. Fix the server-side bottleneck, or make the client call asynchronous if the result is not needed immediately.

A watch shows a black screen when waking from ambient mode for a moment. Which flows are involved?

Several: the input or wake gesture (wrist-raise or button) is detected by the sensor hub or input subsystem; PowerManagerService changes display state; the display HAL moves the panel from low-power mode to normal; SurfaceFlinger and the watch face must produce a new interactive frame. Capture Perfetto with display, power and SurfaceFlinger tracks, and measure wake-to-first-frame. Common causes are the watch face taking too long to render its interactive frame, display power-mode transitions, or CPU still being at a low frequency after resume.

How would you answer "a feature is broken, and you do not know which layer" in an interview?

Say the method out loud: identify which flow is involved, reproduce it reliably, and add a measurement. Then walk the layers from top to bottom and name the tap for each (app logs, framework dumpsys, HAL logs, kernel dmesg or ftrace, firmware or modem logs). Find the last layer where the data is correct and the first where it is wrong; if it is a regression, bisect builds too. Finally hand the evidence to the owning team and add a test so the problem cannot come back silently.