Android Platform

Android Frameworks Internals

The Android framework is the layer of Java/Kotlin system services and app-facing APIs that sits between apps and the native, HAL and kernel layers. This page explains how that layer really works: system_server and its services, Zygote and app start, lifecycles and process priority, the Handler/Looper threading core, ANRs and the watchdog, PackageManager, and the Treble boundary to vendor hardware.

~95 min read 0 interview questions
In 30 seconds
  • Almost all system services (AMS/ATMS, WMS, PMS, PowerManager, InputManager, ...) live as threads inside one privileged process, system_server. If it dies, the whole framework restarts (a "soft reboot").
  • Every app process is forked from Zygote, a warm process with ART and common classes preloaded, so apps start fast and share memory copy-on-write.
  • AMS ranks processes with an oom_adj score; LMKD, driven by kernel PSI memory pressure, kills the least important first.
  • The main thread is a Looper draining a MessageQueue. Block it and you get jank, and after a timeout (input 5 s, broadcast 10 s, service 20 s) an ANR.
  • Apps talk to services, and services talk to vendor HALs, over Binder (see Binder IPC & AIDL). Treble keeps the framework/vendor boundary stable and versioned (HIDL, now stable AIDL, checked by VINTF).

The framework map

When people say "the Android framework" they mean two things together: the system services (Java/Kotlin code that owns activities, windows, packages, power, input and so on) and the SDK API layer that apps call (ActivityManager, PackageManager, PowerManager, ...). Those SDK classes are thin client-side wrappers. The real work happens in other processes, reached over Binder IPC.

The heart of the framework is one privileged process, system_server, which hosts most system services. Around it sit per-app processes forked from Zygote, native daemons (SurfaceFlinger, servicemanager, installd, netd, logd, mediaserver) and, below them, vendor HAL processes that drive the hardware.

Apps (each in its own process, own ART runtime, own Linux UID)
   │  SDK APIs (ActivityManager, PackageManager, PowerManager ...) = thin client proxies
   │  ↕ Binder IPC  (/dev/binder)
system_server  ── AMS/ATMS · WMS · PMS · PowerMS · DisplayMS · InputMS · SensorService · ... (~100 services)
   │  ↕ Binder
Native daemons  ── SurfaceFlinger · servicemanager · installd · netd · logd · mediaserver · lmkd
   │  ↕ stable AIDL HAL (Binder) / legacy HIDL HAL (HwBinder)
Vendor HALs  ── sensors · display (HWC) · audio · power · thermal · radio · camera
   │  ioctl / sysfs / netlink
Kernel drivers  ──▶  hardware

App process

Your APK's code runs here, forked from Zygote with a warm ART runtime and preloaded classes. It reaches the system only through Binder calls to service proxies.

system_server

Forked by Zygote at boot. Hosts the manager services as threads. A crash or watchdog kill restarts the whole framework. The single most important process to understand.

servicemanager

The Binder name registry. Services call addService("activity", binder); clients call getService("activity") to get a handle. The phone book for IPC.

Native daemons

Performance-critical or privileged work in C++: SurfaceFlinger (composition), installd (file ops for packages), netd (networking), lmkd (memory killing).

Analogy

Think of a large office building. Apps are tenant companies, each in its own locked suite. system_server is the building management office where many departments (security, mail room, elevators, electricity) share one floor. The reception desk directory is servicemanager, and the internal phone system is Binder. The basement contractors who actually fix the boilers are the vendor HALs. Tenants never walk into the basement; they phone the management office, which calls the contractor through an agreed work-order form (the HAL interface).

Where the code lives in AOSP

PathWhat is there
frameworks/base/core/java/android/Public SDK classes: Activity, Handler, Looper, Binder, Parcel, managers
frameworks/base/services/core/java/com/android/server/System services: am/ (AMS), wm/ (WMS, ATMS), pm/ (PMS), power/, input/, Watchdog.java
frameworks/base/services/java/com/android/server/SystemServer.javaEntry point of system_server; starts every service
frameworks/base/core/java/com/android/internal/os/ZygoteInit.javaZygote startup, preload, socket loop
frameworks/native/SurfaceFlinger, servicemanager, libbinder, InputFlinger (native input)
system/memory/lmkd/Low Memory Killer Daemon
hardware/interfaces/AOSP HAL interface definitions (.aidl and legacy .hal)
Interview angle A common opener is "draw the Android stack and tell me what crosses each boundary." A strong answer names one concrete component per layer and the mechanism at each arrow: Binder between app and framework, Binder (stable AIDL) or HwBinder (HIDL) between framework and HAL, syscalls/ioctl between HAL and kernel. Mentioning that SurfaceFlinger is a separate native process, not part of system_server, shows real understanding.

system_server and its key services

system_server is the first Java process Zygote forks at boot. Its SystemServer.run() method creates a SystemServiceManager and starts services in three waves, then tells each service when the system reaches certain boot phases.

  1. startBootstrapServices() Services everything else depends on: Installer (connection to installd), ActivityManagerService + ActivityTaskManagerService, PowerManagerService, DisplayManagerService, PackageManagerService, UserManagerService.
  2. startCoreServices() Essential but less entangled services: BatteryService, UsageStatsService, WebViewUpdateService, BinderCallsStatsService.
  3. startOtherServices() The long tail: WindowManagerService, InputManagerService, AlarmManagerService, ConnectivityService, NotificationManagerService, JobSchedulerService, telephony registry, and many more. Newer releases also have startApexServices() for services shipped in Mainline modules.
  4. Boot phases SystemServiceManager.startBootPhase() broadcasts phases such as PHASE_SYSTEM_SERVICES_READY (500), PHASE_ACTIVITY_MANAGER_READY (550), PHASE_THIRD_PARTY_APPS_CAN_START (600) and PHASE_BOOT_COMPLETED (1000). Each service reacts in onBootPhase().
  5. systemReady AMS.systemReady() starts persistent apps (SystemUI, the phone process) and launches the Home activity. When Home draws, WMS stops the boot animation and the system user is marked booted: sys.boot_completed=1 is set and ACTION_LOCKED_BOOT_COMPLETED is sent to Direct Boot-aware apps (device-encrypted storage is available). ACTION_BOOT_COMPLETED is sent later, after the user unlocks credential-encrypted (CE) storage. On a device with no lock screen the two broadcasts can fire in quick succession; they are not the same event. See boot and FBE.
ServiceOwnsWhy a platform engineer cares
ActivityManagerService (AMS)Processes, services, broadcasts, content providers, oom_adj scoring, ANR detection for broadcasts/services/providersCentre of "why did my process die / ANR / not start" triage
ActivityTaskManagerService (ATMS)Activities, tasks, back stack, recents, multi-window. Split out of AMS in Android 10 and lives in the wm packageLaunch, task and lifecycle bugs
WindowManagerService (WMS)Windows, z-order, focus, rotation, transitions, window tokens; drives SurfaceFlinger through SurfaceControl transactionsJank, focus/input routing, rotation and display bugs
PackageManagerService (PMS)Parse and install APKs, signatures, components, intent resolution, UIDsInstall failures, permission issues, preinstalled-app configuration
PowerManagerServiceWakelocks, screen on/off, user activity, doze, suspend decisionsBattery drain and "device never sleeps" triage
DisplayManagerServiceLogical/physical displays, brightness, ambient/AOD stateDisplay power tuning
InputManagerServiceJava side of input; native InputReader and InputDispatcher threads read /dev/input and deliver events to the focused windowTouch latency, input ANRs
SensorServiceSensor HAL client, rates, batching, delivery to apps (native service hosted in system_server)Sensor delivery and power
JobScheduler / AlarmManagerDeferred background work and alarms under Doze/App Standby rulesBackground work and battery
ConnectivityService / TelephonyRegistryNetwork selection and state, telephony state callbacksConnectivity and radio integration

How SurfaceFlinger relates

SurfaceFlinger is not inside system_server. It is a native daemon started by init. WMS decides policy (which window is where, which is on top, which gets focus) and sends layer changes to SurfaceFlinger via SurfaceControl.Transaction. Apps render into buffers (via a BufferQueue / BLASTBufferQueue) and SurfaceFlinger composes all layers, using the Hardware Composer (HWC) HAL, onto the display at each vsync.

App RenderThread ──buffers──▶ BufferQueue ──▶ SurfaceFlinger ──▶ HWC HAL ──▶ panel
                                                  ▲
WMS (system_server) ──SurfaceControl transactions─┘  (position, z-order, visibility)
Analogy

system_server is a city hall with many departments under one roof. AMS is the department of residents (who is alive, who gets evicted), ATMS is the scheduler of public events (which activity is on stage), WMS is the town planner deciding where every building (window) stands, PMS is the permits office, and PowerManager is the utility company deciding when the lights can go off. SurfaceFlinger is the printing press across the street: the planner sends it layout instructions, and it prints the final picture.

Tip Every service can be inspected live with dumpsys <service>; service list shows all registered Binder services. Knowing which dumpsys to reach for is a strong signal of framework fluency (see the dumpsys section).
Common pitfall Saying "system_server has one thread per service." It has a handful of shared, named threads (android.fg, android.ui, android.io, android.display, android.anim, the main thread) plus a Binder thread pool raised to 31 threads. Services mostly run their work on these shared Handler threads.
Interview angle Expect "What is system_server, and what happens if it crashes?" The answer: it hosts most framework services; if it crashes or the watchdog kills it, Zygote (which is its parent) also restarts, all apps die, and the framework comes back up without a kernel reboot. Bonus: mention the three start waves and boot phases, and that AMS and ATMS were split in Android 10.

Zygote and how an app process starts

Zygote is a process started by init from init.zygote64.rc (or init.zygote64_32.rc on mixed 64/32 devices) via the app_process64 binary. The init service name is always zygote, even on 64-bit devices; the "64" is the rc file and the binary, not the service. It boots the ART runtime once, preloads the classes listed in /system/etc/preloaded-classes, common resources, shared libraries and graphics drivers, then waits for fork requests on a Unix domain socket (/dev/socket/zygote). On mixed-ABI devices a second service named zygote_secondary runs app_process32 and listens on /dev/socket/zygote_secondary.

  1. init starts Zygote app_process → AndroidRuntime.start() → ZygoteInit.main().
  2. Preload Classes, resources, libandroid, EGL/Vulkan drivers, WebView resources. This takes seconds, but only once per boot.
  3. Fork system_server forkSystemServer() creates the system_server child with UID 1000 (system).
  4. Select loop ZygoteServer.runSelectLoop() waits on the socket for "start a process" commands.
  5. Fork and specialize For each request, Zygote.forkAndSpecialize() calls fork(), then the child sets its UID/GIDs, SELinux context, seccomp filter, mount namespace, cgroups and nice value, and calls the requested entry class, normally android.app.ActivityThread.main().
init
 └─ zygote (app_process64)          preload ART + classes + resources
     ├─ fork ─▶ system_server        (uid 1000)
     ├─ fork ─▶ com.android.systemui (persistent)
     ├─ fork ─▶ com.android.phone    (persistent)
     └─ fork ─▶ com.example.app      (uid 10xxx)   ...one per app process

Why fork from Zygote

  • Startup skips VM boot and class loading: tens of ms instead of seconds.
  • Preloaded pages are shared copy-on-write across every app, saving a lot of RAM.
  • Every app starts from an identical known state.

Why Zygote uses a socket, not Binder

  • Binder needs a thread pool; fork() in a multithreaded process only copies the calling thread and can leave locks held forever in the child.
  • Zygote therefore stays single-threaded before forking and uses a simple local socket.
  • The Zygote socket is protected by SELinux so only system_server can request forks.

Related Zygote features

  • USAP pool (Unspecialized App Process, Android 10): Zygote pre-forks a few processes so the fork cost is paid ahead of time.
  • WebView zygote and app zygote: separate zygotes for isolated WebView renderer processes and for apps that want to preload their own code for isolated services.
  • 64/32-bit zygotes: the primary init service is named zygote and runs app_process64. A 32-bit helper, if present, is the separate service zygote_secondary (not zygote32 and not a second service also called zygote). 32-bit-only devices use a single zygote running app_process32.

From fork to a running app

AMS / ATMS: startActivity()  ─ process for this app exists? ─ no ─▶
  ProcessList.startProcessLocked()
    → Process.start() → ZygoteProcess writes args to /dev/socket/zygote
      → Zygote: forkAndSpecialize()      child pid returned to AMS
        → child: RuntimeInit → ActivityThread.main()
          → Looper.prepareMainLooper()
          → ActivityThread.attach() → IActivityManager.attachApplication(appThread)   (Binder)
            → AMS: bindApplication() back to the app (via IApplicationThread)
              → app: create Application, install ContentProviders, Application.onCreate()
            → ATMS: realStartActivityLocked() → ClientTransaction (LaunchActivityItem, ResumeActivityItem)
              → app: Activity.onCreate / onStart / onResume
          → Looper.loop()   (main thread now serves messages forever)

Two Binder interfaces make this a two-way conversation: the app calls AMS through IActivityManager, and AMS calls back into the app through IApplicationThread, a Binder object the app hands over in attachApplication(). The app's ActivityThread.H Handler turns those callbacks into main-thread messages.

Analogy

Zygote is a bakery that keeps a pre-made, half-baked dough base ready. When an order arrives, it just tears off a piece and adds the customer's toppings instead of mixing flour from scratch. The dough base is the preloaded runtime and classes, tearing off a piece is fork(), the toppings are the app's UID, SELinux label and code, and because every piece comes from the same base, identical ingredients are shared (copy-on-write pages).

Interview angle "Why does Android use Zygote?" and "Why does Zygote use a socket instead of Binder?" are classic. The second one separates people who memorized the diagram from people who understand fork() semantics with threads. Also be ready to explain IApplicationThread as the reverse channel from AMS into the app.

Cold, warm and hot app start

App startup time is a core performance KPI. Android distinguishes three start types based on how much state already exists.

TypeWhat already existsWhat must happenRelative cost
ColdNothing; no processZygote fork, bindApplication, Application.onCreate, Activity create, inflate, first frameSlowest
WarmProcess is alive, but the Activity was destroyed (back pressed, or activity reclaimed)Activity onCreate (maybe with saved state), inflate, first frameMedium
HotProcess and Activity are in memory (just stopped)onRestart/onStart/onResume, redrawFastest

Metrics and how to measure

  • TTID (time to initial display): from launch intent to the first frame. Logged by the system as ActivityTaskManager: Displayed com.example/.MainActivity: +480ms.
  • TTFD (time to full display): until the app calls reportFullyDrawn() after loading its real content.
  • adb shell am start -W -n com.example/.MainActivity prints LaunchState (COLD/WARM/HOT), TotalTime and WaitTime.
  • Perfetto has an app-startup track showing bindApplication, activityStart, inflate, Choreographer#doFrame and Binder/lock waits on the main thread.
adb shell am force-stop com.example
adb shell am start -W -n com.example/.MainActivity
# Status: ok
# LaunchState: COLD
# Activity: com.example/.MainActivity
# TotalTime: 612
# WaitTime: 630

What makes startup slow, and fixes

Heavy Application.onCreate

SDK initialisation, DI graphs, disk reads. Fix: lazy init, App Startup library, move work off the main thread.

Content providers

Every provider in the manifest is created before Application.onCreate. Many libraries auto-init through providers; merge or remove them.

Class loading / JIT

Cold code runs interpreted. Fix: Baseline Profiles / cloud profiles so hot startup code is AOT-compiled at install.

Main-thread I/O and Binder

SharedPreferences loads, synchronous Binder calls to busy services, lock contention. Visible in Perfetto as "blocked" main-thread slices.

Analogy

Starting a car. A cold start is a car parked overnight in winter: you must unlock, start the engine and warm it up. A warm start is a car whose engine is still running but you got out and must sit back down and adjust the seat. A hot start is a car at a red light: just press the accelerator. The engine is the process, the seat adjustment is recreating the Activity, and pressing the accelerator is resuming it.

Interview angle Interviewers ask you to "walk through a cold start" and then "how would you make it faster?". Name the stages (fork, bindApplication, providers, Application.onCreate, Activity create, inflate, first frame via Choreographer and SurfaceFlinger), say how you would measure (am start -W, Perfetto), and give targeted fixes rather than generic "optimize code".

The four app components and their lifecycles

An app is a set of components declared in AndroidManifest.xml. The system, not the app, instantiates them and drives their callbacks, which always run on the app's main thread (except ContentProvider data calls, which run on Binder threads).

Activity

onCreate → onStart → onResume → [RUNNING, in foreground]
                         ▲             │ another activity comes in front
                         │             ▼
                     onResume ◀──── onPause ──▶ onStop ──▶ onDestroy
                                               │    ▲
                                  (user returns)└─▶ onRestart → onStart
  • onCreate: inflate UI, restore state. onStart: visible. onResume: in foreground and interactive.
  • onPause must be quick: the next activity is not resumed until it returns. onStop: no longer visible; release heavier resources.
  • Configuration change (rotation, locale, font size, dark mode): by default the activity is destroyed and recreated. Keep state in a ViewModel (survives config change) and onSaveInstanceState / SavedStateHandle (survives process death).
  • Lifecycle is driven from ATMS through ClientTransaction items (LaunchActivityItem, ResumeActivityItem, PauseActivityItem, ...) sent over IApplicationThread.

Service

KindStarted byLifecycleNotes
StartedstartService()onCreate → onStartCommand ... stopSelf() → onDestroyBackground starts restricted since Android 8.0 (IllegalStateException if app is in background)
ForegroundstartForegroundService() then startForeground()Same as started, but shows a notificationMust call startForeground() promptly or the app is crashed with ForegroundServiceDidNotStartInTimeException. The timeout is release-dependent (public docs have said 5 s or 10 s; treat it as a few seconds, not a fixed number). Android 14 requires a declared foreground service type
BoundbindService()onCreate → onBind ... all clients unbind → onUnbind → onDestroyReturns an IBinder to clients; the client-server pattern, often backed by AIDL (see Binder & AIDL)

onStartCommand return values control restart after a kill: START_STICKY (recreate, null intent), START_NOT_STICKY (do not recreate), START_REDELIVER_INTENT (recreate and redeliver the last intent). For deferrable work, WorkManager/JobScheduler is preferred over long-running services.

BroadcastReceiver

  • Manifest-declared receivers can start the app; since Android 8.0 most implicit broadcasts can no longer be received this way (exceptions such as BOOT_COMPLETED, LOCALE_CHANGED).
  • Context-registered receivers (registerReceiver) live as long as the registering context. Android 14 requires RECEIVER_EXPORTED/RECEIVER_NOT_EXPORTED flags.
  • onReceive() runs on the main thread and must finish within the broadcast timeout (10 s foreground queue, 60 s background queue). For longer work, call goAsync() and finish quickly on another thread, or hand off to WorkManager.
  • Ordered broadcasts go to receivers one by one by priority, each can modify or abort; normal broadcasts are delivered to all receivers without ordering guarantees.
  • A process that is only running a receiver is considered foreground during onReceive, then drops in priority right after.

ContentProvider

  • Exposes structured data through content://authority/path URIs with query/insert/update/delete and openFile.
  • onCreate() runs on the main thread before Application.onCreate(), during bindApplication. Slow providers slow every cold start.
  • Data methods are called on Binder threads, so the implementation must be thread-safe.
  • Large results move through CursorWindow backed by shared memory (ashmem), avoiding the Binder transaction size limit.
  • Access control: android:readPermission/writePermission, exported, and temporary URI grants (FLAG_GRANT_READ_URI_PERMISSION).
  • If a client is using a provider, AMS links their lifetimes: the provider's process priority is raised, and if the provider process dies, a client holding a stable reference is killed too.
Analogy

A restaurant. The Activity is the dining room where customers interact; a Service is the kitchen that keeps cooking even when no one is looking into it (a foreground service is the kitchen with a visible "open" sign); a BroadcastReceiver is the doorbell that rings when a delivery arrives and must be answered quickly; a ContentProvider is the pantry with a service window where other restaurants can request ingredients using a standard order form (the URI). The manager who opens and closes each room is the system (AMS/ATMS), not the staff.

Common pitfall Assuming onDestroy is always called. When the process is killed by LMKD, no callbacks run at all. Persist critical state in onPause/onStop and design for process death.
Interview angle Expect to recite the Activity lifecycle cold, then probe edge cases: what happens on rotation, on process death, when a dialog appears (only onPause if it is a dialog-themed activity; nothing for a plain Dialog), started vs bound service, why ContentProvider onCreate affects startup, and the foreground-service rules on recent Android versions.

Process priority, oom_adj, LMKD and PSI

Android keeps processes alive after the user leaves them, as a cache, so returning to an app is a hot or warm start. When memory runs low, something must be killed. AMS decides how important each process is; LMKD (the Low Memory Killer Daemon) decides when to kill and picks the least important victim.

Importance levels

Level (high to low)Typical oom_score_adjExampleKill likelihood
Native / system-1000 / -900init, native daemons, system_serverNever (by LMKD)
Persistent-800 / -700SystemUI, phone processPractically never
Foreground0Resumed activity, receiver in onReceive, service executing a callbackOnly in extreme pressure
Visible100Visible but not focused (behind a translucent activity)Rare
Perceptible200Foreground service (music playback), IMERare
Backup / heavy-weight300 / 400App in backupLow
Service500Started background service (recent)Possible under pressure
Home600LauncherKept if possible
Previous700The app the user was just inKept if possible
Service B800Old, long-running servicesLikely
Cached900 - 999Background activities, empty processes (LRU order)First to go

AMS recomputes these values in OomAdjuster.updateOomAdjLocked() whenever something relevant happens (activity resumed, service bound, broadcast delivered) and writes them to /proc/<pid>/oom_score_adj. Bindings propagate importance: if a foreground app binds to your service with BIND_AUTO_CREATE, your process is raised close to the client's level. AMS also tracks a separate process state (PROCESS_STATE_TOP, FOREGROUND_SERVICE, CACHED_EMPTY, ...) used for scheduling groups, network and background restrictions.

LMKD and PSI

  • The old in-kernel lowmemorykiller driver (static free-memory thresholds) was removed from mainline Linux in 4.12. Android moved killing to userspace lmkd.
  • Since Android 10, lmkd uses PSI (Pressure Stall Information) monitors on /proc/pressure/memory. PSI measures the share of time tasks are stalled waiting for memory (reclaim, refaults), which is a much better signal than "free pages below X".
  • When stall time crosses thresholds, lmkd picks a victim with the highest oom_score_adj (and largest footprint within that band) and kills it with SIGKILL (via pidfd). Tuning is via ro.lmk.* properties (for example ro.lmk.psi_partial_stall_ms, ro.lmk.swap_free_low_percentage).
  • The kernel OOM killer is the last resort if lmkd fails to free memory in time.
  • zRAM (compressed swap in RAM) and kswapd reclaim soften pressure before kills.
  • Cached apps freezer: introduced in Android 11 as an opt-in feature and turned on by default later (around Android 12L/13, via cached_apps_freezer). Cached processes are frozen with the cgroup v2 freezer so they use no CPU. A synchronous Binder call to a frozen process fails immediately (the driver returns BR_FROZEN_REPLY; the caller sees a frozen or dead-object style error). The call does not unfreeze the target. Oneway (async) transactions are queued in the driver until AMS unfreezes the process because it became important again (activity, broadcast, bind).
adb shell dumpsys activity processes     # oom_adj, procstate, LRU list
adb shell dumpsys activity lru            # compact LRU with adj values
adb shell cat /proc/pressure/memory       # some avg10=0.52 ... full avg10=0.10 ...
adb logcat -b events | grep am_kill       # AMS kills
adb logcat -s lowmemorykiller             # lmkd kill reports
Analogy

A lifeboat with limited seats. The captain (AMS) hands everyone a priority ticket: the pilot and crew (system, persistent) keep their seats no matter what, passengers currently rowing (foreground) are next, and sleeping passengers at the back (cached) have the lowest tickets. The lookout (lmkd) watches how much the boat is struggling (PSI, time spent stalled), not just how many seats are free, and when it gets bad he asks the lowest-ticket passenger to leave. The ticket number is oom_score_adj.

Common pitfall A classic systemic bug is a background service that pins itself at foreground priority (a fake foreground notification or a bind from a persistent process) to avoid being killed. It drains battery and starves real foreground work. On RAM-tight devices such as wearables and low-end phones, policing this is part of platform integration.
Interview angle "How does the system decide what to kill?" A strong answer covers both halves (AMS computes importance and writes oom_score_adj; lmkd decides when using PSI and kills by highest adj), explains how bindings raise priority, and mentions zRAM, the cached-app freezer and the kernel OOM killer as backstop. Interviewers may ask why PSI beats free-memory thresholds.

Handler, Looper and MessageQueue

Almost every framework thread that handles events (the app main thread, and most system_server threads) runs the same pattern: a Looper loops forever, pulling Messages from a MessageQueue and dispatching each one to its target Handler. This keeps UI code single-threaded and event-driven, so views need no locks.

// Simplified Looper.loop()
for (;;) {
    Message msg = queue.next();        // blocks in nativePollOnce() → epoll_wait()
    if (msg == null) return;           // queue quit
    msg.target.dispatchMessage(msg);   // target is the Handler
    msg.recycleUnchecked();            // back to the global Message pool
}

// Handler.dispatchMessage order:
// 1. msg.callback (the Runnable from post())
// 2. Handler's mCallback.handleMessage()
// 3. Handler subclass handleMessage()

Looper

One per thread, stored in a ThreadLocal. Created with Looper.prepare(), run with Looper.loop(). The main thread's Looper is created by ActivityThread.main() via prepareMainLooper().

MessageQueue

A singly linked list sorted by when (uptime ms). next() sleeps in native code until the head is due or a new message wakes it.

Handler

Bound to one Looper. post(), postDelayed(), sendMessage() enqueue work; handleMessage() runs it on the Looper's thread. Any thread can post; only the Looper thread executes.

Message

Carries what, arg1, arg2, obj, when, target. Get it with Message.obtain() to reuse pooled objects.

How the queue sleeps without busy-waiting

MessageQueue.next() calls nativePollOnce(ptr, timeout), which runs epoll_wait() on an eventfd (older releases used a pipe). The timeout equals the delay until the next message is due, or infinite if empty. Enqueueing a message that becomes the new head calls nativeWake(), which writes to the eventfd and wakes the thread. The same epoll set can also watch other file descriptors, which is how input events (the InputChannel socket) and vsync signals are delivered into the main Looper.

Why Looper.loop() does not itself cause an ANR

The main thread sits in an infinite for (;;) loop for the life of the process. That is not a busy-wait and it is not an ANR. While the queue is empty (or the next message is in the future) the thread is asleep in the kernel inside epoll_wait, using no CPU. An ANR is raised only when the system handed the app a specific piece of work (input, broadcast, service start, provider publish) and that work was not acknowledged before a deadline. The idle loop is waiting to do that work; it is not failing to do it. What does cause an ANR is a single message that runs too long, so later messages (including "I finished this input event") never run in time.

Advanced queue features

  • Sync barrier: a message with a null target. While it is at the head, normal (synchronous) messages are held back and only asynchronous messages run. ViewRootImpl.scheduleTraversals() posts a barrier so the next frame's traversal, posted as async by Choreographer, runs before other queued work.
  • IdleHandler: queue.addIdleHandler() runs a callback when the queue is empty, useful for deferred low-priority init.
  • HandlerThread: a ready-made Thread with its own Looper, used widely in system_server (android.fg, android.io, etc.) and for serial background work in apps.
  • Choreographer: receives vsync from SurfaceFlinger (via DisplayEventReceiver) and runs input, animation, traversal and commit callbacks on the main Looper, once per frame (16.6 ms at 60 Hz, 11.1 ms at 90 Hz).
  • Looper.setMessageLogging() and Looper slow-dispatch logs ("Slow dispatch took ...ms") help find long messages; tools like BlockCanary use the same hook.

Binder threads are different

Incoming Binder calls do not run on the main Looper. Each process has a Binder thread pool (by default up to 15 extra threads plus the main Binder thread; system_server raises the limit to 31). A service implementing a Binder method must therefore be thread-safe and fast, and must post to a Handler if it needs to touch main-thread state. Details are in Binder IPC & AIDL.

Analogy

A single barista with an order rail. Customers (any thread) clip order tickets (Messages) onto the rail through the cashier (Handler). The barista (Looper thread) takes tickets one at a time, in time order, and makes each drink. When the rail is empty, the barista sits down (epoll sleep) and a bell (eventfd) wakes him when a new ticket arrives. A sync barrier is a manager putting up a "VIP only" sign so the frame-drawing ticket jumps the queue. If one ticket asks for a 10-minute drink, every customer behind it waits: that is main-thread blocking, and eventually an ANR.

Common pitfall A non-static inner Handler class holds an implicit reference to its Activity. A delayed message keeps the Handler, and so the Activity, alive after onDestroy: a memory leak. Use a static class with a WeakReference, or remove callbacks in onDestroy. Also, new Handler() without a Looper is deprecated; pass Looper.getMainLooper() explicitly.
Interview angle Beyond the basic loop, interviewers probe: how does the queue block without spinning (epoll + eventfd), why does the main thread not burn CPU in an infinite loop, what is a sync barrier, how does Choreographer fit in, and why Binder callbacks run on other threads. Being able to write a HandlerThread example from memory helps.

ANR detection and debugging

An ANR (Application Not Responding) is raised when an app fails to finish a specific piece of work within a deadline, almost always because its main thread is blocked. The framework arms a timer when it hands work to the app and disarms it when the app reports completion. If the timer fires first, the ANR path runs.

TriggerTimeout (AOSP defaults)Who detects it
Input event not acknowledged (key or touch) while a focused window exists5 sNative InputDispatcher in system_server
Input event with no focused window (app has a focused application but no window yet)5 sInputDispatcher ("no focused window")
BroadcastReceiver onReceive10 s foreground queue / 60 s background queueAMS BroadcastQueue
Service onCreate/onStartCommand/onBind20 s foreground / 200 s backgroundAMS ActiveServices
startForegroundService() without startForeground()a few seconds (docs have said 5 s or 10 s; varies by release and DeviceConfig)AMS ActiveServices
ContentProvider publish during process start10 sAMS
JobService callbacks (newer releases)several secondsJobScheduler

What happens when an ANR fires

  1. Detection InputDispatcher or AMS notices the deadline has passed and calls into AnrHelper / ProcessErrorStateRecord.appNotResponding().
  2. Evidence capture The system sends SIGQUIT (signal 3) to the app and to key processes (system_server, related native processes). ART's signal catcher thread dumps all Java thread stacks; native stacks are collected via debuggerd.
  3. Trace files Stacks are written to /data/anr/anr_<timestamp> files (older releases used a single /data/anr/traces.txt). CPU usage is logged to logcat, and the event goes to am_anr in the events log and DropBox (data_app_anr).
  4. User handling A foreground app shows the "App isn't responding" dialog; background ANRs are usually killed silently.

Debug flow

  1. Get the evidence adb bugreport (includes ANR files, logcat, dumpsys) or adb pull /data/anr/. Note the ANR reason line, e.g. "Input dispatching timed out" or "executing service".
  2. Read the main thread first Find "main" prio=5 tid=1. Its state tells you the class of problem: Blocked (waiting for a monitor held by another thread), Waiting/TimedWaiting, Native (often inside a Binder call or I/O), Runnable (busy computing).
  3. Follow the chain If blocked on a lock, the dump says "waiting to lock <0x...> held by thread N". Read thread N. If in BinderProxy.transact, find which process and thread served the call (system_server stacks are in the same dump, or use /sys/kernel/debug/binder/ state).
  4. Check the environment CPU load at the time ("CPU usage from ..." in logcat), memory pressure and lmkd kills, iowait, thermal throttling. An idle main thread with high system load points to starvation, not app code.
  5. Fix and prevent Move I/O and heavy work off the main thread, avoid synchronous Binder calls to slow services from the UI thread, shrink lock scopes, and add StrictMode and ANR-rate regression gates.
"main" prio=5 tid=1 Blocked
  | group="main" sCount=1 ucsCount=0 flags=1 obj=0x72f1a0c8 self=0xb400...
  | sysTid=4821 nice=-10 cgrp=top-app sched=0/0 handle=0x7f...
  at com.example.cache.Store.get(Store.java:88)
  - waiting to lock <0x0a1b2c3d> (a java.lang.Object) held by thread 23
  at com.example.ui.MainActivity.onResume(MainActivity.java:52)
  ...
"DiskWriter" prio=5 tid=23 Native
  at java.io.FileOutputStream.write(Native method)
  at com.example.cache.Store.flush(Store.java:120)
  - locked <0x0a1b2c3d> (a java.lang.Object)

Here the main thread waits for a lock that a background thread holds while doing disk I/O. The fix is not "make the disk faster" but "never hold the shared lock during I/O".

Main-thread I/O

Disk or network on the UI thread. Catch with StrictMode.

Lock contention / deadlock

Main thread waiting on a monitor, or two threads each holding what the other needs.

Slow Binder call

Synchronous call into a busy service (system_server lock contention, exhausted Binder pool, frozen or dead-slow peer).

System starvation

Heavy CPU, iowait, memory thrashing or thermal throttling; the app's code is fine but never gets scheduled.

Analogy

A referee with a shot clock. Every time the system hands the app a ball (an input event, a broadcast, a service start), the referee starts a clock. If the player does not shoot before it runs out, the referee blows the whistle, takes a photo of every player's position (the stack dump from SIGQUIT), and files a report (the ANR trace). Debugging means studying the photo: who was the main player holding, and who was he waiting on?

Common pitfall Blaming the top frame of the main thread. In an input ANR, the stack is captured at the moment of the timeout, which may be after the real long operation already finished. Look for the pattern across multiple ANRs, check Perfetto traces around the event, and correlate with system load before concluding.
Interview angle "You get an ANR report, how do you debug it?" is near-certain. Structure the answer: identify the ANR type and timeout, open the trace, read the main thread, follow lock or Binder chains to the root thread or process, check system-wide load, then fix and add regression protection. Knowing that input ANRs are detected by InputDispatcher (not AMS) and that stacks come from SIGQUIT is a strong signal.

The system_server watchdog

ANRs protect against unresponsive apps. The Watchdog (com.android.server.Watchdog) protects against an unresponsive system_server. If a core service thread or lock is stuck, the whole device would freeze, so the watchdog kills system_server and forces a framework restart instead.

  1. Checkers The watchdog thread owns a list of HandlerCheckers, one per critical thread (main, android.fg, android.ui, android.io, android.display, android.anim, ...). Services can also register monitors (Watchdog.Monitor) whose monitor() method simply grabs and releases their main lock (AMS, WMS, PowerManager, InputManager do this).
  2. Probe Every half-timeout (30 s by default) it posts a check message to each thread. The foreground checker also runs all monitors.
  3. Half-way dump If a checker has not completed after 30 s, the watchdog dumps stacks as an early warning.
  4. Kill After the full timeout (60 s by default), it dumps stacks of system_server and important native processes to /data/anr, logs WATCHDOG KILLING SYSTEM PROCESS: Blocked in handler on ... / Blocked in monitor ..., writes a DropBox entry, and kills system_server.
  5. Recovery Because system_server is Zygote's child and critical, Zygote restarts, all apps die, and the framework boots again (a soft reboot). Repeated crashes of critical processes during boot can trigger rescue party / recovery mode.
W Watchdog: *** WATCHDOG KILLING SYSTEM PROCESS: Blocked in monitor
    com.android.server.am.ActivityManagerService on foreground thread (android.fg),
    Blocked in handler on main thread (main)
W Watchdog: foreground thread stack trace:
W Watchdog:     at com.android.server.am.ActivityManagerService.monitor(...)
W Watchdog:     - waiting to lock <0x0c3a...> (a com.android.server.am.ActivityManagerService)
    held by thread 112 (Binder:1234_7)

Triage is the same as an ANR: find the thread named in the message, find the lock it waits for, find the owner thread (often a Binder thread stuck in an outgoing call to a HAL or another process, or a deadlock between two service locks such as AMS and WMS), and fix the lock ordering or move the slow call out of the locked region.

Note Other watchdogs exist at other layers: the hardware/kernel watchdog (reboots the SoC if the kernel stops petting it), init restarting critical services, and modem-side watchdogs that trigger subsystem restart. Be precise about which one you mean.
Analogy

A night security guard doing rounds in the city hall. Every 30 minutes he knocks on each department's door (posts a check message) and tries each master key (the monitor locks). If a department does not answer after two rounds, he assumes someone has collapsed inside, photographs the scene (stack dump), and triggers a full building evacuation and reopening (system_server restart), because a silent building is worse than a brief closure.

Interview angle Interviewers ask "What is the difference between an ANR and a watchdog?" or show a watchdog log and ask for root cause. Say who is watched (app vs system_server), the timeouts, the result (dialog or kill vs soft reboot), and that the root is usually a lock held across a slow Binder or HAL call.

PackageManager, the app sandbox and permissions

PackageManagerService (PMS) knows every installed package: its components, intent filters, signatures, permissions, UID and code paths. At boot it scans /system/app, /system/priv-app, /product, /vendor, /system_ext, APEX modules and /data/app, then restores state from /data/system/packages.xml. A slow scan shows up directly in boot time.

Install flow

  1. Session An installer (Play Store, adb install, pm install) opens a PackageInstaller session and streams the APK(s) into a staging directory.
  2. Parse PMS parses AndroidManifest.xml (PackageParser2): package name, version, components, permissions, SDK levels, native ABIs.
  3. Verify Checks the signature (APK Signature Scheme v2/v3/v4 over the whole file). An update must be signed by the same certificate (or a rotated one via v3 lineage) or it fails with INSTALL_FAILED_UPDATE_INCOMPATIBLE. Package verifiers may also run.
  4. Assign identity A new app gets an app ID (10000+). Its Linux UID is userId * 100000 + appId, so the same app has a different UID per Android user or work profile.
  5. Prepare files Through installd (a privileged native daemon) PMS creates /data/data/<pkg> (really /data/user/<id>/<pkg>), extracts native libs, and triggers dexopt (dex2oat, now via artd) guided by profiles.
  6. Commit Components are registered for intent resolution, permissions are granted as appropriate, packages.xml is written, and ACTION_PACKAGE_ADDED/REPLACED is broadcast.

Intent resolution

Explicit intents name the component class directly. Implicit intents describe an action, data URI/type and categories; PMS matches them against registered intent filters (queryIntentActivities) and the system shows a chooser or uses a default. Since Android 11, package visibility rules filter what an app can even see unless it declares <queries>.

Sandbox and permission model

  • UID sandbox (DAC): each app is its own Linux user, so its files and processes are isolated by normal kernel permission checks.
  • SELinux (MAC): every process has a domain (untrusted_app, platform_app, system_server, hal_radio_default, ...). Policy decides which files, sockets and Binder services each domain may use, including binder_call and service_manager find rules. Denials appear as avc: denied in logcat/dmesg and are a very common integration snag.
  • Android permissions are checked in the framework, usually in the service's Binder method with the caller's UID.
Protection levelGranted howExample
normalAutomatically at installINTERNET, VIBRATE
dangerous (runtime)User grants at runtime (Android 6.0+), revocable; groupedCAMERA, ACCESS_FINE_LOCATION, READ_CONTACTS
signatureOnly if the requesting app is signed with the same key as the declaring package (e.g. the platform key)BIND_ACCESSIBILITY_SERVICE-style system permissions
privileged (signature|privileged)Pre-installed in priv-app and allowlisted in /etc/permissions/privapp-permissions-*.xmlMODIFY_PHONE_STATE, INSTALL_PACKAGES
special / appopUser toggles in Settings; enforced via AppOpsManagerSYSTEM_ALERT_WINDOW, MANAGE_EXTERNAL_STORAGE
// Typical enforcement inside a system service Binder method
@Override
public void setSomething(int value) {
    mContext.enforceCallingOrSelfPermission(
            android.Manifest.permission.MODIFY_PHONE_STATE, "setSomething");
    final int uid = Binder.getCallingUid();
    final long token = Binder.clearCallingIdentity();
    try {
        doWorkAsSystem(value);   // now runs with system_server's identity
    } finally {
        Binder.restoreCallingIdentity(token);
    }
}

Permission grant state and runtime checks were split out of PMS into PermissionManagerService gradually across Android 10–12 (the service exists from 10; more of the store and the check path moved in 11–12). PMS remains the entry point for package state. Missing entries in the privapp allowlist are a frequent boot failure on new builds: the device logs "Privileged permission ... not in privapp-permissions allowlist" and may refuse to boot if enforcement is on.

Analogy

An apartment complex. PMS is the leasing office: it reads each tenant's application (the manifest), checks their ID (the signing certificate), assigns an apartment number (UID), and keeps the directory of which tenant offers which service (intent filters). Normal permissions are the mailbox key everyone gets; dangerous permissions are keys the tenant must personally approve each time; signature permissions are master keys only given to family members of the building owner (same signing key). SELinux is the security guard who checks everyone's badge at every door, even if they have a key.

Interview angle Expect "what happens when you install an APK", "how is an app sandboxed", and "how does a system service check permissions". Mention both UID DAC and SELinux MAC, Binder.getCallingUid() with clearCallingIdentity, the four protection levels, and the privapp allowlist for preinstalled apps.

The framework to HAL boundary: Treble, HIDL, AIDL and VINTF

Before Android 8.0, the framework loaded vendor HAL libraries directly (hw_get_module() opening .so files in-process). Upgrading Android meant the SoC vendor had to rebuild every HAL for the new framework. Project Treble (Android 8.0) split the OS into a system side (AOSP framework, updatable) and a vendor side (SoC/OEM HALs and kernel), connected only by stable, versioned HAL interfaces over IPC.

system partition (framework, updated by Android release)
   system_server services, e.g. SensorService, RIL.java in the phone process
        │  stable AIDL HAL (versioned, frozen, declared in VINTF)
        │  ↕ /dev/binder  (legacy HIDL: /dev/hwbinder via hwservicemanager)
vendor partition (SoC / OEM, updated on its own schedule)
   HAL service process, e.g. android.hardware.sensors-service, vendor radio daemon
        │  ioctl / sysfs / QMI / shared memory
kernel driver ──▶ hardware / DSP / modem

Evolution of HAL interfaces

EraMechanismTransportNotes
Pre-8.0 legacyC structs, hardware/libhardware, dlopenIn-processNo stability; vendor rebuild each release
8.0 - 12HIDL (.hal files, android.hardware.foo@1.2)/dev/hwbinder, hwservicemanager; or passthrough (same process) for some HALsVersioned by major.minor; minor versions extend by inheritance
11+ (standard from 13)Stable AIDL (@VintfStability interfaces)/dev/binder, normal servicemanagerOne IPC language for framework and HALs; versions frozen as API snapshots. New HIDL HALs are no longer accepted from Android 13/14

The three Binder domains

Device nodeContext managerUsed between
/dev/binderservicemanagerFramework and apps; framework and stable-AIDL HALs
/dev/hwbinderhwservicemanagerFramework and HIDL HALs (legacy)
/dev/vndbindervndservicemanagerVendor process to vendor process (never system)

The wire-level details of these domains are covered in Binder IPC & AIDL.

VINTF: checking that system and vendor fit together

  • The device manifest (/vendor/etc/vintf/manifest.xml plus fragments) lists which HALs and versions the vendor image provides, e.g. android.hardware.radio.voice.IRadioVoice/slot1, version 2.
  • The framework compatibility matrix lists which HALs and versions the framework requires or can use.
  • Symmetrically, the framework manifest and device compatibility matrix describe what the system provides and what the vendor needs.
  • VintfObject checks compatibility at build time, during OTA and at boot. servicemanager also refuses to register a @VintfStability service that is not declared in the manifest.
  • VNDK is the set of system libraries vendor code may link against with a stable ABI; GSI (Generic System Image) plus VTS tests prove a vendor image works with a pure AOSP system image.

Starting HAL services

HAL services are declared in vendor init.rc files with an interface line. Many are lazy: servicemanager asks init to start them on first lookup (ctl.interface_start) and they may exit when unused. A classic bring-up bug: VINTF declares a HAL but no matching init.rc service exists, so the lookup waits forever and the framework sees a null proxy ("no radio", "no sensors"), which looks like a hardware fault but is a configuration error.

# vendor init.rc fragment
service vendor.radio-default /vendor/bin/hw/android.hardware.radio-service.example
    class hal
    user radio
    group radio inet misc
    interface aidl android.hardware.radio.voice.IRadioVoice/slot1
    interface aidl android.hardware.radio.data.IRadioData/slot1
Analogy

A wall socket standard. Before Treble, every appliance was hard-wired into the house, so renovating the house meant rewiring every appliance. Treble installed standard sockets: the house (framework) and the appliances (vendor HALs) can be replaced independently as long as both follow the socket standard (the frozen HAL interface). VINTF is the inspector who checks that every appliance's plug matches a socket in the house before the power goes on, and HIDL to AIDL is the country switching from an old plug shape to a newer universal one.

Interview angle Typical questions: "What did Treble change and why?", "HIDL vs AIDL for HALs", "What is VINTF and when is it checked?", "What are hwbinder and vndbinder for?". A strong answer ties Treble to faster Android upgrades (system-only updates, GSI), explains that stable AIDL won because it unified the IPC stack and tooling, and mentions the lazy-HAL / VINTF mismatch failure mode.

dumpsys and the framework debugging toolbox

dumpsys asks a Binder service to write its internal state: it looks the service up in servicemanager and calls its dump() method (protected by the DUMP permission). Knowing which dump answers which question is how you localize a systemic issue quickly.

CommandWhat it answers
dumpsys activity activitiesTasks, back stack, resumed/focused activity
dumpsys activity processes / lruProcesses, oom_adj, process state, LRU order
dumpsys activity services <pkg>Running/bound services and their clients
dumpsys activity broadcastsBroadcast queues, pending and recent broadcasts
dumpsys window (windows, displays)Window list, focus, z-order, visibility
dumpsys inputInput devices, InputDispatcher state, focused window, recent ANRs
dumpsys SurfaceFlingerLayers, composition type (GPU vs HWC), refresh rate
dumpsys gfxinfo <pkg>Frame timings, janky frames percentage
dumpsys package <pkg>Version, UID, paths, permissions granted, components, signatures
dumpsys powerWakelocks held, screen/doze state, suspend blockers
dumpsys batterystatsPer-app power attribution; input to Battery Historian
dumpsys meminfo <pkg>PSS/RSS, Java heap, native heap, graphics memory
dumpsys cpuinfoRecent CPU usage by process
dumpsys sensorserviceActive sensor connections, rates, batching
dumpsys alarm / jobschedulerScheduled alarms and jobs, who wakes the device
dumpsys connectivity / telephony.registryNetworks, default network, telephony state
dumpsys -l / service listAll registered services
lshalHIDL HALs and their clients (legacy)

Beyond dumpsys

  • Perfetto (and legacy systrace): timeline of CPU scheduling, thread states, Binder transactions, main thread and RenderThread slices, memory counters. The best tool for jank, startup and ANR root cause.
  • bugreport: adb bugreport bundles logcat, all dumpsys output, ANR/tombstone files, kernel log and system properties.
  • logcat buffers: main, system, events (am_proc_start, am_kill, am_anr, wm_on_resume_called), crash, radio.
  • Stack dumps on demand: kill -3 <pid> for Java stacks, debuggerd -b <pid> for native backtraces.
  • am / cmd: am start -W, am force-stop, am kill, cmd package, cmd activity.
  • Tombstones in /data/tombstones/ for native crashes; DropBox (dumpsys dropbox) for system_server crashes, ANRs and watchdogs.
Analogy

A hospital's diagnostic tests. dumpsys is asking each department for its patient chart (a snapshot of state), Perfetto is a continuous ECG recording (what happened over time), a bugreport is the full medical file, and a stack dump is an X-ray taken at one instant. A good doctor picks the right test for the symptom instead of ordering everything.

Tip Many dumps accept filters: dumpsys activity -h lists sub-commands. On a busy device, dumpsys without arguments can take minutes; always name the service.
Interview angle "Which dumpsys commands do you reach for and why?" Answer by symptom: process died or ANR (activity, input), jank (gfxinfo, SurfaceFlinger, Perfetto), battery (power, batterystats, alarm, jobscheduler), install/permission (package), memory (meminfo, activity lru, PSI). Naming the right one immediately shows hands-on experience.

Launch modes, tasks and task affinity

ATMS does not start a new activity in isolation. It places the activity in a task (a back stack the Recents screen shows as one card) according to the launch mode on the activity, the flags on the incoming Intent, and android:taskAffinity. Getting this wrong is a frequent source of "two copies of my activity" and "Back does not do what I expect" bugs.

Launch modeWhat ATMS doesTypical use
standard (default)Always create a new instance in the caller's task (unless Intent flags say otherwise)Most screens
singleTopIf the target is already the top of the task, reuse it and deliver onNewIntent(); otherwise create a new instanceSearch, notification taps that should not stack duplicates
singleTaskLook for an existing task with this activity at the root. If found, bring that task forward and clear everything above the activity, then onNewIntent(). If not, start a new taskApp entry points, launchers
singleInstanceLike singleTask, but the task holds only this activity; nothing else is allowed into itSystem-style isolated UIs (rare in apps)
singleInstancePerTask (Android 12+)Only one instance per task, but more than one such task may existMulti-window / multiple-instance apps

Intent flags that override the manifest

  • FLAG_ACTIVITY_NEW_TASK: start in a new task (required when starting an activity from a non-Activity context).
  • FLAG_ACTIVITY_CLEAR_TOP: if the activity already exists in the task, finish everything above it and reuse it (onNewIntent).
  • FLAG_ACTIVITY_SINGLE_TOP: same as singleTop for this launch only.
  • FLAG_ACTIVITY_CLEAR_TASK + NEW_TASK: replace the whole task with this activity (common from notifications and shortcuts).
  • FLAG_ACTIVITY_REORDER_TO_FRONT: bring an existing instance forward without clearing what is above it in the stack history.

Task affinity

android:taskAffinity is a string (default: the package name) that names which task an activity wants to belong to. It matters when combined with allowTaskReparenting, FLAG_ACTIVITY_NEW_TASK, or singleTask: ATMS can move the activity into (or start) a task whose affinity matches. Two activities with different affinities can therefore live in two Recents cards even though they are in the same app. A common OEM bug is giving a system activity the wrong affinity so it appears as a stray Recents entry or steals the back stack of another app.

Launcher starts MainActivity (standard, affinity = com.example)
  → task T1: [Main]
Main starts SettingsActivity (standard)
  → task T1: [Main, Settings]
Notification fires with NEW_TASK | CLEAR_TASK → MainActivity
  → task T1 replaced: [Main]     (Settings finished)
singleTask RootActivity already exists under T2
  → T2 brought forward; activities above Root cleared; onNewIntent()
Common pitfall Starting an activity from an Application or Service context without FLAG_ACTIVITY_NEW_TASK throws AndroidRuntimeException: Calling startActivity() from outside of an Activity context requires the FLAG_ACTIVITY_NEW_TASK flag. The other frequent mistake is using singleTask on an inner screen, which unexpectedly clears the back stack.
Interview angle Recite the four classic launch modes and what onNewIntent means, then explain that Intent flags can override the manifest and that affinity decides which task the activity joins. Bonus: dumpsys activity activities prints the real task stacks.

Activity vs Application context, window tokens and BadTokenException

A Context is not "the app". There are several implementations with different lifetimes and different abilities. Using the wrong one is a classic source of leaks and of WindowManager$BadTokenException.

ContextLifetimeHas a window token?Safe for
Activity (and ContextThemeWrapper)Until the activity is destroyedYes, after the window is attachedUI: inflate with the activity theme, show dialogs, start activities, bind views
ApplicationThe processNoSingletons, caches, WorkManager, starting a service, anything that must outlive a screen
ServiceUntil the service is destroyedNo (unless it is a special system window)Background work, notifications. Not dialogs
ContextImpl from createPackageContext / createDeviceProtectedStorageContextAs long as you hold itNoTalking to another package's resources, or DE storage before user unlock

Window tokens

Every window WMS will accept must be associated with a window token (IBinder) that WMS already knows about: an activity's token (created when ATMS starts the activity), an application token, or a special token for system windows (TYPE_APPLICATION_OVERLAY and friends, gated by SYSTEM_ALERT_WINDOW). The token is how WMS ties a ViewRootImpl / SurfaceControl to a task, decides focus, and knows when to remove the windows (activity finish → token removed → windows gone).

BadTokenException: Unable to add window -- token null is not valid (or "token is not valid; is your activity running?") means you asked WindowManager.addView() (a Dialog, a PopupWindow, a Toast on older releases, a custom overlay) with a token WMS does not accept: typically an Application context, or an Activity that has already finished / whose window is not yet attached (onCreate before onAttachedToWindow, or after onDestroy).

  • Show dialogs from the Activity, after it is resumed, and dismiss them in onDestroy.
  • Never store an Activity in a static field or a long-lived singleton; that leaks the whole window.
  • For process-wide work use getApplicationContext(). For UI that needs a theme or a token, use the Activity.
  • LayoutInflater from the Application context will not apply the activity theme; inflate from the Activity (or a ContextThemeWrapper).
Analogy

The Application context is a company badge that gets you into the building but not into a meeting room. An Activity context is a booked room with a door key (the window token). Trying to hang a projector screen (a Dialog) in the lobby with only the company badge is BadTokenException: WMS will not put a window where there is no room.

Interview angle "When do you use Application context vs Activity context?" and "What is a BadTokenException?" A strong answer names the window token, says WMS requires a token it already knows, and gives the two usual causes: Application context, and a finished or not-yet-attached Activity.

Background restrictions and PendingIntent mutability

Each recent Android release tightened what a background app may do. Interviewers expect the timeline, not a vague "background is restricted".

ReleaseRestrictionWhat you do instead
8.0Most implicit broadcasts cannot be received from the manifest; background startService() throwsJobScheduler / WorkManager; explicit broadcasts; startForegroundService()
9–10App Standby buckets, background location limits, Activity starts from the background restrictedForeground service for ongoing user-visible work; a notification trampoline only when the user tapped something
12Apps cannot call startForegroundService() while in the background except for a short allowlist (user interaction, high-priority FCM, exact alarm, being a foreground-service owner already, ...). PendingIntents must declare mutabilityUse an exempted path, or a user-visible activity, or a job. Mark every PendingIntent FLAG_IMMUTABLE or FLAG_MUTABLE
13Runtime permission POST_NOTIFICATIONS. Without it, FGS notifications may be hidden and the FGS is less usefulRequest the permission; do not assume a posted notification is visible
14Every FGS must declare a foregroundServiceType in the manifest and pass matching types to startForeground(). Several types need a special permission or a special allowlist. Short-service and data-sync types have their own time limitsPick the real type (mediaPlayback, location, microphone, specialUse, ...); do not fake mediaPlayback to stay alive

PendingIntent.FLAG_IMMUTABLE (Android 12+)

A PendingIntent is a token another process (NotificationManager, AlarmManager, App Widget host, another app) can send later, acting as your app. Until Android 12 the extras and component on that Intent could be filled in by the holder. That is a confused-deputy hole: a malicious app that received your mutable PendingIntent could retarget it. From Android 12, creating a PendingIntent without FLAG_IMMUTABLE or FLAG_MUTABLE throws IllegalArgumentException on apps targeting 12+.

  • Use FLAG_IMMUTABLE unless you genuinely need the holder to fill in extras (inline reply on a notification, a mutable widget click). Almost all notification and alarm PendingIntents should be immutable.
  • Immutability is about the Intent identity (component, action, extras the sender is allowed to change), not about whether the PendingIntent can be sent more than once.
  • Combine with FLAG_UPDATE_CURRENT or FLAG_CANCEL_CURRENT as needed; those flags are orthogonal.
PendingIntent pi = PendingIntent.getActivity(
        context, 0, intent,
        PendingIntent.FLAG_UPDATE_CURRENT | PendingIntent.FLAG_IMMUTABLE);
Interview angle Walk 8 → 12 → 13 → 14 without notes. Then: "Why FLAG_IMMUTABLE?" Because a PendingIntent is a capability you hand to another process; mutable extras let that process change what your app will do when the token is sent.

How to add a system service

A platform interview often asks you to add a first-party service that apps reach through a manager class. The pieces are always the same; missing any one of them is a bring-up failure.

  1. AIDL Define IFooService.aidl under frameworks/base/core/java (or an APEX). The build generates Stub and Proxy. Hide it from the public SDK if it is privileged (@hide or a SystemApi).
  2. SystemService subclass Implement com.android.server.foo.FooService extending SystemService (or a raw Binder if it is a very old-style service). In onStart() construct the Binder implementation and call publishBinderService("foo", mBinder), which is ServiceManager.addService() with extra bookkeeping.
  3. SystemServer Start it in the right wave: bootstrap only if something in bootstrap needs it; otherwise startOtherServices(). Optionally react to boot phases in onBootPhase().
  4. Manager and Context.getSystemService Add a public FooManager that looks up the Binder (ServiceManager.getService("foo") / IFooService.Stub.asInterface) and caches it. Register the manager in SystemServiceRegistry so apps call context.getSystemService(FooManager.class).
  5. SELinux Add a service name in service_contexts (label e.g. foo_service), allow system_server to add it, and allow client domains (platform_app, priv_app, untrusted_app as appropriate) to find and binder_call. Without this, addService or getService fails with avc: denied.
  6. Permission and identity In every Binder method, enforceCallingPermission (or AppOps) and wrap privileged work in clearCallingIdentity / restoreCallingIdentity.
  7. Dump and tests Implement dump() for dumpsys foo, and add CTS or unit tests against the AIDL.
IFooService.aidl  →  FooService (system_server thread)  →  publishBinderService("foo")
                                                              │
apps / FooManager  ←── ServiceManager.getService("foo")  ←────┘
SELinux: service_contexts + binder_call / service_manager add|find

Vendor HALs are a different path (VINTF + init.rc interface aidl + NDK registration). Do not confuse "add a framework service" with "add a HAL". See Binder registration and the HAL boundary.

Interview angle List the artifacts in order: AIDL, SystemService, publishBinderService, SystemServer start site, manager + SystemServiceRegistry, SELinux service_contexts, permission checks. Naming SELinux last-and-always is what separates platform people from app people.

trimMemory, onLowMemory and ApplicationExitInfo

The framework does not only kill processes. It also asks them to shrink, and after a death it records why they died so you are not guessing from logcat alone.

onTrimMemory vs onLowMemory

ComponentCallbacks2.onTrimMemory(int)onLowMemory()
WhenWhenever AMS wants apps to release caches, with a level hintA last-resort callback when the system is critically low on memory (roughly the old "empty process / complete" path)
LevelsTRIM_MEMORY_RUNNING_MODERATE/LOW/CRITICAL (you are in the foreground), UI_HIDDEN, BACKGROUND, MODERATE, COMPLETENo levels; treat it as "release everything you can"
Modern useThe callback you should implement. Drop image caches, close unused files, trim LruCaches. UI_HIDDEN is the usual moment to drop UI bitmapsStill invoked on some paths for compatibility, often together with TRIM_MEMORY_COMPLETE. Do not rely on it as the only signal; it is not guaranteed before an LMK kill

Neither callback runs if the process is frozen or if lmkd sends SIGKILL. Design caches so losing the process is safe, and use onTrimMemory to make a kill less likely.

ApplicationExitInfo (Android 11+)

After a process dies, AMS keeps a ring buffer of ApplicationExitInfo records per package, readable via ActivityManager.getHistoricalProcessExitReasons() (and dumpsys activity exit-info). Each record has a reason, a status, a description, a timestamp, the pid, and sometimes a trace (ANR) or tombstone fd.

Reason (subset)Meaning
REASON_EXIT_SELFSystem.exit / process called exit()
REASON_SIGNALEDNative crash or signal (including SIGKILL from lmkd)
REASON_LOW_MEMORYKilled specifically for memory pressure
REASON_CRASH / REASON_CRASH_NATIVEUncaught Java exception or native crash
REASON_ANRANR; a trace file may be attached
REASON_DEPENDENCY_DIEDA provider or other dependency died and AMS killed the client
REASON_OTHER (and extras such as package update, permission change, user request)Force-stop, app update, am kill, freezer-related kills, and similar

This is the first tool for "why did my process die?" — better than scraping am_kill by hand. Combine it with dumpsys activity lru (who was cached) and Perfetto (what memory was doing).

Interview angle "onLowMemory vs onTrimMemory" is a filter question: trimMemory is the real API, with levels; onLowMemory is the legacy last-ditch callback and is not a promise you will see it before a kill. Follow-up: "How do you know why a process died after the fact?" → ApplicationExitInfo.

Quick revision

  • Stack: apps → (Binder) → system_server services → native daemons → (stable AIDL / HIDL) vendor HALs → kernel drivers.
  • system_server hosts AMS, ATMS, WMS, PMS, PowerMS, DisplayMS, InputMS, SensorService and about 100 more; if it dies the framework soft-reboots.
  • SystemServer starts services in waves: bootstrap, core, other (and APEX); services react to boot phases up to PHASE_BOOT_COMPLETED (1000).
  • ATMS (activities, tasks) was split from AMS (processes, services, broadcasts, providers) in Android 10.
  • SurfaceFlinger is a separate native process; WMS sets window policy and sends SurfaceControl transactions to it.
  • Zygote preloads ART, classes and resources, then forks every app process; pages are shared copy-on-write.
  • Zygote uses a Unix socket, not Binder, because forking a multithreaded process is unsafe.
  • New app process runs ActivityThread.main(), prepares the main Looper, calls attachApplication(); AMS calls back through IApplicationThread.
  • Cold start = no process; warm = process alive, activity recreated; hot = activity just resumed.
  • Measure startup with am start -W, the "Displayed" log (TTID), reportFullyDrawn() (TTFD) and Perfetto.
  • ContentProvider onCreate runs before Application.onCreate; provider data calls run on Binder threads.
  • Activity lifecycle: onCreate, onStart, onResume, onPause, onStop, onDestroy (onRestart when returning); config change recreates by default.
  • Services: started, foreground (must call startForeground() within a few seconds; the timeout is 5 s or 10 s depending on the release), bound (returns an IBinder).
  • onReceive() runs on the main thread; use goAsync() for short async work, WorkManager for long work.
  • AMS computes oom_score_adj (0 foreground ... 900-999 cached) in OomAdjuster; bindings raise the server's priority.
  • lmkd is a userspace daemon using PSI (/proc/pressure/memory) to decide when to kill; kills the highest adj first.
  • Cached apps can be frozen with the cgroup v2 freezer (feature in Android 11, default-on later); a sync Binder call to a frozen process fails with BR_FROZEN_REPLY and does not unfreeze it; oneway is queued. The kernel OOM killer is the last resort.
  • Looper loops over a time-sorted MessageQueue; next() sleeps in epoll_wait on an eventfd, so no busy-waiting.
  • Sync barriers let Choreographer's asynchronous frame messages jump ahead of normal messages.
  • Binder calls run on the Binder thread pool (15 + 1 by default, 31 in system_server), never on the main Looper.
  • ANR timeouts: input 5 s, broadcast 10 s fg / 60 s bg, service 20 s fg / 200 s bg, provider publish 10 s.
  • Input ANRs are detected by native InputDispatcher; broadcast/service ANRs by AMS.
  • On ANR the system sends SIGQUIT and writes stacks to /data/anr/anr_* (formerly traces.txt); read the main thread first.
  • Main causes of ANRs: main-thread I/O, lock contention/deadlock, slow synchronous Binder calls, system starvation.
  • The Watchdog checks key system_server threads and monitor locks every 30 s; blocked 60 s means it kills system_server.
  • PMS parses manifests, verifies signatures, assigns UID = userId * 100000 + appId, and uses installd/artd for files and dexopt.
  • Permission levels: normal, dangerous (runtime), signature, privileged (allowlisted), special/appop.
  • Services check callers with Binder.getCallingUid(); wrap system work in clearCallingIdentity()/restoreCallingIdentity().
  • Sandbox = per-app UID (DAC) plus SELinux domains (MAC); avc: denied logs show policy blocks.
  • Treble (8.0) separates system and vendor via stable HAL interfaces; HIDL (hwbinder) is legacy, stable AIDL (binder) is current.
  • Binder domains: /dev/binder (framework, AIDL HALs), /dev/hwbinder (HIDL), /dev/vndbinder (vendor to vendor).
  • VINTF manifests and compatibility matrices must match; lazy HAL declared in VINTF without an init.rc entry gives a null proxy forever.
  • Key dumps: activity, window, input, SurfaceFlinger, gfxinfo, package, power, batterystats, meminfo; Perfetto for timelines.
  • Launch modes: standard, singleTop, singleTask, singleInstance (plus singleInstancePerTask); Intent flags can override; taskAffinity names the preferred task.
  • Application context has no window token; dialogs and addView need an Activity token or you get BadTokenException.
  • Background timeline: 8 implicit-broadcast / background-service limits; 12 FGS-from-background and PendingIntent mutability; 13 notification permission; 14 FGS types.
  • PendingIntent on apps targeting 12+ must set FLAG_IMMUTABLE or FLAG_MUTABLE; prefer immutable.
  • Adding a system service: AIDL, SystemService, publishBinderService, SystemServer start site, manager + SystemServiceRegistry, SELinux service_contexts, permission checks.
  • Looper.loop() sleeps in epoll_wait when idle; the infinite loop is not an ANR. ANR is a missed deadline on work the system handed you.
  • onTrimMemory is the real cache-trim callback (with levels); onLowMemory is a legacy last-ditch hook and is not guaranteed before a kill.
  • ApplicationExitInfo (Android 11+) is the first place to look for why a process died (LMK, ANR, crash, dependency, self-exit).

Glossary

ActivityThread
The class whose main() is the entry point of every app process; owns the main Looper and handles lifecycle messages from the system.
AMS (ActivityManagerService)
System service managing processes, services, broadcasts, providers, oom_adj scoring and several ANR types.
ANR
Application Not Responding: raised when an app misses a deadline for input, a broadcast, a service or a provider.
ApplicationExitInfo
Per-package record of why a process died (reason, status, optional ANR/tombstone), from Android 11.
APEX
Container format for updatable system components (Mainline modules), mounted at boot.
ATMS (ActivityTaskManagerService)
Service managing activities, tasks and the back stack; split from AMS in Android 10.
BadTokenException
WMS rejected addView because the window token is null, finished, or not an Activity/system token.
Binder thread pool
Threads in each process that execute incoming Binder calls; default maximum 15 plus the main Binder thread.
Choreographer
Per-thread coordinator that runs input, animation and drawing callbacks once per vsync.
Cold / warm / hot start
App starts that need a new process, a new activity in an existing process, or only a resume, respectively.
Copy-on-write (COW)
Memory pages shared after fork until one side writes, at which point that page is copied.
dexopt
Ahead-of-time compilation of app dex code to native code (dex2oat), guided by profiles.
dumpsys
Tool that calls a Binder service's dump() to print its internal state.
eventfd
Linux kernel object used as a lightweight wake-up signal; MessageQueue sleeps on it through epoll.
FLAG_IMMUTABLE
PendingIntent flag (required to choose mutability from Android 12) that prevents the holder from changing the Intent.
Handler
Object bound to a Looper that enqueues Messages/Runnables and handles them on that Looper's thread.
HandlerThread
A Thread that creates its own Looper, used for serialized background work.
HIDL
HAL Interface Definition Language (Android 8-12), carried over hwbinder; now deprecated in favour of stable AIDL.
IApplicationThread
Binder interface an app gives AMS so the system can call back into the app (bind, launch, pause, trim memory).
InputDispatcher
Native thread in system_server that sends input events to windows and detects input ANRs.
installd
Privileged native daemon that performs file operations for PMS (create data dirs, dexopt triggers, delete).
Launch mode
Manifest (or Intent-flag) rule for whether ATMS creates a new activity instance and which task it joins.
lmkd
Low Memory Killer Daemon; userspace process that kills apps under memory pressure using PSI signals.
Looper
Per-thread loop that takes Messages from a MessageQueue and dispatches them.
MessageQueue
Time-ordered list of Messages for one Looper, with native blocking via epoll.
oom_score_adj
Per-process value (-1000 to 1000) in /proc/<pid>/ set by AMS; higher means killed sooner.
PMS (PackageManagerService)
Service that installs, parses and tracks packages, components, signatures and UIDs.
PowerManagerService
Service owning wakelocks, screen state and doze/suspend decisions.
PSI
Pressure Stall Information: kernel metric of time tasks stall on memory, CPU or I/O.
SELinux
Mandatory access control in the kernel; every process runs in a domain with explicit allow rules.
servicemanager
Binder context manager that maps service names to Binder handles.
Soft reboot
Restart of the framework (system_server and Zygote) without restarting the kernel.
SurfaceFlinger
Native compositor process that combines window buffers and sends them to the display via HWC.
Sync barrier
Special message that blocks synchronous messages so asynchronous ones (frame work) run first.
system_server
The privileged Java process hosting most Android system services.
Task affinity
String naming the task an activity prefers to live in; default is the package name.
Treble
Android 8.0 architecture that separates the framework from vendor code with stable HAL interfaces.
TTID / TTFD
Time to initial display (first frame) and time to full display (reportFullyDrawn()).
USAP
Unspecialized App Process: a pre-forked Zygote child waiting to become an app.
VINTF
Vendor Interface object: manifests and compatibility matrices that declare and check HAL versions.
VNDK
Vendor Native Development Kit: system libraries vendor code may use with a stable ABI.
Watchdog
system_server thread that kills system_server if critical threads or locks are stuck for 60 s.
Window token
An IBinder WMS already knows (usually the activity token) that authorizes adding windows.
WMS (WindowManagerService)
Service managing windows, focus, z-order, rotation and transitions.
zRAM
Compressed swap device in RAM that delays memory pressure.
Zygote
Preloaded process that forks system_server and every app process. The init service is named zygote; a 32-bit helper is zygote_secondary.

Interview questions

Fundamentals

What is the "Android framework" and how is it layered?

It is the Java/Kotlin layer of system services plus the SDK APIs that apps call. Apps run in their own processes and call thin manager classes (ActivityManager, PackageManager), which forward over Binder to services in system_server. Below that are native daemons (SurfaceFlinger, servicemanager, installd), then vendor HALs reached over stable AIDL or legacy HIDL, then kernel drivers. Each boundary has a defined mechanism: Binder, HAL IPC, and syscalls/ioctl.

What is system_server and what happens if it crashes?

It is the privileged Java process, forked by Zygote at boot, that hosts most system services (AMS, ATMS, WMS, PMS, PowerManager, InputManager, and many more). If it crashes or the watchdog kills it, Zygote restarts too, every app process dies, and the framework boots again. The kernel and some native daemons keep running, so this is called a soft reboot. One bad lock in a core service is therefore a whole-device stability event.

Name the key system services and what each owns.
  • AMS: processes, services, broadcasts, providers, oom_adj, several ANR types.
  • ATMS: activities, tasks, back stack, recents.
  • WMS: windows, focus, z-order, rotation, transitions.
  • PMS: packages, components, signatures, UIDs, intent resolution.
  • PowerManagerService: wakelocks, screen and doze state.
  • InputManagerService: input devices and event dispatch.
  • DisplayManagerService, SensorService, AlarmManager, JobScheduler, ConnectivityService, NotificationManagerService.
What is Zygote and why does Android use it?

Zygote is a process started by init that loads the ART runtime, preloads common framework classes, resources and libraries, and then waits for requests to fork new processes. Every app process (and system_server) is forked from it. Forking skips VM startup and class loading, so apps start quickly, and the preloaded pages are shared copy-on-write, which saves a lot of RAM across dozens of processes.

What is servicemanager?

It is the Binder context manager, the process that owns handle 0. Services register a name and a Binder object with addService(); clients look them up with getService() or waitForService() and receive a handle they then call directly. It also enforces SELinux rules on who may register or find each service. More in Binder IPC & AIDL.

Recite the Activity lifecycle.

onCreate (set up UI, restore state) → onStart (visible) → onResume (foreground, interactive). When another activity covers it: onPause → onStop. When the user returns: onRestart → onStart → onResume. When finished or reclaimed: onDestroy. onPause must be quick because the next activity waits for it; if the process is killed, no further callbacks are delivered.

What happens to an Activity on rotation?

Rotation is a configuration change. By default the activity is destroyed and recreated with the new configuration so resources can be reloaded. State in a ViewModel survives this; UI state is saved with onSaveInstanceState. An app can opt out with android:configChanges and handle onConfigurationChanged() itself, but that is usually discouraged.

What are the four app components?

Activity (a UI screen), Service (background or bound work without UI), BroadcastReceiver (reacts to system or app broadcasts), and ContentProvider (exposes structured data through content:// URIs). They are declared in the manifest, instantiated by the system, and their lifecycle callbacks run on the main thread.

Started service vs bound service vs foreground service?

A started service is launched with startService() and runs until it calls stopSelf(). A bound service is created by bindService(), returns an IBinder for clients to call, and is destroyed when all clients unbind. A foreground service is a started service that shows an ongoing notification and gets high priority; it must call startForeground() soon after startForegroundService() and, on Android 14+, declare a foreground service type. A service can be both started and bound.

What is a Looper, a Handler and a MessageQueue?

A Looper is a per-thread loop that repeatedly takes the next Message from its MessageQueue and dispatches it. The MessageQueue is a list of Messages sorted by their due time. A Handler is bound to one Looper; any thread can use it to post Runnables or send Messages, and they execute on the Looper's thread. The main thread of every app runs such a loop.

Why must you not block the main thread?

The main thread processes input, lifecycle callbacks and frame rendering one message at a time. If one message takes long (disk I/O, network, a slow Binder call, waiting on a lock), all others wait: frames are dropped (jank, at 16.6 ms per frame at 60 Hz) and, if it lasts long enough, the system raises an ANR (5 s for input). Heavy work belongs on worker threads, coroutines or a HandlerThread, with results posted back.

What is an ANR and what are the main timeouts?

Application Not Responding: the app failed to complete work the system handed it before a deadline. Defaults: input event not handled in 5 s; BroadcastReceiver.onReceive 10 s (foreground queue) or 60 s (background); service lifecycle calls 20 s (foreground) or 200 s (background); startForeground() a few seconds after startForegroundService() (docs have said 5 s or 10 s; it varies by release); ContentProvider publish 10 s.

Where are ANR traces stored?

In /data/anr/. Modern Android writes one file per ANR (anr_YYYY-MM-DD-HH-MM-SS-mmm); older releases used a single traces.txt. They are included in adb bugreport, and the event is also logged as am_anr in the events buffer and stored in DropBox.

What is oom_adj?

A per-process importance score that AMS computes from the process's state (foreground activity, visible, service, cached, etc.) and writes to /proc/<pid>/oom_score_adj. Values range from -1000 (never kill) to 1000; foreground apps are 0 and cached apps 900-999. lmkd kills processes with the highest value first when memory is low.

What is LMKD?

The Low Memory Killer Daemon, a userspace process that watches memory pressure and kills app processes to free memory before the system thrashes. It replaced the old in-kernel lowmemorykiller driver. Modern lmkd uses PSI (Pressure Stall Information) to decide when to act and picks victims by oom_score_adj.

What is Project Treble?

An Android 8.0 re-architecture that separates the system (framework) side from the vendor (SoC/OEM) side with stable, versioned HAL interfaces over IPC. The framework can be updated without rebuilding vendor HALs, which speeds up Android upgrades and allows Generic System Images. The boundary is checked by VINTF manifests and compatibility matrices.

HIDL vs stable AIDL for HALs?

HIDL (Android 8-12) defined HALs in .hal files and used /dev/hwbinder with hwservicemanager. Stable AIDL (supported for HALs from Android 11, standard from 13) uses the same AIDL language and /dev/binder as the framework, with frozen versioned snapshots for stability. HIDL is deprecated; new HALs must be AIDL. The mental model (Stub/Proxy, async callbacks) is the same.

What does PackageManagerService do?

It scans and parses installed packages, verifies signatures, assigns each app a UID, records components and intent filters, resolves intents to components, and manages install/update/uninstall (using installd for file operations and dexopt). Permission grant state and runtime checks moved into PermissionManagerService gradually across Android 10–12; PMS is still the entry point for package state. Its state is persisted in /data/system/packages.xml.

How is an app sandboxed?

Each app runs as its own Linux UID, so the kernel's normal permission checks isolate its files and processes. On top of that, SELinux assigns the process a domain such as untrusted_app that restricts what files, sockets and Binder services it may touch, even if file permissions would allow it. Android permissions are then enforced by services when the app calls them over Binder, using the caller's UID.

What are the permission protection levels?

Normal (auto-granted at install), dangerous/runtime (user must grant, revocable), signature (only for apps signed with the declaring package's key), privileged (preinstalled priv-apps listed in the privapp allowlist), and special/appop permissions toggled in Settings such as draw-over-apps.

Is SurfaceFlinger part of system_server?

No. SurfaceFlinger is a separate native daemon started by init. WMS in system_server decides window policy and sends layer changes to SurfaceFlinger through SurfaceControl transactions; apps send their rendered buffers through BufferQueues; SurfaceFlinger composes everything (with the HWC HAL) each vsync.

What is dumpsys?

A command-line tool that looks up a Binder service by name and calls its dump() method, printing that service's internal state. Examples: dumpsys activity, dumpsys window, dumpsys package <pkg>, dumpsys power, dumpsys meminfo. It requires the DUMP permission (available to the shell user).

What is the difference between cold, warm and hot start?

Cold: no process exists, so the system forks one from Zygote, binds the application and creates the activity. Warm: the process is alive but the activity must be recreated. Hot: the activity is still in memory and is simply brought to the front and resumed. Cold is slowest and is the usual startup benchmark.

What are the activity launch modes?

standard always creates a new instance in the caller's task. singleTop reuses the instance if it is already at the top and delivers onNewIntent(). singleTask finds or starts a task whose root is that activity, clears anything above it, and delivers onNewIntent(). singleInstance is singleTask but the task may hold only that activity. Android 12 added singleInstancePerTask. Intent flags such as NEW_TASK, CLEAR_TOP and CLEAR_TASK can override the manifest for one launch.

Activity context vs Application context?

An Activity context is tied to that screen: it has the activity theme and a window token WMS will accept for dialogs and other windows, and it dies with the activity. The Application context lives for the process, has no window token, and is the right owner for singletons, caches and services. Using Application to show a Dialog causes BadTokenException; storing an Activity in a singleton leaks the window.

What is PendingIntent.FLAG_IMMUTABLE and why did Android 12 require it?

A PendingIntent is a token another process can send later as your app. If it is mutable, the holder can change extras or even the component (a confused-deputy bug). From Android 12, apps targeting 12+ must pass FLAG_IMMUTABLE or FLAG_MUTABLE when creating one. Use immutable unless the holder must fill in extras (inline reply). The flag is orthogonal to FLAG_UPDATE_CURRENT.

onTrimMemory vs onLowMemory?

onTrimMemory(int) is the real API: AMS asks you to drop caches and passes a level (running moderate/low/critical, UI hidden, background, complete). onLowMemory() is a legacy last-ditch callback for critical pressure and is often paired with TRIM_MEMORY_COMPLETE. Neither runs if the process is frozen or SIGKILL'd by lmkd, so do not rely on onLowMemory as the only signal.

Going deeper

Walk me through a cold app start in detail.
  1. Launcher calls startActivity(); ATMS resolves the intent (via PMS) and finds no process for the app.
  2. AMS ProcessList.startProcessLocked() → Process.start() sends arguments over the Zygote socket.
  3. Zygote forks; the child sets UID, SELinux context and namespaces, then runs ActivityThread.main(), which prepares the main Looper.
  4. The app calls attachApplication(IApplicationThread) on AMS over Binder.
  5. AMS calls bindApplication(); the app creates the Application, installs ContentProviders, then runs Application.onCreate().
  6. ATMS sends a ClientTransaction to launch and resume the activity: onCreate/onStart/onResume.
  7. ViewRootImpl schedules a traversal; Choreographer at the next vsync measures, lays out and draws; RenderThread submits the buffer; SurfaceFlinger composes it. That first frame marks TTID.
Why does Zygote use a Unix socket instead of Binder?

fork() copies only the calling thread. If Zygote had a Binder thread pool, the child could inherit locks or driver state held by threads that no longer exist, leading to deadlocks or corruption. So Zygote stays single-threaded before forking and uses a simple local socket (/dev/socket/zygote) protected by SELinux so only system_server can request forks. The child sets up its own Binder state after forking.

What is IApplicationThread and why is it needed?

It is a Binder interface implemented inside each app (ActivityThread.ApplicationThread) and handed to AMS in attachApplication(). It is the reverse channel that lets the system drive the app: bindApplication, scheduleTransaction (activity lifecycle), service create/bind, broadcast delivery, scheduleTrimMemory. Calls arrive on Binder threads and are forwarded to the main thread via the ActivityThread.H Handler. Most are oneway so a slow app cannot block system_server.

How does system_server start its services?

SystemServer.main() → run() prepares the main Looper, loads libandroid_servers, creates the system context and SystemServiceManager, then calls startBootstrapServices() (Installer, AMS/ATMS, PowerManager, DisplayManager, PMS...), startCoreServices() (Battery, UsageStats, WebViewUpdate...), startOtherServices() (WMS, InputManager, Connectivity, Notification...) and startApexServices(). It advances boot phases, and finally AMS.systemReady() starts persistent apps and Home. Each service is a SystemService subclass that publishes its Binder via publishBinderService().

When is BOOT_COMPLETED sent versus LOCKED_BOOT_COMPLETED?

On file-based-encryption devices they are different events. After the system user finishes booting, sys.boot_completed=1 is set and ACTION_LOCKED_BOOT_COMPLETED goes to Direct Boot-aware apps that can use device-encrypted (DE) storage. ACTION_BOOT_COMPLETED is sent only after the user unlocks, when credential-encrypted (CE) storage is available. On a device with no lock screen they can fire close together. Do not treat "Home is drawn" as "every app may touch CE storage".

How does AMS compute process priority and how do bindings affect it?

OomAdjuster walks each process's components: a resumed activity gives foreground (0), visible activity 100, foreground service or perceptible work 200, started service 500, and so on down to cached (900+). Client-server relationships propagate importance: if a foreground client binds to a service with BIND_AUTO_CREATE or holds a content provider, the server is raised near the client's level (flags like BIND_NOT_FOREGROUND or BIND_WAIVE_PRIORITY limit this). The result is written to oom_score_adj, and a separate process state decides cgroups and restrictions.

Why is PSI better than free-memory thresholds for killing apps?

Free memory is a poor indicator: Linux deliberately uses spare RAM for page cache, so "low free memory" is normal and harmless, while real trouble is when tasks spend time stalled on reclaim and refaults. PSI directly measures that stall time (some and full percentages over 10/60/300 s windows) and lets lmkd register triggers in the kernel. This reduces both unnecessary kills and late kills that cause jank or thrashing.

How does MessageQueue block without spinning the CPU?

MessageQueue.next() calls nativePollOnce(), which runs epoll_wait() on an eventfd with a timeout equal to the time until the next message is due (or infinite). The thread sleeps in the kernel using no CPU. When a new message becomes the head of the queue, enqueueMessage() calls nativeWake(), writing to the eventfd and waking the thread. This is why the main thread's infinite loop does not burn battery.

What is a sync barrier in MessageQueue?

A barrier is a special Message with no target placed into the queue with postSyncBarrier(). While it is at the head, next() skips synchronous messages and only returns messages marked asynchronous. ViewRootImpl.scheduleTraversals() posts a barrier and Choreographer's vsync callbacks are asynchronous, so drawing the next frame is not delayed by other queued work. The barrier is removed after the traversal; forgetting to remove one freezes the thread.

What is Choreographer and how does it relate to vsync?

Choreographer receives vsync signals from SurfaceFlinger through a DisplayEventReceiver file descriptor watched by the Looper. On each vsync it runs callbacks in order: input, animation, insets animation, traversal (measure/layout/draw), commit. Work is aligned to the display refresh (16.6 ms at 60 Hz). If the main thread is busy when vsync arrives, the frame is late and "Skipped N frames" may be logged.

How do Binder threads relate to the main thread?

They are separate. Incoming Binder calls execute on threads from the process's Binder thread pool (up to 15 extra by default; system_server uses 31). The main Looper never receives Binder calls directly. So Binder method implementations must be thread-safe, and if they need UI or main-thread state they post to a Handler. Conversely, an app making a synchronous Binder call from the main thread blocks the main thread until the remote side returns.

How are input ANRs detected?

The native InputDispatcher in system_server sends each event to the focused window over an InputChannel (a socket pair) and waits for the app to send back a "finished" signal after handling it. If the oldest unacknowledged event is older than the dispatching timeout (5 s by default), or there is a focused app but no focused window for 5 s, it declares an ANR and notifies AMS to collect stacks. So an input ANR requires a pending event; a blocked main thread with no input will not produce an input ANR.

How are broadcast and service ANRs detected?

AMS arms a timeout message when it dispatches the work. For broadcasts, BroadcastQueue starts a timer when delivering to a receiver; the app reports completion with finishReceiver(). For services, ActiveServices sets a timeout when it calls scheduleCreateService/scheduleServiceArgs/bind; the app reports serviceDoneExecuting(). If completion does not arrive in time, AMS triggers the ANR path.

What does the system do when an ANR fires?

It logs the reason and CPU usage, sends SIGQUIT to the app (and to system_server and some other relevant processes) so ART's signal-catcher thread dumps all Java stacks, collects native stacks through debuggerd for some processes, writes a file in /data/anr/, adds a DropBox entry and an am_anr event, and then shows the "isn't responding" dialog for foreground apps or kills background ones.

What is the system_server watchdog and how does it differ from an ANR?

The Watchdog is a thread in system_server that every 30 s posts a check to key threads (main, android.fg, android.ui, android.io, display, animation) and runs monitors that try to take core service locks (AMS, WMS, PowerManager...). If any check is stuck for 60 s it dumps stacks and kills system_server, causing a soft reboot. ANRs are about apps missing deadlines and result in a dialog or app kill; the watchdog is about the system itself hanging.

What happens when you install an APK?

The installer opens a PackageInstaller session and writes the APK to staging. PMS parses the manifest, verifies the signature (and that updates match the existing certificate), checks SDK and ABI compatibility, assigns or reuses the app ID, has installd create data directories and extract native libraries, triggers dexopt, registers components and permissions, writes packages.xml, and broadcasts PACKAGE_ADDED or PACKAGE_REPLACED.

How does a system service check the caller's permission?

Inside the Binder method it uses the caller identity delivered by the Binder driver: Binder.getCallingUid()/getCallingPid(), and helpers like Context.enforceCallingPermission() or checkCallingOrSelfPermission(). If the service then needs to act with its own identity (for example to call another service), it wraps that code in Binder.clearCallingIdentity() and restoreCallingIdentity() in a finally block, so downstream checks see system_server rather than the app.

How is an app's UID computed with multiple users?

Each package gets an app ID (10000 to 19999 for normal apps). The actual Linux UID is userId * 100000 + appId. For user 0 it is 10xxx; the same app in a work profile (user 10) is 10010xxx. This gives each user's copy of the app separate files and process identity while the package is installed once.

What are hwbinder and vndbinder?

They are separate Binder device nodes created by Treble. /dev/hwbinder with hwservicemanager carried HIDL HAL calls between framework and vendor. /dev/vndbinder with vndservicemanager lets vendor processes talk to each other using AIDL without touching the framework's /dev/binder namespace. Stable AIDL HALs use /dev/binder itself. Separate domains keep system and vendor namespaces and SELinux policy cleanly separated.

What is VINTF and when is it checked?

VINTF (vendor interface object) is a set of XML manifests and compatibility matrices. The device manifest says which HALs and versions the vendor provides; the framework compatibility matrix says what the framework needs; the reverse pair covers what the system provides to vendor. VintfObject checks compatibility at build time, before applying an OTA, and at boot. servicemanager also refuses to register VINTF-stable HAL services that are not declared.

What is a lazy HAL?

A HAL service that is not started at boot but on demand: when a client looks it up, servicemanager asks init (ctl.interface_start) to start the service whose init.rc entry declares that interface. It can exit when it has no clients, saving memory and power. If the manifest declares it but the init.rc entry is missing or misnamed, the start fails and clients wait forever or get null.

How do ContentProviders affect app startup?

All providers declared by the app are instantiated and their onCreate() runs on the main thread during bindApplication, before Application.onCreate(). Many libraries use a provider to auto-initialize themselves, so each adds startup time even if the app never uses it. Fixes: merge initializers with the App Startup library, remove unnecessary providers, keep onCreate() trivial.

What does goAsync() do in a BroadcastReceiver?

It returns a PendingResult and tells the system the receiver is not finished when onReceive() returns. You can then do short work on a background thread and call finish() when done. The broadcast timeout still applies (10 s foreground), and the process keeps elevated priority until finish(). For longer work, schedule WorkManager or a job.

What is the cached apps freezer?

AMS can freeze processes that have been cached for a short while using the cgroup v2 freezer, so they get no CPU until they are needed again or killed. The feature landed in Android 11 as opt-in and became default-on later (around 12L/13). A synchronous Binder call to a frozen process fails immediately (BR_FROZEN_REPLY); it does not unfreeze the target. Oneway (async) transactions are queued in the driver until AMS unfreezes the process because it became important again. The system avoids freezing processes that are in the middle of a Binder transaction or that hold certain resources.

Which dumpsys commands would you use, and for what?
  • dumpsys activity processes/lru: oom_adj, process state, why something was killed.
  • dumpsys activity activities: task stack, resumed activity.
  • dumpsys input: focused window, dispatcher queues, recent ANRs.
  • dumpsys window, SurfaceFlinger, gfxinfo: windows, layers and jank.
  • dumpsys package <pkg>: install info, permissions, components.
  • dumpsys power, batterystats, alarm, jobscheduler: wakelocks and drain.
  • dumpsys meminfo: memory footprint.
What is the role of installd?

installd is a privileged native daemon that performs filesystem operations PMS is not allowed to do directly: creating and deleting app data directories with the right owner and SELinux labels, moving code, computing sizes, and (historically) running dex2oat. PMS talks to it through the Installer service over Binder. Separating it keeps system_server's own privileges smaller.

What is task affinity?

android:taskAffinity is a string (default: the package name) that names the task an activity prefers. Combined with FLAG_ACTIVITY_NEW_TASK, singleTask or allowTaskReparenting, ATMS can put the activity in (or start) a task with a matching affinity. Different affinities in one app produce multiple Recents cards. Inspect the real stacks with dumpsys activity activities.

What is a window token, and what is BadTokenException?

A window token is an IBinder WMS already knows about — usually the activity token created when ATMS starts the activity, or a special token for system overlays. WindowManager.addView (Dialog, PopupWindow, custom overlay) must carry a valid token. BadTokenException means you passed null, an Application context, or an Activity that has finished or is not yet attached. Show dialogs from a resumed Activity and dismiss them in onDestroy.

Walk the background-restriction timeline from Android 8 to 14.
  • 8: most implicit manifest broadcasts gone; background startService() throws; use jobs or startForegroundService().
  • 9–10: App Standby buckets, background location limits, background activity starts restricted.
  • 12: starting an FGS from the background is blocked except for a short allowlist; PendingIntents must declare FLAG_IMMUTABLE or FLAG_MUTABLE.
  • 13: runtime POST_NOTIFICATIONS; without it the FGS notification may be hidden.
  • 14: every FGS must declare and use a foregroundServiceType; several types need extra permissions or have their own time limits.
Why does the main-thread Looper.loop() not cause an ANR?

The infinite for (;;) is not a busy-wait. When the queue is empty the thread sleeps in epoll_wait on an eventfd and uses no CPU. An ANR is raised only when the system handed the app a specific piece of work (input, broadcast, service start) and that work was not finished before a deadline. The idle loop is waiting to do that work. What causes an ANR is a single message that runs so long that the deadline callback never runs in time.

What is ApplicationExitInfo?

From Android 11, AMS keeps a ring buffer of why each package's processes died, readable via ActivityManager.getHistoricalProcessExitReasons() or dumpsys activity exit-info. Reasons include self-exit, signaled (including LMK SIGKILL), low memory, Java/native crash, ANR, and dependency died. Records can attach an ANR trace or tombstone. It is the first tool for "why did my process die?" after the fact.

Advanced

What happens when Zygote specializes a forked child into an app?

After fork(), SpecializeCommon in native code: sets the supplementary GIDs, resource limits, and UID/GID (dropping root); mounts the app's storage view in its own mount namespace; sets capabilities to none; applies the seccomp filter; sets the SELinux context based on seinfo (for example untrusted_app); joins the right cgroups and sets the nice value; sets the process name. Then Java code closes the Zygote socket, and RuntimeInit starts the Binder thread pool and invokes ActivityThread.main(). The order matters: dropping privileges before running any app code.

Why did Google split ATMS out of AMS?

AMS had grown into a huge class guarded by one global lock, and activity/task management was tightly entangled with window management. In Android 10 activity and task logic moved to ActivityTaskManagerService in the com.android.server.wm package, sharing the WMS global lock, while AMS kept processes, services, broadcasts and providers. This reduced contention on the AMS lock and made the window/activity hierarchy (task, activity records, window containers) one coherent model.

Explain lock contention in system_server and why it causes device-wide jank.

Core services protect state with big locks (the AMS lock, the WMS global lock, the PowerManager lock). Many Binder threads, from many apps, call into these services concurrently. If one thread holds a lock while doing something slow (I/O, an outgoing Binder call to an app or HAL, heavy computation), every other thread needing that lock waits. The waits propagate to apps making synchronous calls, including their main threads, causing jank and ANRs everywhere, and in the worst case the watchdog fires. Perfetto's lock contention slices ("monitor contention with owner ...") reveal the owner and the blocked threads.

How can a deadlock across processes happen with Binder, and how do you avoid it?

Example: thread A in process X holds lock L and makes a synchronous call into process Y; Y's handler calls back synchronously into X, and that callback runs on an X Binder thread that needs lock L. Neither side can progress. Binder does handle recursive calls on the same thread (the callback is routed to the waiting thread itself), but not callbacks that land on a different thread needing the same lock. Avoid it by never holding locks across outgoing Binder calls, using oneway callbacks, and posting callbacks to a Handler instead of processing them under the lock.

What are the watchdog's HandlerCheckers and Monitors in detail?

A HandlerChecker wraps a Handler for a critical thread. Each round the watchdog posts a runnable at the front of that thread's queue; if the runnable has not run by the next check, the thread is considered blocked. The foreground thread's checker additionally calls monitor() on each registered Monitor; each monitor simply does synchronized (mLock) {}, so a held lock blocks the check. The watchdog uses a 60 s timeout with a 30 s check interval, dumps at half-time, and kills at full time. Some threads can be temporarily paused from checking (pauseWatchingCurrentThread) for known long operations.

How does input flow from the kernel to an app view?

The touch controller raises an IRQ; the kernel input driver reports events to /dev/input/eventX (evdev). In system_server, EventHub reads them, InputReader converts raw events into motion/key events, and InputDispatcher finds the target window (using window info from WMS/SurfaceFlinger) and writes the event to that window's InputChannel socket. The app's main Looper wakes on the socket FD, ViewRootImpl's input stages process it, and View.dispatchTouchEvent() reaches your onClick. The app then sends a finished signal back to the dispatcher.

Explain the Handler memory leak and how to fix it.

A non-static inner or anonymous Handler (or Runnable) holds an implicit reference to the enclosing Activity. A delayed Message references its Handler through msg.target, and the MessageQueue holds the Message, so the chain Looper → queue → Message → Handler → Activity keeps the Activity alive after it is destroyed. Fixes: make the Handler static with a WeakReference to the Activity, call handler.removeCallbacksAndMessages(null) in onDestroy(), or use lifecycle-aware coroutines.

How does ART dump stacks on SIGQUIT, and why is that safe?

ART blocks SIGQUIT in all threads except a dedicated "Signal Catcher" thread that waits for it. On receipt, the catcher suspends all threads at safe points (a checkpoint/suspend-all), walks each thread's stack, prints thread state, held and awaited monitors, and native frames, then resumes them. Because dumping happens at safe points in a dedicated thread, the app is not killed by the signal. Threads in native code (for example in a Binder ioctl) are shown as Native with their native backtrace.

What is the difference between oom_score_adj and process state?

oom_score_adj is the kernel-visible kill priority used by lmkd. Process state (ActivityManager.PROCESS_STATE_*, like TOP, BOUND_FOREGROUND_SERVICE, CACHED_EMPTY) is a framework-level classification used for other policies: which cgroup/scheduling group (top-app, foreground, background), whether background restrictions and network blocking apply, whether an app is idle for App Standby, and whether it can be frozen. They are computed together in OomAdjuster but serve different consumers.

How do stable AIDL HAL versions work, and how is an interface evolved?

A @VintfStability AIDL interface is developed as "current" and then frozen: the build stores an API snapshot in aidl_api/<name>/<version>/ with a hash. A frozen version can never change. New versions may only append methods, parcelable fields and enum values. The framework calls getInterfaceVersion() to know which methods a given vendor HAL supports, and a newer client talking to an older server receives UNKNOWN_TRANSACTION for methods that do not exist. See Binder IPC & AIDL for details.

Passthrough vs binderized HALs in HIDL?

Binderized HALs run in their own vendor process and are reached over hwbinder, giving full isolation. Passthrough HALs wrapped legacy libhardware implementations so they could be loaded in-process (getService() with passthrough mode dlopened the -impl.so); this was a migration aid and was only allowed for specific HALs such as graphics mapper. With AIDL, same-process use is possible but HALs are normally separate processes.

What is a GSI and how does it prove Treble compliance?

A Generic System Image is a pure AOSP system partition build. If a device's vendor image, kernel and boot images can boot a GSI and pass VTS (Vendor Test Suite) and CTS-on-GSI, then its vendor side depends only on stable interfaces (HALs, VNDK, sysprops) and not on private framework details. This is how Google enforces that system and vendor can be updated independently.

How does the framework prevent one app from flooding system_server with Binder calls?

Several mechanisms: per-process limits on Binder proxies (the system kills apps that hold too many, e.g. more than several thousand proxies), per-UID tracking via BinderCallsStats, oneway call handling that queues per target node so one sender's async flood mainly delays itself, rate limits in specific services (broadcast registration limits, toast limits, listener count limits), and SELinux restricting which services an app domain can find at all.

How does PMS scan affect boot time, and what optimizations exist?

On each boot PMS must discover all packages on read-only partitions and /data/app. Parsing thousands of manifests is expensive, so it caches parsed results (package cache in /data/system/package_cache), parses in parallel, and skips full rescans when the fingerprint is unchanged. After an OTA (fingerprint change) it rescans and may dexopt system apps, which is why the first boot after an update is slower.

What is the privapp permission allowlist and why does it break bring-up?

Privileged permissions are only granted to apps in priv-app directories that are explicitly listed in /system/etc/permissions/privapp-permissions-*.xml (or the equivalent on product/vendor/system_ext). When ro.control_privapp_permissions=enforce, a priv-app requesting an unlisted privileged permission causes PMS to fail boot. On a new device build or after adding a system app, a missing allowlist entry shows up as a boot loop with a clear log message.

How does WMS interact with SurfaceFlinger at a low level?

Each window has a SurfaceControl (a layer in SurfaceFlinger). WMS builds SurfaceControl.Transactions that set position, size, crop, z-order, alpha, visibility and parent relationships, and applies them atomically; SurfaceFlinger latches them on vsync. Apps get their own child SurfaceControl and use BLASTBufferQueue to submit buffers, which SurfaceFlinger composes with other layers, using HWC overlays where possible and GPU composition otherwise.

What are some reasons a framework process can be killed that are not LMKD?
  • AMS kills: excessive CPU in background, too many cached processes (max cached limit), package update or force-stop, ANR in background, crash of a provider it depends on.
  • Kernel OOM killer if lmkd could not act in time.
  • Watchdog killing system_server (which kills everything).
  • Binder proxy limit exceeded.
  • Phantom-process limits for child processes (Android 12+).

The events log am_kill entries include the reason string, which is the first thing to check.

How are persistent processes handled?

Apps with android:persistent="true" that are system apps (for example SystemUI and the phone process) are started by AMS at systemReady, given very high priority (persistent adj around -800), and restarted automatically if they die. A crash loop of a persistent process is a visible stability problem, and repeated crashes early in boot can trigger rescue party mitigation.

Why are most IApplicationThread calls oneway?

If system_server made synchronous calls into apps, a single misbehaving or frozen app could block a system_server Binder thread, possibly while holding a service lock, and stall the entire system. Oneway calls return immediately after the driver queues the transaction, and deadlines (ANR timers) are enforced separately by waiting for the app's completion callback. This asymmetry, apps call the system synchronously but the system calls apps asynchronously, is a core robustness rule of the framework. A sync call to a frozen app would also fail with BR_FROZEN_REPLY rather than unfreeze it.

How do you add a new system service in AOSP?
  1. Define IFooService.aidl and generate Stub/Proxy.
  2. Implement a SystemService whose onStart() calls publishBinderService("foo", binder).
  3. Start it from SystemServer in the right wave; react to boot phases if needed.
  4. Add a FooManager and register it in SystemServiceRegistry so apps use getSystemService.
  5. Add the name to service_contexts and SELinux add/find/binder_call rules.
  6. Enforce permissions on every Binder method and wrap privileged work in clearCallingIdentity.
  7. Implement dump() for dumpsys foo.

This is a framework service. A vendor HAL is a different path (VINTF, init.rc interface, NDK registration).

How does Android keep background apps from draining the battery at the framework level?

Doze (deep and light) defers jobs, alarms, syncs and network access when the device is idle; App Standby buckets (active, working set, frequent, rare, restricted) limit how often each app can run jobs and alarms; background service limits (Android 8) and foreground service type rules restrict long-running work; the cached apps freezer stops cached processes from running; and PowerManagerService and batterystats attribute wakelocks. Platform engineers verify these with dumpsys deviceidle, dumpsys power and Battery Historian.

Scenario & debugging

You get an ANR report. How do you debug it end to end?
  1. Read the ANR reason: input, broadcast (which action), service (which component) or provider.
  2. Open the trace from the bugreport or /data/anr/ and go to the "main" thread.
  3. Classify the state: Blocked (find the lock owner thread), Native in BinderProxy.transact (find the remote service thread), Runnable (heavy compute), Waiting (on a future or condition).
  4. Follow the chain to the root cause, including system_server stacks if the Binder call went there.
  5. Check system context: CPU usage at ANR time, lmkd kills, iowait, thermal state.
  6. Fix the design issue (off-main-thread I/O, shorter lock scope, async Binder) and add StrictMode checks and ANR-rate gates.
An ANR trace shows the main thread in BinderProxy.transactNative. What next?

The app is waiting for a synchronous Binder reply. Identify the interface from the Java frames above (for example IPackageManager$Stub$Proxy.getPackageInfo) to know the target service. Then look at the target process's stacks in the same dump (system_server is usually included) for a Binder thread handling that call, and see what it is blocked on, often a service lock or an outgoing HAL call. Binder state in /sys/kernel/debug/binder/transactions (or Perfetto's binder tracks) connects client and server threads. The fix may be in the service (lock contention) or in the app (don't make that call on the main thread).

The device soft-rebooted and the log shows "WATCHDOG KILLING SYSTEM PROCESS". How do you triage?

Read the message to see which thread or monitor was blocked ("Blocked in monitor ... on foreground thread" or "Blocked in handler on ui thread"). Open the watchdog stack dump in /data/anr or DropBox (system_server_watchdog). Find the blocked thread's "waiting to lock ... held by thread N", then read thread N: often a Binder thread stuck in an outgoing call to a HAL or app, a slow file operation under a lock, or a lock-order deadlock between two service locks. Check whether the HAL process was hung (its own stacks via debuggerd). Fix by removing the slow call from the locked region or fixing lock ordering.

An app's cold start regressed from 600 ms to 1100 ms in a new build. How do you find the cause?

Reproduce with am force-stop + am start -W over several runs on both builds to confirm. Capture Perfetto app-startup traces on both and compare the main thread: bindApplication (providers, Application.onCreate), activityStart, inflate and first doFrame. Look for new slices, Binder waits, lock contention, class-loading/JIT (missing baseline profile or dexopt state, check dumpsys package dexopt status) and system-level differences (CPU frequency, thermal, background load). Bisect the change once the stage is known.

A background music app keeps getting killed. How do you investigate?

Check logcat -b events for am_kill reasons and lowmemorykiller logs. Check dumpsys activity processes for its adj and process state while playing: if it is not a foreground service with the media playback type, it sits at service or cached priority and is a prime lmkd victim. Also check whether the app is restricted by battery optimization/App Standby, or the OEM has aggressive background policies. The fix is usually a proper foreground service with notification and correct type, and handling restarts gracefully.

Users report jank when scrolling. What is your approach?

Measure with dumpsys gfxinfo <pkg> framestats for janky frame percentage, then capture Perfetto around the scroll. Look at the main thread and RenderThread per frame against the vsync deadline: long doFrame (layout, bind, inflate), main-thread I/O or Binder, GC pauses, RenderThread GPU waits, or SurfaceFlinger missing deadlines (dumpsys SurfaceFlinger, GPU vs HWC composition). Also check CPU frequency and thermal throttling. Fix by moving work off the main thread, avoiding overdraw, prefetching, and reducing layout depth.

After adding a new privileged system app, the device boot-loops. What do you check?

Get logcat from the boot (or adb logcat -b all during the loop) and look for PMS errors, especially "Privileged permission ... for package ... not in privapp-permissions allowlist" with enforcement on. Also check for SELinux denials for the new app's domain, a crash of the app if it is persistent, and signature or sharedUserId mismatches. Fix by adding the allowlist XML to the correct partition, the right seinfo/SELinux rules, and verifying signing.

On a new build the phone shows "no service" and logs show the radio HAL proxy is null. What happened?

Before blaming the modem, check the HAL registration path. Look for init: Could not find '...IRadio...' for ctl.interface_start or servicemanager "Could not find ... in the VINTF manifest" messages. Common causes: VINTF manifest declares the interface but no init.rc service has a matching interface aidl line (lazy start fails), a name/instance mismatch (slot1 vs default), SELinux denying registration (avc: denied { add }), or the HAL binary crashing on start (check tombstones). Fix the configuration; the modem was fine.

The device never enters suspend overnight and battery drains fast. How do you debug from the framework side?

Start with dumpsys power for currently held wakelocks and dumpsys batterystats (or Battery Historian on a bugreport) to see which UIDs held partial wakelocks, scheduled frequent alarms or jobs, and when the CPU was awake. Check dumpsys alarm and dumpsys jobscheduler for frequent wakeups, and /sys/kernel/debug/wakeup_sources for kernel wakelocks (drivers, modem). Check Doze state with dumpsys deviceidle. Fix the offending app/service or driver, and add a standby-drain regression test.

A BroadcastReceiver ANR happens only on boot. Why and how do you fix it?

At boot many apps receive BOOT_COMPLETED and LOCKED_BOOT_COMPLETED at the same time as the system is busiest, so CPU and I/O contention is high. A receiver doing even moderate work (disk, network, database migration) in onReceive can exceed the timeout, and the background queue ANR (60 s) is possible if the system is starved. Fix by doing minimal work in onReceive, scheduling WorkManager jobs with constraints, and checking system boot-time load in Perfetto.

An app crashes with ForegroundServiceDidNotStartInTimeException. Explain.

The app called startForegroundService(), which promises the system that the service will call startForeground() within a short window (a few seconds; public docs have said 5 s or 10 s depending on the release). It did not, perhaps because onCreate/onStartCommand did heavy work first, the main thread was blocked, or it returned early on an error path. The system raises an ANR and crashes the app. Fix by calling startForeground() at the very start of onStartCommand on every path, with the correct service type.

system_server memory keeps growing over days. How do you investigate?

Track dumpsys meminfo system_server over time to see if Java heap, native heap, or Binder-related counts are growing. Check dumpsys meminfo object counts (Binders, proxies, death recipients) for leaks from registered listeners that are never unregistered. Take a heap dump (am dumpheap system_server) and analyse dominators. Common causes: callback lists (RemoteCallbackList misuse), unbounded caches, leaked death recipients from apps that crash and re-register, and native leaks in HAL clients.

An app is fast on a flagship but ANRs frequently on a low-RAM device. What is likely?

On low-RAM devices lmkd and zRAM are much more active: page faults and refaults cost time, kswapd competes for CPU, and the app's own pages may be reclaimed, so the main thread spends time stalled on memory (visible in PSI and Perfetto as uninterruptible sleep / D state). The same code that takes 200 ms on a flagship can take several seconds. Check /proc/pressure/memory, lmkd kill logs and Perfetto thread states; reduce memory footprint and move work off the main thread.

How would you prove whether an ANR is the app's fault or the system's?

If the app's main thread is doing its own work (compute, I/O, waiting on its own lock), it is the app. If it is waiting on a Binder call and the server thread in system_server is blocked on a service lock or a HAL, it is a system problem. If the main thread is idle or runnable but not scheduled, and CPU usage or memory pressure is high system-wide, it is starvation. Perfetto thread states (Running, Runnable, Sleeping, Uninterruptible) over the ANR window make this clear.

A service implementation hangs callers intermittently. You suspect Binder thread pool exhaustion. How do you confirm?

Take stack dumps of the service process during the hang (kill -3 or debuggerd). If all Binder threads (Binder:pid_N) are busy in the same slow path, for example waiting on a lock, a HAL call or network, the pool is exhausted and new calls queue in the driver. /sys/kernel/debug/binder/proc/<pid> shows threads and pending transactions. Fix by making handlers fast, moving long work to a worker with async callbacks, using oneway where no reply is needed, and removing nested blocking calls. See Binder debugging.

An activity loses user input after returning from the background. How do you debug?

The process was likely killed while in the background (cached) and the activity was recreated from saved state, but the app did not save or restore that input. Confirm with am_kill/am_proc_died events and by reproducing with "Don't keep activities" in developer options or am kill <pkg> while backgrounded. Fix by storing UI state via onSaveInstanceState/SavedStateHandle and persistent data in storage.

After an OTA, first boot takes several minutes longer. What is happening, and what can be tuned?

A fingerprint change makes PMS rescan all packages without the cache, may re-verify, and triggers dexopt of system and possibly user apps, plus other first-boot migrations. Options: ship pre-compiled odex/vdex for system apps in the image, use cloud/baseline profiles, run dexopt in the background after boot (pm.dexopt.boot filters), and reduce the number of preinstalled apps. On A/B devices, much of the compilation can run during the OTA itself (otapreopt) before reboot.