Wearables & Platform Integration

Power, Thermal & Battery

Android power saving rests on one idea: keep the application processor suspended as much as possible and wake it only for real work. This page covers the full stack, from kernel suspend and wakelocks through DVFS, EAS and Power HAL / ADPF hints, then Doze, job scheduling, sensor offload, thermal throttling and battery charging, plus the tools to debug drain and the KPIs a platform lead gates a release on.

~72 min read 0 interview questions
In 30 seconds
  • Battery life is roughly capacity divided by average current, and average current is dominated by how much time the AP spends suspended.
  • Linux suspend-to-RAM is the deepest normal state; kernel wakeup sources (2.6.37) and Android wakelocks block it, and wake interrupts end it.
  • While the AP is awake, cpufreq / DVFS, cpuidle, Energy Aware Scheduling and Power HAL / ADPF hints decide how expensive that awake time is.
  • Doze and App Standby buckets defer apps' network, jobs and alarms into batched windows; WorkManager and JobScheduler are the cooperative way to run background work.
  • Always-on sensing belongs on the low-power sensor hub, with batched delivery, so the AP can sleep for minutes.
  • Thermal management is driven by skin temperature on handheld and wearable devices: thermal zones, the thermal HAL and a thermal engine throttle CPU, GPU, charging and modem.
  • Debug with batterystats and Battery Historian, Perfetto (including power rails), dumpsys power, deviceidle, thermalservice and wakeup_sources. Gate every release on power, stability and performance KPIs. See also Wear OS Platform.

The power stack at a glance

A device is always in one of a few power states, and the difference in current between them is enormous. Almost all battery engineering is about spending more time in the low states and less in the high ones. The stack that decides this spans hardware (PMIC, clocks, co-processor), the kernel (cpuidle, cpufreq, suspend, wakeup sources, thermal), the HALs (power, thermal, health, power stats) and the framework (PowerManagerService, DeviceIdleController, JobScheduler, AlarmManager, ThermalManagerService, BatteryStatsService).

Analogy

Think of a house electricity bill. The fridge humming all night costs more over a month than a kettle boiling for three minutes, because it never stops. In the real system the "fridge" is anything that keeps the processor awake or a radio active continuously (a leaked wakelock, a chatty sensor, a searching modem), while the "kettle" is a short, intense burst of real work. Long, low drains usually matter more than brief peaks.

StateWhat is happeningRough current (watch)
Active, screen onCPU running, display on, maybe radios activeTens of mA or more
Awake, screen off (wakelock held)CPU can idle between tasks but the system cannot suspendSeveral mA: the silent killer
CPU idle (cpuidle)Cores in C-states between tasks; system not suspendedLower, but clocks, DDR and rails still on
System suspendedProcesses frozen, devices suspended, non-boot CPUs off, DDR in self-refreshAround 1 mA or less on a watch
Co-processor onlyAP suspended; sensor hub runs sensing and AOD helpVery low: the always-on work lives here
Framework   PowerManagerService  DeviceIdleController  JobScheduler  AlarmManager  ThermalManagerService  BatteryStats
               |                    |                      |             |               |                     |
HALs        SystemSuspend (suspend HAL)   Power HAL (hints)   Thermal HAL   Health HAL   Power Stats HAL
               |                                              |             |              |
Kernel      suspend / autosleep / wakeup sources   cpuidle   cpufreq   thermal zones   power_supply   rail monitors
               |                                                                              |
Hardware    PMIC, clocks, regulators, CPUs, DDR, always-on co-processor, fuel gauge, charger, thermistors
Interview angle Interviewers often start with "how does Android save power?" A strong answer names the states, says that battery life is roughly time spent suspended, covers what happens while awake (DVFS, EAS, Power HAL / ADPF), then walks wakelocks and Doze to job scheduling, and says how you would measure it.

Suspend and resume: the sleep machine

Android power saving is built on Linux suspend-to-RAM. When nothing needs the CPU, the whole system suspends: tasks are frozen, devices are put to sleep through their driver callbacks, non-boot CPUs go offline, clocks and regulators are gated, and RAM stays in self-refresh. Something must actively block suspend (a wakeup source or wakelock), and something must trigger resume (a wake-capable interrupt).

Analogy

Suspend is like a shop closing for the night: staff stop taking new jobs (freeze tasks), each department shuts its equipment in order (driver suspend callbacks), the lights go out except the alarm system (DDR self-refresh and wake interrupts), and anyone still working holds the door open with a "do not lock" sign (a wakeup source). In the real system the kernel refuses to suspend while any wakeup source is active, and a configured wake interrupt (RTC alarm, button, sensor, modem page) is the doorbell that reopens the shop.

The suspend sequence

  1. Decide to suspend On Android the userspace SystemSuspend service (the suspend HAL) waits until no wakelocks are held, reads /sys/power/wakeup_count, writes it back, then writes mem to /sys/power/state.
  2. Freeze The kernel freezes userspace tasks and freezable kernel threads.
  3. Device suspend Drivers run their prepare, suspend, suspend_late and suspend_noirq callbacks. Any driver returning an error aborts the whole suspend.
  4. CPUs down Non-boot CPUs are taken offline, syscore callbacks run, and the platform enters its deepest state (on Arm through PSCI firmware; on Qualcomm SoCs rails and the crystal oscillator can be shut down).
  5. Wake A wake-capable interrupt fires. Resume runs the callbacks in reverse order, thaws tasks, and the kernel records the wakeup reason.
no wakelocks held
   -> SystemSuspend: read wakeup_count, write it back (fails if a wakeup event happened in between)
   -> echo mem > /sys/power/state
   -> freeze tasks -> suspend devices -> offline CPUs -> platform low-power state
                                                         |
                          wake IRQ (RTC, button, sensor, modem, BT) 
                                                         v
   <- thaw tasks  <- resume devices  <- online CPUs  <- platform resume
   -> wakeup reason logged; the event usually takes a wakeup source briefly so the work can run

Autosleep, suspend blockers and wakeup sources

  • Suspend blockers were Android's original (2009) kernel mechanism: "opportunistic suspend" whenever no blocker was held. They were controversial upstream and never merged as-is.
  • Wakeup sources are the mainline Linux object that replaced the idea. They were introduced in kernel 2.6.37 (2010). A driver creates one and calls __pm_stay_awake / __pm_relax, or pm_wakeup_event with a timeout, to block suspend while it handles an event.
  • Autosleep and userspace wakelocks landed later, in kernel 3.5 (2012). /sys/power/autosleep lets the kernel itself suspend whenever no wakeup source is active; /sys/power/wake_lock and wake_unlock expose a userspace wakelock interface compatible with Android's old API.
  • Modern Android (Android 10 and later) usually does not use autosleep. The SystemSuspend service tracks userspace wakelocks itself, uses the wakeup_count handshake to avoid races, and triggers suspend. Kernel wakeup sources still block suspend from drivers.
  • Mem sleep modes: /sys/power/mem_sleep selects between s2idle (suspend-to-idle, CPUs in deepest idle) and deep (full suspend-to-RAM), depending on platform support.

cpuidle is not suspend

cpuidle (idle states)

  • Per-CPU, between tasks, microseconds to milliseconds.
  • System keeps running; timers fire; DDR and shared rails on.
  • Governors (menu, TEO) choose depth by predicted idle time.

System suspend

  • Whole system, seconds to hours.
  • Tasks frozen; only wake IRQs resume it.
  • Deepest normal state; blocked by any wakeup source.
Common pitfall Assuming "screen off" means "suspended". A screen-off device with one partial wakelock held never suspends and may draw ten times the suspend current. Always check suspend residency, not screen state.
Interview angle Expect "what happens when an Android device goes to sleep?" and "what is the difference between a wakelock and a wakeup source?" Walk the sequence, name /sys/power/state and wakeup_count, and mention that one failing driver suspend callback aborts the whole cycle. The Linux Kernel & BSP page covers the driver side.

Wakelocks: what keeps the device awake

A wakelock is a request to keep the CPU (and optionally the screen) from sleeping. In apps it is a PowerManager.WakeLock; PowerManagerService tracks them and asks SystemSuspend to hold a native wakelock while any are held. Drivers use kernel wakeup sources. Held too long or too often, both destroy standby battery.

Analogy

A wakelock is like a "meeting in progress, keep the lights on" sign on an office door. It is fine for a 30-second chat, but if someone leaves the sign up and goes home, the building's lights burn all night. In the real system a partial wakelock that is never released, or held across a network call that stalls, keeps the AP out of suspend for hours while the screen looks off.

TypeLevelKeeps onNotes
Partial wakelockPARTIAL_WAKE_LOCKCPU onlyThe dangerous one: invisible drain. Always use a timeout.
Screen wakelocksSCREEN_DIM, SCREEN_BRIGHT, FULL_WAKE_LOCKCPU plus screenDeprecated; use the keep-screen-on window flag instead.
ProximityPROXIMITY_SCREEN_OFF_WAKE_LOCKTurns screen off near the faceUsed during calls.
Kernel wakeup sourcewakeup_source_register, __pm_stay_awakeBlocks system suspendHeld by drivers (sensors, modem, Bluetooth, USB, charger). Visible in wakeup_sources.
Native userspace wakelockSystemSuspend acquireWakeLockBlocks system suspendHeld by native daemons and by PowerManagerService on behalf of apps.

Kinds of leaks

  • Never released: exception path skips release().
  • Held during stalled I/O: a network or Binder call that hangs while the lock is held.
  • Re-acquired in a loop: short locks acquired thousands of times, each waking the AP.
  • Kernel side: a driver takes a wakeup source on every interrupt of a chatty device, or never relaxes it.
  • Fake work: a foreground service kept alive to avoid process death, holding a lock for "just in case" work.
// Correct pattern: always use a timeout and release in finally
val wl = powerManager.newWakeLock(PowerManager.PARTIAL_WAKE_LOCK, "myapp:sync")
wl.acquire(10 * 60 * 1000L)   // timeout as a safety net
try {
    doShortWork()
} finally {
    if (wl.isHeld) wl.release()
}
Tip Wakelock tags should follow the app:component pattern so batterystats and Perfetto attribute drain clearly. Android vitals and platform tools flag excessive partial wakelocks, so tagging matters for triage.
Common pitfall Classic systemic bug: after an upstream merge, standby drain doubles because a HAL or service takes a partial wakelock and stops batching, so the AP never suspends. You find it in wakeup_sources or Battery Historian, assign it to the owning team, and add a wakelock-budget regression gate.
Interview angle "What is a wakelock and why is it the number one wearable power bug?" is standard. Explain partial wakelocks, kernel wakeup sources, how to find the holder (dumpsys power, wakeup_sources, batterystats), and how a platform prevents recurrence (timeouts, budgets, gates).

CPU power while awake: DVFS, cpufreq, cpuidle and EAS

Suspend decides whether the application processor is on. The next question is how expensive that awake time is. Dynamic voltage and frequency scaling (DVFS), the cpufreq and cpuidle governors, and Energy Aware Scheduling (EAS) decide which cores run, at what frequency and voltage, and how deeply idle cores sleep between tasks. A device that never suspends but stays on little cores at a low OPP can still last a day; the same workload pinned to a big core at max frequency will not.

Analogy

Think of a workshop with slow, cheap benches and a few fast, power-hungry benches. The manager (EAS) sends small jobs to the cheap benches and only fires up a fast bench when a job needs it. The dimmer switch on each bench (DVFS / cpufreq) turns brightness (frequency and voltage) up or down to match the work. When nobody is at a bench, that bench's lights go off (cpuidle) without closing the whole shop (suspend). In the real system, little cores are the cheap benches, big or prime cores are the fast ones, OPP tables are the dimmer steps, and closing the shop is system suspend.

DVFS and cpufreq

Dynamic voltage and frequency scaling changes a CPU cluster's clock and its supply voltage together. Dynamic power is roughly proportional to C·V2·f, so dropping voltage with frequency saves more than frequency alone. Each cluster has an operating-performance-point (OPP) table (device tree or firmware) of legal frequency and voltage pairs. The cpufreq subsystem picks an OPP; cooling devices and thermal policy can cap the maximum.

GovernorHow it picks frequencyWhere you see it
schedutilSets frequency from scheduler utilization (PELT) and uclamp hints. The modern Android default.Current mainline and GKI kernels.
interactiveAndroid's historic governor: jumps to a "hispeed" frequency on load or input, then ramps. Tunables such as hispeed_freq and go_hispeed_load were a common bring-up fight.Older Android kernels; replaced by schedutil.
ondemand / conservativeClassic load-based ramps; conservative is slower to go up.Rare on phones now.
performance / powersave / userspacePin to max, min, or a userspace-chosen OPP.Lab, thermal cap, or a vendor daemon.
adb shell cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor
adb shell cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_cur_freq
adb shell cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_available_frequencies

cpuidle: C-states between tasks

When a CPU has no runnable task it enters an idle state. Each state has an exit latency and a target residency: if the predicted idle time is shorter than the residency, a shallower state is cheaper because the entry and exit energy would not be paid back. Governors: menu (predicts idle duration from recent history) and teo (Timer Events Oriented, the usual modern default). States typically run from WFI (wait-for-interrupt, shallow, fast exit), through cluster clock gating, to cluster power collapse and rail collapse (deepest, longest wake). This is per-CPU idle, not system suspend: timers still fire, DDR and shared rails usually stay on. See the compare box under Suspend and resume.

Energy Aware Scheduling

EAS uses an energy model of each CPU (capacity and power at each OPP, from the device tree or firmware) plus PELT utilization to place a waking task on the CPU that can meet its needs with the least energy. Small tasks stay on little cores; heavy tasks go to big or prime cores. It works with schedutil. When the system is overutilized (not enough spare capacity), EAS steps aside and ordinary load balancing takes over so performance is not sacrificed. Android steers it with:

  • uclamp (utilization clamp): uclamp.min boosts a task (used for top-app and launches) so schedutil raises frequency; uclamp.max caps background work.
  • cpusets / cgroups: top-app may use all clusters; background and cached apps are often limited to little cores.

EAS replaced earlier big.LITTLE placement (HMP and vendor scheduler hooks) on modern Android Common Kernels. It was developed on Android kernels first and mainlined around Linux 5.0. Scheduler detail lives on the Linux Kernel & BSP page; the power takeaway is that placement plus frequency decides energy per unit of work, and a wrong boost (everything on big cores at max OPP) shows up as heat and drain, not as a functional bug.

Tip In a Perfetto trace, look at power/cpu_frequency, power/cpu_idle and the scheduler tracks together. A thread that is runnable on a little core at a mid OPP is cheap; the same thread migrated to a prime core at max frequency for a 2 ms touch boost is expensive if the boost never ends.
Interview angle Expect "how does Android save power while the CPU is awake?" Name DVFS (V2f), cpufreq governors (interactive then schedutil), cpuidle residencies, EAS plus the energy model, and uclamp / cpusets. Then say how you would prove it in Perfetto. Follow-up is usually Power HAL and ADPF, next section.

Doze, App Standby and standby buckets

Wakelocks are the mechanism; Doze and App Standby are the policy. They stop apps from waking the device whenever they like by deferring network access, jobs, syncs and alarms into shared maintenance windows. Batching many apps' work into one wakeup is far cheaper than letting each app wake the device separately.

Analogy

Doze is like an apartment building where the mail is delivered twice a day instead of each sender ringing the doorbell whenever they like. The longer residents are away, the less often the mail carrier comes. Standby buckets are like giving frequent residents a key and rare visitors a buzzer that only works at set hours. In the real system Doze defers apps' work into maintenance windows that grow further apart the longer the device is idle, and buckets limit how often each app's jobs and alarms can run based on how recently the user used it.

Light Doze and deep Doze

Light Doze (Android 7+)Deep Doze (Android 6+)
Enters whenScreen off, on battery, a short time has passed; device may be movingScreen off, on battery, and stationary for a longer period (motion checked with sensors)
RestrictionsNetwork access and jobs/syncs deferredAlso: standard alarms deferred, app wakelocks ignored, Wi-Fi scans stopped, sync adapters and jobs deferred
Maintenance windowsFrequentGrow further apart the longer the device stays idle
Exits whenScreen on or plugged inMotion, screen on, plugged in, or an alarm-clock alarm
screen off, on battery
  ACTIVE -> INACTIVE -> IDLE_PENDING -> SENSING -> LOCATING -> IDLE
                                                                |  ^
                                                   maintenance  v  |  window gaps grow over time
                                                         IDLE_MAINTENANCE
motion / screen on / charger  ->  back to ACTIVE

What still gets through Doze

  • setExactAndAllowWhileIdle() and setAndAllowWhileIdle() alarms, rate-limited per app.
  • setAlarmClock() alarms (user-visible alarms); the system leaves idle shortly before they fire.
  • High-priority FCM messages, which grant the app a short temporary exemption to do work; they should be used only for user-visible events.
  • Apps on the battery-optimization allowlist (for example, companion or accessibility apps with a strong reason).

App Standby buckets (Android 9+)

BucketMeaningEffect
ActiveIn use now or very recentlyNo restrictions
Working setUsed regularlyMild limits on jobs and alarms
FrequentUsed often but not dailyStronger job and alarm limits
RareRarely usedStrict limits on jobs, alarms and high-priority messages
Restricted (Android 11+, API 30)Barely used or misbehavingJobs about once a day in a batched window; very limited alarms and high-priority messages

Adaptive Battery can predict usage and assign buckets. Android 12 did not invent Restricted; it made unused or misbehaving apps more likely to land there, and added app hibernation (unused apps have permissions revoked and their cache cleared). Android 12 also freezes cached processes so they do not run. Android 13 added Low Power Standby: after a long unused period (screen off, not charging), the device applies even tighter idle restrictions than deep Doze. Separately, background execution limits (Android 8+) stop apps from running background services freely, Android 12+ restricts starting foreground services from the background, and Android 14+ requires a declared foreground service type. Exact alarms need the SCHEDULE_EXACT_ALARM or USE_EXACT_ALARM permission on recent releases; allow-while-idle exact alarms are rate-limited to roughly one every nine minutes while idle. Battery Saver mode adds further restrictions on top of buckets and Doze.

# Force Doze for testing
adb shell dumpsys battery unplug
adb shell dumpsys deviceidle force-idle       # jump straight to deep idle
adb shell dumpsys deviceidle step             # or step through states
adb shell dumpsys deviceidle get deep
adb shell dumpsys deviceidle unforce
adb shell dumpsys battery reset

# Standby buckets
adb shell am get-standby-bucket com.example.app
adb shell am set-standby-bucket com.example.app rare
Tip Align periodic work with Doze maintenance windows and batch it. Fighting Doze with exact wakeup alarms or high-priority pushes is the fastest way to blow the battery budget, and a very common cause of "the watch dies overnight" tickets. On wearables these policies are usually tuned even more aggressively than on phones.
Interview angle Interviewers check whether you know the difference between Doze (device-wide, based on idle state) and App Standby (per-app, based on usage), what can break through Doze, and how to test with dumpsys deviceidle. Mentioning buckets (Restricted since Android 11), hibernation, Low Power Standby and exact-alarm permissions shows you are current.

Deferred work: WorkManager, JobScheduler and AlarmManager

Apps and system components should never run their own timers to wake the device. Android offers scheduling APIs that let the system batch work, respect Doze and buckets, and pick a good moment (charging, on unmetered network, idle).

Analogy

The scheduling APIs are like the choices for sending a parcel. WorkManager is a standard courier that guarantees delivery but picks the route and time (cheap and reliable). JobScheduler is the courier's internal dispatch system. AlarmManager exact alarms are a special same-hour delivery that costs extra and should be used only when the time really matters. A foreground service is hiring a dedicated driver who waits with the engine running. In the real system each option trades timing precision against power, and the default should be the cheapest one that meets the requirement.

APIUse forPower behavior
WorkManager (Jetpack)Deferrable, guaranteed background work: sync, upload, cleanupPersists across reboot, supports constraints (network type, charging, battery not low, idle, storage), chaining, expedited work; built on JobScheduler on modern devices. Periodic work minimum is 15 minutes. The recommended default.
JobSchedulerPlatform job API underneath WorkManagerBatches jobs by constraints; Doze- and bucket-aware; quotas per bucket.
AlarmManager inexact (set, setWindow, setInexactRepeating)Approximate time triggersSystem batches them; deferred in Doze.
AlarmManager exact (setExact, setExactAndAllowWhileIdle, setAlarmClock)True time-critical events: user alarms, calendar remindersWakes the AP at that time, can punch through Doze; needs exact-alarm permission. Use rarely.
Foreground serviceUser-visible ongoing work: workout tracking, navigation, media, callsKeeps process alive with a notification; typed on Android 14+. Costs power; justify it and stop it promptly.
FCM high priorityServer events the user must see nowBrief Doze exemption; quotas apply; do not use for silent sync.

Wakeup versus non-wakeup alarms

Alarm types ending in _WAKEUP (RTC_WAKEUP, ELAPSED_REALTIME_WAKEUP) wake the device from suspend; the others fire only when the device is next awake. Prefer ELAPSED_REALTIME types for intervals, since wall-clock time can change.

Common pitfall Using Handler.postDelayed or a Java timer for long background delays. Handler time uses uptime, which stops while the device is suspended, so a "5 minute" delay may take hours. Conversely, a wakelock held to keep such a timer running wastes power. Use WorkManager or AlarmManager instead.
Tip Rule of thumb: defer, batch, constrain. Any code path that wakes the AP on a tight period, ignores constraints, or holds a wakelock through I/O is a power regression waiting to happen.
Interview angle "When would you use AlarmManager versus WorkManager versus JobScheduler?" is very common. Answer by requirement: guaranteed but flexible timing means WorkManager; exact user-visible time means an exact alarm; continuous user-visible work means a foreground service; and explain the power cost of each.

Sensor-hub and always-on offload

The single biggest wearable power lever, and an important phone one, is to run always-on work on low-power hardware so the application processor stays suspended for minutes at a time. The sensor hub (always-on co-processor) is the main example, but the same idea appears in audio, Wi-Fi, Bluetooth, GNSS and cellular data paths.

Analogy

Offload is like a building's security guard watching the cameras overnight so the manager can sleep. The guard handles routine events and calls the manager only for something that needs a decision. In the real system the sensor hub (or audio DSP, or Wi-Fi firmware) handles routine samples and packets, and raises an interrupt to wake the AP only when a batch is full or a meaningful event occurs.

What lives on the sensor hub

  • Continuous step counting, heart-rate sampling, sleep tracking.
  • Sensor fusion (accelerometer plus gyroscope), orientation, activity recognition.
  • Gesture and wrist-raise detection, significant motion, off-body detection.
  • Always-on display help, and on some designs Bluetooth keep-alive.
  • Hardware FIFO batching: buffer samples, deliver in bursts.

How data reaches apps efficiently

  • Hub batches, wakes the AP briefly, then the Sensors HAL, SensorService or Health Services, then the app.
  • Batch latency trades freshness for sleep: bigger batches mean more AP sleep.
  • Non-wake-up sensors wait for the AP to wake for another reason; wake-up sensors wake it.
  • Direct high-rate sensor access from an app bypasses this and wrecks battery.
Sensors -> [ always-on hub: fusion + FIFO batching, AP asleep ] -> batch full or event -> wake IRQ
                                                                   -> AP resumes briefly
                                                                   -> HAL -> framework -> app
                                                                   -> AP suspends again
Battery win = AP asleep time, maximized by batch size and offload coverage

Other offloads worth naming

OffloadWhat it does
Context Hub (CHRE)Android's framework for small "nanoapps" running on the sensor hub, managed through the Context Hub HAL. Lets custom always-on logic run without the AP.
Audio DSPLow-power hotword detection and offloaded audio playback, so the AP sleeps during music.
Wi-Fi packet filteringThe Android Packet Filter (APF) and similar firmware filters drop uninteresting packets in the Wi-Fi chip so they do not wake the AP.
Bluetooth controllerBLE scan filtering and batching in the controller; the AP wakes only for matching advertisements.
GNSS batchingLocation fixes stored in the GNSS chip and delivered in batches.
Cellular dataHardware packet accelerators route bulk data without CPU involvement; modem DRX and eDRX let the radio sleep between paging slots.
Interview angle Interviewers probe whether you can explain why offload saves power (the AP subsystem and DDR can be off) and what breaks it (an app or HAL forcing high-rate delivery, wrong wake-up flags, or misconfigured batch latency). See Wear OS Platform for the watch-specific sensor path and Trace a Path Through the Android Stack for the sensor-sample walk.

Thermal: skin temperature and throttling

Every watt consumed becomes heat. Handheld and wearable devices have no fans, so heat must spread through the case into the air and the user's skin. Thermal management keeps components below their limits and, most importantly for wearables and phones, keeps the surface the user touches comfortable and safe.

Analogy

Thermal management is like a thermostat-controlled oven in a small kitchen. The oven (the SoC) could go hotter, but the kitchen (the device case) and the cook standing next to it (the user's skin) set the real limit, so the controller turns the heat down before the room gets uncomfortable. In the real system temperature sensors feed thermal zones, and when trip points are crossed the thermal governors and thermal engine reduce CPU and GPU frequency, charging current or modem power.

Why it is acute on wearables

  • Worn against the skin for hours, so skin-temperature limits for prolonged contact (in the low 40s degrees Celsius, set by product and safety standards) bind before silicon junction limits.
  • Tiny thermal mass heats up quickly during workouts with GPS, LTE calls, and charging (charging itself produces heat in the battery and charger).
  • Heat and power feed each other: hotter silicon leaks more current, which produces more heat.

The thermal stack

Sensors: on-die sensors (CPU, GPU, modem), board thermistors, battery thermistor
   -> Kernel thermal framework: thermal zones + trip points + governors
        cooling devices: cpufreq, GPU devfreq, charge current, modem, backlight
   -> Fast hardware limiters: CPU frequency limits, battery current limiting on voltage droop
   -> Userspace thermal engine / vendor policy: virtual skin sensor, multi-sensor rules
   -> Thermal HAL (AIDL IThermal): temperatures, thresholds, cooling devices, callbacks
   -> ThermalManagerService -> PowerManager thermal status and headroom -> apps back off
   -> At SHUTDOWN severity the framework shuts the device down cleanly
ConceptDetail
Thermal zone/sys/class/thermal/thermal_zoneN/ with type, temp (millidegrees C), trip_point_N_temp, trip_point_N_type (passive, active, hot, critical) and policy.
Cooling device/sys/class/thermal/cooling_deviceN/ with type, cur_state, max_state; bound to zones.
Governorsstep_wise (step cooling up or down per trip), power_allocator (a controller that divides a power budget between devices), user_space (defer to a daemon), and others.
Skin temperatureUsually a virtual sensor: a weighted model of several thermistors that estimates the case surface temperature.
Thermal engineVendor userspace daemon or policy (on Qualcomm platforms historically thermal-engine) that implements multi-sensor rules and mitigation.
Thermal HALHIDL, then AIDL IThermal (Android 13+). Reports temperatures by type (CPU, GPU, battery, skin, USB port, and more) with severity: NONE, LIGHT, MODERATE, SEVERE, CRITICAL, EMERGENCY, SHUTDOWN.
App APIsPowerManager.getCurrentThermalStatus(), addThermalStatusListener() (Android 10+), getThermalHeadroom(seconds) (Android 11+) so apps can reduce frame rate or quality before throttling.
ChargingJEITA rules reduce charge current or voltage when the battery is cold or hot; thermal policy can further limit charging when the device is hot.
adb shell dumpsys thermalservice                       # HAL temperatures, status, cooling devices
adb shell 'for z in /sys/class/thermal/thermal_zone*; do echo "$(cat $z/type) $(cat $z/temp)"; done'
adb shell cat /sys/class/thermal/cooling_device*/cur_state
adb shell cmd thermalservice override-status 3         # simulate SEVERE for app testing
adb shell cmd thermalservice reset
Common pitfall Tuning thermal limits only on junction temperature. A device can pass silicon limits and still feel too hot on the wrist. Validate with case-surface measurements (thermocouples, thermal camera) in realistic conditions: on a wrist or skin model, in warm ambient, during charging plus use.
Interview angle Expect "why is thermal harder on a watch?" and "how does Android throttle?" Walk the stack from sensors to trip points to cooling devices to the thermal HAL and app APIs, and explain the performance versus skin temperature versus battery trade-off you would tune and gate.

Power HAL and ADPF performance hints

The kernel governors react to load they already see. Launching an activity, handling a touch, or rendering a game frame needs the frequency and core placement to rise before the load shows up in PELT, otherwise the first 50 to 200 ms feel slow. The Power HAL and the Android Dynamic Performance Framework (ADPF) are how the framework and apps tell the SoC that work is coming, and how they back off before thermal throttling hits.

Analogy

Governors are like a kitchen that turns the burners up after orders pile up. Hints are the waiter calling "eight mains incoming" so the burners are already hot. In the real system a touch or launch hint (Power HAL) or a per-thread target duration (ADPF hint session) raises uclamp or frequency immediately; reporting the actual work time lets the vendor policy settle on the lowest OPP that still meets the deadline.

Power HAL

The Power HAL (android.hardware.power; HIDL 1.x, then AIDL from Android 11) is implemented by the vendor and called from PowerManagerService, the activity manager, input, HWUI and others. It does not run the CPUs itself; it programs the same knobs EAS and cpufreq already expose (uclamp, schedtune, cpusets, temporary max-freq boosts, often a vendor "boost" driver).

SurfaceWhat it asks for
setBoost (AIDL) / legacy POWER_HINT_*Short bursts: INTERACTION (touch, fling), DISPLAY_UPDATE_IMMINENT, and historically LAUNCH. Typically tens to hundreds of milliseconds.
setModeLonger modes: INTERACTIVE (screen on), LOW_POWER (Battery Saver), SUSTAINED_PERFORMANCE, FIXED_PERFORMANCE, display inactive, VR, double-tap-to-wake, and vendor modes.
Launch boostActivityManager / PowerManagerService asks for a launch boost so the new process and first frame get big cores and high OPPs. Missing this is a classic "cold start regressed after a HAL change" bug.

On a wearable the same HAL is where you tune "boost hard enough that wrist-raise to face is fast, then drop immediately so AOD current stays in budget." Over-long interaction boosts are a common silent drain.

ADPF

ADPF (Android 12+) is the app-facing, closed-loop version of the same idea. PerformanceHintManager.createHintSession(tids, targetNanos) tells the platform "these threads should finish one unit of work every N nanoseconds." The app then calls reportActualWorkDuration() each cycle. The vendor uses the error (actual versus target) to raise or lower frequency instead of a blind timed boost. HWUI uses sessions for RenderThread; games and camera pipelines should too.

  • Thermal headroom (PowerManager.getThermalHeadroom(seconds), Android 11+) is part of the same story: forecast how close the device is to throttling so the app can drop frame rate or quality first. See Thermal.
  • Thermal status listeners (Android 10+) report the stepped severity the Thermal HAL publishes.
  • Later releases add more ADPF surfaces (GPU hints and extra headroom APIs). The interview point is the loop: target duration, actual duration, thermal headroom, then back off before the kernel thermal governor slams frequency.
// Sketch: one ADPF session for a render loop
val hintMgr = getSystemService(PerformanceHintManager::class.java)
val session = hintMgr.createHintSession(intArrayOf(Process.myTid()), 8_333_333L) // 120 Hz
// each frame:
val start = SystemClock.elapsedRealtimeNanos()
drawFrame()
session?.reportActualWorkDuration(SystemClock.elapsedRealtimeNanos() - start)
Common pitfall Treating Power HAL boosts as a substitute for efficient work, or leaving a launch / interaction boost asserted. Hints only move energy earlier; they do not create it. On traces, a boost that never returns to the little-core mid OPP after the gesture is a HAL or hint-session bug.
Interview angle "How does Android make launch and touch feel fast without destroying battery?" Name the Power HAL (setBoost / setMode, launch and interaction), then ADPF hint sessions (target versus actual duration) and thermal headroom. Say which process calls the HAL (framework) versus who creates sessions (app or HWUI), and how you would confirm a boost in Perfetto (frequency, uclamp, cpu_frequency).

Battery: fuel gauge, charging and health

The battery subsystem answers three questions: how much charge is left (fuel gauge), how to refill it safely and quickly (charger), and how to keep the cell healthy for years (charging policy). Mistakes here show up as sudden shutdowns at 15%, percentage jumps, slow charging, swelling, or capacity fading too fast.

Analogy

The fuel gauge is like a car's fuel meter that has to guess the tank level from both the float sensor and how much fuel has gone through the pipe, while the tank slowly shrinks with age. In the real system the fuel gauge combines a voltage-based model (open-circuit voltage tables) with coulomb counting (integrating current through a sense resistor), corrected for temperature, load and aging, to report state of charge.

Fuel gauge

Usually part of the PMIC or a dedicated chip. Voltage-based estimates are good at rest but poor under load; coulomb counting is good short-term but drifts. Hybrid gauges combine them and learn the cell's real capacity over cycles. Outputs: state of charge, voltage, current, temperature, full-charge capacity, cycle count, state of health.

Charging

Lithium-ion charging goes pre-charge (deeply empty), constant current, then constant voltage until the current tapers to a termination level. JEITA temperature zones reduce current or voltage when cold or hot. Watches usually charge through pogo pins or wireless coils at low power.

Battery health

Capacity fades with cycles, high temperature and long time at high state of charge. Adaptive or overnight charging holds at a lower level and finishes just before the user wakes; some products offer a charge limit (for example 80%).

Shutdown behavior

The system shuts down at a low state of charge or low voltage. A poorly calibrated gauge, or high current pulses causing voltage droop, can cause early shutdowns while the percentage still looks healthy.

Android battery plumbing

Fuel gauge + charger (PMIC)  ->  kernel power_supply class  (/sys/class/power_supply/battery/...)
   -> Health HAL (HIDL 2.x; AIDL android.hardware.health from Android 13)  ->  BatteryService
   -> ACTION_BATTERY_CHANGED broadcast, BatteryManager properties, Settings battery UI
   -> BatteryStatsService attributes drain to apps and components (batterystats)
         using ODPM / Power Stats HAL rails when present, else power_profile.xml models

Without on-device rail monitors, batterystats estimates energy from a device-specific power_profile.xml (CPU cluster current at each frequency, screen, radios, GPS, DSP, and similar). A profile copied from another board makes Settings "battery usage" lie; treat it as a model and confirm with a lab monitor or Power Stats HAL rails.

adb shell dumpsys battery                                   # level, status, temperature, plugged
adb shell cat /sys/class/power_supply/battery/capacity      # percent
adb shell cat /sys/class/power_supply/battery/current_now   # usually microamps, sign convention varies
adb shell cat /sys/class/power_supply/battery/voltage_now   # microvolts
adb shell cat /sys/class/power_supply/battery/cycle_count
Common pitfall Reading current_now once and calling it the device's power draw. It is noisy, sampled, and has platform-specific units and sign conventions. For real numbers, average over time, use the Power Stats HAL rails in Perfetto, or use a lab power monitor in place of the battery.
Interview angle Interviewers ask how Android knows the battery percentage, why percentages jump, and why devices shut down at non-zero percent. Mention the hybrid fuel gauge, temperature and aging compensation, voltage droop under load, and the Health HAL path to BatteryService.

The power debug toolbox

Power debugging is evidence work: quantify the drain, find what is awake and why, attribute it to an owner, fix, and re-measure. Each tool answers a different question, and a good engineer knows which one to reach for first.

Analogy

Debugging battery drain is like auditing a household's electricity bill. The monthly bill (batterystats) tells you how much was used and roughly by whom; the smart meter graph (Battery Historian timeline) shows when; plugging each appliance into a watt meter (Perfetto power rails or a lab power monitor) tells you exactly how much each draws; and checking which lights were left on (wakeup_sources, dumpsys power) finds the culprit. In the real system you combine these views on one timeline to go from symptom to owner.

QuestionTool
How much did we drain, and which app or component used it?dumpsys batterystats, Battery Historian
Is the device suspending? Why did it wake?/sys/kernel/debug/suspend_stats, /sys/kernel/wakeup_reasons/last_resume_reason, kernel log, Perfetto
Who is holding it awake?dumpsys power (app wakelocks), /sys/kernel/debug/wakeup_sources or /sys/class/wakeup/ (kernel), dumpsys suspend_control_internal on recent releases
Is Doze working? Which alarms and jobs fire?dumpsys deviceidle, dumpsys alarm, dumpsys jobscheduler, dumpsys usagestats
Which interrupt keeps firing?/proc/interrupts deltas, wakeup reasons
Which sensors are active, at what rate?dumpsys sensorservice
How much power does each rail (CPU, GPU, display, modem) draw?Perfetto with power rails (Power Stats HAL / on-device power monitor), lab power monitor
Are CPUs at a sensible frequency and idle state?Perfetto power/cpu_frequency and power/cpu_idle, /sys/devices/system/cpu/cpu*/cpufreq/ and cpuidle/
Is a Power HAL boost or ADPF session stuck?Perfetto power/vendor boost traces, uclamp and cpuset changes around launch or touch
Is it too hot, and what is throttled?dumpsys thermalservice, thermal zones and cooling devices in sysfs

Battery Historian workflow

# Start clean, on battery, with a fixed use-case
adb shell dumpsys batterystats --reset
adb shell dumpsys batterystats --enable full-wake-history   # richer wakelock history
# unplug USB or use Wi-Fi adb, run the scenario for a fixed time
adb bugreport bugreport.zip
# Load bugreport.zip into a locally hosted Battery Historian instance
# (open source; run it inside the company network, never upload bugreports to public sites)

Battery Historian turns a bugreport into a timeline of wakelocks, wakeup reasons, CPU running, screen, radio, Doze state, jobs, alarms and per-app statistics. It is useful for "battery regressed" triage, although newer work relies more on Perfetto.

Perfetto with power data

# Perfetto config fragment: battery counters and power rails
data_sources {
  config {
    name: "android.power"
    android_power_config {
      battery_poll_ms: 250
      battery_counters: BATTERY_COUNTER_CURRENT
      battery_counters: BATTERY_COUNTER_CHARGE
      collect_power_rails: true
    }
  }
}
data_sources { config { name: "linux.ftrace"
  ftrace_config { ftrace_events: "power/suspend_resume" ftrace_events: "power/cpu_idle"
                  ftrace_events: "power/cpu_frequency" ftrace_events: "power/wakeup_source_activate" } } }

Power rails come from the Power Stats HAL, backed on supported devices by on-device power monitors (ODPM) that measure energy per rail. Combined with scheduling, frequency, idle and suspend events in one trace, you can see exactly which process woke the device and what it cost.

Kernel-level checks

adb shell cat /sys/kernel/debug/wakeup_sources    # active_count, event_count, total_time, active_since
adb shell cat /sys/kernel/debug/suspend_stats     # success, fail, last_failed_dev, last_failed_errno
adb shell cat /sys/kernel/wakeup_reasons/last_resume_reason
adb shell cat /sys/power/mem_sleep
adb shell cat /proc/interrupts                    # take two snapshots and diff
adb shell dmesg | grep -iE "PM: suspend|wakeup|resume"

A triage method that works

  1. Quantify Measure drain per hour or average current for a fixed scenario on several devices, before and after.
  2. Is it suspending? Suspend residency and suspend failures. If not suspending, find the holder.
  3. Why does it wake? Wakeup reasons and interrupt deltas; alarms, jobs and network in Doze.
  4. What runs when awake? Per-process CPU, sensors, radio activity, rails in Perfetto.
  5. Attribute and bisect Map to an owner; bisect builds or components if it is a regression.
  6. Fix and prevent Fix, re-measure with the same scenario, and add or tighten a regression gate.
Common pitfall Measuring with USB connected. The cable powers the device, can keep it awake and disables Doze. Use Wi-Fi adb or disconnect during the run, or wire a lab power monitor in place of the battery. Also control screen brightness, radio conditions (a shield box or fixed location), accounts and installed apps, or run-to-run noise will hide real regressions.
Interview angle "Battery regressed after a platform update; walk me through your debug" is the most common power question. Interviewers want a method, not a list of commands: quantify, check suspend, find the waker, attribute, bisect, fix, re-measure, gate.

Product KPIs and how a platform lead gates on them

A platform lead is accountable for delivering a product that is competitive on functionality, stability, power and performance. That only works if each is turned into measurable KPIs with budgets, tracked on every integration drop, and enforced as promotion gates rather than discovered in the field.

Analogy

Release gating is like airport security for code: every passenger (each integration drop) goes through the same scanners (power, stability and performance tests) with fixed thresholds, and anyone who sets off the alarm is stopped before boarding, not after take-off. In the real system each build runs a standard test matrix, results are compared against budgets and the last good baseline, and a regression beyond threshold blocks promotion until fixed or explicitly waived.

CategoryKPIHow it is measured
PowerStandby current (mA) or drain per hourIdle device, screen off or AOD, connected to phone, fixed radio conditions; lab power monitor or battery counters
PowerDays of battery (typical use)Scripted day-in-the-life profile (notifications, workouts, glances, AOD hours) run to empty or extrapolated
PowerActive-use drain per use caseCurrent budget for workout with GPS and heart rate, LTE call, music streaming, navigation
PowerSuspend residency, wakeups per hour, wakelock timebatterystats, Perfetto, wakeup sources
ThermalPeak and sustained skin temperature; throttling timeThermal chamber, skin model, worst-case workloads and charging
StabilityCrash rate, ANR rate, crash-free sessions or devices, watchdog and subsystem restarts, kernel panicsLong-run stress labs, dogfood and field telemetry
StabilityBoot success rate, OTA success rateRepeated boot cycles, OTA matrix including rollback
PerformanceBoot time (cold boot to face)Boot markers in logs and traces, many iterations
PerformanceWake-to-render latency, app and Tile launch time, jank (missed frames)Perfetto frame timeline, dumpsys gfxinfo, high-speed camera for end to end
FunctionalityFeature completeness, connectivity reliability, health-sensor accuracyFeature test matrix, certification suites, reference-device comparisons

How to gate well

  1. Set budgets Agree targets per use case (for example standby current, workout drain, boot time) with product and hardware teams, and split them into per-component budgets.
  2. Control the measurement Fixed devices, firmware, accounts, brightness, radio environment; several devices and repeated runs so the noise band is known.
  3. Compare with baseline Every drop is compared with the last good build; a change larger than the noise band and the threshold is a regression.
  4. Block or waive A regression blocks promotion into the main image unless explicitly waived with an owner and a fix date.
  5. Attribute fast Automated bisection across the drop's changes, with dashboards showing which component moved.
  6. Watch the field After release, field telemetry (battery drain distributions, crash and ANR rates, thermal events) confirms lab results and catches long-tail issues.
Tip Use distributions, not single numbers. A median standby current can look fine while 5% of devices drain three times faster because of one app, one carrier or one hardware revision. Track high percentiles as well as medians.
Interview angle Lead-level interviews ask "what would you gate a wearable release on?" and "how do you stop power regressions from reaching users?" A strong answer names KPIs in all four categories, explains controlled measurement and noise, and describes the gate-and-waiver process on every integration drop.

Quick revision

  • Battery life is roughly capacity divided by average current; average current is dominated by suspend residency.
  • Suspend-to-RAM freezes tasks, suspends devices, offlines CPUs and puts DDR in self-refresh; only wake interrupts resume it.
  • One driver's failing suspend callback aborts the whole suspend; check suspend_stats and last_failed_dev.
  • Modern Android triggers suspend from the SystemSuspend service using the wakeup_count handshake and /sys/power/state.
  • Suspend blockers were Android's original (2009) mechanism. Mainline wakeup sources arrived in kernel 2.6.37 (2010); autosleep and /sys/power/wake_lock followed in 3.5 (2012).
  • cpuidle is per-CPU idling between tasks; system suspend is the whole device sleeping.
  • DVFS scales frequency and voltage together; dynamic power is roughly C·V2·f.
  • cpufreq governors: historic Android interactive, modern default schedutil (frequency from scheduler utilization).
  • cpuidle governors (menu, TEO) pick a C-state by predicted idle time versus target residency.
  • EAS places tasks using a per-CPU energy model; it yields to ordinary load balancing when the system is overutilized.
  • Android steers EAS with uclamp (min boosts, max caps) and cpusets (top-app versus background).
  • The Power HAL (AIDL from Android 11) implements short setBoost hints (interaction, launch) and longer setMode (interactive, low power, sustained performance).
  • ADPF (Android 12+) hint sessions close the loop: target work duration versus actual, plus thermal headroom so apps back off before throttle.
  • A partial wakelock keeps the CPU on with the screen off; it is the number one invisible drain.
  • Always acquire wakelocks with a timeout and release in a finally block; tag them as app:component.
  • Kernel wakeup sources are visible in /sys/kernel/debug/wakeup_sources; app wakelocks in dumpsys power.
  • Light Doze starts soon after screen off on battery; deep Doze requires the device to be stationary.
  • In deep Doze, network, jobs, syncs, standard alarms and Wi-Fi scans are deferred and app wakelocks are ignored.
  • Exact allow-while-idle alarms, alarm-clock alarms, high-priority FCM and allowlisted apps get through Doze.
  • App Standby buckets: active, working set, frequent, rare, restricted (Restricted added in Android 11, API 30).
  • Android 12 adds app hibernation and cached-app freezer; Android 13 adds Low Power Standby for long unused periods.
  • Allow-while-idle exact alarms are rate-limited to roughly one every nine minutes while idle.
  • WorkManager is the default for deferrable guaranteed work; periodic minimum is 15 minutes.
  • Exact alarms wake the AP and need special permission; use them only for user-visible times.
  • Handler.postDelayed uses uptime, which stops during suspend.
  • The sensor hub runs always-on sensing with FIFO batching so the AP stays suspended.
  • Other offloads: CHRE nanoapps, audio DSP hotword, Wi-Fi packet filtering, BLE scan filtering, GNSS batching, modem DRX.
  • Skin temperature, not junction temperature, is usually the binding thermal limit on phones and watches.
  • Thermal zones have trip points; cooling devices (cpufreq, GPU, charge current, modem) are applied by governors.
  • The thermal HAL reports severities from NONE to SHUTDOWN; apps use thermal status listeners and headroom.
  • Heat increases leakage, which increases power, which increases heat.
  • batterystats uses Power Stats HAL rails when present, otherwise a device power_profile.xml model.
  • Fuel gauges combine voltage models and coulomb counting, corrected for temperature and aging.
  • Lithium-ion charging is constant current then constant voltage; JEITA zones limit charging when hot or cold.
  • Voltage droop under load plus an aging cell causes shutdowns at non-zero percent.
  • Battery Historian turns a bugreport into a drain timeline; run it locally and keep bugreports internal.
  • Perfetto with power rails shows energy per rail alongside scheduling, frequency and suspend events.
  • Never measure power with USB attached; use a power monitor or pure battery runs in a controlled setup.
  • Triage: quantify, check suspend, find the waker, attribute, bisect, fix, re-measure, gate.
  • Gate each integration drop on power, thermal, stability, performance and functionality KPIs against a baseline and noise band.

Glossary

ADPF
Android Dynamic Performance Framework: hint sessions (target versus actual work duration) and thermal headroom so apps and HWUI can hold a frame rate without over-boosting.
Adaptive Battery
Android feature that predicts app usage to place apps in standby buckets.
App hibernation
Android 12+ state for unused apps: permissions revoked and cache cleared until the user opens the app again.
AlarmManager
Android API for time-based triggers, inexact or exact, optionally waking the device.
ANR
Application Not Responding: the main thread was blocked too long; a stability KPI.
App Standby bucket
Per-app category (active, working set, frequent, rare, restricted) that limits background jobs and alarms.
Autosleep
Kernel mode in which the system suspends automatically whenever no wakeup source is active.
Battery Historian
Open-source tool that visualizes batterystats from a bugreport as a timeline.
batterystats
Framework service and dumpsys output that attributes battery use to apps and components.
CHRE
Context Hub Runtime Environment: runs small nanoapps on the low-power sensor hub.
Cooling device
Kernel thermal object that can reduce heat, such as CPU frequency, GPU frequency or charge current.
Coulomb counting
Measuring charge in and out by integrating current over time.
cpufreq
Kernel subsystem that sets CPU frequency and voltage (DVFS) via a governor such as schedutil or interactive.
cpuidle
Kernel subsystem that puts idle CPUs into low-power idle states between tasks, chosen by latency and target residency.
Doze
Device-wide idle mode that defers apps' network, jobs and alarms into maintenance windows.
DRX / eDRX
Modem sleep cycles between paging occasions; extended DRX gives longer sleep.
DVFS
Dynamic voltage and frequency scaling: clock and supply voltage move together; dynamic power scales about as V2f.
EAS
Energy Aware Scheduling: places tasks using a per-CPU energy model and PELT utilization; pairs with schedutil.
Foreground service
Service with a visible notification that the system keeps alive for user-visible work.
Fuel gauge
Hardware and algorithm that estimate battery state of charge and health.
Health HAL
HAL that reports battery and charger information to BatteryService.
JEITA
Industry guideline that adjusts charge current and voltage by battery temperature.
JobScheduler
Android system service that runs background jobs when constraints are met, batched and Doze-aware.
Low Power Standby
Android 13+ deeper idle after a long unused period, with tighter restrictions than ordinary deep Doze.
Maintenance window
Short period during Doze when deferred work is allowed to run.
ODPM
On-device power monitor: hardware that measures energy per power rail, exposed through the Power Stats HAL.
OPP
Operating performance point: a legal frequency and voltage pair for a CPU or GPU cluster.
Partial wakelock
Wakelock that keeps the CPU running while allowing the screen to turn off.
PerformanceHintManager
ADPF API that creates a hint session for a set of threads with a target work duration.
Perfetto
Android's system-wide tracing tool, including power rails, battery counters and suspend events.
PMIC
Power management IC: regulators, charger, fuel gauge and power sequencing.
Power HAL
Vendor HAL (android.hardware.power) that applies boosts and modes (interaction, launch, low power) by programming cpufreq, uclamp and related knobs.
Power Stats HAL
HAL that exposes per-rail energy and per-subsystem residency data.
power_profile.xml
Device overlay of component currents used by batterystats when measured rails are not available.
PowerManagerService
Framework service that manages wakelocks, screen state and requests to suspend.
schedutil
cpufreq governor that sets frequency from scheduler utilization; the modern Android default.
Sensor hub
Always-on low-power processor that samples, fuses and batches sensor data.
Skin temperature
Estimated temperature of the device surface the user touches, often a virtual sensor.
State of charge (SOC)
Remaining battery charge as a percentage of current full capacity.
State of health (SOH)
Current full capacity relative to the original design capacity.
Suspend blocker
Android's original kernel mechanism for preventing opportunistic suspend; replaced by wakeup sources.
Suspend residency
Share of time the system spends suspended; the key standby power metric.
Suspend-to-RAM
System sleep state with DDR in self-refresh and most of the SoC powered down.
SystemSuspend
Android userspace service (suspend HAL) that tracks native wakelocks and triggers kernel suspend.
Thermal engine
Vendor thermal policy daemon or component implementing multi-sensor mitigation rules.
Thermal HAL
HAL that reports temperatures, thresholds and throttling severity to the framework.
Thermal zone
Kernel object for a temperature sensor with trip points and bound cooling devices.
Trip point
Temperature threshold in a thermal zone that triggers cooling or shutdown.
uclamp
Utilization clamp: per-task or per-cgroup min and max utilization that steers EAS and schedutil.
Wakelock
Request to keep the CPU (and optionally the screen) from sleeping.
Wakeup source
Kernel object that blocks system suspend while active; the mainline form of a wakelock.
WorkManager
Jetpack library for deferrable, guaranteed background work with constraints.

Interview questions

Fundamentals

Why does battery life roughly equal time spent suspended?

Suspend current is one or two orders of magnitude lower than active or awake current, so total energy is dominated by how long and how often the AP is awake. Maximizing suspend residency (few, short wakeups) is the main lever, followed by display and radio power.

What is suspend-to-RAM?

A system sleep state where tasks are frozen, devices are suspended through their driver callbacks, non-boot CPUs are offline, most clocks and rails are off, and DDR is kept in self-refresh so memory contents survive. Only a configured wake interrupt (RTC alarm, button, sensor, modem, Bluetooth) resumes it.

What is a wakelock?

A request to keep the CPU (and optionally the screen) from sleeping. Apps use PowerManager.WakeLock; PowerManagerService tracks them and holds a native wakelock in SystemSuspend while any are active. While any wakelock or kernel wakeup source is held, the system cannot suspend.

What is a partial wakelock and why is it dangerous?

It keeps the CPU running while the screen is off. The user sees a sleeping device, but the system never suspends, so drain can be ten times the suspend current. It is the most common invisible battery bug.

What is a wakeup source?

The kernel's mechanism for blocking suspend, used by drivers while they handle an event. A driver registers one and calls stay-awake and relax functions, or signals a wakeup event with a timeout. Statistics appear in /sys/kernel/debug/wakeup_sources and /sys/class/wakeup/.

What is the difference between cpuidle and system suspend?

cpuidle puts individual CPUs into idle states between tasks for microseconds to milliseconds while the system keeps running. System suspend freezes everything and powers down most of the SoC for seconds to hours, and is blocked by wakeup sources. Both save power, but suspend saves far more.

What is Doze?

A device-wide idle mode that starts when the screen is off and the device is on battery. It defers apps' network access, jobs, syncs and (in deep Doze) alarms and wakelocks into periodic maintenance windows, which get further apart the longer the device stays idle.

What is App Standby?

A per-app policy based on how recently the user used the app. Apps are placed in buckets (active, working set, frequent, rare, restricted), and the lower the bucket, the stricter the limits on jobs, alarms and high-priority messages. The Restricted bucket was added in Android 11 (API 30); Android 12 made unused apps more likely to land there and added hibernation.

What is DVFS?

Dynamic voltage and frequency scaling. The SoC changes a cluster's clock and its supply voltage together, using an OPP table of legal pairs. Dynamic power is roughly C·V2·f, so dropping voltage with frequency saves more than frequency alone. cpufreq chooses the OPP; thermal policy can cap the maximum.

What is Energy Aware Scheduling?

A scheduler extension that places each waking task on the CPU that can meet its utilization with the least energy, using a per-CPU energy model and PELT. Small tasks stay on little cores; heavy tasks go to big or prime cores. It works with the schedutil governor. When the system is overutilized, EAS steps aside and ordinary load balancing takes over. Android steers it with uclamp and cpusets. See Linux Kernel & BSP for the scheduler side.

What is the Power HAL?

The vendor HAL (android.hardware.power, AIDL from Android 11) that applies short boosts and longer modes. Framework code (PowerManagerService, activity manager, input, HWUI) calls setBoost (interaction, display update, historically launch) and setMode (interactive, low power, sustained performance). The HAL programs cpufreq, uclamp, cpusets or a vendor boost driver; it does not replace the governors.

What is ADPF?

The Android Dynamic Performance Framework. Apps (and HWUI) create a PerformanceHintManager session for a set of threads with a target work duration, then report the actual duration each cycle so the vendor can pick the lowest OPP that still meets the deadline. Thermal headroom is the other half: forecast throttling and drop quality first. Sessions arrived in Android 12; headroom in Android 11.

What is the difference between Doze and App Standby?

Doze is device-wide and triggered by the device being idle (screen off, on battery, possibly stationary); it affects all non-exempt apps. App Standby is per-app and triggered by the user not using that app, even while the device is in use.

When would you use WorkManager, JobScheduler or AlarmManager?

WorkManager for deferrable, guaranteed background work that must survive process death and reboot (the default). JobScheduler is the platform API it uses; use it directly mainly in platform code. AlarmManager exact alarms only for true time-critical, user-visible events like alarm clocks and reminders, because they wake the AP and can bypass Doze.

What is a foreground service and what does it cost?

A service with a visible notification that the system keeps alive for user-visible ongoing work, such as workouts, navigation, calls or media. It prevents the process from being cached or killed and usually involves wakelocks and sensors, so it costs power; it should be declared with a type (Android 14+) and stopped as soon as the work ends.

What is sensor-hub offload?

Running always-on sensing, fusion and batching on a low-power co-processor so the AP stays suspended. The hub buffers samples in a FIFO and wakes the AP only when a batch is ready or a meaningful event occurs.

Why is thermal management harder on a watch than on a laptop?

No fan, tiny thermal mass, and the device is worn against the skin for hours, so skin-temperature limits bind early. Heat builds quickly during workouts, LTE use and charging.

What is a thermal zone?

A kernel object representing a temperature sensor (real or virtual) with trip points. When the temperature crosses a trip point, the zone's governor applies bound cooling devices such as CPU frequency limits.

What is a fuel gauge?

The hardware and algorithm that estimate battery state of charge, capacity and health, usually by combining a voltage model with coulomb counting and correcting for temperature and aging.

What does batterystats do?

BatteryStatsService records power-relevant events (wakelocks, CPU time, screen, radio, sensors, jobs, alarms) and attributes estimated energy to apps and components. dumpsys batterystats prints it, and a bugreport includes it for Battery Historian.

What is Battery Historian?

An open-source Google tool that parses a bugreport and shows a timeline of wakelocks, wakeup reasons, CPU, screen, radio, Doze, jobs and alarms, plus per-app statistics. It is used for battery regression triage and should be hosted locally so bugreports stay internal.

Which dumpsys commands matter for power?

dumpsys batterystats (attribution), dumpsys power (wakelocks, wake state), dumpsys deviceidle (Doze), dumpsys alarm, dumpsys jobscheduler, dumpsys usagestats (buckets), dumpsys sensorservice, dumpsys thermalservice and dumpsys battery.

What KPIs define a wearable's power performance?

Standby current or drain per hour, days of battery under a typical-use profile, active drain per use case (workout, LTE call, music), AOD current, suspend residency and wakeups per hour, and wakelock time.

Why should you never measure power with USB attached?

USB supplies power, so the battery reading does not reflect real drain; it often holds the device awake, disables Doze (the device counts as charging) and adds adb traffic. Measure on battery with Wi-Fi adb disconnected, or with a lab power monitor wired in place of the battery.

Going deeper

Walk through what happens when an Android device suspends and resumes.

When PowerManagerService releases its last wakelock and no native wakelocks remain, SystemSuspend reads /sys/power/wakeup_count, writes it back (which fails if a wakeup event occurred in between), then writes mem to /sys/power/state. The kernel freezes tasks, runs driver suspend callbacks in phases (prepare, suspend, late, noirq), offlines non-boot CPUs and enters the platform low-power state through firmware. A wake interrupt starts resume: reverse callbacks, CPUs online, tasks thawed, wakeup reason recorded, and a wakeup source is typically held briefly so the event can be handled.

What is the purpose of the wakeup_count handshake?

It closes a race: if a wakeup event happens after userspace decides to suspend but before the kernel actually suspends, the event could be lost. Userspace reads the count, checks there are no wakelocks, and writes the count back; the kernel rejects the write if the count changed, and userspace retries later. This keeps suspend from swallowing events.

What were suspend blockers and what replaced them?

Suspend blockers (called wakelocks in early Android kernels) were Android's 2009 mechanism for opportunistic suspend: suspend whenever no blocker is held. Upstream rejected that design. Wakeup sources landed first, in kernel 2.6.37 (2010). Autosleep and the userspace /sys/power/wake_lock / wake_unlock interface were merged later, in kernel 3.5 (2012). Modern Android uses kernel wakeup sources for drivers and the SystemSuspend service for userspace wakelocks; it usually does not use autosleep.

How do the interactive and schedutil cpufreq governors differ?

interactive was Android's historic governor: on a load spike or input event it jumps to a configured hispeed frequency, then ramps using tunables such as go_hispeed_load. schedutil is the modern default; it sets frequency from scheduler utilization (PELT) and uclamp, so EAS placement and frequency are one policy. Interactive needed a lot of per-device tuning; schedutil needs a good energy model and sane uclamp from the Power HAL and cgroups.

How does a Power HAL interaction or launch hint actually save or spend power?

It spends power on purpose for a short window so the work finishes sooner and the AP can return to idle or suspend. The HAL typically raises uclamp.min or a cluster max frequency for tens to hundreds of milliseconds. That is a win if the boost matches the work; it is a drain if the boost stays asserted (a leaked hint) or if every small binder call triggers a full launch boost. Confirm in Perfetto with cpu_frequency and the Power HAL traces around the gesture or startActivity.

How does an ADPF hint session differ from a Power HAL boost?

A Power HAL boost is open-loop and timed: "go fast for N ms." An ADPF session is closed-loop: the app names the threads and a target duration, then reports actual duration each cycle. The vendor raises or lowers frequency from the error. Boosts are the right tool for one-shot events (touch, launch); sessions are the right tool for steady periodic work (frames, camera pipelines). Use thermal headroom with either so you do not fight the thermal governor.

What is power_profile.xml and when is it wrong?

A device overlay that lists typical currents for CPU clusters at each frequency, the screen at brightness steps, radios, GPS and other components. BatteryStats uses it to estimate app energy when the Power Stats HAL / ODPM rails are missing. It is wrong when copied from another board, when voltages or process corners differ, or when a new rail (for example a sensor hub) is omitted. Treat Settings battery percentages as a model; confirm with a lab monitor or measured rails.

What are app hibernation, the cached-app freezer and Low Power Standby?

Hibernation (Android 12): unused apps have permissions revoked and cache cleared until the user opens them. Cached-app freezer (Android 12, earlier experiments in 11): cached processes are frozen so they do not run until they become active again. Low Power Standby (Android 13): after a long unused period (screen off, not charging), the device applies idle restrictions even tighter than deep Doze. All three sit on top of buckets and Doze; none replaces a leaked kernel wakeup source.

How do you find which wakelock is draining the battery?

For app wakelocks: dumpsys power for current holders, batterystats or Battery Historian for totals and timeline per tag and uid. For kernel: /sys/kernel/debug/wakeup_sources and look for growing active_count, total_time or a non-zero active_since. Perfetto with wakeup-source ftrace events shows exactly when each was taken. Then map the tag or source to its owner.

What restrictions apply in deep Doze and what is exempt?

Restricted: network access, app wakelocks ignored, standard alarms deferred, Wi-Fi scans stopped, sync adapters and jobs deferred to maintenance windows. Exempt or partly exempt: setExactAndAllowWhileIdle and setAndAllowWhileIdle (rate-limited), setAlarmClock, high-priority FCM (brief exemption), and apps on the battery-optimization allowlist.

How do light and deep Doze differ?

Light Doze starts shortly after screen off on battery, even if the device moves, and restricts network and jobs with frequent maintenance windows. Deep Doze requires the device to be stationary for a longer time (checked with motion sensors), adds alarm and wakelock restrictions, and spaces maintenance windows further apart over time.

How do you test an app or component under Doze?

Run dumpsys battery unplug, then dumpsys deviceidle force-idle (or step repeatedly) to enter deep idle, exercise the feature, and observe whether work waits for a maintenance window. Use am set-standby-bucket to test buckets. Reset with deviceidle unforce and battery reset.

What changed with exact alarms in recent Android releases?

Android 12 introduced the SCHEDULE_EXACT_ALARM permission, which users can revoke; Android 13 added USE_EXACT_ALARM for apps whose core function is alarms or calendars; Android 14 denies SCHEDULE_EXACT_ALARM by default for most newly installed apps. The intent is to force most apps onto inexact, batched scheduling.

What are _WAKEUP alarm types?

RTC_WAKEUP and ELAPSED_REALTIME_WAKEUP alarms wake the device from suspend when they fire. RTC and ELAPSED_REALTIME alarms are delivered only when the device next wakes for another reason. ELAPSED_REALTIME counts time since boot, including sleep, and is not affected by wall-clock changes, so it is preferred for intervals.

Why is Handler.postDelayed unreliable for long background delays?

Handler timing uses SystemClock.uptimeMillis(), which does not advance while the device is suspended. A 10-minute delay may fire hours later, or never if the process is killed. Holding a wakelock to keep it on time wastes power. Use WorkManager or AlarmManager.

How does WorkManager decide when to run work?

It persists work requests in its database and schedules them through JobScheduler on modern devices, passing constraints (network type, charging, battery not low, device idle, storage not low) and backoff policy. The system runs them when constraints are met, respecting Doze, buckets and quotas, and batches them with other work. Periodic work has a 15-minute minimum interval; expedited work runs sooner within quotas.

What is the Android Packet Filter and why does it save power?

APF is a small bytecode program the framework installs in Wi-Fi firmware to drop uninteresting packets (for example irrelevant multicast or broadcast traffic) while the AP is suspended. Without it, each packet could wake the AP. It is an example of offloading decisions to low-power hardware.

What is CHRE?

The Context Hub Runtime Environment: a framework for running small nanoapps on the low-power context hub (sensor hub), managed through the Context Hub HAL and ContextHubManager. It lets OEMs run custom always-on logic, such as gesture or activity detection, without waking the AP.

Explain the kernel thermal framework.

Thermal zones represent sensors and have trip points of types passive, active, hot and critical. Cooling devices (cpufreq, GPU devfreq, charge current, modem, backlight) are bound to zones. A governor (step_wise, power_allocator, user_space) decides cooling states as temperature crosses trips. A critical trip triggers an orderly shutdown. Everything is visible under /sys/class/thermal.

What does the Thermal HAL provide to the framework?

Current temperatures by type (CPU, GPU, battery, skin, USB port and more), thresholds, cooling device states, and callbacks when throttling severity changes. Severity levels run NONE, LIGHT, MODERATE, SEVERE, CRITICAL, EMERGENCY, SHUTDOWN. ThermalManagerService consumes it and exposes status and headroom to apps; at SHUTDOWN the framework powers off cleanly.

How should an app respond to thermal status?

Register a thermal status listener or poll thermal headroom, and reduce work before the system throttles hard: lower frame rate, resolution or quality, pause non-essential work, reduce sensor rates. Games and camera apps use headroom forecasts to stay under the limit smoothly instead of hitting it abruptly.

What is a virtual skin temperature sensor?

There is usually no sensor on the case surface itself, so the platform estimates skin temperature from several board thermistors using a weighted model calibrated in the lab against thermocouples on the case. The thermal policy uses this estimate for comfort and safety limits.

How does a lithium-ion charger work?

Pre-charge at low current if the cell is deeply discharged, then constant current until the cell reaches its target voltage, then constant voltage while current tapers, then termination. JEITA zones reduce current or voltage when the battery is cold or hot, and thermal policy can further limit current if the device is hot.

Why do battery percentages jump or the device shut down at 10 to 20%?

The fuel gauge's model may be miscalibrated for the cell or not yet learned its real capacity; aging reduces capacity; cold temperatures increase internal resistance; and high current pulses (LTE, GPS start, screen) cause voltage to droop below the shutdown threshold even though charge remains. Fixes include gauge profile tuning, capacity learning, and limiting peak current at low charge.

How does battery data flow from hardware to the framework?

The fuel gauge and charger drivers publish properties in the kernel power_supply class. The Health HAL reads them and reports to BatteryService in system_server, which broadcasts ACTION_BATTERY_CHANGED and answers BatteryManager queries. BatteryStats uses battery level and charge counter for attribution.

What can you get from Perfetto that batterystats cannot give you?

Precise timing: exactly which thread ran after each wakeup, CPU frequency and idle state, suspend and resume events, wakeup-source activations, and on supported devices energy per power rail and battery current counters on the same timeline. batterystats gives aggregated attribution; Perfetto gives causality.

What is the Power Stats HAL?

A HAL that exposes energy consumed per power rail (from on-device power monitors), energy consumers such as display or modem, and residency in low-power states per subsystem. Perfetto and batterystats use it for more accurate attribution than models based on time and current constants.

How would you set up a reliable power measurement?

Use a lab power monitor wired in place of the battery, or long battery runs; fixed hardware and build; fixed accounts, apps and brightness; controlled radio conditions (shield box or fixed location); airplane mode variants to isolate radios; several devices and repeated runs to know the noise band; and a warm-up period so post-boot work settles.

Advanced

A driver's suspend callback returns -EBUSY intermittently. What is the effect and how do you handle it?

The whole suspend aborts, devices already suspended are resumed, and userspace retries later. Intermittent failures cause bursts of failed attempts, each costing power, and low suspend residency. Find it in suspend_stats (last_failed_dev, errno) and kernel logs. Fix the driver: finish or cancel pending work before suspend, hold a wakeup source while busy instead of failing, and make the callback idempotent.

How do you distinguish "cannot suspend" from "suspends but wakes too often"?

Cannot suspend: very low suspend residency, few or no suspend entries, a wakelock or wakeup source constantly active, or repeated suspend failures. Wakes too often: many successful suspend and resume cycles per hour, with wakeup reasons pointing to an interrupt (RTC alarms, sensor, modem, Bluetooth) and short awake periods. The fixes differ: release the holder versus reducing wake events by batching or filtering.

How can a single interrupt line drain the battery, and how would you find it?

A misconfigured wake-capable interrupt (for example a floating GPIO, a sensor with a wrong threshold, or a modem that keeps signaling) wakes the AP repeatedly, and each wake may take a wakeup source that also delays the next suspend. Find it with last_resume_reason and wakeup-reason stats, /proc/interrupts deltas during idle, and Perfetto. Fix the pin configuration, threshold, or disable wake capability on that line.

Explain the trade-off between batch latency and power for sensors.

With sampling period P and max report latency L, the AP wakes about once every L instead of every P. Energy per wake (resume, process, suspend) is roughly fixed, so fewer wakes save power almost linearly until other consumers dominate. The cost is staleness (data arrives up to L late) and FIFO size limits (if L times the rate exceeds the FIFO, the hub must wake the AP earlier or drop data for non-wake-up sensors). Choose L per use case: long for passive tracking, short for a visible workout screen.

What is the power allocator (IPA) thermal governor?

A closed-loop controller that computes a total power budget from the difference between the current temperature and a target, then divides it among cooling devices (CPU clusters, GPU) according to their requested power and weights, converting budgets into frequency caps. It gives smoother, more efficient throttling than stepwise control, but needs accurate power models per device.

Why do hardware limiters exist in addition to software thermal governors?

Software governors act on tens to hundreds of milliseconds. Sudden current spikes can cause junction hotspots or battery voltage droop in microseconds. Hardware limiters cap CPU frequency or current almost instantly (for example on voltage droop or peak current), preventing brownouts and resets. They should rarely trigger; frequent triggers mean the software policy or power budget is wrong.

How does temperature affect power consumption?

Transistor leakage grows roughly exponentially with temperature, so a hot SoC consumes more power at the same frequency, which heats it further. This positive feedback is why sustained workloads must be throttled early, and why power measurements must be taken at a controlled temperature.

How would you build a per-component power budget for a watch?

Start from the product target (for example days of battery under a typical-use profile), convert it to an average current budget, and split it across states (suspend, AOD, interactive, workouts) weighted by expected time in each. Then split each state across components (SoC, display, co-processor, radios, sensors) using rail measurements on reference hardware. Each owning team gets a budget, and integration gates check actual rail energy against it.

How do you make power regression gates statistically sound?

Measure the noise band with repeated runs on several devices of the same build. Use enough samples to detect the smallest regression you care about, compare distributions (not single runs) with a significance test or confidence interval, control temperature and radio environment, and re-run suspected regressions before blocking. Track high percentiles as well as medians, since tail issues matter.

How can radio behavior dominate standby power, and what are the levers?

A cellular modem searching in poor coverage, frequent small data transfers each keeping the radio in a high-power state for seconds (tail energy), BLE with a short connection interval, or Wi-Fi with frequent beacons and multicast can outweigh the AP. Levers: DRX and eDRX, power saving mode, batching network traffic, longer BLE connection intervals, packet filtering in Wi-Fi firmware, and preferring Bluetooth to the phone over LTE on watches.

What is radio tail energy and why does batching help?

After a data transfer, a cellular radio stays in a high-power connected state for a while (seconds) before dropping to idle, in case more data follows. Many small transfers therefore each pay the tail. Batching them into one burst pays the tail once, often saving more energy than the transfers themselves cost.

How would you prevent a system service from holding a wakelock across a Binder call to a slow HAL?

Do not hold a wakelock across unbounded calls. Use asynchronous (oneway) calls with a callback that takes a short wakelock only to process the result, apply timeouts on both the wakelock and the call, and log or alert when a call exceeds its expected latency. In review, flag any wakelock acquired before I/O without a timeout.

What is the difference between s2idle and deep suspend?

s2idle (suspend-to-idle) freezes userspace and suspends devices, but CPUs just enter their deepest idle state and resume is fast; savings depend on the platform reaching deep idle states. Deep suspend (suspend-to-RAM) offlines CPUs and lets firmware power down more of the SoC, saving more but with longer entry and exit. /sys/power/mem_sleep shows which is used.

How do charging policy and thermal interact on a watch?

Charging adds heat in the battery and charger at the same time the user may be using the watch. Thermal policy reduces charge current when skin or battery temperature rises, and JEITA rules limit charging when the battery is too hot or cold. Adaptive charging can hold at a lower level overnight to reduce heat and aging. The trade-off is charge time versus temperature and battery lifespan.

How would you estimate days of battery from lab measurements?

Define a typical-use profile (hours of AOD, number of wrist-raises, notifications, a workout, sleep tracking). Measure average current for each state or activity with a power monitor, weight by time in the profile to get a daily energy, and divide usable battery capacity (accounting for shutdown reserve and aging margin) by it. Validate with real run-down tests.

How can foreground services be abused and how does the platform respond?

Apps start a foreground service with a minimal notification just to avoid being killed and to keep running background work, often with a wakelock. Android responded with background start restrictions (Android 12), mandatory foreground service types with matching permissions (Android 14), time limits on some types, and battery usage visibility. A platform can also flag long-running services in batterystats review.

When does EAS stop being energy-aware, and why does that matter for power?

When the system is overutilized (not enough spare capacity at current OPPs), EAS yields to ordinary load balancing so tasks get CPU even if that means big cores and high frequencies. That is correct for performance, but a device that is "always overutilized" (too-high uclamp.min, a leaked launch boost, or background work uncapped) never uses the energy model and will run hot. Check uclamp and cgroup placement in traces before blaming the energy model.

How do you tell a cpuidle problem from a suspend problem?

If the system is not suspending, wakeup sources and wakelocks dominate. If it is suspending but awake current is high, look at idle residencies and frequency: CPUs that never enter the deeper C-states (a chatty timer, a polling thread, nohz issues) or that sit at a high OPP while "idle" waste milliamps. Perfetto power/cpu_idle and power/cpu_frequency plus /sys/devices/system/cpu/cpu*/cpuidle/ residency counters separate the two.

Scenario & debugging

The battery drains overnight while the device is idle. How do you debug?

Quantify the drain per hour. Check whether it suspends at all (suspend residency, suspend_stats, kernel logs). If not, find the holder in dumpsys power and wakeup_sources. If it suspends but wakes often, read wakeup reasons and interrupt deltas, check dumpsys alarm and jobscheduler for frequent wakers, confirm Doze reached deep idle, and check radios (poor coverage search, Bluetooth reconnect loops). Attribute to an owner, fix, and re-measure.

Standby drain doubled after an upstream merge. Walk the triage.

Reproduce with the fixed standby profile on several devices. Reset batterystats, run, take a bugreport and Perfetto trace, and compare with the last good build: suspend residency, wakeups per hour, wakelock time, top wakeup sources, per-uid CPU. Check wakeup_sources for a new kernel source. Bisect the merge by component or change. Typical culprits: a new or leaked wakelock, a HAL that stopped batching, or an alarm defeating Doze. Fix and add a wakelock and power regression gate.

batterystats shows high "kernel" or "Android system" drain with no obvious app. What next?

Look below the framework: wakeup_sources for a driver holding a source, wakeup reasons for a noisy interrupt, suspend_stats for failed suspends, and Perfetto for kernel threads running after wake. Also check system_server wakelocks by tag (for example alarm, job or sync wakelocks acting on behalf of apps) and native daemons' wakelocks in SystemSuspend stats.

A device never enters deep Doze. Why might that be?

Something keeps resetting the idle state machine: motion detected (a noisy accelerometer or a significant-motion sensor misfiring), the device counted as charging (USB, a faulty charger detect), the screen turning on briefly (notifications waking the display), or a system component keeping the device active. dumpsys deviceidle shows the state and why it exited; step through states manually to see where it fails.

Users report the device gets hot and slow during a video call. How do you investigate?

Reproduce in a thermal chamber with the same app. Log thermal zones, skin estimate, cooling device states and CPU and GPU frequency over time with Perfetto and dumpsys thermalservice. Identify the heat source (camera, encoder, modem in poor signal, display brightness) with power rails. Options: tune thermal thresholds, use hardware encoders, lower resolution or frame rate via thermal headroom APIs, or improve heat spreading.

The device shuts down at 15% battery in cold weather. What is the cause and fix?

Cold raises the battery's internal resistance, so current peaks cause voltage droop below the shutdown threshold even though charge remains. Also the gauge may not compensate well for temperature. Fixes: tune the gauge's temperature model, reserve more capacity at low temperatures (report 0% earlier), limit peak current when voltage is low (throttle CPU, delay radio bursts), and use hardware current limiting.

Charging is slower than specified. Where do you look?

Check the charger type detected and input current limit, the charge current actually being applied (dumpsys battery, power_supply sysfs), thermal limits (charge current cooling device active, JEITA zone because the battery is warm), adaptive charging holding at a level, a weak or incompatible charger, and whether heavy use during charging leaves little current for the battery.

An app wakes the device every minute. How do you prove it and fix it?

dumpsys alarm lists alarms by app with wakeup counts; batterystats shows wakeup alarms and wakelock tags per uid; Perfetto shows the process running after each wake. Fix by moving the work to WorkManager with constraints, using inexact alarms, lengthening the interval, or, if it is a platform component, batching with other periodic work. If a third-party app, its bucket and exact-alarm permission are levers.

After enabling a new sensor feature, standby current rose by 2 mA. How do you investigate?

dumpsys sensorservice shows which clients registered which sensors at what rate and latency. Check whether it uses a wake-up sensor with short latency, whether the hub is doing the processing or streaming raw data to the AP, and whether the co-processor rail itself grew (Perfetto rails). Move processing onto the hub, lengthen latency, or switch to event-based wake-up only on meaningful changes.

A watch's always-on display is using more power than the spec. What would you check?

Whether the AP suspends between ambient updates, how often the face or an app redraws (more than once a minute is wrong), ambient brightness and panel low-power mode and refresh, the fraction of lit pixels (OLED power scales with it), and any complications requesting frequent updates. Compare against a reference system face to separate face issues from platform issues. See Wear OS Platform.

Battery drain is fine in the lab but bad in the field for some users. What do you do?

Look at field telemetry distributions to find the segment: carrier, region, hardware revision, installed apps, companion phone model, usage pattern. Common gaps are poor cellular coverage, a specific third-party app, a Wi-Fi network with heavy multicast, or a Bluetooth reconnect loop with certain phones. Reproduce with that condition in the lab, then extend the test matrix to include it.

Suspend residency is good but average current is still high. Why?

The AP is not the main consumer. Check the display (AOD, brightness), the co-processor and sensors (high sampling rates, PPG LEDs), radios (cellular in poor signal, GNSS left on, Wi-Fi scanning), and a peripheral left powered during suspend (a regulator not turned off). Power rails in Perfetto or a lab monitor with rail breakdown identify which.

A game gets throttled after two minutes and frame rate halves. How would you improve it?

Confirm with thermal traces which zone trips and which cooling device acts. On the app side, use thermal headroom to lower quality or cap frame rate before throttling, reducing peak power. On the platform side, check whether governor tuning is too aggressive or cooling is applied in large steps, and whether a power allocator policy could give smoother sustained performance. Also check the device's heat spreading.

An OTA made boot time 5 seconds longer. How do you find why?

Compare boot traces and boot-time markers between builds: bootloader, kernel, init service start times, Zygote preload, system_server service start times and the first frame. Look for a new service blocking the boot, a slow driver probe, an added firmware load, or app compilation after the update. See Android Boot.

A power regression is found two days before release. How do you decide what to do?

Quantify the impact against the budget and the user-visible effect (for example hours of battery lost), identify the change and its owner, and assess fix options: revert, targeted fix, or configuration change, each with risk. If the fix is low risk, take it with extra verification; if not, weigh shipping with a documented waiver and a planned fix in the next update. Communicate the decision and data clearly to stakeholders.

Crash rate jumped after an integration drop, but power is fine. How do you gate and triage?

Stability gates should block the drop on the crash or ANR threshold regardless of power. Cluster crashes by signature and process, identify the top clusters, map them to the changes in the drop, and bisect if unclear. Assign owners, fix or revert, and add the reproducer to the test suite. Track crash-free rate and ANR rate as first-class KPIs alongside power.

A kernel wakeup source from the charger driver is active all the time. What might be wrong?

The driver may take a wakeup source on each charger or USB interrupt and not relax it, a floating detection pin may generate constant interrupts, or the driver may intentionally stay awake while it believes a charger is present because detection is wrong. Check the driver's interrupt counts, charger detection state and code paths, and fix the release logic or pin configuration.

After a Power HAL change, cold start is 200 ms slower and standby current rose. What happened?

Two different knobs often move together in one HAL change. Slower start: launch or interaction boost is shorter, weaker, or never applied, so the first frame runs on a little core at a low OPP. Higher standby: a boost or INTERACTIVE mode stays asserted after screen-off, or uclamp.min for system processes never returns to zero. Compare Perfetto around startActivity (frequency, uclamp, cpuset) and a 10-minute screen-off idle (boost leftovers, wakeup sources) between the two HALs.

A game hits 60 fps for a minute, then thermal-throttles hard. How would ADPF have helped?

Without ADPF the game (or the platform) often over-boosts, so junction and skin temperature race to the trip point and the governor slams frequency. With a hint session the game reports actual frame time and can hold the lowest OPP that still meets 16.6 ms; with thermal headroom it can drop to 45 or 30 fps before SEVERE. Platform-side, check that the Power HAL is not applying a long sustained-performance mode that fights the thermal policy.