Platform Integration & Release
A phone or watch ships as one software image built from Google's AOSP, the chip vendor's BSP and HALs, firmware for other processors, and OEM customisation. This page explains how that image is composed, how upstream releases are integrated into it, how branches, merges, compliance tests and quality gates keep it healthy, and how systemic issues and releases are driven to closure across distributed teams and customers.
- The HLOS image is the Android or Wear OS side (kernel plus user space); non-HLOS images are firmware for the modem, DSPs, TrustZone and boot chain. A meta build ties them together into one flashable, versioned release.
- Upstream integration lands each new AOSP tag onto vendor trees and per-chipset baselines while carrying vendor changes forward; Treble, VINTF and GKI make this tractable. AOSP source lives on android.googlesource.com; check current release notes for tag cadence.
- Day-to-day: Soong/
Android.bpandlunchproduct images, branching and merge strategy, then promotion gates (build, boot, smoke, CTS/VTS/GTS, power, stability, performance). OTAs ship throughupdate_enginewith A/B or virtual A/B. - Systemic issues are driven by one owner through reproduce, instrument, cluster, localise, fix, verify and prevent, with root-cause analysis and a regression gate at the end.
What an HLOS image is
HLOS stands for High-Level Operating System. On a Qualcomm-style SoC it means the rich OS running on the main application processor: the Linux kernel plus Android or Wear OS user space. It is contrasted with the non-HLOS software: firmware and real-time operating systems that run on the other processors inside the chip. The HLOS image is the full shippable Android build, and leading HLOS image activities means owning how it is composed, branched, integrated, validated and shipped for each chipset generation.
HLOS IMAGE = AOSP / Wear OS framework and system apps
+ vendor BSP: Linux kernel (GKI + vendor modules), device tree, drivers
+ vendor HALs (AIDL / HIDL): sensors, display, audio, radio, power, thermal ...
+ proprietary vendor libraries and HLOS-loaded firmware blobs
+ OEM customisation, overlays, configuration and branding
──▶ branched per SoC generation ──▶ validated against gates ──▶ shipped per OEM product
Think of a car. The HLOS is the infotainment and driver-facing software everyone sees; the non-HLOS firmware is the engine control unit, brake controller and airbag module, each with its own dedicated computer. A car only ships when all of them are the right versions and work together. In a device, the infotainment is Android on the application processor, the dedicated controllers are the modem, DSPs, TrustZone and boot firmware, and the "vehicle build sheet" that fixes all versions together is the meta build.
HLOS versus non-HLOS
| Component | Runs on | HLOS or non-HLOS | Typical image |
|---|---|---|---|
| Android / Wear OS user space | Application processor (Arm Cortex-A) | HLOS | system, system_ext, product, vendor, odm inside super |
| Linux kernel and ramdisk | Application processor | HLOS | boot, vendor_boot, init_boot, dtbo, vbmeta |
| Android bootloader (ABL) | Application processor, before the kernel | Usually delivered with HLOS | abl |
| Primary and secondary boot loaders (PBL in ROM, XBL) | Application processor at reset | Non-HLOS | xbl, xbl_config |
| Modem firmware (MPSS) | Modem DSP | Non-HLOS | modem image (NON-HLOS.bin style FAT image) |
| Audio and compute DSP firmware (ADSP, CDSP) | Hexagon DSPs | Non-HLOS | adsp, cdsp images |
| Sensor low-power island (SLPI) or sensor hub | Low-power DSP or MCU | Non-HLOS | sensor firmware image |
| TrustZone / trusted execution environment | Secure world of the application processor | Non-HLOS | tz, hyp, trusted apps |
| Always-on power and resource manager | Small power-management processor | Non-HLOS | aop or equivalent |
| Wi-Fi and Bluetooth firmware | Connectivity chip or subsystem | Non-HLOS (often loaded by HLOS drivers) | firmware blobs |
The meta build
Each non-HLOS component is built by its own team, on its own schedule and often with its own toolchain. A meta build is the integration record that pins one version of every component (HLOS build plus each firmware build) into a single tested combination, together with the partition layout and flashing instructions. Tools such as fastboot or vendor download tools flash the meta build as a unit. Integration bugs frequently come from mismatched combinations, for example a new HLOS driver expecting a firmware interface the old modem build does not have.
Where HLOS and non-HLOS meet
At HAL implementations, kernel drivers for remote processors (remoteproc and subsystem restart), shared memory, and message protocols such as QMI and glink. Changes on either side must keep these contracts stable.
Why wearables are special
Watches add sensor-hub firmware and an always-on co-processor as first-class components, and the image must fit tighter RAM, flash and power budgets.
Image variants
The same source produces user, userdebug and eng builds; only user builds ship and are certified, but most debugging happens on userdebug.
Build fingerprint
ro.build.fingerprint identifies brand, product, device, Android version, build ID, variant and signing keys; certification and OTA targeting depend on it.
AOSP upstream integration
AOSP source is published on android.googlesource.com. Google tags platform releases there; the public tag and branch cadence has changed over the years (monthly, quarterly and other rhythms have all existed), so treat any fixed public cadence as historical and check current Android release notes for the tags you must track. Separately, Android still publishes a monthly security bulletin; that is a patch list, not a promise that a full AOSP drop appears on a fixed calendar. A chip vendor tracks the tags it cares about and lands them onto its internal trees and per-chipset baselines, carrying its own changes (drivers, HALs, performance features, fixes) forward, so that OEMs receive current Android plus the vendor's value-add. Qualcomm calls its upstream-tracking program Keystone; other vendors run equivalent programs.
Google AOSP upstream (release tag N, security patches) │ track, plan the drop, read release notes and API/HAL changes ▼ Vendor internal trees ── merge or rebase upstream, carry vendor deltas forward │ conflicts? ──▶ dependency analysis ──▶ assign owner (framework / BSP / HAL / apps) ▼ Per-chipset baselines (SoC gen A, B, C) ── apply BSP deltas, bump HAL versions, update VINTF │ gates: build health ▸ boot ▸ smoke ▸ CTS/VTS/GTS ▸ power / stability / performance ▼ Promote to OEM-consumable baseline ── OEM branches, customisation ▼ OEM product image ──▶ certification ──▶ carrier / retail release ──▶ OTA
Upstream integration is like a restaurant chain adopting the head office's new seasonal menu while keeping each branch's local specialities. Every season the head office sends a new recipe book; each branch must merge it with its own additions, check nothing clashes, train staff and pass a hygiene inspection before serving customers. The head office is Google's AOSP, the branches are per-chipset baselines, local specialities are vendor deltas, and the hygiene inspection is the set of promotion gates and compliance suites.
What makes it tractable
| Mechanism | What it gives integration |
|---|---|
| Project Treble | Separates the framework (system) from vendor code (vendor, odm) behind stable HAL interfaces, so a framework update does not require rewriting vendor code. |
| VINTF | Device and framework manifests plus compatibility matrices declare which HAL versions each side provides and needs; mismatches are caught at build time and boot. |
| Stable AIDL HALs | Versioned, frozen interfaces; the framework can support several versions so vendor code upgrades independently. |
| GKI (Generic Kernel Image) | One Google-built kernel core per Android common kernel branch, with vendor code in loadable modules against a stable KMI (kernel module interface). Within the same frozen KMI, core-kernel updates do not require vendor modules to rebuild; a rebuild is needed when the KMI itself changes (new ACK branch or ABI) or a module uses a new symbol. |
| Mainline (APEX / APK modules) | Some system components update through Google Play system updates, which reduces what the vendor must integrate but adds version-compatibility checks. |
| Generic System Image (GSI) | A pure AOSP system image used to prove the vendor side follows Treble (VTS and CTS-on-GSI). |
Typical sources of integration conflicts
- HAL interface bumps: a new Android release requires a newer HAL version (for example a HIDL to AIDL migration); vendor HAL owners must implement it before the drop can land.
- Framework API or behaviour changes that vendor patches in framework code relied on, such as a refactored service or new permission checks.
- SELinux policy changes that deny vendor services previously allowed.
- Build system changes (new Soong modules or flags, leftover Make variables, product makefile churn). Google's userspace Bazel migration was halted around 2023; Soong remains the primary userspace build. Kleaf/Bazel for kernel builds is separate and still real.
- Kernel branch moves to a new Android common kernel (a new KMI). Modules must be rebuilt against that new KMI; ordinary same-KMI GKI updates do not force a module rebuild.
- Toolchain updates (new Clang version, stricter warnings treated as errors).
SoC bring-up and where the image comes from
Before an image can be integrated and stabilised, a new chip must be brought up layer by layer. An integration lead does not write every driver, but must understand the sequence to triage "stuck at boot" and readiness issues. Boot chain detail is in Android boot.
| Stage | What happens | Integration angle |
|---|---|---|
| Boot chain | Boot ROM (PBL) ▸ XBL/SBL (DDR, clocks) ▸ bootloader (ABL) ▸ kernel ▸ init ▸ framework | Know the chain to triage "stuck at boot" systemic issues and verified-boot failures |
| Kernel, device tree and drivers | Kernel, device tree, clocks, regulators, pin control, peripheral drivers | Consume BSP outputs into the image and track driver readiness per feature |
| HALs and framework | Vendor HALs (sensors, display, audio, connectivity, radio) wired to the framework | Own the framework-up integration; VINTF and SELinux must be right |
| Stabilisation | Power, thermal, stability and KPI tuning to the ship gate | Drive systemic issue triage to closure; this is where schedules slip |
Bringing up a new chip is like opening a new building: foundations and utilities first, then floors, then offices, then furniture and finishing. You cannot test the elevators before the power is on. In a device, the foundations are the boot chain and DDR, utilities are clocks and regulators, floors are the kernel and drivers, offices are the HALs and framework, and the finishing is power and performance tuning.
Branching models for a platform
A platform codebase is hundreds of Git repositories stitched together by a repo manifest, built for several chipsets and many OEM products at once. The branching model decides where changes land first and how they flow between streams.
aosp upstream (android-N release tags, security bulletins)
│ merge / rebase per drop
▼
vendor main / dev ───●────●────●────●────●──────────▶ (next Android version work)
│ │
│ branch cut │ branch cut
▼ ▼
release-N-socA ──●──●──▶ release-N-socB ──●──▶ (stabilise, cherry-picks only)
│
▼
oem-X-product ──●──▶ (customisation, customer fixes, OTA)
| Model | How it works | Good for | Risk |
|---|---|---|---|
| Trunk-based | Everyone commits small changes to one main branch; unfinished features hidden behind flags | Fast integration, fewer merges; Google's AOSP moved to "trunk stable" development with aconfig feature flags | Needs strong presubmit CI and flag discipline |
| Release branches | Cut a branch from main at feature freeze; only fixes go in after that | Stabilising a specific Android version or chipset while main moves on | Fixes must be ported to several branches |
| Feature branches | Long-lived branch per large feature, merged back later | Isolating risky, large work | Big, late, painful merges ("merge hell") |
| Per-SoC / per-OEM branches | Branch per chipset baseline and per customer product | Different BSPs, schedules and customer changes | Branch explosion; fixes drift between branches |
A branching model is like a river system. The main river (trunk) keeps flowing; at certain points canals are dug off it (release branches) to feed specific towns (chipsets and customers). Water can be pumped from the river into a canal (cherry-picking a fix), but once a canal is cut you must keep pumping, or the towns fall behind. In a platform, the pumping is porting fixes, and the discipline is keeping the number of canals small and deciding clearly which fixes go where.
Branch hygiene rules that scale
- Fix on the oldest supported branch that needs it, then merge forward, or fix on main and cherry-pick back, but pick one direction per organisation and track it.
- Freeze stages: feature freeze (no new features), code freeze or "lockdown" (only approved fixes), release candidate (only blocker fixes with sign-off).
- Change tracking: every change references a bug or change request; release branches accept only tracked, approved changes.
- Automated forward-merge checks detect fixes that landed on a release branch but not on main, which otherwise reappear as regressions next release.
- Retire branches on a published schedule; every live branch costs CI capacity and security-patch effort.
Merge, rebase and cherry-pick strategies
Bringing an upstream drop into a tree that carries vendor changes can be done in several ways. The right choice depends on history requirements, the size of the carried delta and how many teams work on the tree.
| Strategy | What happens | Pros | Cons |
|---|---|---|---|
| Merge | Create a merge commit joining upstream history with the vendor branch | Preserves exact history; one place to resolve conflicts; SHAs of existing commits unchanged, so downstream branches are not disrupted | History becomes hard to read; the carried delta is not visible as a clean patch set |
| Rebase | Replay every vendor commit on top of the new upstream | Linear history; the vendor delta stays a clean, reviewable patch stack; easier to upstream patches | Rewrites SHAs, breaking anyone who built on the old commits; conflicts resolved commit by commit |
| Cherry-pick | Copy selected individual commits onto another branch | Precise; ideal for porting fixes to release branches | Creates duplicate commits with different SHAs; easy to miss dependent commits |
| Squash | Collapse a series of commits into one | Tidy history for a feature | Loses granular history; makes bisecting and reverting harder |
Imagine updating a cookbook you have annotated. Merging is stapling the new edition to your annotated copy and writing a note on how they fit together. Rebasing is taking a fresh new edition and rewriting all your annotations into it, page by page. Cherry-picking is copying one specific annotation into a friend's copy. In Git terms, the annotations are the vendor delta, the edition is the upstream release, and the choice decides whether your history stays intact (merge) or your delta stays clean (rebase).
# sync a multi-repo platform tree
repo init -u <manifest-url> -b <branch> -m <manifest>.xml
repo sync -c -j8
# merge an upstream tag into a vendor branch (one project)
git fetch aosp android-15.0.0_r1
git merge --no-ff android-15.0.0_r1
# or rebase the vendor patch stack onto the new tag
git rebase --onto android-15.0.0_r1 android-14.0.0_r1 vendor-main
# port one fix to a release branch, recording the source commit
git cherry-pick -x <sha>
# remember how a conflict was resolved so repeated merges reuse it
git config rerere.enabled true
# find the change that introduced a regression
git bisect start <bad> <good>
git bisect run ./run_smoke_test.sh
Conflict resolution at scale
- Pre-scan Do a trial merge early (for example on beta or developer-preview tags) to list conflicting repositories and files before the real drop.
- Classify Group conflicts by domain (framework, BSP, HAL, apps, build) and by type (textual, semantic, API or HAL version change).
- Assign owners Each conflict goes to the team that owns the carried change, with a deadline; the integration lead owns the tracking list.
- Resolve with understanding Read both sides' intent: why upstream changed the code and why the vendor patch exists. Sometimes the right resolution is dropping the vendor patch because upstream fixed the same problem.
- Build and test the resolution A textual conflict fixed without compiling and testing is the most common source of silent regressions.
- Record Document non-obvious resolutions in the commit message and reuse them with
rerereon the next drop.
-x helps traceability, the difference between textual and semantic conflicts, and a process with owners and a tracking list rather than one hero resolving everything.Compliance: CTS, VTS, GTS and friends
Android devices must pass Google's compatibility test suites to be called Android-compatible and to ship Google Mobile Services (GMS). Together they are often called xTS. They run on the Trade Federation (Tradefed) test harness.
| Suite | Tests | Why it matters |
|---|---|---|
| CTS (Compatibility Test Suite) | Public Android APIs and behaviour required by the Compatibility Definition Document (CDD) | Apps behave the same on every compatible device |
| CTS Verifier | Manual or semi-automated tests for things automation cannot check (sensors, camera, audio, NFC) | Hardware-dependent behaviour |
| VTS (Vendor Test Suite) | HAL interfaces, kernel (including GKI and KMI requirements), VINTF, vendor partition behaviour | The Treble contract between framework and vendor holds |
| CTS-on-GSI | Runs CTS with a Generic System Image on the vendor's device | Proves the vendor side works with a pure AOSP framework |
| GTS (GMS Test Suite) | Google Mobile Services requirements (Play services, Play Store, Google apps behaviour) | Required under the GMS licence; not part of AOSP |
| STS (Security Test Suite) | Security patches listed in the monthly security bulletin | Validates the security patch level claimed by the build |
| Form-factor suites | Extra requirements for Wear OS, Android TV, Automotive and others | A watch must also pass its Wear OS-specific requirements |
xTS is like the safety and emissions tests a car must pass before it can be sold. CTS checks that the car behaves like any other car to the driver (pedals, indicators); VTS checks the internal parts fit the standard connectors; GTS is the extra inspection a brand partner requires before letting you use their logo. In Android, the driver is the app developer, the standard connectors are HAL and kernel interfaces, and the brand partner is Google's GMS licence.
Running xTS in practice
# run a full CTS plan with Tradefed
./android-cts/tools/cts-tradefed run cts -s <serial>
# run one module or test while debugging a failure
cts-tradefed run cts -m CtsSensorTestCases -t android.hardware.cts.SensorTest
# retry only the failures from a previous session
cts-tradefed run retry --retry <session-id>
- Test often, not only at the end: run subsets of xTS in daily CI and full suites on promotion candidates, so failures are bisectable to a small change range.
- Triage every failure as device bug, test bug, test environment problem (network, SIM, lab setup) or flaky test; only device bugs go to engineering teams.
- Waivers: known test bugs can be waived through Google's process, but a waiver needs evidence and is not a way to hide device bugs.
- Final submission uses the exact
userbuild with the release fingerprint; results are tied to that fingerprint.
GMS, MADA and the licence path
CTS and VTS prove Android compatibility. Shipping the Play Store, Play services and Google apps is a separate, licensed path:
- GMS (Google Mobile Services) is the licensed set of Google apps and services. It is not in AOSP.
- MADA (Mobile Application Distribution Agreement) is the commercial contract under which an OEM may preload GMS. Exact terms are confidential; in interviews, speak only at this level: licence, placement and update requirements, and a test bar.
- GTS is the technical test suite that licensees run against those requirements. Passing CTS + VTS + GTS (and related suites) is necessary but not sufficient; the licence and a formal submission still apply.
- CTS-Verifier covers behaviour that automation cannot fully judge: sensors, camera, audio, haptics, NFC, accessibility, and similar hardware-in-the-loop checks. Budget lab time and operators; it is often the long pole late in a program.
Carrier certification (high level)
Cellular devices also need operator and industry certification, which is independent of Google's xTS:
| Track | What it is | Integration angle |
|---|---|---|
| PTCRB | North-American operator certification forum for cellular devices | Plan lab time, SIM/eSIM variants, RF and protocol cases; failures often need modem plus HLOS owners |
| GCF | Global Certification Forum; similar industry bar used widely outside PTCRB operators | Same idea: a shared test catalogue, operator deltas on top |
| Operator acceptance | Each carrier's own lab and field cases (voice, data, SMS, IMS, emergency, roaming) | Track as a release gate with a named DRI; daily loop on blockers |
CarrierConfig | Per-carrier Android configuration (feature flags, timers, IMS and telephony behaviour) loaded by MCC/MNC or carrier app | Many "works on one SIM, fails on another" bugs are config, not modem silicon |
Keep employer and customer names out of answers. Describe the model: SoC vendor supplies the modem and RF package, OEM owns the product image and GMS path, carriers own their acceptance bar, and the integration lead keeps one status across all three.
CarrierConfig) is a separate gate. Bonus: CDD, CTS-on-GSI, STS and security patch level, and how you handle flaky tests and waivers.Promotion gates and quality criteria
A promotion gate is a set of pass or fail criteria that a build must meet before it moves to the next stage: from integration branch to baseline, from baseline to OEM delivery, from release candidate to shipping. Gates turn quality from an opinion into a measured decision.
| Gate | Checks | Why |
|---|---|---|
| Build health | All targets and variants build across SoC baselines; no new warnings-as-errors | No broken trees downstream |
| Boot | Boots to home screen on every SKU; boot success rate over many cycles | Catch catastrophic breakage immediately |
| VINTF and compatibility | Framework and vendor HAL versions compatible; VTS core passes | Treble contract must hold |
| Smoke / basic acceptance test (BAT) | Core UX, calls, data, Wi-Fi, Bluetooth, sensors, camera, OTA | Catch gross functional breakage early |
| Compliance | CTS, VTS, GTS pass rate at or above target; no new failures | Certification and GMS licence |
| Power regression | Standby drain, suspend residency, wake lock and wakeup budget, per-use-case current against last good | Ship blocker on wearables; see Power and thermal |
| Stability | Crash-free rate, ANR rate, kernel panics, watchdog resets, subsystem restarts, long-run monkey tests | Field quality |
| Performance | Boot time, app launch latency, jank percentage, memory footprint | Competitive user experience |
| Open defects | Zero open P0/P1 regressions; P2 counts under agreed limits with owners | Known risk is explicit and accepted |
Promotion gates are like airport security checkpoints: you pass check-in, then security, then passport control, then boarding, and each stage checks different things. Failing one stops you from boarding even if you passed the rest. In a platform, check-in is build health, security is boot and smoke, passport control is compliance, and boarding is the KPI and defect criteria signed off by the release owner.
Good gate criteria
- Measurable and automated where possible
- Compared to a known-good baseline, not absolute numbers only
- Owned: someone signs off each gate
- Stable thresholds that do not move under schedule pressure
- Published in advance so teams can pre-check
Bad gate criteria
- "Looks OK" manual judgement
- Thresholds so loose that nothing fails
- Flaky tests that fail randomly, training people to ignore failures
- Criteria changed after the build fails
- No link between the gate and a field-quality outcome
CI/CD for a platform
Platform CI is harder than app CI: a full Android build can take hours, tests need real devices, and hundreds of repositories must be consistent. The goal is still the same: find each breakage as close to the change that caused it as possible.
developer change ──▶ Gerrit review │ presubmit: build affected targets ▸ unit tests ▸ static analysis ▸ quick device smoke ▼ (Code-Review +2 and Verified +1) submit to branch │ postsubmit / continuous: full builds on all targets ▸ boot tests ▸ broader test tiers ▼ nightly / candidate build ──▶ full smoke ▸ xTS subsets ▸ power and performance runs ▸ stability soak ▼ promotion candidate ──▶ full xTS ▸ KPI comparison vs last good ▸ sign-off ──▶ baseline / release
Platform CI is like quality control in a car factory: some checks happen at each workstation (a bolt is torqued correctly), some at the end of the line (the car starts and drives), and some on a test track overnight (long durability runs). Catching a bad bolt at the workstation is cheap; finding it after delivery is expensive. Workstation checks are presubmit tests, end-of-line checks are postsubmit builds and boot tests, and the test track is nightly soak, xTS and KPI runs.
| Practice | Why it matters |
|---|---|
| Topic or multi-repo change submission | A change touching framework and vendor repos must submit atomically, or the tree is broken in between (Gerrit topics, repo-level "submit together"). |
| Build caching and incremental builds | Remote build caching and distributed execution cut presubmit time so developers get feedback in minutes, not hours. |
| Test tiers | Fast tests in presubmit; slower device tests postsubmit; full xTS and soak nightly or per candidate. |
| Device labs | Pools of real devices per SKU, with automated flashing and health checks; unhealthy devices are removed automatically. |
| Flaky test management | Detect flakiness statistically, quarantine flaky tests with an owner and deadline, and never let them block or hide real failures. |
| Automated bisection | When a postsubmit or nightly test fails, CI builds intermediate change sets to find the culprit and notifies its author. |
| Build breakage policy | "Revert first, ask later": a change that breaks the build is reverted immediately so others are not blocked. |
| Artifacts and traceability | Every build stores its manifest snapshot (repo manifest -r), images, symbols and test results, so any build can be reproduced and debugged. |
Triage and driving systemic issues to closure
A systemic issue is a problem that spans several layers or teams, or keeps recurring: battery drain after an upstream drop, random reboots across SKUs, call drops on one carrier, a boot-time regression. No single team owns it by default, so it needs a driver. This is the core skill of a platform integration lead.
1. OWN IT ──▶ one directly responsible individual (DRI); stop ownership ping-pong
2. REPRODUCE ──▶ fixed test profile; scope severity, affected SKUs/builds, KPI impact
3. INSTRUMENT ──▶ logs across layers (app, framework, HAL, kernel, firmware) on one timeline:
bug reports, logcat, dumpsys, Perfetto, batterystats, dmesg, modem logs
4. LOCALISE ──▶ bisect builds to the introducing change; bisect layers to the failing one
5. DRIVE ──▶ one triage thread with the right experts; named owners, dates, cadence
6. DECIDE ──▶ fix vs risk vs schedule; communicate to internal teams and customers
7. CLOSE ──▶ land fix ▸ verify on the original repro ▸ regression test/gate ▸ post-mortem
Driving a systemic issue is like an emergency-room doctor leading a patient's care. One doctor is responsible, orders tests from several departments, reads all the results together, decides the treatment and does not discharge the patient until follow-up care is arranged. In platform work, the doctor is the DRI, departments are domain teams (kernel, HAL, framework, modem), the tests are traces and logs across layers, and the follow-up care is the regression gate and post-mortem.
Deciding who owns a cross-domain bug
- Triage by layer and by evidence: reproduce, capture logs across boot, kernel, HAL and framework, bisect against the introducing change.
- Map the failing component to its owning domain. The owner is the team whose code must change, not the team that first saw the symptom.
- If it is truly shared (for example a contract between a HAL and firmware), assign one DRI and have the others co-own actions.
- Remove ambiguity fast, so engineers spend time fixing instead of arguing about ownership.
Root-cause analysis
Root-cause analysis (RCA) finds why the problem happened, not only what broke, so that the fix prevents recurrence. Two common tools:
5 Whys
Ask "why?" repeatedly (about five times) until you reach a process or design cause. Each answer must be backed by evidence, and there may be several branches.
Fishbone (Ishikawa)
List possible causes by category (code, configuration, hardware, test coverage, process, tools) to avoid fixating on the first theory.
Timeline
Reconstruct when the defect was introduced, when it could have been detected, and when it was detected. The gap shows which gate was missing.
Blameless post-mortem
Focus on systems and processes, not people. Output: root cause, contributing factors, detection gap, and dated corrective actions with owners.
A worked 5 Whys example
- Problem Watch standby battery drain doubled after the latest upstream drop.
- Why 1 The AP is waking about 40 times a minute during standby (batterystats, wakeup sources).
- Why 2 A system service registered an accelerometer listener without batching.
- Why 3 The upstream drop changed a default, and a vendor patch that set the report latency no longer applied, although the merge had no conflict.
- Why 4 The semantic change was not caught because no test checks sensor batching after a merge.
- Why 5 The promotion gate measured active-use power but not overnight standby drain.
- Corrective actions Fix the report latency; add a sensor-batching check to VTS-style vendor tests; add standby drain per hour as a mandatory power gate with a threshold against last good.
| Severity | Typical definition | Response |
|---|---|---|
| P0 / blocker | Device does not boot, data loss, security hole, emergency calls broken, widespread crash | Stop the line; daily or continuous triage; blocks promotion and release |
| P1 / critical | Major feature broken or large KPI regression on many devices | Blocks promotion unless an explicit waiver is signed off |
| P2 / major | Feature degraded, workaround exists, limited scope | Fix in the current release if possible; tracked against limits |
| P3 / minor | Cosmetic or rare | Backlog, prioritised with other work |
Crash clustering and triage at scale
Field and lab stability work is not one bug at a time. At fleet scale you cluster first, then spend people only on the clusters that move the crash-free rate.
- Collect Tombstones (
/data/tombstones), ANR traces, DropBox, kernel panic/pstore, watchdog and subsystem-restart reasons, plus a build fingerprint and a small set of device properties. Bug reports for high-value clusters. - Symbolize Native stacks need matching symbols for that exact build; a missing
symbolstree makes every cluster look unique. - Cluster Group by crashing process, faulting frame (or a stable stack signature), exception type, and often build or SoC. Ignore absolute addresses; they change with ASLR.
- Rank Volume × user impact (system_server and surfaceflinger beat a rare app). Watch new clusters after an OTA before chasing old long-tail noise.
- Own Assign a DRI per top cluster; the owner is the team whose code must change. Attach a representative report, not fifty duplicates.
- Close the loop A cluster is done when the signature disappears (or drops below a threshold) on the next build and a regression test or gate exists.
Release management
Release management turns a stream of integrated builds into dated, supported releases for customers, with predictable quality. For a platform it covers Android version releases, maintenance releases, monthly security updates and OTAs.
Plan ──▶ Develop / integrate ──▶ Feature freeze ──▶ Stabilise ──▶ Code freeze
──▶ Release candidate(s) ──▶ Certification (xTS, carrier, regulatory) ──▶ Release
──▶ Staged OTA rollout (1% ▸ 10% ▸ 50% ▸ 100%) ──▶ Maintenance: SMRs, hotfixes ──▶ End of support
Release management is like running a railway timetable. Trains (releases) leave at published times; passengers (features and fixes) that are not on the platform before the doors close wait for the next train rather than delaying everyone. Maintenance crews keep old lines running safely until they are retired. On a platform, the doors closing is the freeze, the next train is the next release or maintenance update, and maintenance crews deliver security patches and hotfixes on supported branches.
| Milestone | Entry criteria | What changes are allowed |
|---|---|---|
| Feature complete / feature freeze | All planned features merged behind flags or enabled | Bug fixes; no new features |
| Code freeze | Promotion gates green on the stabilisation branch | Approved fixes only, reviewed by a change control board |
| Release candidate | No open blockers; KPIs within targets | Blocker fixes only; every change forces a new RC and re-run of key tests |
| Release / golden build | Certification passed on the exact build; sign-offs complete | None; further fixes go to a maintenance release |
| Maintenance | Release shipped | Security patches (monthly security maintenance releases), critical field fixes |
Key release practices
- Change control board (CCB): after code freeze, a small group approves each change by weighing benefit against regression risk.
- Go / no-go meeting: owners of each gate report status against published criteria; the release owner decides and records the decision.
- Release notes: what changed, known issues, fixed issues, security patch level, required firmware versions.
- Staged rollout and monitoring: ship OTA to a small percentage first, watch crash, ANR, OTA success and battery telemetry, then widen or halt.
- Hotfix path: a fast, pre-agreed path for critical field fixes: fix on the release branch, minimal testing matrix, forward-port to main.
- Security patch cadence: Android publishes a monthly security bulletin; OEMs and vendors integrate patches and set
ro.build.version.security_patch, verified by STS. That bulletin cadence is not the same as a public AOSP source-drop cadence.
OTA mechanics: A/B, virtual A/B and update_engine
Shipping the golden build is only half of release management. Devices in the field take an OTA through update_engine, which downloads a signed payload, writes it safely, and reboots into the new slot.
| Idea | What it means |
|---|---|
| A/B (seamless) updates | Two slots (A and B). The device runs one slot while update_engine writes the other. A reboot flips ro.boot.slot_suffix. If the new slot fails to boot, the bootloader rolls back to the old slot. |
| Virtual A/B | Dynamic partitions in super are snapshotted (copy-on-write) instead of storing a full second copy of every large partition. Saves flash, but needs free space for snapshots; low-storage devices are a classic OTA failure class. |
update_engine | The on-device daemon that fetches metadata, verifies signatures, applies the payload and reports success or failure. Sideload and recovery paths still exist for lab and brick recovery. |
| Full versus incremental | A full payload can install from a wide set of source builds. An incremental (delta) payload is smaller but is built for a specific source fingerprint. A missed source build in the delta set is a common "OTA not offered" bug. |
# which slot is running, and the build you will target OTAs at
adb shell getprop ro.boot.slot_suffix
adb shell getprop ro.build.fingerprint
# watch apply progress and errors (service name varies by build; log is the reliable tap)
adb logcat -s update_engine update_engine_client
Staged rollout (1% → 10% → 50% → 100%) sits on top of this: halt on boot failure, crash-rate or OTA-success regressions, then ship a fixed payload. A/B rollback protects devices that fail to boot; it does not protect users from a boot-loop they already experienced, so pause the cohort.
Delivering with distributed teams and customers
Platform programs span several sites and time zones, many domain teams that the integration lead does not manage directly, and external customers such as OEMs and carriers. Delivery depends on influence, clarity and a steady operating rhythm.
Running a distributed program is like an air-traffic control network. Each tower (site) handles its own airspace, but flights (work items) cross between them, so handovers must be precise, written and standardised. Nobody shouts across the country; everyone reads the same flight plan. In a program, the flight plan is the single status of record, handovers are written follow-the-sun notes, and the controllers are DRIs with clear ownership.
Async-first communication
One source of truth for status, decisions written down, clear DRIs. Meetings are for decisions and unblocking, not status reading.
Follow-the-sun triage
Handoff notes at the end of each site's day (what was tried, what is next, what is blocked) so critical issues progress around the clock.
Influence without authority
Domain teams (camera, display, modem) do not report to the integration lead. Influence comes from owning the timeline, the dependency map and fair, specific asks.
Escalation
Escalate early with a clear ask, owner and date, not a status dump. Bad news travels the same day.
OEM, ODM and SoC: who owns what
A shippable device is a three-party (sometimes four-party) software stack. Interviewers for integration leads expect you to place yourself in this model without naming your employer.
| Party | Typical ownership | What they wait on |
|---|---|---|
| SoC vendor | Silicon, BSP, kernel modules, vendor HALs, modem and DSP firmware, per-chipset baselines | AOSP tags, HAL and VINTF changes, GKI/KMI moves, vendor API freeze windows |
| ODM | Board, RF, sensors, factory image, sometimes a thin software overlay | A bring-up-ready BSP and a meta-build recipe that matches the hardware SKU |
| OEM | Brand, product apps and UX, GMS/MADA path, carrier SKUs, OTAs, customer support | A stable vendor baseline, known-issue lists, merge windows, honest dates |
| Carriers / labs | PTCRB/GCF and operator acceptance, field issues | A certifiable user build and a DRI who answers with logs |
The integration lead rarely manages these organisations. Influence comes from one status of record, a dependency map, and fair, specific asks. When a bug bounces, name the contract that failed (HAL, KMI, QMI, CarrierConfig, meta-build pin) rather than the company.
Working with internal and external customers
| Stakeholder | What they care about | How to serve them |
|---|---|---|
| Program and product management | One status, dates, risks | Weekly red/amber/green against milestones and KPIs; exceptions highlighted |
| Domain engineering teams | Clear priorities, clean bug reports, less ping-pong | Triaged bugs with logs and a localised layer; named owners; fair deadlines |
| OEM customers | Stable baselines, fast answers on field issues, merge windows | Defined delivery cadence, a single point of contact, change-request tracking, transparent known-issue lists |
| ODM and factory | A flashable combination that matches the SKU | Meta-build pins, partition layout, factory vs user variants |
| Carriers and certification labs | Certification pass, field issues, honest dates | Daily loop on blockers during certification, logs rather than promises |
- Customer change requests (CRs): track every customer-reported issue with severity, owner, target release and status; review aging CRs weekly.
- Expectation management: never promise a date you do not control; give a date for the next update instead.
- Status that executives read: what changed since last week, what is blocked, and what you need from them, in a few lines.
Metrics and KPIs
Two families of metrics matter: product KPIs (how good the device is) and delivery and integration metrics (how healthy the process is). A lead tracks both and makes regressions visible early.
Product KPIs for a wearable
Power
Battery life in days, standby drain per hour, active-use drain, wake locks and wakeups, suspend residency, always-on display draw, sensor-hub offload ratio, current (mA) budget per use case.
Stability
Crash and ANR rate, kernel panics, watchdog and subsystem restarts, boot success rate, OTA success rate, field crash-free percentage.
Performance and UX
Boot time, app and tile launch latency, jank and frame drops, touch and wrist-raise responsiveness, wake-to-render latency.
Connectivity
Bluetooth reconnection time and drop rate, Wi-Fi and LTE reliability, call drop rate, notification delivery latency.
Integration and delivery metrics
| Metric | What it shows |
|---|---|
| Upstream integration lead time | Days from an AOSP tag release to a promoted baseline containing it |
| Carried delta size | Number of vendor patches on top of upstream; lower is cheaper to maintain |
| Conflict count and resolution time per drop | How painful each drop is; trend should go down |
| Build green percentage and time to fix breakage | CI health and team discipline |
| Test pass rate and flaky-test rate | Signal quality of the test system |
| xTS pass rate per baseline | Certification readiness over time |
| Defect inflow and closure rate, open P0/P1 count | Whether the release is converging |
| Mean time to resolve (MTTR) and CR aging | Responsiveness to customers and systemic issues |
| Escaped defects | Bugs found by customers or in the field that gates should have caught; each one should lead to a new gate |
| On-time delivery | Percentage of baselines and releases delivered on the committed date with the quality bar met |
Metrics are the dashboard of a car: the speedometer shows how fast you are going (delivery speed), the fuel gauge shows what is left (schedule and capacity), and warning lights show problems before the engine fails (regressions). Driving without them is guessing. On a platform, the speedometer is integration lead time, the fuel gauge is open defects against time to release, and the warning lights are KPI regressions against the last good build.
Build system: Soong, Make, lunch and product packages
An integration lead does not have to write every Android.bp, but must know how a product image is selected, what gets packaged, and which build system owns userspace versus the kernel. Wrong claims here are a common interview fail.
Soong (Android.bp)
- Primary userspace build for modules: binaries, libraries, apps, HALs, APEX
- Declarative Blueprint modules (
cc_binary,cc_library,android_app,aidl_interface) - Faster incremental analysis than Make; generates Ninja
- Google's attempt to migrate this userspace graph to Bazel was halted around 2023; Soong remains primary
Make and product config
- Product and board configuration still live in Make:
device.mk,AndroidProducts.mk,BoardConfig.mk - Legacy
Android.mkmodules still exist; new modules should be Soong PRODUCT_PACKAGES(andPRODUCT_PACKAGES_DEBUG) is how a product decides what is on the image- Do not say "Android is moving to Bazel" as a current userspace fact
Kleaf is the Bazel-based build for the Android kernel / GKI. That path is real and separate from the halted userspace Bazel migration. In an interview, split the two: Soong for userspace modules, Kleaf/Bazel for kernel images and modules, Make for product composition.
# select a product, a release configuration and a variant
# modern lunch: PRODUCT-RELEASE-VARIANT
lunch aosp_cf_x86_64_phone-trunk_staging-userdebug
# older trees used PRODUCT-VARIANT (two tokens), for example aosp_arm64-userdebug
# always check the tree you are in; do not memorise one combo as eternal
# what the product actually ships
# device.mk / *.mk
PRODUCT_PACKAGES += \
FooService \
vendor.bar.hal-service
PRODUCT_PACKAGES_DEBUG += \
FooDebugTool
| Token | Meaning |
|---|---|
| Product | The device or lunch target: board, partitions, overlays, PRODUCT_PACKAGES (for example aosp_cf_x86_64_phone). |
| Release config | Which aconfig / release set is enabled (trunk, a named Android release). This is the third token on recent trees. |
| Variant | user (ship and certify), userdebug (debuggable, usual day-to-day), eng (fastest iterate, not a ship image). |
// Android.bp — a vendor HAL binary
cc_binary {
name: "android.hardware.foo-service",
srcs: ["main.cpp"],
shared_libs: ["libbinder_ndk", "android.hardware.foo-V1-ndk"],
vendor: true,
init_rc: ["android.hardware.foo-service.rc"],
relative_install_path: "hw",
}
Soong modules are the recipes for each dish; PRODUCT_PACKAGES is the menu for tonight's restaurant (the product); lunch is seating you in a specific restaurant, on a specific seasonal menu, in kitchen or dining-room mode (the variant). Bazel/Kleaf is a different kitchen used for the kernel, not a replacement of the userspace kitchen.
PRODUCT_PACKAGES it will not be on the image.lunch (product-release-variant), PRODUCT_PACKAGES, Soong versus leftover Make, and where kernel builds sit (Kleaf). Then say who owns a missing binary: the module's Android.bp, or the product makefile that forgot to package it.OS upgrades, vendor API level and freeze windows
A major Android upgrade is not "merge AOSP and ship." Treble already lets a new framework talk to an older vendor implementation if the declared HAL and VINTF versions still match. Android later made that contract more explicit with a vendor API level and a freeze window (often discussed as a GRF-style Google requirements freeze): vendor code can stay on a frozen vendor API while the system image moves forward for a documented number of releases. Exact year counts and brand names of the program change; in an interview, state the mechanism and say you would check the current CDD and vendor-API release notes rather than quoting a memorised policy year.
system / product image ── newer Android, newer SDK, Mainline modules
│
│ VINTF + vendor API level (ro.vendor.api_level)
│ freeze window: vendor may stay put
▼
vendor / odm image ── HALs, kernel modules, sepolicy, firmware pins
│
│ KMI freeze on this ACK branch
▼
GKI + vendor modules ── same KMI: no module rebuild; new ACK: rebuild
- Vendor API level is the API the vendor partition was built against. It can lag the system API level inside the freeze window. Mismatches outside that window fail VINTF / OTA compatibility checks.
- Why it exists: OEMs want a yearly OS upgrade without a full SoC vendor rebase of every HAL and driver. The SoC vendor wants a stable contract so one vendor image supports several system years.
- What still moves on an upgrade: framework behaviour, CDD requirements, CTS/VTS/GTS for the new release, SELinux, Mainline modules, and any HAL the new compatibility matrix newly requires.
- What can stay: vendor HAL implementations, kernel modules on the same KMI, and firmware pins, if the freeze and VINTF allow it and tests still pass.
Think of a rented shop (the vendor image) inside a mall that renovates every year (the system image). A freeze window is a lease that says the shopfront sockets stay the same for N renovations, so the tenant does not have to rebuild the shop each year. When the mall changes the sockets (new vendor API or new KMI), the tenant rebuilds. VINTF is the inspector who checks the plugs before opening day.
Quick revision
- HLOS is the high-level OS on the application processor: Linux kernel plus Android or Wear OS user space.
- Non-HLOS is firmware on other processors: modem (MPSS), ADSP and CDSP, sensor hub (SLPI), TrustZone, XBL boot loaders, power management processor.
- The HLOS image is AOSP framework plus vendor BSP (kernel, device tree, drivers), vendor HALs, proprietary libraries and OEM customisation.
- A meta build pins one tested version of every HLOS and non-HLOS component into a flashable release.
- HLOS and non-HLOS meet at HALs, remote-processor drivers, shared memory and protocols such as QMI.
- AOSP source is published on android.googlesource.com; public tag cadence has changed over the years — check current release notes. The monthly security bulletin is a separate patch list.
- Upstream integration lands each AOSP tag onto vendor trees and per-chipset baselines while carrying vendor deltas forward.
- Treble, VINTF, stable AIDL HALs and GKI are what make frequent upstream integration feasible.
- Within the same frozen KMI, GKI kernel updates do not require vendor modules to rebuild; a new ACK/KMI does.
- Soong +
Android.bpis the primary userspace build; Make still composes products (PRODUCT_PACKAGES). Userspace Bazel migration was halted ~2023; Kleaf/Bazel is for kernel builds. lunchselects product-release-variant;userships,userdebugis the usual debug image. A built module is not on the image unless it is inPRODUCT_PACKAGES.- Common conflict sources: HAL version bumps, framework API changes, SELinux policy, build system changes, kernel branch (KMI) moves, toolchain updates.
- Upstream-first reduces the carried delta and the cost of every future drop.
- Branching models: trunk-based with feature flags, release branches, feature branches, per-SoC and per-OEM branches.
- Freeze stages: feature freeze, code freeze, release candidate, golden build.
- Merge preserves history and SHAs; rebase keeps a clean, linear patch stack but rewrites SHAs.
- Cherry-pick ports individual fixes; use
-xto record the source commit. - Semantic conflicts apply cleanly but are wrong; only builds, tests and informed review catch them.
git rererereuses conflict resolutions;git bisectfinds the introducing change.- CTS tests app-facing APIs against the CDD; VTS tests HALs, kernel and vendor interfaces; GTS tests GMS requirements; CTS-Verifier covers manual hardware checks; STS tests security patches.
- GMS is licensed (MADA); GTS is the licensee test suite. Compatibility tests are not the same as a GMS licence.
- Carrier cert (PTCRB/GCF plus operator labs) is independent of xTS; many per-SIM differences are
CarrierConfig. - CTS-on-GSI proves the vendor side works with a generic AOSP system image.
- Vendor API level plus a GRF-style freeze window can let a vendor image pair with newer system images; confirm current CDD/release notes rather than memorising a year count.
- OTA:
update_engineapplies a signed payload to the inactive A/B slot (virtual A/B snapshots dynamic partitions). Full payloads are source-flexible; incremental payloads are smaller but source-specific. - SoC / ODM / OEM / carrier own different layers; the integration lead owns the contracts and the one status of record, not the org chart.
- Crash triage at scale: symbolize, cluster by stable stack signature, rank by volume × impact, one DRI per top cluster, done when the signature is gone.
- xTS runs on Tradefed; run subsets continuously and full suites on promotion candidates.
- Promotion gates: build health, boot, VINTF, smoke, compliance, power, stability, performance, open defects.
- Gates must be measurable, compared against last good, owned, and stable under schedule pressure.
- Platform CI: presubmit build and quick tests, postsubmit full builds, nightly xTS, KPI and soak runs.
- Keep bisection possible: stage large drops, store manifest snapshots, automate culprit finding.
- Systemic triage: own it, reproduce, instrument, localise, drive, decide, close.
- The owner of a bug is the team whose code must change, not the team that saw the symptom.
- RCA tools: 5 Whys, fishbone diagrams, defect timelines, blameless post-mortems with dated actions.
- An issue is closed only after verification, a regression test or gate, and tracked corrective actions.
- Release practices: change control board, go/no-go against published criteria, release notes, staged OTA rollout, hotfix path.
- Scope, date and quality: with fixed people, make the trade-off explicit.
- Distributed delivery: async-first, single status of record, DRIs, follow-the-sun handoffs, early escalation.
- Track product KPIs (power, stability, performance, connectivity) and process metrics (lead time, build health, escaped defects, MTTR).
Glossary
- A/B update
- Seamless OTA that writes the inactive slot and reboots into it, with bootloader rollback if the new slot fails.
- ADSP / CDSP
- Qualcomm Hexagon digital signal processors for audio and compute workloads; they run non-HLOS firmware.
- Android.bp
- Soong module definition file for userspace binaries, libraries, apps and interfaces.
- AOSP
- Android Open Source Project; Google's open-source Android code base, published on android.googlesource.com.
- Baseline
- A promoted, tested snapshot of the platform for a chipset that downstream teams and customers build on.
- BSP
- Board Support Package; kernel, device tree, drivers and low-level software for a specific SoC and board.
- CarrierConfig
- Per-carrier Android configuration (feature flags and values) keyed by MCC/MNC or a carrier app.
- CCB
- Change Control Board; group that approves changes after code freeze.
- CDD
- Compatibility Definition Document; the written requirements an Android-compatible device must meet.
- Cherry-pick
- Applying a single commit from one branch onto another.
- Code freeze
- Milestone after which only approved fixes may enter a release branch.
- CR
- Change Request; a tracked issue or request, often from a customer.
- CTS
- Compatibility Test Suite; tests public Android APIs and behaviour.
- CTS-Verifier
- Mostly manual CTS tests for hardware-dependent behaviour that automation cannot fully judge.
- DRI
- Directly Responsible Individual; the single person accountable for driving an item to closure.
- Escaped defect
- A bug found by customers or in the field that internal gates missed.
- Feature flag
- A runtime or build-time switch that hides unfinished features so code can merge early (aconfig in AOSP).
- GCF
- Global Certification Forum; industry cellular-device certification catalogue used widely outside PTCRB.
- Gerrit
- Web-based code review system used by AOSP and most platform teams.
- GKI
- Generic Kernel Image; Google-built kernel core with vendor code in modules against a stable kernel module interface. Same-KMI updates do not require those modules to rebuild.
- GMS
- Google Mobile Services; licensed Google apps and services such as Play. Not part of AOSP.
- Golden build
- The exact build approved for release and certification.
- GRF
- Google Requirements Freeze (also discussed as a vendor-API freeze window): vendor implementations may stay on a frozen vendor API while the system image moves forward for a documented period. Check current release notes for the window length.
- GSI
- Generic System Image; pure AOSP system image used to validate Treble compliance.
- GTS
- GMS Test Suite; tests requirements for shipping Google Mobile Services.
- HLOS
- High-Level Operating System; the Linux kernel plus Android or Wear OS on the application processor.
- Keystone
- Qualcomm's name for its program that tracks AOSP upstream releases onto its chipset baselines.
- Kleaf
- Bazel-based build for the Android kernel / GKI; separate from the halted userspace Bazel migration.
- KMI
- Kernel Module Interface; the stable set of kernel symbols vendor modules may use with GKI. Frozen for the life of an ACK branch.
- lunch
- The AOSP command that selects a product, release configuration and build variant before compiling.
- MADA
- Mobile Application Distribution Agreement; the commercial contract under which an OEM may preload GMS.
- Merge
- Joining two branches' histories with a merge commit.
- Meta build
- A versioned bundle pinning one tested build of every HLOS and non-HLOS component.
- MPSS
- Modem Processor Subsystem; the modem's firmware and processor.
- MTTR
- Mean Time To Resolve (or Repair); average time from report to fix.
- Non-HLOS
- Firmware and RTOS images for processors other than the application processor.
- ODM
- Original Design Manufacturer; typically owns the board, RF and factory image, sometimes a thin software overlay.
- OEM
- Original Equipment Manufacturer; owns the brand, product software, GMS path, carrier SKUs and OTAs.
- Post-mortem
- A blameless review after an incident that records root cause and corrective actions.
- Presubmit
- Checks run on a change before it is merged.
- PRODUCT_PACKAGES
- Make variable listing modules installed on a product image. Built is not the same as packaged.
- Promotion gate
- Pass or fail criteria a build must meet to move to the next stage.
- PTCRB
- North-American operator forum for cellular device certification.
- RCA
- Root-Cause Analysis; finding why a problem happened, not only what failed.
- Rebase
- Replaying commits on top of a new base, producing a linear history with new commit IDs.
- repo
- Google's tool for managing many Git repositories from a single manifest.
- rerere
- Git's "reuse recorded resolution" feature, which re-applies earlier conflict resolutions.
- SMR
- Security Maintenance Release; an OEM or vendor update that delivers security bulletin patches. Cadence is a product choice; the Android security bulletin itself is monthly.
- Soong
- Primary AOSP userspace build system; modules are declared in
Android.bp. - STS
- Security Test Suite; checks that security bulletin patches are present.
- Tradefed
- Trade Federation; the Android test harness used to run xTS and device tests.
- Treble
- Android architecture that separates framework and vendor code behind stable interfaces.
- TrustZone
- Arm's secure-world technology; runs the trusted execution environment as non-HLOS firmware.
- update_engine
- On-device daemon that downloads, verifies and applies OTA payloads to the inactive slot.
- Upstream-first
- Policy of contributing changes to AOSP or Linux before, or instead of, carrying them privately.
- Vendor API level
- The API level the vendor partition was built against (
ro.vendor.api_level); may lag the system API inside a freeze window. - Virtual A/B
- A/B updates that snapshot dynamic partitions with copy-on-write instead of storing a full second copy of
super. - VINTF
- Vendor Interface object; manifests and compatibility matrices that declare and check HAL versions.
- VTS
- Vendor Test Suite; tests HALs, kernel and vendor interfaces.
- xTS
- Collective name for CTS, VTS, GTS, STS and related suites.
Interview questions
Fundamentals
What is an HLOS image?
HLOS means High-Level Operating System: the Linux kernel plus Android or Wear OS user space running on the main application processor. The HLOS image is the complete Android software build for a device: the AOSP framework and apps, the vendor BSP (kernel, device tree, drivers), vendor HALs, proprietary libraries and OEM customisation. It is distinct from the non-HLOS firmware that runs on the modem, DSPs and other processors.
What is non-HLOS software? Give examples.
Non-HLOS software is firmware or a real-time OS running on processors other than the application processor, or in the secure world. Examples: modem firmware (MPSS), audio and compute DSP firmware (ADSP, CDSP), sensor hub firmware (SLPI), TrustZone and the hypervisor, the XBL boot loaders, the always-on power management processor, and Wi-Fi and Bluetooth firmware. Each is built and versioned by its own team.
What is a meta build?
A meta build is the integration record that combines one specific, tested version of every software component (the HLOS build and each non-HLOS firmware build) into a single release, along with the partition layout and flashing information. It ensures that the combination flashed onto a device is known to work together. Many hard integration bugs are caused by mismatched component versions, which the meta build is designed to prevent.
What does upstream integration mean in the Android context?
It means taking a new release from the upstream source (Google's AOSP, or the Linux kernel) and bringing it into your own code base, which contains your changes on top. You merge or rebase the new upstream code, resolve conflicts with your carried changes, update interfaces, build and test, and promote the result as a new baseline for your products.
What is Project Treble and why does it help integration?
Treble, introduced in Android 8, separates the Android framework (system partition) from vendor implementation (vendor partition) behind stable, versioned HAL interfaces. The framework can then be updated without rewriting vendor code, as long as the HAL versions remain compatible. This makes upstream drops and OS upgrades much cheaper and is verified by VTS and by booting a Generic System Image.
What is VINTF?
VINTF (Vendor Interface object) is a set of manifests and compatibility matrices. The device manifest lists the HALs and versions the vendor provides; the framework compatibility matrix lists what the framework requires, and vice versa. The build and the OTA system check that they match, and the device refuses incompatible updates. A HAL version mismatch is a classic cross-team dependency during an upstream drop.
What is GKI?
GKI, the Generic Kernel Image, is a kernel core built by Google from the Android Common Kernel for each supported branch. Vendors put SoC and board-specific code into loadable kernel modules that use only the stable Kernel Module Interface (KMI). This separates kernel updates from vendor changes. Within the same frozen KMI, a core-kernel update does not require those modules to be rebuilt. A rebuild is needed when the KMI changes (new ACK branch) or a module needs a new symbol.
What is the difference between git merge and git rebase?
A merge combines two branches by creating a merge commit that has both histories as parents; existing commit IDs are unchanged. A rebase takes your commits and replays them one by one onto a new base, creating new commits with new IDs and a linear history. Merge is safer for shared branches; rebase gives a cleaner history and a tidy patch stack but rewrites history that others may depend on.
What is a cherry-pick and when do you use it?
A cherry-pick applies the changes of one specific commit onto another branch as a new commit. It is used to port a bug fix from main to a release branch (or the reverse) without bringing other changes. Use git cherry-pick -x to record the original commit ID in the message for traceability, and check for dependent commits that must also be picked.
What is the repo tool?
repo is Google's wrapper around Git that manages the hundreds of Git repositories that make up an Android tree. A manifest XML file lists each project, its path, remote and branch or revision. repo init selects a manifest, repo sync fetches all projects, and repo manifest -r produces a snapshot of exact revisions so a build can be reproduced.
What is CTS?
CTS, the Compatibility Test Suite, is Google's automated test suite for the public Android APIs and behaviours required by the Compatibility Definition Document. Passing it shows that apps written against the Android SDK will behave correctly on the device. Passing CTS is required to be Android-compatible and to license Google Mobile Services.
What is VTS?
VTS, the Vendor Test Suite, tests the vendor side of the Treble boundary: HAL implementations, VINTF compliance, kernel configuration and GKI requirements, and vendor partition behaviour. It ensures the vendor implementation keeps the contract the framework relies on, so a generic framework can run on the device.
What is GTS?
GTS, the GMS Test Suite, checks requirements for devices that ship Google Mobile Services, such as Google Play services and Play Store behaviour, preloaded Google apps and related configuration. It is distributed to GMS licensees, not published in AOSP. A device must pass CTS, VTS and GTS (plus others such as STS) to ship with GMS.
What is a promotion gate?
A promotion gate is a set of pass or fail criteria a build must meet to move to the next stage, for example from an integration branch to a baseline delivered to customers. Typical criteria are successful builds on all targets, boot on all SKUs, smoke tests, xTS pass rates, power, stability and performance within thresholds compared to the last good build, and no open P0 or P1 regressions.
What is the difference between presubmit and postsubmit testing?
Presubmit tests run on a change before it is merged, usually a targeted build and fast tests, so obviously broken changes never land. Postsubmit tests run after merging, on the combined tree, and include full builds on all targets and slower device tests. Presubmit protects the branch; postsubmit catches interactions between changes and anything too slow for presubmit.
What is root-cause analysis?
Root-cause analysis (RCA) is a structured investigation into why a problem happened, going past the immediate technical fault to the design, process or test gap that allowed it. Its output is the root cause, contributing factors, why it was not detected earlier, and corrective actions with owners and dates. Tools include 5 Whys, fishbone diagrams and timelines.
Explain the 5 Whys technique.
You state the problem and ask "why did this happen?", then ask "why?" about each answer, typically about five times, until you reach a cause that, if fixed, prevents the problem from recurring (often a process or test gap). Each answer should be supported by evidence, not guesses, and there can be multiple branches. It stops teams from fixing only the surface symptom.
What is a DRI?
A DRI (Directly Responsible Individual) is the one named person accountable for driving an issue or deliverable to completion. They do not have to do all the work, but they coordinate, track, escalate and make sure it closes. Having a DRI prevents problems from bouncing between teams with nobody driving them.
What is a code freeze?
A code freeze is a milestone after which only approved changes (usually bug fixes for release-blocking issues) may enter the release branch. It stabilises the code so testing results remain valid. Changes after freeze typically need a change control board's approval and may trigger re-running key tests.
What is Soong, and how does it differ from Make?
Soong is the primary AOSP userspace build. Modules are declared in Android.bp (Blueprint) and compiled via Ninja. Make still composes the product: device.mk, BoardConfig.mk and PRODUCT_PACKAGES decide what is on the image, and some leftover Android.mk modules remain. Google's userspace Bazel migration was halted around 2023; Soong is still primary. Kernel builds use Kleaf (Bazel), which is a different path.
What does lunch select?
The build target. Recent trees use PRODUCT-RELEASE-VARIANT (for example a Cuttlefish phone, a trunk or named release config, and userdebug). Older trees used two tokens, PRODUCT-VARIANT. The variant is user (ship and certify), userdebug (usual debug image) or eng. Always check the tree you are in rather than memorising one combo.
What is PRODUCT_PACKAGES?
The Make list of modules installed on that product image. A module can build successfully and still be absent from the device if nobody added it to PRODUCT_PACKAGES (or PRODUCT_PACKAGES_DEBUG for debug-only tools). When a binary is "missing on the image," check the product makefile before rewriting the Android.bp.
What is GMS versus MADA?
GMS is the licensed Google apps and services (Play Store, Play services and related apps). MADA is the commercial agreement under which an OEM may preload them. GTS is the technical test suite for that licence. CTS/VTS prove Android compatibility; they do not grant GMS. Do not quote confidential placement rules; say you would check current licensee docs and the CDD.
What is the difference between a full OTA and an incremental OTA?
A full payload can apply from a wide set of source builds. An incremental (delta) payload is smaller and is built for a specific source fingerprint. update_engine writes the inactive A/B slot (virtual A/B uses snapshots for dynamic partitions). If a device's source build is not in the delta set, it will not be offered that incremental and needs a full payload or another delta.
Going deeper
Walk me through integrating a Google upstream release into a vendor chipset baseline.
- Track the AOSP release tag and read release notes for API, HAL, SELinux, build and kernel changes.
- Do an early trial merge (for example on preview tags) to estimate conflicts and dependencies.
- Plan the merge or rebase strategy per repository and sequence the drop across affected SoC baselines.
- Run dependency analysis: assign each conflict or breakage to its owning domain (framework, BSP, HAL, apps, build) with a date.
- Update HAL implementations and VINTF manifests for required interface versions.
- Gate on build health, boot, smoke, VINTF, xTS, and power, stability and performance against last good.
- Promote the baseline, publish release notes and known issues, and hand off to OEM teams on the agreed schedule.
Throughout, keep one status of record and coordinate multimedia, camera, connectivity, display, power and security stakeholders across sites.
What goes into composing an HLOS image for a wearable, and how does it differ from a phone?
The ingredients are the same: AOSP and Wear OS framework, vendor BSP, vendor HALs, proprietary libraries, sensor hub firmware and OEM customisation, branched per chipset. The differences are the constraints: smaller RAM, flash and display, a dominant power budget, an always-on co-processor handling sensors and ambient display, health sensors as first-class components, a companion and Bluetooth connectivity model (and eSIM or LTE for standalone watches), and tiles, complications and watch faces as the main user surface. The power architecture shapes what is included and how it is tuned to the ship gate.
When would you choose merge over rebase for an upstream drop, and vice versa?
Choose merge when the branch is shared by many teams and downstream branches, because a merge does not rewrite commit IDs and conflict resolution happens once. Choose rebase when you want to keep the vendor delta as a clean, reviewable patch stack on top of upstream (common for kernel trees or small repositories), which makes it easy to see and upstream your changes. Many organisations mix them: merge for large shared framework repositories, rebase for patch-stack style kernel or HAL repositories.
What is a semantic conflict and how do you catch it?
A semantic conflict is when two changes merge without any textual conflict, but the combined code is wrong: for example upstream changes a default value, renames a behaviour or moves a permission check, and a vendor patch that depended on the old behaviour still applies. Git cannot detect this. You catch it with builds, unit and integration tests, xTS, KPI regression runs, and reviewers who understand why each carried patch exists. Reviewing the upstream diff for areas touched by carried patches helps too.
How do you resolve hundreds of merge conflicts on a tight schedule?
Do not have one person resolve everything. Pre-scan conflicts early, classify them by domain and type, and assign each group to the team that owns the carried change, with deadlines and a tracking list. Resolve easy mechanical conflicts centrally, and send semantic or interface conflicts to domain experts. Build and test each resolution, reuse earlier resolutions with rerere, and drop vendor patches that upstream has made unnecessary. Report progress daily against the list.
Which branching model would you use for a platform serving several chipsets and OEMs?
A main development branch that receives upstream drops and new features, preferably with feature flags to keep it close to trunk-based. Release branches cut from main per Android version and chipset baseline at feature freeze, which accept only fixes. OEM or product branches for customer-specific customisation, kept as thin as possible. Fixes land on main and are cherry-picked to supported release branches (or the reverse, but consistently), with automated checks that nothing is missed. Branches are retired on a published schedule.
How do you make sure a fix on a release branch is not lost on main?
Require every fix to reference a bug, and track per-branch status on that bug. Use cherry-pick -x or Change-Id matching so tools can compare branches, and run an automated forward-merge or "missing fixes" report that lists changes present on a release branch but not on main. Review the report as part of the release checklist. Without this, the same bug reappears in the next release.
What promotion gates would you enforce before an upstream drop lands in the baseline?
- Build health across all SoC baselines and build variants.
- Boot success on all SKUs, including repeated boot cycles.
- VINTF compatibility and core VTS.
- Smoke or BAT covering calls, data, Wi-Fi, Bluetooth, sensors, display, camera, OTA.
- CTS, VTS and GTS pass rate at or above the previous baseline.
- Power: standby drain, suspend residency, wake lock and wakeup budget versus last good.
- Stability: crash, ANR, panic, watchdog and subsystem restart rates.
- Performance: boot time, launch latency, jank.
- No open P0 or P1 regressions without a signed-off waiver.
How do you handle a CTS failure found close to release?
First triage: is it a device bug, a test bug, a test environment problem (network, SIM, lab setup) or flakiness? Re-run the single module with Tradefed to confirm, and compare with a previous passing build to find the introducing change. If it is a device bug, assign it to the owning team as a blocker. If it is a genuine test bug, collect evidence and request a waiver through the official process. Either way, record it and add the module to continuous CI so it is caught earlier next time.
What is CTS-on-GSI and why run it?
CTS-on-GSI means flashing Google's Generic System Image (a pure AOSP system partition) on the vendor's device and running CTS. If the vendor implementation follows Treble correctly, the generic framework should work with the vendor partition. Failures show that the vendor side relies on system partition modifications or breaks the interface contract, which would make future framework updates expensive.
How would you set up CI for a large multi-repository platform?
Use Gerrit with presubmit that builds affected targets and runs fast tests and static analysis, with atomic submission for changes that span repositories (topics). Postsubmit builds all targets continuously and runs boot and smoke tests on real devices. Nightly or per-candidate builds run xTS subsets, power and performance benchmarks and stability soaks. Add remote build caching, automated culprit finding, flaky-test quarantine, a revert-first policy for breakages, and store manifest snapshots and artifacts for every build.
How do you deal with flaky tests in platform CI?
Measure flakiness automatically (for example a test that fails and then passes on retry without code changes). Quarantine flaky tests from blocking gates, but keep running them and assign an owner and a deadline to fix or delete them. Track the flaky rate as a metric. Never simply retry until green, because that hides real intermittent bugs such as race conditions, which are often real device defects.
How do you decide which team owns a cross-domain bug?
By layer and evidence. Reproduce the bug, capture logs across boot, kernel, HAL and framework on one timeline, and bisect builds to the introducing change. Map the failing component to the team whose code must change. If it is genuinely shared, such as an interface contract between a HAL and firmware, assign one DRI and have the other teams co-own actions rather than letting the bug bounce. The goal is to remove ambiguity quickly so engineers fix instead of argue.
A systemic issue is reported by an external customer late in the program. How do you drive it to closure?
Take ownership as the DRI and acknowledge the customer quickly with a time for the next update. Reproduce and scope the severity and KPI impact. Pull the right cross-domain experts into one triage thread, build a cross-layer log timeline, and bisect to root cause. Weigh fix versus risk versus schedule and agree the plan with the customer. Communicate on a fixed cadence to internal teams and the customer. Close with the verified fix, a regression test or gate, and a post-mortem.
What makes a good post-mortem?
It is blameless and factual: a timeline of when the defect was introduced, when it could have been detected, and when it was detected and fixed. It identifies the root cause and contributing factors (technical and process), the detection gap (which gate or test was missing), and specific corrective actions with owners and dates. It is shared widely, and the actions are tracked to completion rather than forgotten.
What happens at a go/no-go meeting?
Each gate owner reports status against the published release criteria: build, xTS, KPIs, open defects, certification, and customer or carrier sign-offs. Known risks and waivers are reviewed explicitly. The release owner makes the decision (go, no-go, or go with conditions) and records it with reasons. A good meeting is short because the data was prepared in advance; debates about criteria belong before the meeting, not in it.
How do you run a staged OTA rollout?
Release the update to a small percentage of devices first (for example 1 percent), then widen in steps (10, 50, 100 percent) if health metrics are good. Monitor OTA success rate, boot success, crash and ANR rates, battery telemetry and customer-reported issues against the previous build. Define halt criteria in advance, and be ready to pause the rollout and ship a fix. A/B updates with rollback make each step safer.
How do you keep a geographically distributed program on track?
Work async-first with one source of truth for status, clear DRIs and written decisions. Use follow-the-sun triage with structured handoff notes so critical issues move forward around the clock. Keep a regular status cadence with internal teams and external customers, highlighting what changed, what is blocked and what help is needed. Escalate blockers early with a clear ask, owner and date. Rotate meeting times so the same site is not always inconvenienced.
Which KPIs would you gate a wearable release on?
Battery life and standby drain per hour against the last good build, suspend residency and wakeup counts, wake-to-render latency, crash and ANR rates, kernel panic and watchdog reset rates, boot and OTA success rates, thermal ceiling under sustained load, and connectivity reliability (Bluetooth reconnection, notification delivery). Define regression thresholds in advance and hold promotion if any exceed them.
How do you ramp up quickly on an unfamiliar platform or domain?
Map the system and its interfaces first: the image composition, branches, build and test pipelines, and KPI dashboards. Find the two or three people who hold the most context and learn from them. Sit in triage meetings to learn the current problems. Land a small real change early to learn the pipeline end to end. Keep a list of unknowns and burn it down deliberately, and write down what you learn so the next person ramps faster.
Advanced
How do you sequence one upstream drop across several SoC generations?
Start with the chipset that is most representative and best staffed (often the newest, which will ship the release first) as the lead baseline. Resolve common framework and HAL conflicts there once, then apply the same resolutions to other baselines, handling only their BSP-specific deltas. Older chipsets may stay on an earlier kernel branch or HAL version, so check VINTF and GKI support per chipset. Stagger promotions so teams are not overloaded, and publish the schedule to OEMs.
How does GKI change kernel integration work for a vendor?
Before GKI, each vendor carried a large, forked kernel with thousands of patches, and every kernel update meant a painful forward-port. With GKI, the core kernel image comes from Google's Android Common Kernel, and vendor code lives in modules that may only use symbols in the KMI symbol list. Integration work shifts to keeping modules compatible with the frozen KMI, requesting new symbols through the upstream process, and validating with VTS kernel tests. Same-KMI security and bug-fix kernel updates can be taken without rebuilding vendor modules. Modules rebuild when you move to a new ACK/KMI, or when you need a new symbol or a KMI break (which ABI tooling should reject on a frozen branch).
What is a HAL interface bump, and how do you manage it during an upstream drop?
A new Android release may require a newer version of a HAL (or migration from HIDL to AIDL) to support new features, as defined in the framework compatibility matrix. The vendor HAL owner must implement the new version, update the device manifest, and pass VTS for it. As integration lead, identify these requirements from the release notes and compatibility matrix early, create tracked items per HAL with owners and dates, and decide whether the drop can land with the old version (if still allowed) while the new one is completed.
How do feature flags change branching and release strategy?
Feature flags let unfinished features merge into main early but remain disabled, so fewer long-lived feature branches are needed and integration happens continuously. AOSP's trunk-stable model uses aconfig flags with release configurations that decide which flags are enabled in each release. The costs are flag management discipline, testing both flag states where it matters, and removing old flags. For release management, flags allow disabling a risky feature late instead of reverting code.
How do you keep a large merge bisectable?
Avoid squashing an upstream drop into one change; keep upstream history so git bisect can walk individual commits. Land the drop in stages where possible (by repository or subsystem) with builds and tests between stages. Store repo manifest -r snapshots for every CI build so any build can be reproduced. When a regression appears, bisect first between CI builds (coarse) and then between commits within the suspect repositories (fine).
How do you design promotion gate thresholds that are strict but not noisy?
Base thresholds on the distribution of results from known-good builds, not on a single run: measure run-to-run variance and set limits outside normal noise (for example the mean plus a margin). Compare against last good on the same hardware and test setup. Require several runs for noisy KPIs such as power and performance. Separate hard blockers (boot, P0 regressions) from soft limits that require a documented waiver. Review thresholds periodically, but never lower them to pass a failing build.
How would you measure power regressions reliably in CI?
Use dedicated devices with external power monitors or on-device power rails, fixed test profiles (screen off standby, ambient mode, workout, music), controlled radio conditions (shielded boxes or fixed SIM and network), and consistent battery and thermal starting states. Run each profile several times, report mean and variance, and compare with last good. Automatically attach batterystats, wakeup sources and Perfetto traces to failures so they are debuggable. Make standby drain per hour a gate.
How do you handle a regression introduced by a carried vendor patch that upstream now conflicts with?
First ask whether the patch is still needed: upstream may have fixed the same problem differently, in which case drop the patch and verify the original issue stays fixed. If still needed, re-implement it against the new upstream code with the owning team, adding a test that captures the original intent. Consider upstreaming it so it stops being a carried delta. Record the decision in the commit message.
How do you balance quality gates against schedule pressure from customers?
Make the trade-off explicit and data-driven. Quantify the gate failure (which KPI, by how much, which users), lay out options (slip the date, ship with a documented known issue and a dated fix, disable the feature with a flag, take a targeted fix with focused re-test), and state the risk of each. Decide with stakeholders and record the decision. Some gates are non-negotiable, such as boot success, security, and life-critical paths like emergency calling; others can be waived with sign-off and a follow-up plan.
What is an escaped defect, and how do you use it to improve the process?
An escaped defect is a bug found by a customer or in the field that the internal gates should have caught. For each one, do an RCA focused on the detection gap: which test or gate was missing, why the existing tests did not cover it, and what the cheapest reliable way to catch it earlier is. Add that test or gate, and track the escaped defect rate as a process metric. Over time, the gate set becomes a record of lessons learned.
How do HLOS and non-HLOS version mismatches cause bugs, and how do you prevent them?
HLOS drivers and HALs talk to firmware through message protocols and shared memory layouts. If a new HLOS build expects a new firmware message or field that an older modem or DSP build does not support (or vice versa), you get failures such as features silently not working, subsystem crashes and restarts, or boot hangs. Prevent them with a meta build that pins tested combinations, versioned interfaces with capability negotiation, compatibility checks at boot, and integration tests that run the exact combination being released.
How would you reduce integration lead time from an AOSP release to a promoted baseline?
- Start early: integrate developer previews and betas continuously instead of one big drop.
- Shrink the carried delta by upstreaming patches and removing obsolete ones.
- Automate: trial merges, conflict reports, dependency tracking and CI gates.
- Standardise conflict ownership so work is parallel across domain teams.
- Reuse resolutions (
rerere) and make HAL work predictable using the compatibility matrix. - Measure the lead time per phase to see where time goes.
How do you structure an RCA for an intermittent issue that takes days to reproduce?
Increase the reproduction rate first: stress conditions, run many devices in parallel, and automate detection so failures are captured without a person watching. Add targeted instrumentation (always-on ring-buffer tracing, extra logs around the suspected area) and make sure logs survive reboots (pstore, persistent logs). Collect a large sample and look for correlations (build, SKU, temperature, uptime, network). Form hypotheses, test them one at a time, and use the 5 Whys once the technical cause is found.
How do you manage security patch integration alongside feature work?
Security patches follow the monthly Android Security Bulletin and vendor bulletins, with embargo rules before public disclosure. Maintain a dedicated path: patches go into every supported branch on a fixed schedule, verified by STS and a focused regression test set, with the security patch level property updated. Keep this path independent of feature branches so security updates are never delayed by feature integration. Track patch level lag per branch as a metric.
What metrics would you present to leadership about platform integration health?
A short set with trends: upstream integration lead time, carried delta size, build green percentage and time to fix breakages, xTS pass rate per baseline, open P0 and P1 counts and defect convergence toward release, customer CR aging and MTTR, escaped defects, and on-time delivery rate. Pair each with a one-line interpretation and the action being taken, and highlight red items and the help needed.
What does upstream-first mean in practice, and what are its trade-offs?
Upstream-first means that changes to shared code (AOSP framework, Linux kernel) are contributed upstream and, ideally, merged there before or instead of being carried privately. Benefits: less carried delta, easier future drops, community review and testing. Trade-offs: upstream review takes time, the change must be generic enough to be accepted, and schedules may require carrying the patch temporarily. Use a clear policy: carry temporarily only with a tracked upstream submission.
How do you design a change-control process after code freeze that does not become a bottleneck?
Publish clear criteria for what is accepted (for example blocker bugs, security fixes, certification failures). Require each request to include the bug, root cause, risk assessment, test evidence and the branches affected. Meet frequently in short sessions, allow asynchronous approval for low-risk items, and keep the board small with authority to decide. Track approved changes and re-run the relevant tests automatically after they land.
Is Android moving its userspace build to Bazel?
Not as a current fact. Google explored a userspace migration from Soong to Bazel and halted it around 2023; Soong remains the primary userspace build. Kleaf, the Bazel-based kernel build, is still real and is what GKI/kernel teams use. A strong answer splits the two paths and does not treat a 2021-era migration slide as today's architecture.
What is vendor API level, and what does a GRF-style freeze change about OS upgrades?
Vendor API level (ro.vendor.api_level) is the API the vendor partition was built against. A freeze window (often called GRF / Google Requirements Freeze) lets that vendor image pair with newer system images for a documented number of releases, as long as VINTF still matches. The OEM can take a yearly OS upgrade without a full SoC rebase of every HAL. Exact window lengths change; say you would check the current CDD and vendor-API notes. What still moves: new matrix-required HALs, CTS/VTS/GTS for the new release, CDD, and any new KMI.
How do you triage crashes at fleet scale?
Do not debug one tombstone at a time. Symbolize stacks for that exact build, cluster by process plus a stable stack signature (not ASLR addresses), rank by volume times user impact, and assign a DRI to each top cluster with one representative report. A cluster is closed when the signature is gone (or below threshold) on the next build and a regression gate exists. Pair crash-free rate with the top-N signatures so a lucky week cannot hide a new system_server cluster.
Scenario & debugging
Battery life regressed after a platform update. Walk me through your debug.
Reproduce and quantify on a fixed profile, comparing drain per hour with the previous build. Pull a bug report and load batterystats into Battery Historian to spot new wake locks, wakeups, jobs or alarms. Use Perfetto and /sys/kernel/debug/wakeup_sources to find what is keeping the CPU awake, and check suspend residency. Bisect the change set or components to find the culprit. Fix it (remove the wake lock, batch the work, move it to WorkManager, restore sensor batching), re-measure, and add a power KPI gate so it cannot recur silently.
A systemic issue spans app, framework, HAL and kernel, and every team says it is not theirs. What do you do?
Take ownership as the DRI and move the discussion to one triage thread. Get a reliable repro, then capture logs and traces at every layer on one timeline (logcat, dumpsys, HAL logs, dmesg, Perfetto). Bisect to find the layer where the data first goes wrong and the change that introduced it. Present the evidence and assign the owner whose code must change, with a date. Keep a single status of record, track to closure, and add a regression gate.
After an upstream drop, the device does not boot on one chipset. How do you approach it?
Find where it stops. No splash or bootloader output suggests XBL, ABL, AVB or partition layout problems. Splash then reboot loop suggests a kernel panic or init failure, so read the UART console, last_kmsg or pstore. Stuck at the boot animation suggests system_server or a critical service crashing, so read logcat. Compare with the working chipsets to see what differs (BSP, kernel branch, HAL versions, VINTF, SELinux denials). Bisect the drop by repository if needed, and treat it as a P0 blocking promotion for that chipset.
CTS pass rate dropped from 99.8 to 97 percent on the nightly build. What do you do?
Group the new failures by module to see if they share a cause (one broken service can fail hundreds of tests). Check the lab first: network, SIM, device health and test suite version, because environment problems often cause mass failures. Compare with the previous nightly build and its manifest to list the changes in between, and re-run a sample of failures to confirm. Bisect to the culprit change, revert it if it blocks others, and assign a fix. Add the relevant modules to presubmit if they are cheap enough.
A customer reports random reboots in the field on a released product. How do you drive it?
Acknowledge quickly and set an update cadence. Collect data: reboot reasons, kernel panic logs, ramdumps, watchdog and subsystem restart logs, and field telemetry (which builds, SKUs, regions, conditions). Look for patterns and try to reproduce with stress tests under similar conditions. Localise to the layer (kernel panic, modem crash, system_server watchdog), assign the owner and drive root cause. Agree a fix and OTA plan with the customer, roll out staged, confirm the reboot rate drops, then run a post-mortem and add a stability gate.
Two days before release, a P1 regression is found. What do you do?
Quantify impact: which users, how often, how severe, and whether a workaround exists. Check whether a low-risk fix is available and how much re-testing it needs. Lay out options: slip the release, ship with a documented known issue and a dated maintenance fix, disable the feature by flag, or take a targeted fix with focused re-test. Present the options with risks to the release owner and stakeholders, decide transparently, record the decision, and communicate it to customers honestly.
A merged upstream drop builds and boots, but camera start-up got 400 ms slower. How do you find the cause?
Confirm with repeated measurements on both builds under the same conditions. Capture Perfetto traces of camera launch on both builds and compare the phases: app start, camera service connect, HAL open, sensor configuration, first preview frame. Identify the phase that grew, then look at changes in that area in the drop (framework camera service, HAL interface version, SELinux, scheduling). Bisect commits within the suspect repositories. Assign to the camera or framework owner with the traces, and add launch latency to the performance gate.
Your team cherry-picked a fix to three release branches, but one branch still shows the bug. Why might that be?
Possible reasons: the fix depends on an earlier commit that exists on the other branches but not this one; the code on that branch differs and the conflict was resolved incorrectly; the bug on that branch has a different root cause; the fix is in a component that is built from a different repository or prebuilt on that branch; or the tested build did not actually include the change (check the build's manifest snapshot). Verify the change is in the build, compare the code paths, and re-investigate the root cause for that branch.
An OEM customer wants a feature that is not in the current baseline, added a week before code freeze. How do you respond?
Do not refuse in the meeting; understand the business need and the deadline behind it. Assess scope, risk, affected components and test effort with the owning team. Return with options: deliver in the next maintenance release, deliver a limited version behind a flag, or take it now with explicit trade-offs (another item moves, or freeze moves). Make the trade-off visible to program management and the customer, decide together, and document it.
Build breakages on the main branch happen several times a day and block everyone. How would you fix this?
Measure first: which targets break, which kinds of changes cause it, and how long breakages last. Strengthen presubmit to build the affected targets and variants, and enforce atomic submission for multi-repository changes. Adopt a revert-first policy with an on-call build sheriff. Add automated culprit finding to notify authors quickly. Track build green percentage and time to repair as team metrics and review them weekly.
A watch baseline passes all gates, but the customer's product build shows poor standby battery. How do you investigate?
Compare the customer's build against the baseline: OEM apps and services, overlays, configuration, preinstalled watch faces and firmware versions (a different meta build combination). Reproduce with the customer's build on the same power profile and collect batterystats, wakeup sources and Perfetto. Often the cause is an OEM app holding wake locks, a watch face updating too often in ambient mode, or a different sensor or connectivity configuration. Share evidence with the customer, help fix it, and offer them the same power gate you use internally.
You inherit a program with no clear gates and frequent escaped defects. What do you do in the first 90 days?
In the first 30 days, map the image, branches, teams, KPIs and current red items; sit in triage; do not reorganise yet. By 60 days, introduce one status of record, written promotion gates based on the most common escaped defect types, named DRIs for systemic issues, and work-in-progress limits. By 90 days, run the first gated drop, report gate results and escaped defect trends, set a customer communication cadence, and plan the next set of gate improvements.
A modem firmware update in the meta build breaks VoLTE, but the HLOS team says nothing changed on their side. How do you handle it?
Confirm by testing combinations: old modem with new HLOS and new modem with old HLOS, which isolates the component. If the new modem alone breaks it, collect modem logs, IMS and RIL logs and a SIP trace on both firmware versions, and give the modem team the exact failing step. Check whether the new firmware changed an interface (a QMI message or a configuration item) that the HLOS side must adapt to. Pin the old combination in the meta build until a fix is ready, and add a VoLTE call test to the meta build gate.
Tell me how you would handle a disagreement with a domain architect about a fix approach.
Move the discussion from opinions to data. Define what matters (failure rate, performance, risk, how many branches or customers are affected), run a quick experiment or stress test on both options, and write a short one-page comparison with a rollback plan. Present it to the architect and agree on the decision criteria before discussing the choice. Once a decision is made, commit fully and help carry it out, even if it was not your preferred option.
An integration drop is two weeks late because one HAL team keeps missing dates. What do you do?
Understand why: capacity, unclear requirements, technical blockers or competing priorities. Break the remaining work into smaller tracked pieces with daily visibility. Offer help: pair them with engineers from other teams, clarify the minimum required for promotion, or allow the drop to land with the old HAL version if the compatibility matrix permits. If priorities conflict, escalate to management with a clear ask and the impact of each choice. Communicate the revised plan to stakeholders promptly.
An OTA rollout shows a higher boot failure rate than the previous release after reaching 10 percent. What do you do?
Pause the rollout immediately; A/B rollback protects devices that fail to boot, but the failure still harms users. Collect data from affected devices: models, previous build, storage state, and logs from failed boot attempts. Check whether failures correlate with a specific source build (delta payload problem), storage condition (for example low free space for virtual A/B snapshots), or hardware variant. Fix, test on the affected configurations, then resume the rollout from a small percentage.
A new Android release requires migrating several HIDL HALs to AIDL. How would you plan it?
List every HAL that must change using the framework compatibility matrix and deprecation notices. For each, identify the owner, effort and dependencies (framework clients, vendor clients, tests). Prioritise HALs that block the release, and check which can remain on HIDL for now. Plan the migration so both versions can coexist during the transition where possible, write or update VTS tests for the AIDL version, and track progress in the integration status. Start early, ideally on preview releases.
Describe how you would lead a platform through ambiguous requirements and unclear ownership.
Create structure quickly: define the deliverable and the quality bar, list the components and interfaces, and name an owner for each, even if temporary. Set up a single status of record, promotion gates and a regular cadence. Drive cross-team triage on systemic issues and convert unknowns into a tracked list that you burn down. Communicate clearly and early with stakeholders when assumptions change. In an interview, use the STAR format and quantify the result (on-time delivery, defects closed, KPIs held).
A HAL binary builds but is missing from the userdebug image. Where do you look?
First confirm the module exists in the build graph (Android.bp name, vendor: true, no broken required deps). Then check the product makefile: it must be in PRODUCT_PACKAGES (or pulled in by another packaged module). Also check the variant (debug-only packages go in PRODUCT_PACKAGES_DEBUG and will be absent on user), the partition (vendor vs system), and whether a conditional ifeq excluded that SKU. Built is not installed.
An OEM can take next year's Android on last year's vendor image. What must still be true?
The new system must be inside the vendor-API freeze window for that vendor API level, VINTF must still match (no newly required HAL the vendor does not provide), GKI/KMI must still be compatible if the kernel stays, and the device must pass the new release's CTS/VTS/GTS and CDD. Firmware pins in the meta build must still satisfy the new HLOS. If any of those fail, it is not a free upgrade: someone must move vendor code, kernel branch or firmware.
A carrier lab fails a call case that passes on the open-market SKU. How do you start?
Do not start in the modem C-core. Compare CarrierConfig for that MCC/MNC, IMS and APN overlays, the exact meta-build firmware pins, and whether the lab SIM exercises a different feature flag (VoLTE, VoWiFi, 5G NSA/SA). Reproduce on the same carrier config. If config matches, then collect RIL, IMS and modem logs. PTCRB/GCF failures need the same split: protocol/RF versus HLOS policy. Keep one DRI and one status; name the failed contract, not the company.