Wearables & Platform Integration

Platform Integration & Release

A phone or watch ships as one software image built from Google's AOSP, the chip vendor's BSP and HALs, firmware for other processors, and OEM customisation. This page explains how that image is composed, how upstream releases are integrated into it, how branches, merges, compliance tests and quality gates keep it healthy, and how systemic issues and releases are driven to closure across distributed teams and customers.

~90 min read 0 interview questions
In 30 seconds
  • The HLOS image is the Android or Wear OS side (kernel plus user space); non-HLOS images are firmware for the modem, DSPs, TrustZone and boot chain. A meta build ties them together into one flashable, versioned release.
  • Upstream integration lands each new AOSP tag onto vendor trees and per-chipset baselines while carrying vendor changes forward; Treble, VINTF and GKI make this tractable. AOSP source lives on android.googlesource.com; check current release notes for tag cadence.
  • Day-to-day: Soong/Android.bp and lunch product images, branching and merge strategy, then promotion gates (build, boot, smoke, CTS/VTS/GTS, power, stability, performance). OTAs ship through update_engine with A/B or virtual A/B.
  • Systemic issues are driven by one owner through reproduce, instrument, cluster, localise, fix, verify and prevent, with root-cause analysis and a regression gate at the end.

What an HLOS image is

HLOS stands for High-Level Operating System. On a Qualcomm-style SoC it means the rich OS running on the main application processor: the Linux kernel plus Android or Wear OS user space. It is contrasted with the non-HLOS software: firmware and real-time operating systems that run on the other processors inside the chip. The HLOS image is the full shippable Android build, and leading HLOS image activities means owning how it is composed, branched, integrated, validated and shipped for each chipset generation.

HLOS IMAGE  =  AOSP / Wear OS framework and system apps
             +  vendor BSP: Linux kernel (GKI + vendor modules), device tree, drivers
             +  vendor HALs (AIDL / HIDL): sensors, display, audio, radio, power, thermal ...
             +  proprietary vendor libraries and HLOS-loaded firmware blobs
             +  OEM customisation, overlays, configuration and branding
   ──▶ branched per SoC generation ──▶ validated against gates ──▶ shipped per OEM product
Analogy

Think of a car. The HLOS is the infotainment and driver-facing software everyone sees; the non-HLOS firmware is the engine control unit, brake controller and airbag module, each with its own dedicated computer. A car only ships when all of them are the right versions and work together. In a device, the infotainment is Android on the application processor, the dedicated controllers are the modem, DSPs, TrustZone and boot firmware, and the "vehicle build sheet" that fixes all versions together is the meta build.

HLOS versus non-HLOS

ComponentRuns onHLOS or non-HLOSTypical image
Android / Wear OS user spaceApplication processor (Arm Cortex-A)HLOSsystem, system_ext, product, vendor, odm inside super
Linux kernel and ramdiskApplication processorHLOSboot, vendor_boot, init_boot, dtbo, vbmeta
Android bootloader (ABL)Application processor, before the kernelUsually delivered with HLOSabl
Primary and secondary boot loaders (PBL in ROM, XBL)Application processor at resetNon-HLOSxbl, xbl_config
Modem firmware (MPSS)Modem DSPNon-HLOSmodem image (NON-HLOS.bin style FAT image)
Audio and compute DSP firmware (ADSP, CDSP)Hexagon DSPsNon-HLOSadsp, cdsp images
Sensor low-power island (SLPI) or sensor hubLow-power DSP or MCUNon-HLOSsensor firmware image
TrustZone / trusted execution environmentSecure world of the application processorNon-HLOStz, hyp, trusted apps
Always-on power and resource managerSmall power-management processorNon-HLOSaop or equivalent
Wi-Fi and Bluetooth firmwareConnectivity chip or subsystemNon-HLOS (often loaded by HLOS drivers)firmware blobs

The meta build

Each non-HLOS component is built by its own team, on its own schedule and often with its own toolchain. A meta build is the integration record that pins one version of every component (HLOS build plus each firmware build) into a single tested combination, together with the partition layout and flashing instructions. Tools such as fastboot or vendor download tools flash the meta build as a unit. Integration bugs frequently come from mismatched combinations, for example a new HLOS driver expecting a firmware interface the old modem build does not have.

Where HLOS and non-HLOS meet

At HAL implementations, kernel drivers for remote processors (remoteproc and subsystem restart), shared memory, and message protocols such as QMI and glink. Changes on either side must keep these contracts stable.

Why wearables are special

Watches add sensor-hub firmware and an always-on co-processor as first-class components, and the image must fit tighter RAM, flash and power budgets.

Image variants

The same source produces user, userdebug and eng builds; only user builds ship and are certified, but most debugging happens on userdebug.

Build fingerprint

ro.build.fingerprint identifies brand, product, device, Android version, build ID, variant and signing keys; certification and OTA targeting depend on it.

Interview angle Interviewers open with "what is an HLOS image?" to see whether you can define it crisply and place it in the full system. A strong answer defines HLOS versus non-HLOS, lists what composes the image, explains the meta build as the tested combination of all components, and names the contracts at the boundary (HALs, QMI, shared memory, subsystem restart).

AOSP upstream integration

AOSP source is published on android.googlesource.com. Google tags platform releases there; the public tag and branch cadence has changed over the years (monthly, quarterly and other rhythms have all existed), so treat any fixed public cadence as historical and check current Android release notes for the tags you must track. Separately, Android still publishes a monthly security bulletin; that is a patch list, not a promise that a full AOSP drop appears on a fixed calendar. A chip vendor tracks the tags it cares about and lands them onto its internal trees and per-chipset baselines, carrying its own changes (drivers, HALs, performance features, fixes) forward, so that OEMs receive current Android plus the vendor's value-add. Qualcomm calls its upstream-tracking program Keystone; other vendors run equivalent programs.

Google AOSP upstream (release tag N, security patches)
   │  track, plan the drop, read release notes and API/HAL changes
   ▼
Vendor internal trees  ── merge or rebase upstream, carry vendor deltas forward
   │  conflicts? ──▶ dependency analysis ──▶ assign owner (framework / BSP / HAL / apps)
   ▼
Per-chipset baselines (SoC gen A, B, C)  ── apply BSP deltas, bump HAL versions, update VINTF
   │  gates: build health ▸ boot ▸ smoke ▸ CTS/VTS/GTS ▸ power / stability / performance
   ▼
Promote to OEM-consumable baseline  ── OEM branches, customisation
   ▼
OEM product image ──▶ certification ──▶ carrier / retail release ──▶ OTA
Analogy

Upstream integration is like a restaurant chain adopting the head office's new seasonal menu while keeping each branch's local specialities. Every season the head office sends a new recipe book; each branch must merge it with its own additions, check nothing clashes, train staff and pass a hygiene inspection before serving customers. The head office is Google's AOSP, the branches are per-chipset baselines, local specialities are vendor deltas, and the hygiene inspection is the set of promotion gates and compliance suites.

What makes it tractable

MechanismWhat it gives integration
Project TrebleSeparates the framework (system) from vendor code (vendor, odm) behind stable HAL interfaces, so a framework update does not require rewriting vendor code.
VINTFDevice and framework manifests plus compatibility matrices declare which HAL versions each side provides and needs; mismatches are caught at build time and boot.
Stable AIDL HALsVersioned, frozen interfaces; the framework can support several versions so vendor code upgrades independently.
GKI (Generic Kernel Image)One Google-built kernel core per Android common kernel branch, with vendor code in loadable modules against a stable KMI (kernel module interface). Within the same frozen KMI, core-kernel updates do not require vendor modules to rebuild; a rebuild is needed when the KMI itself changes (new ACK branch or ABI) or a module uses a new symbol.
Mainline (APEX / APK modules)Some system components update through Google Play system updates, which reduces what the vendor must integrate but adds version-compatibility checks.
Generic System Image (GSI)A pure AOSP system image used to prove the vendor side follows Treble (VTS and CTS-on-GSI).

Typical sources of integration conflicts

  • HAL interface bumps: a new Android release requires a newer HAL version (for example a HIDL to AIDL migration); vendor HAL owners must implement it before the drop can land.
  • Framework API or behaviour changes that vendor patches in framework code relied on, such as a refactored service or new permission checks.
  • SELinux policy changes that deny vendor services previously allowed.
  • Build system changes (new Soong modules or flags, leftover Make variables, product makefile churn). Google's userspace Bazel migration was halted around 2023; Soong remains the primary userspace build. Kleaf/Bazel for kernel builds is separate and still real.
  • Kernel branch moves to a new Android common kernel (a new KMI). Modules must be rebuilt against that new KMI; ordinary same-KMI GKI updates do not force a module rebuild.
  • Toolchain updates (new Clang version, stricter warnings treated as errors).
Tip Upstream-first reduces future pain: every vendor change that is contributed to AOSP or the Linux kernel is one less patch to carry and rebase on every future drop. Track the size of your carried delta as a health metric.
Interview angle "Walk me through integrating a Google upstream release" is the core question for integration roles. Cover: tracking the tag, planning the merge or rebase, sequencing across SoC baselines, dependency analysis with named owners, the gates you run before promotion, schedule management, and the stakeholder domains involved (multimedia, camera, connectivity, display, power, security). Mention Treble, VINTF and GKI to show you know why it is feasible at all.

SoC bring-up and where the image comes from

Before an image can be integrated and stabilised, a new chip must be brought up layer by layer. An integration lead does not write every driver, but must understand the sequence to triage "stuck at boot" and readiness issues. Boot chain detail is in Android boot.

StageWhat happensIntegration angle
Boot chainBoot ROM (PBL) ▸ XBL/SBL (DDR, clocks) ▸ bootloader (ABL) ▸ kernel ▸ init ▸ frameworkKnow the chain to triage "stuck at boot" systemic issues and verified-boot failures
Kernel, device tree and driversKernel, device tree, clocks, regulators, pin control, peripheral driversConsume BSP outputs into the image and track driver readiness per feature
HALs and frameworkVendor HALs (sensors, display, audio, connectivity, radio) wired to the frameworkOwn the framework-up integration; VINTF and SELinux must be right
StabilisationPower, thermal, stability and KPI tuning to the ship gateDrive systemic issue triage to closure; this is where schedules slip
Analogy

Bringing up a new chip is like opening a new building: foundations and utilities first, then floors, then offices, then furniture and finishing. You cannot test the elevators before the power is on. In a device, the foundations are the boot chain and DDR, utilities are clocks and regulators, floors are the kernel and drivers, offices are the HALs and framework, and the finishing is power and performance tuning.

Interview angle Expect "outline the SoC boot and bring-up sequence at a high level." Give the chain in order, then show judgment about who owns which stage and what signals readiness (boots to shell, boots to home screen, passes smoke, passes xTS, meets KPIs).

Branching models for a platform

A platform codebase is hundreds of Git repositories stitched together by a repo manifest, built for several chipsets and many OEM products at once. The branching model decides where changes land first and how they flow between streams.

                aosp upstream (android-N release tags, security bulletins)
                      │ merge / rebase per drop
                      ▼
 vendor main / dev ───●────●────●────●────●──────────▶  (next Android version work)
                           │              │
                           │ branch cut   │ branch cut
                           ▼              ▼
             release-N-socA ──●──●──▶   release-N-socB ──●──▶   (stabilise, cherry-picks only)
                        │
                        ▼
             oem-X-product ──●──▶  (customisation, customer fixes, OTA)
ModelHow it worksGood forRisk
Trunk-basedEveryone commits small changes to one main branch; unfinished features hidden behind flagsFast integration, fewer merges; Google's AOSP moved to "trunk stable" development with aconfig feature flagsNeeds strong presubmit CI and flag discipline
Release branchesCut a branch from main at feature freeze; only fixes go in after thatStabilising a specific Android version or chipset while main moves onFixes must be ported to several branches
Feature branchesLong-lived branch per large feature, merged back laterIsolating risky, large workBig, late, painful merges ("merge hell")
Per-SoC / per-OEM branchesBranch per chipset baseline and per customer productDifferent BSPs, schedules and customer changesBranch explosion; fixes drift between branches
Analogy

A branching model is like a river system. The main river (trunk) keeps flowing; at certain points canals are dug off it (release branches) to feed specific towns (chipsets and customers). Water can be pumped from the river into a canal (cherry-picking a fix), but once a canal is cut you must keep pumping, or the towns fall behind. In a platform, the pumping is porting fixes, and the discipline is keeping the number of canals small and deciding clearly which fixes go where.

Branch hygiene rules that scale

  • Fix on the oldest supported branch that needs it, then merge forward, or fix on main and cherry-pick back, but pick one direction per organisation and track it.
  • Freeze stages: feature freeze (no new features), code freeze or "lockdown" (only approved fixes), release candidate (only blocker fixes with sign-off).
  • Change tracking: every change references a bug or change request; release branches accept only tracked, approved changes.
  • Automated forward-merge checks detect fixes that landed on a release branch but not on main, which otherwise reappear as regressions next release.
  • Retire branches on a published schedule; every live branch costs CI capacity and security-patch effort.
Common pitfall Letting per-customer branches accumulate unique fixes. Each time the next upstream drop arrives, those fixes must be rediscovered and re-applied. Push customer fixes back to the shared baseline whenever they are not customer-specific.
Interview angle Interviewers ask "what branching model would you use and why?" There is no single right answer; they want trade-offs. Explain how many streams you need (upstream, main, per-release, per-customer), where fixes land first, how they propagate, when branches freeze, and how you prevent drift.

Merge, rebase and cherry-pick strategies

Bringing an upstream drop into a tree that carries vendor changes can be done in several ways. The right choice depends on history requirements, the size of the carried delta and how many teams work on the tree.

StrategyWhat happensProsCons
MergeCreate a merge commit joining upstream history with the vendor branchPreserves exact history; one place to resolve conflicts; SHAs of existing commits unchanged, so downstream branches are not disruptedHistory becomes hard to read; the carried delta is not visible as a clean patch set
RebaseReplay every vendor commit on top of the new upstreamLinear history; the vendor delta stays a clean, reviewable patch stack; easier to upstream patchesRewrites SHAs, breaking anyone who built on the old commits; conflicts resolved commit by commit
Cherry-pickCopy selected individual commits onto another branchPrecise; ideal for porting fixes to release branchesCreates duplicate commits with different SHAs; easy to miss dependent commits
SquashCollapse a series of commits into oneTidy history for a featureLoses granular history; makes bisecting and reverting harder
Analogy

Imagine updating a cookbook you have annotated. Merging is stapling the new edition to your annotated copy and writing a note on how they fit together. Rebasing is taking a fresh new edition and rewriting all your annotations into it, page by page. Cherry-picking is copying one specific annotation into a friend's copy. In Git terms, the annotations are the vendor delta, the edition is the upstream release, and the choice decides whether your history stays intact (merge) or your delta stays clean (rebase).

# sync a multi-repo platform tree
repo init -u <manifest-url> -b <branch> -m <manifest>.xml
repo sync -c -j8

# merge an upstream tag into a vendor branch (one project)
git fetch aosp android-15.0.0_r1
git merge --no-ff android-15.0.0_r1

# or rebase the vendor patch stack onto the new tag
git rebase --onto android-15.0.0_r1 android-14.0.0_r1 vendor-main

# port one fix to a release branch, recording the source commit
git cherry-pick -x <sha>

# remember how a conflict was resolved so repeated merges reuse it
git config rerere.enabled true

# find the change that introduced a regression
git bisect start <bad> <good>
git bisect run ./run_smoke_test.sh

Conflict resolution at scale

  1. Pre-scan Do a trial merge early (for example on beta or developer-preview tags) to list conflicting repositories and files before the real drop.
  2. Classify Group conflicts by domain (framework, BSP, HAL, apps, build) and by type (textual, semantic, API or HAL version change).
  3. Assign owners Each conflict goes to the team that owns the carried change, with a deadline; the integration lead owns the tracking list.
  4. Resolve with understanding Read both sides' intent: why upstream changed the code and why the vendor patch exists. Sometimes the right resolution is dropping the vendor patch because upstream fixed the same problem.
  5. Build and test the resolution A textual conflict fixed without compiling and testing is the most common source of silent regressions.
  6. Record Document non-obvious resolutions in the commit message and reuse them with rerere on the next drop.
Common pitfall Semantic conflicts do not show up as Git conflicts. Upstream may rename a behaviour, change a default or move a check, and the vendor patch still applies cleanly but is now wrong. Only builds, tests and reviewers who understand both sides catch these.
Interview angle Expect "merge or rebase, and why?" and "how do you resolve hundreds of conflicts on a schedule?" Show you know the SHA-rewrite cost of rebasing shared branches, why cherry-pick with -x helps traceability, the difference between textual and semantic conflicts, and a process with owners and a tracking list rather than one hero resolving everything.

Compliance: CTS, VTS, GTS and friends

Android devices must pass Google's compatibility test suites to be called Android-compatible and to ship Google Mobile Services (GMS). Together they are often called xTS. They run on the Trade Federation (Tradefed) test harness.

SuiteTestsWhy it matters
CTS (Compatibility Test Suite)Public Android APIs and behaviour required by the Compatibility Definition Document (CDD)Apps behave the same on every compatible device
CTS VerifierManual or semi-automated tests for things automation cannot check (sensors, camera, audio, NFC)Hardware-dependent behaviour
VTS (Vendor Test Suite)HAL interfaces, kernel (including GKI and KMI requirements), VINTF, vendor partition behaviourThe Treble contract between framework and vendor holds
CTS-on-GSIRuns CTS with a Generic System Image on the vendor's deviceProves the vendor side works with a pure AOSP framework
GTS (GMS Test Suite)Google Mobile Services requirements (Play services, Play Store, Google apps behaviour)Required under the GMS licence; not part of AOSP
STS (Security Test Suite)Security patches listed in the monthly security bulletinValidates the security patch level claimed by the build
Form-factor suitesExtra requirements for Wear OS, Android TV, Automotive and othersA watch must also pass its Wear OS-specific requirements
Analogy

xTS is like the safety and emissions tests a car must pass before it can be sold. CTS checks that the car behaves like any other car to the driver (pedals, indicators); VTS checks the internal parts fit the standard connectors; GTS is the extra inspection a brand partner requires before letting you use their logo. In Android, the driver is the app developer, the standard connectors are HAL and kernel interfaces, and the brand partner is Google's GMS licence.

Running xTS in practice

# run a full CTS plan with Tradefed
./android-cts/tools/cts-tradefed run cts -s <serial>

# run one module or test while debugging a failure
cts-tradefed run cts -m CtsSensorTestCases -t android.hardware.cts.SensorTest

# retry only the failures from a previous session
cts-tradefed run retry --retry <session-id>
  • Test often, not only at the end: run subsets of xTS in daily CI and full suites on promotion candidates, so failures are bisectable to a small change range.
  • Triage every failure as device bug, test bug, test environment problem (network, SIM, lab setup) or flaky test; only device bugs go to engineering teams.
  • Waivers: known test bugs can be waived through Google's process, but a waiver needs evidence and is not a way to hide device bugs.
  • Final submission uses the exact user build with the release fingerprint; results are tied to that fingerprint.

GMS, MADA and the licence path

CTS and VTS prove Android compatibility. Shipping the Play Store, Play services and Google apps is a separate, licensed path:

  • GMS (Google Mobile Services) is the licensed set of Google apps and services. It is not in AOSP.
  • MADA (Mobile Application Distribution Agreement) is the commercial contract under which an OEM may preload GMS. Exact terms are confidential; in interviews, speak only at this level: licence, placement and update requirements, and a test bar.
  • GTS is the technical test suite that licensees run against those requirements. Passing CTS + VTS + GTS (and related suites) is necessary but not sufficient; the licence and a formal submission still apply.
  • CTS-Verifier covers behaviour that automation cannot fully judge: sensors, camera, audio, haptics, NFC, accessibility, and similar hardware-in-the-loop checks. Budget lab time and operators; it is often the long pole late in a program.
Note Do not invent current GMS placement rules or 2025/2026 policy details. If asked, say you would check the current licensee documentation and CDD for the Android version being shipped.

Carrier certification (high level)

Cellular devices also need operator and industry certification, which is independent of Google's xTS:

TrackWhat it isIntegration angle
PTCRBNorth-American operator certification forum for cellular devicesPlan lab time, SIM/eSIM variants, RF and protocol cases; failures often need modem plus HLOS owners
GCFGlobal Certification Forum; similar industry bar used widely outside PTCRB operatorsSame idea: a shared test catalogue, operator deltas on top
Operator acceptanceEach carrier's own lab and field cases (voice, data, SMS, IMS, emergency, roaming)Track as a release gate with a named DRI; daily loop on blockers
CarrierConfigPer-carrier Android configuration (feature flags, timers, IMS and telephony behaviour) loaded by MCC/MNC or carrier appMany "works on one SIM, fails on another" bugs are config, not modem silicon

Keep employer and customer names out of answers. Describe the model: SoC vendor supplies the modem and RF package, OEM owns the product image and GMS path, carriers own their acceptance bar, and the integration lead keeps one status across all three.

Common pitfall Leaving xTS until the week before release. Failures then come in hundreds, span many changes and cannot be bisected easily. Treat CTS and VTS pass rate as a continuous metric on every baseline. The same applies to CTS-Verifier and the first carrier lab entry.
Interview angle Interviewers check that you know what each suite covers (CTS for app-facing behaviour, VTS for vendor and HAL, GTS for GMS, CTS-Verifier for manual hardware checks), that MADA/GMS is a licence plus tests, and that carrier cert (PTCRB/GCF plus CarrierConfig) is a separate gate. Bonus: CDD, CTS-on-GSI, STS and security patch level, and how you handle flaky tests and waivers.

Promotion gates and quality criteria

A promotion gate is a set of pass or fail criteria that a build must meet before it moves to the next stage: from integration branch to baseline, from baseline to OEM delivery, from release candidate to shipping. Gates turn quality from an opinion into a measured decision.

GateChecksWhy
Build healthAll targets and variants build across SoC baselines; no new warnings-as-errorsNo broken trees downstream
BootBoots to home screen on every SKU; boot success rate over many cyclesCatch catastrophic breakage immediately
VINTF and compatibilityFramework and vendor HAL versions compatible; VTS core passesTreble contract must hold
Smoke / basic acceptance test (BAT)Core UX, calls, data, Wi-Fi, Bluetooth, sensors, camera, OTACatch gross functional breakage early
ComplianceCTS, VTS, GTS pass rate at or above target; no new failuresCertification and GMS licence
Power regressionStandby drain, suspend residency, wake lock and wakeup budget, per-use-case current against last goodShip blocker on wearables; see Power and thermal
StabilityCrash-free rate, ANR rate, kernel panics, watchdog resets, subsystem restarts, long-run monkey testsField quality
PerformanceBoot time, app launch latency, jank percentage, memory footprintCompetitive user experience
Open defectsZero open P0/P1 regressions; P2 counts under agreed limits with ownersKnown risk is explicit and accepted
Analogy

Promotion gates are like airport security checkpoints: you pass check-in, then security, then passport control, then boarding, and each stage checks different things. Failing one stops you from boarding even if you passed the rest. In a platform, check-in is build health, security is boot and smoke, passport control is compliance, and boarding is the KPI and defect criteria signed off by the release owner.

Good gate criteria

  • Measurable and automated where possible
  • Compared to a known-good baseline, not absolute numbers only
  • Owned: someone signs off each gate
  • Stable thresholds that do not move under schedule pressure
  • Published in advance so teams can pre-check

Bad gate criteria

  • "Looks OK" manual judgement
  • Thresholds so loose that nothing fails
  • Flaky tests that fail randomly, training people to ignore failures
  • Criteria changed after the build fails
  • No link between the gate and a field-quality outcome
Tip Make power and stability first-class promotion gates, regression-tested against the last good build for every upstream drop. That catches drain and crashes at integration time, not in the field, where they cost ten times more to fix.
Interview angle "What gates a build promotion?" is a near-certain question. List the gates in order, give example thresholds, and explain what happens when one fails: hold and triage rather than promote, or promote with an explicit, signed-off waiver and a dated fix plan. Interviewers listen for whether you would hold the line under schedule pressure.

CI/CD for a platform

Platform CI is harder than app CI: a full Android build can take hours, tests need real devices, and hundreds of repositories must be consistent. The goal is still the same: find each breakage as close to the change that caused it as possible.

developer change ──▶ Gerrit review
   │ presubmit: build affected targets ▸ unit tests ▸ static analysis ▸ quick device smoke
   ▼  (Code-Review +2 and Verified +1)
submit to branch
   │ postsubmit / continuous: full builds on all targets ▸ boot tests ▸ broader test tiers
   ▼
nightly / candidate build ──▶ full smoke ▸ xTS subsets ▸ power and performance runs ▸ stability soak
   ▼
promotion candidate ──▶ full xTS ▸ KPI comparison vs last good ▸ sign-off ──▶ baseline / release
Analogy

Platform CI is like quality control in a car factory: some checks happen at each workstation (a bolt is torqued correctly), some at the end of the line (the car starts and drives), and some on a test track overnight (long durability runs). Catching a bad bolt at the workstation is cheap; finding it after delivery is expensive. Workstation checks are presubmit tests, end-of-line checks are postsubmit builds and boot tests, and the test track is nightly soak, xTS and KPI runs.

PracticeWhy it matters
Topic or multi-repo change submissionA change touching framework and vendor repos must submit atomically, or the tree is broken in between (Gerrit topics, repo-level "submit together").
Build caching and incremental buildsRemote build caching and distributed execution cut presubmit time so developers get feedback in minutes, not hours.
Test tiersFast tests in presubmit; slower device tests postsubmit; full xTS and soak nightly or per candidate.
Device labsPools of real devices per SKU, with automated flashing and health checks; unhealthy devices are removed automatically.
Flaky test managementDetect flakiness statistically, quarantine flaky tests with an owner and deadline, and never let them block or hide real failures.
Automated bisectionWhen a postsubmit or nightly test fails, CI builds intermediate change sets to find the culprit and notifies its author.
Build breakage policy"Revert first, ask later": a change that breaks the build is reverted immediately so others are not blocked.
Artifacts and traceabilityEvery build stores its manifest snapshot (repo manifest -r), images, symbols and test results, so any build can be reproduced and debugged.
Common pitfall Merging a huge upstream drop as one change that CI can only test as a whole. When it fails, nobody knows which of thousands of commits caused it. Stage large drops (for example repository by repository, or in chunks) and keep bisection possible.
Interview angle Interviewers ask how you would set up or improve CI for a platform, or how to reduce build breakages. Talk about presubmit versus postsubmit, test tiers, multi-repo atomic submits, bisection, flaky-test handling, revert policy and metrics such as build-green percentage and time to fix a broken build.

Triage and driving systemic issues to closure

A systemic issue is a problem that spans several layers or teams, or keeps recurring: battery drain after an upstream drop, random reboots across SKUs, call drops on one carrier, a boot-time regression. No single team owns it by default, so it needs a driver. This is the core skill of a platform integration lead.

1. OWN IT       ──▶ one directly responsible individual (DRI); stop ownership ping-pong
2. REPRODUCE    ──▶ fixed test profile; scope severity, affected SKUs/builds, KPI impact
3. INSTRUMENT   ──▶ logs across layers (app, framework, HAL, kernel, firmware) on one timeline:
                    bug reports, logcat, dumpsys, Perfetto, batterystats, dmesg, modem logs
4. LOCALISE     ──▶ bisect builds to the introducing change; bisect layers to the failing one
5. DRIVE        ──▶ one triage thread with the right experts; named owners, dates, cadence
6. DECIDE       ──▶ fix vs risk vs schedule; communicate to internal teams and customers
7. CLOSE        ──▶ land fix ▸ verify on the original repro ▸ regression test/gate ▸ post-mortem
Analogy

Driving a systemic issue is like an emergency-room doctor leading a patient's care. One doctor is responsible, orders tests from several departments, reads all the results together, decides the treatment and does not discharge the patient until follow-up care is arranged. In platform work, the doctor is the DRI, departments are domain teams (kernel, HAL, framework, modem), the tests are traces and logs across layers, and the follow-up care is the regression gate and post-mortem.

Deciding who owns a cross-domain bug

  • Triage by layer and by evidence: reproduce, capture logs across boot, kernel, HAL and framework, bisect against the introducing change.
  • Map the failing component to its owning domain. The owner is the team whose code must change, not the team that first saw the symptom.
  • If it is truly shared (for example a contract between a HAL and firmware), assign one DRI and have the others co-own actions.
  • Remove ambiguity fast, so engineers spend time fixing instead of arguing about ownership.

Root-cause analysis

Root-cause analysis (RCA) finds why the problem happened, not only what broke, so that the fix prevents recurrence. Two common tools:

5 Whys

Ask "why?" repeatedly (about five times) until you reach a process or design cause. Each answer must be backed by evidence, and there may be several branches.

Fishbone (Ishikawa)

List possible causes by category (code, configuration, hardware, test coverage, process, tools) to avoid fixating on the first theory.

Timeline

Reconstruct when the defect was introduced, when it could have been detected, and when it was detected. The gap shows which gate was missing.

Blameless post-mortem

Focus on systems and processes, not people. Output: root cause, contributing factors, detection gap, and dated corrective actions with owners.

A worked 5 Whys example

  1. Problem Watch standby battery drain doubled after the latest upstream drop.
  2. Why 1 The AP is waking about 40 times a minute during standby (batterystats, wakeup sources).
  3. Why 2 A system service registered an accelerometer listener without batching.
  4. Why 3 The upstream drop changed a default, and a vendor patch that set the report latency no longer applied, although the merge had no conflict.
  5. Why 4 The semantic change was not caught because no test checks sensor batching after a merge.
  6. Why 5 The promotion gate measured active-use power but not overnight standby drain.
  7. Corrective actions Fix the report latency; add a sensor-batching check to VTS-style vendor tests; add standby drain per hour as a mandatory power gate with a threshold against last good.
SeverityTypical definitionResponse
P0 / blockerDevice does not boot, data loss, security hole, emergency calls broken, widespread crashStop the line; daily or continuous triage; blocks promotion and release
P1 / criticalMajor feature broken or large KPI regression on many devicesBlocks promotion unless an explicit waiver is signed off
P2 / majorFeature degraded, workaround exists, limited scopeFix in the current release if possible; tracked against limits
P3 / minorCosmetic or rareBacklog, prioritised with other work

Crash clustering and triage at scale

Field and lab stability work is not one bug at a time. At fleet scale you cluster first, then spend people only on the clusters that move the crash-free rate.

  1. Collect Tombstones (/data/tombstones), ANR traces, DropBox, kernel panic/pstore, watchdog and subsystem-restart reasons, plus a build fingerprint and a small set of device properties. Bug reports for high-value clusters.
  2. Symbolize Native stacks need matching symbols for that exact build; a missing symbols tree makes every cluster look unique.
  3. Cluster Group by crashing process, faulting frame (or a stable stack signature), exception type, and often build or SoC. Ignore absolute addresses; they change with ASLR.
  4. Rank Volume × user impact (system_server and surfaceflinger beat a rare app). Watch new clusters after an OTA before chasing old long-tail noise.
  5. Own Assign a DRI per top cluster; the owner is the team whose code must change. Attach a representative report, not fifty duplicates.
  6. Close the loop A cluster is done when the signature disappears (or drops below a threshold) on the next build and a regression test or gate exists.
Tip A stability dashboard that only shows "crash-free percentage" without cluster IDs trains the team to celebrate a lucky week. Pair the rate with the top-N signatures and whether each is new, regressed or fixed.
Common pitfall Closing a systemic issue when a fix lands. It is only closed when the fix is verified on the original repro, the regression test or gate exists, and the corrective actions from the post-mortem are tracked. Otherwise the same class of bug returns on the next drop.
Interview angle "A systemic issue is reported by an external customer late in the program. How do you drive it to closure?" Interviewers want structure (ownership, reproduction, cross-layer evidence, bisection), people skills (one thread, named owners, cadence, escalation) and judgment (fix versus risk versus schedule, honest customer communication), ending with prevention (regression gate and post-mortem). For field stability, add clustering: signature, volume, owner, and "gone on the next OTA." Link the technical taps to Trace a Path Through the Android Stack.

Release management

Release management turns a stream of integrated builds into dated, supported releases for customers, with predictable quality. For a platform it covers Android version releases, maintenance releases, monthly security updates and OTAs.

Plan ──▶ Develop / integrate ──▶ Feature freeze ──▶ Stabilise ──▶ Code freeze
     ──▶ Release candidate(s) ──▶ Certification (xTS, carrier, regulatory) ──▶ Release
     ──▶ Staged OTA rollout (1% ▸ 10% ▸ 50% ▸ 100%) ──▶ Maintenance: SMRs, hotfixes ──▶ End of support
Analogy

Release management is like running a railway timetable. Trains (releases) leave at published times; passengers (features and fixes) that are not on the platform before the doors close wait for the next train rather than delaying everyone. Maintenance crews keep old lines running safely until they are retired. On a platform, the doors closing is the freeze, the next train is the next release or maintenance update, and maintenance crews deliver security patches and hotfixes on supported branches.

MilestoneEntry criteriaWhat changes are allowed
Feature complete / feature freezeAll planned features merged behind flags or enabledBug fixes; no new features
Code freezePromotion gates green on the stabilisation branchApproved fixes only, reviewed by a change control board
Release candidateNo open blockers; KPIs within targetsBlocker fixes only; every change forces a new RC and re-run of key tests
Release / golden buildCertification passed on the exact build; sign-offs completeNone; further fixes go to a maintenance release
MaintenanceRelease shippedSecurity patches (monthly security maintenance releases), critical field fixes

Key release practices

  • Change control board (CCB): after code freeze, a small group approves each change by weighing benefit against regression risk.
  • Go / no-go meeting: owners of each gate report status against published criteria; the release owner decides and records the decision.
  • Release notes: what changed, known issues, fixed issues, security patch level, required firmware versions.
  • Staged rollout and monitoring: ship OTA to a small percentage first, watch crash, ANR, OTA success and battery telemetry, then widen or halt.
  • Hotfix path: a fast, pre-agreed path for critical field fixes: fix on the release branch, minimal testing matrix, forward-port to main.
  • Security patch cadence: Android publishes a monthly security bulletin; OEMs and vendors integrate patches and set ro.build.version.security_patch, verified by STS. That bulletin cadence is not the same as a public AOSP source-drop cadence.

OTA mechanics: A/B, virtual A/B and update_engine

Shipping the golden build is only half of release management. Devices in the field take an OTA through update_engine, which downloads a signed payload, writes it safely, and reboots into the new slot.

IdeaWhat it means
A/B (seamless) updatesTwo slots (A and B). The device runs one slot while update_engine writes the other. A reboot flips ro.boot.slot_suffix. If the new slot fails to boot, the bootloader rolls back to the old slot.
Virtual A/BDynamic partitions in super are snapshotted (copy-on-write) instead of storing a full second copy of every large partition. Saves flash, but needs free space for snapshots; low-storage devices are a classic OTA failure class.
update_engineThe on-device daemon that fetches metadata, verifies signatures, applies the payload and reports success or failure. Sideload and recovery paths still exist for lab and brick recovery.
Full versus incrementalA full payload can install from a wide set of source builds. An incremental (delta) payload is smaller but is built for a specific source fingerprint. A missed source build in the delta set is a common "OTA not offered" bug.
# which slot is running, and the build you will target OTAs at
adb shell getprop ro.boot.slot_suffix
adb shell getprop ro.build.fingerprint

# watch apply progress and errors (service name varies by build; log is the reliable tap)
adb logcat -s update_engine update_engine_client

Staged rollout (1% → 10% → 50% → 100%) sits on top of this: halt on boot failure, crash-rate or OTA-success regressions, then ship a fixed payload. A/B rollback protects devices that fail to boot; it does not protect users from a boot-loop they already experienced, so pause the cohort.

Tip Scope, date and quality form an iron triangle: with fixed people, you can hold two. Make the trade-off explicit (move scope, move the date, or accept a documented risk) instead of letting quality silently lose. Life-critical paths such as emergency calling are never negotiable scope.
Interview angle Interviewers ask "a P1 appears the day before release; what do you do?" or "how do you decide go or no-go?" A strong answer uses the published criteria, quantifies the impact, lays out options (slip, ship with a known issue and a dated fix, disable the feature by flag, fix with a targeted change and partial re-test), and makes the decision transparently with stakeholders.

Delivering with distributed teams and customers

Platform programs span several sites and time zones, many domain teams that the integration lead does not manage directly, and external customers such as OEMs and carriers. Delivery depends on influence, clarity and a steady operating rhythm.

Analogy

Running a distributed program is like an air-traffic control network. Each tower (site) handles its own airspace, but flights (work items) cross between them, so handovers must be precise, written and standardised. Nobody shouts across the country; everyone reads the same flight plan. In a program, the flight plan is the single status of record, handovers are written follow-the-sun notes, and the controllers are DRIs with clear ownership.

Async-first communication

One source of truth for status, decisions written down, clear DRIs. Meetings are for decisions and unblocking, not status reading.

Follow-the-sun triage

Handoff notes at the end of each site's day (what was tried, what is next, what is blocked) so critical issues progress around the clock.

Influence without authority

Domain teams (camera, display, modem) do not report to the integration lead. Influence comes from owning the timeline, the dependency map and fair, specific asks.

Escalation

Escalate early with a clear ask, owner and date, not a status dump. Bad news travels the same day.

OEM, ODM and SoC: who owns what

A shippable device is a three-party (sometimes four-party) software stack. Interviewers for integration leads expect you to place yourself in this model without naming your employer.

PartyTypical ownershipWhat they wait on
SoC vendorSilicon, BSP, kernel modules, vendor HALs, modem and DSP firmware, per-chipset baselinesAOSP tags, HAL and VINTF changes, GKI/KMI moves, vendor API freeze windows
ODMBoard, RF, sensors, factory image, sometimes a thin software overlayA bring-up-ready BSP and a meta-build recipe that matches the hardware SKU
OEMBrand, product apps and UX, GMS/MADA path, carrier SKUs, OTAs, customer supportA stable vendor baseline, known-issue lists, merge windows, honest dates
Carriers / labsPTCRB/GCF and operator acceptance, field issuesA certifiable user build and a DRI who answers with logs

The integration lead rarely manages these organisations. Influence comes from one status of record, a dependency map, and fair, specific asks. When a bug bounces, name the contract that failed (HAL, KMI, QMI, CarrierConfig, meta-build pin) rather than the company.

Working with internal and external customers

StakeholderWhat they care aboutHow to serve them
Program and product managementOne status, dates, risksWeekly red/amber/green against milestones and KPIs; exceptions highlighted
Domain engineering teamsClear priorities, clean bug reports, less ping-pongTriaged bugs with logs and a localised layer; named owners; fair deadlines
OEM customersStable baselines, fast answers on field issues, merge windowsDefined delivery cadence, a single point of contact, change-request tracking, transparent known-issue lists
ODM and factoryA flashable combination that matches the SKUMeta-build pins, partition layout, factory vs user variants
Carriers and certification labsCertification pass, field issues, honest datesDaily loop on blockers during certification, logs rather than promises
  • Customer change requests (CRs): track every customer-reported issue with severity, owner, target release and status; review aging CRs weekly.
  • Expectation management: never promise a date you do not control; give a date for the next update instead.
  • Status that executives read: what changed since last week, what is blocked, and what you need from them, in a few lines.
Interview angle Behavioural questions here ("how do you keep a distributed program on track?", "a customer is escalating, what do you do?") are scored on concrete mechanisms: one status of record, DRIs, cadence, written decisions, early escalation, and honest communication. Use the STAR format with a measurable result.

Metrics and KPIs

Two families of metrics matter: product KPIs (how good the device is) and delivery and integration metrics (how healthy the process is). A lead tracks both and makes regressions visible early.

Product KPIs for a wearable

Power

Battery life in days, standby drain per hour, active-use drain, wake locks and wakeups, suspend residency, always-on display draw, sensor-hub offload ratio, current (mA) budget per use case.

Stability

Crash and ANR rate, kernel panics, watchdog and subsystem restarts, boot success rate, OTA success rate, field crash-free percentage.

Performance and UX

Boot time, app and tile launch latency, jank and frame drops, touch and wrist-raise responsiveness, wake-to-render latency.

Connectivity

Bluetooth reconnection time and drop rate, Wi-Fi and LTE reliability, call drop rate, notification delivery latency.

Integration and delivery metrics

MetricWhat it shows
Upstream integration lead timeDays from an AOSP tag release to a promoted baseline containing it
Carried delta sizeNumber of vendor patches on top of upstream; lower is cheaper to maintain
Conflict count and resolution time per dropHow painful each drop is; trend should go down
Build green percentage and time to fix breakageCI health and team discipline
Test pass rate and flaky-test rateSignal quality of the test system
xTS pass rate per baselineCertification readiness over time
Defect inflow and closure rate, open P0/P1 countWhether the release is converging
Mean time to resolve (MTTR) and CR agingResponsiveness to customers and systemic issues
Escaped defectsBugs found by customers or in the field that gates should have caught; each one should lead to a new gate
On-time deliveryPercentage of baselines and releases delivered on the committed date with the quality bar met
Analogy

Metrics are the dashboard of a car: the speedometer shows how fast you are going (delivery speed), the fuel gauge shows what is left (schedule and capacity), and warning lights show problems before the engine fails (regressions). Driving without them is guessing. On a platform, the speedometer is integration lead time, the fuel gauge is open defects against time to release, and the warning lights are KPI regressions against the last good build.

Common pitfall Measuring what is easy rather than what matters, or letting a metric become a target that people game (for example closing bugs as "cannot reproduce" to improve closure rate). Pair each metric with a quality check and review the trend, not a single number.
Interview angle "Which KPIs would you track?" Interviewers want both product KPIs with example thresholds (for example standby drain per hour against last good) and process metrics (lead time, build health, escaped defects), and how you would act on them: which ones gate promotion and which ones drive improvement work.

Build system: Soong, Make, lunch and product packages

An integration lead does not have to write every Android.bp, but must know how a product image is selected, what gets packaged, and which build system owns userspace versus the kernel. Wrong claims here are a common interview fail.

Soong (Android.bp)

  • Primary userspace build for modules: binaries, libraries, apps, HALs, APEX
  • Declarative Blueprint modules (cc_binary, cc_library, android_app, aidl_interface)
  • Faster incremental analysis than Make; generates Ninja
  • Google's attempt to migrate this userspace graph to Bazel was halted around 2023; Soong remains primary

Make and product config

  • Product and board configuration still live in Make: device.mk, AndroidProducts.mk, BoardConfig.mk
  • Legacy Android.mk modules still exist; new modules should be Soong
  • PRODUCT_PACKAGES (and PRODUCT_PACKAGES_DEBUG) is how a product decides what is on the image
  • Do not say "Android is moving to Bazel" as a current userspace fact

Kleaf is the Bazel-based build for the Android kernel / GKI. That path is real and separate from the halted userspace Bazel migration. In an interview, split the two: Soong for userspace modules, Kleaf/Bazel for kernel images and modules, Make for product composition.

# select a product, a release configuration and a variant
# modern lunch: PRODUCT-RELEASE-VARIANT
lunch aosp_cf_x86_64_phone-trunk_staging-userdebug

# older trees used PRODUCT-VARIANT (two tokens), for example aosp_arm64-userdebug
# always check the tree you are in; do not memorise one combo as eternal

# what the product actually ships
# device.mk / *.mk
PRODUCT_PACKAGES += \
    FooService \
    vendor.bar.hal-service

PRODUCT_PACKAGES_DEBUG += \
    FooDebugTool
TokenMeaning
ProductThe device or lunch target: board, partitions, overlays, PRODUCT_PACKAGES (for example aosp_cf_x86_64_phone).
Release configWhich aconfig / release set is enabled (trunk, a named Android release). This is the third token on recent trees.
Variantuser (ship and certify), userdebug (debuggable, usual day-to-day), eng (fastest iterate, not a ship image).
// Android.bp — a vendor HAL binary
cc_binary {
    name: "android.hardware.foo-service",
    srcs: ["main.cpp"],
    shared_libs: ["libbinder_ndk", "android.hardware.foo-V1-ndk"],
    vendor: true,
    init_rc: ["android.hardware.foo-service.rc"],
    relative_install_path: "hw",
}
Analogy

Soong modules are the recipes for each dish; PRODUCT_PACKAGES is the menu for tonight's restaurant (the product); lunch is seating you in a specific restaurant, on a specific seasonal menu, in kitchen or dining-room mode (the variant). Bazel/Kleaf is a different kitchen used for the kernel, not a replacement of the userspace kitchen.

Common pitfall Claiming every kernel update forces a vendor-module rebuild, or that "Android is moving to Bazel" in userspace. Same-KMI GKI updates keep existing modules; Soong is still the userspace build. Also: a package that is built is not necessarily installed — if it is missing from PRODUCT_PACKAGES it will not be on the image.
Interview angle "How do you produce a product image?" Walk lunch (product-release-variant), PRODUCT_PACKAGES, Soong versus leftover Make, and where kernel builds sit (Kleaf). Then say who owns a missing binary: the module's Android.bp, or the product makefile that forgot to package it.

OS upgrades, vendor API level and freeze windows

A major Android upgrade is not "merge AOSP and ship." Treble already lets a new framework talk to an older vendor implementation if the declared HAL and VINTF versions still match. Android later made that contract more explicit with a vendor API level and a freeze window (often discussed as a GRF-style Google requirements freeze): vendor code can stay on a frozen vendor API while the system image moves forward for a documented number of releases. Exact year counts and brand names of the program change; in an interview, state the mechanism and say you would check the current CDD and vendor-API release notes rather than quoting a memorised policy year.

system / product image     ── newer Android, newer SDK, Mainline modules
        │
        │  VINTF + vendor API level (ro.vendor.api_level)
        │  freeze window: vendor may stay put
        ▼
vendor / odm image         ── HALs, kernel modules, sepolicy, firmware pins
        │
        │  KMI freeze on this ACK branch
        ▼
GKI + vendor modules       ── same KMI: no module rebuild; new ACK: rebuild
  • Vendor API level is the API the vendor partition was built against. It can lag the system API level inside the freeze window. Mismatches outside that window fail VINTF / OTA compatibility checks.
  • Why it exists: OEMs want a yearly OS upgrade without a full SoC vendor rebase of every HAL and driver. The SoC vendor wants a stable contract so one vendor image supports several system years.
  • What still moves on an upgrade: framework behaviour, CDD requirements, CTS/VTS/GTS for the new release, SELinux, Mainline modules, and any HAL the new compatibility matrix newly requires.
  • What can stay: vendor HAL implementations, kernel modules on the same KMI, and firmware pins, if the freeze and VINTF allow it and tests still pass.
Analogy

Think of a rented shop (the vendor image) inside a mall that renovates every year (the system image). A freeze window is a lease that says the shopfront sockets stay the same for N renovations, so the tenant does not have to rebuild the shop each year. When the mall changes the sockets (new vendor API or new KMI), the tenant rebuilds. VINTF is the inspector who checks the plugs before opening day.

Interview angle "How would you plan an OS upgrade for a device already in the field?" Cover: which vendor API the device is frozen at, whether the new system is inside the freeze window, which HALs the new matrix newly requires, GKI/KMI branch, xTS delta, GMS/CDD delta, OTA payload type (full vs incremental), and a staged rollout. Do not invent a specific 2025/2026 freeze length.

Quick revision

  • HLOS is the high-level OS on the application processor: Linux kernel plus Android or Wear OS user space.
  • Non-HLOS is firmware on other processors: modem (MPSS), ADSP and CDSP, sensor hub (SLPI), TrustZone, XBL boot loaders, power management processor.
  • The HLOS image is AOSP framework plus vendor BSP (kernel, device tree, drivers), vendor HALs, proprietary libraries and OEM customisation.
  • A meta build pins one tested version of every HLOS and non-HLOS component into a flashable release.
  • HLOS and non-HLOS meet at HALs, remote-processor drivers, shared memory and protocols such as QMI.
  • AOSP source is published on android.googlesource.com; public tag cadence has changed over the years — check current release notes. The monthly security bulletin is a separate patch list.
  • Upstream integration lands each AOSP tag onto vendor trees and per-chipset baselines while carrying vendor deltas forward.
  • Treble, VINTF, stable AIDL HALs and GKI are what make frequent upstream integration feasible.
  • Within the same frozen KMI, GKI kernel updates do not require vendor modules to rebuild; a new ACK/KMI does.
  • Soong + Android.bp is the primary userspace build; Make still composes products (PRODUCT_PACKAGES). Userspace Bazel migration was halted ~2023; Kleaf/Bazel is for kernel builds.
  • lunch selects product-release-variant; user ships, userdebug is the usual debug image. A built module is not on the image unless it is in PRODUCT_PACKAGES.
  • Common conflict sources: HAL version bumps, framework API changes, SELinux policy, build system changes, kernel branch (KMI) moves, toolchain updates.
  • Upstream-first reduces the carried delta and the cost of every future drop.
  • Branching models: trunk-based with feature flags, release branches, feature branches, per-SoC and per-OEM branches.
  • Freeze stages: feature freeze, code freeze, release candidate, golden build.
  • Merge preserves history and SHAs; rebase keeps a clean, linear patch stack but rewrites SHAs.
  • Cherry-pick ports individual fixes; use -x to record the source commit.
  • Semantic conflicts apply cleanly but are wrong; only builds, tests and informed review catch them.
  • git rerere reuses conflict resolutions; git bisect finds the introducing change.
  • CTS tests app-facing APIs against the CDD; VTS tests HALs, kernel and vendor interfaces; GTS tests GMS requirements; CTS-Verifier covers manual hardware checks; STS tests security patches.
  • GMS is licensed (MADA); GTS is the licensee test suite. Compatibility tests are not the same as a GMS licence.
  • Carrier cert (PTCRB/GCF plus operator labs) is independent of xTS; many per-SIM differences are CarrierConfig.
  • CTS-on-GSI proves the vendor side works with a generic AOSP system image.
  • Vendor API level plus a GRF-style freeze window can let a vendor image pair with newer system images; confirm current CDD/release notes rather than memorising a year count.
  • OTA: update_engine applies a signed payload to the inactive A/B slot (virtual A/B snapshots dynamic partitions). Full payloads are source-flexible; incremental payloads are smaller but source-specific.
  • SoC / ODM / OEM / carrier own different layers; the integration lead owns the contracts and the one status of record, not the org chart.
  • Crash triage at scale: symbolize, cluster by stable stack signature, rank by volume × impact, one DRI per top cluster, done when the signature is gone.
  • xTS runs on Tradefed; run subsets continuously and full suites on promotion candidates.
  • Promotion gates: build health, boot, VINTF, smoke, compliance, power, stability, performance, open defects.
  • Gates must be measurable, compared against last good, owned, and stable under schedule pressure.
  • Platform CI: presubmit build and quick tests, postsubmit full builds, nightly xTS, KPI and soak runs.
  • Keep bisection possible: stage large drops, store manifest snapshots, automate culprit finding.
  • Systemic triage: own it, reproduce, instrument, localise, drive, decide, close.
  • The owner of a bug is the team whose code must change, not the team that saw the symptom.
  • RCA tools: 5 Whys, fishbone diagrams, defect timelines, blameless post-mortems with dated actions.
  • An issue is closed only after verification, a regression test or gate, and tracked corrective actions.
  • Release practices: change control board, go/no-go against published criteria, release notes, staged OTA rollout, hotfix path.
  • Scope, date and quality: with fixed people, make the trade-off explicit.
  • Distributed delivery: async-first, single status of record, DRIs, follow-the-sun handoffs, early escalation.
  • Track product KPIs (power, stability, performance, connectivity) and process metrics (lead time, build health, escaped defects, MTTR).

Glossary

A/B update
Seamless OTA that writes the inactive slot and reboots into it, with bootloader rollback if the new slot fails.
ADSP / CDSP
Qualcomm Hexagon digital signal processors for audio and compute workloads; they run non-HLOS firmware.
Android.bp
Soong module definition file for userspace binaries, libraries, apps and interfaces.
AOSP
Android Open Source Project; Google's open-source Android code base, published on android.googlesource.com.
Baseline
A promoted, tested snapshot of the platform for a chipset that downstream teams and customers build on.
BSP
Board Support Package; kernel, device tree, drivers and low-level software for a specific SoC and board.
CarrierConfig
Per-carrier Android configuration (feature flags and values) keyed by MCC/MNC or a carrier app.
CCB
Change Control Board; group that approves changes after code freeze.
CDD
Compatibility Definition Document; the written requirements an Android-compatible device must meet.
Cherry-pick
Applying a single commit from one branch onto another.
Code freeze
Milestone after which only approved fixes may enter a release branch.
CR
Change Request; a tracked issue or request, often from a customer.
CTS
Compatibility Test Suite; tests public Android APIs and behaviour.
CTS-Verifier
Mostly manual CTS tests for hardware-dependent behaviour that automation cannot fully judge.
DRI
Directly Responsible Individual; the single person accountable for driving an item to closure.
Escaped defect
A bug found by customers or in the field that internal gates missed.
Feature flag
A runtime or build-time switch that hides unfinished features so code can merge early (aconfig in AOSP).
GCF
Global Certification Forum; industry cellular-device certification catalogue used widely outside PTCRB.
Gerrit
Web-based code review system used by AOSP and most platform teams.
GKI
Generic Kernel Image; Google-built kernel core with vendor code in modules against a stable kernel module interface. Same-KMI updates do not require those modules to rebuild.
GMS
Google Mobile Services; licensed Google apps and services such as Play. Not part of AOSP.
Golden build
The exact build approved for release and certification.
GRF
Google Requirements Freeze (also discussed as a vendor-API freeze window): vendor implementations may stay on a frozen vendor API while the system image moves forward for a documented period. Check current release notes for the window length.
GSI
Generic System Image; pure AOSP system image used to validate Treble compliance.
GTS
GMS Test Suite; tests requirements for shipping Google Mobile Services.
HLOS
High-Level Operating System; the Linux kernel plus Android or Wear OS on the application processor.
Keystone
Qualcomm's name for its program that tracks AOSP upstream releases onto its chipset baselines.
Kleaf
Bazel-based build for the Android kernel / GKI; separate from the halted userspace Bazel migration.
KMI
Kernel Module Interface; the stable set of kernel symbols vendor modules may use with GKI. Frozen for the life of an ACK branch.
lunch
The AOSP command that selects a product, release configuration and build variant before compiling.
MADA
Mobile Application Distribution Agreement; the commercial contract under which an OEM may preload GMS.
Merge
Joining two branches' histories with a merge commit.
Meta build
A versioned bundle pinning one tested build of every HLOS and non-HLOS component.
MPSS
Modem Processor Subsystem; the modem's firmware and processor.
MTTR
Mean Time To Resolve (or Repair); average time from report to fix.
Non-HLOS
Firmware and RTOS images for processors other than the application processor.
ODM
Original Design Manufacturer; typically owns the board, RF and factory image, sometimes a thin software overlay.
OEM
Original Equipment Manufacturer; owns the brand, product software, GMS path, carrier SKUs and OTAs.
Post-mortem
A blameless review after an incident that records root cause and corrective actions.
Presubmit
Checks run on a change before it is merged.
PRODUCT_PACKAGES
Make variable listing modules installed on a product image. Built is not the same as packaged.
Promotion gate
Pass or fail criteria a build must meet to move to the next stage.
PTCRB
North-American operator forum for cellular device certification.
RCA
Root-Cause Analysis; finding why a problem happened, not only what failed.
Rebase
Replaying commits on top of a new base, producing a linear history with new commit IDs.
repo
Google's tool for managing many Git repositories from a single manifest.
rerere
Git's "reuse recorded resolution" feature, which re-applies earlier conflict resolutions.
SMR
Security Maintenance Release; an OEM or vendor update that delivers security bulletin patches. Cadence is a product choice; the Android security bulletin itself is monthly.
Soong
Primary AOSP userspace build system; modules are declared in Android.bp.
STS
Security Test Suite; checks that security bulletin patches are present.
Tradefed
Trade Federation; the Android test harness used to run xTS and device tests.
Treble
Android architecture that separates framework and vendor code behind stable interfaces.
TrustZone
Arm's secure-world technology; runs the trusted execution environment as non-HLOS firmware.
update_engine
On-device daemon that downloads, verifies and applies OTA payloads to the inactive slot.
Upstream-first
Policy of contributing changes to AOSP or Linux before, or instead of, carrying them privately.
Vendor API level
The API level the vendor partition was built against (ro.vendor.api_level); may lag the system API inside a freeze window.
Virtual A/B
A/B updates that snapshot dynamic partitions with copy-on-write instead of storing a full second copy of super.
VINTF
Vendor Interface object; manifests and compatibility matrices that declare and check HAL versions.
VTS
Vendor Test Suite; tests HALs, kernel and vendor interfaces.
xTS
Collective name for CTS, VTS, GTS, STS and related suites.

Interview questions

Fundamentals

What is an HLOS image?

HLOS means High-Level Operating System: the Linux kernel plus Android or Wear OS user space running on the main application processor. The HLOS image is the complete Android software build for a device: the AOSP framework and apps, the vendor BSP (kernel, device tree, drivers), vendor HALs, proprietary libraries and OEM customisation. It is distinct from the non-HLOS firmware that runs on the modem, DSPs and other processors.

What is non-HLOS software? Give examples.

Non-HLOS software is firmware or a real-time OS running on processors other than the application processor, or in the secure world. Examples: modem firmware (MPSS), audio and compute DSP firmware (ADSP, CDSP), sensor hub firmware (SLPI), TrustZone and the hypervisor, the XBL boot loaders, the always-on power management processor, and Wi-Fi and Bluetooth firmware. Each is built and versioned by its own team.

What is a meta build?

A meta build is the integration record that combines one specific, tested version of every software component (the HLOS build and each non-HLOS firmware build) into a single release, along with the partition layout and flashing information. It ensures that the combination flashed onto a device is known to work together. Many hard integration bugs are caused by mismatched component versions, which the meta build is designed to prevent.

What does upstream integration mean in the Android context?

It means taking a new release from the upstream source (Google's AOSP, or the Linux kernel) and bringing it into your own code base, which contains your changes on top. You merge or rebase the new upstream code, resolve conflicts with your carried changes, update interfaces, build and test, and promote the result as a new baseline for your products.

What is Project Treble and why does it help integration?

Treble, introduced in Android 8, separates the Android framework (system partition) from vendor implementation (vendor partition) behind stable, versioned HAL interfaces. The framework can then be updated without rewriting vendor code, as long as the HAL versions remain compatible. This makes upstream drops and OS upgrades much cheaper and is verified by VTS and by booting a Generic System Image.

What is VINTF?

VINTF (Vendor Interface object) is a set of manifests and compatibility matrices. The device manifest lists the HALs and versions the vendor provides; the framework compatibility matrix lists what the framework requires, and vice versa. The build and the OTA system check that they match, and the device refuses incompatible updates. A HAL version mismatch is a classic cross-team dependency during an upstream drop.

What is GKI?

GKI, the Generic Kernel Image, is a kernel core built by Google from the Android Common Kernel for each supported branch. Vendors put SoC and board-specific code into loadable kernel modules that use only the stable Kernel Module Interface (KMI). This separates kernel updates from vendor changes. Within the same frozen KMI, a core-kernel update does not require those modules to be rebuilt. A rebuild is needed when the KMI changes (new ACK branch) or a module needs a new symbol.

What is the difference between git merge and git rebase?

A merge combines two branches by creating a merge commit that has both histories as parents; existing commit IDs are unchanged. A rebase takes your commits and replays them one by one onto a new base, creating new commits with new IDs and a linear history. Merge is safer for shared branches; rebase gives a cleaner history and a tidy patch stack but rewrites history that others may depend on.

What is a cherry-pick and when do you use it?

A cherry-pick applies the changes of one specific commit onto another branch as a new commit. It is used to port a bug fix from main to a release branch (or the reverse) without bringing other changes. Use git cherry-pick -x to record the original commit ID in the message for traceability, and check for dependent commits that must also be picked.

What is the repo tool?

repo is Google's wrapper around Git that manages the hundreds of Git repositories that make up an Android tree. A manifest XML file lists each project, its path, remote and branch or revision. repo init selects a manifest, repo sync fetches all projects, and repo manifest -r produces a snapshot of exact revisions so a build can be reproduced.

What is CTS?

CTS, the Compatibility Test Suite, is Google's automated test suite for the public Android APIs and behaviours required by the Compatibility Definition Document. Passing it shows that apps written against the Android SDK will behave correctly on the device. Passing CTS is required to be Android-compatible and to license Google Mobile Services.

What is VTS?

VTS, the Vendor Test Suite, tests the vendor side of the Treble boundary: HAL implementations, VINTF compliance, kernel configuration and GKI requirements, and vendor partition behaviour. It ensures the vendor implementation keeps the contract the framework relies on, so a generic framework can run on the device.

What is GTS?

GTS, the GMS Test Suite, checks requirements for devices that ship Google Mobile Services, such as Google Play services and Play Store behaviour, preloaded Google apps and related configuration. It is distributed to GMS licensees, not published in AOSP. A device must pass CTS, VTS and GTS (plus others such as STS) to ship with GMS.

What is a promotion gate?

A promotion gate is a set of pass or fail criteria a build must meet to move to the next stage, for example from an integration branch to a baseline delivered to customers. Typical criteria are successful builds on all targets, boot on all SKUs, smoke tests, xTS pass rates, power, stability and performance within thresholds compared to the last good build, and no open P0 or P1 regressions.

What is the difference between presubmit and postsubmit testing?

Presubmit tests run on a change before it is merged, usually a targeted build and fast tests, so obviously broken changes never land. Postsubmit tests run after merging, on the combined tree, and include full builds on all targets and slower device tests. Presubmit protects the branch; postsubmit catches interactions between changes and anything too slow for presubmit.

What is root-cause analysis?

Root-cause analysis (RCA) is a structured investigation into why a problem happened, going past the immediate technical fault to the design, process or test gap that allowed it. Its output is the root cause, contributing factors, why it was not detected earlier, and corrective actions with owners and dates. Tools include 5 Whys, fishbone diagrams and timelines.

Explain the 5 Whys technique.

You state the problem and ask "why did this happen?", then ask "why?" about each answer, typically about five times, until you reach a cause that, if fixed, prevents the problem from recurring (often a process or test gap). Each answer should be supported by evidence, not guesses, and there can be multiple branches. It stops teams from fixing only the surface symptom.

What is a DRI?

A DRI (Directly Responsible Individual) is the one named person accountable for driving an issue or deliverable to completion. They do not have to do all the work, but they coordinate, track, escalate and make sure it closes. Having a DRI prevents problems from bouncing between teams with nobody driving them.

What is a code freeze?

A code freeze is a milestone after which only approved changes (usually bug fixes for release-blocking issues) may enter the release branch. It stabilises the code so testing results remain valid. Changes after freeze typically need a change control board's approval and may trigger re-running key tests.

What is Soong, and how does it differ from Make?

Soong is the primary AOSP userspace build. Modules are declared in Android.bp (Blueprint) and compiled via Ninja. Make still composes the product: device.mk, BoardConfig.mk and PRODUCT_PACKAGES decide what is on the image, and some leftover Android.mk modules remain. Google's userspace Bazel migration was halted around 2023; Soong is still primary. Kernel builds use Kleaf (Bazel), which is a different path.

What does lunch select?

The build target. Recent trees use PRODUCT-RELEASE-VARIANT (for example a Cuttlefish phone, a trunk or named release config, and userdebug). Older trees used two tokens, PRODUCT-VARIANT. The variant is user (ship and certify), userdebug (usual debug image) or eng. Always check the tree you are in rather than memorising one combo.

What is PRODUCT_PACKAGES?

The Make list of modules installed on that product image. A module can build successfully and still be absent from the device if nobody added it to PRODUCT_PACKAGES (or PRODUCT_PACKAGES_DEBUG for debug-only tools). When a binary is "missing on the image," check the product makefile before rewriting the Android.bp.

What is GMS versus MADA?

GMS is the licensed Google apps and services (Play Store, Play services and related apps). MADA is the commercial agreement under which an OEM may preload them. GTS is the technical test suite for that licence. CTS/VTS prove Android compatibility; they do not grant GMS. Do not quote confidential placement rules; say you would check current licensee docs and the CDD.

What is the difference between a full OTA and an incremental OTA?

A full payload can apply from a wide set of source builds. An incremental (delta) payload is smaller and is built for a specific source fingerprint. update_engine writes the inactive A/B slot (virtual A/B uses snapshots for dynamic partitions). If a device's source build is not in the delta set, it will not be offered that incremental and needs a full payload or another delta.

Going deeper

Walk me through integrating a Google upstream release into a vendor chipset baseline.
  1. Track the AOSP release tag and read release notes for API, HAL, SELinux, build and kernel changes.
  2. Do an early trial merge (for example on preview tags) to estimate conflicts and dependencies.
  3. Plan the merge or rebase strategy per repository and sequence the drop across affected SoC baselines.
  4. Run dependency analysis: assign each conflict or breakage to its owning domain (framework, BSP, HAL, apps, build) with a date.
  5. Update HAL implementations and VINTF manifests for required interface versions.
  6. Gate on build health, boot, smoke, VINTF, xTS, and power, stability and performance against last good.
  7. Promote the baseline, publish release notes and known issues, and hand off to OEM teams on the agreed schedule.

Throughout, keep one status of record and coordinate multimedia, camera, connectivity, display, power and security stakeholders across sites.

What goes into composing an HLOS image for a wearable, and how does it differ from a phone?

The ingredients are the same: AOSP and Wear OS framework, vendor BSP, vendor HALs, proprietary libraries, sensor hub firmware and OEM customisation, branched per chipset. The differences are the constraints: smaller RAM, flash and display, a dominant power budget, an always-on co-processor handling sensors and ambient display, health sensors as first-class components, a companion and Bluetooth connectivity model (and eSIM or LTE for standalone watches), and tiles, complications and watch faces as the main user surface. The power architecture shapes what is included and how it is tuned to the ship gate.

When would you choose merge over rebase for an upstream drop, and vice versa?

Choose merge when the branch is shared by many teams and downstream branches, because a merge does not rewrite commit IDs and conflict resolution happens once. Choose rebase when you want to keep the vendor delta as a clean, reviewable patch stack on top of upstream (common for kernel trees or small repositories), which makes it easy to see and upstream your changes. Many organisations mix them: merge for large shared framework repositories, rebase for patch-stack style kernel or HAL repositories.

What is a semantic conflict and how do you catch it?

A semantic conflict is when two changes merge without any textual conflict, but the combined code is wrong: for example upstream changes a default value, renames a behaviour or moves a permission check, and a vendor patch that depended on the old behaviour still applies. Git cannot detect this. You catch it with builds, unit and integration tests, xTS, KPI regression runs, and reviewers who understand why each carried patch exists. Reviewing the upstream diff for areas touched by carried patches helps too.

How do you resolve hundreds of merge conflicts on a tight schedule?

Do not have one person resolve everything. Pre-scan conflicts early, classify them by domain and type, and assign each group to the team that owns the carried change, with deadlines and a tracking list. Resolve easy mechanical conflicts centrally, and send semantic or interface conflicts to domain experts. Build and test each resolution, reuse earlier resolutions with rerere, and drop vendor patches that upstream has made unnecessary. Report progress daily against the list.

Which branching model would you use for a platform serving several chipsets and OEMs?

A main development branch that receives upstream drops and new features, preferably with feature flags to keep it close to trunk-based. Release branches cut from main per Android version and chipset baseline at feature freeze, which accept only fixes. OEM or product branches for customer-specific customisation, kept as thin as possible. Fixes land on main and are cherry-picked to supported release branches (or the reverse, but consistently), with automated checks that nothing is missed. Branches are retired on a published schedule.

How do you make sure a fix on a release branch is not lost on main?

Require every fix to reference a bug, and track per-branch status on that bug. Use cherry-pick -x or Change-Id matching so tools can compare branches, and run an automated forward-merge or "missing fixes" report that lists changes present on a release branch but not on main. Review the report as part of the release checklist. Without this, the same bug reappears in the next release.

What promotion gates would you enforce before an upstream drop lands in the baseline?
  • Build health across all SoC baselines and build variants.
  • Boot success on all SKUs, including repeated boot cycles.
  • VINTF compatibility and core VTS.
  • Smoke or BAT covering calls, data, Wi-Fi, Bluetooth, sensors, display, camera, OTA.
  • CTS, VTS and GTS pass rate at or above the previous baseline.
  • Power: standby drain, suspend residency, wake lock and wakeup budget versus last good.
  • Stability: crash, ANR, panic, watchdog and subsystem restart rates.
  • Performance: boot time, launch latency, jank.
  • No open P0 or P1 regressions without a signed-off waiver.
How do you handle a CTS failure found close to release?

First triage: is it a device bug, a test bug, a test environment problem (network, SIM, lab setup) or flakiness? Re-run the single module with Tradefed to confirm, and compare with a previous passing build to find the introducing change. If it is a device bug, assign it to the owning team as a blocker. If it is a genuine test bug, collect evidence and request a waiver through the official process. Either way, record it and add the module to continuous CI so it is caught earlier next time.

What is CTS-on-GSI and why run it?

CTS-on-GSI means flashing Google's Generic System Image (a pure AOSP system partition) on the vendor's device and running CTS. If the vendor implementation follows Treble correctly, the generic framework should work with the vendor partition. Failures show that the vendor side relies on system partition modifications or breaks the interface contract, which would make future framework updates expensive.

How would you set up CI for a large multi-repository platform?

Use Gerrit with presubmit that builds affected targets and runs fast tests and static analysis, with atomic submission for changes that span repositories (topics). Postsubmit builds all targets continuously and runs boot and smoke tests on real devices. Nightly or per-candidate builds run xTS subsets, power and performance benchmarks and stability soaks. Add remote build caching, automated culprit finding, flaky-test quarantine, a revert-first policy for breakages, and store manifest snapshots and artifacts for every build.

How do you deal with flaky tests in platform CI?

Measure flakiness automatically (for example a test that fails and then passes on retry without code changes). Quarantine flaky tests from blocking gates, but keep running them and assign an owner and a deadline to fix or delete them. Track the flaky rate as a metric. Never simply retry until green, because that hides real intermittent bugs such as race conditions, which are often real device defects.

How do you decide which team owns a cross-domain bug?

By layer and evidence. Reproduce the bug, capture logs across boot, kernel, HAL and framework on one timeline, and bisect builds to the introducing change. Map the failing component to the team whose code must change. If it is genuinely shared, such as an interface contract between a HAL and firmware, assign one DRI and have the other teams co-own actions rather than letting the bug bounce. The goal is to remove ambiguity quickly so engineers fix instead of argue.

A systemic issue is reported by an external customer late in the program. How do you drive it to closure?

Take ownership as the DRI and acknowledge the customer quickly with a time for the next update. Reproduce and scope the severity and KPI impact. Pull the right cross-domain experts into one triage thread, build a cross-layer log timeline, and bisect to root cause. Weigh fix versus risk versus schedule and agree the plan with the customer. Communicate on a fixed cadence to internal teams and the customer. Close with the verified fix, a regression test or gate, and a post-mortem.

What makes a good post-mortem?

It is blameless and factual: a timeline of when the defect was introduced, when it could have been detected, and when it was detected and fixed. It identifies the root cause and contributing factors (technical and process), the detection gap (which gate or test was missing), and specific corrective actions with owners and dates. It is shared widely, and the actions are tracked to completion rather than forgotten.

What happens at a go/no-go meeting?

Each gate owner reports status against the published release criteria: build, xTS, KPIs, open defects, certification, and customer or carrier sign-offs. Known risks and waivers are reviewed explicitly. The release owner makes the decision (go, no-go, or go with conditions) and records it with reasons. A good meeting is short because the data was prepared in advance; debates about criteria belong before the meeting, not in it.

How do you run a staged OTA rollout?

Release the update to a small percentage of devices first (for example 1 percent), then widen in steps (10, 50, 100 percent) if health metrics are good. Monitor OTA success rate, boot success, crash and ANR rates, battery telemetry and customer-reported issues against the previous build. Define halt criteria in advance, and be ready to pause the rollout and ship a fix. A/B updates with rollback make each step safer.

How do you keep a geographically distributed program on track?

Work async-first with one source of truth for status, clear DRIs and written decisions. Use follow-the-sun triage with structured handoff notes so critical issues move forward around the clock. Keep a regular status cadence with internal teams and external customers, highlighting what changed, what is blocked and what help is needed. Escalate blockers early with a clear ask, owner and date. Rotate meeting times so the same site is not always inconvenienced.

Which KPIs would you gate a wearable release on?

Battery life and standby drain per hour against the last good build, suspend residency and wakeup counts, wake-to-render latency, crash and ANR rates, kernel panic and watchdog reset rates, boot and OTA success rates, thermal ceiling under sustained load, and connectivity reliability (Bluetooth reconnection, notification delivery). Define regression thresholds in advance and hold promotion if any exceed them.

How do you ramp up quickly on an unfamiliar platform or domain?

Map the system and its interfaces first: the image composition, branches, build and test pipelines, and KPI dashboards. Find the two or three people who hold the most context and learn from them. Sit in triage meetings to learn the current problems. Land a small real change early to learn the pipeline end to end. Keep a list of unknowns and burn it down deliberately, and write down what you learn so the next person ramps faster.

Advanced

How do you sequence one upstream drop across several SoC generations?

Start with the chipset that is most representative and best staffed (often the newest, which will ship the release first) as the lead baseline. Resolve common framework and HAL conflicts there once, then apply the same resolutions to other baselines, handling only their BSP-specific deltas. Older chipsets may stay on an earlier kernel branch or HAL version, so check VINTF and GKI support per chipset. Stagger promotions so teams are not overloaded, and publish the schedule to OEMs.

How does GKI change kernel integration work for a vendor?

Before GKI, each vendor carried a large, forked kernel with thousands of patches, and every kernel update meant a painful forward-port. With GKI, the core kernel image comes from Google's Android Common Kernel, and vendor code lives in modules that may only use symbols in the KMI symbol list. Integration work shifts to keeping modules compatible with the frozen KMI, requesting new symbols through the upstream process, and validating with VTS kernel tests. Same-KMI security and bug-fix kernel updates can be taken without rebuilding vendor modules. Modules rebuild when you move to a new ACK/KMI, or when you need a new symbol or a KMI break (which ABI tooling should reject on a frozen branch).

What is a HAL interface bump, and how do you manage it during an upstream drop?

A new Android release may require a newer version of a HAL (or migration from HIDL to AIDL) to support new features, as defined in the framework compatibility matrix. The vendor HAL owner must implement the new version, update the device manifest, and pass VTS for it. As integration lead, identify these requirements from the release notes and compatibility matrix early, create tracked items per HAL with owners and dates, and decide whether the drop can land with the old version (if still allowed) while the new one is completed.

How do feature flags change branching and release strategy?

Feature flags let unfinished features merge into main early but remain disabled, so fewer long-lived feature branches are needed and integration happens continuously. AOSP's trunk-stable model uses aconfig flags with release configurations that decide which flags are enabled in each release. The costs are flag management discipline, testing both flag states where it matters, and removing old flags. For release management, flags allow disabling a risky feature late instead of reverting code.

How do you keep a large merge bisectable?

Avoid squashing an upstream drop into one change; keep upstream history so git bisect can walk individual commits. Land the drop in stages where possible (by repository or subsystem) with builds and tests between stages. Store repo manifest -r snapshots for every CI build so any build can be reproduced. When a regression appears, bisect first between CI builds (coarse) and then between commits within the suspect repositories (fine).

How do you design promotion gate thresholds that are strict but not noisy?

Base thresholds on the distribution of results from known-good builds, not on a single run: measure run-to-run variance and set limits outside normal noise (for example the mean plus a margin). Compare against last good on the same hardware and test setup. Require several runs for noisy KPIs such as power and performance. Separate hard blockers (boot, P0 regressions) from soft limits that require a documented waiver. Review thresholds periodically, but never lower them to pass a failing build.

How would you measure power regressions reliably in CI?

Use dedicated devices with external power monitors or on-device power rails, fixed test profiles (screen off standby, ambient mode, workout, music), controlled radio conditions (shielded boxes or fixed SIM and network), and consistent battery and thermal starting states. Run each profile several times, report mean and variance, and compare with last good. Automatically attach batterystats, wakeup sources and Perfetto traces to failures so they are debuggable. Make standby drain per hour a gate.

How do you handle a regression introduced by a carried vendor patch that upstream now conflicts with?

First ask whether the patch is still needed: upstream may have fixed the same problem differently, in which case drop the patch and verify the original issue stays fixed. If still needed, re-implement it against the new upstream code with the owning team, adding a test that captures the original intent. Consider upstreaming it so it stops being a carried delta. Record the decision in the commit message.

How do you balance quality gates against schedule pressure from customers?

Make the trade-off explicit and data-driven. Quantify the gate failure (which KPI, by how much, which users), lay out options (slip the date, ship with a documented known issue and a dated fix, disable the feature with a flag, take a targeted fix with focused re-test), and state the risk of each. Decide with stakeholders and record the decision. Some gates are non-negotiable, such as boot success, security, and life-critical paths like emergency calling; others can be waived with sign-off and a follow-up plan.

What is an escaped defect, and how do you use it to improve the process?

An escaped defect is a bug found by a customer or in the field that the internal gates should have caught. For each one, do an RCA focused on the detection gap: which test or gate was missing, why the existing tests did not cover it, and what the cheapest reliable way to catch it earlier is. Add that test or gate, and track the escaped defect rate as a process metric. Over time, the gate set becomes a record of lessons learned.

How do HLOS and non-HLOS version mismatches cause bugs, and how do you prevent them?

HLOS drivers and HALs talk to firmware through message protocols and shared memory layouts. If a new HLOS build expects a new firmware message or field that an older modem or DSP build does not support (or vice versa), you get failures such as features silently not working, subsystem crashes and restarts, or boot hangs. Prevent them with a meta build that pins tested combinations, versioned interfaces with capability negotiation, compatibility checks at boot, and integration tests that run the exact combination being released.

How would you reduce integration lead time from an AOSP release to a promoted baseline?
  • Start early: integrate developer previews and betas continuously instead of one big drop.
  • Shrink the carried delta by upstreaming patches and removing obsolete ones.
  • Automate: trial merges, conflict reports, dependency tracking and CI gates.
  • Standardise conflict ownership so work is parallel across domain teams.
  • Reuse resolutions (rerere) and make HAL work predictable using the compatibility matrix.
  • Measure the lead time per phase to see where time goes.
How do you structure an RCA for an intermittent issue that takes days to reproduce?

Increase the reproduction rate first: stress conditions, run many devices in parallel, and automate detection so failures are captured without a person watching. Add targeted instrumentation (always-on ring-buffer tracing, extra logs around the suspected area) and make sure logs survive reboots (pstore, persistent logs). Collect a large sample and look for correlations (build, SKU, temperature, uptime, network). Form hypotheses, test them one at a time, and use the 5 Whys once the technical cause is found.

How do you manage security patch integration alongside feature work?

Security patches follow the monthly Android Security Bulletin and vendor bulletins, with embargo rules before public disclosure. Maintain a dedicated path: patches go into every supported branch on a fixed schedule, verified by STS and a focused regression test set, with the security patch level property updated. Keep this path independent of feature branches so security updates are never delayed by feature integration. Track patch level lag per branch as a metric.

What metrics would you present to leadership about platform integration health?

A short set with trends: upstream integration lead time, carried delta size, build green percentage and time to fix breakages, xTS pass rate per baseline, open P0 and P1 counts and defect convergence toward release, customer CR aging and MTTR, escaped defects, and on-time delivery rate. Pair each with a one-line interpretation and the action being taken, and highlight red items and the help needed.

What does upstream-first mean in practice, and what are its trade-offs?

Upstream-first means that changes to shared code (AOSP framework, Linux kernel) are contributed upstream and, ideally, merged there before or instead of being carried privately. Benefits: less carried delta, easier future drops, community review and testing. Trade-offs: upstream review takes time, the change must be generic enough to be accepted, and schedules may require carrying the patch temporarily. Use a clear policy: carry temporarily only with a tracked upstream submission.

How do you design a change-control process after code freeze that does not become a bottleneck?

Publish clear criteria for what is accepted (for example blocker bugs, security fixes, certification failures). Require each request to include the bug, root cause, risk assessment, test evidence and the branches affected. Meet frequently in short sessions, allow asynchronous approval for low-risk items, and keep the board small with authority to decide. Track approved changes and re-run the relevant tests automatically after they land.

Is Android moving its userspace build to Bazel?

Not as a current fact. Google explored a userspace migration from Soong to Bazel and halted it around 2023; Soong remains the primary userspace build. Kleaf, the Bazel-based kernel build, is still real and is what GKI/kernel teams use. A strong answer splits the two paths and does not treat a 2021-era migration slide as today's architecture.

What is vendor API level, and what does a GRF-style freeze change about OS upgrades?

Vendor API level (ro.vendor.api_level) is the API the vendor partition was built against. A freeze window (often called GRF / Google Requirements Freeze) lets that vendor image pair with newer system images for a documented number of releases, as long as VINTF still matches. The OEM can take a yearly OS upgrade without a full SoC rebase of every HAL. Exact window lengths change; say you would check the current CDD and vendor-API notes. What still moves: new matrix-required HALs, CTS/VTS/GTS for the new release, CDD, and any new KMI.

How do you triage crashes at fleet scale?

Do not debug one tombstone at a time. Symbolize stacks for that exact build, cluster by process plus a stable stack signature (not ASLR addresses), rank by volume times user impact, and assign a DRI to each top cluster with one representative report. A cluster is closed when the signature is gone (or below threshold) on the next build and a regression gate exists. Pair crash-free rate with the top-N signatures so a lucky week cannot hide a new system_server cluster.

Scenario & debugging

Battery life regressed after a platform update. Walk me through your debug.

Reproduce and quantify on a fixed profile, comparing drain per hour with the previous build. Pull a bug report and load batterystats into Battery Historian to spot new wake locks, wakeups, jobs or alarms. Use Perfetto and /sys/kernel/debug/wakeup_sources to find what is keeping the CPU awake, and check suspend residency. Bisect the change set or components to find the culprit. Fix it (remove the wake lock, batch the work, move it to WorkManager, restore sensor batching), re-measure, and add a power KPI gate so it cannot recur silently.

A systemic issue spans app, framework, HAL and kernel, and every team says it is not theirs. What do you do?

Take ownership as the DRI and move the discussion to one triage thread. Get a reliable repro, then capture logs and traces at every layer on one timeline (logcat, dumpsys, HAL logs, dmesg, Perfetto). Bisect to find the layer where the data first goes wrong and the change that introduced it. Present the evidence and assign the owner whose code must change, with a date. Keep a single status of record, track to closure, and add a regression gate.

After an upstream drop, the device does not boot on one chipset. How do you approach it?

Find where it stops. No splash or bootloader output suggests XBL, ABL, AVB or partition layout problems. Splash then reboot loop suggests a kernel panic or init failure, so read the UART console, last_kmsg or pstore. Stuck at the boot animation suggests system_server or a critical service crashing, so read logcat. Compare with the working chipsets to see what differs (BSP, kernel branch, HAL versions, VINTF, SELinux denials). Bisect the drop by repository if needed, and treat it as a P0 blocking promotion for that chipset.

CTS pass rate dropped from 99.8 to 97 percent on the nightly build. What do you do?

Group the new failures by module to see if they share a cause (one broken service can fail hundreds of tests). Check the lab first: network, SIM, device health and test suite version, because environment problems often cause mass failures. Compare with the previous nightly build and its manifest to list the changes in between, and re-run a sample of failures to confirm. Bisect to the culprit change, revert it if it blocks others, and assign a fix. Add the relevant modules to presubmit if they are cheap enough.

A customer reports random reboots in the field on a released product. How do you drive it?

Acknowledge quickly and set an update cadence. Collect data: reboot reasons, kernel panic logs, ramdumps, watchdog and subsystem restart logs, and field telemetry (which builds, SKUs, regions, conditions). Look for patterns and try to reproduce with stress tests under similar conditions. Localise to the layer (kernel panic, modem crash, system_server watchdog), assign the owner and drive root cause. Agree a fix and OTA plan with the customer, roll out staged, confirm the reboot rate drops, then run a post-mortem and add a stability gate.

Two days before release, a P1 regression is found. What do you do?

Quantify impact: which users, how often, how severe, and whether a workaround exists. Check whether a low-risk fix is available and how much re-testing it needs. Lay out options: slip the release, ship with a documented known issue and a dated maintenance fix, disable the feature by flag, or take a targeted fix with focused re-test. Present the options with risks to the release owner and stakeholders, decide transparently, record the decision, and communicate it to customers honestly.

A merged upstream drop builds and boots, but camera start-up got 400 ms slower. How do you find the cause?

Confirm with repeated measurements on both builds under the same conditions. Capture Perfetto traces of camera launch on both builds and compare the phases: app start, camera service connect, HAL open, sensor configuration, first preview frame. Identify the phase that grew, then look at changes in that area in the drop (framework camera service, HAL interface version, SELinux, scheduling). Bisect commits within the suspect repositories. Assign to the camera or framework owner with the traces, and add launch latency to the performance gate.

Your team cherry-picked a fix to three release branches, but one branch still shows the bug. Why might that be?

Possible reasons: the fix depends on an earlier commit that exists on the other branches but not this one; the code on that branch differs and the conflict was resolved incorrectly; the bug on that branch has a different root cause; the fix is in a component that is built from a different repository or prebuilt on that branch; or the tested build did not actually include the change (check the build's manifest snapshot). Verify the change is in the build, compare the code paths, and re-investigate the root cause for that branch.

An OEM customer wants a feature that is not in the current baseline, added a week before code freeze. How do you respond?

Do not refuse in the meeting; understand the business need and the deadline behind it. Assess scope, risk, affected components and test effort with the owning team. Return with options: deliver in the next maintenance release, deliver a limited version behind a flag, or take it now with explicit trade-offs (another item moves, or freeze moves). Make the trade-off visible to program management and the customer, decide together, and document it.

Build breakages on the main branch happen several times a day and block everyone. How would you fix this?

Measure first: which targets break, which kinds of changes cause it, and how long breakages last. Strengthen presubmit to build the affected targets and variants, and enforce atomic submission for multi-repository changes. Adopt a revert-first policy with an on-call build sheriff. Add automated culprit finding to notify authors quickly. Track build green percentage and time to repair as team metrics and review them weekly.

A watch baseline passes all gates, but the customer's product build shows poor standby battery. How do you investigate?

Compare the customer's build against the baseline: OEM apps and services, overlays, configuration, preinstalled watch faces and firmware versions (a different meta build combination). Reproduce with the customer's build on the same power profile and collect batterystats, wakeup sources and Perfetto. Often the cause is an OEM app holding wake locks, a watch face updating too often in ambient mode, or a different sensor or connectivity configuration. Share evidence with the customer, help fix it, and offer them the same power gate you use internally.

You inherit a program with no clear gates and frequent escaped defects. What do you do in the first 90 days?

In the first 30 days, map the image, branches, teams, KPIs and current red items; sit in triage; do not reorganise yet. By 60 days, introduce one status of record, written promotion gates based on the most common escaped defect types, named DRIs for systemic issues, and work-in-progress limits. By 90 days, run the first gated drop, report gate results and escaped defect trends, set a customer communication cadence, and plan the next set of gate improvements.

A modem firmware update in the meta build breaks VoLTE, but the HLOS team says nothing changed on their side. How do you handle it?

Confirm by testing combinations: old modem with new HLOS and new modem with old HLOS, which isolates the component. If the new modem alone breaks it, collect modem logs, IMS and RIL logs and a SIP trace on both firmware versions, and give the modem team the exact failing step. Check whether the new firmware changed an interface (a QMI message or a configuration item) that the HLOS side must adapt to. Pin the old combination in the meta build until a fix is ready, and add a VoLTE call test to the meta build gate.

Tell me how you would handle a disagreement with a domain architect about a fix approach.

Move the discussion from opinions to data. Define what matters (failure rate, performance, risk, how many branches or customers are affected), run a quick experiment or stress test on both options, and write a short one-page comparison with a rollback plan. Present it to the architect and agree on the decision criteria before discussing the choice. Once a decision is made, commit fully and help carry it out, even if it was not your preferred option.

An integration drop is two weeks late because one HAL team keeps missing dates. What do you do?

Understand why: capacity, unclear requirements, technical blockers or competing priorities. Break the remaining work into smaller tracked pieces with daily visibility. Offer help: pair them with engineers from other teams, clarify the minimum required for promotion, or allow the drop to land with the old HAL version if the compatibility matrix permits. If priorities conflict, escalate to management with a clear ask and the impact of each choice. Communicate the revised plan to stakeholders promptly.

An OTA rollout shows a higher boot failure rate than the previous release after reaching 10 percent. What do you do?

Pause the rollout immediately; A/B rollback protects devices that fail to boot, but the failure still harms users. Collect data from affected devices: models, previous build, storage state, and logs from failed boot attempts. Check whether failures correlate with a specific source build (delta payload problem), storage condition (for example low free space for virtual A/B snapshots), or hardware variant. Fix, test on the affected configurations, then resume the rollout from a small percentage.

A new Android release requires migrating several HIDL HALs to AIDL. How would you plan it?

List every HAL that must change using the framework compatibility matrix and deprecation notices. For each, identify the owner, effort and dependencies (framework clients, vendor clients, tests). Prioritise HALs that block the release, and check which can remain on HIDL for now. Plan the migration so both versions can coexist during the transition where possible, write or update VTS tests for the AIDL version, and track progress in the integration status. Start early, ideally on preview releases.

Describe how you would lead a platform through ambiguous requirements and unclear ownership.

Create structure quickly: define the deliverable and the quality bar, list the components and interfaces, and name an owner for each, even if temporary. Set up a single status of record, promotion gates and a regular cadence. Drive cross-team triage on systemic issues and convert unknowns into a tracked list that you burn down. Communicate clearly and early with stakeholders when assumptions change. In an interview, use the STAR format and quantify the result (on-time delivery, defects closed, KPIs held).

A HAL binary builds but is missing from the userdebug image. Where do you look?

First confirm the module exists in the build graph (Android.bp name, vendor: true, no broken required deps). Then check the product makefile: it must be in PRODUCT_PACKAGES (or pulled in by another packaged module). Also check the variant (debug-only packages go in PRODUCT_PACKAGES_DEBUG and will be absent on user), the partition (vendor vs system), and whether a conditional ifeq excluded that SKU. Built is not installed.

An OEM can take next year's Android on last year's vendor image. What must still be true?

The new system must be inside the vendor-API freeze window for that vendor API level, VINTF must still match (no newly required HAL the vendor does not provide), GKI/KMI must still be compatible if the kernel stays, and the device must pass the new release's CTS/VTS/GTS and CDD. Firmware pins in the meta build must still satisfy the new HLOS. If any of those fail, it is not a free upgrade: someone must move vendor code, kernel branch or firmware.

A carrier lab fails a call case that passes on the open-market SKU. How do you start?

Do not start in the modem C-core. Compare CarrierConfig for that MCC/MNC, IMS and APN overlays, the exact meta-build firmware pins, and whether the lab SIM exercises a different feature flag (VoLTE, VoWiFi, 5G NSA/SA). Reproduce on the same carrier config. If config matches, then collect RIL, IMS and modem logs. PTCRB/GCF failures need the same split: protocol/RF versus HLOS policy. Keep one DRI and one status; name the failed contract, not the company.