Core Computer Science

Linux Kernel & BSP

The Board Support Package is the layer that makes a generic Linux kernel and Android run on one chip and board: bootloader, device tree, kernel config and drivers. Classic OS theory (processes, VM, locks, deadlock, filesystems) is on Operating System Concepts. This page is how those ideas look in Linux and on Android.

~130 min read 0 interview questions
In 30 seconds
  • BSP = bootloader + device tree + kernel config + drivers for one SoC/board; Android userspace sits on top.
  • Linux is a monolithic kernel with loadable modules; the device tree describes hardware, and drivers bind to it through compatible strings and probe().
  • Interrupt handlers do the minimum (top half) and defer the rest (threaded IRQ, softirq, tasklet, workqueue).
  • Know which contexts may sleep: process context yes; hard IRQ, softirq and spinlock-held code no.
  • Virtual memory (MMU, page tables, TLB, page faults, copy-on-write) and scheduling (CFS/EEVDF, RT classes, EAS) are asked in almost every interview.
  • GKI and Treble split Google-owned code from vendor code so each can be updated independently.
  • Debug with evidence: dmesg, ftrace, Perfetto, pstore, ramdumps, addr2line, /proc and /sys.

The big picture: where the BSP lives

A Board Support Package (BSP) is everything that makes a generic operating system run on this specific silicon and this specific board: the bootloader, the device tree, the kernel configuration (defconfig), and the drivers for every peripheral. Android then layers its userspace (init, HALs, Zygote, system_server, apps) on top of that Linux kernel.

Analogy

Think of a franchise restaurant. The recipe book (Linux and Android) is the same everywhere, but each branch has a different building: different kitchen layout, different ovens, different wiring. The BSP is the branch-specific setup manual that tells staff where every oven is and how to switch it on. In real terms: the kernel and framework are generic, the SoC and board are unique, and the BSP (device tree + drivers + bootloader + config) bridges the two.

+-------------------------------------------------------------+
|  Apps (Java/Kotlin, NDK)                                    |
+-------------------------------------------------------------+
|  Android framework: system_server (AMS, WMS, PMS, Power...) |
+-------------------------------------------------------------+
|  Native daemons + vendor HALs (AIDL/HIDL over Binder)       |  <- "userspace BSP"
+=============================================================+  syscalls / ioctl
|  Linux kernel: scheduler, MM, VFS, net, Binder driver       |
|  + BSP: device tree, drivers, clk/pinctrl/regulator, PM     |  <- "kernel BSP"
+=============================================================+  MMIO / IRQ / DMA
|  Bootloader (ROM -> SBL/XBL -> ABL/U-Boot), firmware        |
+-------------------------------------------------------------+
|  Hardware: SoC (CPU cores, GIC, MMU, DDR ctrl, clocks),     |
|  PMIC, storage (UFS/eMMC), peripherals (I2C/SPI/UART/USB)   |
+-------------------------------------------------------------+

Hardware + SoC

CPU cores, the GIC (interrupt controller), MMU, memory controller, clocks/PLLs, regulators from the PMIC, and peripheral IP blocks (I2C, SPI, UART, USB, MMC, display, camera).

BSP + Linux kernel

Bootloader loads the kernel image and a device tree describing the board; drivers bind to devices; subsystems (clk, pinctrl, regulator, PM) expose them to userspace via /dev, /sys and /proc.

Android userspace

init starts vendor HALs, native daemons, Zygote and system_server, then apps. It talks to the kernel through syscalls and to vendor hardware through HALs (AIDL, formerly HIDL) over Binder.

What "BSP work" means day to day

  • Bring up a new board: power, console, DDR, storage, then peripherals one by one.
  • Port or patch the bootloader.
  • Write or fix device-tree nodes and drivers.
  • Tune the kernel defconfig and module list.
  • Debug boot, power (suspend/resume, idle current), thermal and peripheral issues using dmesg, ftrace, Perfetto and ramdumps.

Boot in five lines

The full boot chain is covered on the Android boot page. The short version:

  1. BootROM Immutable code in the SoC verifies and loads the first-stage bootloader (root of trust in eFuses).
  2. Bootloader Trains DDR, initializes storage and PMIC, picks the A/B slot, runs Android Verified Boot, loads kernel + ramdisk + DTB.
  3. Kernel Sets up the MMU, parses the device tree, initializes memory and the scheduler, probes drivers, then runs /init as PID 1.
  4. init Mounts partitions, starts ueventd, property service, servicemanager and HALs, then launches Zygote.
  5. Zygote and system_server Zygote preloads the runtime and forks system_server, which starts framework services; the launcher appears and BOOT_COMPLETED is sent.
Interview angle Interviewers open with "what is a BSP?" or "where does the kernel sit in Android?" to see if you have a layered mental model. A strong answer names the layers, says what crosses each boundary (syscalls/ioctl between userspace and kernel, MMIO/IRQ/DMA between kernel and hardware, Binder between framework and HALs), and gives concrete BSP tasks (DT nodes, drivers, defconfig, bring-up, power debug).

Kernel architecture, user space and system calls

Linux is a monolithic kernel: the scheduler, memory manager, filesystems, network stack and drivers all run in one privileged address space and call each other directly. It is also modular: drivers and features can be compiled as loadable kernel modules (.ko) and inserted at runtime. This gives the speed of direct function calls with some of the flexibility of a microkernel.

Monolithic (Linux)

  • All core services and drivers in kernel space.
  • Fast: plain function calls, no message passing.
  • A buggy driver can crash the whole system.
  • Loadable modules add flexibility.

Microkernel (QNX, seL4, Zircon)

  • Kernel keeps only scheduling, IPC and basic memory.
  • Drivers and filesystems run as user processes.
  • Better isolation and fault recovery.
  • Extra IPC and context-switch cost.
Analogy

A monolithic kernel is an open-plan office where every department sits in one room and just shouts across the desk: fast, but one noisy person disrupts everyone. A microkernel is a building of locked offices that pass memos through a mail room: slower, but a fire in one office stays there. In Linux, all subsystems share one address space (the open room), so a bad pointer in any driver can panic the whole kernel.

User space vs kernel space

The CPU has privilege levels. On ARM64, apps and daemons run at EL0 (user), the kernel runs at EL1, a hypervisor at EL2 and secure firmware at EL3 (x86 uses ring 3 and ring 0). User code cannot touch hardware or kernel memory directly; each process gets its own virtual address space, and the kernel's mappings are protected.

AspectUser spaceKernel space
PrivilegeEL0 / ring 3EL1 / ring 0
MemoryPrivate virtual address space per processShared kernel address space, can access all memory
Crash impactProcess dies (SIGSEGV), others keep runningOops or panic; can take down the whole system
StackLarge, growable (MBs)Small, fixed (16 KB on arm64)
Librarieslibc (bionic on Android), any libraryNo libc; kernel APIs only (printk, kmalloc)
Floating pointFree to useAvoided (needs explicit save/restore)

How a system call works

  1. Library wrapper The app calls read(fd, buf, n) in libc, which puts the syscall number and arguments in registers (arm64: number in x8, args in x0-x5).
  2. Trap libc executes svc #0 (arm64) or syscall (x86-64). The CPU switches to EL1 and jumps to the kernel's exception vector.
  3. Dispatch The kernel saves user registers, looks up the handler in the syscall table and calls, for example, ksys_read() which goes through the VFS to the driver's file_operations.read.
  4. Safe copy Data crosses the boundary only through copy_to_user() / copy_from_user(), which validate the user pointer and handle faults.
  5. Return The result (or -errno) goes in x0; eret returns to EL0. libc converts negative values to -1 and sets errno.
user:   app -> libc read() -> svc #0 --------------------+
                                                         | exception (EL0 -> EL1)
kernel: el0_svc -> syscall table -> ksys_read -> vfs_read -> driver->read()
                                                         | copy_to_user(buf, ...)
user:   <- return value in x0 <- eret <-------------------+
Tip ioctl() is the "everything else" syscall for drivers: a command number plus an argument pointer. Binder, GPU, camera and many vendor drivers use it. Always validate the command and use copy_from_user on the argument.
Common pitfall Dereferencing a user pointer directly in kernel code. It may be invalid, unmapped, or point into kernel memory (a security hole). Modern arm64 kernels enforce this with PAN (Privileged Access Never), so a direct access faults.

Building the kernel and modules

# Cross-compile an arm64 kernel
export ARCH=arm64 CROSS_COMPILE=aarch64-linux-gnu-     # or LLVM=1 (Android uses Clang)
make vendor_board_defconfig        # writes .config from a defconfig
make menuconfig                    # optional: tweak options (Kconfig)
make -j$(nproc) Image dtbs modules # kernel image, device tree blobs, .ko modules

# Module handling on target
insmod mydrv.ko        # load a module file (no dependency resolution)
modprobe mydrv         # load with dependencies (uses modules.dep)
rmmod mydrv            # unload
lsmod                  # list loaded modules
modinfo mydrv.ko       # show license, params, aliases, vermagic
  • Kconfig defines options; Kbuild (Makefiles with obj-y / obj-m) decides what is built.
  • CONFIG_FOO=y builds it into the kernel image; =m builds a loadable module.
  • A module must match the kernel's version/ABI (vermagic, symbol CRCs with CONFIG_MODVERSIONS).
#include <linux/module.h>
#include <linux/init.h>

static int __init hello_init(void)
{
    pr_info("hello: loaded\n");
    return 0;                 /* non-zero aborts the load */
}

static void __exit hello_exit(void)
{
    pr_info("hello: unloaded\n");
}

module_init(hello_init);
module_exit(hello_exit);
MODULE_LICENSE("GPL");        /* needed to use EXPORT_SYMBOL_GPL symbols */
Interview angle Expect "what happens when you call read()?", "monolithic vs microkernel?", and "why can't the kernel just dereference a user pointer?". Strong answers mention the trap instruction, privilege switch, syscall table, VFS dispatch, copy_to_user, and the cost of a mode switch (much cheaper than a full context switch).

Device tree and overlays

On ARM systems most peripherals sit at fixed addresses on the SoC and cannot be discovered by probing (unlike PCI or USB, which announce themselves). Something has to tell the kernel "there is an I2C controller at this address, using this interrupt, and an accelerometer at address 0x1d on it". That is the device tree (DT): a data structure describing the hardware, kept separate from kernel code so one kernel image can support many boards.

Analogy

A device tree is the floor plan handed to a new building manager. The manager (kernel) already knows how to run a lift or an air conditioner (drivers), but needs the plan to know which floors have lifts and where the fuse box is. Overlays are sticky notes on the plan for a variant ("this flat has an extra room"). Each node in the plan names a device, its address, its interrupt and its power/clock supplies, and the kernel uses it to decide which driver to start for which piece of hardware.

Files and flow

soc.dtsi  (SoC-wide: CPUs, GIC, controllers, mostly status = "disabled")
   ^ #include
board.dts (board: enables controllers, adds sensors, regulators, pin config)
   |  dtc (device tree compiler)
   v
board.dtb  ----> bootloader passes its address to the kernel (x0 on arm64)
   +  variant.dtbo (overlay) applied by the bootloader at boot
   v
kernel unflattens it -> creates platform devices -> matches drivers -> probe()
  • .dts is the board source, .dtsi are shared includes (SoC, PMIC). The DTC compiles them into a binary .dtb (flattened device tree).
  • Overlays (.dtbo) patch a base DTB for board variants (different display panel, sensor, memory). On Android they live in the dtbo partition and the bootloader merges the right one.
  • On Android 11+ with GKI, the DTB ships in vendor_boot, not in the generic boot image.
  • x86 and servers usually use ACPI tables instead of a device tree for the same purpose.

Anatomy of a node

&i2c_3 {                               /* reference a node defined in the .dtsi */
    status = "okay";                   /* enable the controller on this board */
    pinctrl-names = "default", "sleep";
    pinctrl-0 = <&i2c3_active>;
    pinctrl-1 = <&i2c3_sleep>;

    accel@1d {                         /* node-name@unit-address */
        compatible = "bosch,bmi160";   /* the key used to match a driver */
        reg = <0x1d>;                  /* I2C address (for MMIO: base + size) */
        interrupt-parent = <&tlmm>;    /* GPIO/pin controller acting as irq chip */
        interrupts = <86 IRQ_TYPE_EDGE_RISING>;
        vdd-supply = <&pm_l17>;        /* regulator phandle */
        clocks = <&gcc 42>;
        reset-gpios = <&tlmm 23 GPIO_ACTIVE_LOW>;
    };
};
PropertyMeaning
compatible"vendor,model" string(s), most specific first; matched against a driver's of_match_table.
regAddress (and size) of the device on its parent bus; for MMIO the register window.
interrupts / interrupt-parentIRQ number and trigger type, and which interrupt controller it belongs to.
clocks, *-supply, *-gpios, pinctrl-*References (phandles) to clock, regulator, GPIO and pin-mux providers.
status"okay" enables the node; "disabled" means no device is created and no probe happens.
#address-cells / #size-cellsHow many 32-bit cells child reg entries use.
/chosenSpecial node for boot arguments (bootargs), initrd location, console.

In the driver, properties are read with the OF/fwnode API:

u32 rate;
if (of_property_read_u32(dev->of_node, "sample-rate", &rate))
    rate = 100;                                   /* default */
gpio = devm_gpiod_get(dev, "reset", GPIOD_OUT_HIGH);  /* reads reset-gpios */
vdd  = devm_regulator_get(dev, "vdd");                /* reads vdd-supply  */
Common pitfall A typo in the compatible string, or forgetting status = "okay", means the driver silently never probes. Check /proc/device-tree/ (or /sys/firmware/devicetree/base/) on the target to see the tree the kernel actually received.
Interview angle "What is a device tree and why does Linux use it?" is near-certain. Say: it describes non-discoverable hardware separately from code; before DT, ARM had per-board C "board files" that did not scale; the kernel matches compatible to of_match_table and calls probe(); overlays handle variants. Bonus: DT is a description of hardware, not configuration of software policy.

The driver model: bus, device, driver and probe

The Linux driver model is built from three objects: a bus (I2C, SPI, PCI, USB, or the virtual "platform" bus), devices on that bus, and drivers that know how to run them. When a device and a driver on the same bus match, the core calls the driver's probe(). When either goes away, it calls remove(). Everything shows up in sysfs under /sys/bus, /sys/devices and /sys/class.

Analogy

A staffing agency: the bus is the agency, devices are job openings, drivers are workers with listed skills. The agency matches an opening's requirement ("needs bosch,bmi160") with a worker's skill list and sends the worker in for their first day (probe()), where they set up their desk (map registers, request IRQ, enable power). When the job ends, remove() cleans up. In the kernel the bus's match() function does the pairing, usually by comparing compatible strings or device IDs.

      bus_type (platform / i2c / spi / usb / pci)
        /                       \
   devices                    drivers
   (from DT, ACPI,            (registered by modules:
    enumeration)               platform_driver_register)
        \                       /
         bus->match(dev, drv)  ?
                 | yes
                 v
         drv->probe(dev)  -> returns 0 (bound) / -EPROBE_DEFER (retry later) / error

A minimal platform driver

Platform devices are on-SoC blocks that cannot be enumerated (UART, I2C controller, GPIO controller). They come from device-tree nodes and bind to a platform_driver.

static int foo_probe(struct platform_device *pdev)
{
    struct device *dev = &pdev->dev;
    struct foo *foo;
    int irq, ret;

    foo = devm_kzalloc(dev, sizeof(*foo), GFP_KERNEL);
    if (!foo)
        return -ENOMEM;

    foo->base = devm_platform_ioremap_resource(pdev, 0);  /* map "reg" */
    if (IS_ERR(foo->base))
        return PTR_ERR(foo->base);

    foo->clk = devm_clk_get(dev, NULL);
    if (IS_ERR(foo->clk))
        return dev_err_probe(dev, PTR_ERR(foo->clk), "no clock\n"); /* handles -EPROBE_DEFER */

    irq = platform_get_irq(pdev, 0);
    if (irq < 0)
        return irq;
    ret = devm_request_threaded_irq(dev, irq, foo_hardirq, foo_thread,
                                    IRQF_ONESHOT, "foo", foo);
    if (ret)
        return ret;

    platform_set_drvdata(pdev, foo);
    return clk_prepare_enable(foo->clk);
}

static const struct of_device_id foo_of_match[] = {
    { .compatible = "acme,foo-v2" },
    { }
};
MODULE_DEVICE_TABLE(of, foo_of_match);    /* lets modprobe autoload by alias */

static struct platform_driver foo_driver = {
    .probe  = foo_probe,
    .remove = foo_remove,
    .driver = {
        .name = "foo",
        .of_match_table = foo_of_match,
        .pm = &foo_pm_ops,                 /* suspend/resume, runtime PM */
    },
};
module_platform_driver(foo_driver);
  • devm_* ("managed") resources are freed automatically when probe fails or the device is removed, which removes most error-path leaks.
  • -EPROBE_DEFER: if a dependency (clock, regulator, GPIO provider) is not ready yet, probe returns this and the core retries later. /sys/kernel/debug/devices_deferred lists devices still waiting.
  • Newer kernels use fw_devlink to order probes automatically from DT phandle dependencies.

Driver types

TypeUserspace viewKey structureExamples
Character/dev/xxx, byte stream, open/read/write/ioctl/mmapstruct file_operations, cdev or miscdeviceserial, input (evdev), Binder, sensors, GPU
Block/dev/sda, /dev/block/..., random-access fixed-size blocks, goes through page cache and block layerstruct gendisk, blk-mq request queuesUFS, eMMC, NVMe, loop, zram
NetworkNo /dev node; interfaces like wlan0, rmnet0 used via socketsstruct net_device, sk_buff, NAPIWi-Fi, Ethernet, cellular data
static const struct file_operations foo_fops = {
    .owner          = THIS_MODULE,
    .open           = foo_open,
    .read           = foo_read,
    .write          = foo_write,
    .unlocked_ioctl = foo_ioctl,
    .mmap           = foo_mmap,
    .poll           = foo_poll,
    .release        = foo_release,
};

static struct miscdevice foo_misc = {
    .minor = MISC_DYNAMIC_MINOR,
    .name  = "foo",            /* ueventd creates /dev/foo */
    .fops  = &foo_fops,
};
/* in probe: misc_register(&foo_misc); */

Subsystems a BSP engineer touches daily

clk

Clock tree: enable/disable, set rates. /sys/kernel/debug/clk/clk_summary.

pinctrl + GPIO

Pin muxing (is this pin UART TX or GPIO?), pull-ups, drive strength; GPIO lines as gpiod.

regulator

Voltage rails from the PMIC; consumers call regulator_enable().

I2C / SPI

Controller drivers plus client drivers for sensors, touch, PMICs, codecs.

IOMMU / SMMU

Gives devices their own virtual address space for DMA and isolates them.

thermal, cpufreq, cpuidle, devfreq

Temperature zones and trip points, CPU/GPU frequency scaling, idle states.

PM / runtime PM

System suspend/resume callbacks and per-device idle power control.

input, IIO, regmap

Input events, industrial I/O for sensors/ADCs, register-map abstraction over I2C/SPI/MMIO.

How a /dev node appears

  1. Register The driver registers a device (e.g. misc_register or device_create).
  2. uevent The kernel emits a uevent over netlink; devtmpfs may also create a basic node.
  3. ueventd On Android, ueventd receives it and creates the node with owner, mode and SELinux label from ueventd.rc.
  4. Access A HAL opens it; SELinux policy must allow that domain to access the device type, otherwise you see avc: denied.
Interview angle Interviewers ask you to "write the skeleton of a platform driver" or "explain what happens between DT parsing and probe()". Mention of_match_table, resource mapping, devm_ helpers, IRQ request, deferred probe, file_operations for a char device, and how /dev nodes get created and labeled. Knowing -EPROBE_DEFER signals real hands-on experience.

Interrupts and deferred work

An interrupt is a hardware signal that makes the CPU stop what it is doing and run a handler. On ARM the GIC (Generic Interrupt Controller) collects interrupts from peripherals and routes them to CPU cores. Because an interrupt handler preempts everything, Linux splits the work in two: a tiny top half that runs immediately and a bottom half that finishes later in a gentler context.

Analogy

A hospital emergency room. The triage nurse (top half) sees every arrival instantly, checks vital signs, attaches a wristband and writes a ticket, taking seconds per patient so nobody at the door waits. The doctors (bottom half) do the long treatment later from the queue. If the triage nurse started doing surgery, the queue at the door would overflow. In the kernel, the hard IRQ handler acknowledges the hardware and grabs urgent data with interrupts masked, then hands the heavy work to a threaded IRQ, softirq, tasklet or workqueue.

device raises line -> GIC -> CPU exception vector
   |
   v
TOP HALF (hard IRQ context: cannot sleep, this IRQ line masked, keep it microseconds)
   - read/ack status register, copy urgent data
   - return IRQ_WAKE_THREAD / schedule tasklet / queue_work / raise softirq
   |
   v
BOTTOM HALF (later)
   softirq   : static, per-CPU, very high rate (NET_RX, BLOCK, TIMER); cannot sleep
   tasklet   : built on softirq, one instance runs at a time; cannot sleep (being phased out)
   workqueue : runs in a kworker kernel thread (process context); CAN sleep
   threaded IRQ : handler runs in a dedicated irq/N-name kthread; CAN sleep
MechanismContextCan sleep?Use when
Hard IRQ (top half)InterruptNoAck hardware, read FIFO, timestamp
SoftirqSoftirq (atomic)NoHigh-frequency core subsystems (networking, block completion, timers, RCU)
TaskletSoftirq (atomic)NoLegacy simple deferral; same tasklet never runs on two CPUs at once
WorkqueueProcess (kworker)YesAnything needing mutexes, I2C/SPI transfers, allocations with GFP_KERNEL
Threaded IRQProcess (irq thread)YesPreferred for most device drivers, especially devices behind I2C/SPI
static irqreturn_t foo_hardirq(int irq, void *data)
{
    struct foo *foo = data;
    u32 st = readl(foo->base + STATUS);
    if (!(st & FOO_IRQ_PENDING))
        return IRQ_NONE;              /* not ours (shared line) */
    writel(st, foo->base + STATUS);   /* ack */
    return IRQ_WAKE_THREAD;           /* run foo_thread() */
}

static irqreturn_t foo_thread(int irq, void *data)
{
    struct foo *foo = data;
    mutex_lock(&foo->lock);           /* allowed: process context */
    foo_read_samples_over_i2c(foo);
    mutex_unlock(&foo->lock);
    return IRQ_HANDLED;
}

/* IRQF_ONESHOT keeps the line masked until the thread finishes */
devm_request_threaded_irq(dev, irq, foo_hardirq, foo_thread,
                          IRQF_ONESHOT, "foo", foo);

Other interrupt facts worth knowing

  • Edge vs level triggered: edge fires once on a transition (can be missed if not latched); level keeps asserting until the source is cleared (must ack or you get an interrupt storm).
  • Shared IRQs (IRQF_SHARED): each handler must check whether its device raised it and return IRQ_NONE otherwise.
  • Affinity: /proc/irq/N/smp_affinity selects which CPUs handle an IRQ; /proc/interrupts shows counts per CPU.
  • Wake IRQs: enable_irq_wake() marks an interrupt as able to wake the system from suspend.
  • Types in the GIC: SGI (software, inter-processor), PPI (per-CPU, e.g. arch timer), SPI (shared peripheral), LPI (message-based).
  • NAPI in network drivers switches from interrupt-per-packet to polling under load to avoid interrupt storms.
  • Disabling: local_irq_save/restore masks interrupts on the current CPU; disable_irq() masks one line and waits for running handlers.
Common pitfall Calling anything that may sleep in the top half or a tasklet: mutex_lock, msleep, kmalloc(GFP_KERNEL), copy_to_user, or an I2C transfer. The result is "BUG: scheduling while atomic" or a deadlock. Use a threaded IRQ or workqueue.
Interview angle "Top half vs bottom half" and "tasklet vs workqueue" are staples. A strong answer covers what can sleep, why the top half must be short (interrupt latency, dropped interrupts), and that threaded IRQs are the modern default (and what PREEMPT_RT forces). Follow-ups: how do you share data between an IRQ handler and process context? (spin_lock_irqsave.)

Kernel memory allocation, DMA and buffer sharing

The kernel manages physical memory in pages (usually 4 KB; Android is moving to 16 KB on arm64). Several allocators sit on top of each other: the buddy allocator hands out power-of-two blocks of contiguous pages, the slab allocator (SLUB today) carves pages into caches of small same-sized objects, and kmalloc / vmalloc are the general-purpose APIs drivers call.

Analogy

A warehouse. The buddy allocator rents out whole shelves, always in sizes 1, 2, 4, 8 shelves next to each other. The slab allocator buys a shelf and divides it into labeled bins for one kind of small part (all inode objects, all task_struct objects), so grabbing one is instant. kmalloc is picking a bin from the right-size rack; vmalloc is collecting scattered free shelves and giving you a map that makes them look like one row. The key trade-off: kmalloc memory is physically contiguous, vmalloc memory is only contiguous in virtual addresses.

          kmalloc / kmem_cache_alloc        vmalloc        alloc_pages / DMA API
                   |                            |                    |
             SLUB (object caches)               |                    |
                   \____________________________|____________________/
                                                |
                                    buddy allocator (per zone, orders 0..10)
                                                |
                                   physical page frames (struct page)
APIContiguitySizeUse for
kmalloc / kzallocPhysically and virtually contiguous (direct map)Small to a few MB; large sizes may fail under fragmentationMost driver structures, small buffers
kmem_cache_create/allocContiguous objects from a slab cacheFixed object sizeMany allocations of one struct type
vmalloc / kvmallocVirtually contiguous onlyLargeBig tables or firmware images the CPU alone touches; kvmalloc tries kmalloc first
alloc_pages / __get_free_pagesPhysically contiguous 2^order pagesPage granularityLow-level page users
dma_alloc_coherentDevice-addressable, coherentAny, from CMA if largeDescriptor rings, buffers shared continuously with hardware

GFP flags: can this allocation sleep?

  • GFP_KERNEL: normal; may sleep to reclaim memory. Only in process context without spinlocks held.
  • GFP_ATOMIC: never sleeps, can dip into reserves; for IRQ, softirq or spinlock-held code. More likely to fail, so always check for NULL.
  • GFP_NOWAIT, GFP_NOIO, GFP_NOFS: variants that avoid sleeping or avoid recursing into I/O or filesystem code (used inside block/FS paths).
  • __GFP_ZERO, GFP_DMA32: zero the memory, or restrict to low physical addresses for limited devices.

DMA and cache coherency

DMA (Direct Memory Access) lets a device read or write RAM without the CPU copying each byte. Two complications: the device sees bus/IOVA addresses, not kernel virtual addresses (an IOMMU/SMMU may translate them), and CPU caches may hold stale data if the bus is not cache-coherent.

  • Coherent (consistent) mapping: dma_alloc_coherent() gives memory both CPU and device see consistently (often uncached on non-coherent ARM systems). Good for long-lived control structures.
  • Streaming mapping: dma_map_single() / dma_map_sg() on existing buffers for one transfer. The API does the cache clean (write back before device reads) or invalidate (discard stale lines before CPU reads what the device wrote). Call dma_unmap_* or dma_sync_* before the CPU touches the buffer again.
  • CMA (Contiguous Memory Allocator) reserves a region at boot that can be lent to movable pages and reclaimed when a driver needs a large contiguous buffer (camera, display).
  • Scatter-gather lets a device work on physically scattered pages, removing the need for large contiguous blocks, especially with an IOMMU.
dma_addr_t dma;
dma = dma_map_single(dev, buf, len, DMA_TO_DEVICE);   /* cache clean */
if (dma_mapping_error(dev, dma))
    return -ENOMEM;
start_hw_transfer(dma, len);
/* ... wait for completion interrupt ... */
dma_unmap_single(dev, dma, len, DMA_TO_DEVICE);

Sharing buffers across drivers and processes: ION to DMA-BUF heaps

  • DMA-BUF is the kernel framework for sharing one buffer between devices (camera, GPU, display, video codec) and processes. A buffer is represented by a file descriptor, so it can be passed over Binder or Unix sockets, giving zero-copy pipelines.
  • ION was Android's older allocator for such buffers. It was removed from mainline (Linux 5.11) and replaced by DMA-BUF heaps (/dev/dma_heap/system, vendor heaps), which Android 12+ GKI devices use.
  • ashmem (anonymous shared memory) is being replaced by memfd in newer Android releases for CPU-only shared memory.
  • Debug: /sys/kernel/debug/dma_buf/bufinfo, dmabuf_dump on Android.
Common pitfall Passing a vmalloc or stack buffer to DMA, or forgetting to sync caches. Symptoms are corrupted or stale data that appears only sometimes and disappears when you add prints (which change cache state).
Interview angle "kmalloc vs vmalloc" is the classic opener, then "which GFP flag in an interrupt handler?" and "how do you do DMA safely on a non-coherent ARM SoC?". Show you know physical vs virtual contiguity, why DMA needs physically contiguous or IOMMU-mapped memory, cache clean/invalidate, and DMA-BUF as the zero-copy sharing mechanism.

Virtual memory, paging, MMU and TLB

Every process sees its own virtual address space. The MMU (Memory Management Unit) translates virtual addresses to physical addresses using page tables that the kernel maintains. This gives isolation (one process cannot see another's memory), lets memory be allocated lazily, lets files be mapped into memory, and lets processes share read-only pages.

Analogy

Hotel room numbers. Guests (processes) each think they are in "room 101", but the front desk (MMU) keeps a register (page table) mapping each guest's room 101 to a different real room. The desk clerk keeps a small sticky-note list of recent lookups (the TLB) so common requests are instant. If a guest asks for a room that has not been assigned yet, the clerk calls the manager (a page fault) who either assigns a room on the spot (demand paging) or throws the guest out for asking for a room they never booked (segmentation fault).

Address translation

virtual address (48-bit, 4 KB pages, 4-level table on arm64)
+--------+--------+--------+--------+-------------+
| L0 idx | L1 idx | L2 idx | L3 idx | page offset |
|  9 bit |  9 bit |  9 bit |  9 bit |   12 bit    |
+--------+--------+--------+--------+-------------+
    |
    |  1) check TLB (tagged with ASID) --- hit ---> physical address (fast)
    |  2) miss: hardware table walk from TTBR0 (user) / TTBR1 (kernel)
    v
 L0 table -> L1 table -> L2 table -> L3 table -> PTE (phys frame + permissions)
    |  3) PTE valid     -> fill TLB, access proceeds
    |  4) PTE invalid / permission mismatch -> PAGE FAULT exception to kernel
  • Page table entries hold the physical frame number plus bits: valid, read/write/execute permissions, user/kernel, accessed, dirty, cacheability.
  • TLB (Translation Lookaside Buffer) caches recent translations. A miss costs a page-table walk (several memory reads).
  • ASID (Address Space ID) tags TLB entries per process so a context switch does not have to flush the whole TLB.
  • Huge pages (2 MB / 1 GB blocks, transparent huge pages) cover more memory per TLB entry, reducing misses.
  • On arm64, user space uses TTBR0_EL1, the kernel uses TTBR1_EL1; the kernel has a linear (direct) map of all RAM.

Process address space

high  +------------------------+
      | kernel (not accessible)|
      +------------------------+
      | stack (grows down)     |
      |          v             |
      | mmap region: shared    |  libc.so, libart.so, mapped files,
      | libraries, anon mmaps  |  thread stacks, Java heap
      |          ^             |
      | heap (brk, grows up)   |
      | .bss  (zeroed globals) |
      | .data (init globals)   |
      | .text (code, r-x)      |
low   +------------------------+

The kernel tracks each region as a VMA (vm_area_struct) with its permissions and backing (anonymous memory or a file). cat /proc/<pid>/maps lists them; /proc/<pid>/smaps adds RSS, PSS, dirty and swap per region.

Page faults

  1. Trap The MMU cannot translate an address or permissions do not allow the access; the CPU raises a fault and the kernel's handler runs with the faulting address.
  2. Find the VMA The kernel checks whether the address falls in a valid VMA with suitable permissions.
  3. Minor fault The page is valid and already in RAM (page cache, zero page, or copy-on-write needed): the kernel allocates or copies a page and updates the PTE. No disk I/O.
  4. Major fault The data must be read from storage (file-backed page not cached, or swapped out). The task sleeps while I/O happens.
  5. Invalid access No VMA or wrong permission: user space gets SIGSEGV (or SIGBUS); in kernel space it is an oops ("Unable to handle kernel NULL pointer dereference" / "paging request").

Demand paging and copy-on-write

  • Demand paging: malloc or mmap just creates a VMA; physical pages are allocated on first touch. That is why virtual size (VSS) is much bigger than resident size (RSS).
  • Copy-on-write (COW): fork() copies page tables, not pages. Both processes map the same pages read-only; the first write to a page faults and the kernel copies just that page. This is why forking from Zygote is fast and why preloaded framework classes stay shared across all apps.
  • mmap of a file maps page-cache pages directly; reads trigger faults that load file data. Shared mappings write back to the file.

Reclaim and out-of-memory

  • Page cache keeps file data in RAM; clean pages can be dropped instantly, dirty pages must be written back first.
  • kswapd reclaims in the background when free memory falls below watermarks; direct reclaim happens in the allocating task when it cannot wait (causes latency spikes).
  • Swap / ZRAM: Android usually has no disk swap; it uses ZRAM, a compressed block device in RAM, to swap out anonymous pages.
  • OOM killer (kernel) kills a process by oom_score when memory truly runs out.
  • LMKD (Android userspace low memory killer daemon) acts earlier. It watches PSI (Pressure Stall Information, how long tasks stall waiting for memory) and kills processes by oom_score_adj, which ActivityManager sets from process importance (foreground, visible, service, cached). It replaced the old in-kernel lowmemorykiller driver.
  • Thrashing: when the working set does not fit in RAM, the system spends its time faulting pages in and out; PSI memory "full" values rise sharply.
MetricMeaning
VSSVirtual set size: all mapped virtual memory, mostly meaningless for usage.
RSSResident set size: physical pages mapped, counting shared pages fully in each process.
PSSProportional set size: shared pages divided among sharers; sums correctly across processes. Android's main metric.
USSUnique set size: pages only this process uses; what you free by killing it.
Interview angle "Walk me through a page fault", "what is the TLB and what happens to it on a context switch?", "how is fork cheap?", and "how does Android decide which app to kill?" are all common. Strong answers distinguish minor/major/invalid faults, mention ASIDs, tie COW to Zygote, and explain PSI-driven LMKD plus oom_score_adj.

Processes, threads and scheduling

A process is a running program with its own address space and resources (open files, signal handlers, credentials). A thread is an independent flow of execution that shares its process's address space. Linux represents both as tasks (struct task_struct); clone() flags decide what is shared. A "thread" is just a task created with CLONE_VM | CLONE_FILES | CLONE_SIGHAND | CLONE_THREAD.

Analogy

A process is a household with its own house (address space), furniture and bills (files, credentials). Threads are the family members living in that house: they share the fridge and the TV (memory, file descriptors), which makes cooperation cheap but means they must agree who uses the TV (locking). Starting a new household (fork) is more work than a new family member joining (thread), and a fire in one house does not spread to the next (process isolation).

Process

  • Own address space and page tables.
  • Crash isolated from other processes.
  • Communication needs IPC (pipes, sockets, Binder, shared memory).
  • Creation: fork() (COW) plus exec().
  • Switch may change page tables (TTBR0, ASID).

Thread

  • Shares address space, heap, fds, signal handlers.
  • Own stack, registers, thread-local storage, scheduling state.
  • Communication through shared memory, needs locks.
  • Creation: pthread_create() = clone() with sharing flags.
  • Switch within a process keeps the same page tables, so the TLB stays warm.

Process lifecycle and states

          fork()/clone()
               |
               v
   +------ RUNNABLE (R) <-----------------------+
   |      (on run queue)                         |
   |  scheduled |   ^ preempted / time slice up  | woken up
   |            v   |                            |
   |         RUNNING ------ blocks on I/O, lock --> SLEEPING
   |            |                                  S = interruptible (signals wake)
   |            | exit()                           D = uninterruptible (usually I/O)
   |            v
   |         ZOMBIE (Z): exited, parent has not called wait() yet
   |            |
   +--------> reaped by parent (or init if orphaned)   T = stopped (SIGSTOP, ptrace)
  • fork() creates a copy (COW); exec() replaces the program image; wait() reaps a child.
  • vfork() shares memory with the parent until the child calls exec or _exit; the parent is suspended. Rarely needed now.
  • Zombie: exited process whose parent has not read its exit status. Orphan: parent died; the child is re-parented to init (PID 1) or a subreaper.
  • D state tasks ignore signals; many D-state tasks usually mean stuck I/O or a driver waiting forever. The kernel's hung task detector reports tasks stuck in D for over 120 s.

Scheduling classes

Class / policyPriorityBehavior
SCHED_DEADLINEHighestEarliest Deadline First with runtime/period/deadline reservations.
SCHED_FIFORT 1-99Runs until it blocks, yields, or a higher RT task arrives. No time slice.
SCHED_RRRT 1-99Like FIFO but round-robins among equal priorities with a time slice.
SCHED_NORMAL (SCHED_OTHER)nice -20..+19Fair share: CFS, replaced by EEVDF in Linux 6.6.
SCHED_BATCH / SCHED_IDLELowestThroughput jobs / only when nothing else wants the CPU.
  • CFS (Completely Fair Scheduler) gives each task a virtual runtime (vruntime) that grows as it runs, scaled by its weight (from nice; each nice step is about 10% CPU share). Tasks sit in a red-black tree per CPU ordered by vruntime; the scheduler picks the leftmost (least vruntime). Sleeping tasks fall behind and get CPU quickly when they wake, which favors interactive work.
  • EEVDF (Earliest Eligible Virtual Deadline First), default since Linux 6.6, keeps the fairness idea but also gives each task a virtual deadline based on its requested slice, so latency-sensitive tasks can run sooner without extra priority.
  • RT throttling: by default RT tasks may use 95% of each second (sched_rt_runtime_us) so a runaway FIFO task cannot lock up the system completely.
  • Load balancing migrates tasks between per-CPU run queues; CPU affinity (sched_setaffinity, cpusets) restricts where tasks run.

Android specifics: big.LITTLE, EAS, cgroups, uclamp

  • Phone SoCs mix small efficient cores and big fast cores (big.LITTLE / DynamIQ).
  • EAS (Energy Aware Scheduling) uses an energy model of each CPU and PELT (per-entity load tracking) utilization to place each task on the CPU that meets its needs with the least energy.
  • schedutil cpufreq governor picks CPU frequency from scheduler utilization.
  • cgroups and cpusets group tasks: top-app, foreground, background, system-background. The framework moves app processes between groups as they change state.
  • uclamp (utilization clamping) boosts or caps a group's apparent utilization (e.g. top-app gets a minimum so it lands on big cores). It replaced the older out-of-tree schedtune.

Context switch

  1. Trigger The current task blocks, its time slice ends (timer tick sets need_resched), a higher-priority task wakes, or it yields.
  2. Pick next schedule() asks each scheduling class in priority order for the next task.
  3. Switch memory If the next task is in another process, switch page tables (TTBR0 + ASID); threads of the same process skip this.
  4. Switch registers Save the old task's callee-saved registers, stack pointer and PC; restore the new task's (cpu_switch_to on arm64). FP/SIMD state is saved as needed.
  5. Resume The new task continues where it stopped. The indirect cost (cold caches and TLB) is often larger than the direct cost of a few microseconds.

Preemption models

  • User preemption: always possible when returning from kernel to user space.
  • Kernel preemption (CONFIG_PREEMPT): a task in kernel mode can be preempted whenever it does not hold a spinlock or have preemption disabled. Android kernels are preemptible.
  • PREEMPT_RT (merged in 6.12): converts most spinlocks to sleeping rt-mutexes and forces threaded IRQs for hard real-time latency.
ps -A -T -o PID,TID,PRI,NI,PCY,STAT,NAME   # Android toybox: threads, priority, policy
top -H                                     # per-thread CPU
chrt -f -p 50 <pid>                        # set SCHED_FIFO priority 50
renice -n 10 -p <pid>
cat /proc/<pid>/sched                      # vruntime, switches, wait time
cat /proc/<pid>/status                     # State, Threads, voluntary_ctxt_switches
cat /dev/cpuset/top-app/tasks              # tasks in an Android cpuset
Interview angle Expect "process vs thread", "what exactly happens in a context switch", "how does CFS pick the next task", "FIFO vs RR", and for Android roles "how does the scheduler handle big.LITTLE". Mention task_struct/clone, vruntime and the red-black tree, EEVDF as the modern successor, ASIDs, EAS with the energy model, and uclamp/cpusets.

Synchronization: locks, atomics and RCU

Kernel code runs concurrently on many CPUs and can be interrupted at almost any point by IRQs, softirqs or preemption. Any data touched by more than one of these paths needs protection. The right primitive depends on two questions: can the code holding the lock sleep? and which contexts touch the data?

Analogy

A single-occupancy restroom. A spinlock is standing at the door jiggling the handle until it opens: fine if the person inside will be out in seconds, wasteful otherwise. A mutex is taking a number and sitting down (sleeping) until called. A semaphore is a restroom with N stalls and a counter. A reader-writer lock is a museum room: many visitors can look at once, but the cleaner needs it empty. RCU is replacing a notice-board poster: you pin up the new version, people already reading the old one finish, and only when everyone has moved on do you throw the old one away. Each maps directly to its kernel primitive and the context in which waiting is acceptable.

PrimitiveWaitingContextUse for
spinlock_tBusy-wait; disables preemptionAny, including IRQVery short sections; data shared with IRQ handlers (spin_lock_irqsave)
struct mutexSleepsProcess onlyLonger sections; one owner, must be unlocked by the owner
struct semaphoreSleepsProcess (up() allowed from IRQ)Counting resources; rarely used for plain mutual exclusion now
rwlock_t / rw_semaphoreSpin / sleepAny / processRead-mostly data with multiple concurrent readers
seqlock_tReaders retry, writers spinAnySmall, frequently read data (e.g. time) where writers must not starve
RCUReaders never blockAny reader; writer waits a grace periodRead-mostly linked structures (routing tables, task lists)
atomic_t, refcount_tNone (single instruction / LL-SC)AnyCounters, reference counts, flags
completion, wait queuesSleepsWaiter in process context"Wait until event X happens" (e.g. DMA done)
per-CPU variablesNoneWith preemption disabledAvoid sharing altogether (statistics, caches)

Choosing a spinlock variant

spin_lock(&l);                  /* data never touched from IRQ/softirq */
spin_lock_bh(&l);               /* shared with softirq/tasklet: disables bottom halves */
spin_lock_irq(&l);              /* shared with hard IRQ; IRQs known enabled before */
spin_lock_irqsave(&l, flags);   /* shared with hard IRQ; safest, restores prior state */
spin_unlock_irqrestore(&l, flags);

Why disable IRQs? If process code holds lock L and an IRQ on the same CPU tries to take L, the IRQ spins forever because the holder can never run again: a self-deadlock.

Mutex vs semaphore vs spinlock

Mutex

  • Binary, has an owner.
  • Only the owner may unlock.
  • Can sleep; process context only.
  • Supports optimistic spinning and debugging (lockdep).
  • Priority inheritance available via rt_mutex.

Semaphore

  • Counter; N holders allowed.
  • No owner: any task (or an IRQ) may call up().
  • Useful as a signaling mechanism.
  • No priority inheritance possible.
  • Prefer mutex for mutual exclusion, completion for signaling.

Atomics and memory ordering

  • Atomic ops (atomic_inc, atomic_cmpxchg) compile to single atomic instructions (ARMv8.1 LSE like LDADD, or LDXR/STXR loops).
  • CPUs and compilers reorder memory accesses. ARM is weakly ordered, so another CPU may see writes in a different order. Memory barriers (smp_mb(), smp_wmb(), smp_rmb(), smp_store_release() / smp_load_acquire()) enforce ordering. Locks include the necessary barriers.
  • READ_ONCE() / WRITE_ONCE() stop the compiler from tearing, merging or caching accesses to shared variables.
  • volatile is not a synchronization primitive: it stops the compiler caching a value but gives no atomicity and no CPU ordering. It is appropriate for memory-mapped registers (the kernel uses readl/writel) and signal-handler flags, not for thread synchronization.

RCU in one picture

reader:  rcu_read_lock(); p = rcu_dereference(gp); use(p); rcu_read_unlock();
                 (no lock, no atomic write, never blocks the writer)

writer:  new = copy(old); modify(new);
         rcu_assign_pointer(gp, new);     // publish: new readers see new
         synchronize_rcu();               // wait: all pre-existing readers finished
         kfree(old);                      // safe to reclaim   (or call_rcu / kfree_rcu)

Userspace locking: futex

pthread_mutex and Java monitors are built on the futex (fast userspace mutex) syscall. The uncontended path is a single atomic compare-and-swap in user space with no syscall; only when there is contention does the thread call futex(FUTEX_WAIT) to sleep in the kernel. PI-futexes add priority inheritance.

Common pitfall Sleeping while holding a spinlock (mutex, GFP_KERNEL, copy_to_user, msleep), taking a spinlock in process context without disabling IRQs when an IRQ handler also takes it, or using volatile instead of a lock. Enable CONFIG_PROVE_LOCKING (lockdep) and CONFIG_DEBUG_ATOMIC_SLEEP during development.
Interview angle "Spinlock vs mutex vs semaphore, and when do you use each in a driver?" is asked in nearly every kernel interview. Answer by context (can I sleep?), duration, IRQ sharing (_irqsave), ownership (mutex has an owner, semaphore does not), and give RCU for read-mostly data. Follow-ups often go to memory barriers and why volatile is not enough.

Deadlock, livelock, starvation and priority inversion

A deadlock is a set of tasks each waiting for a resource held by another in the set, so none can proceed. It needs all four Coffman conditions at once; breaking any one prevents it.

Analogy

Four cars arrive at a four-way intersection at the same moment, each waiting for the car on its right to go first: nobody moves, forever. Add a rule that everyone follows (lock ordering: "north goes first") and the gridlock cannot form. In code, the cars are threads, the lanes are locks, and the fixed rule is a global lock-acquisition order.

ConditionMeaningHow to break it
Mutual exclusionResource held by one task at a timeUse lock-free structures, RCU, or read-sharing where possible
Hold and waitHold one lock while waiting for anotherAcquire all locks at once, or release before waiting
No preemptionLocks cannot be taken awaytrylock with back-off and retry; timeouts
Circular waitA waits for B, B waits for AGlobal lock ordering (most common fix; lockdep enforces it)
/* ABBA deadlock */
Thread 1: mutex_lock(&A); mutex_lock(&B);   /* holds A, wants B */
Thread 2: mutex_lock(&B); mutex_lock(&A);   /* holds B, wants A */

/* Fix: always A then B, e.g. order by address */
if (a < b) { mutex_lock(a); mutex_lock_nested(b, 1); }
else       { mutex_lock(b); mutex_lock_nested(a, 1); }
  • Kernel-specific deadlocks: taking the same spinlock twice (self-deadlock; Linux spinlocks are not recursive), IRQ handler contending with process context that did not disable IRQs, waiting on a workqueue item from inside the same workqueue, flush_work under a lock the work also needs.
  • Detection: lockdep records lock-order dependencies at runtime and warns about potential cycles before they happen; hung-task detector and soft/hard lockup detectors catch real hangs; in a ramdump, look for tasks in D state and who owns each lock.
  • Avoidance (theory): the Banker's algorithm grants requests only if the system stays in a safe state. Rarely used in real kernels because needs must be known in advance.
  • Livelock: tasks keep changing state in response to each other but make no progress (two people stepping aside in a corridor). Randomized back-off helps.
  • Starvation: a task never gets the resource or CPU (e.g. writers starved by constant readers). Fair locks (ticket/queued spinlocks) and aging help.

Priority inversion

time -->
Low (L)   : [lock M ]...preempted.....................[unlock M]
Medium (M):             [ runs, CPU-bound, long ...  ]
High (H)  :        [wants M -> BLOCKED ...........................][runs]
            H effectively waits for M, a task it outranks.

With priority inheritance: L temporarily gets H's priority, M cannot preempt L,
L releases the lock quickly, H runs.
  • Priority inheritance: the lock owner inherits the highest waiter's priority until it releases. Linux rt_mutex and PI-futexes (PTHREAD_PRIO_INHERIT) implement it.
  • Priority ceiling: a lock has a ceiling priority and any holder runs at that ceiling.
  • Famous case: the Mars Pathfinder lander (1997) kept resetting from a watchdog because of priority inversion; enabling priority inheritance on the mutex fixed it.
  • Android: Binder propagates caller priority to the server thread (priority inheritance across IPC) so a foreground call is not served at background priority.
Interview angle You will be asked to list the four deadlock conditions and give prevention strategies, then maybe to spot an ABBA deadlock in code. For platform roles, add lockdep, watchdogs and ramdump analysis. Explaining priority inversion with a timeline and naming priority inheritance (and Binder's version of it) is a strong signal.

Inter-process communication (IPC)

Processes are isolated, so they need kernel-provided channels to exchange data or signals. Linux offers many; Android adds Binder as its primary mechanism.

Analogy

Ways neighbors communicate: a pipe is a tube between two houses (one direction, in order); a signal is ringing the doorbell (no message, just "something happened"); shared memory is a shared garden both can reach into (fastest, but you must agree on rules); a socket is a phone line (two-way, can even reach another town); Binder is a concierge who knows every resident's name, checks your ID, delivers your request and brings back the answer. The kernel is the town that builds and polices all these channels.

MechanismDataNotes
Pipe / FIFOByte stream, one-waypipe() between related processes; named FIFO via filesystem path. Two copies (in and out of a kernel buffer).
SignalSignal number onlyAsync notification (SIGKILL, SIGSEGV, SIGTERM). Handlers may call only async-signal-safe functions.
Shared memoryArbitraryPOSIX shm_open/mmap, System V shmget, memfd. Zero-copy; needs separate synchronization.
Message queueDiscrete messagesPOSIX mq_* or System V; preserves message boundaries.
Unix domain socketStream or datagramLocal, fast, can pass file descriptors (SCM_RIGHTS) and credentials. Android uses them for logd, zygote, lmkd, property_service.
TCP/UDP socketStream or datagramWorks across machines.
NetlinkMessages kernel <-> userUevents, routing, Wi-Fi (nl80211) configuration.
eventfd / futexCounter / wait wordLightweight notification and locking primitives.
BinderParcels, objects, fdsAndroid RPC: /dev/binder, /dev/hwbinder, /dev/vndbinder.

Binder in brief

  1. Setup Each process opens /dev/binder and mmaps a receive buffer; it also starts a Binder thread pool.
  2. Call The client proxy marshals arguments into a Parcel and calls ioctl(BINDER_WRITE_READ).
  3. One copy The driver copies the data once, from the client's memory directly into the server's mmap'd buffer, and wakes a server Binder thread.
  4. Dispatch The server stub (onTransact) unmarshals and runs the method; the reply comes back the same way. oneway calls do not wait for a reply.
  • Adds what raw Linux IPC lacks: caller UID/PID delivered by the kernel (unforgeable identity for permission checks), object references with reference counting, death notifications, priority inheritance, fd passing, and a name registry (servicemanager).
  • Large buffers go through DMA-BUF / memfd fds instead of the Binder buffer (transaction buffer is about 1 MB per process).
  • See the dedicated AIDL/Binder page for the full framework view.
Interview angle Typical questions: "list IPC mechanisms and compare them", "fastest IPC?" (shared memory, but you need synchronization), "why does Android use Binder instead of sockets?", and "what is the copy count of a pipe vs Binder vs shared memory?" (2 vs 1 vs 0). Strong answers mention identity, lifetime and security, not just speed.

Filesystems and the storage stack

The VFS (Virtual File System) gives every filesystem the same interface: open, read, write, mmap. Under it, concrete filesystems (ext4, f2fs, erofs, tmpfs, procfs, sysfs) store data; below those, the block layer schedules I/O to the storage driver (UFS, eMMC).

Analogy

A library. The VFS is the front desk that accepts the same request form for every collection. Each filesystem is a different archive system behind the desk (card index, database, microfilm). Inodes are catalogue cards (who owns the book, where its pages are), dentries are the signs pointing from names to cards, and the page cache is the reading room table where recently used books stay out for fast access. The block layer is the stack of carts moving books to and from the storage basement.

app: read(fd) / mmap
        |
       VFS  -- file (open instance, offset) -> dentry (name) -> inode (metadata, blocks)
        |
   page cache (hit: return from RAM; miss: read from storage; writes are dirty pages)
        |
   filesystem: ext4 | f2fs | erofs | tmpfs | proc/sysfs (virtual, no storage)
        |
   block layer: bio -> blk-mq queues -> I/O scheduler (mq-deadline, bfq, none)
        |
   storage driver: UFS / eMMC / NVMe  --DMA--> flash (with its own FTL)

Key objects

  • inode: metadata (owner, mode, size, timestamps, block map) but not the name. Hard links are extra names for the same inode.
  • dentry: maps a path component to an inode; cached in the dcache for fast lookups.
  • file: an open instance with current offset and flags; a file descriptor is an index into the process's fd table pointing to it.
  • superblock: one per mounted filesystem, describes it.

Android filesystems

FilesystemWhere on AndroidWhy
ext4Historically /system, /vendor, /data; still commonMature journaling FS; metadata journal protects consistency after power loss
f2fs/data on most modern devicesFlash-Friendly FS: log-structured, sequential writes, fits NAND/FTL behavior, good random-write performance
erofs/system, /vendor, /product (read-only)Enhanced Read-Only FS: compressed, saves space, fast random reads, works well with dm-verity
tmpfs/dev, /mnt, /apex stagingRAM-backed
procfs, sysfs, debugfs, configfs, tracefs/proc, /sys, /sys/kernel/debug, /config, /sys/kernel/tracingVirtual windows into kernel state
FUSE (MediaProvider)Emulated shared storage (/sdcard, /storage/emulated) since Android 11Userspace filesystem for scoped storage and per-app permissions. sdcardfs was the in-kernel wrapper used on Android 6โ€“10 and was removed; do not treat them as current alternatives.
  • Durability: write() only dirties the page cache. fsync() forces data and metadata to storage (with cache flush / FUA commands). Without it, a power cut can lose recent writes.
  • Dynamic partitions: system, vendor, product live inside one super partition using device-mapper (dm-linear); dm-verity checks read-only blocks against a hash tree; file-based encryption (fscrypt) protects /data.
  • Block I/O scheduler: for fast flash, mq-deadline or none is typical; bfq favors interactivity.
Interview angle Questions include "what is an inode?", "hard link vs soft link?", "what happens when you write() and then pull the battery?", "why f2fs for /data and erofs for /system?", "mmap vs read?", and "is /sdcard FUSE or sdcardfs?". Tie answers to page cache, writeback, fsync, flash characteristics, and the Android 11 switch back to MediaProvider FUSE.

Kernel power management: suspend, resume and wakeup sources

Battery life on a phone or watch is roughly proportional to how much time the application processor spends suspended. Linux offers system suspend (whole system to a low-power state) and runtime PM (individual devices idle while the system runs), plus CPU idle states and frequency scaling.

Analogy

Closing an office for the night. The manager checks that nobody has put up a "still working" sign (wakelock), sends everyone home in order (freeze tasks), turns off each department's equipment (driver suspend()), and locks up, leaving only the alarm system on (wakeup IRQs). An alarm or a phone call reopens the office in reverse order. Runtime PM is each department switching off its own lights when empty during the day. In the kernel, a held wakeup source blocks suspend, and a failing driver suspend() aborts the whole sequence.

  1. Trigger Android's PowerManagerService sees no wakelocks held and lets the system suspend (autosuspend writes mem to /sys/power/state).
  2. Freeze User tasks and freezable kernel threads are frozen.
  3. Suspend devices Each driver's .suspend, then .suspend_late, .suspend_noirq callbacks run, children before parents.
  4. Non-boot CPUs off Secondary CPUs are taken offline; the last CPU enters the deepest state via PSCI firmware. RAM stays in self-refresh.
  5. Wake A wake-capable IRQ (RTC alarm, modem, button, sensor hub) fires; resume runs in reverse order and tasks are thawed.
  • Wakeup sources (kernel wakelocks): pm_stay_awake() / __pm_wakeup_event() or userspace /sys/power/wake_lock block suspend. Framework wakelocks from apps are aggregated by PowerManagerService into a kernel wakeup source.
  • Runtime PM: pm_runtime_get_sync() / pm_runtime_put_autosuspend() count users; when idle, the device's runtime_suspend gates its clocks and regulators.
  • cpuidle picks C-states when a CPU has nothing to run; cpufreq/devfreq scale frequency; thermal framework throttles when trip points are crossed.
cat /sys/power/state                     # supported states (freeze mem ...)
cat /sys/kernel/debug/wakeup_sources     # active_count, total_time, active_since per source
cat /sys/kernel/debug/suspend_stats      # success/fail counts, last_failed_dev
cat /sys/power/wake_lock                 # userspace-held kernel wakelocks
dmesg | grep -iE "PM: suspend|wakeup|Freezing"
dumpsys power                            # framework wakelocks
Interview angle "How does suspend work?", "what is a wakelock at the kernel level?" and "the device will not suspend, how do you debug it?" are frequent in BSP and wearable interviews. Walk the suspend sequence, name wakeup_sources and suspend_stats, and separate framework wakelocks from kernel wakeup sources and wake IRQs.

GKI, vendor modules, Treble and VINTF

Android used to ship a different, heavily patched kernel for every device, and every OS upgrade required the SoC vendor to rework everything. Google split things up at two levels: Project Treble (Android 8) separated the Android framework from vendor userspace, and GKI (Generic Kernel Image) separated the core kernel from vendor kernel code. GKI 1.0 shipped optionally with Android 11; GKI 2.0 is required for new devices launching with Android 12+ on kernel 5.10 or newer. Devices that launched earlier can still upgrade Android on a non-GKI vendor kernel.

Analogy

A standard wall socket. Appliance makers (vendors) build plugs to the standard, and the power company (Google) can upgrade the grid without anyone rewiring their toaster. Treble defines the socket between the Android framework and vendor HALs (stable AIDL interfaces checked by VINTF). GKI defines the socket inside the kernel: a stable Kernel Module Interface (KMI) that vendor modules plug into, so Google can update the core kernel independently.

          Google-owned / updatable             Vendor-owned
  +---------------------------------+  +-----------------------------------+
  | system, system_ext, product     |  | vendor, odm partitions            |
  | (framework, GSI)                |  | HAL services, vendor daemons      |
  +---------------------------------+  +-----------------------------------+
               |  stable AIDL (HIDL legacy) over binder / hwbinder / vndbinder
  +---------------------------------+  +-----------------------------------+
  | GKI kernel: boot.img            |  | vendor modules (.ko):             |
  | (core kernel, built by Google)  |<-| vendor_dlkm, vendor_boot ramdisk  |
  |   KMI = stable symbol list      |  | + DTB in vendor_boot, DTBO        |
  +---------------------------------+  +-----------------------------------+

GKI essentials

  • The GKI kernel is built by Google from the Android Common Kernel (ACK) branch for an Android release and kernel version (for example android14-6.1) and signed.
  • Vendor modules contain SoC and board support (drivers, pinctrl, clocks). They load from vendor_boot (early, needed for mounting) and vendor_dlkm (later).
  • The KMI (Kernel Module Interface) is the set of exported symbols and data structure layouts that modules may use; it is frozen for the life of a branch and checked with ABI tooling (symbol lists, STG/libabigail).
  • Vendors cannot patch core kernel code. Needed hooks are added through vendor hooks (tracepoint-based) upstreamed to ACK.
  • init_boot (Android 13+) holds the generic ramdisk, separating it from the kernel in boot.

Treble essentials

  • HALs are stable interfaces between framework and vendor code. HIDL (Android 8-10 style) is deprecated; stable AIDL is the current standard.
  • VINTF (Vendor Interface) objects: the device manifest lists HALs the vendor provides; the framework compatibility matrix lists HALs and versions the framework needs. They are checked at build time and on OTA.
  • VNDK (Vendor NDK) defined which system libraries vendor code could link; it is being deprecated in favor of stable interfaces.
  • GSI (Generic System Image) plus VTS and CTS verify that a pure AOSP system boots on the vendor implementation.
  • Three Binder domains keep the split clean: /dev/binder (framework-app), /dev/hwbinder (HIDL), /dev/vndbinder (vendor-vendor).

init, properties, ueventd and SELinux (userspace BSP)

  • init parses *.rc files: services (with class, user, group, seclabel, onrestart), actions triggered by events (on boot, on property:sys.boot_completed=1).
  • Properties: ro.* (read-only, set once), persist.* (survive reboot), sys.*/vendor.*; read with getprop, set with setprop.
  • ueventd handles kernel uevents, creates /dev nodes with permissions and SELinux labels, and loads firmware.
  • SELinux is enforcing: every process and file has a context; denials show as avc: denied in dmesg/logcat. Vendor policy lives in the vendor partition.
service vendor.foo-hal /vendor/bin/hw/android.hardware.foo-service
    class hal
    user system
    group system input
    onrestart restart vendor.bar

on property:sys.boot_completed=1
    write /sys/devices/platform/foo/enable 1
Interview angle Expect "what problem does GKI solve?", "what is the KMI?", "how do vendor drivers work with GKI?", and "what is Treble / VINTF?". Tie both to update speed and security patches: Google can ship kernel and framework updates without waiting for every vendor to rebase. Mention that vendor code must be modules using only KMI symbols.

Debugging and tracing tools

Kernel bugs are rarely solved by staring at code. You need evidence: logs, traces, crash dumps and live state. Pick the tool by the question: "what happened?" (logs), "when and in what order?" (traces), "where is time spent?" (profilers), "what was the state when it died?" (dumps).

Analogy

Investigating an aircraft incident. dmesg is the cockpit log, ftrace and Perfetto are the flight data recorder with timestamps for every control input, perf is the fuel-consumption report showing which engine burned most, pstore/ramoops is the black box that survives the crash, and a ramdump analyzed in crash or Trace32 is the full wreckage reconstruction. /proc and /sys are the live instrument panel you can read while still flying.

Kernel logs

  • printk with levels KERN_EMERG (0) to KERN_DEBUG (7); prefer pr_err(), dev_err(dev, ...), dev_dbg(). /proc/sys/kernel/printk sets the console level.
  • Dynamic debug: enable pr_debug/dev_dbg at runtime: echo 'file foo.c +p' > /sys/kernel/debug/dynamic_debug/control.
  • dmesg -w (follow), logcat -b kernel; pstore/ramoops keeps the previous boot's console and panic log in reserved RAM: /sys/fs/pstore/console-ramoops-0 (older devices: /proc/last_kmsg).
  • Early boot: earlycon on the kernel command line gives UART output before the normal console driver loads.

ftrace

cd /sys/kernel/tracing                     # or /sys/kernel/debug/tracing
cat available_tracers                      # function function_graph irqsoff preemptoff wakeup nop
echo function_graph > current_tracer
echo 'foo_*' > set_ftrace_filter           # only trace matching functions
echo 1 > events/sched/sched_switch/enable  # static tracepoints
echo 1 > events/irq/enable
echo 1 > tracing_on; sleep 5; echo 0 > tracing_on
cat trace > /data/local/tmp/trace.txt
echo nop > current_tracer                  # reset

# kprobe: trace an arbitrary kernel function with arguments
echo 'p:myprobe do_sys_openat2 dfd=%x0' >> kprobe_events
echo 1 > events/kprobes/myprobe/enable

trace-cmd is a friendly front end; trace_printk() writes to the trace buffer cheaply from hot paths. The irqsoff and preemptoff tracers find the longest regions with interrupts or preemption disabled.

Perfetto, systrace and perf

  • Perfetto is Android's system-wide tracer (successor to systrace). It collects ftrace events (sched, irq, freq, idle, binder), atrace userspace markers, process stats, memory counters and heap profiles into one timeline viewed at ui.perfetto.dev (runs locally in the browser).
  • adb shell perfetto -o /data/misc/perfetto-traces/trace -t 10s sched freq idle am wm gfx binder_driver
  • systrace is the older Python/HTML front end over atrace + ftrace; same data, now superseded.
  • perf samples CPU hotspots and hardware counters: perf top, perf record -g, perf report; simpleperf is Android's version.
  • eBPF (bpftrace, BCC; used by Android for network and GPU stats) attaches small safe programs to kprobes and tracepoints for custom tracing.

Crash and post-mortem tools

ToolPurpose
addr2line -e vmlinux <addr>Convert a kernel address to file:line (needs the unstripped vmlinux with debug info). For modules, use the .ko with the offset.
scripts/decode_stacktrace.sh vmlinux < oops.txtSymbolize a whole oops call trace.
gdb vmlinux + list *(func+0x48)Show source around a func+offset from an oops.
kdump / ramdumpFull memory image captured after a panic or watchdog bite (via a crash kernel or the SoC's bootloader/firmware dump mode).
crash utilityAnalyze a ramdump with vmlinux: bt, ps, log, struct, kmem, foreach bt.
Lauterbach Trace32 (T32)JTAG debugger and ramdump analyzer common in SoC bring-up; can halt and inspect a hung CPU.
kgdb / kdbLive source-level kernel debugging over serial or network; set breakpoints in the kernel.
KASAN, KFENCE, UBSANDetect use-after-free, out-of-bounds, undefined behavior (KFENCE is a low-overhead sampling version usable in production).
lockdep, KCSANLock-order violations and data races.
kmemleakReport unreferenced kernel allocations (/sys/kernel/debug/kmemleak).
Soft/hard lockup detectors, hung task detector, watchdogDetect CPUs stuck with preemption or interrupts off, and tasks stuck in D state.

Reading a kernel oops

Unable to handle kernel NULL pointer dereference at virtual address 0000000000000010
Mem abort info: ESR = 0x96000005 ...
Internal error: Oops: 96000005 [#1] PREEMPT SMP
CPU: 3 PID: 412 Comm: kworker/3:1 Tainted: G  O  6.1.25-android14 ...
pc : foo_read_samples+0x48/0x1a0 [foo]      <- faulting function + offset / size, module
lr : foo_thread+0x2c/0x60 [foo]              <- return address (caller)
Call trace:
 foo_read_samples+0x48/0x1a0 [foo]
 foo_thread+0x2c/0x60 [foo]
 irq_thread_fn+0x2c/0xa0
 irq_thread+0x16c/0x278
 kthread+0x110/0x120
  • Faulting address 0x10 means a NULL struct pointer plus a field at offset 16.
  • Tainted: O means an out-of-tree module is loaded; G means only GPL modules.
  • An oops kills the current task and the kernel may continue (in a possibly broken state); a panic halts or reboots. The kernel default is to keep running after an oops (panic_on_oops=0) unless it is built with CONFIG_PANIC_ON_OOPS. Android typically writes 1 to /proc/sys/kernel/panic_on_oops from init, so an oops becomes a panic and reboot (or ramdump) rather than leaving a tainted, half-dead kernel.

Userspace and general tools

  • ps -A, top -H, vmstat, iostat, free, /proc/<pid>/status, lsof.
  • strace (syscalls), ltrace (library calls), gdb/lldb, debuggerd -b <pid> (native stack dump), ndk-stack / addr2line to symbolize tombstones.
  • logcat, bugreport, dumpsys (power, meminfo, activity, cpuinfo).
  • Boot timing: dmesg timestamps, initcall_debug on the command line, bootchart, Perfetto boot traces, bootanalyze scripts in AOSP.

The /proc and /sys windows

PathWhat it shows
/proc/<pid>/{status,maps,smaps,stack,sched,fd,wchan}Per-process state, memory map, kernel stack, scheduler stats, open files, what it waits on
/proc/meminfo, /proc/slabinfo, /proc/vmstat, /proc/buddyinfoMemory breakdown, slab caches, VM counters, free blocks per order (fragmentation)
/proc/interrupts, /proc/softirqsIRQ and softirq counts per CPU
/proc/pressure/{cpu,memory,io}PSI stall information
/proc/cmdline, /proc/version, /proc/device-tree/Kernel command line, version, live device tree
/sys/class/, /sys/bus/, /sys/devices/Device and driver model; bind/unbind drivers
/sys/devices/system/cpu/cpu*/cpufreq/Frequency, governor, available frequencies
/sys/class/thermal/thermal_zone*/Temperatures and trip points
/sys/kernel/debug/ (debugfs)clk_summary, regulator_summary, gpio, pinctrl, wakeup_sources, dma_buf, devices_deferred
/config (configfs)Runtime configuration objects such as USB gadget functions
Tip Always ship symbol files (vmlinux, System.map, unstripped .ko files) with each build. Without them a ramdump or oops from the field is nearly useless.
Interview angle Interviewers ask "how would you debug X?" and listen for tools matched to the question. Name the exact file or command (/sys/kernel/debug/wakeup_sources, addr2line, ftrace function_graph, Perfetto sched tracks), explain what the output would tell you, and describe a method: localize the layer, gather evidence, hypothesize, confirm, fix, add a regression check.

Practical scenarios

Scenario questions turn theory into a structured answer. The pattern interviewers like: narrow the layer, gather evidence, form a hypothesis, confirm, fix, prevent regression. Say the commands and files out loud.

Analogy

A doctor does not prescribe before diagnosing: symptoms first, then tests that split the possibilities in half (blood test, X-ray), then a diagnosis, then treatment and a follow-up visit. Debugging a board is the same: the symptom tells you which stage to look at, each command rules out half the possible causes, and a fix is only done when a test proves it and a regression check keeps it fixed.

1. Device stuck at the boot logo

  1. Which stage died? The logo means the bootloader ran, so the failure is in the kernel or init/userspace.
  2. Kernel log Capture the UART console or /sys/fs/pstore/console-ramoops-0. Look for a panic, a driver hanging in probe(), or "VFS: Unable to mount root fs".
  3. Userspace If the kernel is fine, check logcat -b all: a critical service crash-looping (init: Service 'x' ... restarting), an avc: denied blocking a daemon, a failed mount, or system_server crashing.
  4. Bisect Does a known-good build or the other A/B slot boot? Diff DT, defconfig and driver changes since then.

2. Kernel panic or oops: decode it

  1. Read the header Fault type ("NULL pointer dereference", "paging request"), faulting address, CPU, task (Comm), taint flags.
  2. Locate pc is the faulting instruction, lr the caller; the top of the call trace is where it died. Symbolize with addr2line -e vmlinux or decode_stacktrace.sh.
  3. Classify NULL deref (uninitialized pointer, probe ordering, missing error check), use-after-free, "scheduling while atomic", "BUG: spinlock lockup", stack overflow.
  4. Reproduce with sanitizers Enable KASAN, lockdep and DEBUG_ATOMIC_SLEEP to catch the root cause at the first bad access instead of the later crash.

3. A peripheral is not detected (I2C sensor missing)

  1. Did the driver probe? dmesg | grep foo, ls /sys/bus/i2c/drivers/foo/. No probe means DT mismatch (compatible), status = "disabled", module not loaded, or stuck in deferred probe (/sys/kernel/debug/devices_deferred).
  2. Is it on the bus? i2cdetect -y <bus>: does the address ACK? If not, check wiring, wrong bus number, reset GPIO, or the supply is off.
  3. Power and pins cat /sys/kernel/debug/regulator/regulator_summary for the vdd-supply; /sys/kernel/debug/pinctrl/ to confirm SDA/SCL mux; clk_summary for the controller clock.
  4. Interrupt Right GPIO, trigger type and pull; watch /proc/interrupts while moving the sensor.

Order: DT match, probe, power, bus ACK, IRQ.

4. Bring up a new board

  1. Power and clocks PMIC rails, reset, oscillators. Get a UART console alive first.
  2. DDR Training in the bootloader; then boot the kernel with a board DT derived from the reference design.
  3. Storage UFS/eMMC so rootfs can mount; then USB (adb/fastboot).
  4. Peripherals one by one Display, touch, audio, sensors, modem, Wi-Fi: add the DT node, confirm probe(), validate with sysfs or a test tool.
  5. Power tuning Suspend/resume, idle current, thermal limits, then performance.

5. High idle battery drain: find the wakeup source

  1. Is it suspending at all? dmesg for "PM: suspend entry/exit"; suspend_stats success count.
  2. Who blocks it? /sys/kernel/debug/wakeup_sources: the source with growing active_since or total_time. dumpsys power for app partial wakelocks.
  3. Aborts suspend_stats last_failed_dev: a driver whose suspend() fails aborts every attempt.
  4. Who wakes it? Compare /proc/interrupts before and after a sleep period; dmesg "wakeup IRQ"; Perfetto or Battery Historian timelines. Common culprits: chatty sensor, modem, misconfigured GPIO wake IRQ, app wakelock.

6. Boot is too slow

  1. Measure per stage Bootloader logs, kernel dmesg timestamps, "Freeing unused kernel memory" (kernel to init), init service start markers, sys.boot_completed time, Perfetto boot trace.
  2. Kernel initcall_debug to find slow initcalls; make drivers modules loaded later or use asynchronous probe (PROBE_PREFER_ASYNCHRONOUS); remove unnecessary delays (msleep in probe).
  3. Userspace Parallelize init services, lazy-start HALs, reduce system_server and Zygote preload work, avoid blocking on wait_for_prop.
  4. Storage Read-ahead tuning, erofs compression, verify dm-verity is not the bottleneck.

Optimize the biggest measured contributor, not guesses.

7. Random reboots or watchdog resets

  1. Read the reboot reason Bootloader log, ro.boot.bootreason, pstore. Watchdog, thermal, kernel panic, PMIC fault?
  2. Watchdog bite A CPU stopped petting the watchdog: stuck in a spinlock or IRQs-off loop, or deadlock. Collect the ramdump and inspect each CPU's stack and lock owners (crash, T32).
  3. Thermal Check /sys/class/thermal trip points and throttling logs.
  4. Hardware/power PMIC fault registers, voltage droop under load (brown-out).

8. "BUG: scheduling while atomic"

  1. Cause A sleeping function (mutex_lock, kmalloc(GFP_KERNEL), msleep, copy_to_user, I2C transfer) called while holding a spinlock, with preemption/IRQs disabled, or in a hard IRQ/softirq.
  2. Find The stack trace in the splat shows the sleeping call and the atomic context.
  3. Fix Use GFP_ATOMIC, restructure to drop the spinlock before sleeping, or move work to a threaded IRQ / workqueue.
  4. Prevent Build debug kernels with lockdep and CONFIG_DEBUG_ATOMIC_SLEEP.

9. Memory leak or OOM kills under sustained use

  1. Kernel or userspace? /proc/meminfo: growing Slab (SUnreclaim), KernelStack, VmallocUsed point to the kernel; growing app PSS points to userspace.
  2. Kernel kmemleak, slabtop / /proc/slabinfo for the growing cache, DMA-BUF totals; look for alloc in probe/IRQ/ioctl paths without a matching free.
  3. Userspace dumpsys meminfo <pkg>, showmap, heapprofd in Perfetto, malloc debug or ASan/HWASan.
  4. Confirm Trend MemAvailable over a long run after the fix.

10. UART console is silent on a new board

  1. earlycon Is earlycon on the command line with the correct UART base address? Is stdout-path set in /chosen?
  2. Pins and clocks Pinmux of TX/RX, baud rate, UART clock and regulator enabled.
  3. Pre-kernel? If the bootloader is silent too, the problem is earlier: bootloader UART config, DDR or clocks not up.
  4. Fallback Toggle a GPIO/LED at known points, or attach JTAG, to see how far boot gets.

11. High interrupt latency or UI jank traced to the kernel

  1. Trace Perfetto with sched, irq and freq events: are UI threads Runnable but not Running (CPU contention), sleeping in D state (I/O or lock), or running on a little core at low frequency?
  2. IRQs-off hogs ftrace irqsoff/preemptoff tracers show the longest disabled regions and their call sites.
  3. Fix Move long work out of hard IRQ, split long spinlock sections, adjust IRQ affinity, fix uclamp/cpuset placement.
Interview angle For scenario questions, interviewers care more about your method than the final answer. State your first question ("at what stage does it stop?"), name specific commands and files, explain what each result would rule in or out, and end with prevention (a test, a debug config, a regression gate).

Cheat sheet

Commands and paths worth memorizing.

Analogy

A mechanic's most-used wrench set, laid out in order on the bench. You do not need every tool in the shop to fix most cars, just these, within reach. The table below is that bench for kernel and BSP work.

CommandUse
dmesg -w / logcat -b kernelLive kernel log
cat /sys/fs/pstore/console-ramoops-0Previous boot's kernel log after a crash
cat /proc/interruptsIRQ counts per CPU
cat /sys/kernel/debug/wakeup_sourcesWhat blocks suspend
cat /sys/kernel/debug/suspend_statsSuspend failures and last failed device
i2cdetect -y NScan an I2C bus
echo function_graph > /sys/kernel/tracing/current_tracerftrace call graph
perfetto -t 10s sched freq idle binder_driverSystem trace
top -H -b -n1Per-thread CPU snapshot
addr2line -e vmlinux ADDRSymbolize an oops address
lsmod, modprobe, modinfoKernel modules
getprop / setpropAndroid properties
dumpsys power, dumpsys meminfoWakelocks and PM state, memory by process
strace -f -p PIDSyscalls of a running process
echo 'file foo.c +p' > /sys/kernel/debug/dynamic_debug/controlTurn on dev_dbg for a file
echo t > /proc/sysrq-triggerDump all task stacks to dmesg (w = blocked tasks, c = crash)
PathWhat
/proc/meminfo, /proc/slabinfo, /proc/buddyinfoMemory breakdown and fragmentation
/proc/pressure/memoryPSI (what LMKD watches)
/proc/<pid>/maps, smaps, stack, statusProcess memory, kernel stack, state
/proc/device-tree/Device tree the kernel booted with
/sys/power/stateTrigger or inspect suspend
/sys/kernel/debug/clk/clk_summaryClock tree state and rates
/sys/kernel/debug/regulator/regulator_summaryRegulator state and consumers
/sys/kernel/debug/devices_deferredDevices waiting in deferred probe
/sys/class/thermal/Thermal zones and trip points
/sys/devices/system/cpu/cpu*/cpufreq/CPU frequency and governor
/dev/binder, /dev/hwbinder, /dev/vndbinderBinder IPC domains
/dev/dma_heap/DMA-BUF heaps (ION replacement)
init.rc, *.rc, ueventd.rcService/action definitions, device node permissions
vmlinux, System.mapSymbols for debugging
ContextMay sleep?Allocation flagLock choice
Process context (syscall, kthread, workqueue, threaded IRQ)YesGFP_KERNELmutex, rw_semaphore, spinlock
Softirq / tasklet / timer callbackNoGFP_ATOMICspinlock (_bh from process side)
Hard IRQ handlerNoGFP_ATOMICspinlock (_irqsave from process side)
Holding a spinlock / preemption disabledNoGFP_ATOMICOnly other spinlocks, in a fixed order
Interview angle Being able to quote exact paths like wakeup_sources, clk_summary or devices_deferred without hesitation is one of the quickest ways to show real hands-on experience. The context table above answers a whole family of "can I call X here?" questions.

Clocks, pinctrl, regulators and wait queues

Most "the driver probed but the device does nothing" bugs are not in the driver logic. They are a disabled clock, a pin left in the wrong mux, or a rail that never came up. These three frameworks, plus wait queues for sleeping on events, are daily BSP tools.

Analogy

A workshop machine. The regulator is the wall socket (voltage must be present before anything else), the clock is the motor speed control (the block is gated or running at a rate), and pinctrl is the gearbox that decides whether a shaft is driving the lathe or sitting as a spare handle (the same SoC ball is UART TX or GPIO). A wait queue is the "ding" when a batch finishes: the worker sleeps instead of spinning, and the interrupt handler rings the bell.

The three providers

  • clk: devm_clk_get, clk_prepare_enable, clk_set_rate, clk_disable_unprepare. Prepare may sleep (it can turn on a parent PLL); enable is the cheap atomic gate. Dump the tree with /sys/kernel/debug/clk/clk_summary. A clock left enabled is a classic idle-current leak.
  • pinctrl: a pin has a mux function (UART, I2C, GPIO, unused), pull, and drive strength. DT uses pinctrl-0 / pinctrl-names = "default" and often a "sleep" state. Confirm mux in /sys/kernel/debug/pinctrl/. GPIO via gpiod_get is for lines used as GPIO, not for "this pad is I2C SDA".
  • regulator: PMIC rails. Consumers declare vdd-supply = <&ldo9> and call regulator_enable. Check /sys/kernel/debug/regulator/regulator_summary for voltage, enable count and who holds it.

Order in probe() is almost always: pinctrl default state, enable supplies, enable clocks, deassert reset, then talk to the device. Reverse that in remove / suspend. Missing any one of them is the usual reason i2cdetect sees no ACK.

Wait queues and completions

  • Wait queue: wait_event_interruptible(wq, condition) sleeps in process context until the condition is true. The waker sets the condition and calls wake_up (from IRQ or another thread). Use the _interruptible or _killable variants so SIGKILL can break the wait; a plain wait_event is how processes get stuck in D state.
  • Completion: a one-shot "this finished" signal (wait_for_completion / complete). Prefer it over a binary semaphore for "wait until the IRQ says DMA is done".
Interview angle "Walk me through probe()" should name pinctrl, regulator and clk in that order, plus -EPROBE_DEFER if a provider is missing. "A process cannot be killed" should name uninterruptible wait_event. Quote clk_summary, regulator_summary and pinctrl debugfs as your first three dumps.

Quick revision

  • BSP = bootloader + device tree + kernel defconfig + drivers that adapt Linux/Android to one SoC and board.
  • Linux is monolithic (all core code in one privileged address space) with loadable modules (.ko, =m in Kconfig).
  • User code runs at EL0, the kernel at EL1; a syscall traps via svc, dispatches through the syscall table and returns with eret.
  • Kernel code must use copy_to_user/copy_from_user for user pointers.
  • The device tree describes non-discoverable hardware; compatible matches a driver's of_match_table, then probe() runs.
  • .dts/.dtsi compile to .dtb; overlays (.dtbo) patch board variants; status = "okay" enables a node.
  • Driver model: bus + device + driver; -EPROBE_DEFER retries when a dependency is not ready; devm_* frees resources automatically.
  • Character drivers expose file_operations; block drivers go through the block layer; network drivers expose net_device, no /dev node.
  • Top half: minimal, cannot sleep. Bottom half: softirq and tasklet cannot sleep; workqueue and threaded IRQ can.
  • Threaded IRQs with IRQF_ONESHOT are the modern default for device drivers.
  • Sleeping allowed only in process context without spinlocks held; otherwise "scheduling while atomic".
  • kmalloc = physically contiguous (DMA-able); vmalloc = virtually contiguous only; GFP_ATOMIC in atomic context.
  • Buddy allocator hands out 2^order pages; SLUB caches small objects on top.
  • DMA on non-coherent ARM needs cache clean before device reads and invalidate before CPU reads; use the DMA API.
  • ION is gone; DMA-BUF heaps share zero-copy buffers by fd.
  • MMU walks page tables; TLB caches translations; ASIDs avoid full TLB flushes on context switch.
  • Page faults: minor (in RAM), major (needs I/O), invalid (SIGSEGV or kernel oops).
  • fork uses copy-on-write; that is why Zygote forking is fast and shares preloaded memory.
  • Android has no disk swap; ZRAM compresses anonymous pages; LMKD kills by oom_score_adj using PSI.
  • Linux schedules tasks (task_struct); threads are tasks sharing an address space via clone flags.
  • CFS picks the task with the smallest vruntime from a red-black tree; EEVDF replaced it in Linux 6.6.
  • RT classes (SCHED_FIFO, SCHED_RR, 1-99) always preempt normal tasks; SCHED_DEADLINE is above them.
  • EAS uses an energy model to place tasks on big.LITTLE cores; uclamp and cpusets steer Android app groups.
  • Spinlock = busy-wait, short, any context; mutex = sleeps, owner, process context; semaphore = counter, no owner.
  • Share data with an IRQ handler using spin_lock_irqsave; RCU for read-mostly data.
  • volatile is not synchronization; use atomics, barriers, READ_ONCE/WRITE_ONCE or locks.
  • Deadlock needs mutual exclusion, hold-and-wait, no preemption and circular wait; lock ordering breaks it; lockdep detects it.
  • Priority inversion is fixed by priority inheritance (rt_mutex, PI futex, Binder priority propagation).
  • Binder: one copy, kernel-verified caller identity, reference counting, death notification, thread pool.
  • Android filesystems: f2fs for /data, erofs for read-only partitions, ext4 still common; fsync for durability.
  • Suspend: freeze tasks, suspend devices, CPUs off, wake on wake IRQ; wakeup sources block it.
  • GKI 1.0 optional in Android 11; GKI 2.0 required for new Android 12+ devices on kernel 5.10+. Google-built core kernel, stable KMI, vendor modules in vendor_boot/vendor_dlkm.
  • Android typically sets /proc/sys/kernel/panic_on_oops to 1, so an oops becomes a panic; the kernel default is to keep running.
  • Emulated /sdcard is MediaProvider FUSE since Android 11; sdcardfs was the Android 6โ€“10 in-kernel wrapper and is gone.
  • Probe order: pinctrl, regulators, clocks, then talk to the device. Debug with clk_summary, regulator_summary and pinctrl debugfs.
  • Prefer wait_event_interruptible / wait_event_killable or a completion; plain wait_event can leave a task unkilleable in D state.
  • Treble: stable AIDL HALs between system and vendor partitions, checked by VINTF manifests and compatibility matrices.
  • Debug kit: dmesg, pstore, ftrace, Perfetto, perf, addr2line, crash/T32 ramdumps, KASAN, lockdep, kmemleak.

Glossary

ACK
Android Common Kernel: Google's kernel branches from which GKI builds are made.
ASID
Address Space Identifier: tag on TLB entries so translations from different processes can coexist.
AVB
Android Verified Boot: verifies boot images via signed vbmeta; dm-verity checks read-only partitions at runtime.
big.LITTLE
ARM design mixing energy-efficient and high-performance CPU cores in one SoC.
Binder
Android's kernel-backed RPC/IPC mechanism with one-copy transfer and caller identity.
BSP
Board Support Package: bootloader, device tree, config and drivers for a specific board.
Buddy allocator
Kernel page allocator that splits and merges power-of-two blocks of physical pages.
CFS
Completely Fair Scheduler: vruntime-based fair scheduler for normal tasks (replaced by EEVDF in 6.6).
CMA
Contiguous Memory Allocator: reserved region lent to movable pages and reclaimed for large contiguous DMA buffers.
Context switch
Saving one task's CPU state and restoring another's, possibly switching page tables.
COW
Copy-On-Write: pages shared read-only after fork and copied only when written.
Deferred probe
Retrying a driver's probe later when it returns -EPROBE_DEFER because a dependency is missing.
Device tree (DT)
Data structure describing non-discoverable hardware; compiled from .dts to .dtb.
DMA
Direct Memory Access: a device transfers data to or from RAM without the CPU copying it.
DMA-BUF
Kernel framework for sharing buffers between devices and processes via file descriptors.
DTBO
Device Tree Blob Overlay: patch applied to a base DTB for a board variant.
EAS
Energy Aware Scheduling: places tasks using a per-CPU energy model.
EEVDF
Earliest Eligible Virtual Deadline First: Linux's default fair scheduler since 6.6.
Completion
One-shot kernel signal: one side wait_for_completion, the other complete.
erofs
Enhanced Read-Only File System: compressed read-only filesystem used for Android system partitions.
FUSE
Filesystem in Userspace. Android 11+ uses MediaProvider FUSE for emulated shared storage; it replaced sdcardfs.
f2fs
Flash-Friendly File System: log-structured filesystem used for /data.
ftrace
Kernel's built-in tracer for functions, tracepoints and latency.
Futex
Fast userspace mutex: syscall used by pthread locks only when contention occurs.
GFP flags
"Get Free Pages" flags telling the allocator whether it may sleep, do I/O, etc.
GIC
ARM Generic Interrupt Controller: routes interrupts to CPU cores.
GKI
Generic Kernel Image: Google-built Android core kernel with vendor code in modules. Optional in Android 11; required for new Android 12+ devices on 5.10+.
HAL
Hardware Abstraction Layer: stable interface between the Android framework and vendor code.
inode
Filesystem object holding a file's metadata and data location, but not its name.
IOMMU / SMMU
Translates device DMA addresses and isolates devices from each other's memory.
KASAN
Kernel Address Sanitizer: detects out-of-bounds and use-after-free accesses.
KMI
Kernel Module Interface: the stable set of symbols and types GKI vendor modules may use.
LMKD
Low Memory Killer Daemon: Android userspace process killer driven by PSI and oom_score_adj.
lockdep
Kernel lock validator that detects lock-order problems and misuse.
MMU
Memory Management Unit: hardware that translates virtual to physical addresses.
Oops / panic
Kernel error that kills the current task / fatal error that halts or reboots the system.
Page fault
Exception raised when a virtual address has no valid translation or wrong permissions.
Priority inversion
High-priority task waits on a lock held by a low-priority task that a medium task keeps preempting.
PSI
Pressure Stall Information: time tasks spend stalled on CPU, memory or I/O.
pstore / ramoops
Reserved RAM that preserves kernel logs across a crash and reboot.
Ramdump
Full memory image captured after a crash for offline analysis.
RCU
Read-Copy-Update: synchronization with lock-free readers and deferred reclamation.
Softirq
Statically defined bottom-half mechanism for high-rate kernel work; runs in atomic context.
Spinlock
Busy-waiting lock for short critical sections; holder must not sleep.
Tasklet
Simple deferred-work mechanism built on softirqs; cannot sleep.
TLB
Translation Lookaside Buffer: cache of recent virtual-to-physical translations.
Treble
Android architecture separating the framework (system) from vendor implementation.
ueventd
Android daemon that creates /dev nodes and loads firmware in response to kernel uevents.
VINTF
Vendor Interface objects: manifests and compatibility matrices checked between system and vendor.
Wakeup source
Kernel object (wakelock) that prevents system suspend while active.
Workqueue
Deferred work executed by kernel worker threads in process context; can sleep.
ZRAM
Compressed RAM block device used as swap on Android.

Interview questions

Fundamentals

What is a BSP and what does it contain?

A Board Support Package is the software that makes a generic OS run on a specific SoC and board. It typically contains the bootloader (or its board port), the device tree describing the board's hardware, the kernel configuration (defconfig), and drivers for every peripheral (clocks, pinctrl, regulators, storage, display, sensors, modem interfaces). On Android it also extends into userspace: vendor HALs, init .rc files, ueventd.rc permissions and SELinux vendor policy. BSP work means board bring-up, driver and DT development, and debugging boot, power and peripheral issues.

What is the role of an operating system kernel?

The kernel is the privileged core that manages hardware and shares it safely among programs. Its main jobs are: process and thread management and CPU scheduling; memory management (virtual memory, allocation, protection); device management through drivers; filesystems and I/O; networking; inter-process communication; and security (permissions, isolation). Programs request these services through system calls rather than touching hardware directly.

Monolithic kernel vs microkernel: which is Linux and what are the trade-offs?

Linux is monolithic: scheduler, memory management, filesystems, networking and drivers run together in kernel space and call each other directly, which is fast. The downside is weaker isolation: a bug in any driver can crash the whole system. A microkernel (QNX, seL4, Fuchsia's Zircon) keeps only scheduling, IPC and basic memory in the kernel and runs drivers and filesystems as user processes, giving better fault isolation at the cost of IPC overhead. Linux adds flexibility with loadable modules, so it is often called a modular monolithic kernel.

What is the difference between user space and kernel space?

They are separate privilege levels and address-space regions. User space (EL0 on ARM64, ring 3 on x86) runs applications with no direct hardware access and a private virtual address space per process. Kernel space (EL1 / ring 0) runs the kernel with full access to memory and devices. A bug in user space kills only that process; a bug in kernel space can oops or panic the whole system. User code enters the kernel only through defined entry points: system calls, exceptions (like page faults) and interrupts.

What is a system call and how does it work?

A system call is the controlled way user code asks the kernel for a service (open a file, allocate memory, create a process). The libc wrapper places the syscall number and arguments in registers and executes a trap instruction (svc #0 on ARM64, syscall on x86-64). The CPU switches to kernel mode and jumps to the exception vector; the kernel saves registers, looks up the handler in the syscall table, runs it (copying data with copy_from_user/copy_to_user) and returns the result with eret. Errors come back as negative values that libc turns into -1 plus errno.

What is a device tree and why does Linux use it?

A device tree is a data structure (.dts source compiled to a .dtb blob) that describes a board's hardware: buses, register addresses, interrupts, clocks, regulators, GPIOs. It is kept separate from kernel code so one kernel image can support many boards. At boot the kernel creates devices from DT nodes and matches each node's compatible string to a driver's of_match_table, then calls probe(). Before DT, ARM used per-board C files, which did not scale. Overlays (.dtbo) patch the base tree for board variants.

What is a kernel module? How do you load and unload one?

A kernel module (.ko) is code that can be linked into the running kernel on demand, typically a driver. It has an init function (module_init) and exit function (module_exit). Load with insmod file.ko (no dependency handling) or modprobe name (resolves dependencies via modules.dep); unload with rmmod; list with lsmod; inspect with modinfo. In Kconfig, =y builds code into the kernel image and =m builds it as a module. With GKI, all vendor drivers must be modules.

What are the main types of Linux device drivers?
  • Character drivers: byte-stream devices accessed via /dev nodes and file_operations (open, read, write, ioctl, mmap), e.g. serial ports, input, sensors, Binder.
  • Block drivers: random-access storage in fixed-size blocks, going through the page cache and block layer, e.g. UFS, eMMC, NVMe.
  • Network drivers: packet interfaces (net_device) used via sockets, no /dev node, e.g. Wi-Fi, Ethernet.

Orthogonally, drivers sit on buses: platform (on-SoC, from DT), I2C, SPI, USB, PCI.

What is a platform driver?

A platform driver handles a device on the "platform bus", a virtual bus for hardware that cannot be discovered by enumeration, typically IP blocks inside the SoC (UART, I2C controller, GPIO controller, watchdog). The devices are created from device-tree nodes (or ACPI, or board code). The driver registers a struct platform_driver with probe, remove and an of_match_table; when a matching node exists, probe() gets the platform_device, maps its registers (devm_platform_ioremap_resource), gets its IRQ (platform_get_irq) and clocks, and initializes the hardware.

What is an interrupt and how is it different from polling?

An interrupt is a hardware signal that makes the CPU pause its current work and run a handler, so the device notifies the CPU when something happens. Polling means the CPU repeatedly checks a status register. Interrupts save CPU time and power when events are infrequent; polling can be better at very high event rates (it avoids per-event interrupt overhead) or when latency must be tightly controlled. Linux's network NAPI mixes both: interrupt to start, then poll while traffic is heavy.

What are the top half and bottom half of an interrupt?

The top half is the hard IRQ handler: it runs immediately in interrupt context, cannot sleep, and should do the minimum (acknowledge the hardware, read urgent data, schedule the rest). The bottom half does the remaining work later in a less time-critical context: a softirq (high-frequency, e.g. networking), a tasklet (simple, built on softirq), a workqueue (kernel thread, can sleep), or a threaded IRQ (dedicated kernel thread, can sleep). The split keeps interrupts-off time short so other interrupts are not delayed or lost.

What is the difference between a process and a thread?

A process is a running program with its own address space and resources (open files, credentials, signal handlers). A thread is an execution flow within a process that shares its address space and resources but has its own stack, registers and scheduling state. Threads communicate cheaply through shared memory but need locking; processes are isolated and need IPC, but a crash in one does not corrupt another. In Linux both are tasks (task_struct) created with clone(); the flags decide what is shared.

What are the states of a process in Linux?
  • R running or runnable (on a run queue).
  • S interruptible sleep: waiting for an event, can be woken by signals.
  • D uninterruptible sleep: usually waiting for I/O; ignores signals.
  • T stopped (SIGSTOP) or traced by a debugger.
  • Z zombie: exited but not yet reaped by its parent.

You can see them in ps output or /proc/<pid>/status.

What is a zombie process and what is an orphan process?

A zombie has exited, but its parent has not yet called wait()/waitpid(), so the kernel keeps its PID and exit status in the process table. It uses no memory beyond that entry, but many zombies can exhaust PIDs; fix the parent to reap children (or handle SIGCHLD). An orphan is a running process whose parent died; it is re-parented to init (PID 1) or a designated subreaper, which reaps it when it exits.

What is virtual memory and why is it used?

Virtual memory gives each process its own address space that the MMU translates to physical memory using page tables. Benefits: isolation (processes cannot access each other's memory), protection (read-only code, no-execute data), simpler programming (each process sees a large contiguous space), lazy allocation (pages allocated on first use), sharing (libraries and COW pages mapped into many processes), memory-mapped files, and the ability to swap or compress rarely used pages.

What is paging and what is a page table?

Paging divides virtual and physical memory into fixed-size pages (typically 4 KB). A page table maps each virtual page to a physical frame and stores permission bits (read, write, execute, user/kernel) and status bits (valid, accessed, dirty). To save space, page tables are multi-level (4 levels on ARM64 with 48-bit addresses); the MMU walks them on a TLB miss. The kernel builds and updates the tables; the hardware reads them.

What is a TLB?

The Translation Lookaside Buffer is a small, fast cache in the MMU holding recent virtual-to-physical translations. A TLB hit translates an address in about a cycle; a miss requires a page-table walk of several memory accesses. When mappings change, the kernel must invalidate affected entries (TLB shootdown on SMP, using broadcast TLBI on ARM). ASIDs tag entries per address space so switching processes does not require flushing the whole TLB. Huge pages increase TLB coverage.

What is a page fault?

A page fault is an exception raised when the MMU cannot complete an access: the page is not present or the access violates permissions. The kernel handler checks whether the address belongs to a valid VMA. If yes, it resolves the fault: allocate a zero page, load a file page, swap in, or copy a COW page (minor fault if no I/O, major if I/O). If not, the process gets SIGSEGV, or the kernel oopses if it happened in kernel code. Faults are normal and frequent; excessive major faults indicate memory pressure or thrashing.

What is a context switch?

A context switch is the kernel stopping one task and running another on the same CPU. It saves the current task's registers, stack pointer and program counter, picks the next task, switches the address space if the new task is in a different process (page table base plus ASID), and restores the new task's registers. It is triggered by blocking, preemption (time slice expiry or higher-priority wakeup) or yielding. Direct cost is a few microseconds; the indirect cost of cold caches and TLB often dominates.

What is a mutex? What is a semaphore?

A mutex is a lock for mutual exclusion: only one holder at a time, it has an owner, only the owner can unlock it, and waiters sleep. A semaphore is a counter: down() decrements it and sleeps if it would go negative, up() increments it and wakes a waiter. A counting semaphore with N allows N concurrent holders (e.g. a pool of N buffers); a binary semaphore resembles a mutex but has no owner, so any task can signal it. In the Linux kernel, prefer mutex for exclusion and completions for signaling.

What is a spinlock?

A spinlock is a lock where a waiting CPU busy-loops until the lock becomes free, instead of sleeping. It is used for very short critical sections and in contexts that cannot sleep (interrupt handlers, softirqs). In Linux, taking a spinlock disables preemption; the holder must not sleep. Variants like spin_lock_irqsave also disable local interrupts to prevent deadlock with an IRQ handler that takes the same lock. Spinning is only efficient if the hold time is shorter than the cost of sleeping and waking.

What is a deadlock and what are the four necessary conditions?

A deadlock is when a set of tasks each wait for a resource held by another task in the set, so none can proceed. All four Coffman conditions must hold: mutual exclusion (resources are held exclusively), hold and wait (tasks hold resources while waiting for more), no preemption (resources cannot be forcibly taken), and circular wait (a cycle in the wait-for graph). Breaking any one prevents deadlock; the most practical is a global lock ordering to prevent circular wait.

What IPC mechanisms does Linux provide?
  • Pipes and named pipes (FIFOs): one-way byte streams.
  • Signals: asynchronous notifications carrying only a number.
  • Shared memory (POSIX shm_open/mmap, System V, memfd): fastest, needs separate synchronization.
  • Message queues (POSIX, System V): discrete messages.
  • Sockets: Unix domain (local, can pass fds) and TCP/UDP (network).
  • Netlink: kernel-userspace messaging.
  • eventfd, futex, semaphores for notification and synchronization.

Android adds Binder as its main IPC.

What are /proc and /sys?

Both are virtual filesystems generated by the kernel on the fly, with no storage behind them. /proc (procfs) exposes process information (/proc/<pid>/status, maps, fd) and kernel state (/proc/meminfo, /proc/interrupts, /proc/cmdline). /sys (sysfs) exposes the device model: devices, drivers, buses and classes, with one value per attribute file, many writable (e.g. cpufreq governor, power controls). debugfs (/sys/kernel/debug) is for developer-only debug data.

What is dmesg?

dmesg prints the kernel ring buffer: messages written with printk by the kernel and drivers, such as boot progress, driver probe results, errors, warnings, oopses and SELinux denials. Each line has a timestamp since boot and a log level. dmesg -w follows new messages. On Android, logcat -b kernel shows the same messages; logs from the previous boot after a crash are in pstore (/sys/fs/pstore/console-ramoops-0).

What is the difference between kmalloc and vmalloc?

kmalloc returns memory that is physically contiguous and in the kernel's linear map. It is fast, suitable for DMA, but large requests can fail when memory is fragmented. vmalloc returns memory that is only virtually contiguous, built from scattered physical pages mapped into a separate kernel virtual range. It suits large buffers the CPU alone uses but is slower (page-table setup, more TLB pressure) and not directly DMA-able. kvmalloc tries kmalloc and falls back to vmalloc. In atomic context use kmalloc(..., GFP_ATOMIC); vmalloc may sleep.

What is copy-on-write?

Copy-on-write lets two processes share the same physical pages until one of them writes. After fork(), the kernel copies the page tables and marks shared pages read-only in both. A write triggers a page fault; the kernel then copies just that page, gives the writer a private writable copy, and resumes. This makes fork() fast and memory-efficient, especially when followed by exec(). Android's Zygote relies on it: apps forked from Zygote share its preloaded classes and resources.

What does fork() do, and how is it different from exec()?

fork() creates a new child process that is a copy of the parent (same code, same memory contents via COW, duplicated file descriptors). It returns twice: 0 in the child, the child's PID in the parent. exec() replaces the current process's program image with a new executable, keeping the PID and (unless close-on-exec) open file descriptors. The classic pattern for launching a program is fork then exec in the child, with the parent calling wait().

What do the clk, pinctrl and regulator frameworks do?

They are the three providers almost every SoC driver depends on. clk gates and rates clocks (clk_prepare_enable, clk_set_rate); prepare may sleep, enable is the cheap gate. pinctrl muxes SoC balls to a function (UART, I2C, GPIO) and sets pull and drive; GPIO via gpiod is only for lines used as GPIO. regulator turns PMIC rails on and sets voltage. In probe() the usual order is pinctrl default, enable supplies, enable clocks, deassert reset, then talk to the hardware. First dumps: clk_summary, regulator_summary, /sys/kernel/debug/pinctrl/.

How does Android implement /sdcard? FUSE or sdcardfs?

It changed twice. Early Android used a userspace FUSE daemon (sdcard) so emulated storage could enforce per-app permissions on top of /data/media. Android 6 replaced that with sdcardfs, an in-kernel wrapfs, for much lower overhead. Android 11 removed sdcardfs and went back to FUSE, this time driven by MediaProvider, to implement scoped storage and richer permission checks. On a current device, /sdcard is FUSE; saying "we use sdcardfs" is an Android 6โ€“10 answer.

What is Project Treble?

Treble (Android 8) is an architecture that separates the Android framework (system partition) from the vendor implementation (vendor partition) through stable, versioned HAL interfaces (HIDL originally, stable AIDL now) over Binder. VINTF manifests and compatibility matrices check that system and vendor are compatible. The benefit is that Google or the OEM can update the framework without the SoC vendor rebuilding its code, which speeds up OS upgrades and is verified by running a Generic System Image with VTS/CTS.

What is GKI?

GKI (Generic Kernel Image) is Google's approach to fragmentation in Android kernels. Google builds and signs one core kernel per Android release and kernel version from the Android Common Kernel. SoC and board support is delivered as vendor modules loaded from vendor_boot and vendor_dlkm, and they may use only the stable Kernel Module Interface (KMI). This lets the core kernel receive security fixes without vendors rebasing a fork. GKI 1.0 appeared optionally in Android 11; GKI 2.0 is required for new devices launching with Android 12+ on kernel 5.10+. The OEM still tests and ships the combined image; "independent update" means the core binary and vendor modules can move on different cadences as long as the KMI holds.

Going deeper

Spinlock vs mutex: when do you use each?

Use a spinlock for very short critical sections and whenever the code cannot sleep: interrupt handlers, softirqs, or data shared with them. The holder must not sleep, and preemption is disabled while it is held. Use a mutex for longer sections in process context where waiting by sleeping is cheaper than spinning; the holder may sleep (e.g. do I2C transfers or GFP_KERNEL allocations). Calling a sleeping function while holding a spinlock causes "BUG: scheduling while atomic". If data is shared with a hard IRQ handler, process context must use spin_lock_irqsave so the IRQ cannot interrupt the holder on the same CPU and deadlock.

Mutex vs semaphore vs spinlock: which do you use in a kernel driver?
  • Mutex: sleepable, single owner, process context only; the default for protecting driver state touched only from syscalls, workqueues or threaded IRQs.
  • Spinlock: non-sleeping, usable in IRQ context (with _irqsave), short sections; use in the top half and for any data shared with it.
  • Semaphore: counting, no owner, up() can be called from IRQ; rarely used now for exclusion; completions or wait queues are preferred for signaling, and mutexes for exclusion.

Rule: in a hard IRQ handler use a spinlock; for long-held resources use a mutex in process context.

Tasklet vs workqueue vs threaded IRQ vs softirq: how do you choose?

Softirqs are statically defined and reserved for core high-rate subsystems (network RX/TX, block completion, timers, RCU); drivers do not add new ones. Tasklets are dynamically created on top of softirqs; they cannot sleep and the same tasklet never runs concurrently on two CPUs; they are deprecated in favor of alternatives. Workqueues run in kernel worker threads in process context, so they can sleep, take mutexes and do I2C/SPI transfers. Threaded IRQs give the interrupt its own kernel thread, also able to sleep, and are the modern default for device drivers, especially for devices on slow buses. Choose by whether you need to sleep and how latency-sensitive the work is.

How does a driver get bound to a device (from DT to probe)?
  1. The bootloader passes the DTB; the kernel unflattens it into device_node structures.
  2. The OF core creates platform_devices for nodes under simple buses with status = "okay"; bus controllers (I2C, SPI) create child devices when they probe.
  3. Drivers register with their bus (platform_driver_register), providing an of_match_table.
  4. When a device and driver are both present, the bus's match() compares compatible strings; on a match the core calls probe().
  5. Probe returns 0 (bound), -EPROBE_DEFER (retried later when dependencies appear), or an error.
  6. For modules, udev/ueventd uses the MODALIAS from MODULE_DEVICE_TABLE to autoload the right module.
What is deferred probing and why is it needed?

Drivers often depend on other providers: clocks, regulators, GPIO or pinctrl controllers, PHYs. If a provider has not probed yet, the consumer's probe() gets -EPROBE_DEFER from calls like devm_clk_get() and returns it. The driver core puts the device on a deferred list and retries whenever another driver binds successfully. This avoids fragile hand-tuned init ordering. /sys/kernel/debug/devices_deferred lists devices still waiting, which is a quick way to find a missing dependency. fw_devlink now orders probing automatically from DT phandles, reducing deferral churn.

What are devm_ (managed) resources?

devm_* functions (e.g. devm_kzalloc, devm_ioremap_resource, devm_request_irq, devm_clk_get) register the resource with the device, and the driver core releases them automatically, in reverse order, when probe fails or the device is unbound. This eliminates most error-path cleanup code and leaks. The caution: release order matters, for example a devm-requested IRQ is freed after remove() runs, so the handler must not use resources remove() already freed.

How do you write a simple character driver?

Implement a struct file_operations with the callbacks you need (open, read, write, unlocked_ioctl, poll, mmap, release), then register it. The simplest way is a miscdevice with a dynamic minor, which creates /dev/name; the fuller way is alloc_chrdev_region + cdev_init/cdev_add + class_create/device_create. In read/write, move data with copy_to_user/copy_from_user, return bytes transferred or a negative errno, protect shared state with a mutex, and support blocking I/O with a wait queue plus poll.

static ssize_t foo_read(struct file *f, char __user *buf, size_t n, loff_t *off)
{
    struct foo *foo = f->private_data;
    size_t len;
    if (wait_event_interruptible(foo->wq, foo->count))
        return -ERESTARTSYS;
    mutex_lock(&foo->lock);
    len = min(n, foo->count);
    if (copy_to_user(buf, foo->data, len)) { mutex_unlock(&foo->lock); return -EFAULT; }
    foo->count -= len;
    mutex_unlock(&foo->lock);
    return len;
}
Wait queue vs completion: which do you use in a driver?

A wait queue sleeps until a condition becomes true: wait_event_interruptible(wq, foo->ready), with the IRQ or worker setting the flag and calling wake_up. Use it for recurring "data available" or "state changed" waits. Prefer the interruptible or killable variant so signals (including SIGKILL) can break the sleep; a plain wait_event is how tasks get stuck in D state. A completion is a one-shot "this finished" event (wait_for_completion / complete), ideal for "wait until this DMA or probe helper is done". Do not use a binary semaphore for that pattern; completions are the modern API.

What is CMA and when do you need it?

The Contiguous Memory Allocator reserves a region of physical memory that can be lent to movable pages when idle and reclaimed for large physically contiguous DMA buffers when a device needs them (camera, display, codecs). Use it when a device cannot use the IOMMU / scatter-gather and needs big contiguous buffers that kmalloc cannot satisfy under fragmentation. Cost: that memory is less flexible for the rest of the system. On modern SoCs with an SMMU, prefer IOMMU + scatter-gather and keep CMA small.

What is ioctl and when should you use it?

ioctl(fd, cmd, arg) is a system call for device-specific operations that do not fit read/write, such as configuring a mode, starting a DMA, or querying capabilities. Commands are encoded with _IO, _IOR, _IOW, _IOWR macros that include a magic number, command number, direction and argument size. The driver implements unlocked_ioctl (and compat_ioctl for 32-bit apps on a 64-bit kernel), validates the command, and copies arguments with copy_from_user. For simple attributes, sysfs files are often cleaner; for streaming data, read/write or mmap.

How do you share data between an interrupt handler and process context safely?

Protect it with a spinlock. In the hard IRQ handler use spin_lock() (interrupts for that line are already masked on this CPU). In process context use spin_lock_irqsave(&lock, flags) / spin_unlock_irqrestore(), which disables local interrupts while holding the lock. Otherwise, if the IRQ fires on the same CPU while process context holds the lock, the handler spins forever. For single counters or flags, atomics may suffice; for producer/consumer data, a lock-free kfifo with one reader and one writer works without locking. Keep the locked region short and never sleep inside it.

Edge-triggered vs level-triggered interrupts?

An edge-triggered interrupt fires on a signal transition (rising or falling). It is raised once per event; if another edge happens while the first is being handled and the controller does not latch it, it can be missed. A level-triggered interrupt stays asserted as long as the line is at the active level; the handler must clear the source in the device, otherwise the interrupt fires again immediately (an interrupt storm). Level triggering is safer for shared lines. The DT interrupts property specifies the type, and getting it wrong is a common bring-up bug.

What are GFP_KERNEL and GFP_ATOMIC and when do you use each?

They are allocation flags. GFP_KERNEL is the normal flag: the allocator may sleep to reclaim memory, write back pages, or compact. Use it in process context when no spinlock is held and preemption is enabled. GFP_ATOMIC never sleeps and may use emergency reserves; use it in interrupt handlers, softirqs, timers, or while holding a spinlock. Atomic allocations fail more easily, so always handle NULL and avoid large atomic allocations. GFP_NOIO/GFP_NOFS prevent recursion into I/O or filesystem code during reclaim.

What are the buddy allocator and the slab allocator?

The buddy allocator manages physical memory as blocks of 2^order contiguous pages per zone. To allocate, it splits a larger free block into two "buddies" until it has the requested size; on free, it merges a block with its free buddy back into a larger block. That keeps external fragmentation manageable; /proc/buddyinfo shows free blocks per order. The slab allocator (SLUB in modern kernels) sits on top: it takes pages from the buddy allocator and carves them into caches of same-sized objects (e.g. task_struct, dentry), giving fast allocation, reduced internal fragmentation and per-CPU caching. kmalloc uses generic size-class slab caches. /proc/slabinfo and slabtop show usage.

How does DMA work, and what is cache coherency in that context?

With DMA, the driver gives the device a bus address and length; the device reads or writes RAM directly and interrupts when done, freeing the CPU. Problems arise because the CPU caches data: if the bus is not cache-coherent, the device may read stale memory (CPU's writes still in cache) or the CPU may read stale cache lines after the device wrote RAM. The DMA API handles this: streaming mappings (dma_map_single) clean caches before a device read and invalidate before a CPU read; coherent allocations (dma_alloc_coherent) use memory that is uncached or hardware-coherent. The API also translates to device addresses through the IOMMU if present.

What is an IOMMU (SMMU) and why is it useful?

An IOMMU is an MMU for devices: it translates device-visible addresses (IOVAs) to physical addresses using per-device page tables. Benefits: devices can use large buffers that are virtually contiguous but physically scattered (no need for big contiguous allocations); devices can be isolated so a buggy or malicious device cannot DMA into arbitrary memory; and 32-bit devices can reach memory above 4 GB. On ARM it is called SMMU. Faults show up as "smmu context fault" messages, which usually indicate a driver using an unmapped or freed buffer.

What happens on a page fault, step by step?
  1. The MMU fails to translate an address (not present) or detects a permission violation and raises an exception with the faulting address (FAR_EL1 on ARM64) and cause.
  2. The kernel's fault handler finds the VMA containing the address.
  3. If there is no VMA, or the access violates VMA permissions, it sends SIGSEGV to the user process (or oopses if in kernel mode without an exception fixup).
  4. If valid, handle_mm_fault() resolves it: allocate a zeroed anonymous page, map a page-cache page (reading from storage if needed = major fault), swap in from ZRAM, or break COW by copying.
  5. It updates the PTE, and the faulting instruction is restarted.
Minor vs major page fault?

A minor fault is resolved without I/O: the page is already in memory (page cache, shared with another process) or can be created (zero page, COW copy); the kernel only updates page tables. A major fault requires reading data from storage or swap, so the task sleeps for I/O and it is much slower (microseconds vs milliseconds). /proc/<pid>/stat fields and /usr/bin/time -v show counts. Many major faults during app start or scrolling indicate memory pressure or cold file cache.

How does the CFS scheduler decide which task runs next?

CFS tracks a virtual runtime (vruntime) per task: actual CPU time scaled inversely by the task's weight (derived from its nice value, with each nice level about a 10% share change). Runnable tasks sit in a per-CPU red-black tree ordered by vruntime; the scheduler picks the leftmost (smallest vruntime), i.e. the task that has received the least fair share. A running task is preempted when another task's vruntime is smaller by more than a granularity. Waking tasks get a vruntime close to the minimum so they run soon, which favors interactive tasks. Periodic load balancing moves tasks between CPUs. Since Linux 6.6, EEVDF replaces CFS's pick logic, adding virtual deadlines for better latency control.

What are the Linux scheduling policies?
  • SCHED_DEADLINE: EDF with runtime/period/deadline reservations; highest priority.
  • SCHED_FIFO: real-time priority 1-99; runs until it blocks or a higher-priority RT task arrives.
  • SCHED_RR: like FIFO with a time slice among tasks of equal priority.
  • SCHED_NORMAL/SCHED_OTHER: fair scheduling with nice -20 to +19.
  • SCHED_BATCH: CPU-bound, less frequent preemption.
  • SCHED_IDLE: runs only when nothing else wants the CPU.

Set with sched_setscheduler() or chrt; RT tasks are throttled to 95% of CPU by default so they cannot hang the system completely.

Why is a thread switch cheaper than a process switch?

Both save and restore registers and kernel stack. A process switch additionally changes the address space: it loads a new page table base (TTBR0 on ARM64, CR3 on x86), which may invalidate TLB entries (ASIDs/PCIDs mitigate a full flush) and leaves caches holding the old process's data. Threads of the same process share the address space, so page tables stay the same and the TLB and caches remain warm. Typical costs are a few microseconds for a process switch and noticeably less for a thread switch, plus indirect cache-miss costs afterwards.

What is the difference between fork(), vfork() and clone()?

fork() creates a child with a copy-on-write copy of the parent's address space and duplicated file descriptors. vfork() creates a child that borrows the parent's address space without copying page tables; the parent is suspended until the child calls exec() or _exit(). It is faster but dangerous if the child modifies memory; COW made it mostly unnecessary (though posix_spawn may use it internally). clone() is the underlying primitive: flags select what is shared (CLONE_VM, CLONE_FILES, CLONE_FS, CLONE_SIGHAND, CLONE_THREAD, namespaces). fork and pthread_create are both implemented with it.

What is priority inversion and how does the kernel handle it?

Priority inversion happens when a high-priority task waits for a lock held by a low-priority task, and a medium-priority task preempts the low-priority one, so the high-priority task is effectively blocked by a less important task for an unbounded time. The fix is priority inheritance: the lock holder temporarily inherits the highest waiter's priority until it releases the lock. Linux implements it in rt_mutex and PI futexes (PTHREAD_PRIO_INHERIT); PREEMPT_RT converts most kernel locks to rt_mutexes. Priority ceiling is an alternative. The Mars Pathfinder resets in 1997 are the classic example; Android's Binder also propagates caller priority to the server thread.

How do you prevent and detect deadlocks in practice?

Prevention: define and document a global lock order and always acquire in that order; keep lock scopes small; avoid calling unknown code (callbacks) with locks held; use trylock with back-off when ordering is impossible; avoid holding locks across sleeping calls that may need the same lock (e.g. flush_work). Detection: lockdep in debug kernels records lock-dependency chains and reports potential cycles even if the deadlock never actually happened; hung-task and soft-lockup detectors flag real hangs; echo w > /proc/sysrq-trigger dumps blocked tasks; in a ramdump, map lock owners to waiting tasks. In userspace, /data/anr traces and Java lock holders show similar cycles.

What is RCU and when would you use it?

Read-Copy-Update is a synchronization mechanism for read-mostly data. Readers enter a read-side critical section (rcu_read_lock()) that costs almost nothing and never blocks; they access data through rcu_dereference(). A writer makes a new copy of the element, updates it, publishes it atomically with rcu_assign_pointer(), then waits for a grace period (synchronize_rcu(), or call_rcu/kfree_rcu asynchronously) until all readers that might see the old version have finished, and only then frees it. It scales extremely well for readers on many CPUs. Use it for lookup tables, lists and configuration that change rarely; writers still need their own lock among themselves.

What are memory barriers and why are they needed?

Both compilers and CPUs reorder memory operations for performance. On a single CPU this is invisible, but on SMP another CPU may observe writes in a different order. ARM is weakly ordered, so this happens in practice. Memory barriers constrain ordering: smp_wmb() orders stores, smp_rmb() orders loads, smp_mb() orders both; acquire/release operations (smp_load_acquire, smp_store_release) give one-way ordering suited to flag-and-data patterns. For device registers, readl/writel include the needed I/O barriers and dma_wmb() orders descriptor writes before ringing a doorbell. Locks and atomic RMW operations with return values already imply barriers.

Why is volatile not enough for multithreaded code in C?

volatile only tells the compiler not to cache the variable in a register and not to optimize away or reorder accesses to that variable relative to other volatile accesses. It does not make operations atomic (x++ is still read-modify-write), does not prevent the CPU from reordering memory operations, and does not order non-volatile accesses around it. Use C11 atomics, kernel atomics, locks, or READ_ONCE/WRITE_ONCE plus barriers. volatile is appropriate for memory-mapped I/O registers, variables modified by signal handlers (volatile sig_atomic_t) and values changed by a debugger.

What is Binder and why does Android use it instead of standard Linux IPC?

Binder is Android's IPC and RPC mechanism, implemented as a kernel driver (/dev/binder). A client's proxy marshals arguments into a Parcel and calls ioctl(BINDER_WRITE_READ); the driver copies the data once into the server's mmap'd buffer and wakes a thread from the server's Binder thread pool, which runs the method and replies. Compared to pipes or sockets, it provides: caller UID/PID filled in by the kernel (unforgeable identity for permission checks), object references with reference counting across processes, death notifications, priority inheritance, fd passing, synchronous and one-way calls, and a name service (servicemanager). It is the backbone of app-framework and framework-HAL communication.

How many data copies do a pipe, Binder and shared memory need?

A pipe needs two copies: from the writer's buffer into a kernel buffer, then from the kernel buffer into the reader's buffer. Binder needs one copy: the driver copies from the sender's user memory directly into a buffer that is mmap'd into the receiver, so the receiver reads it without a second copy. Shared memory needs zero copies after setup, since both processes map the same physical pages, but it requires separate synchronization (futex, semaphore, eventfd). For large data such as graphics buffers, Android passes a DMA-BUF or memfd file descriptor over Binder to get zero-copy with Binder's security.

What is the page cache and how do writes reach storage?

The page cache holds file data in RAM in page-sized units, indexed per file. Reads hit the cache if possible; misses read from storage (with read-ahead for sequential access). write() copies data into the page cache and marks pages dirty, then returns; background writeback threads flush dirty pages to storage later based on age and dirty ratio thresholds. fsync()/fdatasync() force a file's dirty data (and metadata) to stable storage and issue a device cache flush. O_DIRECT bypasses the cache. Clean cache pages are the first thing reclaimed under memory pressure.

What is an inode? Hard link vs symbolic link?

An inode is the filesystem structure that holds a file's metadata (type, permissions, owner, size, timestamps, link count) and where its data blocks are, but not its name. Names live in directory entries that map a name to an inode number. A hard link is another directory entry pointing to the same inode; the file exists until the link count drops to zero and no process has it open; hard links cannot span filesystems or (normally) point to directories. A symbolic link is a separate small file containing a path; it can cross filesystems and point to directories, but becomes dangling if the target is removed.

Why does Android use f2fs for /data and erofs for /system?

/data gets many small random writes. f2fs (Flash-Friendly File System) is log-structured: it turns random writes into sequential ones, which suits NAND flash and its flash translation layer, improves write performance and reduces wear; it also supports file-based encryption and compression. /system and /vendor are read-only and verified by dm-verity. erofs (Enhanced Read-Only File System) compresses data in fixed-size output blocks, saving significant space while keeping fast random reads, and its immutable layout fits verified boot. ext4 remains in use on many devices and was the historical default for both.

mmap vs read: what are the differences?

read() copies data from the page cache into a user buffer on each call; it is simple and efficient for sequential streaming. mmap() maps the file's page-cache pages directly into the process's address space; data is loaded lazily by page faults and accessed with normal pointers, avoiding the extra copy and suiting random access to large files and sharing between processes. Costs of mmap: page-fault overhead, TLB pressure, SIGBUS if the file shrinks, and harder error handling. Drivers also implement mmap to expose device memory or DMA buffers to user space.

How does Linux suspend-to-RAM work, and what is a wakeup source?

When suspend is requested (on Android, autosleep once no wakelocks are held), the kernel freezes user tasks and freezable kernel threads, calls every driver's suspend callbacks (suspend, suspend_late, suspend_noirq) in child-before-parent order, takes secondary CPUs offline, and the last CPU enters a deep power state through PSCI firmware while RAM stays in self-refresh. A configured wake interrupt (RTC, button, modem, sensor hub) triggers resume in reverse order. A wakeup source is a kernel object (a kernel wakelock) that, while active, prevents suspend; drivers use pm_stay_awake/pm_relax or __pm_wakeup_event. If any driver's suspend callback fails, the whole suspend is aborted.

What is runtime PM?

Runtime power management lets individual devices power down while the system is running. Drivers call pm_runtime_get_sync() before using the hardware and pm_runtime_put() / pm_runtime_put_autosuspend() when done; the core keeps a usage count and, when it reaches zero (after an optional autosuspend delay), calls the driver's runtime_suspend to gate clocks and regulators, and runtime_resume when needed again. Power domains (genpd) can turn off a whole domain when all its devices are idle. Runtime PM is key to low active and idle current on mobile devices.

What is ueventd, and how do device nodes get their permissions on Android?

When a driver registers a device, the kernel sends a uevent over netlink containing the device's name, major/minor numbers and attributes. On Android, ueventd (part of init) receives it, creates the /dev node, and sets owner, group and mode according to ueventd.rc files (e.g. /dev/foo 0660 system input) and the SELinux label from file_contexts. It also handles firmware loading requests. If a HAL cannot open a node, check the node's permissions, SELinux label, and avc: denied messages.

What is the difference between an oops and a panic?

An oops is the kernel's report of a serious error (such as a bad memory access) in kernel code: it prints registers and a stack trace, kills the offending task, and the kernel may keep running, although it may be in an inconsistent state (for example a lock held forever) and is marked tainted. A panic is a fatal error after which the kernel stops; it halts or reboots after panic_timeout. The kernel default is panic_on_oops=0 (continue). Android typically sets /proc/sys/kernel/panic_on_oops to 1 from init (and ACK builds often also enable CONFIG_PANIC_ON_OOPS), so an oops becomes a panic and a ramdump or reboot rather than a half-dead system.

Advanced

Which kernel contexts may sleep, and why can't interrupt context sleep?

Process context (syscalls, kernel threads, workqueue items, threaded IRQ handlers) may sleep, as long as no spinlock is held and preemption or interrupts are not disabled. Hard IRQ, softirq and tasklet context, and any code holding a spinlock, must not sleep. Interrupt context borrows whatever task was running; it has no task of its own that the scheduler could put to sleep and wake later, and sleeping would leave the interrupted task and possibly held locks stuck. With a spinlock held, preemption is disabled, so sleeping could deadlock if the next task needs the same lock. might_sleep() annotations with CONFIG_DEBUG_ATOMIC_SLEEP catch violations at runtime.

How does the kernel implement a spinlock on ARM64?

Modern Linux uses queued spinlocks (qspinlock) on arm64; older kernels used ticket spinlocks. A ticket lock has "next" and "owner" counters: a CPU atomically takes a ticket and spins until owner equals its ticket, which gives FIFO fairness. Qspinlock keeps a 32-bit word for the fast uncontended path and, under contention, queues waiters in per-CPU MCS nodes, so each spins on its own cache line instead of hammering the shared lock line. Atomics use ARMv8.1 LSE instructions (like CAS, LDADD) or LDXR/STXR exclusive loops, and waiters use WFE to reduce power while spinning; acquire/release semantics provide the memory ordering. spin_lock also disables preemption.

What does PREEMPT_RT change in the kernel?

PREEMPT_RT (fully merged in Linux 6.12) makes almost all kernel code preemptible to achieve bounded latency. Key changes: most spinlock_t locks become sleeping rt_mutexes with priority inheritance (raw_spinlock_t remains true spinning for the few places that need it); interrupt handlers are forced into threads so they can be prioritized and preempted; softirqs run in thread context; and long non-preemptible sections are broken up. The trade-off is slightly lower throughput for much better worst-case latency. Drivers must use raw_spinlock_t only where truly needed (e.g. inside irqchip code).

Explain Energy Aware Scheduling on big.LITTLE systems.

EAS adds an energy model (power and capacity per performance state per CPU cluster, from the DT or firmware) to the scheduler. When a task wakes, EAS estimates its utilization with PELT (per-entity load tracking) and evaluates candidate CPUs: which placement meets the task's capacity needs with the lowest total energy, considering the frequency each cluster would have to run at. Small tasks stay on little cores; heavy tasks go to big cores. It works with the schedutil governor, which sets frequency from utilization. It is active only when the system is not overutilized; otherwise normal load balancing takes over. Android steers it with uclamp (min/max utilization clamps per cgroup, e.g. boosting top-app) and cpusets.

What is the difference between CFS and EEVDF?

Both aim for fair CPU sharing weighted by nice. CFS picks the task with the smallest vruntime and relied on heuristics (wakeup preemption granularity, sleeper credits) to give interactive tasks low latency. EEVDF (Earliest Eligible Virtual Deadline First, default since 6.6) computes a "lag" per task (how much service it is owed); only tasks with non-negative lag are eligible, and among them it picks the one with the earliest virtual deadline, where the deadline is eligible time plus requested slice divided by weight. A task requesting a shorter slice gets an earlier deadline and so runs sooner, without getting more total CPU. This replaces many CFS heuristics with a cleaner model and gives latency-sensitive tasks a proper knob.

How does the kernel handle memory pressure, from kswapd to OOM kill, and how does Android differ?

Each memory zone has min/low/high watermarks. When free pages drop below low, kswapd wakes and reclaims in the background until high is reached: dropping clean page-cache pages, writing back dirty pages, and swapping anonymous pages (to ZRAM on Android). The LRU lists (active/inactive, or multi-gen LRU in newer kernels) choose victims. If an allocation cannot be satisfied, the allocating task does direct reclaim and compaction, causing latency stalls. If that fails, the OOM killer selects a process by oom_score (memory usage adjusted by oom_score_adj) and kills it. Android avoids reaching that point: LMKD monitors PSI memory stall levels and kills cached/background apps by oom_score_adj set by ActivityManager, keeping the foreground responsive.

What is PSI and how does LMKD use it?

Pressure Stall Information (/proc/pressure/cpu, memory, io) reports the percentage of time over 10, 60 and 300 second windows that some tasks ("some") or all non-idle tasks ("full") were stalled waiting for that resource. Unlike free-memory numbers, it measures the actual impact of pressure. Userspace can register triggers such as "notify me when memory some-stall exceeds 70 ms in a 1 s window". Android's LMKD registers PSI triggers; when they fire, it picks a victim with the highest oom_score_adj above a threshold that depends on the pressure level and kills it, replacing the old in-kernel lowmemorykiller that used free-memory thresholds.

How are ARM64 kernel and user virtual address spaces laid out and translated?

ARM64 has two translation table base registers: TTBR0_EL1 for the lower range (user space, per process, tagged with an ASID) and TTBR1_EL1 for the upper range (kernel, shared). With 4 KB pages and 48-bit virtual addresses, translation uses four levels (9 bits each plus a 12-bit offset); block mappings at level 1 or 2 give 1 GB or 2 MB pages. The kernel has a linear map of all physical RAM, a vmalloc area, a fixmap, module space and so on. On a context switch only TTBR0 and the ASID change. Security features include PAN (kernel cannot access user memory without explicit uaccess), PXN/UXN (execute-never), KASLR, and on some systems KPTI (unmapping the kernel while in user space).

How does the kernel copy data to and from user space safely?

copy_from_user/copy_to_user (and get_user/put_user) first check that the pointer range lies in user space (access_ok), then perform the copy with privileged access temporarily enabled (on ARM64, clearing PAN with uaccess_enable). If the user page is not present, a normal page fault occurs and may sleep to bring it in, which is why these functions can sleep. If the address is invalid, the fault handler finds the faulting instruction in the exception table and jumps to a fixup that makes the function return the number of bytes not copied, so the driver returns -EFAULT instead of oopsing. Checks against kernel objects (hardened usercopy) catch overflows.

How would you implement mmap in a driver to expose a buffer to user space?

Implement file_operations.mmap(struct file *f, struct vm_area_struct *vma). For device registers or physically contiguous memory, validate the requested size and offset, set appropriate page protection (e.g. pgprot_noncached or pgprot_writecombine for MMIO), and call remap_pfn_range() or io_remap_pfn_range(). For DMA memory from dma_alloc_coherent, use dma_mmap_coherent(). For scattered pages, use vm_insert_page or a vm_operations_struct with a fault handler that supplies pages on demand. Never allow mapping beyond the buffer, and keep the buffer alive until the VMA is closed (reference counting in vm_ops->open/close). For sharing between devices, exporting a DMA-BUF is usually better.

What is the GKI KMI, and how do vendor modules stay compatible?

The Kernel Module Interface is the set of exported kernel symbols plus the layouts of the data types they use that GKI vendor modules may rely on. For each GKI branch (e.g. android14-6.1), Google freezes the KMI after a stabilization period: symbols listed in the vendor symbol lists are guaranteed, and ABI tooling compares every build against the frozen ABI representation to reject incompatible changes. Vendors compile modules against the GKI source and headers; at load time, symbol CRCs (CONFIG_MODVERSIONS) must match. When vendors need a new symbol or a hook into core code, they upstream it to the Android Common Kernel (adding symbols or vendor hooks based on tracepoints) rather than patching the core kernel locally.

What are Android vendor hooks and why do they exist?

With GKI, vendors cannot modify core kernel code, but they sometimes need to change behavior in scheduler, memory or other core paths (e.g. custom task placement or thermal decisions). Vendor hooks are special tracepoints (DECLARE_HOOK / DECLARE_RESTRICTED_HOOK) added to the Android Common Kernel at agreed points; vendor modules register handlers that run there. Restricted hooks can sleep and attach once, normal hooks behave like regular tracepoints. They provide a controlled extension point that keeps the core kernel generic while preserving a stable KMI. They must be proposed and accepted into ACK.

How does the GIC deliver an interrupt to Linux, and what is an irq domain?

A peripheral asserts a line into the GIC distributor (SPI for shared peripherals, PPI for per-CPU sources like the arch timer, SGI for inter-processor interrupts, LPI for MSI-style message interrupts via the ITS). The distributor routes it to a target CPU's redistributor/CPU interface based on priority and affinity. The CPU takes an IRQ exception, reads the interrupt ID from the acknowledge register, and Linux's generic IRQ layer maps the hardware ID to a Linux IRQ number through an irq domain, then calls the flow handler (e.g. handle_fasteoi_irq) and the driver's handler, and finally signals end of interrupt. irq domains also chain controllers: a GPIO controller or PMIC can be an irqchip whose domain translates its pins into Linux IRQs cascaded under a GIC interrupt, which is what interrupt-parent in the DT expresses.

How do you find and fix a data race in the kernel?

Symptoms include rare corrupted state, list corruption (list_add corruption warnings), refcount underflows and crashes that move around. Tools: KCSAN detects data races dynamically by watching concurrent unsynchronized accesses; KASAN catches the resulting use-after-free; lockdep verifies locking assumptions (use lockdep_assert_held() in functions that require a lock). Review which contexts touch the data: process, softirq, hard IRQ, other CPUs. Fix by protecting all accesses with the same lock of the right type, converting counters to atomic_t/refcount_t, using RCU for read-mostly structures, or READ_ONCE/WRITE_ONCE for intentionally lockless flags with appropriate barriers.

How do futexes make userspace locks fast?

A futex is a 32-bit integer in user memory plus a kernel wait queue keyed by its address. Acquiring an uncontended pthread_mutex is a single atomic compare-and-swap in user space, with no syscall. Only when the lock is contended does the thread mark it as contended and call futex(FUTEX_WAIT), which sleeps if the value is still as expected (avoiding lost-wakeup races). The unlocking thread, seeing the contended state, calls futex(FUTEX_WAKE). PI futexes (FUTEX_LOCK_PI) store the owner TID so the kernel can apply priority inheritance. Java monitors and ART locks are also built on futexes.

What causes interrupt latency and how do you measure it?

Latency comes from regions where interrupts are disabled (local_irq_save, spin_lock_irqsave), long hard IRQ handlers of other devices, higher-priority interrupts, softirq processing, preemption-disabled sections delaying threaded handlers, CPU idle exit latency from deep C-states, and frequency ramp-up. Measure with the ftrace irqsoff and preemptoff tracers (report the maximum disabled section and its call site), cyclictest for scheduling latency, IRQ and sched events in Perfetto, or a GPIO toggle measured on an oscilloscope for end-to-end latency. Fix by shortening disabled regions, moving work into threaded handlers, adjusting IRQ affinity, raising the priority of the IRQ thread, or limiting deep idle states when latency matters.

How does kernel tracing (ftrace, tracepoints, kprobes) work internally?

With CONFIG_FUNCTION_TRACER, the compiler inserts a call (or NOPs with patchable entries) at the start of every function; at runtime these are NOPs and ftrace patches selected sites to jump to the tracer, so overhead is near zero when disabled. The function_graph tracer also hooks returns to record durations. Tracepoints are static markers placed in code (TRACE_EVENT) that call registered probes via a static key when enabled, recording structured events into per-CPU lockless ring buffers exposed in tracefs. Kprobes dynamically insert a breakpoint at almost any instruction and run a handler; fprobes/fentry and eBPF programs can attach to these points too. Perfetto reads the same ftrace ring buffers plus userspace atrace markers.

How would you analyze a ramdump from a hung device?

Load the dump with the matching vmlinux (and module symbols) into crash or Trace32. First, read the kernel log buffer (log) for the last messages, watchdog reports or lockup warnings. Check each CPU's backtrace (bt -a): is one spinning on a lock, stuck with IRQs disabled, or in a tight loop? List tasks (ps) and focus on those in D state (foreach UN bt) to see what they wait on. For a mutex, inspect its owner field (struct mutex) and follow the owner's stack to find a lock cycle. Examine relevant driver structures with struct, memory usage with kmem -i, and runqueues with runq. The goal is to identify the stuck resource and which task or CPU holds it.

What is the difference between SLAB, SLUB and SLOB, and what does SLUB debugging catch?

They were three implementations of the kernel's object allocator. SLAB was the original, with complex per-CPU and per-node queues; SLOB was a minimal allocator for tiny systems; SLUB simplified the design with per-CPU slabs and little metadata and became the default. SLOB and SLAB have since been removed, so SLUB is the only one in current kernels. With slub_debug (options like F sanity checks, Z red zones, P poisoning, U user tracking), SLUB detects buffer overruns into red zones, use-after-free via poison patterns, and double frees, and records allocation/free call sites; KASAN and KFENCE provide more precise detection.

What changed from ION to DMA-BUF heaps?

ION was an Android-specific allocator (/dev/ion) that allocated buffers from heaps (system, carveout, CMA) and exported them as DMA-BUF file descriptors. Its single ioctl interface with heap IDs and flags was hard to keep ABI-stable and vendors added many incompatible changes. Mainline Linux adopted DMA-BUF heaps instead: each heap is its own character device under /dev/dma_heap/ (e.g. system, linux,cma, vendor heaps as modules), with a simple allocation ioctl returning a DMA-BUF fd. ION was removed from mainline in Linux 5.11 and deprecated in Android's GKI kernels; Android 12+ devices use DMA-BUF heaps through libdmabufheap. Cache maintenance and sharing semantics come from the standard DMA-BUF framework.

How is the VINTF compatibility check performed, and what breaks it?

The vendor provides a device manifest listing HALs it implements (name, interface, version, instance) plus kernel requirements; the framework provides a compatibility matrix listing HALs and versions it needs and kernel config requirements. The build system checks that the vendor manifest satisfies the framework matrix and vice versa (the device matrix vs framework manifest), and VintfObject checks again at OTA time and boot. Failures happen when the framework requires a newer HAL version than the vendor provides, a required HAL is missing, a HAL is declared but not registered at runtime, or kernel configs/versions do not meet the matrix. Symptoms include OTA rejection, VTS failures, or a service failing to find a HAL.

Scenario & debugging

The device shows the boot logo but never reaches the launcher. How do you debug it?

The logo means the bootloader worked, so the problem is in the kernel or userspace. First get the kernel log: UART console, or /sys/fs/pstore/console-ramoops-0 after a reboot. Look for a panic, a driver hanging in probe, or a root mount failure. If the kernel reached init, check logcat -b all (via adb if available) for a critical service crash-looping (init: Service ... restarting), avc: denied blocking a daemon, a failed mount of /data or /vendor, or system_server/zygote crashing. Then bisect: try the other A/B slot or a known-good build, and diff DT, defconfig, modules and rc changes. For the full stage-by-stage method, see the Android boot page.

You get "Unable to handle kernel NULL pointer dereference" at boot. Walk through your analysis.
  1. Read the header: faulting virtual address (a small value like 0x10 means NULL plus a struct field offset), the CPU, the task (Comm), taint flags.
  2. Look at pc (e.g. foo_probe+0x48/0x1a0 [foo]) and lr, and the call trace.
  3. Symbolize: addr2line -e vmlinux for built-in code, or gdb foo.ko then list *(foo_probe+0x48) for modules; scripts/decode_stacktrace.sh for the whole trace.
  4. Identify which pointer was NULL and why: missing error check (e.g. of_get_property returned NULL), a dependency not yet probed, drvdata not set before the IRQ fired, or an early IRQ calling into uninitialized state.
  5. Fix and harden: check returns, request the IRQ only after initialization, use -EPROBE_DEFER. Reproduce with KASAN if the pointer might be a freed object.
An I2C sensor is not detected on a new board. How do you debug it?
  1. Probe: dmesg | grep for the driver; check /sys/bus/i2c/devices/ for the device and whether it is bound to a driver. No device means the DT node is missing, disabled, or under the wrong bus; device but no driver means a compatible mismatch or module not loaded; deferred means a missing dependency (/sys/kernel/debug/devices_deferred).
  2. Bus: i2cdetect -y <bus> to see if the address ACKs. No ACK means wrong address, bus, wiring, or the chip is unpowered or held in reset.
  3. Power: regulator_summary to confirm the vdd-supply is enabled at the right voltage; check the reset GPIO polarity.
  4. Pins and clocks: pinctrl state for SDA/SCL, controller clock in clk_summary; a logic analyzer or scope on the bus if still unclear.
  5. Interrupt: correct GPIO, trigger type and pull; watch /proc/interrupts.
Idle battery drain is high: the device does not seem to enter deep sleep. What do you check?
  1. Is it suspending? dmesg for "PM: suspend entry/exit"; /sys/kernel/debug/suspend_stats for success and failure counts.
  2. Suspend aborted? suspend_stats shows last_failed_dev and the failing step; a driver returning an error from suspend() blocks every attempt.
  3. Who holds it awake? /sys/kernel/debug/wakeup_sources: sort by active_since/total_time; dumpsys power for app partial wakelocks.
  4. Who wakes it? /proc/interrupts deltas across a sleep period, "wakeup IRQ" messages, Battery Historian or a Perfetto trace. Suspects: chatty sensor or modem IRQ, a GPIO configured as wake source with a floating line, an RTC alarm set too often, an app wakelock.
  5. Also check residency of CPU idle states and whether peripherals are runtime-suspended; a rail left on can drain even while suspended.
A driver prints "BUG: scheduling while atomic". What does it mean and how do you fix it?

Code called something that can sleep while in atomic context: holding a spinlock, with preemption or interrupts disabled, or in a hard IRQ, softirq or tasklet. Typical offenders are mutex_lock, kmalloc(GFP_KERNEL), msleep, copy_to_user, wait_event, or a blocking I2C/SPI transfer inside an IRQ handler. The splat's stack trace shows the sleeping call and the preempt count. Fix by moving the work to a threaded IRQ or workqueue, using GFP_ATOMIC, releasing the spinlock before the sleeping call, or using a mutex when all users are in process context. Prevent recurrence by running debug builds with CONFIG_DEBUG_ATOMIC_SLEEP and lockdep.

The device reboots randomly under load. How do you find the cause?
  1. Start from the recorded reboot reason: ro.boot.bootreason, bootloader logs, PMIC reset reason registers, pstore console and panic logs. Never guess.
  2. Kernel panic: decode the oops from pstore.
  3. Watchdog bite: a CPU stopped petting the watchdog, typically stuck in a spinlock, an IRQs-disabled loop, or a deadlock. Collect a ramdump and inspect per-CPU stacks and lock owners.
  4. Thermal: check thermal zone logs and trip points in /sys/class/thermal; a critical trip triggers shutdown.
  5. Power: voltage droop or PMIC over-current under load (brown-out); correlate with CPU/GPU frequency peaks and battery state.
  6. Reproduce with a stress test and bisect recent kernel, DT or firmware changes.
Free memory shrinks over hours and apps get killed. How do you find a leak?

First decide kernel vs userspace using /proc/meminfo over time. Growing SUnreclaim (slab), KernelStack, VmallocUsed or DMA-BUF totals point to the kernel; growing app PSS (from dumpsys meminfo) points to userspace. For kernel leaks, use slabtop or /proc/slabinfo to find the growing cache, enable kmemleak to list unreferenced allocations with their stacks, and inspect /sys/kernel/debug/dma_buf/bufinfo for leaked buffers (often an fd not closed in userspace). For userspace leaks, use heapprofd in Perfetto, malloc debug, or HWASan, and showmap for mapping growth. Confirm the fix by trending MemAvailable in a long run.

The UART console is completely silent on a newly bring-up board. What do you do?

Determine whether any stage prints. If even the bootloader is silent, check the bootloader's UART configuration, pinmux, clock, baud rate, and whether DDR and clocks come up at all (JTAG can tell you where the CPU is). If the bootloader prints but the kernel does not, add earlycon to the kernel command line with the correct UART type and base address, confirm stdout-path in /chosen and the UART DT node (address, clocks, pinctrl), and check that the serial driver is built in, not a module. If the kernel crashes before the console, early printk or reading the log buffer from memory via JTAG helps. As a last resort, toggle a GPIO or LED at known points.

Boot takes 40 seconds and the target is under 20. How do you approach it?
  1. Measure each stage: bootloader timestamps, kernel dmesg timestamps up to "Freeing unused kernel memory" and init start, init service start times, sys.boot_completed, and a Perfetto boot trace or bootchart.
  2. Kernel: boot with initcall_debug to find slow initcalls; move non-critical drivers to modules loaded later; enable asynchronous probe; remove long msleeps and firmware load waits in probe; trim unused config.
  3. Userspace: parallelize services, start non-critical HALs lazily, avoid serial wait_for_prop chains, reduce system_server and Zygote preload work, and check dexopt state.
  4. Storage: tune read-ahead, use erofs compression, check that dm-verity and fs-verity are not bottlenecks.
  5. Always attack the biggest measured contributor and re-measure after each change.
A process is stuck and cannot be killed with kill -9. What is happening?

It is almost certainly in uninterruptible sleep (state D), waiting inside the kernel for something like I/O completion, a mutex, or a driver event that uses wait_event (not the interruptible or killable variant). Signals, including SIGKILL, are handled only when the task returns toward user space, so it cannot die until the wait ends. Check cat /proc/<pid>/stack and /proc/<pid>/wchan to see where it waits, or echo w > /proc/sysrq-trigger to dump all blocked tasks. The root cause is typically a hung storage device, a network filesystem, or a driver that never completes a request. Drivers should use wait_event_killable or timeouts where possible. A zombie (state Z) also cannot be killed, but that is fixed by its parent reaping it.

UI jank appears only on certain devices. How would you check whether the kernel scheduler is involved?

Capture a Perfetto trace with sched, freq, idle, irq and binder events along with app atrace markers. For the UI thread and RenderThread in the janky frames, check the thread state: Runnable but not Running means CPU contention or poor placement (look at which CPU and what else ran there); Running on a little core at low frequency means utilization estimation, uclamp or cpuset issues; Uninterruptible sleep means I/O or lock waits; Sleeping on Binder means the slowness is in another process. Also look for long IRQ or softirq bursts on the same CPU and thermal throttling (frequency caps). Fixes include correcting cgroup/cpuset assignment, uclamp boosts for top-app, IRQ affinity changes, or removing priority inversions.

After a kernel update, a vendor module fails to load with "disagrees about version of symbol". What happened?

The module was built against a kernel whose exported symbol CRCs (from CONFIG_MODVERSIONS) differ from the running kernel, meaning a function signature or a data structure used by that symbol changed, or the module was built against different headers or config. With GKI, this indicates either the module was not rebuilt against the new GKI release, or a KMI break (which the ABI tooling should prevent on a frozen branch). Check modinfo for vermagic, compare with uname -r, rebuild the module against the exact GKI kernel source and config, and ensure the symbols it uses are in the vendor symbol list. Never force-load with mismatched CRCs; that risks memory corruption.

A HAL cannot open /dev/foo even though the driver probed. What do you check?
  1. Does the node exist? ls -lZ /dev/foo. If not, check that the driver registered the char/misc device and that ueventd processed it.
  2. Permissions: owner, group, mode from ueventd.rc; is the HAL's user or group allowed?
  3. SELinux: the node's label (from file_contexts) and avc: denied messages in dmesg or logcat for the HAL's domain; add correct type and allow rules in vendor policy (not permissive mode).
  4. The driver's open() may itself return an error (e.g. runtime PM resume failing, device busy); check dmesg and the errno the HAL logs.
A peripheral stops working after suspend and resume. How do you debug it?

Suspend may have powered off the device's regulator or power domain, losing its register state, while the driver's resume callback does not restore it. Check whether the driver implements .suspend/.resume (and runtime PM callbacks) and re-initializes the hardware, restores pinctrl state (sleep vs default), re-enables clocks and regulators in the right order, and re-arms interrupts. Use dmesg with pm_debug_messages or initcall_debug to see resume callback order and errors, and compare register dumps before and after. Dependency ordering matters: the device must resume after its parent bus and supplies (device links help). Also verify the firmware or co-processor state if the peripheral runs its own firmware.

You see intermittent data corruption in buffers received from a DMA-capable peripheral. What could be wrong?
  • Missing cache maintenance: the CPU reads stale cache lines because the buffer was not invalidated (dma_unmap or dma_sync_single_for_cpu not called) on a non-coherent system.
  • The buffer shares cache lines with other data (not cache-line aligned), so CPU writes to neighbors write back stale data over what the device wrote.
  • Using a stack or vmalloc buffer for DMA.
  • The CPU touches the buffer while the device still owns it (ownership rule violated), or the buffer is freed or reused too early.
  • Missing dma_wmb() before ringing a doorbell, so the device reads a descriptor before it is fully written.
  • IOMMU mapping errors (check for SMMU faults).

Debug with CONFIG_DMA_API_DEBUG, which warns on API misuse, and add checksums at producer and consumer.

A system hangs with no panic and the watchdog eventually resets it. How do you find the stuck code?

Enable the soft-lockup and hard-lockup detectors and the hung-task detector with panic options so the kernel dumps stacks and panics before the hardware watchdog fires, producing a ramdump. If the hang is a hard lockup (a CPU with interrupts disabled), the arm64 pseudo-NMI (or the SoC's watchdog pre-timeout / FIQ) can capture that CPU's stack. With the ramdump, check all CPU stacks: a CPU spinning in queued_spin_lock_slowpath points to a lock whose owner you then find; a loop in a driver with IRQs off points to a missing timeout while polling hardware. sysrq (l for all CPUs' backtraces) helps if the console is still responsive. Lockdep in a debug build can reveal the ordering problem before it reproduces.

How would you bring up a new board that uses an existing SoC?
  1. Start from the vendor reference board: its bootloader config, DT and defconfig are ground truth; diff the schematics to list what changed.
  2. Power and clocks first: PMIC rails, reset sequencing, crystals. Get a UART console working.
  3. DDR configuration and training in the bootloader; then boot the kernel with a new board DT derived from the reference.
  4. Storage (UFS/eMMC) so the root filesystem mounts; then USB for adb/fastboot.
  5. Peripherals one at a time: display, touch, audio, sensors, cameras, modem, Wi-Fi/BT. For each: DT node, pinctrl, supplies, probe, validation via sysfs or test tools.
  6. Power and thermal: suspend/resume, idle and active current, thermal zones; then performance tuning.
  7. Keep a working ramdump path early so any crash can be analyzed.
How would you ramp up quickly on an unfamiliar SoC or kernel codebase?

Treat the vendor reference board, its BSP and defconfig as the source of truth, and read the SoC technical reference manual for the blocks you will touch (clocks, interrupt controller, the peripherals in question). Set up tooling first: a UART console, ramdump collection, symbol files and a fast build-flash loop, so every experiment produces evidence. Learn the code top-down from the DT: follow a node's compatible to its driver, then to the subsystem it registers with. Use dmesg, deferred-probe lists, ftrace function_graph on the driver, and diffs against the reference DT to understand behavior. Approach every bug the same way: which stage, what evidence, smallest reproducer, fix, regression test.

An RT audio thread occasionally misses its deadline. How do you investigate?

Trace with Perfetto or ftrace (sched_switch, sched_wakeup, irq, softirq events) around a glitch. Measure wakeup latency: time from wakeup to running. If it is Runnable for long, find what occupies the CPU: a higher-priority RT task, a long IRQ or softirq burst, or an IRQs-off or preemption-off region (use the irqsoff/preemptoff tracers). If it is blocked, look for a lock held by a lower-priority thread (priority inversion; use PI mutexes), a page fault on memory that was not locked (mlock, prefaulting), or a Binder call. Also check CPU frequency ramping and deep idle exit latency on its core, and whether RT throttling kicked in. Fixes include pinning to a quiet core, raising IRQ thread priorities appropriately, avoiding allocations and locks in the audio callback, and PI locking.

A device sits in devices_deferred forever and never probes. How do you find the missing dependency?

cat /sys/kernel/debug/devices_deferred lists the device and often the reason string (missing clock, regulator, GPIO, PHY). Cross-check the DT: is the provider's node status = "okay", is the phandle correct, is the provider itself a module that never loaded? fw_devlink should order probes from DT phandles; if it does not, look for a missing clocks, *-supply or resets property. Also check that the provider driver is in the vendor module list (GKI) and that it did not fail probe earlier. Fix the provider or the DT; do not paper over it with a long msleep in the consumer.

You must add support for a new sensor with an existing Linux driver. What are the steps?
  1. Check the upstream or vendor kernel for a driver matching the part and its DT binding documentation (Documentation/devicetree/bindings/).
  2. Enable the driver in the vendor defconfig as a module (GKI) and add it to the vendor module list so it is packaged in vendor_dlkm or vendor_boot.
  3. Add the DT node under the correct I2C/SPI bus with compatible, reg, interrupt, supplies, reset GPIO and pinctrl as the binding specifies.
  4. Verify probe in dmesg, the IIO or input device in sysfs, and raw readings.
  5. Add ueventd.rc permissions and SELinux labels for any device nodes, then wire up the sensor HAL.
  6. Validate suspend/resume, interrupt wake behavior and power consumption.