Binder IPC & AIDL
Binder is the IPC mechanism that glues Android together: nearly every call between an app and the framework, and between the framework and vendor HALs, is a Binder transaction. AIDL is the language you use to define those interfaces; the build turns it into a client Proxy and a server Stub that move Parcels through the Binder kernel driver.
- Define a contract in
.aidl; the build generates a Proxy (client side) and a Stub (server side). - A call writes arguments into a Parcel, the Proxy calls
transact(), the driver copies the data once into the server's mmap'd buffer and wakes a thread in its Binder thread pool, andStub.onTransact()runs the real method. - servicemanager is the name registry (handle 0). The driver stamps each call with the caller's UID/PID for permission checks, and linkToDeath tells clients when a server dies.
onewaycalls do not wait for a reply. The Radio HAL is fully async: requests carry a serial; results come back on separate Response and Indication interfaces.- Stable AIDL replaced HIDL for HALs; versions are frozen and can only be extended. Keep payloads small: the per-process buffer is about 1 MB, shared by all in-flight calls, or you get
TransactionTooLargeException.
Why Android needs IPC, and why Binder
Every Android app runs in its own process with its own Linux UID and a private address space. System services run in other processes (system_server, the phone process, vendor HAL daemons). So a simple call such as "give me the current network state" must cross a process boundary. Android could have used standard Linux IPC, but it built Binder because the framework needed things plain IPC does not give.
| Mechanism | Copies per message | Caller identity | Object model / RPC | Lifetime tracking | Comment |
|---|---|---|---|---|---|
| Pipes / FIFOs | 2 (user → kernel → user) | No | Byte stream only | No | One-way, simple, parent/child |
| Unix domain sockets | 2 | Yes, via SO_PEERCRED (per connection) | Byte stream; you build your own protocol | Connection close only | Used by Zygote, InputChannel, logd, netd |
| System V / POSIX message queues | 2 | No | Messages | No | Rarely used on Android |
| Shared memory (ashmem, memfd, dma-buf) | 0 | No | None; needs separate synchronization | No | Great for large buffers, used together with Binder |
| Binder | 1 | Yes, per call, set by the kernel | Object-oriented RPC with method codes | Reference counting + death notifications | Synchronous by default, oneway available, built-in thread pool |
Performance
One copy per transaction instead of two, and no connection setup per client.
Security
The kernel attaches the caller's UID and PID to every transaction; they cannot be forged. SELinux also mediates which domains may call which.
Object model
Binder references are capabilities: if you hold a handle to an object you may call it, and handles can be passed to other processes inside Parcels.
Lifetime
The driver reference-counts objects across processes and delivers death notifications when a hosting process dies.
Pipes and sockets are like shouting across a street or talking on a raw phone line: it works, but you do not know for sure who is speaking and you must invent your own conversation rules. Binder is like a company's internal mail system with a trusted mail room: every envelope is stamped by the mail room with the sender's employee ID (UID/PID), the mail room copies the letter straight into the recipient's inbox tray (one copy), and if a department is closed down, everyone who had its address is told (death notification). Holding an address card (a Binder handle) is what lets you send it mail at all.
The Binder driver and its three domains
Binder is implemented as a kernel driver (drivers/android/binder.c). User space talks to it with open(), mmap() and ioctl() on a device node. There is no read()/write(); everything goes through the BINDER_WRITE_READ ioctl, which both sends commands and receives work in one system call.
Device nodes (Binder domains)
| Node | Context manager | Who talks over it | IDL |
|---|---|---|---|
/dev/binder | servicemanager | Apps and framework; framework and stable-AIDL vendor HALs | AIDL |
/dev/hwbinder | hwservicemanager | Framework and legacy HIDL HALs | HIDL |
/dev/vndbinder | vndservicemanager | Vendor process to vendor process only | AIDL |
Each domain is a separate namespace with its own context manager, so a vendor service registered on vndbinder is invisible to system processes. On newer kernels the nodes are provided by binderfs (mounted at /dev/binderfs, with the classic paths as symlinks). SELinux policy controls which domains may open which node.
Key driver objects
- binder_proc: one per process that opened the device. Holds its threads, nodes, references and the mmap'd buffer.
- binder_node: represents a Binder object (a service) living in its owning process.
- binder_ref: a reference from another process to a node. User space sees it as an integer handle, which is only meaningful inside that process. Handle 0 always means the context manager.
- binder_thread: a user thread that has entered the driver; tracks its transaction stack for replies and recursion.
- binder_buffer: a chunk of the receiver's mmap'd area holding one transaction's data.
Important ioctls and protocol commands
| Name | Direction | Meaning |
|---|---|---|
BINDER_WRITE_READ | ioctl | Send a buffer of BC_* commands and receive BR_* returns |
BINDER_SET_MAX_THREADS | ioctl | How many extra pool threads the driver may ask this process to spawn |
BINDER_SET_CONTEXT_MGR | ioctl | Used once by servicemanager to become handle 0 |
BC_TRANSACTION / BR_TRANSACTION | client → driver / driver → server | A call and its delivery |
BC_REPLY / BR_REPLY | server → driver / driver → client | The reply and its delivery |
BR_TRANSACTION_COMPLETE | driver → sender | Driver accepted the transaction (for oneway, this is all the sender gets) |
BC_REQUEST_DEATH_NOTIFICATION / BR_DEAD_BINDER | both | Register for, and receive, death notifications |
BR_SPAWN_LOOPER / BC_REGISTER_LOOPER | driver → process / thread → driver | Ask the process to start another pool thread; new thread registers |
BR_DEAD_REPLY / BR_FAILED_REPLY | driver → client | Target died, or delivery failed (for example no buffer space) |
BR_FROZEN_REPLY | driver → client | Target process is frozen (cached-apps freezer). Sync calls fail immediately; the target is not unfrozen |
Client process Kernel (binder driver) Server process ┌──────────────┐ ioctl ┌───────────────────────┐ wakes ┌──────────────┐ │ Proxy │ BC_TRANSACTION │ find target node │ BR_TRANSACTION│ Binder thread│ │ handle = 7 │──────────────▶ │ ref(7) → node(svc) │─────────────▶│ Stub.onTrans │ │ │ │ copy data into server │ │ act() │ │ blocked in │ BR_REPLY │ mmap buffer, stamp │ BC_REPLY │ │ │ ioctl │◀────────────── │ uid/pid, queue work │◀─────────────│ write reply │ └──────────────┘ └───────────────────────┘ └──────────────┘
A hotel switchboard. Each guest room is a process; the switchboard operator is the Binder driver. Guests never dial each other's real room numbers; they have speed-dial numbers (handles) that only make sense on their own phone, and the operator maps them to the right room (node). The front desk (handle 0) is servicemanager. There are three separate switchboards in the building: guests and management (binder), management and old contractors (hwbinder), and contractors among themselves (vndbinder), so calls never leak between them.
BINDER_WRITE_READ and the BC/BR command pairs signals depth.One copy: how mmap makes Binder efficient
When a process opens Binder, libbinder's ProcessState calls mmap() on the device to reserve a receive area (by default 1 MB minus two pages, about 1016 KB). The driver allocates physical pages for this area on demand and maps them twice: into the kernel's address space and, read-only, into the process's address space.
- Sender writes The Proxy builds a Parcel in the sender's own memory.
- One copy In the
BC_TRANSACTIONhandling, the driver allocates abinder_bufferinside the receiver's mmap area and does a singlecopy_from_user()from the sender's Parcel into it. - Receiver reads in place Because those same pages are mapped into the receiver, it reads the data directly, with no second copy to user space.
- Free When the receiver's Parcel is destroyed, libbinder sends
BC_FREE_BUFFERso the driver can reuse that space.
Classic IPC (socket/pipe): sender buf ──copy──▶ kernel buf ──copy──▶ receiver buf (2 copies)
Binder: sender Parcel ──copy_from_user──▶ physical pages
│ mapped in kernel
└ mapped (read-only) in receiver
receiver reads here (1 copy)
Shared memory (ashmem/memfd/dma-buf): pass an fd over Binder, both sides map it (0 copies)
Why not zero copy for everything?
- Zero copy needs the sender to write directly into shared pages and both sides to synchronize; unsafe for general RPC.
- One copy isolates the sender: after the copy, the sender can reuse its memory freely.
How large data is handled
- Send a file descriptor (ashmem, memfd, dma-buf) in the Parcel; the driver installs a duplicate fd in the receiver.
- Examples:
CursorWindowfor provider queries,SharedMemory,HardwareBufferand graphics buffers, FMQ in HALs.
Posting a letter normally means you hand it to the post office (copy one) and the postman drops a photocopy in the recipient's box (copy two). With Binder, the recipient has installed a glass mailbox that the post office can reach into from its side: the clerk writes your letter once, directly into that mailbox, and the recipient reads it through the glass. For a huge parcel, you do not copy it at all; you hand over a key to a shared storage unit (a file descriptor to shared memory).
The Binder thread pool
Incoming transactions are executed on Binder threads in the server process, not on its main thread. libbinder manages these threads.
ProcessState::startThreadPool()starts the first ("main") Binder thread; Java processes do this automatically when they start (Zygote children call it inonZygoteInit).BINDER_SET_MAX_THREADStells the driver how many additional threads it may request. The default is 15, so a normal app has up to 16 Binder threads. system_server raises its maximum to 31. Native services callProcessState::self()->setThreadPoolMaxThreadCount(n).- When a transaction arrives and no thread is idle, the driver returns
BR_SPAWN_LOOPERand libbinder creates a new thread (namedBinder:<pid>_N) until the limit is reached. - If all threads are busy and the limit is reached, new transactions wait in the process's queue inside the driver. Callers are blocked meanwhile: this is thread-pool exhaustion.
- Native daemons without a pool can instead call
IPCThreadState::joinThreadPool()on their own thread, or poll the Binder fd from a Looper (setupPolling/handlePolledCommands).
Recursion and priority
- Recursive calls: if process A calls B synchronously and B calls back into A while handling it, the driver routes the callback to the very thread in A that is waiting, instead of a new pool thread. This prevents simple self-deadlocks and keeps thread-local context consistent.
- Priority inheritance: the driver passes the caller's scheduling priority (nice value, and real-time policy where the node allows it) to the server thread for the duration of the call, so a high-priority caller is not starved by a low-priority server thread.
- oneway ordering: oneway transactions to the same Binder object are queued and delivered one at a time, in order, so a single object never runs two oneway calls concurrently.
// Minimal native service main()
int main() {
ProcessState::self()->setThreadPoolMaxThreadCount(4);
sp<MyService> svc = sp<MyService>::make();
defaultServiceManager()->addService(String16("my.service"), svc);
ProcessState::self()->startThreadPool();
IPCThreadState::self()->joinThreadPool(); // main thread also serves calls
return 0;
}
A call centre. Incoming calls (transactions) are answered by agents (Binder threads), not by the manager (the main thread). When every agent is busy, the switchboard hires another temp (BR_SPAWN_LOOPER) up to a fixed headcount (max threads). If all agents are stuck on long calls, new callers sit on hold, and everyone who called feels the service is frozen. A VIP caller's priority is passed to the agent who takes the call (priority inheritance).
From .aidl to generated Stub and Proxy
AIDL (Android Interface Definition Language) describes the methods a service offers. The aidl compiler generates all marshalling code so that a remote call looks like a local method call.
// IFoo.aidl
package com.example;
interface IFoo {
int add(in int a, in int b);
void getValues(out int[] values); // server fills the array
oneway void ping(); // fire-and-forget
}
| Generated artifact | Used by | Role |
|---|---|---|
IFoo (interface extending IInterface) | Both | Method signatures both sides agree on |
IFoo.Stub (extends Binder) | Server | Abstract base: implement the methods; onTransact() unpacks and dispatches |
IFoo.Stub.Proxy | Client | Wraps a remote IBinder; each method packs a Parcel and calls transact() |
DESCRIPTOR | Both | Interface name string ("com.example.IFoo") written as the interface token |
TRANSACTION_add etc. | Both | Method codes: FIRST_CALL_TRANSACTION + 0, + 1, ... in declaration order |
IFoo.Stub.asInterface(IBinder) | Client | Returns the local object if in the same process, otherwise a new Proxy |
Server implementation
public class FooService extends IFoo.Stub {
@Override public int add(int a, int b) {
return a + b; // runs on a Binder thread
}
@Override public void getValues(int[] v) {
v[0] = 42;
}
@Override public void ping() { }
}
// Publish (system code):
ServiceManager.addService("foo", new FooService());
// Or return it from Service.onBind() for bound servicesClient usage
IBinder b = ServiceManager.getService("foo");
IFoo foo = IFoo.Stub.asInterface(b);
try {
int sum = foo.add(2, 3); // looks local, is IPC
} catch (RemoteException e) {
// server died or transaction failed
}What the generated code looks like (simplified)
// Proxy side
public int add(int a, int b) throws RemoteException {
Parcel data = Parcel.obtain();
Parcel reply = Parcel.obtain();
try {
data.writeInterfaceToken(DESCRIPTOR);
data.writeInt(a);
data.writeInt(b);
mRemote.transact(TRANSACTION_add, data, reply, 0);
reply.readException(); // rethrows server-side exceptions
return reply.readInt();
} finally {
reply.recycle();
data.recycle();
}
}
// Stub side
@Override
protected boolean onTransact(int code, Parcel data, Parcel reply, int flags)
throws RemoteException {
switch (code) {
case TRANSACTION_add: {
data.enforceInterface(DESCRIPTOR);
int a = data.readInt();
int b = data.readInt();
int result = this.add(a, b);
reply.writeNoException();
reply.writeInt(result);
return true;
}
// ...
}
return super.onTransact(code, data, reply, flags);
}
// asInterface
public static IFoo asInterface(IBinder obj) {
if (obj == null) return null;
IInterface local = obj.queryLocalInterface(DESCRIPTOR);
if (local instanceof IFoo) return (IFoo) local; // same process: direct call
return new IFoo.Stub.Proxy(obj); // remote: wrap in Proxy
}
Types and direction tags
- Supported types: primitives,
String,CharSequence,List/Mapof supported types, arrays,Parcelabletypes, other AIDL interfaces (passed as Binder objects),ParcelFileDescriptor, and in structured AIDL, parcelables, enums and unions defined in.aidlitself. - Direction for non-primitive parameters:
in(client to server, default for primitives),out(server fills it and it is copied back),inout(both ways, costs twice). Useinunless you really need the others. - Annotations:
@nullable,@utf8InCpp(C++ strings),@VintfStability(HAL interfaces),@Backingfor enums,@JavaPassthrough. - Backends: Java, NDK (C++ against the stable
libbinder_ndk, required for vendor/APEX code), CPP (platform-internallibbinder), and Rust. - Exceptions: server-side exceptions of certain types (
SecurityException,IllegalArgumentException,ServiceSpecificException, ...) are written into the reply and rethrown in the client; others are logged and the client gets a generic failure. HAL interfaces useServiceSpecificException/ScopedAStatusfor error codes.
A contract drafted by lawyers. The .aidl file is the agreed contract listing each service and its terms. The build hires two translators from it: one sits with the client (the Proxy) and turns each spoken request into a standard form in the right order; the other sits with the server (the Stub) and reads each form back into a request the server understands. Both translators work from the same contract, so the field order always matches. asInterface is the receptionist who notices when the client and server are in the same building and just walks the request over without forms.
asInterface does, and to sketch generated code. Mention the interface token check (enforceInterface), method codes from FIRST_CALL_TRANSACTION, readException for exception propagation, and why reordering methods in an interface breaks compatibility.Parcel: the transaction payload
A Parcel is a flat, sequential container for one transaction's data. Values are written and read in exactly the same order; there are no field names or type tags in the plain stream, which makes it fast but means both sides must agree exactly on the layout (that agreement is what AIDL generates).
- Primitives and strings are written as aligned 4-byte words (strings as length plus UTF-16 in Java, or UTF-8 with
@utf8InCpp). - Binder objects (
writeStrongBinder) are written as aflat_binder_object. The driver translates it: a local object becomes a new reference (handle) in the receiver, a handle to a third process becomes the receiver's own handle to that node. This is how capabilities are passed. - File descriptors (
writeFileDescriptor,ParcelFileDescriptor) are translated by the driver into a new fd in the receiving process pointing at the same open file. This is how shared memory is shared. - Interface token: the first thing in a call. Besides the descriptor string it carries StrictMode policy and a work-source header, so the server can apply the caller's StrictMode and attribute power usage.
- Parcelable: a class that implements
writeToParcel()and aCREATORto read itself back. Structured parcelables declared in AIDL get this code generated. - Bundle: a key/value map that is itself parcelled; used for Intent extras and saved instance state.
public final class Point implements Parcelable {
public int x, y;
@Override public void writeToParcel(Parcel out, int flags) {
out.writeInt(x);
out.writeInt(y); // order matters
}
public static final Creator<Point> CREATOR = new Creator<>() {
public Point createFromParcel(Parcel in) {
Point p = new Point();
p.x = in.readInt();
p.y = in.readInt(); // same order
return p;
}
public Point[] newArray(int n) { return new Point[n]; }
};
@Override public int describeContents() { return 0; }
}
Parcelable
- Hand-written or generated read/write code; no reflection.
- Fast; designed for IPC.
- Not a stable storage format across versions.
Serializable
- Java reflection-based; creates many temporary objects.
- Much slower for IPC.
- Still requires versioning care for persistence.
A Parcel is a shipping box packed in a strict agreed order: first the label (interface token), then item one, item two, and so on. The receiver unpacks in exactly the same order and never looks at labels on individual items. Special items, such as a key to another office (a Binder handle) or a key to a storage unit (a file descriptor), are swapped by the courier (the driver) for keys that work in the receiver's building.
flat_binder_object entries into receiver-local handles/fds, which is what makes Binder a capability system.One call end to end: transact, onTransact and oneway
Here is what happens when a client calls foo.add(2, 3) on a remote service.
- Proxy marshals
Proxy.add()writes the interface token, then 2 and 3 into thedataParcel. - transact
mRemote.transact(TRANSACTION_add, data, reply, 0).BinderProxygoes to nativeIPCThreadState::transact(), which queues aBC_TRANSACTIONand callsioctl(BINDER_WRITE_READ). The client thread now blocks in the kernel waiting for the reply. - Driver delivers The driver resolves the handle to the target node, checks SELinux (
binder_call), allocates a buffer in the server's mmap area, copies the data once, records the caller's euid and pid, and wakes an idle Binder thread in the server (or asks for a new one). - Stub dispatches The server thread returns from its ioctl with
BR_TRANSACTION; libbinder callsBBinder::transact()→ JavaBinder.execTransact()→Stub.onTransact(TRANSACTION_add, ...), which checks the token and reads 2 and 3. - Implementation runs
FooService.add(2, 3)runs on that Binder thread. It may checkBinder.getCallingUid(), take locks, or talk to hardware. - Reply The Stub writes "no exception" and 5 into
reply; libbinder sendsBC_REPLY. The driver copies it into the client's buffer and wakes the waiting client thread withBR_REPLY. - Proxy unmarshals The Proxy reads the exception header and the int, and returns 5. The caller never saw the IPC.
CLIENT (Proxy) SERVER (Stub + Impl)
add(2,3)
│ write token + 2 + 3 into data
│ transact(ADD) ── BC_TRANSACTION ──▶ [driver: copy once, stamp uid/pid]
│ (thread blocks in ioctl) ──▶ BR_TRANSACTION on Binder thread
│ onTransact(ADD)
│ enforceInterface, read 2, 3
│ impl.add(2,3) → 5
│ writeNoException, write 5
│ ◀── BR_REPLY ─── [driver] ◀── BC_REPLY ──┘
│ readException, read 5
return 5
oneway calls
A method (or a whole interface) marked oneway is sent with the FLAG_ONEWAY flag. The caller's transact() returns as soon as the driver has queued the transaction (BR_TRANSACTION_COMPLETE); there is no reply Parcel.
Benefits
- The caller never blocks on a slow or hung server; essential when system_server calls into apps (callbacks,
IApplicationThread). - Calls to the same Binder object are delivered in order, one at a time.
- Natural fit for notifications, listeners and async HAL callbacks.
Trade-offs
- No return value and no exceptions reach the caller; failures are silent unless you design an ack or callback.
- Async transactions may use at most half of the receiver's buffer; a flood fills it and later calls fail.
getCallingPid()is 0 for oneway calls; only the UID is reliable.- If the target is local (same process), a oneway call runs synchronously on the caller's thread.
Useful IBinder methods
| Method | Purpose |
|---|---|
transact(code, data, reply, flags) | Send a transaction (Proxy side) |
onTransact(...) | Handle a transaction (Stub side) |
queryLocalInterface(descriptor) | Get the local object if same process |
pingBinder() / isBinderAlive() | Check if the remote side is alive |
linkToDeath() / unlinkToDeath() | Death notifications |
dump(fd, args) | What dumpsys calls |
getInterfaceDescriptor() | The interface name |
A normal (two-way) call is a phone call where you stay on the line until the other person answers your question. A oneway call is sending a text message: you hit send and carry on with your day, messages to the same person arrive in order, but you never know if they read it unless they text back. That is why the system texts apps (oneway) instead of calling them: a person who never picks up cannot keep the system waiting on the line.
ioctl(BINDER_WRITE_READ), one copy into the target's mmap area, caller identity stamped, pool thread woken, onTransact dispatch, reply returns the same way. Then expect "what changes with oneway?"servicemanager: registration and lookup
A client needs a first handle to talk to anything. Binder solves this bootstrapping problem with a context manager: the process that called BINDER_SET_CONTEXT_MGR becomes reachable by every process at the fixed handle 0. On /dev/binder that is servicemanager, started very early by init.
- Register A service process calls
addService("phone", binder)on handle 0. servicemanager checks SELinux (service_manager addfor that service's label inservice_contexts) and, for@VintfStabilityHALs, that the name is declared in the VINTF manifest. It stores name → reference. - Look up A client calls
getService("phone")/checkService()/waitForService(). servicemanager checksservice_manager findpermission and replies with the Binder object; the driver turns it into a new handle in the client. - Call directly From then on, the client talks to the service directly; servicemanager is not in the path.
- Notifications and lazy start Clients can register for service availability callbacks. For lazy services, a lookup of a declared but not running service makes servicemanager ask init to start it (
ctl.interface_start); the service can later unregister and exit when it has no clients.
Framework services
ServiceManager.addService("phone", binder)at startup (hidden API, system only).- Apps use
Context.getSystemService(TelephonyManager.class), which wrapsgetService()andasInterface()and caches the result. service listordumpsys -lshows everything registered.
Vendor HAL (stable AIDL)
- VINTF manifest declares e.g.
android.hardware.radio.voice.IRadioVoice/slot1. init.rcservice entry withinterface aidl ...starts the daemon (often lazily).- HAL registers with
AServiceManager_addService()orAServiceManager_registerLazyService(). - Client:
AServiceManager_waitForService("...IRadioVoice/slot1"), thenIRadioVoice::fromBinder()(NDK) orIRadioVoice.Stub.asInterface()(Java).
Bound services: the other way to get a Binder
Apps cannot register with servicemanager. Instead, an app exposes a Binder through a bound Service: the client calls bindService(intent, connection, BIND_AUTO_CREATE); AMS starts the service process if needed, calls onBind() to get the Stub, and delivers it to the client's ServiceConnection.onServiceConnected(name, binder). AMS acts as the broker and also links the two processes' priorities (see process priority).
servicemanager is the telephone directory whose own number everyone knows by heart (handle 0). Businesses list themselves by name (addService), after a background check (SELinux, VINTF). Customers look up the name (getService) and then call the business directly; the directory is never on the line again. A lazy service is a business that only opens its shop when someone looks it up. A bound app service is like getting a number through a concierge (AMS) instead of the public directory.
init.rc cannot be started: ctl.interface_start fails and clients wait forever or receive null. Instance names (default vs slot1) must match exactly between manifest, init.rc and the code.waitForService vs getService (the latter may return null or block briefly).Death notifications, caller identity and permissions
linkToDeath and DeadObjectException
Because the driver tracks which process owns each node, it knows when a process dies (its Binder fd is closed). Every other process holding a reference and registered for notification gets BR_DEAD_BINDER, surfaced as DeathRecipient.binderDied().
private final IBinder.DeathRecipient mDeath = new IBinder.DeathRecipient() {
@Override public void binderDied() {
// runs on a Binder thread
mService = null;
mHandler.post(this::reconnect); // re-acquire when the service restarts
}
};
IBinder b = ServiceManager.waitForService("my.service");
b.linkToDeath(mDeath, 0);
mService = IMyService.Stub.asInterface(b);
// later: b.unlinkToDeath(mDeath, 0);
- A call to a dead process throws
DeadObjectException(aRemoteException). Always handle it. - Servers use the same mechanism in reverse to clean up per-client state (registered listeners, wakelocks held on behalf of a client).
RemoteCallbackListdoes this automatically for listener lists. - AMS links to each app's
IApplicationThreadto learn immediately when an app process dies. PowerManagerService links to each wakelock's token so wakelocks are released if the holder dies. - Telephony:
RIL.javalinks a death recipient to each Radio HAL service; onserviceDied()it fails all pending requests withRADIO_NOT_AVAILABLE, clears the request list, resets proxies and re-acquires the HAL when it restarts. HAL death means the vendor daemon died; it is not proof that the modem itself crashed (look for modem subsystem-restart markers at the same time).
Caller identity
For every transaction the driver records the sender's effective UID and PID (and, when configured, SELinux security context). The server reads them with Binder.getCallingUid(), Binder.getCallingPid() (native: IPCThreadState::self()->getCallingUid()). The client cannot fake these values, which is the foundation of Android's permission enforcement.
@Override
public void dial(String number) {
// 1. Who is calling?
mContext.enforceCallingPermission(Manifest.permission.CALL_PHONE, "dial");
int callerUid = Binder.getCallingUid();
// 2. Act as ourselves for downstream calls
final long token = Binder.clearCallingIdentity();
try {
mPhone.dial(number); // downstream checks see our UID, not the app's
} finally {
Binder.restoreCallingIdentity(token);
}
}
clearCallingIdentity()resets the calling identity to the current process so that further checks and outgoing calls are made as the service. ForgettingrestoreCallingIdentity()(always infinally) is a security bug.- For oneway calls the calling PID is 0; rely on UID.
- PIDs can be reused after a process dies; UID is the stable identity. Use
getCallingPid()mainly for logging or combined checks. - SELinux adds mandatory checks before your code runs:
binder_call(may domain A call domain B),binder_transfer(may it pass Binder objects), and service_manageradd/find. Denials appear asavc: denied { call }. - Work source: a caller's attribution (for power blame) can be propagated in the interface token header.
The mail room stamps every envelope with the sender's employee badge number (UID/PID), so the department receiving it can decide whether that person is allowed to ask. When a department handles a request and needs to order supplies from purchasing, it signs the order with its own badge (clearCallingIdentity), then goes back to acting on behalf of the original requester (restoreCallingIdentity). Death notification is HR sending a memo to everyone who had a department's extension when that department is closed.
RemoteCallbackList, and a concrete death-handling example such as RIL's serviceDied recovery.The async HAL pattern: Radio HAL requests, responses and indications
Some HALs talk to slow hardware. The Radio HAL is the classic example: a modem operation such as dialing can take seconds. Instead of blocking a Binder thread for each request, the radio HAL is fully asynchronous and uses three interfaces per domain.
| Interface | Direction | Role |
|---|---|---|
IRadioVoice | Framework → vendor | Requests: dial, hangup, getCurrentCalls, ... each with an int serial. Declared oneway, so it returns at once |
IRadioVoiceResponse | Vendor → framework | Solicited replies: dialResponse(RadioResponseInfo info); info carries serial, type and error |
IRadioVoiceIndication | Vendor → framework | Unsolicited events: callStateChanged, callRing, ... (network, signal, SIM events live in the other domains' indication interfaces) |
Since Android 13 the radio HAL is stable AIDL split by domain: IRadioNetwork, IRadioData, IRadioVoice, IRadioSim, IRadioModem, IRadioMessaging, IRadioIms and IRadioConfig, each with its own Response and Indication interfaces. RIL.java keeps one proxy per domain with a death recipient each.
SETUP (once per HAL connection):
RIL creates Binder objects RadioVoiceResponse + RadioVoiceIndication (framework is SERVER for these)
IRadioVoice.setResponseFunctions(voiceResponse, voiceIndication)
MO DIAL:
[1] RIL ──▶ IRadioVoice.dial(serial=42, dialInfo) Proxy → Stub in vendor radio daemon (oneway)
[2] vendor daemon starts modem dial (QMI / AT) transact already returned
[3] vendor ──▶ IRadioVoiceResponse.dialResponse(info{serial=42, error=NONE}) REVERSE IPC
[4] RIL finds request 42 in its request list, completes it, releases its wakelock count
UNSOLICITED (any time later):
modem event ──▶ vendor ──▶ IRadioVoiceIndication.callStateChanged(type)
──▶ RIL RadioIndication ──▶ RegistrantList ──▶ GsmCdmaCallTracker polls getCurrentCalls
How RIL tracks solicited requests
obtainRequest(RIL_REQUEST_DIAL, resultMsg, workSource)
└─ RILRequest.obtain() // pooled object, assigns unique mSerial
└─ acquireWakeLock(rr, FOR_WAKELOCK) // keep the CPU awake while request is pending
└─ mRequestList.put(serial, rr)
└─ voiceProxy.dial(serial, dialInfo) // oneway HAL call
... modem works ...
IRadioVoiceResponse.dialResponse(info)
└─ processResponse(info) → findAndRemoveRequestFromList(info.serial)
└─ rr.onReceived(result) // sendMessage back to the caller's Handler
└─ decrementWakeLock(rr) // release when count reaches 0
- The serial is a correlation ID: many requests can be outstanding at once, and responses may arrive in any order.
- A wakelock (a partial wakelock, reference-counted manually) is held while any request is pending, with a timeout so a lost response cannot keep the device awake forever.
- Indications may require an acknowledgement (
RadioIndicationType.UNSOLICITED_ACK_EXP), so the vendor can hold its own wakelock until the framework has taken the event; RIL answers withresponseAcknowledgement(). - AIDL runs in both directions: the vendor daemon is the server for requests, the framework is the server for responses and indications.
A dry cleaner with ticket numbers. You drop off a shirt and get ticket 42 (the serial); you do not wait at the counter (oneway). When the shirt is ready, the shop calls you on the number you registered at your first visit (the Response interface, set with setResponseFunctions) and quotes ticket 42 so you know which order it is. The shop also calls you unprompted about things like "we are closed tomorrow" (Indications). If the shop burns down, you are told (death notification) and every open ticket is cancelled (RADIO_NOT_AVAILABLE).
Telephony: several Binder boundaries on one call
One outgoing voice call crosses multiple processes, each hop a Binder interface following the same Stub/Proxy pattern.
| # | Boundary | Interface | Client → Server |
|---|---|---|---|
| 1 | App → system | ITelecomService | Dialer (TelecomManager.placeCall) → Telecom in system_server |
| 2 | Telecom → phone process | IConnectionService (and callback IConnectionServiceAdapter) | Telecom ConnectionServiceWrapper → TelephonyConnectionService in com.android.phone |
| 3 | Apps/Telecom → phone process | ITelephony | Callers of TelephonyManager → PhoneInterfaceManager in com.android.phone |
| 4 | Phone process → vendor | IRadioVoice (+ Response + Indication) | RIL.java → vendor radio daemon |
| 5 | Phone process → IMS vendor | IImsMmTelFeature, IImsCallSession | ImsPhoneCallTracker → vendor ImsService |
| 6 | Telecom → in-call UI | IInCallService | Telecom → Dialer's InCallService (call state updates, oneway) |
Dialer app
→ TelecomManager.placeCall() (Binder: ITelecomService)
→ TelecomService / CallsManager [system_server]
→ bind ConnectionService (Binder: IConnectionService)
→ TelephonyConnectionService [com.android.phone]
→ GsmCdmaPhone.dial() → GsmCdmaCallTracker
→ RIL.dial()
→ IRadioVoice.dial(serial) (Binder: stable AIDL HAL, oneway)
→ vendor radio daemon [vendor partition]
→ QMI / AT → modem
◀── IRadioVoiceResponse.dialResponse / IRadioVoiceIndication.callStateChanged
◀── Connection state → Telecom → IInCallService.onCallAdded / updateCall → in-call UI
For VoLTE, the phone process goes through ImsPhone and the vendor ImsService (an MmTelFeature) over IImsCallSession instead of IRadioVoice.dial, while IRadioIms carries IMS-related modem control. Being able to name each boundary and the process on each side is what interviewers look for.
Sending an international parcel: you hand it to the local post office (Telecom in system_server), which passes it to the national sorting centre (the phone process), which hands it to the foreign carrier (the vendor radio daemon), which delivers it to the recipient (the modem). Each handover uses a standard form (an AIDL interface) and a signed receipt comes back the other way (responses and state callbacks).
Stable AIDL vs HIDL, versioning and freezing
Plain AIDL used inside the platform (between system_server and apps built with the same release) can change freely, because both sides ship together. Interfaces that cross an independently updated boundary, such as system ↔ vendor (HALs) or platform ↔ Mainline module, need a stable contract. That is what stable AIDL provides, replacing HIDL.
| HIDL (legacy, Android 8-12) | Stable AIDL (HALs from 11, standard from 13) | |
|---|---|---|
| Definition files | .hal, packages like android.hardware.radio@1.6 | .aidl with @VintfStability, e.g. android.hardware.radio.voice |
| Transport | /dev/hwbinder, hwservicemanager; passthrough mode possible | /dev/binder, servicemanager |
| Versioning | major.minor; a minor version extends the previous by inheritance (IRadio@1.6 extends @1.5 ...) | Integer versions; each version frozen as an API snapshot; new versions only append |
| Radio HAL shape | One monolithic IRadio growing each version | Split by domain: Voice, Data, Network, Sim, Modem, Messaging, Ims, Config |
| Backends | C++ and Java HIDL | NDK C++, Java, Rust (CPP backend is not allowed for vendor) |
| Tooling | hidl-gen, lshal | aidl, Soong aidl_interface, same tooling as framework |
| Status | Deprecated; no new HIDL HALs accepted | The standard for all new HALs |
Declaring and freezing a stable interface
// Android.bp
aidl_interface {
name: "android.hardware.foo",
vendor_available: true,
srcs: ["android/hardware/foo/*.aidl"],
stability: "vintf",
backend: {
java: { sdk_version: "module_current" },
ndk: { enabled: true },
rust: { enabled: true },
},
versions_with_info: [
{ version: "1", imports: [] },
{ version: "2", imports: [] },
],
frozen: true,
}
// Commands
m android.hardware.foo-update-api // refresh the "current" API dump
m android.hardware.foo-freeze-api // create aidl_api/android.hardware.foo/3/ + .hash
- Each frozen version is stored as a snapshot in
aidl_api/<name>/<N>/with a.hashfile. The build fails if a frozen snapshot is modified. - Allowed changes in a new version: add methods at the end, add fields at the end of parcelables (with defaults), add enum values, add new types. Not allowed: remove or reorder methods or fields, change types or method signatures, change method codes.
- Clients check
getInterfaceVersion()(andgetInterfaceHash()) at runtime and avoid calling methods the server does not have. Calling a missing method returnsUNKNOWN_TRANSACTION(STATUS_UNKNOWN_TRANSACTION). - A parcelable written by a newer side with extra fields is read correctly by an older side: stable parcelables are size-prefixed, so unknown trailing fields are skipped.
- The VINTF manifest lists the version the vendor implements; the framework compatibility matrix lists the versions the framework accepts.
<!-- vendor manifest fragment -->
<manifest version="1.0" type="device">
<hal format="aidl">
<name>android.hardware.radio.voice</name>
<version>2</version>
<fqname>IRadioVoice/slot1</fqname>
</hal>
</manifest>
A published printed edition of a standard. Once edition 2 is printed (frozen), no one may change its pages. Edition 3 may add new chapters at the end, but must keep every existing chapter exactly as it was, so a reader holding edition 2 can still follow edition 3 documents by ignoring chapters they do not know. Asking a supplier for a chapter that only exists in edition 3, when they only have edition 2, gets you "unknown chapter" (UNKNOWN_TRANSACTION); polite clients check the supplier's edition first (getInterfaceVersion).
Size limits and TransactionTooLargeException
Each process's Binder receive buffer is about 1 MB (1 MB minus 8 KB with libbinder defaults) and is shared by all transactions currently in flight to that process. Oneway transactions may use at most half of it. If the driver cannot allocate space for an incoming transaction, it fails, and the Java client sees TransactionTooLargeException (for large payloads) or a DeadObjectException/failed transaction.
- A single transaction does not need to be 1 MB to fail: many concurrent medium-sized transactions, or a receiver that is slow to free buffers, can exhaust the space.
- Common culprits: large
Bundles in Intent extras, bigonSaveInstanceStatedata (Android 7.0+ throws aTransactionTooLargeExceptionwrapped inRuntimeExceptionwhen the saved state is too large), returning a huge list from a service, large bitmaps inRemoteViewsor notifications, andgetInstalledPackages()on devices with many apps. - Framework mitigations:
ParceledListSlicesends large lists in chunks over multiple transactions;CursorWindow,SharedMemoryandParcelFileDescriptormove bulk data through shared memory. - Guideline: keep each transaction well under 100 KB, remembering that other calls share the same buffer, and pass file descriptors for anything large.
- Too many Binder proxies: the system limits how many proxies one process can hold toward system_server (a few thousand) and may kill offenders; leaking listener registrations is the usual cause.
E JavaBinder: !!! FAILED BINDER TRANSACTION !!! (parcel size = 1215048)
W System.err: android.os.TransactionTooLargeException: data parcel size 1215048 bytes
W System.err: at android.os.BinderProxy.transactNative(Native Method)
W System.err: at android.os.BinderProxy.transact(BinderProxy.java:...)
W System.err: at android.app.IActivityManager$Stub$Proxy.startActivity(...)
A shared inbox tray of fixed size on each desk. Every letter currently waiting to be read takes space in it. One enormous letter does not fit, but neither do many medium letters if the owner is slow at reading. For big content, you do not stuff it in the tray; you leave a note saying "the documents are in locker 12" and hand over the locker key (a file descriptor to shared memory).
ParceledListSlice, use shared memory or files via fds.Debugging Binder: tools, latency and common failures
| Tool | What it shows |
|---|---|
/sys/kernel/debug/binder/ (or /dev/binderfs/binder_logs/) | state, stats, transactions (in-flight calls with from/to pid:tid), transaction_log, failed_transaction_log, proc/<pid> (threads, nodes, refs, buffers). Needs root / debug builds |
Perfetto (binder_driver ftrace events, android.binder tracks) | Each transaction as a slice linking client and server threads, with latency, reply, and the server's work; the best tool for Binder latency |
am trace-ipc start / stop --dump-file | Java stack traces of every Binder call made by apps while tracing, grouped by interface |
dumpsys binder_calls_stats | Per-interface, per-method call counts and CPU/latency inside system_server (when enabled) |
service list, dumpsys -l, lshal | Registered services; HIDL HALs and clients |
Stack dumps (kill -3, debuggerd -b, ANR traces) | Which threads are blocked in BinderProxy.transactNative / IPCThreadState::talkWithDriver, and what the server's Binder threads are doing |
| logcat | !!! FAILED BINDER TRANSACTION !!!, DeadObjectException, avc: denied { call }, "Slow Binder call" / slow-dispatch warnings |
StrictMode / Binder.setProxyTransactListener | Detect Binder calls on the main thread; observe outgoing transactions in-process |
Measuring Binder latency
- End-to-end latency of a call = driver overhead (usually tens of microseconds) + time waiting for a free server thread + server execution time + any nested calls. In practice, slow calls are almost always slow server work or lock contention, not the driver.
- In Perfetto, select the client's
binder transactionslice and follow the flow arrow to thebinder replyslice on the server thread; look at what the server thread did (running, blocked on a lock, waiting on I/O, making its own outgoing call). - Watch for priority inversion: a foreground caller waiting on a server thread that is runnable but not scheduled (low priority, or in a restricted cgroup).
- Frozen processes: a synchronous call to a process frozen by the cached-apps freezer fails immediately with
BR_FROZEN_REPLY(the caller sees a frozen or dead-object style error). The call does not unfreeze the target. Oneway transactions are queued in the driver until AMS unfreezes the process because it became important again; they can still fail if the async buffer fills.
Failure catalogue
| Symptom | Likely cause | What to check |
|---|---|---|
| Callers hang, ANRs in several apps | Server Binder pool exhausted or service lock contention | Server stacks: are all Binder:pid_N threads in the same slow path? |
DeadObjectException / frozen error | Server process died, or it is frozen and the call was synchronous (BR_FROZEN_REPLY) | Tombstones / am_proc_died vs freezer state; linkToDeath only helps if the process actually died |
TransactionTooLargeException | Payload too big or buffer full | Parcel size in log; Bundle and list sizes |
SecurityException from service | Missing Android permission | Manifest, runtime grants, dumpsys package |
avc: denied { call } / { find } | SELinux policy | Domain and service labels; add a policy rule, do not disable enforcement |
| Service lookup returns null / waits forever | Service not registered, VINTF or init.rc mismatch, lazy start failure | service list, init and servicemanager logs |
UNKNOWN_TRANSACTION | Client calls a method the server version does not implement | getInterfaceVersion(), VINTF versions |
| Deadlock between two processes | Locks held across synchronous calls with callbacks on other threads | Stacks of both processes; debugfs transactions |
Debugging Binder is like investigating slow deliveries in a courier network. The debugfs files are the dispatcher's live board showing which parcels are in transit between which addresses right now. Perfetto is the GPS tracking history for each parcel, showing how long it sat at each depot. Stack dumps are photos of what each courier was doing at one moment. Most of the time the delay is not the road (the driver) but a warehouse that is short of staff (thread pool) or has a jammed door (a lock).
BinderProxy.transactNative, the question is never "why is Binder slow" but "what is the other side doing". Always get the server's stacks at the same moment.Binder and the cached-apps freezer
When AMS freezes a cached process (cgroup v2 freezer, plus BINDER_FREEZE on the driver), that process must not become a black hole that hangs every caller. The driver therefore treats incoming work specially. This is a favourite follow-up once you have mentioned the freezer on the frameworks page.
| Call type | What the driver does | What the caller sees |
|---|---|---|
| Synchronous (two-way) | Does not deliver the call. Returns BR_FROZEN_REPLY to the sender. Does not unfreeze the target | A failed transaction (frozen / dead-object style error). The caller must not retry in a tight loop |
| Oneway (async) | Queues the transaction on the frozen node until the process is unfrozen, or fails if the async half of the buffer is full | Success once queued; the callback or listener runs later, or a later send fails |
- AMS unfreezes a process when it becomes important again (start an activity, deliver a broadcast, bind a service), not because someone called it synchronously.
- The feature was introduced in Android 11 as opt-in and became default-on later (around 12L/13). Do not say "Android 11 turned the freezer on for everyone".
- Freezing is deferred if the process is in the middle of a Binder transaction, so a call already being handled can finish.
BINDER_GET_FROZEN_INFOreports whether sync or async traffic arrived while frozen; that is how AMS accounts for the process, not a signal to thaw it on every sync call.
BR_FROZEN_REPLY, async waits. Unfreeze is a policy decision in AMS.BR_FROZEN_REPLY, say sync fails and does not thaw, async is queued, and AMS thaws on importance. Cross-link: process priority and the freezer.Messenger vs AIDL, and other ways to get a Binder
Not every cross-process API needs a hand-written .aidl file. Interviews often ask you to pick the right IPC tool.
| Tool | What it is | Use when |
|---|---|---|
| AIDL | Generated Stub/Proxy, typed methods, oneway, parcelables, versioning | A real service API: many methods, Binder objects passed around, HALs, framework services |
| Messenger | A thin wrapper: a Handler on the server, an IMessenger Binder under the hood. Clients send Message objects (what, arg1, obj, a replyTo Messenger) | A small, serial API you already think of as messages. All calls land on one Handler thread, so you get sequencing for free and do not have to make the implementation thread-safe |
| Binder (hand-rolled) | Subclass Binder, implement onTransact yourself | Almost never; AIDL generates this |
| Intent / startActivity / startService | One-shot parcelled Intent, no live object | Fire-and-forget work or navigation; not a call-return API |
| ContentProvider | CRUD + file descriptors over Binder, with URI permission grants | Sharing data or large files with another app, including the permission grant model |
// Server (a Service.onBind)
Messenger messenger = new Messenger(new Handler(Looper.getMainLooper()) {
@Override public void handleMessage(Message msg) {
if (msg.what == MSG_PING && msg.replyTo != null) {
Message reply = Message.obtain(null, MSG_PONG);
try { msg.replyTo.send(reply); } catch (RemoteException e) { }
}
}
});
return messenger.getBinder();
// Client
Messenger server = new Messenger(binder);
Message msg = Message.obtain(null, MSG_PING);
msg.replyTo = new Messenger(clientHandler);
server.send(msg);
Messenger is still Binder: DeathRecipient, TransactionTooLargeException and the 1 MB buffer all apply. You lose typed methods, oneway annotations and generated versioning. AIDL is the default for anything you would write a HAL or a system service in.
Quick revision
- Apps and services live in separate processes with separate address spaces, so calls between them need IPC; on Android that is Binder.
- Binder over sockets/pipes: one copy, kernel-stamped caller UID/PID, object handles as capabilities, reference counting and death notifications.
- Binder is a kernel driver used through
open,mmapandioctl(BINDER_WRITE_READ)with BC_* commands and BR_* returns. - Three domains:
/dev/binder(servicemanager: framework, apps, AIDL HALs),/dev/hwbinder(hwservicemanager: HIDL),/dev/vndbinder(vndservicemanager: vendor to vendor). - Handles are per-process integers referring to nodes in other processes; handle 0 is the context manager.
- One copy: the driver copies the sender's data straight into pages mapped read-only into the receiver.
- The receive buffer is about 1 MB per process, shared by all in-flight transactions; oneway may use half.
- Incoming calls run on Binder pool threads: 15 extra by default (16 total), 31 in system_server;
BR_SPAWN_LOOPERgrows the pool. - Pool exhaustion: all threads stuck in slow calls, new callers block and may ANR.
- Recursive callbacks go to the waiting thread; the driver passes caller priority to the server thread.
.aidlgenerates an interface,Stub(server,onTransact) andStub.Proxy(client,transact).asInterface()returns the local object in the same process, otherwise a Proxy.- Method codes start at
FIRST_CALL_TRANSACTIONin declaration order; the interface token is checked withenforceInterface. - Direction tags
in,out,inout; preferin. - Parcel is a sequential, order-dependent IPC container; not for persistence.
- Binder objects and fds inside a Parcel are translated by the driver into receiver-local handles and fds.
- Parcelable is fast hand-written/generated marshalling; Serializable uses reflection and is slow.
- oneway: returns once queued, no reply or exceptions, ordered per object, calling PID is 0.
- servicemanager:
addService/getService/waitForService, SELinux checks, VINTF checks for HALs, lazy start via init. - Apps expose Binders via bound services; AMS brokers
bindServiceand links priorities. linkToDeathgivesbinderDied(); calls to a dead process throwDeadObjectException.Binder.getCallingUid()for permission checks;clearCallingIdentity()/restoreCallingIdentity()in finally.- SELinux checks
binder_call,binder_transfer, and service_manageradd/find. - Radio HAL: oneway request interface with serials, Response interface for solicited replies, Indication interface for unsolicited events.
- RIL registers callbacks with
setResponseFunctions, tracks requests by serial, holds a wakelock while pending, fails pending requests onserviceDied. - An outgoing call crosses Dialer → system_server (Telecom) → com.android.phone → vendor radio daemon → modem.
- Stable AIDL (
@VintfStability) replaced HIDL; versions are frozen snapshots and only append changes are allowed. - Clients use
getInterfaceVersion(); calling a missing method givesUNKNOWN_TRANSACTION. TransactionTooLargeException: payload or concurrent usage exceeded the buffer; send IDs, paginate or use shared memory.- Debug with
/sys/kernel/debug/binder/*, Perfetto Binder tracks,am trace-ipc, stack dumps and logcat failures. - Slow Binder calls are almost always slow server work or lock contention, not driver overhead.
- Frozen target: sync fails with
BR_FROZEN_REPLYand does not unfreeze; oneway is queued. AMS thaws on importance. Freezer was opt-in in 11, default-on later. - Messenger is Binder with a Handler and
Messages (serialized on one thread). AIDL is typed RPC. ContentProvider is for data/URI grants. Intent is one-shot.
Glossary
- AIDL
- Android Interface Definition Language; defines Binder interfaces and generates marshalling code.
- ashmem / memfd
- Anonymous shared memory regions shared between processes by passing a file descriptor.
- asInterface()
- Generated static method that returns a local implementation or wraps a remote IBinder in a Proxy.
- BC_* / BR_*
- Binder protocol commands sent to the driver (BC) and returns from the driver (BR).
- BR_FROZEN_REPLY
- Driver return to a sync caller when the target process is frozen; the target is not unfrozen.
- Binder
- Android's kernel-mediated, object-oriented IPC and RPC mechanism.
- Binder thread pool
- Threads in a process that execute incoming Binder transactions.
- BINDER_WRITE_READ
- The main Binder ioctl; sends commands and receives work in one system call.
- binderfs
- Filesystem that provides Binder device nodes and debug logs on newer kernels.
- BinderProxy
- Java object representing a remote Binder handle in the client process.
- Context manager
- The process reachable at handle 0 in a Binder domain (servicemanager on /dev/binder).
- DeadObjectException
- RemoteException thrown when calling a Binder whose hosting process has died.
- DeathRecipient
- Callback object registered with
linkToDeath(); itsbinderDied()runs when the remote process dies. - FLAG_ONEWAY
- Transaction flag meaning the caller does not wait for a reply.
- flat_binder_object
- Parcel entry describing a Binder object or fd that the driver translates for the receiver.
- Handle
- Per-process integer that refers to a Binder node owned by another process.
- HIDL
- HAL Interface Definition Language used over hwbinder in Android 8-12; now deprecated.
- hwservicemanager
- Context manager for /dev/hwbinder, used by HIDL HALs.
- IBinder
- Base interface for remotable objects; provides
transact,linkToDeath,pingBinder. - Indication
- Unsolicited event from a HAL to the framework, e.g.
callStateChanged. - Interface token
- Header written first in a call: interface descriptor plus StrictMode and work-source data.
- Lazy HAL
- HAL service started on first lookup and allowed to exit when unused.
- linkToDeath
- Registers for notification when the process hosting a Binder dies.
- Messenger
- Handler-backed Binder wrapper: clients send
Messageobjects; all calls run on one thread. - Node
- Driver object representing a Binder service in its owning process.
- oneway
- AIDL keyword for asynchronous calls without a reply.
- onTransact
- Stub method that decodes an incoming transaction and dispatches it to the implementation.
- Parcel
- Sequential container holding one transaction's arguments or reply.
- Parcelable
- Interface for objects that write themselves to and read themselves from a Parcel.
- ParceledListSlice
- Framework helper that sends large lists across several transactions.
- Proxy
- Generated client-side class that marshals calls and sends transactions.
- RemoteCallbackList
- Helper that holds client callbacks and removes them automatically when clients die.
- Serial
- Request ID in the Radio HAL used to match async responses to requests.
- servicemanager
- Name registry for Binder services on /dev/binder.
- Solicited response
- HAL reply to a specific framework request, correlated by serial.
- Stable AIDL
- AIDL interfaces with frozen versions for boundaries that update independently, such as HALs.
- Stub
- Generated server-side base class that receives transactions.
- TransactionTooLargeException
- Error when a transaction cannot fit in the receiver's Binder buffer.
- UNKNOWN_TRANSACTION
- Status returned when a transaction code is not implemented by the server.
- VINTF
- Vendor interface manifests and compatibility matrices declaring HAL names and versions.
- @VintfStability
- AIDL annotation marking an interface as a stable system/vendor contract.
- vndbinder
- Binder domain (/dev/vndbinder) for vendor-to-vendor IPC, managed by vndservicemanager.
Interview questions
Fundamentals
What is Binder?
Binder is Android's inter-process communication mechanism: a kernel driver (/dev/binder) plus user-space libraries (libbinder, libbinder_ndk, the Java Binder classes) that let one process call methods on an object in another process as if it were local. It provides object handles, one-copy data transfer, a thread pool, caller identity and death notifications. Almost every app-to-framework and framework-to-HAL call uses it.
Why does Android use Binder instead of standard Linux IPC?
- Performance: one data copy instead of two.
- Security: the kernel stamps each transaction with the caller's UID/PID, which cannot be forged.
- Object model: handles to remote objects act as capabilities and can be passed between processes.
- Lifetime: cross-process reference counting and death notifications.
- Built-in RPC semantics (method codes, replies, exceptions) and thread pool management.
What is AIDL?
AIDL (Android Interface Definition Language) describes an interface that can be called across processes. From an .aidl file the build generates a Java/C++/Rust interface, a server-side Stub and a client-side Proxy that handle all marshalling. It is used for app-to-app bound services, framework services and, as stable AIDL, for vendor HALs.
What is the difference between Stub and Proxy?
The Stub is the server half: an abstract class extending Binder that you subclass to implement the methods; its onTransact() unpacks incoming Parcels and calls your implementation. The Proxy is the client half: it wraps a remote IBinder and implements the same interface by packing arguments into a Parcel and calling transact(). Both are generated from the same .aidl file so their layouts match.
What is a Parcel?
A Parcel is the container for one transaction's data: arguments on the way in, return value and exception status on the way out. Values are written and read sequentially in the same order. It can also carry Binder objects and file descriptors, which the driver translates for the receiving process. It is designed for IPC only and is not a stable storage format.
What does oneway mean in AIDL?
A oneway method (or interface) is asynchronous: the caller's transact() returns as soon as the driver queues the transaction, without waiting for the server to run it. There is no return value and server exceptions do not reach the caller. Calls to the same Binder object are delivered in order. It is used for callbacks, notifications and async HAL request/response patterns.
What is servicemanager?
servicemanager is the Binder context manager on /dev/binder, reachable by every process at handle 0. Services register with addService(name, binder); clients get handles with getService(name) or waitForService(name), then call the service directly. It enforces SELinux rules for registration and lookup, validates VINTF declarations for HALs, and can start lazy services through init.
How many data copies does a Binder transaction make?
One. The receiving process has mmap'd a buffer from the driver. The driver copies the sender's Parcel data with copy_from_user() directly into physical pages that are mapped both in the kernel and (read-only) in the receiver, so the receiver reads it in place. Sockets and pipes need two copies. Large data can be shared with zero copies by passing a shared-memory file descriptor.
On which thread does an AIDL method run in the server?
On one of the server process's Binder threads, not the main thread (unless the caller is in the same process, in which case it runs on the caller's thread). Multiple calls can therefore run concurrently, so the implementation must be thread-safe. If it needs to touch UI or main-thread state, it must post to a Handler.
What does asInterface() do?
It converts an IBinder into the typed interface. It calls queryLocalInterface(DESCRIPTOR): if the Binder is a local object in the same process, it returns that object directly so calls are plain method calls; otherwise it wraps the Binder in a new Stub.Proxy that performs IPC. It returns null for a null Binder.
What is linkToDeath?
IBinder.linkToDeath(recipient, flags) registers a DeathRecipient whose binderDied() is called when the process hosting that Binder dies. Clients use it to drop stale references and reconnect; servers use it to clean up state held for clients. It works because the driver knows when a process closes its Binder file descriptor.
What is DeadObjectException?
A subclass of RemoteException thrown when you call a method on a Binder whose hosting process has died. The Binder handle is now useless; the client must obtain a new one after the service restarts. Robust clients catch it, and preferably register a death recipient to react before the next call.
What are the in, out and inout tags?
They specify the direction of non-primitive parameters. in: data goes from client to server only (the default, and the only option for primitives). out: the server fills the object and it is copied back to the client. inout: copied both ways, costing twice as much. Use in unless you need the server to return data through the parameter.
What types can AIDL methods use?
Primitives, String, CharSequence, arrays, List and Map of supported types, Parcelable classes, other AIDL interfaces (passed as Binders), IBinder, ParcelFileDescriptor, and in structured AIDL, parcelables, enums and unions defined in .aidl. Stable AIDL disallows unstructured (hand-written) parcelables.
How do you know who called your Binder service?
Call Binder.getCallingUid() and Binder.getCallingPid() inside the method (native: IPCThreadState::self()->getCallingUid()). The driver sets these from the sending process, so they cannot be forged. Use the UID for permission checks (enforceCallingPermission) because PIDs can be reused and are 0 for oneway calls.
What are /dev/binder, /dev/hwbinder and /dev/vndbinder?
Three separate Binder domains introduced with Treble. /dev/binder (servicemanager) is for framework and apps, and now also stable-AIDL HALs. /dev/hwbinder (hwservicemanager) carried HIDL HAL traffic between framework and vendor. /dev/vndbinder (vndservicemanager) is for vendor processes talking to each other. Separate domains keep system and vendor namespaces and policies apart.
What is HIDL and how does it relate to AIDL?
HIDL was the HAL interface language introduced in Android 8.0 with Treble, using .hal files and hwbinder. Stable AIDL, supported for HALs from Android 11 and the standard from 13, replaced it: same Binder concepts, but one language and one transport for both framework and HAL interfaces. HIDL is deprecated and no new HIDL HALs are accepted.
What is TransactionTooLargeException?
It is thrown when a Binder transaction cannot be delivered because the receiving process's Binder buffer (about 1 MB, shared by all in-flight transactions) has no room for it. Typical causes are large Bundles in Intents, big saved instance state, or returning large lists or bitmaps. Fix by sending less (IDs instead of objects), chunking, or using shared memory or files passed by file descriptor.
How does an app expose an AIDL interface to other apps?
It implements a bound Service whose onBind() returns an instance of the generated Stub, and declares the service in its manifest (usually exported and protected by a permission). Clients call bindService() with an intent and receive the Binder in ServiceConnection.onServiceConnected(), then call IFoo.Stub.asInterface(binder). Both apps need the same .aidl file.
What is a Binder thread pool?
The set of threads in a process that wait in the driver for incoming transactions and execute them. libbinder starts a main pool thread and the driver asks for more (BR_SPAWN_LOOPER) when all are busy, up to the configured maximum (15 extra by default, 31 for system_server). If all are busy, new calls wait.
AIDL vs Messenger?
Both are Binder. AIDL generates a typed Stub/Proxy with method codes, oneway, parcelables and versioning; incoming calls run on the Binder thread pool, so the implementation must be thread-safe. Messenger wraps a Handler: clients send Message objects (what, args, optional replyTo) and they are handled one at a time on that Handler's thread, so you get sequencing for free and a weaker schema. Use AIDL for a real service or HAL; use Messenger for a small message-oriented API. Neither replaces shared memory for large payloads.
What is the Radio HAL request/response/indication pattern?
The framework calls request methods on IRadioX (for example IRadioVoice.dial(serial, info)), which return immediately. The vendor later calls IRadioXResponse methods with the same serial to deliver the result, and calls IRadioXIndication methods for unsolicited events like call state changes. The framework registers its Response and Indication objects with setResponseFunctions().
What is Parcelable vs Serializable?
Parcelable requires explicit (hand-written or generated) writeToParcel and CREATOR code, uses no reflection and is designed for fast IPC. Serializable is a Java marker interface that relies on reflection, creates many temporary objects and is much slower. Use Parcelable for anything passed through Binder or Intents.
Going deeper
Walk me through a Binder transaction at the wire level.
- Client calls
proxy.add(2, 3); the Proxy writes the interface token and arguments into a Parcel. BinderProxy.transact()→IPCThreadState::transact()queuesBC_TRANSACTIONand callsioctl(BINDER_WRITE_READ); the client thread blocks.- The driver resolves the handle to the target node, checks SELinux, allocates a buffer in the target's mmap area, copies the data once, records the caller's UID/PID, and queues the work on an idle Binder thread (asking for a new thread if needed).
- The server thread returns with
BR_TRANSACTION;Stub.onTransact()checks the token, reads the arguments and calls the implementation. - The reply is written and sent with
BC_REPLY; the driver copies it to the client and wakes it withBR_REPLY. - The Proxy reads the exception status and result and returns.
Explain how mmap enables one-copy transfer.
When a process opens Binder, libbinder mmaps about 1 MB of the device. The driver backs this area with physical pages allocated on demand and maps them both in kernel space and read-only in the process's user space. For a transaction to that process, the driver allocates a region of this area and does one copy_from_user() from the sender's buffer into it. Since the receiver already has those pages mapped, it reads the data directly with no second copy. The receiver later frees the region with BC_FREE_BUFFER.
What is a Binder handle and how does the driver translate Binder objects in a Parcel?
A handle is an integer in the client process that refers to a binder_ref, which points to a binder_node owned by another process. When a Parcel contains a Binder object, it is written as a flat_binder_object. The driver rewrites it: a local object (BINDER_TYPE_BINDER) sent out becomes a handle (BINDER_TYPE_HANDLE) in the receiver, creating a node and reference as needed; a handle sent to the node's own process becomes the local object pointer again; a handle sent to a third process becomes that process's own handle. That is how a capability is passed safely.
How does the driver grow the Binder thread pool?
The process sets a maximum with BINDER_SET_MAX_THREADS. When the driver has work for a process and no thread is waiting, and the number of spawned threads is below the maximum, it adds BR_SPAWN_LOOPER to a thread's return buffer. libbinder then creates a new thread that registers with BC_REGISTER_LOOPER and enters the loop. Threads started by the app itself use BC_ENTER_LOOPER. Pool threads are not destroyed when idle.
What is Binder thread-pool exhaustion and how does it show up?
It happens when every Binder thread in a server process is busy, typically all stuck in the same slow path (a lock, disk I/O, a nested synchronous call to a HAL). New transactions queue in the driver, so callers block. Symptoms: many apps hang or ANR in BinderProxy.transactNative calls to the same service; server stack dumps show all Binder:pid_N threads in similar frames. Fix the slow path, make handlers fast, use async callbacks or oneway, and avoid nested blocking calls.
How are exceptions propagated across Binder?
The generated Stub catches exceptions from the implementation. For a set of parcelable exception types (SecurityException, IllegalArgumentException, IllegalStateException, NullPointerException, UnsupportedOperationException, ServiceSpecificException, ...) it writes an exception code and message into the reply with writeException(); the Proxy's readException() rethrows it in the client. Other runtime exceptions are logged in the server and the client receives a generic failure. With oneway calls, nothing reaches the caller.
What are the trade-offs of oneway calls?
Pros: the caller never blocks on the server, which protects system_server when calling apps; ordering is preserved per Binder object. Cons: no return value or exception, so you need a callback to learn results or errors; async transactions can use only half of the target's buffer and a flood can fill it; calling PID is 0; the server executes oneway calls for one object serially, so a slow one delays the rest; and if the object is local, oneway runs synchronously.
Why is getCallingPid() 0 for oneway calls?
For oneway transactions the driver does not keep a link to the sending thread (there is no reply to route back), and it deliberately reports sender PID as 0 because by the time the server processes the call the sender may have exited and its PID been reused. The UID is still delivered. Services must therefore use UID-based checks for oneway methods.
What does clearCallingIdentity() do and when is it needed?
It resets the thread's calling UID/PID to the current process and returns a token holding the original identity. It is needed when a service, while handling a client call, performs work that should be checked against the service's own identity: calling another service, accessing its own content provider, sending a broadcast. Always pair it with restoreCallingIdentity(token) in a finally block; leaving the identity cleared is a privilege escalation bug.
How does servicemanager handle lazy services?
A lazy service's interface is declared in an init.rc service entry with interface aidl <name> (and for HALs in the VINTF manifest) but it is not started at boot. When a client looks it up, servicemanager sets ctl.interface_start, init starts the matching service, and the service registers with registerLazyService. servicemanager tracks clients; when none remain, it notifies the service, which may unregister and exit. Mismatched names mean the start fails and clients never get the service.
getService vs checkService vs waitForService?
checkService returns immediately with the Binder or null. getService historically retried for a few seconds before returning null (and may start lazy services). waitForService blocks until the service is registered, starting it if it is lazy and declared, and is the recommended call for HALs and services known to exist. isDeclared checks whether a VINTF service is declared without starting it.
How does SELinux interact with Binder?
SELinux checks happen in the driver and in servicemanager before any service code runs. binder_call controls which domains may send transactions to which; binder_transfer controls passing Binder objects; binder_set_context_mgr restricts who can become servicemanager. servicemanager checks add, find and list permissions against labels in service_contexts / hwservice_contexts. Denials appear as avc: denied and must be fixed with policy, not by disabling enforcement.
How are file descriptors passed through Binder?
You write an fd into the Parcel (writeFileDescriptor or a ParcelFileDescriptor). It is encoded as a BINDER_TYPE_FD object. The driver installs a new file descriptor in the receiving process referring to the same open file description and rewrites the value in the Parcel. The receiver can then read the file or mmap the shared memory. This is how CursorWindow, SharedMemory, graphics buffers and pipes for large data are shared.
What is the interface token in a Parcel?
The first data written by a Proxy via writeInterfaceToken(DESCRIPTOR). It contains the StrictMode policy of the calling thread (so the server can apply it), a work-source UID for power attribution, a header identifying the system/vendor stability context, and the interface descriptor string. The Stub calls enforceInterface() to verify the descriptor matches, rejecting transactions meant for a different interface.
What is RemoteCallbackList and why use it?
A framework class for servers that keep a list of client callback interfaces. It links to death on each registered callback and removes it automatically when the client dies, identifies callbacks by their underlying Binder (so the same client registering through different proxies is recognized), and provides safe iteration with beginBroadcast()/finishBroadcast() while calling out. Using a plain list leaks callbacks from dead clients and causes DeadObjectExceptions.
How does a client detect a server restart and reconnect?
Register a DeathRecipient with linkToDeath(). In binderDied(), drop the stale proxy and any state tied to it, then (usually on a Handler, not the Binder thread) call waitForService() or rebind, re-register callbacks, and replay required state. For bound app services, ServiceConnection.onServiceDisconnected() and then onServiceConnected() fire automatically when the service restarts. RIL's handling of serviceDied is a real example.
Why does RIL use serials in the Radio HAL?
Because the HAL is asynchronous: a request returns immediately and the answer arrives later on a separate Response interface. Many requests can be outstanding, and responses can arrive in any order. Each request gets a unique serial stored in RIL's request list; the response's RadioResponseInfo.serial identifies which RILRequest to complete, which caller Message to send, and which wakelock count to release.
In the Radio HAL, who is the client and who is the server?
Both are both. For requests (IRadioVoice), the vendor radio daemon is the server and RIL.java is the client. For IRadioVoiceResponse and IRadioVoiceIndication, the framework (phone process) is the server and the vendor daemon is the client. The framework passes its Response and Indication Binder objects to the vendor with setResponseFunctions(), which is possible because Binder objects can be sent inside Parcels.
Solicited vs unsolicited messages in RIL?
Solicited: framework-initiated requests (dial, get signal strength) with a serial; the vendor answers on the matching Response callback, completing the RILRequest; RIL holds a wakelock while they are pending. Unsolicited: modem-initiated events (call state change, network state change, incoming SMS, NITZ) delivered on the Indication interface; RIL converts them to Messages and notifies subscribers via RegistrantList. Some indications require an acknowledgement so the vendor can release its wakelock.
Name the Binder boundaries crossed by an outgoing call.
Dialer app → Telecom in system_server (ITelecomService); Telecom → TelephonyConnectionService in com.android.phone (IConnectionService, with IConnectionServiceAdapter callbacks); phone process → vendor radio daemon (IRadioVoice plus Response/Indication), or for VoLTE, vendor ImsService (IImsCallSession); Telecom → in-call UI (IInCallService). Other apps reach the phone process through ITelephony via TelephonyManager.
What does freezing a stable AIDL interface mean?
Running m <name>-freeze-api copies the current interface into aidl_api/<name>/<N>/ with a hash, making version N immutable. The build then checks that the snapshot never changes and that the next version is a compatible extension. Devices implementing version N can be relied on by any future framework that supports N.
What changes are allowed between stable AIDL versions?
Allowed: adding methods at the end of an interface, adding fields at the end of parcelables (with sensible defaults), adding enum values, adding new types and interfaces. Not allowed: removing, renaming the wire meaning of, or reordering methods and fields; changing parameter or return types; changing the oneway status. Method codes and field positions are part of the wire format.
What is the difference between the NDK, CPP, Java and Rust AIDL backends?
Java is for framework and app code. CPP uses the platform-internal libbinder, whose ABI is not stable, so it is only for code built with the platform. NDK uses libbinder_ndk, a stable C API with C++ wrappers, required for vendor code and APEX modules that must keep working when the system updates. Rust uses the binder crate on top of libbinder_ndk. Stable HALs are typically NDK or Rust on the vendor side and Java or NDK on the framework side.
What causes !!! FAILED BINDER TRANSACTION !!! in logcat?
The driver could not deliver a transaction. Most often the receiver's Binder buffer is full (a large payload, many concurrent transactions, or a slow receiver not freeing buffers), which surfaces as TransactionTooLargeException. It can also mean the target died (DeadObjectException) or the target is frozen (BR_FROZEN_REPLY for a sync call). The log shows the parcel size, which helps separate "one huge call" from "buffer already full".
What is BR_FROZEN_REPLY?
The Binder driver's reply to a synchronous caller when the target process is frozen by the cached-apps freezer. The call is not delivered and the target is not unfrozen. The caller sees a failed transaction. Oneway calls are queued instead. AMS unfreezes the process later when it becomes important (activity, broadcast, bind), not because of the sync call. The freezer was opt-in in Android 11 and default-on later.
How can you see which Binder calls an app makes?
adb shell am trace-ipc start, exercise the app, then am trace-ipc stop --dump-file /data/local/tmp/ipc.txt; the file lists Binder calls with Java stack traces grouped by interface and count. Perfetto with Binder tracing shows each transaction on a timeline with latency. StrictMode can flag Binder calls on the main thread in debug builds.
Advanced
How does Binder handle recursive (nested) calls between two processes?
Each binder_thread keeps a transaction stack. If thread T1 in process A calls process B, and B's handler thread T2 calls back into A synchronously while handling it, the driver sees that T2's current transaction originated from T1 and delivers the callback to T1 itself (which is blocked waiting for its reply) instead of to a random pool thread in A. T1 processes the nested call and returns, then continues waiting for its original reply. This avoids deadlocks from reentrancy and preserves thread-local state, but callbacks landing on other threads needing the same locks can still deadlock.
How does Binder priority inheritance work?
When a transaction is delivered, the driver temporarily sets the server thread's scheduling priority to match the caller's (nice value, and real-time policy/priority if the node permits RT inheritance, which HALs like audio use). After the reply, the server thread's priority is restored. This prevents a high-priority caller (for example the UI thread or an audio thread) from being blocked behind a low-priority server thread, a form of priority inversion. oneway calls do not inherit priority from the caller in the same way; they use the node's minimum priority.
How does Binder reference counting work across processes?
The driver tracks strong and weak references per binder_ref (in clients) and aggregates them on the binder_node (in the owner). User space sends BC_INCREFS/BC_ACQUIRE/BC_RELEASE/BC_DECREFS as proxies are created and destroyed; the driver tells the owner with BR_ACQUIRE/BR_RELEASE etc. so it can keep the local object alive while remote references exist. In Java, BinderProxy finalization and GC drive releases; in native code, sp<IBinder> smart pointers do. Leaked proxies keep remote objects alive.
How is the oneway queue per Binder node managed?
Each node has an async todo list. If a oneway transaction for that node is already being processed, later oneway transactions for the same node wait in that list instead of being given to another thread, so they are executed serially and in order. When the server frees the buffer of the current async transaction, the next is dispatched. This is why one slow oneway method delays all later oneway calls on the same object, while different objects in the same process can proceed in parallel.
Why can a oneway spam from one client affect a server, and what protects against it?
oneway transactions consume buffer space in the server until processed. If a client sends them faster than the server handles them, async space (half the buffer) fills, and further async transactions fail; synchronous callers can also be affected if the buffer is exhausted. The driver detects suspicious async senders and logs "pid X spamming oneway?" when one process uses a large share of async space; newer kernels can flag such transactions. Services should rate-limit, coalesce updates, or switch to synchronous calls or shared memory for bulk data.
How does Java Binder connect to native libbinder?
Java Binder objects have a native peer (JavaBBinder, a BBinder subclass) created lazily when the object is first sent across processes. Incoming transactions arrive on a native Binder thread in BBinder::transact(), which calls into Java via JNI (Binder.execTransact() → onTransact()). On the client side, a BinderProxy Java object wraps a native BpBinder; transact() goes through JNI to BpBinder::transact() → IPCThreadState::transact(). Java Parcels wrap native Parcel objects.
Explain the stability header in Binder (system vs vendor).
libbinder tags each Binder object with a stability level: local (same partition), vendor, system, or VINTF. A Binder marked with @VintfStability may be passed across the system/vendor boundary; a platform-internal (system-local) Binder must not be sent to vendor processes, and vice versa, because their interfaces are not stable. libbinder checks this when objects are sent or used, and interface tokens include a system or vendor header. This enforces Treble rules at runtime, not just at build time.
How would you add a new method to a frozen HAL and keep old vendor implementations working?
- Add the method at the end of the interface in the "current" version and any new parcelable fields at the end.
- Freeze a new version (N+1) and update the framework compatibility matrix to accept both N and N+1.
- In the framework, check
getInterfaceVersion()on the HAL proxy; call the new method only when the version is at least N+1, else use a fallback path. - Vendors implementing N continue to work; those upgrading implement N+1 and bump their VINTF manifest version.
- Test both combinations in VTS/CTS.
How does a stable parcelable stay compatible when fields are added?
Stable (structured) parcelables are written with a size prefix. A reader reads the fields it knows, then uses the size to skip any trailing unknown fields written by a newer version. A newer reader receiving an older, shorter parcelable reads the fields that are present and leaves the rest at their declared defaults. This is why fields may only be appended and why defaults matter.
What are the costs of a Binder call, and when is it too expensive?
A minimal call involves two system calls (send and receive), two context switches, one copy each way, marshalling and unmarshalling, and waking a server thread: on the order of tens of microseconds on modern hardware. That is fine for occasional calls but expensive in tight loops (for example calling getPackageInfo thousands of times, or per-frame calls). Batch requests, cache results on the client, use callbacks instead of polling, and use shared memory or FMQ for streaming data.
What is FMQ and when is it used instead of Binder calls?
Fast Message Queue is a HAL-level ring buffer in shared memory, set up once over Binder (the descriptor is passed in a call) and then used without Binder for each message. Producers and consumers read and write directly, optionally using an event flag (futex) to wake each other. It is used for high-rate, low-latency data such as audio, sensors, and neural network execution, where a Binder call per message would add too much overhead.
How does the cached-apps freezer interact with Binder?
Before freezing a process, AMS asks the driver (BINDER_FREEZE) to freeze its Binder state. While frozen, a synchronous transaction fails immediately with BR_FROZEN_REPLY; the caller gets an error and the target is not unfrozen. Oneway transactions are queued until AMS unfreezes the process because it became important again (activity, broadcast, bind), and can fail if the async buffer fills. If a process is in the middle of a Binder transaction, freezing is deferred. The driver can report sync_recv/async_recv via BINDER_GET_FROZEN_INFO so AMS knows traffic arrived; that is telemetry, not an automatic unfreeze.
What is the proxy limit and why does system_server enforce it?
Each Binder object an app sends to system_server (such as a listener) creates a BinderProxy in system_server. A buggy app that registers listeners in a loop without unregistering can create tens of thousands of proxies, exhausting memory and driver resources. system_server tracks proxies per UID and, above a high watermark (thousands), logs and kills the offending app. It protects the whole system from one app's leak.
Why does the Zygote not use Binder, and how does a forked child get Binder?
Binder requires a thread pool, and fork() in a multithreaded process copies only the calling thread, leaving locks and driver state inconsistent in the child. Zygote therefore stays single-threaded and receives fork commands over a Unix socket. After fork and specialization, the child opens /dev/binder itself (ProcessState), mmaps its buffer and starts its own Binder thread pool in onZygoteInit(), so it has a fresh, clean Binder state.
What is the difference between a Binder node's weak and strong references, and why do both exist?
A strong reference keeps the remote object alive; a weak reference only keeps the node identity valid and can be promoted to strong if the object still exists. Weak references let clients hold a handle without forcing the service object to stay alive (for example caches, or death-notification bookkeeping). libbinder's sp/wp mirror this. For most Java code only strong references matter.
How does AIDL support versioned interfaces at runtime in Java?
Generated Java stubs for versioned interfaces include getInterfaceVersion() and getInterfaceHash() methods (reserved transaction codes near LAST_CALL_TRANSACTION). A client can call these to determine what the server implements. The generated Proxy can also have a default implementation set with setDefaultImpl(), which is called when the server returns UNKNOWN_TRANSACTION for a method it does not implement, giving graceful fallback.
How does Binder deliver a death notification internally?
A client sends BC_REQUEST_DEATH_NOTIFICATION for a handle with a cookie. When the owning process exits, the kernel releases its binder_proc (on fd close), marks its nodes dead, and for every reference with a registered death notification queues BR_DEAD_BINDER to that client process. A client thread picks it up, libbinder calls the registered recipients (binderDied), and acknowledges with BC_DEAD_BINDER_DONE. Any transaction in flight to the dead process returns BR_DEAD_REPLY to its caller.
Can a server call back into a client that is currently blocked on a synchronous call to it?
Yes. A synchronous callback from the server to the waiting client is routed to the client thread that is blocked in the call (the recursion rule), which executes it and returns. The danger is when the callback needs a lock the client thread already holds while calling, or when it is delivered to a different thread (for example a oneway callback) that then waits on that lock. Design rule: do not hold locks while making outgoing Binder calls, and treat callbacks as reentrant.
Why is Binder a good fit for HALs after Treble, compared with shared libraries?
Running HALs in separate processes over Binder gives a stable IPC contract instead of a C ABI, so system and vendor can be built and updated separately. It also isolates faults (a crashing HAL does not crash system_server, and death recipients allow recovery), applies least privilege (each HAL has its own SELinux domain), and lets the same interface be used from Java, C++ and Rust. The cost, per-call IPC overhead, is mitigated with FMQ and shared memory for high-rate data.
How do you define and use a callback interface in AIDL?
// IListener.aidl
oneway interface IListener {
void onEvent(int code);
}
// IService.aidl
interface IService {
void register(IListener l);
void unregister(IListener l);
}The client implements IListener.Stub and passes it to register(); the driver turns it into a proxy in the server. The server stores it in a RemoteCallbackList and calls onEvent() on it later. Making the callback interface oneway ensures a slow client cannot block the server.
Scenario & debugging
Many apps ANR at the same time and all traces show BinderProxy.transactNative into system_server. How do you find the root cause?
Identify which interface they call from the Java frames (for example IActivityManager or IPackageManager). Look at the system_server stacks in the same ANR dump: find Binder threads handling those calls and see what they wait on, typically a global service lock ("waiting to lock ... held by thread N"). Follow thread N: often it holds the lock while doing I/O or an outgoing call to a HAL or app. Check whether all 31 Binder threads are busy (pool exhaustion). Fix the lock holder's slow path; the apps are victims.
A call to your service occasionally takes 2 seconds. How do you investigate?
- Reproduce and capture a Perfetto trace with Binder (binder_driver) and scheduling events.
- Find the slow transaction slice in the client, follow the flow to the server's reply slice.
- Examine the server thread during that time: running (CPU-heavy work), sleeping on a lock (who holds it?), waiting on I/O, or making its own nested Binder call.
- Check if the transaction waited before a server thread picked it up (pool exhaustion) or if the server thread was runnable but not scheduled (priority, CPU contention).
- Fix accordingly and add latency metrics or a regression test.
An app crashes with TransactionTooLargeException when starting an activity. What do you do?
The Intent extras (a Bundle) are too large, often a bitmap, a large list, or a big serialized object. Measure the Bundle size (for example by parceling it and checking dataSize()). Fix by passing an ID or URI and loading the data in the target from a repository, database, file or content provider, and by not storing large data in saved instance state. If it happens in onSaveInstanceState, move large state to a ViewModel or persistent storage.
After a HAL restart, the framework keeps getting DeadObjectException. What is wrong?
The framework client is holding a stale proxy from before the restart and never re-acquired the service. It either did not register a death recipient or its binderDied() handler does not reconnect. Fix: linkToDeath on the HAL binder; in binderDied(), clear the proxy, fail or retry pending requests, call waitForService() to get the new instance, re-register callbacks (for the Radio HAL, call setResponseFunctions() again), and restore required state.
RIL reports serviceDied. Does that mean the modem crashed?
Not necessarily. serviceDied means the vendor radio HAL process died, which RIL detects through its death recipient. The modem (baseband) might be fine, or a modem subsystem restart might have caused the daemon to exit. Check for modem crash or subsystem-restart markers and ramdump logs at the same timestamp, and look for the HAL daemon's tombstone. RIL itself completes all pending requests with RADIO_NOT_AVAILABLE, resets its proxies, and re-acquires the HAL when it comes back.
A new HAL service works on the bench but the framework cannot find it in the full build. How do you debug?
Run service list | grep <name> and check if the process is running. Look in logcat for servicemanager errors such as "not declared in VINTF manifest" or SELinux avc: denied { add }/{ find }, and init errors about ctl.interface_start. Verify the instance name matches exactly across the VINTF manifest, init.rc interface line and the registration call, that the service has a service_contexts label, and that the framework compatibility matrix includes the HAL version.
SELinux blocks a Binder call from your new daemon. How do you fix it properly?
Read the denial: avc: denied { call } for scontext=u:r:mydaemon:s0 tcontext=u:r:system_server:s0 tclass=binder. Add the minimal policy: use macros like binder_use(mydaemon), binder_call(mydaemon, target_domain), and allow find on the target service's label (allow mydaemon foo_service:service_manager find;). Respect Treble neverallow rules (vendor domains cannot call arbitrary system services). Never switch to permissive in production; verify with a CTS/VTS run.
system_server is killed by the watchdog; the blocked thread is a Binder thread calling a HAL. What happened and how do you fix it?
A system_server service took its main lock and then made a synchronous Binder call to a vendor HAL that hung (HAL deadlock, hardware timeout, or a HAL waiting on system_server). Other threads needing the lock blocked until the watchdog timeout, so system_server was killed. Get the HAL process stacks (debuggerd) from the bugreport to find why it hung. Fix both sides: the HAL must not block indefinitely (timeouts), and the framework must not hold global locks across HAL calls (call outside the lock, or use async/oneway with callbacks).
Two processes deadlock with each other through Binder. How do you confirm and prevent it?
Take stack dumps of both processes during the hang: typically thread A1 holds lock LA and is in a synchronous call to B; B's handler thread B1 holds LB and is in a synchronous call to A, where an A pool thread waits for LA. The debugfs transactions file shows the transaction chains between their threads. Prevent by never holding locks during outgoing calls, defining a strict call direction (for example the system calls vendors only asynchronously), and using oneway callbacks posted to a Handler.
After a framework update, a vendor HAL method call returns UNKNOWN_TRANSACTION. Why?
The framework is calling a method added in a newer HAL version, but the vendor implements an older version that does not have that transaction code. The framework should check getInterfaceVersion() before calling newer methods and fall back if needed. Confirm the vendor's declared version in the VINTF manifest and dumpsys/lshal-equivalent output, and fix the framework code path or update the vendor implementation.
Your service handles a burst of requests and some clients time out. What would you change in its design?
Check whether handlers do slow work on Binder threads, exhausting the pool. Move long work to a dedicated executor and respond asynchronously through a callback interface (or return a future-like handle), keeping Binder methods short. Make fire-and-forget methods oneway. Reduce lock scope, avoid nested synchronous calls, and consider raising the pool size only if work is genuinely parallel. Add metrics on handler latency and queue depth.
A listener-heavy app causes system_server memory to grow and eventually the app is killed. What is happening?
The app repeatedly registers new callback Binder objects (for example on every resume) without unregistering. Each creates a proxy and a death recipient in system_server; when it crosses the per-UID proxy limit, system_server kills the app to protect itself. Confirm with logs about too many Binder proxies and dumpsys meminfo system_server object counts. Fix by registering once, unregistering in the matching lifecycle callback, and using the same callback object.
An in-process call through an AIDL interface behaves differently from the remote call (for example a mutated argument). Why?
When client and server are in the same process, asInterface() returns the real object and there is no marshalling: objects are passed by reference, so the server can mutate the client's objects and in semantics are not enforced; oneway methods also run synchronously; and calling identity is the process's own. Remote calls copy data through Parcels. Write implementations that do not rely on either behaviour, and test the remote path.
How would you debug a Radio HAL request that never gets a response?
Find the request's serial in the radio log (RIL logs request and response with serials). Check whether the vendor daemon received it (vendor logs), whether it sent a modem command and got an answer (modem logs), and whether it called the Response method. Confirm the Response object is still registered (a HAL restart without setResponseFunctions would lose callbacks). RIL's wakelock timeout will fire if the response never arrives; look for that log too. Fix the layer where the chain breaks.
Main-thread Binder calls are causing jank in your app. How do you find and fix them?
Enable StrictMode or use Perfetto to find binder transaction slices on the main thread during frames; am trace-ipc lists calls with stack traces. Common offenders: PackageManager queries, getSystemService calls that do IPC, content provider queries, and account or settings reads. Fix by moving calls to background threads, caching results (for example package info that does not change), batching queries, and registering for change callbacks instead of polling.
Oneway callbacks from your service to a client arrive late or pile up. What could be wrong?
oneway calls to the same client Binder object execute serially, so if the client's callback implementation is slow, all later callbacks queue behind it. The client's async buffer can fill, causing failed transactions. Also, a frozen (cached) client does not run oneway work until AMS unfreezes it; a sync call to that client fails with BR_FROZEN_REPLY instead of thawing it. Fixes: keep client callbacks tiny and post work to a Handler, coalesce updates on the server (send only the latest state), and avoid sending high-rate events per callback.