Skip to content

Archive

Process Isolation

8 articles
Cybersecurity 18 Sep 2026 5 min read

Seccomp User Notification Delegates Syscall Execution Across a Privilege Boundary

A confined process can reach a syscall that the kernel would reject under its current credentials, while a separate supervisor has enough privilege to perform an equivalent operation safely on its behalf. Linux seccomp user notification creates a mediation channel for that arrangement: a filter can stop the calling thread, emit a notification to a listener, and wait for a userspace response. That mechanism is more precise than treating the supervisor as a general syscall proxy. The notification carries register-level syscall data and an identifier tied to the pending request. The supervisor can synthesize a return value, inject a file descriptor, or in selected cases tell the kernel to continue the original syscall. Each option places the trust boundary in a different location.

Cybersecurity 18 Sep 2026 6 min read

process_vm_readv Crosses Process Memory Behind ptrace Access Checks

A diagnostic agent may need bytes from another process without stopping that process or attaching a traditional debugger. Linux process_vm_readv() provides that data path: the caller supplies local buffers and address ranges in a target process, and the kernel transfers bytes between the two address spaces. The interface is powerful because the target does not explicitly send the data. Its security boundary therefore sits outside the target’s application protocol. Linux gates the operation with a ptrace access-mode check, while the memory transfer itself remains subject to the target’s changing virtual-memory layout.

Cybersecurity 18 Sep 2026 5 min read

PR_SET_DUMPABLE Changes Linux Process Inspection Boundaries

A service receives credentials into process memory, drops privileges, and continues running under an ordinary account. Another same-account process may still be able to inspect it through interfaces intended for debugging. Linux places an additional process attribute, commonly called dumpable, into several of these access decisions. prctl(PR_SET_DUMPABLE, 0) marks the calling process non-dumpable. The effect is broader than suppressing a core file: Linux also incorporates dumpable state into ptrace access checks and changes ownership behavior for files under /proc/<pid>. These effects form related boundaries, but they are not a single universal ban on process observation.

Cybersecurity 18 Sep 2026 6 min read

pidfds Bind Process Operations to Stable Kernel References

A supervisor records PID 4127, performs unrelated work, then sends a signal to 4127. Between those steps, the original process can exit and the kernel can eventually assign the same numeric PID to another process. The integer still names a process, but not necessarily the process that the supervisor intended to affect. Linux PID file descriptors, commonly called pidfds, move that boundary from repeated numeric lookup to a file descriptor that refers to a particular task. That change is narrow but security-relevant: operations that accept a pidfd can stay bound to the task selected when the reference was acquired rather than resolving a reusable number again.

Cybersecurity 18 Sep 2026 6 min read

Pidfd Process References Separate Identity from Numeric PIDs

A supervisor records a worker PID, performs unrelated work, then sends a signal to that number. If the original worker exited and the kernel reused its numeric PID, a later operation can address a different process. The number identifies an entry in a PID namespace at a moment in time; it is not, by itself, a durable process handle. Linux pidfds add a file-descriptor representation of process identity. A pidfd obtained for a process continues to refer to that process rather than being retargeted when its numeric PID is recycled. This changes the identity boundary for supervision, but it does not grant broad authority over the referenced process.

Cybersecurity 18 Sep 2026 6 min read

close_range Narrows File Descriptor Inheritance Before exec

A service process can accumulate sockets, pipes, directory handles, log files, and control descriptors long before it launches a helper. If those descriptors survive into the new program, the helper receives capabilities that its command-line arguments and environment do not reveal. A connected socket can carry authenticated access; an open directory can preserve reachability to a filesystem location; a pipe can expose another component’s data path. Linux close_range() gives pre-exec code a range operation over file descriptors. Its security value is not that descriptors become harmless. It is that a process can narrow the descriptor set that crosses an execve() boundary without enumerating /proc/self/fd or issuing one close() call per candidate descriptor.

Cybersecurity 17 Sep 2026 5 min read

Close-on-Exec Makes Descriptor Inheritance an Explicit Boundary

A service opens a privileged socket, starts helper programs, and expects those helpers to receive only standard input, output, and error. One descriptor created without close-on-exec can quietly violate that boundary. If it remains present when a new program image is installed, the helper inherits access to the kernel object even when its own credentials could never have opened that object. Linux treats this as descriptor inheritance, not a new authorization event. The security decision made when the object was opened is embodied in the descriptor. FD_CLOEXEC controls whether that established authority crosses a successful execve().

Cybersecurity 17 Sep 2026 5 min read

close_range with UNSHARE Detaches Descriptor Tables Before Bulk Closure

A multithreaded Linux process can reach an awkward boundary just before execve(): one thread wants to discard every file descriptor above standard input, output, and error, while another thread can still create descriptors in the same table. A loop of close() calls treats descriptor numbers individually, but it does not by itself change the fact that the table is shared. close_range() with CLOSE_RANGE_UNSHARE addresses that specific race. The kernel first gives the caller a file descriptor table that is no longer shared with the other users of the old table, then applies the requested bulk closure to the caller’s table. The security property is about table ownership during cleanup, not merely fewer system calls.