Skip to content

Archive

Linux Security

25 articles
Cybersecurity 18 Sep 2026 7 min read

Seccomp User Notification Moves Selected System Calls Behind a Supervisor Decision

A sandboxed process may need an operation that cannot be represented safely as a permanent seccomp allow rule. The operation can depend on runtime policy, external state, or a resource that only a more privileged component should inspect. Allowing the system call unconditionally widens the sandbox, while rejecting it removes required functionality. Linux seccomp user notification provides a mediation point for this case. A seccomp filter can return SECCOMP_RET_USER_NOTIF for selected calls. The kernel then blocks the triggering task and emits a notification through a listener file descriptor. A supervisor reads that notification and sends a response that determines the immediate disposition of the intercepted call.

Cybersecurity 18 Sep 2026 5 min read

Seccomp User Notification Delegates Syscall Execution Across a Privilege Boundary

A confined process can reach a syscall that the kernel would reject under its current credentials, while a separate supervisor has enough privilege to perform an equivalent operation safely on its behalf. Linux seccomp user notification creates a mediation channel for that arrangement: a filter can stop the calling thread, emit a notification to a listener, and wait for a userspace response. That mechanism is more precise than treating the supervisor as a general syscall proxy. The notification carries register-level syscall data and an identifier tied to the pending request. The supervisor can synthesize a return value, inject a file descriptor, or in selected cases tell the kernel to continue the original syscall. Each option places the trust boundary in a different location.

Cybersecurity 18 Sep 2026 6 min read

process_vm_readv Crosses Process Memory Behind ptrace Access Checks

A diagnostic agent may need bytes from another process without stopping that process or attaching a traditional debugger. Linux process_vm_readv() provides that data path: the caller supplies local buffers and address ranges in a target process, and the kernel transfers bytes between the two address spaces. The interface is powerful because the target does not explicitly send the data. Its security boundary therefore sits outside the target’s application protocol. Linux gates the operation with a ptrace access-mode check, while the memory transfer itself remains subject to the target’s changing virtual-memory layout.

Cybersecurity 18 Sep 2026 5 min read

PR_SET_DUMPABLE Changes Linux Process Inspection Boundaries

A service receives credentials into process memory, drops privileges, and continues running under an ordinary account. Another same-account process may still be able to inspect it through interfaces intended for debugging. Linux places an additional process attribute, commonly called dumpable, into several of these access decisions. prctl(PR_SET_DUMPABLE, 0) marks the calling process non-dumpable. The effect is broader than suppressing a core file: Linux also incorporates dumpable state into ptrace access checks and changes ownership behavior for files under /proc/<pid>. These effects form related boundaries, but they are not a single universal ban on process observation.

Cybersecurity 18 Sep 2026 6 min read

Pidfd Process References Separate Identity from Numeric PIDs

A supervisor records a worker PID, performs unrelated work, then sends a signal to that number. If the original worker exited and the kernel reused its numeric PID, a later operation can address a different process. The number identifies an entry in a PID namespace at a moment in time; it is not, by itself, a durable process handle. Linux pidfds add a file-descriptor representation of process identity. A pidfd obtained for a process continues to refer to that process rather than being retargeted when its numeric PID is recycled. This changes the identity boundary for supervision, but it does not grant broad authority over the referenced process.

Cybersecurity 18 Sep 2026 7 min read

OverlayFS Stashed Credentials Separate Overlay Access from Backing Filesystem Access

OverlayFS Stashed Credentials Separate Overlay Access from Backing Filesystem Access A process opens a path through an OverlayFS mount and appears to access one ordinary filesystem object. The kernel may actually consult an upper layer, a lower layer, or both, and a write can trigger copy-up before the requested operation proceeds. That indirection creates an authorization problem: the caller must be permitted to use the object as exposed by the overlay, while the internal access to the backing filesystems must also run under a defined security identity.

Cybersecurity 18 Sep 2026 6 min read

openat2 Makes Path Resolution an Explicit Security Boundary

A privileged service may accept a relative pathname from a less trusted component while intending to access only files below a designated directory. Checking the string for .., rejecting an initial slash, or inspecting symbolic links before a later open() does not bind the check to the kernel lookup that acquires the file. Directory entries can change between operations, symbolic links can redirect traversal, and mount topology can alter the namespace reached by a path.

Cybersecurity 18 Sep 2026 6 min read

Mount Propagation Defines the Filesystem Boundary Between Linux Mount Namespaces

A process can enter a new Linux mount namespace and still observe a later mount created elsewhere. The namespace boundary is intact: the process has its own mount table. The new mount appears because some mounts in the two namespaces remain connected by propagation relationships. This distinction matters in container runtimes, service sandboxes, build systems, and privileged helpers. Creating a mount namespace separates the namespace’s view of the mount table, but it does not by itself make every future mount event local. Shared-subtree state determines whether mount and unmount events cross that boundary.

Cybersecurity 18 Sep 2026 4 min read

Memfd Seals Turn Mutable Anonymous Files into Explicit Handoff Objects

A process prepares a binary payload in memory, passes a file descriptor to another process, and expects the bytes to remain stable after validation. A plain descriptor does not create that guarantee. If some holder still has write authority, the object can change after a consumer has inspected it, and pathname permissions offer no useful boundary when the object has no ordinary filesystem name. Linux memfd_create() provides an anonymous file backed by memory-like filesystem storage, and file seals can constrain later changes to that file. The useful security property is not anonymity by itself. It is the ability to construct a mutable object, apply irreversible restrictions to that object, then hand out descriptors whose backing file can no longer be changed in the prohibited ways.

Cybersecurity 18 Sep 2026 5 min read

MADV_DONTDUMP Excludes Selected Memory Mappings from Linux Core Images

A long-running service may keep credentials, session material, or decrypted state in memory while still relying on core images for crash diagnosis. Disabling core generation for the entire process removes diagnostic state along with sensitive state. Linux provides a narrower control: madvise() with MADV_DONTDUMP marks selected mappings so the kernel omits them from a core image. This mechanism changes core-dump inclusion policy for an address range. It does not make the bytes inaccessible to the process, encrypt them, erase them, or create a general barrier against process inspection. Its security value is specific to one data-exposure path: memory captured through the kernel core-dump mechanism.

Cybersecurity 18 Sep 2026 6 min read

Landlock Rulesets Add a Process-Scoped Filesystem Access Boundary

A service can begin with ordinary filesystem permissions that are broader than the files it needs during steady-state operation. Changing ownership or mount topology may be impractical because the same host resources are shared with other processes. Linux Landlock addresses this gap by letting a process add a kernel-enforced access restriction to itself and, through inheritance, to descendants. Landlock is a Linux Security Module designed for sandboxing. Its rules do not grant filesystem access that DAC, ACLs, capabilities, or another security mechanism would otherwise deny. They add another authorization layer. An operation succeeds only when the other applicable controls and the Landlock policy permit it.

Cybersecurity 18 Sep 2026 6 min read

io_uring Restrictions Freeze an Allowed Operation Surface Before Ring Activation

A service can expose an io_uring instance to code that should perform only a narrow class of asynchronous operations. The ring itself, however, supports many submission opcodes and registration commands. Relying only on application code to avoid unwanted operations leaves the allowed surface as a convention rather than a kernel-enforced property. Linux provides a tighter mechanism through IORING_REGISTER_RESTRICTIONS. A ring created with IORING_SETUP_R_DISABLED can receive a restriction set before it becomes usable for submissions. The process then enables the ring with IORING_REGISTER_ENABLE_RINGS. From that point, the kernel evaluates operations against the registered restrictions.

Cybersecurity 18 Sep 2026 4 min read

Idmapped Mounts Remap File Ownership Without Rewriting Inodes

Idmapped Mounts Remap File Ownership Without Rewriting Inodes A container needs read-write access to a directory whose files carry host ownership values that do not line up with the container’s user namespace. Recursively changing ownership can make the directory usable, but it also mutates persistent inode metadata and can disrupt every other view of the same filesystem. Linux idmapped mounts provide a narrower mechanism: one mount can apply a different identity mapping while the stored ownership remains intact.

Cybersecurity 18 Sep 2026 6 min read

close_range Narrows File Descriptor Inheritance Before exec

A service process can accumulate sockets, pipes, directory handles, log files, and control descriptors long before it launches a helper. If those descriptors survive into the new program, the helper receives capabilities that its command-line arguments and environment do not reveal. A connected socket can carry authenticated access; an open directory can preserve reachability to a filesystem location; a pipe can expose another component’s data path. Linux close_range() gives pre-exec code a range operation over file descriptors. Its security value is not that descriptors become harmless. It is that a process can narrow the descriptor set that crosses an execve() boundary without enumerating /proc/self/fd or issuing one close() call per candidate descriptor.

Cybersecurity 17 Sep 2026 5 min read

Userfaultfd Moves Page-Fault Resolution Into Userspace

Userfaultfd Moves Page-Fault Resolution Into Userspace A thread touches a registered virtual-memory page and stops before the access completes. Instead of resolving the fault entirely inside the kernel, Linux can report the event through a userfaultfd and let another userspace component decide when and with what content execution may continue. That design supports live migration, post-copy memory transfer, checkpointing, and related memory-management systems, but it also places a concurrency-sensitive decision point outside the faulting thread.

Cybersecurity 17 Sep 2026 6 min read

Seccomp User Notification Moves Selected Syscall Decisions to a Supervisor

A confined process issues a system call that its ordinary seccomp policy cannot safely reduce to a static allow-or-deny decision. The arguments may refer to mutable process memory, or the operation may need privileged work performed outside the confined process. Returning a fixed errno is too restrictive, while permitting the call directly gives the target more authority than the deployment intends. Linux seccomp user notification creates a mediation path for this case. A filter can return SECCOMP_RET_USER_NOTIF, causing the kernel to block the triggering task and emit a request on a listener file descriptor. A userspace supervisor receives the request and later supplies a result. This mechanism changes where a selected syscall decision is made, but it does not turn seccomp into a general reference monitor without additional policy and race controls.

Software Engineering 17 Sep 2026 6 min read

Seccomp User Notification Delegates Selected System Calls to a Supervisor

A system call selected by a seccomp filter can stop before kernel execution and appear as a request on a notification file descriptor. With SECCOMP_RET_USER_NOTIF, Linux turns that call into a coordination point between the blocked target and a userspace supervisor. The supervisor can emulate a result, inject a file descriptor for suitable operations, or permit the kernel to continue the original call. This mechanism is deliberately narrower than a general userspace security policy engine. Its strongest boundary is the kernel-mediated suspension and response protocol. Data reached through target-memory pointers can still change around a supervisor’s inspection, and a response that continues the original call re-enters ordinary kernel execution with that race still relevant.

Cybersecurity 17 Sep 2026 7 min read

Seccomp Filters Reduce Syscall Surface Without Forming a Complete Sandbox

Seccomp Filters Reduce Syscall Surface Without Forming a Complete Sandbox A service can run with a short seccomp allowlist and still retain broad authority through file descriptors, filesystem permissions, network endpoints, and credentials. The filter may sharply reduce the kernel interfaces reachable through system calls, yet the process can remain capable of damaging actions through operations that are explicitly allowed. This is the central boundary of seccomp: it filters syscall attempts; it does not define the full security policy of a process.

Cybersecurity 17 Sep 2026 5 min read

Pidfds Turn Process Identity Into a Stable Kernel Reference

A supervisor records PID 1842 for a worker, waits for an asynchronous event, then sends a signal. Between those operations the worker can exit, be reaped, and its numeric PID can later identify another process. The number still looks valid, but the identity it denotes has changed. Linux pidfds move this class of process control away from repeated numeric lookup. A PID file descriptor refers to a particular process, giving userspace a kernel-held reference that can be passed to interfaces such as pidfd_send_signal(), polling APIs, and, under additional permission checks, pidfd_getfd().

Cybersecurity 17 Sep 2026 6 min read

Memfd Seals Turn Shared Memory Into Monotonic File Policy

Memfd Seals Turn Shared Memory Into Monotonic File Policy A process prepares a binary object in memory, passes its file descriptor to another process, and expects the bytes to remain stable after validation. Ordinary shared memory does not provide that property by itself: another holder of writable authority can change the object after a check, resize it, or keep a writable mapping alive. Linux file seals provide a narrower contract. They remove selected mutation operations from a sealable file, and successfully added seals cannot later be removed.

Cybersecurity 17 Sep 2026 7 min read

Landlock Rulesets Restrict Future Path Access, Not Open File Authority

Landlock Rulesets Restrict Future Path Access, Not Open File Authority A process opens a writable configuration file, installs a restrictive Landlock ruleset, and then continues running code that should have access only to a small working directory. The later policy can block a fresh attempt to open that configuration path, yet the descriptor obtained before confinement remains usable. The filesystem view has narrowed, but authority already materialized as an open file has not vanished.

Cybersecurity 17 Sep 2026 6 min read

Landlock Rulesets Add Process-Local Filesystem Denial Boundaries

A service starts with ordinary filesystem access inherited from its credentials, loads configuration, opens several resources, then begins processing data that may be hostile. Changing UID or entering a container can alter the surrounding authority model, but neither action by itself expresses a narrow rule such as “from this point onward, new reads are limited to these hierarchies and writes are limited to that directory.” Linux Landlock provides a process-controlled restriction layer for this boundary. A process creates a ruleset, adds object rules, and enforces the ruleset on itself. The resulting Landlock domain is stacked with existing discretionary access control and other Linux Security Module decisions. Landlock can remove access that those mechanisms would otherwise permit; it does not grant access they deny.

Cybersecurity 17 Sep 2026 7 min read

fs-verity Makes File Data Integrity a Read-Time Property

A host may need to keep independently updated executables, packages, models, or data on a writable filesystem while still detecting modification of file contents after an artifact has been accepted. A one-time userspace hash can identify the bytes at one moment, but it does not make later reads depend on that measurement. The file may be opened again, pages may be evicted and reloaded, and storage below the page cache may return different data.

Cybersecurity 17 Sep 2026 5 min read

execveat Binds Program Execution to an Open File Reference

A launcher selects an executable from a directory, checks attributes or content, and then starts it. If selection and execution each resolve the pathname independently, a rename, symlink change, or directory replacement between those operations can make the executed object differ from the object that was checked. Linux execveat() can move that boundary from a second pathname lookup to an already acquired file reference. With AT_EMPTY_PATH, an empty pathname tells the kernel to execute the object referred to by dirfd. That descriptor may have been opened with O_PATH. The execution decision still passes through normal kernel permission and executable-format checks, but object selection no longer depends on resolving the original pathname again.