Skip to content

Archive

Containers

5 articles
Linux 19 Sep 2026 5 min read

seccomp User Notifications Move Selected System Calls to a Supervisor

seccomp User Notifications Move Selected System Calls to a Supervisor A seccomp filter can do more than allow a system call or reject it in the kernel. When a filter returns SECCOMP_RET_USER_NOTIF, Linux can suspend the calling thread and deliver a description of that attempted system call to a user-space supervisor. The mechanism creates an interposition boundary around selected calls. It is useful when a less-privileged process needs an operation mediated by another process, such as a container manager handling a call that the container cannot perform directly. The boundary is deliberately narrower than a general security-policy engine: notification state can race with mutable target memory, and the kernel documentation warns against treating the supervisor’s inspection as an authorization primitive.

Linux 19 Sep 2026 6 min read

Idmapped Mounts Remap Ownership Without Rewriting Inodes

The same inode can appear with different ownership through two mount points without any recursive chown(). Linux idmapped mounts attach an ID mapping to a mount, so VFS ownership presentation and permission checks can translate user and group IDs for that view while the ownership stored by the filesystem remains unchanged. This property separates persistent inode metadata from the identity view exposed at a particular mount. It is especially useful when a filesystem tree must be shared with a container whose user namespace maps IDs differently from the host.

Linux 19 Sep 2026 4 min read

CLONE_INTO_CGROUP Places a Child in Its Target cgroup at Creation

CLONE_INTO_CGROUP Places a Child in Its Target cgroup at Creation A process created in one cgroup and moved to another has a short but real interval in the original cgroup. During that interval, accounting, resource controls, and freezer state come from the initial placement rather than the destination. Linux provides CLONE_INTO_CGROUP so clone3() can place the child in a cgroup v2 target as part of process creation. This changes the placement boundary. Instead of creating a task and repairing its cgroup membership afterward, the caller identifies the destination before the child exists.

Software Engineering 18 Sep 2026 6 min read

seccomp User Notification Delegates Selected Syscalls to a Supervisor

seccomp User Notification Delegates Selected Syscalls to a Supervisor A seccomp filter can stop a selected system call before the kernel executes it and emit a notification to a user-space supervisor instead. The target thread remains blocked while the supervisor receives the event and returns a disposition. This behavior turns a filter result into a controlled handoff across the kernel/user-space boundary. The mechanism is SECCOMP_RET_USER_NOTIF. It differs from ordinary seccomp actions because the BPF filter does not finish the decision by itself. A listener file descriptor becomes the coordination point for notification receipt, response delivery, and optional file-descriptor injection.

Linux 02 Sep 2026 5 min read

Inspect Process Isolation with Linux Namespaces

Containers rely on several Linux kernel features, but namespaces provide much of the process-level isolation people notice first. They let different groups of processes see different views of resources such as process IDs, mounts, hostnames, and network interfaces. You do not need a container runtime to inspect namespaces. Standard Linux tools and /proc expose the relationships directly. What a namespace changes A namespace virtualizes one class of global system resource.