Skip to content

Archive

Concurrency

172 articles
Linux 23 Sep 2026 4 min read

signalfd Consumes Blocked Signals Through File Descriptor Reads

signalfd Consumes Blocked Signals Through File Descriptor Reads signalfd() turns selected signal notifications into records that can be consumed with read(). A matching pending signal makes the descriptor readable, which allows signal handling to share the same poll(), select(), or epoll path as sockets, timers, and other descriptors. The descriptor does not redirect signals merely because its mask names them. Normal signal-mask rules still apply. In the usual design, the selected signals are blocked before they can be delivered through their ordinary dispositions, then a signalfd reads the pending instances.

Software Engineering 23 Sep 2026 5 min read

Priority Inheritance Bounds Priority Inversion Around Mutexes

Priority Inheritance Bounds Priority Inversion Around Mutexes Priority scheduling does not guarantee that the highest-priority runnable task can always make progress. A high-priority task can block on a mutex held by a lower-priority task. If medium-priority work then preempts the owner, the high-priority task remains blocked even though the medium-priority work has no direct dependency on the mutex. This is priority inversion. The inversion starts with an ordinary dependency: the high-priority task needs a resource owned by a lower-priority task. The damaging part is interference from tasks between those priorities, which can delay the owner and extend the blocking interval.

Linux 23 Sep 2026 5 min read

membarrier Private Expedited Orders Userspace Memory Across Threads

Linux membarrier() can move part of a synchronization cost from a frequently executed path to a less frequent coordination path. With MEMBARRIER_CMD_PRIVATE_EXPEDITED, one thread asks the kernel to establish a memory-ordering point across the running threads in the same process. The caller pays for the system call when coordination is needed instead of requiring every fast-path execution to carry an explicit hardware memory barrier. This is narrower than a general thread rendezvous. membarrier() does not run an application callback on sibling threads, does not wait for application-level acknowledgements, and does not turn ordinary data races into valid synchronization. Its contract concerns ordering of userspace memory accesses around the barrier.

Software Engineering 23 Sep 2026 6 min read

Lock Convoys Turn Short Critical Sections into Long Queues

Lock Convoys Turn Short Critical Sections into Long Queues A mutex can protect a tiny critical section and still become the center of a large latency problem. The code inside the lock may take only microseconds in normal operation, yet one delayed holder can allow several threads to accumulate behind it. Once that queue exists, the lock may remain continuously contended as ownership passes from one waiting thread to another.

Tech 23 Sep 2026 5 min read

Linux RCU Grace Periods Separate Removal from Reclamation

Read-Copy Update, commonly called RCU, is a synchronization family used throughout the Linux kernel for data that is read frequently and changed less often. Its central move is to separate two events that ordinary locking often treats as one: removing an object from a shared structure and reclaiming the memory that stored it. A writer can publish a new state or unlink an old object while readers continue through read-side critical sections. The old storage remains valid until the kernel has established that every reader that could have observed the old reference has passed through a quiescent state. That interval is an RCU grace period.

Software Engineering 23 Sep 2026 6 min read

Hazard Pointers Delay Reclamation Until Readers Release References

Removing a node from a lock-free data structure does not make its memory immediately safe to reuse. Another thread may already hold the node’s address and may still dereference it. If the remover frees that allocation too early, an otherwise correct atomic update can be followed by a use-after-free. Hazard pointers separate logical removal from physical reclamation. A reader publishes the address it intends to access in a designated hazard slot. A remover can unlink a node and place it on a retired list, but reclamation waits until a scan confirms that no hazard slot protects that address.

Linux 23 Sep 2026 5 min read

fcntl OFD Locks Follow Open File Descriptions Instead of Processes

Traditional POSIX record locks have a property that can surprise code with several descriptors for the same file: lock ownership is associated with the process, and closing a descriptor for that file can release the process’s locks on it. Linux open file description locks move ownership to the kernel open file description referenced by a descriptor. The interface still uses fcntl() and struct flock, but the ownership boundary changes. F_OFD_SETLK, F_OFD_SETLKW, and F_OFD_GETLK make byte-range locking track the open file description rather than the process identity.

Linux 23 Sep 2026 5 min read

eventfd Semaphore Mode Turns Counter Values into Single-Unit Reads

Linux eventfd exposes a kernel-maintained 64-bit counter through a file descriptor. Its compact interface hides an important semantic choice: a normal read drains the current counter value, while an EFD_SEMAPHORE read consumes exactly one unit. The write path still adds values to the same counter. That distinction changes the object from an aggregate notification counter into a descriptor-backed source of individually consumable units. The readiness model remains compatible with poll, epoll, and related descriptor multiplexing, so the same object can connect producer accounting with an event loop without adding a separate pipe payload.

Linux 23 Sep 2026 4 min read

EPOLLEXCLUSIVE Limits Wakeups Across epoll Instances

EPOLLEXCLUSIVE Limits Wakeups Across epoll Instances EPOLLEXCLUSIVE changes event distribution when several epoll instances register the same target file description. Without the flag, a readiness event can wake waiters associated with every interested epoll instance. With exclusive registrations, Linux can restrict that fan-out and reduce redundant wakeups. The flag does not make delivery single-consumer, does not assign ownership of the target, and does not replace application-level coordination. Its contract is narrower: among epoll instances that registered a target with EPOLLEXCLUSIVE, one or more receive an event for a wakeup rather than requiring all of them to receive it.

Tech 23 Sep 2026 7 min read

epoll Readiness Modes Change How Event Loops Drain File Descriptors

Linux epoll lets one thread wait on readiness changes across many file descriptors without scanning every descriptor on each iteration. The interface is common in network servers, proxies, runtimes, and other programs that keep large sets of sockets active. The registration mode matters. Level-triggered operation keeps reporting a descriptor while the relevant condition remains ready. Edge-triggered operation reports transitions in readiness and expects the application to consume available work until the descriptor would block.

Linux 23 Sep 2026 5 min read

close_range with UNSHARE Separates File Descriptor Cleanup from Peer Threads

A multithreaded Linux process can share one file descriptor table across its threads. That arrangement is convenient during normal execution: a descriptor opened by one thread becomes available to its peers. It becomes less convenient when one thread is preparing a restricted execution context and wants to discard a broad descriptor range without racing with peers that can still allocate descriptors. close_range() provides a range operation for this boundary. With CLOSE_RANGE_UNSHARE, the kernel first separates the caller from the shared descriptor table and then applies the requested closure to the caller’s resulting table. The operation is close to combining unshare(CLONE_FILES) with a range close, but the kernel can perform less work in common cases.

Linux 22 Sep 2026 5 min read

TFD_TIMER_CANCEL_ON_SET Reports Realtime Clock Discontinuities

A timerfd armed against wall-clock time can cross a discontinuous clock adjustment before its deadline. With TFD_TIMER_CANCEL_ON_SET, Linux exposes that event through the descriptor: a subsequent read() fails with ECANCELED rather than presenting the clock jump as an ordinary timer expiration. The flag is deliberately narrow. It applies only when timerfd_settime() arms an absolute timer on CLOCK_REALTIME or CLOCK_REALTIME_ALARM, and it changes the handling of discontinuous changes to those clocks.

Software Engineering 22 Sep 2026 7 min read

Request Sequence Guards Reject Late UI Responses

Request Sequence Guards Reject Late UI Responses Interactive interfaces often issue a new request before the previous one has finished. Search boxes, filters, route changes, autocomplete fields, and detail panels all create this pattern. Network completion order is not guaranteed to match the order in which the user changed state. That mismatch can produce a subtle race. Request A starts first, request B starts second, B completes first, and the interface renders B. If A then completes and its callback writes without checking context, the screen moves backward to data associated with the older intent.

Software Engineering 22 Sep 2026 7 min read

Request Coalescing Stops Cache Misses from Multiplying Backend Work

Request Coalescing Stops Cache Misses from Multiplying Backend Work A cache miss is usually cheap when one caller causes one backend lookup. The same miss can become expensive when many callers arrive for the same key at nearly the same time. Each caller observes the key as absent, each starts identical work, and the backend receives a burst precisely when the cache is providing no protection for that key. Request coalescing changes that concurrency pattern. The first caller becomes the leader for a key. Later callers join the same in-flight operation and wait for its result rather than starting equivalent work. Once the fill completes, the result can populate the cache and be returned to the waiting callers.

Software Engineering 22 Sep 2026 8 min read

Optimistic Concurrency Rejects Stale Writes Before They Replace Newer State

Optimistic Concurrency Rejects Stale Writes Before They Replace Newer State A read-modify-write flow looks harmless when only one actor touches a record. A client reads state, changes part of it, then writes the result back. With concurrent actors, the interval between the read and the write becomes a race. Another writer can commit a newer value during that interval, and an unconditional update can erase it. Optimistic concurrency control puts a condition on the final write. The client carries a version derived from the state it read, and the storage layer accepts the mutation only if that version is still current. A mismatch becomes a conflict rather than a silent overwrite.

Software Engineering 22 Sep 2026 6 min read

Optimistic Concurrency Control Rejects Stale Writes

Optimistic Concurrency Control Rejects Stale Writes Two clients can read the same record, make different edits, and save seconds apart. If each update blindly replaces the stored value, the later write can erase the earlier one even though both requests succeeded. Optimistic concurrency control prevents that silent overwrite by attaching a condition to the write. The client records a version when it reads the data. Its update succeeds only if that version is still current. A changed version turns the write into a conflict instead of an unnoticed loss.

Software Engineering 22 Sep 2026 6 min read

Fencing Tokens Stop Stale Lease Holders

Fencing Tokens Stop Stale Lease Holders A distributed lease gives one worker temporary permission to act as an owner. The lease eventually expires so another worker can take over after a crash or network failure. That solves availability, but expiry alone does not guarantee that the old worker has stopped. A process can pause long enough for its lease to expire, then resume with stale local state. A long garbage-collection pause, scheduler stall, suspended virtual machine, or delayed network path can create this condition. If the old worker writes after a replacement has taken ownership, two workers can affect the same resource even though the lease service never considered both leases valid at the same instant.

Software Engineering 22 Sep 2026 7 min read

Fencing Tokens Block Stale Lock Holders

Fencing Tokens Block Stale Lock Holders A distributed lock is often used to keep two workers from changing the same resource at once. The difficult case begins when lock ownership depends on a lease. A client can acquire the lease, pause long enough for it to expire, then resume after another client has acquired a new lease. From the old client’s point of view, execution simply continued. From the coordination service’s point of view, ownership already moved. If the protected storage system accepts both clients’ writes, the old holder can overwrite work performed by the current holder.

Linux 22 Sep 2026 5 min read

EFD_SEMAPHORE Makes eventfd Reads Consume One Counter Unit

An eventfd normally turns its entire nonzero counter into one read result and resets the counter to zero. Creating it with EFD_SEMAPHORE changes only the read side: each successful read returns the 64-bit value 1 and subtracts one from the kernel-maintained counter. That difference lets several units accumulated by writers remain separately consumable. The object is still an eventfd, with the same counter, write rules, descriptor lifetime, and readiness integration.

Software Engineering 22 Sep 2026 6 min read

Bulkheads Keep One Saturated Dependency from Consuming Every Worker

Bulkheads Keep One Saturated Dependency from Consuming Every Worker A service can have healthy CPU, available memory, and responsive internal code while still becoming unavailable. One downstream dependency is enough to consume the service’s entire concurrency budget if calls to it become slow and every request is allowed to wait. The failure is not limited to the slow dependency. Shared worker pools, connection pools, semaphores, queues, and request slots turn local saturation into a service-wide outage. Bulkhead isolation limits that blast radius by reserving separate capacity for distinct workloads or dependencies.

Software Engineering 22 Sep 2026 5 min read

Bulkheads Isolate Concurrency Before One Dependency Consumes It All

Bulkheads Isolate Concurrency Before One Dependency Consumes It All A service can have enough CPU and memory yet stop making useful progress because its concurrency is exhausted. Threads, database connections, outbound sockets, worker slots, and in-flight request permits are finite. If one dependency becomes slow, calls to that dependency can occupy the entire shared pool. Bulkhead isolation divides that capacity before saturation occurs. Workloads that can fail independently receive separate concurrency budgets, so pressure in one path does not automatically consume every slot needed by another.

Software Engineering 22 Sep 2026 6 min read

Backpressure Keeps Fast Producers from Overrunning Slow Consumers

Backpressure Keeps Fast Producers from Overrunning Slow Consumers A pipeline is stable only while work leaves each stage at roughly the rate it arrives over a useful time window. When a producer can submit work faster than a consumer can finish it, the difference has to accumulate somewhere. An unbounded queue makes that accumulation easy to miss. Requests continue to be accepted, the producer appears healthy, and the consumer keeps working. Meanwhile queued work consumes memory and ages before execution. Backpressure turns downstream saturation into an upstream signal before the backlog becomes the failure.

Software Engineering 21 Sep 2026 5 min read

Write Skew Breaks Cross-Row Invariants Under Snapshot Isolation

Write Skew Breaks Cross-Row Invariants Under Snapshot Isolation Snapshot isolation gives each transaction a stable database view and commonly prevents concurrent transactions from committing conflicting writes to the same row. That is a strong concurrency property, but it does not make every application invariant serializable. Write skew appears when two transactions read overlapping state, make decisions from the same valid snapshot, then write different records. Since their write sets do not collide, both commits may succeed. The combined state can violate a rule that each transaction checked before writing.

Software Engineering 21 Sep 2026 4 min read

Version Vectors Distinguish Concurrent Updates from Causal Successors

Version Vectors Distinguish Concurrent Updates from Causal Successors Replicated data can receive writes at different nodes while communication between those nodes is delayed. When versions later meet, a scalar revision number can say that two values differ, but it cannot always say whether one descends from the other or both were produced independently. A version vector records progress per replica. Comparing those counters provides a partial order: one version can dominate another, the vectors can be equal, or neither can dominate. The last case identifies concurrent histories that require an explicit reconciliation rule.