A single ready socket can wake several threads when each thread waits on a different epoll instance that watches that socket. Linux provides EPOLLEXCLUSIVE to narrow that wakeup fan-out: among epoll instances that registered the target with the flag, a readiness event wakes one or more rather than all of them.
The distinction is deliberately weaker than “exactly one waiter.” EPOLLEXCLUSIVE changes notification selection across epoll instances. It does not transfer ownership of the file descriptor, serialize all I/O, or guarantee that only one thread can observe useful work.
The flag belongs to a registration
An epoll instance maintains an interest list of watched file descriptors. A target is added with epoll_ctl() and an event mask:
struct epoll_event ev = {
.events = EPOLLIN | EPOLLEXCLUSIVE,
.data.fd = listener,
};
if (epoll_ctl(epfd, EPOLL_CTL_ADD, listener, &ev) == -1)
perror("epoll_ctl");Here the exclusive property applies to the relationship between epfd and listener. Another epoll instance can register the same target independently.
Without EPOLLEXCLUSIVE, when multiple epoll instances are attached to the same target and readiness occurs, all of those epoll instances can receive the event. With exclusive registrations, Linux limits which exclusive instances are awakened. The documented contract is “one or more,” so code must not depend on an exactly-one rule.
This behavior targets the thundering-herd pattern in which many blocked workers become runnable for one readiness transition and most then find no useful work.
Exclusive and ordinary registrations can coexist
Exclusivity does not suppress nonexclusive observers. If a target is registered in several epoll instances, some with EPOLLEXCLUSIVE and some without it, an event is delivered to all nonexclusive instances and at least one exclusive instance.
That property matters for designs that combine worker dispatch with monitoring. An ordinary epoll registration used by a diagnostic or control component still receives readiness even when worker epoll instances use exclusive registrations.
The resulting model is not a global lock on the target:
target fd
|-- ordinary epoll A -> receives event
|-- exclusive epoll B -> eligible exclusive receiver
|-- exclusive epoll C -> eligible exclusive receiver
`-- exclusive epoll D -> eligible exclusive receiverThe kernel narrows the exclusive branch while leaving ordinary registrations outside that selection.
Registration constraints are part of the API
EPOLLEXCLUSIVE is accepted only with EPOLL_CTL_ADD. It cannot be introduced later with EPOLL_CTL_MOD. Once an epfd and target pair was added with EPOLLEXCLUSIVE, a later EPOLL_CTL_MOD for that pair also fails with EINVAL.
The event-mask combinations are restricted. EPOLLIN, EPOLLOUT, EPOLLWAKEUP, and EPOLLET may be specified with EPOLLEXCLUSIVE. EPOLLHUP and EPOLLERR are reported as usual and need not be requested. Other event flags combined with EPOLLEXCLUSIVE can produce EINVAL.
An epoll file descriptor itself cannot be the target of an exclusive registration. Attempting that combination also fails with EINVAL.
These constraints make exclusivity a property selected when the watch is created, rather than a mode intended for arbitrary runtime toggling.
Wakeup selection does not serialize consumption
Suppose several worker threads each own an epoll instance, and each instance registers the same listening socket with EPOLLIN | EPOLLEXCLUSIVE. When a connection arrives, the kernel can avoid waking every blocked epoll waiter.
After a worker wakes, normal socket semantics still apply. Another thread may call accept() through a shared descriptor, more than one connection may already be queued, and another readiness transition may occur while processing continues. EPOLLEXCLUSIVE does not establish a critical section around accept().
The same boundary applies to other descriptor types. Readiness says that an operation could proceed at the time the readiness state was evaluated; it is not a reservation of data for the selected epoll waiter.
Applications that require ownership, ordering, or mutual exclusion need a separate mechanism or a topology that assigns descriptors to workers explicitly.
Edge triggering remains a separate dimension
EPOLLET can be combined with EPOLLEXCLUSIVE, but the flags address different behavior. EPOLLET controls edge-triggered delivery semantics for a watched target. EPOLLEXCLUSIVE controls wakeup selection when competing epoll instances watch that target.
Combining them therefore does not turn exclusive wakeup into exclusive access. Edge-triggered code still has to respect the normal requirement to drain available work appropriately for its I/O mode and event loop design.
Keeping these dimensions separate avoids attributing missed progress to the wakeup-selection flag when the actual issue is edge-triggered consumption logic.
The optimization boundary is competing epoll instances
EPOLLEXCLUSIVE is specifically useful when the same target is watched by multiple epoll instances and redundant wakeups are undesirable. It is not a general scheduler hint and does not promise fair rotation among workers.
The API contract also permits more than one exclusive epoll instance to receive an event. That latitude is important: the guarantee is reduction of wakeup fan-out relative to the ordinary all-observers behavior, not deterministic dispatch to one worker.
For a server architecture, this places the flag at a precise boundary. It can reduce contention caused by waking many event loops for shared readiness, while connection assignment, load distribution, request ordering, and application-level synchronization remain separate concerns.