Skip to content

Archive

Sockets

14 articles
Tech 23 Sep 2026 5 min read

TCP Keepalive Probes Detect Silent Dead Peers

An established TCP connection can remain quiet for a long time. Silence alone does not mean either endpoint has failed: an application may simply have no data to exchange. That property is useful for long-lived sessions, but it also creates an operational problem when a peer disappears without sending FIN or RST. A machine can lose power, a network path can fail, or state in an intermediate device can vanish. The surviving endpoint may retain a socket that still appears established because no packet has arrived to prove otherwise. TCP keepalive provides an optional mechanism for testing such idle connections.

Tech 22 Sep 2026 6 min read

TCP TIME_WAIT Keeps Closed Connections Distinct from Delayed Segments

A TCP connection can finish exchanging application data and still leave a socket record behind. The familiar TIME_WAIT state is part of TCP’s close machinery, not evidence that a process forgot to close a descriptor. The endpoint that performs the active close commonly enters TIME_WAIT after the closing handshake. It keeps enough state for a bounded interval so late segments from the old connection cannot be confused with traffic from a later connection using the same endpoint identity.

Linux 22 Sep 2026 5 min read

SO_PEEK_OFF Gives UNIX Sockets a Stateful Peek Cursor

SO_PEEK_OFF Gives UNIX Sockets a Stateful Peek Cursor A normal recv(..., MSG_PEEK) examines data at the front of a socket receive queue without removing it. On Linux UNIX-domain sockets, SO_PEEK_OFF can replace that repeated front-of-queue view with a cursor that advances after each peek. The cursor is socket state. Peeking moves it forward, while consuming bytes from the front of the queue moves it backward by the amount removed. That interaction lets a process inspect successive queued regions without consuming them, while keeping the peek position tied to the remaining queue.

Linux 17 Sep 2026 5 min read

TCP_NODELAY Disables Nagle Coalescing on a Socket

A TCP socket can hold a small write instead of transmitting it immediately when earlier data remains unacknowledged. This behavior comes from Nagle coalescing: it limits the stream of small TCP segments by allowing outstanding data to influence transmission of newly queued bytes. On Linux, setting TCP_NODELAY disables that coalescing rule for the socket. Small writes become eligible for prompt transmission, subject to the rest of the TCP stack, congestion control, flow control, queue state, and device scheduling.

Software Engineering 17 Sep 2026 6 min read

SO_REUSEPORT Moves TCP Connection Distribution Into the Kernel

With SO_REUSEPORT, several Linux TCP sockets can listen on the same local address and port at the same time. Incoming connections are assigned to a member of that reuseport group before an application calls accept(). The application no longer needs one shared listening socket as the sole handoff point between the network stack and multiple workers. That changes more than bind eligibility. It moves connection distribution into the kernel and gives each listener its own socket identity and accept path. The resulting architecture has different queueing, lifecycle, and routing properties from a design in which many workers compete on one listening socket.

Software Engineering 17 Sep 2026 4 min read

SO_REUSEPORT Forms Kernel-Selected Socket Groups

Multiple Linux sockets can bind the same local address when every participating socket enables SO_REUSEPORT before bind(). Incoming traffic is then assigned to a member of the resulting reuseport group rather than delivered to every socket. The shared address is therefore a kernel selection boundary, not a broadcast endpoint. This behavior supports independent receive or accept loops without forcing all work through one listening descriptor. It also creates a distinct operational property: group membership and the selection policy determine which socket receives a packet or connection.

Linux 17 Sep 2026 6 min read

SO_REUSEPORT Distributes Traffic Across Socket Groups

SO_REUSEPORT changes a local endpoint from a single-socket binding into a socket group. On Linux, multiple TCP or UDP sockets can bind the same local address when every participating socket enables the option before bind() and the bind credentials satisfy the kernel’s reuse rules. That behavior is distinct from merely relaxing address-conflict checks. Incoming traffic must also be assigned to one member of the group. The resulting selection boundary affects listener architecture, queue isolation, process restarts, UDP flow placement, and any design that assumes a port maps to exactly one socket.

Linux 17 Sep 2026 4 min read

SO_RCVLOWAT Raises the Readability Threshold for Linux Sockets

A Linux socket with SO_RCVLOWAT set above one byte can have data queued while poll(), select(), or epoll still reports no normal readable readiness. Since Linux 2.6.28, those readiness interfaces respect the configured receive low-water mark. The option changes the threshold associated with normal receive readiness. It does not define message boundaries, reserve receive-buffer space, or guarantee that a later receive operation returns exactly the configured number of bytes. Readability can require more than one queued byte Socket receive readiness is usually observed with the default low-water mark of one byte. In that state, ordinary queued data is enough to satisfy the data-volume part of the readable condition.

Tech 16 Sep 2026 6 min read

TCP TIME-WAIT Preserves Closed Connection State

A TCP endpoint can finish an application’s close operation while the protocol still retains state for that connection. After an active close completes its FIN exchange, the endpoint normally enters TIME-WAIT instead of discarding the connection record immediately. That retained state has two jobs. It leaves the endpoint able to acknowledge a retransmitted final FIN, and it separates a closed connection from a later incarnation that could use the same local and remote addresses and ports.

Tech 16 Sep 2026 6 min read

TCP Keepalive Probes Test Idle Connections

A TCP connection can remain established while carrying no application data. That is valid behavior: an open connection does not need a continuous stream of packets to remain a TCP connection. Silence creates a practical problem when one endpoint disappears without completing the normal close sequence. A machine can lose power, a network path can fail, or state in an intermediate device can vanish. If the surviving endpoint has no data to send, ordinary retransmission logic has nothing to act on.

Tech 16 Sep 2026 5 min read

Nagle Algorithm Batches Small TCP Writes

TCP applications can issue writes much smaller than the network’s maximum segment size. Sending every tiny write as a separate segment can consume disproportionate header and processing overhead. The Nagle algorithm limits that pattern by allowing one small segment to remain in flight while later small writes wait for an acknowledgment or enough queued data to form a larger segment. This behavior reduces streams of tiny TCP segments. It can also add latency when an application expects each small write to leave immediately.

Linux 16 Sep 2026 5 min read

Linux TCP TIME_WAIT Retains Closed Connection State

A TCP socket can disappear from an application while the kernel still retains state for the closed connection. On Linux, the endpoint that completes the active close commonly enters TIME_WAIT, keeping enough protocol state to protect a later connection from delayed segments associated with the old one. This state is not evidence that a process forgot to close a file descriptor. The application-visible socket can already be gone. TIME_WAIT belongs to TCP’s connection-lifecycle machinery and persists independently of the process that initiated the close.

Linux 16 Sep 2026 4 min read

Linux TCP Autocorking Coalesces Consecutive Small Writes

A small TCP write does not always trigger an immediate packet transmission on Linux. With TCP autocorking enabled, the stack can defer a new small send when an earlier packet from the same flow is still waiting in a qdisc or device transmit queue, giving a following write a chance to join the pending data. The mechanism targets packet count rather than application-visible buffering semantics. A successful write() or sendmsg() still reports bytes accepted by the socket; autocorking influences when queued bytes advance into transmission.

Linux 05 Sep 2026 12 min read

Pass Open File Descriptors Between Linux Processes with SCM_RIGHTS

Processes often need to hand each other access to an already-open resource. A supervisor may accept a client connection and delegate it to a worker. A privileged helper may open a protected file, then give an unprivileged process access without revealing broader filesystem permissions. A service may create an anonymous in-memory file and transfer it to another process. Sending the integer value of a file descriptor does not solve this problem. File descriptor numbers are meaningful only inside one process’s descriptor table. Descriptor 7 in one process can refer to a socket while descriptor 7 in another process refers to an unrelated file.