Skip to content

Archive

Performance

83 articles
Software Engineering 23 Sep 2026 6 min read

Lock Convoys Turn Short Critical Sections into Long Queues

Lock Convoys Turn Short Critical Sections into Long Queues A mutex can protect a tiny critical section and still become the center of a large latency problem. The code inside the lock may take only microseconds in normal operation, yet one delayed holder can allow several threads to accumulate behind it. Once that queue exists, the lock may remain continuously contended as ownership passes from one waiting thread to another.

Tech 22 Sep 2026 7 min read

NUMA First-Touch Placement Ties Physical Pages to the Faulting CPU

On a NUMA machine, reserving virtual address space does not necessarily choose the physical NUMA node that will supply every page. For anonymous memory, physical allocation commonly happens later, when a CPU first faults on a page. Under the default local allocation policy, that fault can make initialization order part of the memory-placement decision. This behavior is often called first-touch placement. The thread that first writes or otherwise faults a page can cause Linux to allocate backing memory near the NUMA node on which that thread is running, subject to the active memory policy, cpuset constraints, available memory, and kernel fallback behavior.

Software Engineering 22 Sep 2026 8 min read

Hedged Requests Trade Duplicate Work for Lower Tail Latency

Hedged Requests Trade Duplicate Work for Lower Tail Latency Most requests may finish quickly while a small fraction take much longer. A busy worker, a transient network queue, a cold cache entry, garbage collection, storage contention, or another local disturbance can stretch one attempt far beyond the median. At scale, those slow outliers become visible in p95, p99, and higher-percentile latency even when average service time looks healthy. A hedged request sends an additional attempt after the original has been outstanding for a chosen delay. Both attempts represent the same logical operation. The caller accepts the first valid result and cancels or ignores the remaining attempt.

Software Engineering 20 Sep 2026 7 min read

Adaptive Concurrency Limits Follow Service Capacity

Adaptive Concurrency Limits Follow Service Capacity A service can become slower before it becomes unavailable. As in-flight work rises, CPU queues grow, connection pools fill, lock contention increases, and downstream calls accumulate. A fixed concurrency ceiling can protect the service, but one number rarely fits every operating condition. Capacity shifts with request mix, cache hit rate, dependency latency, deployment shape, and resource pressure. Adaptive concurrency control treats the admission limit as a value that can move. The controller observes recent service behavior, raises the limit while additional concurrency remains productive, and reduces it when latency indicates growing queues or saturation. The goal is not maximum concurrency. It is enough parallel work to use available capacity without allowing queues to dominate response time.

Artificial Intelligence 19 Sep 2026 8 min read

Where PEGASUS-XSum Inference Time Goes

google/pegasus-xsum is a summarization checkpoint, not a compact text utility. Its latency follows directly from the work performed by a large Transformer encoder-decoder: first encode the source document, then run the decoder repeatedly until the summary is complete. On a CPU, the second phase is usually the part that makes a short output feel disproportionately expensive. The original PEGASUS work describes a Transformer encoder-decoder pretrained with Gap Sentences Generation and reports a 568M-parameter best model. The XSum checkpoint is fine-tuned for highly abstractive single-document summarization. That combination is useful when summary quality matters, but it also means inference has substantially more machinery than extracting a few source sentences or running a small classifier.

Tech 19 Sep 2026 5 min read

NIC Interrupt Moderation Trades Wakeup Rate for Packet Latency

A network adapter does not need to interrupt a CPU for every received packet or completed transmission. Many NICs can hold interrupt delivery briefly and report several completion events together. This interrupt moderation reduces interrupt traffic and CPU entry overhead, but it can also delay the moment software notices newly completed work. The mechanism sits between packet DMA and the driver’s receive or transmit processing. It changes notification timing; it does not change the packet’s wire format, Ethernet ordering rules, or the basic requirement that the driver eventually process completed descriptors.

Tech 19 Sep 2026 5 min read

Cache-Line False Sharing Moves Coherence Ownership Between CPUs

Cache-Line False Sharing Moves Coherence Ownership Between CPUs Two threads can update different variables without sharing a lock or touching the same bytes and still interfere at the hardware level. If those variables occupy the same cache line, a coherent multiprocessor treats their storage as one coherence unit. Repeated writes from different CPUs can therefore move ownership of that line between caches even though the program considers the variables independent.

Tech 17 Sep 2026 6 min read

Write Combining Merges Adjacent Stores Before Memory Traffic

Some memory regions are written far more often than they are read. Frame buffers, device apertures, and streaming output areas are common examples. Sending every small CPU store as a separate memory transaction can waste bus bandwidth and transaction overhead. Write combining gives the processor a temporary place to collect compatible stores. Several writes targeting nearby addresses can be merged into a larger transaction before they leave the CPU. The technique favors sustained write throughput, but it changes the timing and ordering properties that software can safely assume.

Tech 17 Sep 2026 9 min read

NVMe Queue Pairs Separate Command Submission from Completion

NVMe Queue Pairs Separate Command Submission from Completion An NVMe solid-state drive does not need the CPU to hand each storage command directly to a device register and then wait for that command to finish. Instead, NVMe places command and completion records in queues held in host memory. The controller reads pending commands from submission queues and writes results to associated completion queues. That arrangement matches fast PCIe storage well. Modern SSD controllers can process many operations at once across flash channels, internal dies, and controller pipelines. A queue model lets software keep that parallel hardware busy while avoiding a long series of synchronous command handoffs.

Tech 17 Sep 2026 6 min read

NUMA Makes Memory Location Part of Access Cost

NUMA Makes Memory Location Part of Access Cost A large multiprocessor server can expose one physical address space while giving different processors different paths to that memory. A load from a page attached to the processor running a thread can take a shorter route than a load from memory attached to another processor package or NUMA node. This arrangement is called non-uniform memory access, or NUMA. It lets systems scale memory capacity and bandwidth across multiple processor sockets or chiplet groups without forcing every memory request through one centralized controller.

Tech 17 Sep 2026 6 min read

NIC Interrupt Coalescing Trades CPU Overhead for Packet Latency

A network interface can receive packets faster than a CPU should service one hardware interrupt per packet. Interrupt coalescing addresses that mismatch by allowing the adapter to group completion notifications and interrupt the CPU less often. The tradeoff is explicit. Fewer interrupts reduce interrupt handling and scheduling pressure, but a packet may wait longer before software is told that receive work is ready. The best setting depends on packet rate, latency targets, CPU capacity, and the adapter’s coalescing controls.

Tech 17 Sep 2026 6 min read

Memory Fences Constrain Cross-Core Memory Ordering

Memory Fences Constrain Cross-Core Memory Ordering A processor can execute memory operations with more freedom than source-code order suggests. Loads may begin early, stores may wait in buffers, cache-coherence traffic may complete at different times, and independent operations can overlap. These techniques improve throughput, but concurrent software needs precise rules for publishing and observing shared state. A memory fence places an ordering constraint around selected memory operations. It does not normally flush every cache, serialize the entire processor, or make all cores execute one instruction stream. Its role is narrower: it restricts which memory-order outcomes are permitted across a defined boundary.

Tech 16 Sep 2026 7 min read

TLB Caches Recent Virtual Address Translations

Modern processors commonly execute programs in virtual address spaces. A load or store can begin with a virtual address while the memory system ultimately needs a physical location and access permissions. Page tables hold the mapping information, but consulting their hierarchy for every memory reference would add substantial work. A translation lookaside buffer, or TLB, keeps recently used address translations near the processor. A TLB hit supplies cached mapping information without a full page-table walk. A TLB miss triggers additional translation work even when the requested application data is already present in a CPU cache.

Tech 16 Sep 2026 6 min read

TCP Window Scaling Expands the Receive Window for Fast Long Paths

TCP flow control limits how much data a sender may have outstanding according to the receiving endpoint’s available buffer space. The receiver advertises that limit in the TCP Window field so the sender does not deliver data faster than the receiving stack can accept it. The Window field in the TCP header is 16 bits wide. Without an extension, its largest value is 65,535 bytes. That ceiling can be too small on a path that carries data quickly but has a substantial round-trip time.

Tech 16 Sep 2026 5 min read

SSD Garbage Collection Amplifies Host Writes

An SSD can write substantially more data to NAND than the host sends to the device. The extra traffic appears when the controller must relocate still-valid pages before reclaiming flash blocks that contain invalid data. This internal movement is write amplification. It is a consequence of the mismatch between fine-grained logical updates and NAND erase constraints, not an extra write issued by the application. NAND pages cannot be overwritten in place NAND flash is programmed in pages but erased in larger erase blocks. A page that already contains programmed data cannot generally receive an arbitrary in-place replacement. The controller writes the new version to another available page and marks the old physical page as stale in its mapping state.

Tech 16 Sep 2026 3 min read

Receive Side Scaling Distributes Network Flows Across CPU Queues

Receive Side Scaling Distributes Network Flows Across CPU Queues A fast network adapter can receive packets faster than one processor core can handle them efficiently. Receive Side Scaling, commonly abbreviated RSS, spreads incoming traffic across multiple hardware receive queues. Each queue can be associated with a different processor, allowing packet processing to run in parallel. RSS usually assigns packets by flow rather than distributing every packet independently. This preserves useful ordering properties while still spreading many simultaneous connections across available queues.

Tech 16 Sep 2026 5 min read

Network Interrupt Coalescing Batches Packets Before CPU Notification

A network interface can receive packets much faster than a CPU should be interrupted for each individual arrival. At high packet rates, one hardware interrupt per packet would consume substantial processor time in interrupt entry, scheduling, driver work, and return paths. Interrupt coalescing changes that pattern. The adapter waits for a small interval, a packet count, or another implementation-specific threshold before notifying the CPU. Several packet arrivals can then be handled from one notification.

Tech 16 Sep 2026 5 min read

Network Interrupt Coalescing Batches Packet Notifications

A network interface can receive packets far faster than a processor should handle individual hardware interrupts. If every packet immediately triggered an interrupt, high packet rates could consume substantial CPU time in interrupt handling and context transitions. Interrupt coalescing changes that pattern. The network adapter waits for several packets, a short timer, or another configured threshold before notifying the CPU. One interrupt can then cover multiple received packets. The tradeoff is direct: fewer interrupts reduce per-packet CPU overhead, while waiting to form a batch can add latency.

Linux 16 Sep 2026 5 min read

Linux PSI Separates Partial and Total Resource Stalls

A Linux host can report modest CPU utilization while runnable work is delayed, or ample memory capacity while tasks repeatedly stall in reclaim. Utilization counters describe resource activity; pressure stall information records time in which work cannot make progress because a resource is contended. PSI exposes that lost execution opportunity through CPU, memory, and I/O pressure files. Its central distinction is between a stall affecting at least one task and a stall that leaves every non-idle task unable to make progress.

Tech 16 Sep 2026 7 min read

Jumbo Frames Raise Payload Efficiency and MTU Risk

Ethernet networks commonly use an IP maximum transmission unit of 1500 bytes, but many switches, network adapters, and operating systems also support larger frames often called jumbo frames. A larger MTU lets each packet carry more application data before another set of packet headers and per-packet processing is required. That can reduce packet rate for a given throughput. The benefit is most relevant when hosts move large volumes of data and the full path supports the selected frame size.

Linux 16 Sep 2026 5 min read

io_uring Registered Files Bypass Repeated Descriptor Lookup

An io_uring request that uses a normal file descriptor still has to resolve that descriptor through the submitting task’s file table. A registered file takes a different path: the ring holds a reference to the open file, and an SQE names a slot in that ring-local table. That distinction removes repeated descriptor lookup from the request path. It also changes resource lifetime, update semantics, and the meaning of the SQE fd field.

Go 16 Sep 2026 5 min read

Go sync.Pool Items Can Disappear Across Garbage Collection

A value placed in a Go sync.Pool is not guaranteed to remain there until a later Get. The runtime may remove pooled items automatically, so the pool acts as a reuse opportunity rather than durable storage. That property shapes both the performance profile and the correctness boundary of sync.Pool. Code can benefit when an object survives long enough to be reused, but it must remain correct when every Get behaves as if no prior item were available.

Tech 16 Sep 2026 6 min read

DRAM Refresh Restores Charge Before Bits Fade

Dynamic random-access memory stores data in cells whose electrical state does not remain stable indefinitely. Charge leaks from a cell over time, even when software performs no reads or writes. A memory system therefore has to refresh DRAM periodically to preserve stored bits. Refresh is a maintenance operation rather than a request from an application. The memory controller and DRAM device coordinate it alongside ordinary reads and writes. During parts of that work, some memory resources cannot serve normal requests.

Tech 16 Sep 2026 7 min read

CPU Thermal Throttling Reduces Clock Speed at Temperature Limits

A processor can run at a high clock rate only while electrical, power, and temperature limits permit it. Heavy work raises transistor switching activity, which raises power consumption and heat output. If cooling cannot remove that heat quickly enough, the processor approaches a thermal limit. Modern CPUs respond automatically. Control logic reduces performance states so the chip generates less heat. This behavior is commonly called thermal throttling. Thermal throttling is not the same as a processor simply running below its advertised maximum frequency. Boost algorithms already vary clock speed according to workload, active core count, current, power, and temperature. Thermal throttling refers specifically to temperature becoming a limiting condition.