Skip to content

Archive

Performance

83 articles
Tech 16 Sep 2026 6 min read

CPU Store Buffers Decouple Retirement from Cache Writes

A processor does not need every store instruction to finish its cache update before later instructions make progress. Modern cores commonly place completed stores into a store buffer, allowing the instruction to retire while the memory subsystem handles the write afterward. This separation improves throughput because cache ownership, coherence traffic, and other memory activity can take longer than the execution pipeline can afford to wait. The buffer acts as a queue between architectural execution and the cache hierarchy.

Tech 16 Sep 2026 6 min read

CPU Cache Associativity Limits Where Lines Can Reside

Processor caches keep recently used memory close to execution cores, but cache capacity alone does not determine which data can remain resident. Most general-purpose CPU caches divide storage into sets and give each set a fixed number of slots, commonly called ways. A memory block maps to a particular set. It can occupy any way inside that set, but it cannot move into an unrelated set merely because that other set has free space. This placement rule makes hardware lookup practical and fast, while creating a distinct source of misses when too many active blocks compete for the same set.

Tech 15 Sep 2026 5 min read

Thermal Throttling Changes Sustained Processor Speed

A processor can finish a short burst of work at a high clock rate and then settle at a lower rate during a long workload. The change does not necessarily indicate a fault. Modern processors operate inside several limits, and temperature is one of the conditions that can reduce the frequency available over time. Thermal throttling is a control response that keeps a processor within permitted operating conditions. It becomes visible when heat generation exceeds what the cooling system can remove while the workload continues. The resulting clock behavior makes peak specifications a poor substitute for sustained performance measurements.

Tech 15 Sep 2026 7 min read

TCP Window Scaling Expands Receive Capacity

A TCP connection can have plenty of bandwidth available and still transfer data below the path’s capacity. One limit can come from flow control: the receiver tells the sender how much additional data it is prepared to accept, and the sender must respect that boundary. The original TCP header allocates 16 bits to the advertised receive window. That field can represent at most 65,535 bytes directly. TCP window scaling extends its effective range by negotiating a multiplier during connection setup, making much larger receive windows possible without changing the size of the header field.

Tech 15 Sep 2026 7 min read

TCP Delayed ACK Reduces Acknowledgment Traffic

TCP Delayed ACK Reduces Acknowledgment Traffic TCP acknowledgments provide essential feedback, but sending a separate ACK for every incoming data segment is not always necessary. A receiver can briefly defer an acknowledgment so that one ACK covers more than one segment. This behavior is known as delayed acknowledgment, or delayed ACK. The mechanism reduces packet processing and reverse-path traffic during steady data transfer. It also introduces a timing tradeoff: if another segment does not arrive soon enough, the receiver eventually has to send the pending ACK on its own.

Tech 15 Sep 2026 5 min read

SSD Write Cache and Sustained Transfer Speed

An SSD can copy the first part of a large file at high speed, then settle at a much lower rate even though nothing else appears to have changed. That drop can be normal. Many consumer SSDs use part of their NAND as a fast write cache, allowing short bursts to finish before the drive has to sustain writes in its denser storage mode. This behavior makes a single peak transfer number a poor description of every write workload. Cache size, free space, NAND type, controller policy, temperature, and the amount of data already waiting inside the drive can all affect the speed seen during a long transfer.

Tech 15 Sep 2026 7 min read

NVMe Queues Let Storage Handle Many Commands in Parallel

NVMe storage does not send every read or write through one shared command line. The protocol is built around queue pairs: software places commands into a submission queue, and the controller reports finished work through a corresponding completion queue. That structure matters most when several processor cores and application threads are generating storage work at the same time. Multiple queues can distribute command handling across cores, reduce contention around a single software path, and keep a fast solid-state drive supplied with enough outstanding work.

Tech 14 Sep 2026 6 min read

Thermal Throttling Cuts Chip Speed as Heat Rises

A processor can feel fast during a short task and then slow during a long render, game, compile, or benchmark. One common cause is thermal throttling: automatic control that reduces chip activity as temperature approaches a configured limit. The mechanism protects the processor and surrounding hardware while keeping the system operating. It also means that advertised peak clock speeds do not describe sustained performance in every enclosure, workload, or ambient temperature.

Tech 14 Sep 2026 5 min read

Memory Compression Keeps More Active Data in RAM

Modern operating systems can hold compressed copies of memory pages in RAM when physical memory becomes crowded. The technique increases the amount of useful data that fits in a fixed quantity of RAM without changing the installed hardware. Compression is not free capacity. It exchanges processor time and some memory space for a smaller representation of data that would otherwise occupy more RAM or become a candidate for storage-backed paging.

Tech 13 Sep 2026 7 min read

Thermal Throttling and Sustained Performance

A phone, laptop, handheld console, or desktop can begin a demanding task at high speed and settle at a lower speed several minutes later. The processor has not necessarily developed a fault. Modern chips continuously operate within electrical, power, and temperature limits, and their control systems can reduce performance when those limits become restrictive. This behavior is commonly called thermal throttling when temperature is the active constraint. It protects the processor and surrounding components while keeping the device inside its intended operating envelope. The practical result is a gap between brief peak performance and the level a system can sustain during a long workload.

Tech 13 Sep 2026 7 min read

SSD TRIM and Deallocated Storage

Deleting a large file can make free space appear immediately in the operating system, yet an SSD does not treat that event like a hard drive overwriting a fixed physical location. The file system releases its own allocation first. A separate deallocation signal can then tell the SSD that the corresponding logical blocks no longer contain data the host needs. That signal is commonly called TRIM. On NVMe storage, the comparable operation is deallocation through Dataset Management. The names differ across storage interfaces, but the practical idea is similar: the host identifies logical ranges whose previous contents no longer need to be preserved.

Tech 13 Sep 2026 5 min read

SD Card Speed Classes Describe Minimum Write Performance

An SD card can carry several speed marks at once: C10, U3, and V30 are common examples. The numbers can look like competing estimates of the same speed, but they serve a more specific purpose. A speed class states an assured minimum sequential access performance under defined conditions, with a strong emphasis on sustained recording workloads. That makes a class mark different from a large read or write figure printed elsewhere on a card or package. A peak figure can describe a maximum transfer rate reached under particular conditions. A speed class establishes a performance floor for a supported access pattern.

Tech 12 Sep 2026 5 min read

Virtual Memory and Storage Under Memory Pressure

A computer can keep several large applications open even when their combined memory demands exceed the amount of physical RAM available at that moment. The operating system manages this pressure by deciding which memory contents need to remain in RAM and which can be moved elsewhere or discarded and recreated later. Virtual memory is central to that process. It gives software an address space that is separate from the exact layout of physical RAM, allowing the operating system to map memory pages to different physical locations as conditions change.

Tech 12 Sep 2026 7 min read

SSD TRIM and Deleted Storage Space

Deleting a large file can make free space appear immediately in an operating system, yet the solid-state drive underneath may still contain the old data in flash cells for some time. The file system and the SSD are tracking different things. The file system knows which logical blocks are no longer needed; the drive manages where those blocks physically reside in flash. TRIM connects those two views. It lets an operating system tell an SSD that specified logical block addresses no longer contain data that must be preserved. The command does not act like a file shredder, and it does not directly erase every affected flash cell at the moment a file disappears.

Tech 11 Sep 2026 9 min read

Phone Storage Speed: What It Changes in Everyday Use

A phone with plenty of free storage can still feel slow when opening a large app, installing an update, or moving a big video. Capacity tells you how much data fits on the device. Storage speed describes how quickly that data can be read or written. That distinction matters because phones constantly use internal storage. Apps read code and resources from it, the camera writes photos and video to it, and the operating system uses it for updates, caches, and other files.

Software Engineering 10 Sep 2026 8 min read

Stale-While-Revalidate for Responsive Caches

Stale-While-Revalidate for Responsive Caches A cache entry expires just as a request arrives. The cached value is still only seconds old, but the request now has to wait while the application fetches a replacement from a slower dependency. If many entries expire during a busy period, cache refreshes can turn a normally fast read path into a burst of slow work. Stale-while-revalidate changes that trade-off. For data that can safely be slightly out of date, the application may return an expired cached value immediately while refreshing it separately for future requests. The reader gets predictable latency, and the cache still moves toward fresh data.

Go 10 Sep 2026 7 min read

Preallocate Append Capacity in Go with slices.Grow

Sometimes you know a slice is about to receive several elements, even though you don’t have those elements yet. Repeated append calls will grow the slice automatically, but some of those appends may have to allocate a larger backing array and copy the existing elements. slices.Grow lets you reserve enough capacity for a known amount of upcoming growth. It doesn’t add placeholder elements and it doesn’t change the slice’s length. It simply returns a slice that has room for at least the requested number of additional elements.

Software Engineering 10 Sep 2026 11 min read

Negative Caching for Repeated Failures

Negative Caching for Repeated Failures Caching usually brings successful results to mind: load a value once, keep it for a while, and avoid repeating expensive work. But repeated failures can be just as expensive as repeated successes. Suppose a service receives thousands of requests for an object that does not exist. If every request queries the same downstream system, the absence of that object becomes a source of load. The same pattern appears with invalid identifiers, unavailable optional resources, failed name lookups, and other outcomes that are expensive to rediscover but unlikely to change immediately.

Go 10 Sep 2026 9 min read

Build Lazy Iterators in Go with iter.Seq

A Go API that returns a slice is pleasantly simple, but a slice isn’t always the right contract. Sometimes the caller only needs the first matching value. Sometimes producing each value requires work. Sometimes the complete result could be large enough that building it up front is wasteful. Go 1.23 gives those APIs a standard alternative: iter.Seq. It represents a sequence that produces values on demand and works directly with a for range loop. The useful part isn’t just new syntax. An iter.Seq can hide a container’s representation, avoid an intermediate result slice, and stop producing values as soon as the caller stops iterating.

Tech 09 Sep 2026 9 min read

Why SSDs Can Slow Down When They Are Nearly Full

A solid-state drive can still have free space and yet become less consistent at writing data as it fills up. You might notice a large file transfer slowing down, an installation taking longer than expected, or a heavily used computer feeling less responsive when storage is almost exhausted. The reason is not simply that an SSD has to “search harder” for empty space. Flash storage has rules about how data can be written and erased. When plenty of unused space is available, the drive has more flexibility to work within those rules. When little space remains, that work can become more complicated.

Software Engineering 09 Sep 2026 11 min read

Using Hedged Requests to Reduce Tail Latency

Most calls to a dependency may finish quickly while a small fraction take much longer. A request that depends on one of those slow calls inherits the delay even when another healthy instance could have answered sooner. Increasing the timeout does not solve this problem. Retrying only after the timeout may also be too late: by then, the caller has already spent most of its latency budget. A hedged request is a deliberately delayed duplicate of an operation that is still in progress. The original request starts normally. If it has not completed after a chosen delay, the caller sends one additional equivalent request, usually to another eligible instance. The first acceptable result wins, and the remaining work is cancelled or ignored.

Python 09 Sep 2026 10 min read

Run CPU-Bound Python with InterpreterPoolExecutor

For CPU-heavy Python work, I usually reach for ProcessPoolExecutor. Threads are convenient, but ordinary CPython threads do not give CPU-bound Python code the kind of multi-core parallelism people often expect. Python 3.14 adds another option: concurrent.futures.InterpreterPoolExecutor. It looks deliberately familiar. You still submit callables and receive futures, but every worker thread owns a separate Python interpreter. Each interpreter has its own GIL, so Python code in different workers can execute on different CPU cores at the same time.

Python 09 Sep 2026 13 min read

Reduce Filesystem Stat Calls with Python Path.info

Filesystem code often looks cheap until it runs over a directory with hundreds of thousands of entries. A loop that asks whether every path is a file, directory, or symbolic link can translate into a large number of metadata queries. On a local SSD that cost may be tolerable. On network filesystems, container-mounted volumes, or very large trees, repeated metadata lookups can become a noticeable part of runtime. Python 3.14 adds Path.info, a cached file-type information interface on pathlib.Path. It is especially useful when paths come from Path.iterdir(), because Python may initialize the cache with information already obtained while scanning the directory.

Software Engineering 09 Sep 2026 10 min read

Coalescing Duplicate In-Flight Work

A service can receive many requests for the same expensive result at almost the same time. If every request starts identical work, a brief traffic burst can become a much larger burst against a database, remote API, filesystem, or CPU-heavy computation. Caching can help after a result exists. It does not necessarily help when the result is missing and many callers discover that miss together. Request coalescing solves this narrower problem. While an operation for a particular key is already running, later callers for the same key join that operation instead of starting another one. When it finishes, the waiting callers receive the same outcome.