Skip to content

Archive

Performance

83 articles
Database 08 Sep 2026 7 min read

Store JSON Faster with SQLite JSONB

I like SQLite’s JSON functions because they let me keep a small amount of flexible data without immediately turning every property into a column. The trade-off is easy to miss: if I store JSON as text, SQLite has to parse that text before it can navigate the structure. Since SQLite 3.45.0, there is another option. SQLite can persist its binary JSON representation, called JSONB, directly in a BLOB. Here’s the idea: if the database is going to inspect the same JSON repeatedly, I can let SQLite store the representation it already wants to process instead of making it parse the text again.

Tech 07 Sep 2026 7 min read

Why Closing Background Apps Does Not Always Help Your Phone

Open the app switcher on a phone and you may see many apps that you used earlier. It is easy to assume that every card represents an app actively running, consuming processor time, and draining the battery. That is not generally how modern mobile operating systems work. An app can remain available in the recent-apps interface while doing little or no active work. The operating system decides how much memory and background execution each app receives, and it can reclaim resources when they are needed elsewhere.

Python 07 Sep 2026 11 min read

Use Memory-Mapped Files for Random Access in Python

Reading a file with read() gives your program a straightforward model: ask for bytes, receive a bytes object, and let Python manage the buffer. That is often the right choice. Some workloads are different. A program may need to inspect small regions scattered across a large file, search the same file repeatedly, or pass file-backed bytes to APIs that understand the buffer protocol. Repeated seek() and read() calls can work, but they make every access an explicit file operation in your code.

Go 07 Sep 2026 11 min read

Suppress Duplicate Concurrent Work in Go with singleflight

A service can become overloaded even when each individual request is reasonable. Imagine a popular product whose cached record expires. Fifty requests arrive almost together, all observe the same cache miss, and all start the same database query. The problem is not ordinary parallelism. Those requests are doing duplicate work for the same result at the same time. golang.org/x/sync/singleflight provides a small mechanism for suppressing that duplication inside one Go process. For a given key, one caller performs the work while concurrent callers for the same key wait and receive the same result. Calls using different keys can still perform their own work.

Python 07 Sep 2026 8 min read

Build Reliable Worker Queues in Python with queue.Queue

A worker thread is easy to start. A reliable worker queue is harder. The difficult parts appear when production code must answer questions such as: What happens when producers are faster than consumers? How does the main thread know that processing, rather than merely dequeuing, is complete? How do workers stop without abandoning queued work? What happens if processing raises an exception? Python’s queue.Queue provides the synchronization needed to pass work safely between threads, but correct coordination still depends on a few application-level invariants. The most important are to bound work when memory matters, pair every successful get() with exactly one task_done(), and separate “all work is finished” from “workers should exit.”

Python 07 Sep 2026 11 min read

Bound In-Flight Thread Pool Work in Python

A thread pool limits how many functions run at the same time, but it does not automatically limit how much work your producer can queue. That distinction matters when the input is large or unbounded. A loop can submit millions of tasks to a ThreadPoolExecutor while only a handful of worker threads execute them. The remaining tasks are pending Future objects, along with their arguments and other referenced state. If the producer is much faster than the workers, memory use can grow long before CPU or network capacity is exhausted.

Tech 06 Sep 2026 9 min read

Why Your Laptop Fan Speeds Up and Slows Down

A laptop can be almost silent while you read a document, become noticeably louder during a video call, and keep its fan running for a while after the call ends. The changing noise can make it seem as though the computer is working unpredictably. Usually, the opposite is true. The laptop is adjusting its cooling to match changing heat inside the machine. The fan does not simply respond to whether an application looks demanding. It responds to temperatures and to the cooling policy chosen by the laptop’s hardware and software.

Tech 06 Sep 2026 7 min read

Why Background Browser Tabs Sometimes Reload

You leave a web page open in a background tab, work elsewhere for a while, and then return to it. Instead of appearing exactly as you left it, the page briefly goes blank, shows a loading indicator, or fetches its contents again. This can look like a browser failure, but it is often deliberate resource management. An open tab is not necessarily a promise that the entire page will remain active in memory indefinitely. Browsers and operating systems can reduce the resources used by pages you are not currently viewing, especially when memory is needed elsewhere.

Tech 05 Sep 2026 9 min read

Why File Copy Speed Can Change During a Transfer

Copy a large folder to an external drive and the speed may begin high, fall sharply, recover, and then change again. Sometimes the progress window even shows a speed that looks impossible for the device involved. This can make a healthy drive seem unreliable or a fast connection seem slower than expected. The key is that a file copy is not one continuous action at one fixed rate. Data passes through several stages, and the operating system may temporarily hold some of it in fast memory before slower storage finishes the work. The source, destination, connection, file sizes, and other activity on the device can all become the limiting factor at different moments.

Artificial Intelligence 05 Sep 2026 9 min read

Reduce LLM Decoding Latency with Speculative Decoding

Large language models generate text autoregressively: each new token depends on the tokens that came before it. That dependency makes ordinary decoding sequential. Even when a GPU has enough compute to process many token positions in parallel, the model normally discovers only one new token per decoding step. Speculative decoding tries to turn some of that sequential work into parallel verification. A faster draft model proposes several future tokens. The larger target model then scores those proposed positions together and decides which proposals can be accepted. When the draft predicts well, one expensive target-model pass can advance generation by multiple tokens.

Linux 05 Sep 2026 10 min read

Move Data Between Linux File Descriptors with splice()

A conventional file-copy loop reads bytes into a user-space buffer and then writes those bytes somewhere else. That pattern is portable and easy to understand, but sometimes the program does not need to inspect or transform the data at all. It only needs to move bytes from one file descriptor to another. On Linux, splice() can handle that case differently. It transfers data between file descriptors while keeping the transferred data out of a user-space buffer. At least one endpoint must be a pipe, so a pipe can act as the kernel-side bridge between a source and a destination.

Python 05 Sep 2026 9 min read

Buffer Temporary Data in Python with SpooledTemporaryFile

Temporary data often has an awkward size profile. Most requests may produce a few kilobytes, while an occasional import, report, archive, or upload grows to hundreds of megabytes. Using io.BytesIO is convenient for the small case, but its contents stay in memory. Using a temporary file avoids keeping the whole payload in memory, but every payload uses file-system-backed storage even when it is tiny. Python’s tempfile.SpooledTemporaryFile gives you a middle ground. It behaves like a file object while keeping data in memory up to a configured threshold. When the data grows beyond that threshold, it rolls over to a temporary file and continues through the same interface.

Tech 04 Sep 2026 7 min read

What Happens When You Leave Apps Running in the Background

You switch away from a messaging app, open a browser, and later return to find the message exactly where you left it. It is natural to assume the first app has been fully running the whole time. That assumption leads to a common question: should you regularly close every app to save battery or make your device faster? Usually, the answer is no. An app that appears to be open is not necessarily doing active work. Modern operating systems manage apps according to what they need, what the device can afford, and what background activity the platform allows.

Linux 04 Sep 2026 13 min read

Understand Memory-Mapped Files on Linux with mmap

Reading a file usually means calling read() and copying bytes into a buffer that your program manages. That model is explicit and works well for most file I/O. Linux offers another model with mmap(): map a file region into the process’s virtual address space, then access the file through ordinary memory loads and stores. That can simplify workloads such as random access into large files, shared file-backed state, indexes, and binary formats whose access pattern naturally looks like “read bytes at offset N.” It can also avoid an application-managed read buffer for those accesses.

Go 04 Sep 2026 9 min read

Understand Escape Analysis and Heap Allocations in Go

A Go program can create many values without you choosing whether each value lives on a goroutine stack or in heap memory. The compiler usually makes that decision for you. This becomes important when a hot path allocates more than expected. A small helper function may look harmless, yet a value it creates can outlive the function call and require heap storage. More heap allocation can mean more work for the garbage collector, but changing code blindly to avoid the heap can make a program harder to understand without producing a measurable benefit.

Go 03 Sep 2026 11 min read

Use sync.Pool for Temporary Object Reuse in Go

Repeatedly allocating short-lived helper objects can become expensive in a hot path. A formatter may create temporary buffers for every request, an encoder may allocate scratch space for every record, or a parser may repeatedly construct helper objects that are discarded immediately after use. Go’s sync.Pool can reuse some of those temporary objects across independent operations. That can reduce allocation work and garbage-collector pressure when the same kind of object is created frequently under load.

Artificial Intelligence 03 Sep 2026 10 min read

Speculative Decoding for Faster LLM Inference

Autoregressive language models generate text sequentially. After processing the prompt, the model predicts a next token, appends that token to the sequence, and repeats the process. That dependency makes generation difficult to parallelize across time: token 101 cannot normally be generated until token 100 is known. Speculative decoding changes the amount of useful work performed during each expensive target-model step. A cheaper draft process proposes several future tokens, then the target model verifies those proposals together. When enough proposals are accepted, the application can advance by multiple tokens while invoking the large model fewer times.

Python 03 Sep 2026 10 min read

Building Memory-Efficient Iterator Pipelines with Python itertools

Python programs often transform data in stages: read records, discard unwanted items, reshape values, group adjacent entries, and stop after enough output has been produced. A straightforward implementation may build a new list after every stage. That is easy to understand, but it can also allocate intermediate collections that the next stage immediately consumes. Iterator pipelines offer another model. Each stage requests values from the stage before it as needed. Python’s itertools module provides building blocks for this style, including tools for chaining inputs, taking slices from streams, computing running values, grouping consecutive records, and duplicating an iterator when two consumers genuinely need it.

Artificial Intelligence 03 Sep 2026 9 min read

Batch LLM Inference for Better Throughput

An LLM server can receive many requests at the same time, yet processing every request independently is often an inefficient way to use an accelerator. GPUs and similar hardware are designed to perform large amounts of parallel numerical work. A single small request may leave part of that capacity unused. Batching combines work from multiple requests so the model can process more of it together. This can improve total throughput, but it introduces an important trade-off: waiting to form a batch can delay individual requests, and requests with different sequence lengths do not all consume the same amount of work.

Tech 02 Sep 2026 8 min read

What RAM Does and How Much Memory You Actually Need

Random-access memory, usually shortened to RAM, is one of the specifications people see when buying a computer, tablet, or phone. A device might have 8 GB, 16 GB, 32 GB, or more, but the number is easy to misunderstand. RAM is not the same as permanent storage. It is fast working memory that the system uses while applications, documents, browser tabs, and background services are active. More RAM can make a device handle heavier workloads more comfortably, but adding memory does not automatically make every task faster. The useful question is whether your workload regularly needs more memory than the device can provide efficiently.

Database 02 Sep 2026 7 min read

SQLite WAL Mode: Concurrency, Checkpoints, and Operational Pitfalls

SQLite is often chosen because it keeps deployment simple: an application can get transactional storage without operating a separate database server. As workloads become more concurrent, however, the default rollback journal can make read and write activity interfere more than expected. Write-ahead logging (WAL) changes that coordination model. Readers can usually continue while a writer commits changes, but WAL does not turn SQLite into a multi-writer database. Correct operation still depends on short transactions, sensible busy handling, and checkpoints that can make progress.

Python 02 Sep 2026 9 min read

Python memoryview: Zero-Copy Access to Binary Buffers

Binary-processing code often needs only a small region of a larger byte buffer. A normal bytes or bytearray slice is convenient, but it creates a new object containing copied data. When buffers are large or slicing happens repeatedly on a hot path, those copies can become unnecessary allocation and memory traffic. Python’s memoryview provides a different model. It exposes data from an object that supports the buffer protocol and lets Python code work with that data without first copying it into a new bytes object.

Database 02 Sep 2026 7 min read

PostgreSQL Partial Indexes for Focused Query Workloads

A normal PostgreSQL index contains entries for every table row that has indexable values. That is often appropriate, but some workloads repeatedly query a small, well-defined subset of a much larger table. A partial index stores entries only for rows that satisfy an index predicate. When the predicate matches a stable access pattern, the index can be smaller and cheaper to maintain than an equivalent full-table index. The trade-off is specificity: PostgreSQL can use the partial index only when it can determine at planning time that the query condition implies the index predicate.

Software Engineering 02 Sep 2026 7 min read

Load Shedding and Bounded Queues for Overload Control

A service can be healthy at 500 requests per second and unusable at 700. The extra load does not merely make every request 40 percent slower. Queues grow, deadlines expire while work is still waiting, memory usage rises, retries create more traffic, and useful requests compete with work that can no longer finish in time. Overload control keeps that failure mode bounded. Instead of accepting unlimited work, a service limits concurrency and queueing, then rejects or degrades excess work early enough for the remaining requests to succeed.