Skip to content

Archive

File I/O

8 articles
Linux 17 Sep 2026 5 min read

O_APPEND Couples End Positioning with Each Write

O_APPEND changes a write from two separable actions into one coupled operation: Linux positions the open file description at the current end of the file and performs the write as a single atomic step. That property matters when multiple writers target one regular file. A sequence built from lseek(fd, 0, SEEK_END) followed by write(fd, ...) does not carry the same append semantics because another writer can change the file between those two system calls.

Python 07 Sep 2026 11 min read

Use Memory-Mapped Files for Random Access in Python

Reading a file with read() gives your program a straightforward model: ask for bytes, receive a bytes object, and let Python manage the buffer. That is often the right choice. Some workloads are different. A program may need to inspect small regions scattered across a large file, search the same file repeatedly, or pass file-backed bytes to APIs that understand the buffer protocol. Repeated seek() and read() calls can work, but they make every access an explicit file operation in your code.

Python 05 Sep 2026 9 min read

Buffer Temporary Data in Python with SpooledTemporaryFile

Temporary data often has an awkward size profile. Most requests may produce a few kilobytes, while an occasional import, report, archive, or upload grows to hundreds of megabytes. Using io.BytesIO is convenient for the small case, but its contents stay in memory. Using a temporary file avoids keeping the whole payload in memory, but every payload uses file-system-backed storage even when it is tiny. Python’s tempfile.SpooledTemporaryFile gives you a middle ground. It behaves like a file object while keeping data in memory up to a configured threshold. When the data grows beyond that threshold, it rolls over to a temporary file and continues through the same interface.

Linux 04 Sep 2026 8 min read

Write Files Atomically on Linux with rename and fsync

Updating a small file looks simple: open it, truncate it, write the new contents, and close it. That works when nothing interrupts the write. The failure mode appears when a process crashes, the machine loses power, or another process reads the file while it is being rewritten. A reader can observe an empty or partially written file, and a crash can leave the pathname referring to incomplete data. For configuration files, state snapshots, generated metadata, and similar single-file updates, a better pattern is to write a complete replacement beside the original file and then rename it into place.

Python 04 Sep 2026 10 min read

Use Callable-Sentinel Iteration for Chunked Reads in Python

Reading a file in chunks is a small problem that appears in many larger tasks: hashing uploads, copying large files, parsing binary records, compressing streams, and sending data without loading everything into memory. A common solution is a while loop that reads one chunk, checks for end-of-file, processes the chunk, and repeats. That loop is correct when written carefully, but Python has another standard-library pattern that expresses the same control flow as iteration:

Linux 04 Sep 2026 13 min read

Understand Memory-Mapped Files on Linux with mmap

Reading a file usually means calling read() and copying bytes into a buffer that your program manages. That model is explicit and works well for most file I/O. Linux offers another model with mmap(): map a file region into the process’s virtual address space, then access the file through ordinary memory loads and stores. That can simplify workloads such as random access into large files, shared file-backed state, indexes, and binary formats whose access pattern naturally looks like “read bytes at offset N.” It can also avoid an application-managed read buffer for those accesses.

Linux 04 Sep 2026 13 min read

Prevent Overlapping Linux Jobs with Advisory File Locks and flock

A scheduled job often looks harmless until two copies run at the same time. A backup takes longer than usual, a second timer fires, and both processes start writing the same output. A maintenance script overlaps with itself and launches duplicate work. A cache refresh runs concurrently and leaves partially coordinated state behind. The problem is not that Linux started the processes incorrectly. The problem is that the application needs a rule saying, “only one cooperating process may enter this critical section at a time.”