Skip to content

Archive

Standard Library

147 articles
Go 06 Sep 2026 9 min read

Handle Long Lines Safely in Go with bufio.Scanner

bufio.Scanner is one of the simplest ways to process line-oriented input in Go. It is a good fit for log files, command output, newline-delimited JSON, and other formats where one logical record fits in memory. The convenience has an important boundary: a Scanner will stop if the next token grows beyond the amount of buffering it is allowed to use. That default protects a program from growing memory without bound, but it can surprise code that works on small test files and later encounters one unusually long line in production.

Python 05 Sep 2026 12 min read

Parse Binary Records Safely in Python with struct

Binary files and network messages often begin with fixed-width fields: a four-byte signature, a one-byte version, a two-byte payload length, or a four-byte identifier. Those fields are easy to describe on paper but surprisingly easy to parse incorrectly in code. The difficult part is not converting bytes to integers. It is preserving the binary layout contract: exactly which byte belongs to which field, which byte order is used, how wide each value is, and what should happen when the input is incomplete or malformed.

Python 05 Sep 2026 9 min read

Build Reliable Priority Queues in Python with heapq

A priority queue answers one question repeatedly: which pending item should run next? Schedulers, retry systems, graph algorithms, simulations, and background workers all need some version of that operation. A list can hold pending items, but finding the best one by scanning costs linear time each time. Keeping the whole list sorted makes retrieval cheap, but insertion has to preserve that full ordering. Python’s heapq module uses a heap instead. A heap is partially ordered: it guarantees that the smallest item is at heap[0], but it does not keep every element globally sorted. Push and pop operations take logarithmic time, while reading the current minimum is constant time.

Python 05 Sep 2026 9 min read

Buffer Temporary Data in Python with SpooledTemporaryFile

Temporary data often has an awkward size profile. Most requests may produce a few kilobytes, while an occasional import, report, archive, or upload grows to hundreds of megabytes. Using io.BytesIO is convenient for the small case, but its contents stay in memory. Using a temporary file avoids keeping the whole payload in memory, but every payload uses file-system-backed storage even when it is tiny. Python’s tempfile.SpooledTemporaryFile gives you a middle ground. It behaves like a file object while keeping data in memory up to a configured threshold. When the data grows beyond that threshold, it rolls over to a temporary file and continues through the same interface.

Go 05 Sep 2026 8 min read

Bound Untrusted Stream Reads in Go with io.LimitReader

Reading an entire io.Reader is convenient, but convenience can become a resource problem when the reader is not fully under your control. A request body, uploaded file, decompressed stream, subprocess output, or protocol payload may be much larger than expected. Calling io.ReadAll directly asks Go to keep reading until EOF, growing memory as needed. The important mental model is that a size limit should sit in front of the consumer. Instead of trusting every caller to stop at the right point, wrap the source in a reader that exposes only a bounded prefix. Go’s io.LimitReader does exactly that.

Python 04 Sep 2026 10 min read

Use Decimal for Exact Base-10 Arithmetic in Python

A price of 0.10, a tax rate of 8.25%, and a total rounded to cents look like ordinary numbers. But the representation you choose determines which arithmetic rules your program actually follows. Python’s float is binary floating point. It is excellent for measurements, graphics, scientific calculations, and many other workloads where small approximation is expected. The problem appears when your domain requires values and rounding rules expressed in base 10.

Python 04 Sep 2026 10 min read

Use Callable-Sentinel Iteration for Chunked Reads in Python

Reading a file in chunks is a small problem that appears in many larger tasks: hashing uploads, copying large files, parsing binary records, compressing streams, and sending data without loading everything into memory. A common solution is a while loop that reads one chunk, checks for end-of-file, processes the chunk, and repeats. That loop is correct when written carefully, but Python has another standard-library pattern that expresses the same control flow as iteration:

Python 04 Sep 2026 11 min read

Model Combinable Options in Python with enum.Flag and IntFlag

Some values represent one choice from a fixed set. A log level might be INFO, WARNING, or ERROR. Python’s Enum is a natural fit because one value should identify one member. Other values represent a combination of independent options. A file operation may allow reading and writing. A protocol field may enable compression and encryption. A component may expose several capabilities at the same time. Representing those combinations as ordinary booleans works at first:

Python 04 Sep 2026 10 min read

Decode Streaming Text Safely in Python with Incremental Codecs

Network sockets, compressed streams, subprocess pipes, and chunked file reads often deliver bytes in arbitrary pieces. If those bytes represent text, it is tempting to decode each piece immediately: for chunk in byte_chunks: text = chunk.decode("utf-8") process(text) That works only when every chunk happens to end on a character boundary.

Python 04 Sep 2026 10 min read

Build Portable Readiness Loops in Python with selectors

A network service can handle one connection with straightforward blocking calls: accept a client, read a request, write a response, and repeat. The model becomes awkward when one thread must manage many connections at once. The problem is not that sockets are slow. The problem is that a blocking operation can stop the thread while one connection waits, even though other connections are ready for useful work. Python’s selectors module provides a higher-level way to wait for I/O readiness across multiple file objects. Instead of asking one socket to block until something happens, you register many sockets and ask the selector which ones are currently ready.

Python 03 Sep 2026 11 min read

Use Weak References for Non-Owning Object Relationships in Python

Most Python code should use ordinary references. If one object stores another object in an attribute, list, or dictionary, that reference normally means the stored object should remain available for as long as the owner needs it. Some relationships are different. A cache may want to reuse an object only while another part of the program already owns it. A registry may want to discover live objects without extending their lifetime. An observer table may want to remember listeners without becoming the reason those listeners can never be collected.

Python 03 Sep 2026 10 min read

Use Single Dispatch for Type-Based Behavior in Python

A function sometimes needs to perform the same conceptual operation for several unrelated Python types. The straightforward solution is usually an if chain: def format_value(value): if isinstance(value, str): return value if isinstance(value, int): return str(value) if isinstance(value, dict): return ", ".join(f"{key}={item}" for key, item in value.items()) raise TypeError(f"unsupported type: {type(value).__name__}") This is perfectly reasonable when the set of supported types is small and unlikely to grow.

Python 03 Sep 2026 9 min read

Use Decimal for Predictable Base-10 Arithmetic in Python

Many Python programs can use float without trouble. Measurements, graphics, statistics, and scientific calculations often benefit from fast binary floating-point arithmetic. Problems appear when the data itself is defined in decimal terms and exact decimal values matter. A price such as 19.99, a tax rate such as 7.5%, or a quantity rounded to two decimal places may need rules that match decimal arithmetic rather than the binary representation used by float.

Python 03 Sep 2026 10 min read

Read and Write CSV Reliably in Python

CSV looks simple because a small file may resemble plain text with commas between values. That mental model breaks as soon as a field itself contains a comma, quote, or newline. Consider one valid record: 42,"Nguyen, Mai","Line one Line two" Splitting this text on commas cannot recover the three fields correctly. The comma inside the name is data, and the newline inside the quoted field belongs to the same record.

Python 03 Sep 2026 9 min read

Practical Frequency Counting in Python with collections.Counter

Counting repeated values looks simple until the surrounding code starts accumulating special cases. A plain dictionary can tally events, words, status codes, or inventory units, but the implementation also has to initialize missing keys, rank frequent values, merge counts, and decide what zero or negative counts mean. Python’s collections.Counter packages those operations into a dictionary-like type designed for counting hashable objects. It is useful when the problem is fundamentally about frequencies or multisets rather than arbitrary key-value storage.

Python 03 Sep 2026 10 min read

Parse and Render Shell Arguments Safely with Python shlex

Command-line text looks deceptively simple. Splitting on spaces works until an argument contains whitespace. Concatenating strings works until a filename contains shell metacharacters. Logging a list of arguments works, but the result may be difficult for a human to copy and inspect. Python’s shlex module handles a useful middle ground: shell-like lexical analysis for Unix-style command text. Its split(), quote(), and join() helpers let programs move deliberately between a string representation and a sequence of argument tokens.

Python 03 Sep 2026 9 min read

Model Finite States Clearly with Python Enum

Many programs represent a small fixed set of states with plain strings: status = "paid" That looks simple, but the string carries no built-in guarantee that it belongs to the set of states your application actually supports. A typo such as "paied" is still a valid Python string. So is an unexpected value received from a file, database, message, or HTTP request.

Python 03 Sep 2026 9 min read

Model Domain Constants Safely with Python enum

Strings and integers are convenient ways to represent states, modes, result codes, and permissions. They are also easy to mistype, mix with unrelated values, or pass through an API without making their meaning obvious. Python’s enum module lets a program give those values names and a controlled set of members. The benefit is not simply replacing constants with a class. A well-chosen enumeration defines the domain boundary: which values exist, how they compare, whether integer compatibility is intentional, and whether values may be combined.

Python 03 Sep 2026 11 min read

Layer Configuration Safely with Python ChainMap

Applications often build configuration from several sources. Command-line arguments may override environment-derived values, which in turn override built-in defaults. A straightforward implementation copies dictionaries and applies update() repeatedly. That works, but copying hides an important part of the design: configuration is not merely one dictionary. It is a precedence chain of independent sources. Python’s collections.ChainMap makes that relationship explicit. It presents several mappings as one lookup view without merging them first.

Python 03 Sep 2026 10 min read

Building Memory-Efficient Iterator Pipelines with Python itertools

Python programs often transform data in stages: read records, discard unwanted items, reshape values, group adjacent entries, and stop after enough output has been produced. A straightforward implementation may build a new list after every stage. That is easy to understand, but it can also allocate intermediate collections that the next stage immediately consumes. Iterator pipelines offer another model. Each stage requests values from the stage before it as needed. Python’s itertools module provides building blocks for this style, including tools for chaining inputs, taking slices from streams, computing running values, grouping consecutive records, and duplicating an iterator when two consumers genuinely need it.

Go 02 Sep 2026 7 min read

Short Reads and Exact-Length I/O in Go

Reading bytes in Go looks simple: allocate a buffer, call Read, and inspect the error. The subtlety is that io.Reader does not promise to fill the buffer in one call. A valid reader may return fewer bytes than requested even when more data will arrive later. That behavior matters for network protocols, binary file formats, framed messages, and any code that expects an exact number of bytes. Correct stream handling starts by matching the API to the requirement: use ordinary Read when partial progress is acceptable, and use helpers such as io.ReadFull when a fixed-size field must be complete.

Go 02 Sep 2026 4 min read

Reading Streams Correctly in Go with io.Reader

Go’s io.Reader interface is tiny: type Reader interface { Read(p []byte) (n int, err error) } Its small surface hides an important contract: a read is allowed to return fewer bytes than the buffer can hold, and it can return useful bytes together with an error. Correct stream processing must handle both cases.

Python 02 Sep 2026 9 min read

Practical Queues and Sliding Windows with Python deque

Many programs need a sequence that changes at both ends. A worker may append new jobs on the right and consume the oldest job from the left. A monitoring loop may keep only the most recent measurements. An algorithm may need to add or remove candidates from either side while scanning an input stream. A Python list is excellent when random access and operations near the right end dominate. It is a poor fit for a FIFO queue that repeatedly removes index zero, because the remaining list elements must be shifted. The collections.deque type is designed for efficient appends and pops at both ends.

Python 02 Sep 2026 10 min read

Practical Priority Queues in Python with heapq

Many programs need to repeatedly choose the most important pending item rather than process items in insertion order. Schedulers pick the next deadline, graph algorithms choose the lowest-cost candidate, and streaming systems keep only the best few observations seen so far. A sorted list can solve these problems, but maintaining full ordering is often unnecessary. Python’s heapq module provides a heap: a compact data structure that keeps one extreme element immediately available while doing only enough work to preserve that property.