Skip to content

Archive

Data Processing

5 articles
Go 13 Sep 2026 4 min read

Compact Adjacent Slice Values with slices.CompactFunc in Go

slices.CompactFunc removes repeated values only when equivalent elements are adjacent. That detail makes it distinct from general deduplication: the function operates on runs, keeps the first element from each run, and leaves separated matches alone. The custom equality function also allows compaction for structs and for equivalence rules that differ from Go’s == operator. Compaction is based on neighboring values The function has this signature: func CompactFunc[S ~[]E, E any](s S, eq func(E, E) bool) S For each run in which neighboring elements satisfy eq, the first element remains in the result. Consider case-insensitive string comparison:

Python 08 Sep 2026 8 min read

Compare Neighboring Values Lazily with itertools.pairwise

Many data-processing tasks are really questions about transitions: Did a measurement increase? How long was the gap between two events? Did a state change? Is a sequence sorted? These problems need neighboring values, not arbitrary pairs. Python 3.10 added itertools.pairwise() for exactly this pattern. It produces overlapping adjacent pairs lazily, which makes the intent clearer than manual indexing and lets the same code work with lists, generators, files, and other iterables.

Python 08 Sep 2026 9 min read

Batch Python Iterables Lazily with itertools.batched

Processing data in groups is common in Python. An application may send records to an API 100 at a time, insert rows into a database in manageable groups, or divide a stream of identifiers into work units without first loading the whole input into memory. Since Python 3.12, the standard library provides itertools.batched() for this pattern. It consumes an iterable lazily and yields tuples containing up to a requested number of items. Python 3.13 added a strict option for cases where an incomplete final batch should be treated as an error.

Python 05 Sep 2026 12 min read

Parse Binary Records Safely in Python with struct

Binary files and network messages often begin with fixed-width fields: a four-byte signature, a one-byte version, a two-byte payload length, or a four-byte identifier. Those fields are easy to describe on paper but surprisingly easy to parse incorrectly in code. The difficult part is not converting bytes to integers. It is preserving the binary layout contract: exactly which byte belongs to which field, which byte order is used, how wide each value is, and what should happen when the input is incomplete or malformed.

Python 03 Sep 2026 10 min read

Read and Write CSV Reliably in Python

CSV looks simple because a small file may resemble plain text with commas between values. That mental model breaks as soon as a field itself contains a comma, quote, or newline. Consider one valid record: 42,"Nguyen, Mai","Line one Line two" Splitting this text on commas cannot recover the three fields correctly. The comma inside the name is data, and the newline inside the quoted field belongs to the same record.