Skip to content

Archive

Unicode

3 articles
Go 13 Sep 2026 4 min read

Transform Unicode Text with strings.Map in Go

strings.Map applies one function to every rune in a UTF-8 string and builds a string from the returned runes. A mapping function can preserve a rune, replace it, or remove it entirely. That makes the API a compact fit for transformations whose rule is naturally expressed one Unicode code point at a time. mapped := strings.Map(func(r rune) rune { if r == '_' { return '-' } return r }, input) The operation is rune-oriented rather than byte-oriented. ASCII input still follows the same contract, but multibyte UTF-8 sequences arrive at the callback as decoded rune values.

Go 13 Sep 2026 4 min read

Split Text with Custom Rune Boundaries Using strings.FieldsFunc in Go

strings.FieldsFunc treats selected Unicode code points as boundaries and returns the non-empty text between them. That behavior fits inputs where separators belong to a class rather than one fixed substring: commas and semicolons, several punctuation marks, or any rune accepted by a deterministic predicate. fields := strings.FieldsFunc(input, func(r rune) bool { return r == ',' || r == ';' }) For alpha,,beta;gamma;, the result is []string{"alpha", "beta", "gamma"}. Consecutive matching runes form a boundary region, and matching runes at either edge do not produce empty elements.

Python 04 Sep 2026 10 min read

Decode Streaming Text Safely in Python with Incremental Codecs

Network sockets, compressed streams, subprocess pipes, and chunked file reads often deliver bytes in arbitrary pieces. If those bytes represent text, it is tempting to decode each piece immediately: for chunk in byte_chunks: text = chunk.decode("utf-8") process(text) That works only when every chunk happens to end on a character boundary.