Skip to content

Archive

Observability

6 articles
Linux 16 Sep 2026 5 min read

Linux PSI Separates Partial and Total Resource Stalls

A Linux host can report modest CPU utilization while runnable work is delayed, or ample memory capacity while tasks repeatedly stall in reclaim. Utilization counters describe resource activity; pressure stall information records time in which work cannot make progress because a resource is contended. PSI exposes that lost execution opportunity through CPU, memory, and I/O pressure files. Its central distinction is between a stall affecting at least one task and a stall that leaves every non-idle task unable to make progress.

Python 09 Sep 2026 12 min read

Debug Running Asyncio Services with pstree and ps in Python 3.14

An asynchronous service can be alive while making no useful progress. The process still responds to signals. CPU usage may be low. The event loop is still running. Yet a request, worker, or shutdown path appears stuck somewhere inside a chain of coroutines. Traditional stack traces are only part of the answer. An asyncio application is organized around tasks and await relationships, so the useful question is often not merely “where is this thread?” but “which task is waiting for which other task?”

Go 09 Sep 2026 8 min read

Control Structured Log Output in Go with slog.LogValuer

Passing a struct directly to slog is convenient until that struct grows a field that should never appear in logs. An access token, session secret, internal note, or large payload can turn an ordinary diagnostic line into a security problem or an expensive blob of noise. Go’s slog.LogValuer interface gives a type control over its own structured log representation. Instead of teaching every call site which fields are safe, you can define that representation next to the type and let slog use it wherever the value is logged.

Python 09 Sep 2026 12 min read

Attach to Running Python Processes with sys.remote_exec in Python 3.14

A production Python process can be healthy enough to stay alive while still being difficult to understand. Perhaps one thread appears stuck. Memory is growing but the application has no diagnostic endpoint. A profiler was not enabled before startup. Restarting the process would erase the state you need to inspect. Python 3.14 adds a new CPython capability for this situation: sys.remote_exec(). It lets one Python process request that a .py file be executed by another running CPython process. The target executes that file on its main thread at a safe execution point.

Software Engineering 02 Sep 2026 10 min read

Designing Structured Logs for Production Debugging

Production debugging often starts with a deceptively simple question: what happened to this request? Plain-text logs can answer that question in small systems, but they become difficult to search reliably when message wording changes, multiple services participate in one operation, or operators need to aggregate millions of records. Structured logging addresses that problem by representing important context as named fields instead of embedding everything in prose. The goal is not to turn every variable into a log field. A useful log schema captures stable facts about an event, preserves enough correlation context to connect related work, and avoids recording data that creates security or privacy risk.

Go 01 Sep 2026 7 min read

Structured Logging in Go with slog

Plain text logs are easy to print but difficult to query reliably. Once an application runs across multiple processes or containers, operators usually need to filter events by fields such as HTTP status, request path, customer ID, or latency rather than search arbitrary strings. Go’s standard library includes log/slog for structured logging. It was added in Go 1.21, so the examples in this article require Go 1.21 or newer. What structured logging changes A traditional log message often embeds data inside prose: