Skip to content

Archive

Testing

55 articles
Software Engineering 11 Sep 2026 9 min read

Parameterize from Above: Make Dependencies Explicit

Parameterize from Above: Make Dependencies Explicit A method can look self-contained while quietly deciding which clock, repository, client, or filesystem implementation the application must use. That hidden construction becomes painful when a test needs a controlled collaborator or production needs a different implementation. Parameterize from Above is a small design move: instead of constructing a dependency inside the code that uses it, accept that dependency from a caller at a higher level. The caller becomes responsible for choosing and constructing the collaborator.

Software Engineering 11 Sep 2026 9 min read

Mutation Testing: Check Whether Tests Detect Broken Behavior

Mutation Testing: Check Whether Tests Detect Broken Behavior A test suite can execute every line of a function and still miss a defect. The tests may call the right code but make weak assertions, cover only one outcome, or never check a boundary condition. Mutation testing probes that gap by making small changes to production code and running the relevant tests against each changed version. If a test fails, the change is said to be killed. If all tests still pass, the mutation survives and deserves inspection.

Software Engineering 11 Sep 2026 8 min read

Humble Object Pattern for Hard-to-Test Boundaries

Humble Object Pattern for Hard-to-Test Boundaries Some code is difficult to test for reasons that have little to do with its business rules. A screen handler may depend on a UI framework. A file watcher may need operating-system events. A scheduled job may be invoked by infrastructure that is awkward to reproduce in a unit test. A common mistake is to put more logic inside that difficult boundary. Tests then need the framework, filesystem, clock, process, or device just to check an ordinary decision.

Software Engineering 11 Sep 2026 8 min read

Consumer-Driven Contract Tests for Service Compatibility

Consumer-Driven Contract Tests for Service Compatibility Two services can pass their own test suites and still fail when deployed together. A provider may rename a field, narrow an accepted value, change a status code, or remove an endpoint. Its internal tests can remain green because those tests describe the provider’s own view of correct behavior. A consumer can still depend on the old interaction. Consumer-driven contract testing turns selected consumer expectations into executable contracts. The consumer records the interactions it requires. The provider then verifies those contracts against its implementation.

Software Engineering 10 Sep 2026 9 min read

Testing Asynchronous Behavior with Eventual Assertions

Testing Asynchronous Behavior with Eventual Assertions A test starts background work and then checks the result. On a fast machine the work finishes first and the test passes. Under CI load, the assertion runs a few milliseconds earlier and fails. Someone adds sleep(1 second). The failure disappears, but every successful run now pays a full second, and a sufficiently slow run can still fail. The problem is not that the test needs a longer delay. The test doesn’t know exactly when the result will become observable.

Go 10 Sep 2026 7 min read

Test Concurrent Go Code with testing/synctest

Tests for concurrent Go code often become timing tests by accident. A goroutine starts, the test sleeps for 20 milliseconds, then checks whether something happened. That can pass thousands of times and still fail on a loaded CI runner because the sleep never proved the goroutine reached the state you cared about. Go’s testing/synctest package gives these tests a better model. It runs code inside an isolated bubble where time is virtualized, and synctest.Wait can synchronize the test with goroutines in that bubble. The result is especially useful for code built around timers, context deadlines, retries, and background goroutines.

Software Engineering 10 Sep 2026 8 min read

Functional Core, Imperative Shell for Testable Code

Functional Core, Imperative Shell for Testable Code A function that calculates a decision, reads the clock, queries a database, sends a message, and writes a log can be difficult to test for a simple reason: its business rules and its interactions with the outside world are tangled together. Testing one rule may require arranging several dependencies that have nothing to do with that rule. Functional core, imperative shell is a design approach for separating those concerns. The functional core contains deterministic decision-making: given the same explicit inputs, it produces the same result without performing external side effects. The imperative shell handles effects such as reading data, obtaining the current time, calling services, and persisting results.

Software Engineering 10 Sep 2026 9 min read

Functional Core, Imperative Shell for Managing Side Effects

Functional Core, Imperative Shell for Managing Side Effects Business logic is often easy to describe and hard to test because it sits between database reads, network calls, clocks, queues, and file writes. A pricing rule that should be a few comparisons can become tangled with fetching a customer, saving an order, and sending a notification. Functional core, imperative shell is a design approach for separating those concerns. The functional core makes decisions from explicit input values and returns values describing the result. The imperative shell obtains those inputs and performs the required side effects.

Software Engineering 10 Sep 2026 9 min read

Differential Testing for Behavior-Preserving Changes

Differential Testing for Behavior-Preserving Changes Replacing working code is risky when the requirement is “change the implementation, not the behavior.” A rewritten parser may accept a different edge case. A faster pricing engine may round one value differently. A new library may return the same records in a different order. Ordinary tests help, but they only cover cases and assertions someone thought to write. Differential testing adds another source of evidence: run the old and new implementations on the same inputs, compare their observable results, and investigate differences.

Software Engineering 10 Sep 2026 9 min read

Creating Seams to Test Hard-to-Change Code

Creating Seams to Test Hard-to-Change Code A method can contain a simple business rule and still be difficult to test. The difficulty often comes from everything attached to the rule: the current clock, a network client, a filesystem call, a global configuration object, or a constructor that creates its own dependencies. When rewriting the surrounding code would be risky, a seam can give you a smaller move. A seam is a place where you can change which behavior the code uses without changing the code that makes the business decision. In tests, that lets you replace an awkward dependency with controlled behavior.

Software Engineering 10 Sep 2026 12 min read

Choosing Sociable or Solitary Unit Tests

Choosing Sociable or Solitary Unit Tests A unit test fails after a harmless refactor. The production behaviour is still correct, but the test expected an internal collaborator to receive exactly three calls. Elsewhere, a different unit test passes even though two real classes no longer work together, because both sides were replaced with mocks. Both tests are isolated in some sense, yet they give poor feedback for different reasons. The useful question isn’t simply whether unit tests should use mocks. It is where the test boundary should be.

Software Engineering 09 Sep 2026 8 min read

Making Complex Rules Visible with Decision Tables

Conditional code often starts clearly. One condition becomes two, then a special case appears, and eventually nobody can answer a simple question with confidence: have we covered every meaningful combination? The problem is not necessarily that if statements are bad. The problem is that branching code makes a set of rules visible one execution path at a time. When several independent conditions affect one decision, developers must mentally reconstruct the whole rule set from those paths.

Python 09 Sep 2026 13 min read

Make Warning Tests Concurrency-Safe with Context-Aware Warnings

Python’s warnings.catch_warnings() is convenient in tests, compatibility shims, and small diagnostic scopes. It lets code temporarily change warning filters and then restore the previous state. That model becomes harder to reason about when several threads or asynchronous tasks use it at the same time. Historically, catch_warnings() manipulated process-global state in the warnings module. Two overlapping contexts could therefore interfere with each other. Python 3.14 adds an opt-in context-aware mode that changes this behavior. When sys.flags.context_aware_warnings is true, catch_warnings() stores its filtering state in a context variable instead of mutating the same global warning state for every concurrent execution path.

Software Engineering 09 Sep 2026 8 min read

Isolating Hard-to-Test Code with a Humble Object

Some code is difficult to test for reasons that have little to do with the behavior you care about. A screen handler may require a UI framework. A file watcher may need operating-system events. A message consumer may only run inside a broker callback. Tests become slow or fragile because ordinary business decisions are trapped inside code that is expensive to execute in isolation. The Humble Object pattern addresses this by separating the difficult boundary from the logic behind it. The boundary object stays deliberately small: it translates an external event into plain data, calls ordinary code, then translates the result back. The decisions move into code that can be exercised without the framework or environment.

Software Engineering 09 Sep 2026 8 min read

Finding Seams for Safer Code Changes

Sometimes a small code change feels much larger than the requirement. You want to test one decision, replace one dependency, or alter one behavior, but the code gives you no place to do that without executing or editing a large surrounding block. A useful way to reason about this problem is to look for a seam: a place where you can change the behavior of a program without editing the code that uses that behavior. A seam might be a function parameter, an object boundary, a configurable callback, or another point where one implementation can be substituted for another.

Software Engineering 09 Sep 2026 9 min read

Decision Tables for Complex Business Rules

Business rules often start as a few harmless conditions. Then another exception arrives, followed by a special customer type, a threshold, and a fallback. The code still runs, but reviewing it becomes difficult because the real question is no longer “what does this if statement do?” It is “have we handled every meaningful combination of conditions, and do any rules disagree?” A decision table makes those combinations explicit before they are buried in branching code. It lists the conditions that matter, the relevant combinations of those conditions, and the outcome for each combination.

Python 09 Sep 2026 10 min read

Compare Python Syntax Trees with ast.compare

Tools that rewrite Python source often need to answer a deceptively simple question: did this transformation preserve the syntax tree that matters? Before Python 3.14, a common solution was to serialize both trees with ast.dump() and compare the resulting strings. That works in small tests, but it turns a structural question into a formatting contract. Python 3.14 adds ast.compare(), a recursive AST comparison helper that expresses the intent directly. This is especially useful for formatters, codemods, linters, source generators, refactoring tools, and tests that round-trip Python code.

Software Engineering 09 Sep 2026 8 min read

Characterization Tests Before Changing Legacy Code

Changing unfamiliar code creates a difficult question: how do you know a refactor preserved behavior when nobody can state exactly what the current behavior is? Existing unit tests may be sparse. Documentation may describe the intended rules but not the edge cases the system actually implements. Some odd behavior may even have become a dependency for callers. A characterization test helps in this situation. Instead of starting from what the code ought to do, it records what the code does now for a carefully chosen input. That gives you a behavioral reference point before you change the implementation.

Software Engineering 09 Sep 2026 9 min read

Building a Walking Skeleton Before Filling In the System

A team can make steady progress inside individual components and still discover late that the system does not work as a whole. The application starts differently in production, two modules disagree about a contract, a deployment is missing configuration, or the real request path was never exercised until several weeks of work depended on it. A walking skeleton is a small, working path through the system that connects the important architectural pieces before those pieces contain much functionality. It does not prove that the product is complete. It proves that a thin version of the system can travel from an external entry point, through the chosen boundaries, to an observable result.

Software Engineering 08 Sep 2026 9 min read

Use Mutation Testing to Find Weak Tests

A test suite can execute every line of an important function and still fail to detect that the function is wrong. Coverage tells you which code ran. It does not tell you whether the assertions would notice a meaningful defect in that code. Mutation testing approaches the problem from the other direction. A mutation testing tool makes small, deliberate changes to production code and runs the tests. If the tests fail, they detected the change. If the tests still pass, the altered behavior has exposed a possible weakness in the suite.

Software Engineering 08 Sep 2026 9 min read

Testing Test Suites with Mutation Testing

A test suite can execute every line of an important function and still fail to notice that the function is wrong. Coverage tells you which code ran during tests. It does not tell you whether the tests would detect a meaningful mistake in that code. Mutation testing examines that missing question. A mutation testing tool makes small changes to production code, one change at a time, and runs the relevant tests. If the tests fail, they detected the changed behavior. If they still pass, the altered code has exposed something worth investigating.

Software Engineering 08 Sep 2026 9 min read

Testing Resilience with Fault Injection

A service can pass every normal-path test and still behave badly when a dependency times out, a write fails halfway through, or a connection disappears at an inconvenient moment. The problem is often not missing error handling. It is that the team has never observed whether the error handling produces the behavior they expect. Fault injection is the deliberate introduction of a controlled failure into a system or test. Instead of waiting for a real dependency to fail, you make a specific failure happen and observe the consequence.

Software Engineering 08 Sep 2026 9 min read

Testing Hard-to-Test Code with Humble Objects

Some code is difficult to test for reasons that have little to do with the behavior you care about. A user-interface callback depends on a framework event loop. A scheduled job reads the clock, queries a service, writes a file, and decides whether to alert someone. A device handler receives data through an operating-system API before applying a simple rule. When the decision and the awkward environment live in the same unit, every test inherits the environment’s complexity. The test may need framework setup, timing control, filesystem state, or several mocks just to reach a small branch.

Software Engineering 08 Sep 2026 7 min read

Testing Complex Rules with Decision Tables

A rule can be easy to understand one condition at a time and still be difficult to test correctly when several conditions interact. Consider a refund policy. A refund depends on whether the order is within 30 days, whether the item is damaged, and whether it was marked final sale. Writing a few examples from memory can miss an important combination. Adding every possible combination can create noisy tests that repeat the same reasoning.