All articles Engineering

Tests that fail when you break them

A passing test proves nothing until you have watched it fail for the right reason.

Photo: Shixart1985 (CC BY 2.0) / Wikimedia Commons

We write a lot of tests. The discipline that makes them worth anything is not writing them — it is breaking the code on purpose afterwards and confirming the test notices.

It has caught our own tests being useless more than once. Three examples from recent weeks.

A test that passed for the wrong reason

We assert that a crypto payment already visible on-chain cannot be cancelled. The test checked that cancelling was refused.

Then we removed the on-chain guard entirely — and the test still passed. Cancelling was still refused, but by a *different* check further down that happened to catch it.

The assertion was satisfied by a guarantee we had not meant to rely on. If somebody later removed the second check too, the test would have started failing for a reason nobody would connect to the first change.

We tightened it to assert *which* refusal, by matching the message. It now fails as it should.

A test that was measuring the fixture

Another suite asserted "two payments are pending". It passed, until earlier assertions in the same file left extra rows behind, at which point it failed for reasons that had nothing to do with the code under test.

Counted as a delta against a baseline taken at the top, it says what the *decision* did rather than what the fixture happened to contain. It survives reordering, and it survives being run twice — which is a good property, because a suite you cannot run twice is a suite people stop running.

A test that found the wrong file

A guard that fails if any source file types a currency symbol flagged one of our own blog articles — which quotes the symbol while describing that exact bug.

The right answer was not to weaken the rule. It was to exempt that one file with a stated reason, and add a second assertion that the exempted file never formats money itself. An exemption without a boundary is just a hole.

The practice

After writing a test: revert the fix, run it, watch it fail, read the failure. Then restore.

Three things to check in that failure:

  1. Does it fail at all? If not, it is not testing what you think.
  2. Does it fail for the right reason? A different guard catching it is a false pass waiting to happen.
  3. Is the message actionable? In six months that line is all somebody has.

This costs about ninety seconds per test. It is the highest-return ninety seconds in the job, and it is the step almost everybody skips — because the test is green, and green feels like done.

Start building

Lock a key to your domain in about five minutes.

Get started