Skip to content

The Testing Mentality

  • Testing is a mentality, not a phase. It’s the habit of treating every claim — about code, infra, docs, data, or a teammate’s assumption — as a hypothesis you can cheaply falsify.
  • The goal is confidence per unit cost. Many fast, isolated tests; few slow, broad ones. Shape the pyramid, not an ice-cream cone.
  • Test behavior, not implementation. A test that breaks when you refactor (without changing behavior) is a liability, not an asset.
  • In production, testing continues — monitoring, canaries, and chaos experiments are tests you run against the real system.
  • Read four Testing on the Toilet posts; each is a 5-minute mental model.
  • Take one function with example-based tests and add a property-based test. Watch it find an edge case you didn’t think of.
  • For a system you run: write the failure hypothesis you’re most afraid of, then design the smallest experiment that would confirm or refute it.

A test is a falsifiable claim plus a cheap experiment that tries to break it. The “testing mentality” generalizes that beyond unit tests: a type signature tests a claim about shapes; a code review tests a claim about readability; an alert tests a claim about health; a canary tests a claim about a release; an ADR’s “Consequences” section tests a claim about the future. Engineers with the mentality reflexively ask, “what would prove this wrong, and how cheaply can I find out?” — and they ask it about their own beliefs first.

1. The test pyramid (shape your confidence by cost)

Section titled “1. The test pyramid (shape your confidence by cost)”
graph TD
    E2E["End-to-End / System<br/>few · slow · brittle · highest fidelity"]
    INT["Integration / Medium<br/>some · real dependencies via fakes/containers"]
    UNIT["Unit / Small<br/>many · fast · isolated · deterministic"]
    UNIT --> INT --> E2E
    style UNIT fill:#1168bd,color:#fff
    style INT fill:#438dd5,color:#fff
    style E2E fill:#85bbf0,color:#000
  • Small (unit) — no I/O, no clock, no network; milliseconds; run on every save. Most of your tests.
  • Medium (integration) — real dependencies via fakes or Testcontainers (a real Kafka/Postgres in Docker); seconds.
  • Large (E2E) — the whole system; minutes; flaky by nature, so keep few and treat flakes as bugs.

Anti-pattern — the ice-cream cone: mostly E2E tests, few unit tests. Slow, flaky, and they tell you something broke without telling you what. Invert it.

The single highest-leverage rule. A good test states what the code promises, so it survives refactoring and fails only on real regressions.

# BAD -- couples to implementation. Refactoring the cache breaks this test
# even though behavior is unchanged.
def test_calls_redis_setnx_once():
worker.process(msg)
redis_mock.setnx.assert_called_once()
# GOOD -- asserts the promise: a duplicate message is processed exactly once.
def test_duplicate_message_processed_once():
worker.process(msg); worker.process(msg) # same message twice
assert db.count(order_id=msg.order_id) == 1 # observable behavior

This is Hyrum’s Law turned to your advantage: depend only on the contract you intend to keep.

DoubleWhat it isGoogle’s stance
FakeA working lightweight impl (in-memory DB, in-memory Kafka)Preferred — behaves correctly, tests survive refactors
StubReturns canned answersFine for simple inputs
MockAsserts how it was calledUse sparingly — mock-heavy tests are brittle and test implementation

Mock-heavy suites are the most common reason a test breaks during a no-op refactor. Reach for a fake first.

  • Property-based testing — assert invariants over generated inputs (Hypothesis, proptest, QuickCheck). E.g. “encode then decode returns the original,” “the consumer never skips an offset.” Finds the edge cases you’d never enumerate by hand.
  • Fuzzing — feed random/adversarial bytes to parsers and decoders; the bug class behind a huge share of CVEs. cargo fuzz, go test -fuzz, libFuzzer.
  • Golden / snapshot tests — pin a complex output (a rendered diagram, an API response) and diff future runs. Cheap regression net for serializers and generators.
  • Contract tests — the producer and consumer of an API each test against a shared contract, so a breaking change is caught in CI rather than in production (Pact, schema registries for Kafka).

5. Testing things that aren’t application code

Section titled “5. Testing things that aren’t application code”

The mentality applies everywhere knowledge can be wrong:

  • Infrastructure — terraform plan in CI, kubeconform/kubeval to validate manifests, policy tests (OPA/Conftest) for “no container runs as root.” A misconfigured manifest is a bug.
  • Data — pipeline tests with Great Expectations / dbt tests: “no nulls in order_id,” “row count within 3σ of yesterday.” Bad data is a production incident.
  • Documentation — run the code samples (Documentation & Technical Writing). An example that doesn’t compile is a failing test.
  • Diagrams — d2 *.d2 in CI fails if a diagram no longer compiles (Diagramming & the C4 Model).

6. Testing in production (because you already are)

Section titled “6. Testing in production (because you already are)”

You cannot fully reproduce production, so test against it deliberately rather than pretending staging is enough:

  • Canary / progressive rollout — ship to 1% of traffic, watch the metrics, then ramp. A controlled experiment on the real system.
  • Synthetic monitoring — a bot continuously exercises the critical path; the first to know is you, not the customer.
  • Chaos engineering — inject failure (kill a pod, add latency, partition the network) to test the hypothesis “the system tolerates this.” Start with a small blast radius and a defined steady-state metric.
  • SLOs and error budgets — the SLO is a testable claim about reliability; burning the error budget is the test failing in slow motion (see Observability).
TechniqueWhen to apply
Shape the test pyramidDesigning any service’s test strategy
Test behavior, not implementationEvery test you write
Prefer fakes over mocksAny test needing a dependency
Property-based testingEncoders, parsers, invariant-heavy logic
FuzzingAny code that parses untrusted input
Contract testsAny producer/consumer or service boundary
Manifest/policy testsAny IaC or Kubernetes change
Data quality testsAny data pipeline
Canary + synthetic monitoringEvery production release
Chaos experimentsSystems claiming resilience — prove it

In the course this topic hosts the testing ladder: one practice module per rung, each placed in the pass where your system first needs that kind of test. Your own tests are graded by mutation testing (what they catch, course principle P10), so each rung is practised on the system you are building, not on toy code.

ModuleTopicKindPass
craft.03TDD, unit tests, and how you are graded (rungs R0 to R3): includes the mutation-testing primer (what a mutant is, killed vs survived, how the score and required semantic mutants work) before the first mutation gradepractice2
craft.04Property-based tests (R4) with Hypothesis, proptest, rapid, and ol_prop.hpractice3
craft.07Mutation testing in depth: equivalent mutants, semantic mutants from pitfalls, reading survivorspractice4
craft.05Oracles, golden and differential tests, gradcheck as a test (R5)practice5
craft.06Benchmarks and perf gates (R7)practice6
craft.20Contract tests (R6): consumer-driven tests for gateway to enginepractice7
craft.21Resilience tests (R10)practice8
craft.22Model evals as tests (R8)practice9
craft.23Agent evals as tests (R9)practice10
#ModuleChapterKindPass
1craft.03TDD, unit tests, and how your tests are gradedpractice2
2craft.04Property-based tests (R4) with Hypothesis, proptest, rapid, and ol_prop.hpractice3
3craft.05Oracles, golden and differential tests, gradcheck as a test (R5)practice5
4craft.06Benchmarks and perf gates (R7)practice6
5craft.07Mutation testing in depth: equivalent mutants, semantic mutants from pitfalls, reading survivorspractice4
6craft.20Contract tests (R6): consumer-driven tests for gateway to enginepractice7
7craft.21Resilience tests (R10)practice8
8craft.22Model evals as tests (R8)practice9
9craft.23Agent evals as tests (R9)practice10
ConceptConnected TrackHow
Test pyramid, fakes, Hyrum’s LawSoftware Engineering at GoogleGoogle’s testing chapters are the source
Property-based & fuzz testingAlgorithmsInvariants are properties of the algorithm
Canary, SLOs, chaosObservabilityYou test in prod through telemetry
Data quality testsOrchestration & Modelingdbt tests gate the pipeline
Manifest/policy testsContainers, Kubernetes & WorkloadsValidate YAML before it reaches the cluster
CompanyHow This AppearsFocus
GoogleTesting on the Toilet, fakes-first, presubmit testingTesting at scale
NetflixInvented chaos engineering (Chaos Monkey)Resilience by experiment
AmazonOperational readiness reviews, canary deploysTest in prod safely
AnthropicRigorous evals + safety testing of models and systemsFalsify before shipping
Any Staff+ roleYou set the team’s testing bar and quality cultureMentality, not coverage %