Skip to content

Case Studies

Real builds, generalised. Each case study takes a system that was actually designed and shipped under a deadline, strips it to the concepts that transfer, and records the sequence it was built in.

Prerequisites: varies by study. Each one names its own, and every study links back to the track that teaches the underlying theory.

  • What these are: worked builds. Problem, concepts, the decisions and their trade-offs, and the order the thing was actually assembled in.
  • What these are not: project archives. No original source is preserved. Every implementation here was rewritten to be self-contained, dependency-free, and runnable, so it teaches the idea rather than documenting a codebase.
  • Estimated time: 1-2 days each, or an afternoon if you only read.
  • The build order is the lesson. Knowing that a matching engine needs price-time priority is cheap; knowing to write the O(n) version first and earn the heap is what actually separates outcomes under time pressure.
  • Every study has one invariant that makes testing possible. A crossed book, a double-fired event, an unbounded query, a rising inertia, a record that does not replay. Find that property and validation stops being guesswork.
  • The interesting decisions are refusals. Not adding durability, not letting the model be the safety boundary, not shipping exactly-once. Each study states what it deliberately did not build.
  • Read the README, then run the implementation. Every one is standard library only and executes its own tests: python <file>.py.
  • Then delete a safeguard and watch a test fail. Remove the row cap, remove the primary key, skip the zero-quantity removal. The tests exist to catch exactly those, and breaking them on purpose is the fastest way to see why they matter.
  • Read the build log last, and compare it to how you would have sequenced it.
#StudyDomainRunnable
01Order Book MatchingData structures under a latency budgetmatching_engine.py
02Grounded SQL AgentLLM tool loops, grounding, and safetysafety_gate.py
03Exactly-Once Event APIAPI design and processing guaranteesevent_api.py
04K-Means OptimizationOptimising an algorithm in tierskmeans_ladder.py
05Agent Evaluation HarnessGrading an agent that changes stateeval_harness.py

Four of the five studies are the same move applied to different domains: build the obvious correct version, name its bottleneck precisely, then earn each improvement.

StudyBaselineEarned improvementWhat the jump costs
Order bookFlat list, linear scanHeap of price levels, then a tick-indexed ladderMemory, and an assumption about price range
K-meansRandom init, fixed iterationsk-means++, then incremental updates and distance boundsNothing algorithmic; only code complexity
SQL agentOne tool, no groundingSchema discovery, bounded retry, deterministic gateTokens per query, and latency
Event APICheck-then-act dedupConstraint dedup, then atomic claimNothing; the correct version is also simpler
Eval harnessCompare final rowsReplay verification, then oracles and mutantsStorage per call, and a suite to maintain

The last row is worth sitting with. Dedup done properly is less code than dedup done by checking first, and it is correct under concurrency. Not every improvement is a trade-off; some are just the right answer written down.

StudyTheory lives in
Order Book MatchingAlgorithms (heaps, ordering, amortised analysis)
Grounded SQL AgentLLM Evaluation, Retrieval & RAG, Authorization
Exactly-Once Event APIDistributed Workers, Durable Orchestration
K-Means OptimizationML & Statistics, Statistical Learning

For the failure modes these builds were consciously avoiding, see Lessons from Practice.