A System That Is Always Up and Frequently Wrong
System Design

A System That Is Always Up and Frequently Wrong

Subtitle: Uptime is not the same as being right. Left column - Availability - did it answer?: - Uptime: something came back - Counts responses, not truth - A stale cache scores high here - Redundancy is the lever - The green dashboard lives here Right column - Reliability - was it right?: - The answer was actually correct - It did what it claimed to do - Refusing a risky charge scores high - Checks and idempotency are levers - Silent corruption hides here Simple difference: - Availability = good replies / all replies - Reliability = mean time between failures Different remedies: - Redundancy raises availability - Correctness checks raise reliability - A replica cannot make wrong data right Five nines hides this: - 99.999% says nothing about correctness - You can hit the SLO serving bad results - The dashboard stays green throughout - A liveness probe is not a correctness check Honest test - three steps: 1. Make a dependency return a wrong value 2. Watch which alarms actually fire 3. Only latency moved? Never measured reliability