Architecture notebook

High availability / Informed by production work

Staying available while things fail

Availability is not one property but a set of decisions: how many instances run, how a failed one is detected, what happens to in-flight work, and what the system does when a dependency is the thing that broke.

  1. 01Load balancer
  2. 02Stateless replicas
  3. 03Health probes
  4. 04Managed datastore
Probe → drain → reschedule → degrade rather than fail

01 Make instances replaceable

Services hold no session state, so any replica can serve any request and a lost instance costs nothing but capacity. State lives in the datastore or the cache, not in the process.

02 Let the platform detect and replace

Readiness and liveness probes distinguish "not ready yet" from "broken", so traffic is withheld during startup and a wedged container is restarted rather than left in rotation.

03 Deploy without a gap in service

Rolling updates with a surge replica and graceful shutdown let in-flight requests finish while new pods come up. Migrations stay backward compatible so both versions can run at once.

04 Reliability: contain the dependency that fails

Timeouts, bounded retries, and circuit breakers stop a slow dependency from consuming every thread. Where the feature allows it, degrading — a cached answer, a queued action — beats returning an error.

05 Trade-off: redundancy costs money and clarity

Running spare capacity and multi-zone datastores raises both spend and operational surface. The question is always which failures are worth the standing cost of surviving.