7b127 No.1966
just found this breakdown on how to handle alerts without panicking. it argues that ops teams need to answer three specific questions before touching anything, which is basically
the key to avoiding a total meltdown ]. i think the hardest part is keeping ur incident_response_logs clean enough to actually see the pattern, but
dont ignore the architecture side of things. anyone else find that
properly structured services make triage much faster?
https://thenewstack.io/build-resilient-service-architecture/