Resilience, Observability, and Performance at Infrastructure Scale
Synopsis
Infrastructure supporting Cloud and other large-scale applications often demand high resilience against faults and performance agility in responding to varying usage patterns. Understanding and addressing these challenges are mathematical, engineering, and operational disciplines in their own right. Clear, operationally relevant definitions, models, and technologies can simplify parts of the problems of large-scale resiliency, observability, and performance. Nevertheless, mathematical foundations and assistive technologies cannot relieve engineering and operations teams of planning for and working through the wide variety of faults that occur at scale.
Resilience of large-scale infrastructures implies the joint property that the infrastructure is always operational (cannot be irrecoverably broken), and that service disruptions are not only avoided but detected and reported, ideally by those affected. Team central to maintaining this resilience formally define the following essential concepts: Recoverable Fault—non-catastrophic fault in infrastructure layer that can be detected and is independently resolvable. Capacity/fault-resilience—resilience property such that, unless all nodes of a certain type fail simultaneously, no service disruption occurs. Performance Testing and Benchmarking is one of the essential considerations of any system targeting cloud and other large-scale application deployments, whether for customer-facing usage or for internal project development.










