Resilience, Observability, and Performance at Infrastructure Scale

Authors

Rohit Gorle
Senior Software Engineer, Oracle, Austin, Texas, United States

Synopsis

Infrastructure supporting Cloud and other large-scale applications often demand high resilience against faults and performance agility in responding to varying usage patterns. Understanding and addressing these challenges are mathematical, engineering, and operational disciplines in their own right. Clear, operationally relevant definitions, models, and technologies can simplify parts of the problems of large-scale resiliency, observability, and performance. Nevertheless, mathematical foundations and assistive technologies cannot relieve engineering and operations teams of planning for and working through the wide variety of faults that occur at scale.

Resilience of large-scale infrastructures implies the joint property that the infrastructure is always operational (cannot be irrecoverably broken), and that service disruptions are not only avoided but detected and reported, ideally by those affected. Team central to maintaining this resilience formally define the following essential concepts: Recoverable Fault—non-catastrophic fault in infrastructure layer that can be detected and is independently resolvable. Capacity/fault-resilience—resilience property such that, unless all nodes of a certain type fail simultaneously, no service disruption occurs. Performance Testing and Benchmarking is one of the essential considerations of any system targeting cloud and other large-scale application deployments, whether for customer-facing usage or for internal project development.

Downloads

Published

17 August 2026

How to Cite

Gorle, R. . (2026). Resilience, Observability, and Performance at Infrastructure Scale. In The Governed Intelligence Fabric: Architecting Trusted Data Platforms for Infrastructure at Scale (pp. 113-126). Deep Science Publishing. https://doi.org/10.70593/978-81-69589-32-1_9