conference-paper
Open access
A Large Scale Study of Data Center Network Reliability
Research footprint
At a glance
- Citations
- 74
- References
- 85
- Comments
- 0
Paper overview
Abstract
The ability to tolerate, remediate, and recover from network incidents (caused by device failures and fiber cuts, for example) is critical for building and operating highly-available web services. Achieving fault tolerance and failure preparedness requires system architects, software developers, and site operators to have a deep understanding of network reliability at scale, along with its implications on the software systems that run in data centers. Unfortunately, little has been reported on the reliability characteristics of large scale data center network infrastructure, let alone its impact on the availability of services powered by software running on that network infrastructure.
Record transparency
Publication details
- DOI
- 10.1145/3278532.3278566
- OpenAlex
- W2903107001
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.