conference-paper Open access

A Large Scale Study of Data Center Network Reliability

Research footprint

At a glance

Citations
74
References
85
Comments
0
Paper overview

Abstract

The ability to tolerate, remediate, and recover from network incidents (caused by device failures and fiber cuts, for example) is critical for building and operating highly-available web services. Achieving fault tolerance and failure preparedness requires system architects, software developers, and site operators to have a deep understanding of network reliability at scale, along with its implications on the software systems that run in data centers. Unfortunately, little has been reported on the reliability characteristics of large scale data center network infrastructure, let alone its impact on the availability of services powered by software running on that network infrastructure.

Record transparency

Publication details

DOI
10.1145/3278532.3278566
OpenAlex
W2903107001
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.