TY - GEN
T1 - Avicenna
T2 - 2026 European Conference on Computer Systems, EUROSYS 2026
AU - Hodsdon, Christopher
AU - Qin, Zijian
AU - Ngo, Khiem
AU - Sen, Siddhartha
AU - Katz-Bassett, Ethan
AU - Lloyd, Wyatt
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s)
PY - 2026/4/26
Y1 - 2026/4/26
N2 - Geo-distributed replicated state machines (RSMs) are at the heart of many production distributed systems, offering linearizability and fault tolerance via consensus protocols. Most existing protocols target crash fault tolerance, however, and are vulnerable to fail-slow faults, where a single slow replica can significantly degrade system latency. Existing protocols that tolerate fail-slow faults do so with much higher normal-case latency in geo-distributed settings. This paper presents Avicenna, the first consensus protocol for geo-distributed RSMs that maintains low normal-case latency while tolerating a single fail-slow replica. Avicenna uses a single leader to order commands, naturally tolerating a fail-slow follower. To tolerate a fail-slow leader, Avicenna compares the current latency with the counterfactual latency clients would experience if a different replica, the shadow leader, were the leader. When that comparison indicates the current leader might be slow, Avicenna quickly promotes the shadow leader with a fast leader rotation protocol. Our evaluation shows Avicenna has the same normal-case latency as Multi-Paxos while tolerating fail-slow faults.
AB - Geo-distributed replicated state machines (RSMs) are at the heart of many production distributed systems, offering linearizability and fault tolerance via consensus protocols. Most existing protocols target crash fault tolerance, however, and are vulnerable to fail-slow faults, where a single slow replica can significantly degrade system latency. Existing protocols that tolerate fail-slow faults do so with much higher normal-case latency in geo-distributed settings. This paper presents Avicenna, the first consensus protocol for geo-distributed RSMs that maintains low normal-case latency while tolerating a single fail-slow replica. Avicenna uses a single leader to order commands, naturally tolerating a fail-slow follower. To tolerate a fail-slow leader, Avicenna compares the current latency with the counterfactual latency clients would experience if a different replica, the shadow leader, were the leader. When that comparison indicates the current leader might be slow, Avicenna quickly promotes the shadow leader with a fast leader rotation protocol. Our evaluation shows Avicenna has the same normal-case latency as Multi-Paxos while tolerating fail-slow faults.
KW - distributed system
KW - fault tolerance
KW - replicated state machines
UR - https://www.scopus.com/pages/publications/105038439881
UR - https://www.scopus.com/pages/publications/105038439881#tab=citedBy
U2 - 10.1145/3767295.3803615
DO - 10.1145/3767295.3803615
M3 - Conference contribution
AN - SCOPUS:105038439881
T3 - EUROSYS 2026 - Proceedings of the 2026 European Conference on Computer Systems
SP - 1581
EP - 1603
BT - EUROSYS 2026 - Proceedings of the 2026 European Conference on Computer Systems
PB - Association for Computing Machinery, Inc
Y2 - 27 April 2026 through 30 April 2026
ER -