2026-09-01
Paxos works, but explaining it will hurt. Raft was designed in 2013 with an explicit goal that Paxos never had: understandability. Diego Ongaro's paper measured this — students who studied Raft scored 25% higher on comprehension quizzes than those who studied Paxos. That's why etcd, Consul, CockroachDB, TiKV, and Kubernetes all run on Raft.
Raft decomposes consensus into three independent subproblems:
Every node is in one of three states: follower, candidate, or leader. Followers wait for heartbeats. If a follower doesn't hear from the leader within its election timeout (randomized, typically 150–300ms), it becomes a candidate, increments the term number, votes for itself, and asks everyone else for votes. Win a majority? You're leader. Randomizing the timeout prevents split votes — the trick that makes Raft simple where Paxos is subtle.
Real-world example: etcd runs a 3-node Raft cluster behind every Kubernetes control plane. When you kubectl apply a deployment, the API server writes to etcd's leader. The leader appends the entry to its log and sends AppendEntries RPCs to the two followers. Once one follower acknowledges (2 out of 3 = majority), the write is committed and returned to you. If the leader crashes, one of the followers times out within ~200ms, wins an election, and continues serving. Your kubectl command might retry once — you probably won't notice.
Rule of thumb — cluster sizing: A Raft cluster of 2f+1 nodes tolerates f failures. So 3 nodes tolerate 1 failure, 5 tolerate 2, 7 tolerate 3. Always use odd numbers. A 4-node cluster still only tolerates 1 failure (majority is 3) but has more nodes that can fail — strictly worse than 3.
The gotchas:
