Architecture & System Design › Distributed Systems
Byzantine Fault
Nodes that behave arbitrarily or maliciously.
Also known as: byzantine fault, byzantine failure, arbitrary fault
A Byzantine fault is arbitrary misbehaviour: not just crashes, but lies — corrupted messages, malicious nodes, equivocation (telling different stories to different peers). Named for generals with traitors among them, it’s the fault model for adversarial settings: blockchains, cross-organisation consensus, and any system where participants may be compromised rather than merely dead.
crash fault: node stops (detectable silence)
byzantine fault: node lies differently to each peer (needs 3f+1 to tolerate f)
Tolerance costs more than crash tolerance (3f+1 nodes per f traitors, extra protocol rounds), so use it where adversaries are realistic — permissionless or multi-party systems — and stick with crash-fault protocols (Raft/Paxos) inside single trust domains.
The classic mistakes:
- Byzantine protocols for crash problems. Paying 3f+1 and protocol complexity inside one company’s datacenter buys nothing over Raft. Match the fault model to the adversary.
- Crash assumptions in adversarial settings. Majority votes among potentially-lying nodes decide nothing; equivocation defeats naive quorums. Use BFT protocols (or trust boundaries) where lies are possible.
- Ignoring the network adversary. Byzantine models often assume bounded message delay manipulation; real attackers partition, delay and replay. Model the network too.
- Key management gaps. BFT identity rests on keys; compromised key infrastructure voids protocol guarantees. Harden PKI with the same seriousness as the protocol.
- Performance optimism. BFT throughput/latency trails crash protocols significantly; capacity-plan for the protocol’s reality, not its marketing.
- Assuming blockchains equal BFT understanding. Using a chain doesn’t confer its guarantees to surrounding systems (oracles, bridges, clients) — each link needs its own fault model.
- Forgetting recovery. Detected traitors need exclusion and reconfiguration paths; a protocol that tolerates but never expels degrades permanently.
When to assume it: multi-party, permissionless or compromise-plausible settings. Everywhere else, crash-fault tolerance plus good security buys the same safety cheaper.