Contents

Architecture & System Design › System Design Fundamentals

Coordination Service

Systems like ZooKeeper and etcd for configuration, locks and leader election.

Also known as: coordination service, distributed coordination, zookeeper etcd

A coordination service gives distributed processes shared truth: leader election, locks, configuration, group membership, barriers — small, strongly-consistent state every node can trust. ZooKeeper, etcd and Consul play the role: a handful of nodes running consensus, exposing simple primitives (key-value with watches, sequential nodes, leases) that whole fleets build on.

fleet → coordination service: "who is leader?" "hold this lock" "watch /config"

The pattern centralises the hard part (consensus) so services stay simple: instead of every pair negotiating, all consult the same strongly-consistent source. The service itself must be rock-solid and fast — it’s a dependency of everything, so its failure modes (partitions, slowness) become everyone’s.

The classic mistakes:

  • Storing data in it. Coordination services hold kilobytes of metadata well and megabytes badly. Configuration, membership, locks — yes; application data, queues, caches — never.
  • Chatty watches. Thousands of clients watching volatile keys storms the service on every flap. Watch stable keys; debounce reactions; cache locally.
  • No session handling. Locks and leadership tied to sessions evaporate on expiry — holders must verify ownership before acting (see fencing tokens).
  • Single coordination cluster for everything. One cluster’s outage halts all coordination. Isolate critical coordination paths or run dedicated clusters.
  • Assuming speed. Strong consistency means cross-quorum latency on writes; coordination reads can be local, writes cannot. Keep write rates low.
  • Reimplementing consensus. Hand-rolled leader election over a database races and splits. Use the coordination service — that’s what it’s for.
  • Forgetting the operator. The coordination cluster needs its own backup, monitoring and quorum planning. The foundation needs foundations.

When to use it: leader election, distributed locks, service discovery, config distribution — any shared-truth problem. Centralise consensus once, correctly, and let services coordinate through it rather than among themselves.