Distributed Systems
Many machines cooperating over an unreliable network.
Backend Engineer track
Junior
Write correct code, ship small changes safely, ask good questions.
Nothing here yet.
Mid-level
Own a feature end to end without hand-holding.
Core: start here
- CAP TheoremDuring a network partition, you must choose consistency or availability.
- Eventual ConsistencyReplicas converge once updates stop, but reads may be stale in the meantime.
2 more mid-level concepts
- BASEBasically Available, Soft state, Eventual consistency: the counterpart to ACID.
- IdempotenceDoing something twice has the same effect as doing it once.
Senior
Own a system, its failure modes, and its trade-offs.
Core: start here
- Consistency ModelsLinearizable, sequential, causal, eventual: what readers are allowed to see.
- Distributed SystemMany machines cooperating over an unreliable network.
- Fallacies of Distributed ComputingEight false assumptions, starting with "the network is reliable".
- Saga PatternA chain of local transactions with compensating actions instead of a distributed transaction.
- Transactional OutboxSaving events in the same database transaction, then publishing them reliably.
22 more senior concepts
- Change Data CaptureStreaming database changes out as events.
- Choreography vs OrchestrationServices reacting to events vs a central coordinator directing them.
- Clock SkewMachines' clocks disagreeing, which breaks time-based ordering.
- ConsensusGetting nodes to agree on a value despite failures.
- Data Consistency Across ServicesKeeping data correct when each service owns its own database.
- Database per ServiceEach microservice owning its data exclusively.
- Distributed LockMutual exclusion across machines, and its pitfalls.
- Gossip ProtocolNodes spreading information by talking to random peers.
- HeartbeatPeriodic "I'm alive" signals used to detect failed nodes.
- Idempotency in Distributed SystemsMaking retries safe when you can't know whether the first attempt worked.
- Leader ElectionChoosing one node to coordinate the others.
- LeaseA lock or role that expires automatically unless renewed.
- PACELCExtends CAP: even without partitions, you trade latency against consistency.
- QuorumRequiring a majority of nodes to agree on reads or writes.
- RaftAn understandable consensus algorithm, used in etcd.
- Request CoalescingCollapsing identical concurrent requests into one.
- Service DiscoveryHow services find each other's addresses.
- Service MeshInfrastructure handling service-to-service traffic, retries and mTLS.
- SidecarA helper process deployed alongside each service.
- Split BrainTwo nodes both believing they're the leader.
- Strong ConsistencyEvery read sees the latest write.
- Two Generals' ProblemWhy agreement over an unreliable link can't be guaranteed.
Staff
Shape how many teams build, across systems.
Nothing here yet.
Principal
Set technical direction for the organization.
- Byzantine FaultNodes that behave arbitrarily or maliciously.
- Fencing TokenA counter that stops stale lock holders from writing.
- Hedged RequestsSending a duplicate request when the first is slow, to cut tail latency.
- LinearizabilityOperations appear to happen instantly, in real-time order.
- Logical ClocksLamport and vector clocks for ordering events without real time.
- PaxosThe classic, notoriously hard-to-understand consensus algorithm.
- Total Order BroadcastDelivering the same messages in the same order to every node.
Data Engineer track
Junior
Build and fix pipelines from clear specs; write correct SQL.
Nothing here yet.
Mid-level
Own pipelines and models end to end, including their quality.
Core: start here
- Change Data CaptureStreaming database changes out as events.
4 more mid-level concepts
- BASEBasically Available, Soft state, Eventual consistency: the counterpart to ACID.
- CAP TheoremDuring a network partition, you must choose consistency or availability.
- Eventual ConsistencyReplicas converge once updates stop, but reads may be stale in the meantime.
- IdempotenceDoing something twice has the same effect as doing it once.
Senior
Design the platform's storage, processing and modeling choices.
- Choreography vs OrchestrationServices reacting to events vs a central coordinator directing them.
- Clock SkewMachines' clocks disagreeing, which breaks time-based ordering.
- ConsensusGetting nodes to agree on a value despite failures.
- Consistency ModelsLinearizable, sequential, causal, eventual: what readers are allowed to see.
- Data Consistency Across ServicesKeeping data correct when each service owns its own database.
- Database per ServiceEach microservice owning its data exclusively.
- Distributed LockMutual exclusion across machines, and its pitfalls.
- Distributed SystemMany machines cooperating over an unreliable network.
- Fallacies of Distributed ComputingEight false assumptions, starting with "the network is reliable".
- Gossip ProtocolNodes spreading information by talking to random peers.
- HeartbeatPeriodic "I'm alive" signals used to detect failed nodes.
- Idempotency in Distributed SystemsMaking retries safe when you can't know whether the first attempt worked.
- Leader ElectionChoosing one node to coordinate the others.
- LeaseA lock or role that expires automatically unless renewed.
- PACELCExtends CAP: even without partitions, you trade latency against consistency.
- QuorumRequiring a majority of nodes to agree on reads or writes.
- RaftAn understandable consensus algorithm, used in etcd.
- Request CoalescingCollapsing identical concurrent requests into one.
- Saga PatternA chain of local transactions with compensating actions instead of a distributed transaction.
- Service DiscoveryHow services find each other's addresses.
- Service MeshInfrastructure handling service-to-service traffic, retries and mTLS.
- SidecarA helper process deployed alongside each service.
- Split BrainTwo nodes both believing they're the leader.
- Strong ConsistencyEvery read sees the latest write.
- Transactional OutboxSaving events in the same database transaction, then publishing them reliably.
- Two Generals' ProblemWhy agreement over an unreliable link can't be guaranteed.
Staff
Shape how the whole organization produces and uses data.
Nothing here yet.
Principal
Set data strategy and architecture across the company.
- Byzantine FaultNodes that behave arbitrarily or maliciously.
- Fencing TokenA counter that stops stale lock holders from writing.
- Hedged RequestsSending a duplicate request when the first is slow, to cut tail latency.
- LinearizabilityOperations appear to happen instantly, in real-time order.
- Logical ClocksLamport and vector clocks for ordering events without real time.
- PaxosThe classic, notoriously hard-to-understand consensus algorithm.
- Total Order BroadcastDelivering the same messages in the same order to every node.
Frontend Engineer track
Junior
Build UI that works, ship small changes safely, ask good questions.
Nothing here yet.
Mid-level
Own a feature end to end without hand-holding.
- CAP TheoremDuring a network partition, you must choose consistency or availability.
- Eventual ConsistencyReplicas converge once updates stop, but reads may be stale in the meantime.
- IdempotenceDoing something twice has the same effect as doing it once.
Senior
Own an app's architecture, performance, and failure modes.
- Consistency ModelsLinearizable, sequential, causal, eventual: what readers are allowed to see.
- Distributed SystemMany machines cooperating over an unreliable network.
- Fallacies of Distributed ComputingEight false assumptions, starting with "the network is reliable".
- Idempotency in Distributed SystemsMaking retries safe when you can't know whether the first attempt worked.
- Request CoalescingCollapsing identical concurrent requests into one.
- Strong ConsistencyEvery read sees the latest write.
Staff
Shape how many teams build, across apps.
Nothing here yet.
Principal
Set technical direction for the organization.
Nothing here yet.