Architecture & System Design › Distributed Systems
Service Mesh
Infrastructure handling service-to-service traffic, retries and mTLS.
Also known as: service mesh, istio linkerd, sidecar mesh
A service mesh moves service-to-service concerns — discovery, load balancing, retries, mTLS, telemetry, policy — out of application code into a dedicated infrastructure layer: sidecar proxies beside every instance, controlled centrally. Applications call plain HTTP; the mesh makes it resilient, observable and secure.
app → localhost sidecar (retry? TLS? route? record) → network → sidecar → app
The payoff is uniform capability without per-language libraries: retries with budgets, circuit breaking, traffic splitting, mutual TLS and golden metrics everywhere, configured once. The price is operational surface (control plane, proxy fleet, certificate rotation) and per-hop latency/resource overhead.
The classic mistakes:
- Mesh for a monolith. Two services gain nothing from a control plane, sidecars and mTLS rotation. Meshes pay off past dozens of services with polyglot stacks.
- Retry amplification. Every sidecar retrying independently multiplies load geometrically under partial failure. Budgets, hedges and circuit breaking tuned mesh-wide, not per default.
- mTLS without identity story. Encrypted hops with unauthenticated identities add cost without security. Pair with workload identity (SPIFFE-style) and authorisation policy.
- Observability bills. Per-hop spans and metrics at high cardinality explode tracing/storage costs. Sample, aggregate, and budget telemetry like any production expense.
- Control-plane fragility. A downed control plane freezing data-plane updates (or worse, draining config) halts adaptation. Harden and cache; data planes must survive control outages.
- Debugging through layers. App→sidecar→network→sidecar→app multiplies “where did it break” surfaces. Mesh-aware tracing and consistent header propagation are prerequisites, not extras.
- Big-bang adoption. Flag-day mesh rollouts across fleets break in correlated ways. Migrate namespace by namespace with rollback paths.
When to adopt: polyglot microservice fleets needing uniform resilience, security and observability without per-stack libraries. Small or single-stack systems get further with good libraries and gateways.