Contents

Data Engineering › Collection & Instrumentation

Telemetry Pipeline

The path from emitted logs, metrics and events to where they're stored and queried.

Also known as: observability pipeline, telemetry collection, telemetry ingestion

A telemetry pipeline carries logs, metrics and traces from where they are produced to where they are stored and queried. The path usually runs: instrumented service or device → local agent or SDK → collector → buffer or stream → storage → query UI and alerting.

Telemetry is not business data. It is high volume, mostly low value one event at a time, and losing a small fraction is often acceptable. The classic mistake is making the application block on telemetry: writing every log line synchronously to a shared database, or sending spans inline with the request. When the telemetry backend slows down, your users feel it. Emit to a local buffer or a collector and let it handle retries and backpressure.

For backend engineers

  • Instrument once with a vendor-neutral standard such as OpenTelemetry, then route to whichever backend you use.
  • Keep emission non-blocking and bounded. A full buffer should drop or sample, not stall the request.
  • Watch label cardinality: a unique value per request (user id, URL with ids) can blow up metrics cost.
  • Propagate trace context across services so distributed tracing joins spans into one request.

For data engineers

  • The pipeline is a streaming system: collectors, queues and a store built for append-heavy writes.
  • Decide sampling early — head or tail — because you cannot store everything cheaply.
  • Separate hot queryable data from cold archive with a retention and tiering policy.
  • Expect schema drift in logs; parse defensively and keep the raw line.

The three signal types have different shapes; see logs, metrics and traces. The trade-off is cost versus fidelity: sampling and short retention save money but make rare bugs harder to reproduce, so keep raw data long enough to investigate incidents, then down-sample.