Infrastructure & Operations › Observability
Sampling
Recording only a fraction of traces or logs to control cost.
Also known as: sampling, trace sampling, head sampling, tail sampling
Sampling records only a fraction of your telemetry instead of everything. At high volume, keeping every trace and log line is expensive and mostly redundant; sampling keeps a representative subset so you can still see how the system behaves and, ideally, still find the rare problems. It’s a cost-control measure, and done badly it’s a visibility-control measure too.
Two broad approaches:
- Head sampling decides at the start: e.g. keep 10% of requests, chosen as they arrive. It’s simple and cheap, and the decision can be made consistently across a whole trace so traces stay complete. The trade-off: you may drop a rare slow or failing request before you know it matters.
- Tail sampling decides at the end, after the trace is complete: keep all slow traces, all error traces, and a sample of the rest. It’s more useful — you keep the interesting ones — but it requires buffering traces until they finish, which costs memory and complexity.
head: keep 10% up front → simple, may miss rare failures
tail: keep errors + slow + 1% → smarter, needs buffering
The classic mistakes:
- Sampling away the interesting events. A flat 1% sample drops most of the very errors you’d investigate. Prefer strategies that keep errors and outliers (tail sampling, or separate retention).
- Inconsistent sampling. If one service samples 10% and another 1%, the same request gets recorded in one and not the other, and its trace breaks across services. Propagate the sampling decision so a trace is kept or dropped as a whole.
- No sampling at all at scale. Keeping everything can be prohibitively expensive and slow; the bill, not the dashboard, becomes the problem.
- Sampling logs and metrics the same as traces. They have different volume and value; a counter is cheap to keep in full, while raw logs may need aggressive sampling or filtering.
When to sample: whenever the volume of traces or logs exceeds what you need or can afford. Start with a modest head sample for the routine traffic, and add rules to always keep errors and slow requests. Keep an eye on cardinality too — sampling interacts with how much unique data you store. Configure it centrally, often in an OpenTelemetry collector.