Distributed tracing is one of the most useful tools in an SRE's kit, and one of the easiest to overspend on. Once a handful of services turn into a few dozen, trace volume grows faster than traffic. Storage bills climb, queries slow down, and teams quietly start turning tracing off.
Most traces in a distributed tracing setup are boring. A healthy GET /health or a cached product lookup that returns in 8 ms tells you very little, and you may be generating millions of them per hour. The traces that matter are the slow ones, the failed ones, and the ones touching a recently deployed service. The goal of sampling isn't to keep less data, but to keep an honest picture of normal behavior while retaining nearly all of the interesting outliers.
No distributed tracing strategy is perfect, but the worst choice is no sampling at all, followed closely by a policy nobody revisits. Start simple, measure what you're losing, and tighten the rules as you learn which traces your team actually opens during incidents.
Why 100% Retention Rarely Makes Sense
Most traces are boring. A healthy GET /health or a cached product lookup that returns in 8 ms tells you very little, and you may be generating millions of them per hour. The traces that matter are the slow ones, the failed ones, and the ones touching a recently deployed service.
The goal of sampling isn't "keep less data". It's to keep a statistically honest picture of normal behavior while retaining nearly all of the interesting outliers.
Head-Based Sampling
Head-based sampling decides at the start of a trace, usually in the SDK or at the entry-point service, whether to record it. The decision is then propagated to downstream services through trace context.
Strengths
- Cheap and simple. Unsampled requests never generate span data.
- Predictable cost, since you control the percentage directly.
- Works well with OpenTelemetry's built-in samplers, such as
TraceIdRatioBasedandParentBased.
Weaknesses
- The decision is made before you know how the request will turn out. A 1% sample means you keep roughly 1% of your errors too.
- Rare, high-value failures are the most likely to go missing.
Head-based sampling is a reasonable default for high-volume, low-risk traffic, but it's a poor fit if you rely on traces for debugging rare issues.
Tail-Based Sampling
Tail-based sampling waits until a trace is complete, then decides whether to keep it. That lets you write rules based on what actually happened:
- Keep every trace containing an error
- Keep every trace slower than a latency threshold (say, p95)
- Keep a small random percentage of everything else
Strengths
- You retain the traces that matter most.
- Policies can be expressive: by service, endpoint, status code, or custom attributes.
Weaknesses
- It's more expensive to run. All spans for a trace must be buffered somewhere until the decision is made.
- All spans from one trace have to reach the same decision point. In the OpenTelemetry Collector, that means load balancing by trace ID before the tail-sampling processor.
- Late-arriving spans, from async jobs or message queues, can be missed if your decision window is too short.
A Practical Hybrid
Most mature setups end up combining both:
- Drop the obvious noise early. Filter health checks, readiness probes, and synthetic pings at the SDK or Collector level. This is a filtering decision, not sampling.
- Apply light head sampling at the edge to cap raw volume.
- Run tail sampling in a Collector tier with policies for errors, high latency, and a baseline random sample.
- Always keep certain traffic. Traces tied to a feature flag rollout, a canary deployment, or a specific customer under investigation shouldn't be left to chance.
This gives you cost control from the head stage and precision from the tail stage.
Common Pitfalls
Broken traces. If services make independent sampling decisions, you'll get traces with missing spans. Always honor the parent's sampling decision with ParentBased samplers.
Skewed metrics. If you compute request counts or error rates from sampled traces, the numbers will be wrong unless you correct for the sampling rate. Generate metrics from the full stream before sampling, for example with span-metrics in the Collector, and use traces for investigation only.
Set-and-forget policies. A latency threshold chosen last year may no longer reflect reality. Review your tail-sampling rules whenever traffic patterns or SLOs change.
Ignoring the sampler's own health. Collectors that buffer traces can run out of memory under load. Monitor buffer size, dropped spans, and decision latency like any other production system.
Cardinality surprises. Policies keyed on high-cardinality attributes, such as user IDs or full URLs, can blow up memory use in the sampling tier.
Storage and Query Considerations
Sampling is only half the story. How retained traces are stored affects how quickly you can query them:
- Index selectively. Indexing every attribute is expensive. Index the ones you filter on most (service, status, duration, a few business keys) and leave the rest searchable but unindexed.
- Tier your retention. Keep recent traces on fast storage and move older ones to cheaper storage.
- Partition by time. Most investigations focus on a narrow window, so time-based partitioning keeps scans small.
How to Decide
Ask these questions:
- Do I need to debug rare failures? → Tail sampling.
- Is my main constraint cost at very high volume? → Head sampling plus aggressive filtering.
- Do I need accurate rate and error metrics? → Derive them before sampling.
- Do I have the operational capacity to run a stateful Collector tier? If not, start with head sampling and graduate later.
Final Thoughts
No sampling strategy is perfect, but the worst choice is none at all, followed closely by one nobody revisits. Start simple, measure what you're losing, and tighten the policies as you learn which traces your team actually opens during incidents.
What sampling strategy does your team use, and what's the one trace you wish you'd kept? Share your experience in the comments.
Top comments (0)