← All writing

Written October 2026, looking back at Systems4 min read

Keep the failing traces

Adding OpenTelemetry to Watson serving at IBM, and why I biased sampling toward errors and slow requests instead of sampling uniformly.

From IBM · May 2024 – Aug 2025

Part of my work at IBM was adding OpenTelemetry tracing to the Watson serving path. The instrumentation itself was the straightforward part. The interesting question turned out to be which traces to keep.

Why you can't keep everything

A trace records one request's journey through the system: every service it touched, how long each step took, where it failed. They're incredibly useful when something goes wrong. They're also expensive to export, store, and index, so at any meaningful traffic level you can't keep all of them. You sample.

The default approach is head sampling. When a request starts, you roll a die, and if it comes up, you record the whole trace. It's simple, cheap, and consistent across services, because the decision is made once at the start and propagated.

The problem is that the decision happens before you know anything about the request.

Head sampling throws away the interesting ones

Think about what you actually open a trace viewer for. Almost never a healthy request that came back in the usual time. You open it when something failed, or when a request took ten times longer than it should have.

Those requests are rare. That's what makes them interesting, and it's also why uniform sampling at a low rate mostly misses them. You end up with a large pile of healthy traces that all look the same, and when someone asks why a particular call failed, the honest answer is that the trace was probably sampled out.

It's a uniquely frustrating feeling: you go looking for the trace of a slow request and find plenty of fast traces from the same minute, none of the one you need.

Head sampling decides too earlyA head sampler rolls the dice before the request has done anything, so it keeps healthy traces and failures at the same low rate.

Error-biased sampling

So the policy I pushed for is simple to say:

  • keep every trace that ends in an error
  • keep every trace that's slow, above a latency threshold
  • sample the healthy, fast ones at a low rate, enough to know what normal looks like

The catch is that you can only know whether a request failed or was slow at the end. That means the decision has to move from the head to the tail. Spans get buffered until the trace is complete, or close enough, and then a policy decides whether the whole thing is kept.

Deciding at the tailSpans are buffered per trace in the collector, and the keep-or-drop decision happens once the outcome is known.

In OpenTelemetry terms, that's tail-based sampling in the collector rather than a sampler in the SDK. The services still emit spans; the collector holds them briefly, groups them by trace, and applies the rules.

Things I'd watch out for

Tail sampling isn't free, and a few things surprised me.

Memory and timing. The collector has to hold spans until it decides. If a trace is very long, or some spans arrive late, you have to choose how long to wait. Too short and you decide on partial traces; too long and memory grows. There's no perfect setting, only one you've tested against your actual traffic.

All spans for a trace need to land in one place. If you run several collectors, spans from the same trace must be routed to the same one, or each sees a fragment and makes its own decision. This is the kind of thing that works fine in a single-instance test and breaks at scale.

Cost during incidents. Keeping every error trace sounds great until a backend goes down and suddenly everything is an error. Your trace volume spikes at the exact moment the system is under stress. A cap or a rate limit on the "keep errors" rule is worth having, even if it feels like it contradicts the whole idea.

Cardinality. It's tempting to attach every useful-looking attribute to spans. Some of those, like raw IDs or free-form strings, explode the number of distinct values your backend has to index. I learned to be deliberate about which attributes are for searching and which are just context.

What I took from it

The broader lesson is that sampling is a product decision, not just a cost knob. The question isn't "what fraction can we afford" but "which requests will someone need to look at later". Once you frame it that way, biasing toward failures is obvious.

The test I'd apply to any sampling policy: list the last few incidents and ask whether the traces you needed would have survived it. That's a much better test than any dashboard. It's also the same instinct as the Go client: design for the bad day, because the good days take care of themselves.