August 20, 2026 · 7 min read
The Old Pipeline Lost ~47k Spans. It Didn't. We Counted the Wrong Thing.
A side-by-side per-service span count during a parallel-run validation showed the new pipeline losing roughly 0.2% of spans per service across a 60-minute window. Read as a regression in the new pipeline, it would have triggered a rollback. Read correctly, it was the old pipeline ingesting tens of thousands of duplicate rows in a few-second window, with the new pipeline matching it span-for-span once the duplicates were excluded. The lesson is not that the new pipeline was fine. It is that the unit a storage layer exposes as “count” is not the unit the operator thinks it is, when the storage has no unique constraint on the thing being counted.
The first exhibit: the alarming per-service delta
The Traces Explorer’s count() panel reports row count. Per service, the new pipeline was short by roughly 0.2% across the window. The shape of the shortfall was uniform across services, which is the part that read as a real signal. A uniform per-service shortfall is the signature of a systemic issue rather than a single misbehaving component, and the natural reflex was “the new pipeline is dropping spans.” That reflex is what the post is about, because the reflex was wrong in direction. The new pipeline was not dropping. The old pipeline was duplicating, and a duplication in the old pipeline reads, from the new pipeline’s vantage point, exactly like a loss.
The instinct to read a uniform shortfall as a regression is not stupid; it is the right reflex in most failure modes. The reflex is what you should have when one of two pipelines is wrong. The problem with the reflex is that it assumes one of the two pipelines is the source of the discrepancy and treats the other as ground truth. The parallel-run setup made the old pipeline the reference by virtue of being there first, and a uniform shortfall against a reference looks like the test pipeline being off, not the reference being off. That asymmetry is what gets read as a regression.
The second exhibit: the drill-down that inverted the picture
Per-hour counts were computed and compared side by side, against both pipelines’ databases directly. Twenty straight hours matched span-for-span. One hour did not.
For that hour, the old pipeline’s count() read roughly 620k while its uniqExact(spanID) read roughly 580k. The new pipeline’s two counts both read roughly 580k. The shortfall was not in the new pipeline. The shortfall was the old pipeline having tens of thousands of duplicate rows inside the hour, and the duplication was concentrated in a small window inside it. Per service, the new pipeline’s count() equalled the old pipeline’s uniqExact(spanID) exactly, which is the relationship that inverts the diagnosis. The old pipeline was not larger than the new pipeline. The old pipeline was smaller by exactly the count of duplicates it had inserted twice.
The drill-down mattered because the hour-level discrepancy was the level at which the wrong reflex would have stuck. The window-level drill-down is what showed the duplication; the per-service match is what showed the direction. Both pieces were needed; either one alone would have left the diagnosis ambiguous. The shape of the disagreement (old count() over-reports; old uniqExact agrees with new) is the fingerprint of a duplicate-insert path, not the fingerprint of a dropped-span path.
The third exhibit: the log sweep that found the smoking gun
The old collector’s logs were swept for errors inside the window. A MEMORY_LIMIT_EXCEEDED on the error-table insert surfaced, occurring after the index-table insert for the same batch had already landed. The exporter retried the whole batch. The non-transactional multi-table write plus the retry is the duplication mechanism: the index table commits, the error table fails on its memory ceiling, the retry re-commits the index table a second time. The column-oriented store has no unique constraint on spanID and no dedup-on-insert, so the second commit lands as two identical rows. The error log shows hundreds of previous MEMORY_LIMIT_EXCEEDED occurrences on the old collector, which means the duplication is recurring, not anomalous.
The memory ceiling itself is the trigger, and it is workload-dependent. The error-table insert fails when the error batch is large enough to exceed the memory budget, which is more likely under traffic spikes, which is more likely to be the same window where someone is paying attention to per-service span counts. The duplication is not noise; it is correlated with the work pattern that surfaces the discrepancy, which is what makes the symptom so convincing as a regression in the new pipeline.
Why a count disagreed with a uniqueness-aware count
The Traces Explorer’s count() panel reports row count. The store’s spanID column is not a primary key; nothing rejects a duplicate insert. The two numbers, count() and uniqExact(spanID), agree when the pipeline is healthy and diverge when it is not. The divergence is the direction of the bug: count() over-reports because it counts duplicates. The new pipeline’s two counts agreed exactly because the new collector’s insert path does not retry in a way that produces duplicates. The disagreement between the two pipelines was not a regression in the new pipeline. It was the old pipeline’s duplication-driven over-counting being misread as the new pipeline’s under-counting.
This is the part of the post that generalises. A row count and a unique-keyed count are two different measurements of the same storage, and they are only equivalent when the storage rejects duplicates. The instant the storage accepts duplicates, the two measurements diverge, and the divergence is the signal. The direction of the divergence tells you which side is duplicating, not which side is dropping. A short panel against a reference can be caused by the short side dropping, the reference duplicating, or both; the disagreement between count() and uniqExact on the reference side is the diagnostic that tells you which.
The recurring-fault trap
The hundreds of previous MEMORY_LIMIT_EXCEEDED occurrences are what turns a one-off debugging session into a transferable rule. A single batch that retries after a mid-batch failure is a transient; the same batch path failing every time a workload peaks is a property of the old pipeline. The parallel-run will keep showing the “new pipeline losing spans” pattern every time the old pipeline’s error path triggers, and the divergence will accumulate over the validation window.
The reflex to fix is “compare with uniqExact.” That is the symptom treatment. The rule is broader: pick the comparison metric that does not depend on the old pipeline’s retry behaviour. A unique-keyed count is one such metric; a per-row-counts-of-error-events metric would also tell you the duplication happened; a per-row-counts-of-distinct-trace-ids metric would tell you the duplication did not affect the trace graph. Any metric whose definition is invariant under duplicate insertion is robust to this class of bug. The row count is not, and using it as the parallel-run comparison metric guarantees that any window overlapping a duplication event will read as a regression in the test pipeline.
The rule and the trap
When a count and a uniqueness-aware count disagree, the storage layer is telling you which one to use, and the disagreement is the signal. A non-transactional multi-table write plus a retry is a duplication path whether or not it fires; it fires when one of the tables hits its ceiling mid-batch, and once it has fired, the duplication lands permanently in the storage because nothing rejects it on insert.
The next time a side-by-side panel shows one pipeline short by a small, uniform percentage, the question is not “which pipeline is dropping.” The question is “which pipeline is double-counting,” and the answer is the one whose row count exceeds its unique-keyed count for the same window. The two pipelines are not symmetric in a parallel-run: the reference side, the one that was there first, has had more time to accumulate whatever failures its insert path is susceptible to. The short side’s deficit, in the common case, is the reference side’s surplus expressed from the wrong direction.