September 8, 2026 · 8 min read
The Error Was 25 Hours Old. The Client Was Healthy the Whole Time.
A production producer rejected every request for most of a day. Every health signal stayed green the entire time: restart count 0, TCP sessions to the Kafka brokers ESTABLISHED, liveness and readiness probes passing. The error it returned was byte-identical to one its underlying client had emitted 25 hours earlier, elapsed-time value included. The client had recovered within seconds of the incident. The wrapper never noticed.
The pattern that works
Like a lot of platforms, we run Kafka clients behind a shared wrapper library, and the wrapper’s answer to broker loss is deliberately crude: count errors, and once a threshold is crossed, panic. The orchestrator restarts the process, and the service comes back cold and clean.
It is easy to sneer at crash-to-recover until you have debugged a client in a half-broken state, wedged in some corner of its reconnect state machine that no test covers and no runbook names. The panic converts an unbounded space of those states into exactly one state: a fresh process. And it composes with machinery the platform already has. The wrapper does not reimplement recovery; it delegates to the restart, which is the one recovery path everyone has already invested in.
During the incident at the center of this post, the platform rolled its brokers, and the busy services using this wrapper behaved exactly as designed. A stream processor touches Kafka constantly, so a broker roll generates failures several times per second; the threshold crosses within seconds, the process panics, the orchestrator replaces it, and the service is clean within a minute. Over a dozen services on the same library version did precisely that, at the exact second their broker shut down. Crash-to-recover earned its keep that night.
Why idle breaks it
The threshold counts errors, and errors only happen when the service does work. That is the whole flaw.
A stream processor produces errors fast enough to trip any threshold you care to set. A request-driven producer, a service that only touches Kafka when something calls it, and which something calls roughly once a day, produces about one error per day. The counter never crosses anything. And the wrapper’s behavior for a threshold that never trips is not “keep trying”. It is to latch the last error and return it for every subsequent attempt, indefinitely. We watched one do it for 22 hours, and it only ended because a human restarted the pod.
Recovery proportional to traffic is not recovery. The services that need it least get it; the ones that need it most never do. The quiet service is exactly the service whose failure goes unnoticed, and the mechanism built to save it is rate-limited by the one thing it lacks: traffic.
Wrong twice
The route to the right answer ran through two wrong ones, both worth keeping.
The first was correlation wearing the costume of cause. The broker logs showed one user failing authentication, thousands of times a day, flat, for weeks. The audit trail pinned it: an automated cleanup job had deleted that user’s credential. Exactly one credential had been deleted, and exactly one user was failing. I restored the credential, watched the failures stop, and reported the outage root-caused.
It was not the outage. The failing service authenticated as an entirely different user, which one read of its deployment config would have shown before anyone touched anything. The restore was harmless, and it did fix a genuinely broken client, but it was a different client. “Only one thing failed and only one thing changed” is a property of your search, not of the system. Uniqueness is a prompt to verify, not proof.
One detail actively misled that first pass: the service whose logs everyone was reading was not the service producing to Kafka. It had received the producer’s error in the body of an HTTP 500 response and logged it as its own, so a Java stack trace was carrying a C library’s error string. The reporter and the actor were different services, and the log everyone was grepping answered a different question than the one being asked.
The second wrong answer was the version. The healthy service ran a newer client library than the stuck one, so a version bump looked like the fix, and a ticket got cut. Then came a fleet inventory: grep the embedded Go build info out of /proc/1/exe on every pod, and you know exactly which library version each service runs. Dozens of services on the wrapper, most on the same version as the stuck one, and over a dozen of those same-version services had panicked and self-healed during the roll. The version splits nothing.
Neither direction of the usual argument survives that. “Another service on this version was fine” exonerates nothing, and “the healthy one runs a newer version” convicts nothing. Split the fleet by traffic profile before reaching for a version diff. Getting this wrong sends someone off to do an upgrade that cannot fix the problem, and a colleague pushing back, with a crash window from the client’s own logs that my story could not accommodate, is what surfaced it.
The evidence in the digits
The decisive evidence was sitting in the error string the whole time.
The stuck producer’s error was byte-identical to one the underlying client had emitted during the broker roll, and I mean identical down to the elapsed-time value, the “after N milliseconds” fragment. The underlying client had logged nothing since, because it had reconnected normally within minutes. The wrapper had cached the error object and re-served it for a day.
Identical elapsed-time digits cannot be two independent events. An elapsed time is a measurement, and two measurements do not agree to the digit. Once that lands, the question stops being “why does it keep failing” and becomes “why is the error never cleared”, and that question has an answer. When an error message carries a duration or a counter, diff those digits across occurrences. Same digits means one cached error being re-served. It costs nothing, and it flips the investigation.
What a latched process looks like
Nothing on the standard dashboard distinguishes stuck from idle.
Restart count is 0, because the process never needed saving by its own logic. Sockets to the brokers are ESTABLISHED, because the transport genuinely recovered; socket state is evidence from a layer below the claim, and I had called the thing recovered on that evidence once already, which was premature. Liveness and readiness pass, because neither probe exercises a produce. The logs show nothing, because a quiet service logs nothing whether it is broken or merely quiet. And app-level retries get the same cached error a few seconds later, which is why “add retries” is the wrong reflex. You cannot retry your way out of a cache.
What closes it
Three fixes, in descending order of value.
Clear the error state on success. A successful produce, or a successful reconnect, should retire the cached error. The failure to do this is the entire bug; everything else is mitigation.
Make the health probe exercise the produce path. A producer’s readiness claim is “I can produce”, so that is what the probe should test. A producer that reports healthy while rejecting every produce is lying on the one axis anyone is polling.
Make the threshold time-based as well as count-based. A quiet service then eventually recovers the same way a busy one does, by the clock instead of by the counter.
And accept the recurring trigger, because it will not go away. A managed Kafka platform rolls its brokers on a maintenance schedule, monthly in our case, and every roll re-runs this experiment against every client. After any provider-side roll, the request-driven producers are the ones to check, and the only reliable check is pushing a request through them or restarting them. There is a small set of silent low-traffic producers that cannot be distinguished from stuck without pushing a request through; they go on the list for the next roll.
Where this generalises, and the trap
Every recovery mechanism with a rate-based trigger has this failure mode. Circuit breakers need failures to trip. Watchdogs count events. Connection pools validate on checkout, so a pool nobody checks out never validates anything. Health checks that only exercise the happy path pass forever next to a broken one. Low traffic turns each of these into a no-op, and it does so silently: nothing is disabled, nothing is misconfigured, the mechanism is simply waiting for evidence that is not coming.
Two rules leave this page. When a recovery mechanism is rate-based, ask what happens when traffic does not meet the rate. And when an error carries a counter, the counter is evidence.
The trap is to treat recovery as a property of the failure. It is a property of the traffic. A client that only recovers while failing recovers exactly when it fails often enough, and the service too quiet to fail often enough never gets the chance. It will sit there, healthy on every dashboard, serving yesterday’s error.