I Inferred Three Things About a Live System. Two of Them Weren't True.

kafka opentelemetry terraform troubleshooting devops

Three times in one week, working on the same migration, I made a claim about a live system by reading a file. Two of the three claims were wrong. The third was wrong by accident and saved at low cost. The unifying mechanism is that the artefact and the live system are two views of the same thing, and they can agree for reasons the artefact can’t tell you. Only the live view is ground truth.

The seventh consumer

A canary migration had three signals green. The new collector had logged its first traffic. The span processors were writing to the new broker. The six services that turn those metrics into database rows were getting fed. The fix was shipped. Only afterwards did a broader sweep find a seventh service, a billing consumer, that had been disconnected the same way the whole week, on the same mechanism, and missed by the service-name scan that drove the cleanup.

The two truths held at once. Every signal in the room was genuine, and the system under change had a hole the cleanup had papered over. That gap is the post.

Why the scan missed it

The fix was scoped by a consumer-name prefix. Six services matched. The seventh didn’t carry that prefix, so the sweep didn’t see it. The mechanism is worth naming out loud: the topics are the contract, the names are a convention, and a convention can be wrong. Matching consumer subscriptions by service-name prefix treats the namespace as the contract. The namespace lied.

The cost hides in plain sight, because the match succeeds on the cases it sees, and the case it doesn’t see is exactly the one that won’t show up in the report. A green sweep is not the same as a complete sweep. It is the proof that the cases your naming convention covers were covered. It is silent about the cases the convention skipped, which is a smaller set that has no obligation to be empty.

A retraction that wasn’t a swap

Same week, different artefact, same failure mode. The collector’s Kafka consumer groups had been read off configuration templates to conclude they carried no stage prefix and would therefore be denied once broker-side ACLs were enforced. Enumerated on the live cluster, every one carried a prefix and was covered.

The general principle still held. One billing-side consumer is genuinely unprefixed and genuinely would have broken once the rules tightened. The error wasn’t speculative; it was sourced from a specific claim about these N groups that the live data did not back. Templates look like truth because they’re the artefact the platform hands you, but they’re somebody’s rendering of intent, not a contract with the broker. The reason the retraction caught itself was that the broader picture had already broken somewhere else in the same migration, and reading the live cluster was already the reflex.

The interesting shape of the mistake is that it would have been cheap to make and easy to ship: a confident paragraph in a runbook, a checkmark on a review, a fixed-naming-convention change applied to a fleet. The cost of being wrong was zero until the moment a different consumer query landed on the unprefixed group and the rule that hadn’t actually been enforced closed it silent. The cost of being right was a paragraph no one would have read twice.

The cheap third

Argued against inlining a Kafka SASL username in a values file because SSM stores secrets encrypted, a sensible-looking objection, retracted on a one-line check: the codebase already carried exactly such a literal in another values file. The convention existed; a policy had been inferred from a default. This one was cheap to catch and almost cheaper to miss, because it was the third time in three days pattern-matching had been the move, and the third miss of a week is the one a tired operator stops flagging to themselves, because the brain is already triaging them.

The shape of every one of these is the same: a confident claim about the live system, sourced from the artefact, that the artefact cannot actually answer. The same shape wears three different clothes over a week.

Why “I reasoned from the file” doesn’t feel like a finished sentence

The artefact and the live system are two views of the same thing, and two views can agree for very different reasons. They can agree because they are derived from each other, which is the canonical case and the only one where “I read the file” is a complete answer. They can agree because one of them was edited and the other wasn’t, which is the three-way-merge shape and which a terraform plan flattens into a single diff that looks like normal drift. Or they can agree because one of them was edited in a way the other couldn’t see, which is the byte-identical cousin of this post: a config faithful to a source that has since been read by a different binary.

Reading the artefact gives you what it was supposed to be. Only the live system gives you what it is. Both feel like ground truth. One of them is right far less often than either feels.

What the checks actually look like

The principle in the last section has to land in something you can type. Three small reads, one per layer of the Kafka system, none of which depend on a template, a prefix, or a config file. Each is paired below with what it is buying you, which matters more than the flag set.

Enumerate consumer groups against the cluster. The shape that finds the seventh consumer:

kafka-consumer-groups.sh \
  --bootstrap-server "$BROKER" \
  --describe --all-groups

The --all-groups flag is the part that catches the missed consumer, because it does not pattern-match on a name prefix; it lists every group the cluster knows about, including any whose names do not fit the convention. Pair this with --all-topics if the broker is shared across stages, so the read spans every partition of every topic rather than the ones your consumer config happens to declare.

This is the read that replaces the service-name prefix sweep. Both passes look like inventory. Only one of them takes the cluster as ground truth.

Pull the topic list as the contract. When the artefact said the consumers were misaligned, the first thing to check is what topics actually exist on the cluster and what their partition counts are:

kafka-topics.sh \
  --bootstrap-server "$BROKER" \
  --list --exclude-internal

Followed by --describe for the topic or topics of interest when the count matters:

kafka-topics.sh \
  --bootstrap-server "$BROKER" \
  --describe --topic "spans.${STAGE}"

These read the contract from the broker, where it lives. Config templates tell you what should exist, which is what gets you into the byte-identical cousin of this post; the topic-list commands tell you what does. The two will frequently diverge, and when they do, the broker is right.

Read the ACLs from the live cluster. When the second miss said “the consumer groups will be denied once ACLs tighten,” the read that would have settled it in one command:

kafka-acls.sh \
  --bootstrap-server "$BROKER" \
  --list --topic "spans.${STAGE}"

This is the read the template-based inference was trying to substitute for, and it does not need a stage prefix or a config block to answer the question. Whatever the cluster says this principal is allowed to do, that is what the rule actually does. Pair it with --resource-pattern-type prefixed if your environment uses prefix-style ACLs rather than literal resource names, since --list defaults to literal patterns and silently hides the prefixed grants.

Three reads. None of them asks for permission beyond what an operator with read-only cluster access already has. None of them takes more than a few seconds. None of them is hard to remember. The hard part is remembering to reach for them before reaching for the file.

The rule and the check

Infer from the artefact, verify against the live system. Treat the first as a hypothesis, not as evidence.

The seventh consumer is what the rule buys you, and the check is the one above that caught it: enumerate consumer groups across the cluster, all topics, all groups, no permission ask, five seconds of wall-clock. The fragment reads almost like the prefix-sweep it replaces; it just does not take the convention’s word for which subscriptions exist. The prefix was a guess about a contract; the topic was the contract.

The check is what survives into the next week. The example is what makes it memorable. Naming both together is the part that earns the next three “wait, is this one of those” moments when the artefact and the live system first stop agreeing, before the brain has had time to fall back on the reflex.

$ cat OPENTELEMETRY .md
· 6 min read

The Config Was Byte-Identical. That Proved the Wrong Thing.

We re-derived a collector config, confirmed it was byte-for-byte identical to the one running in production, and shipped it with a stack of green checks behind it. It crash-looped on startup. Byte-identical is a real property; it just answers a question nobody was asking.

opentelemetry kafka terraform devops troubleshooting
$ cat TERRAFORM .md
· 9 min read

The Plan Was Green Because Two of Three Views Agreed. The Third One Was the Truth.

A terraform plan came back clean for a managed observability stack after an incident fix had been applied by hand. The fix was live, the fix was in the file, and state alone was behind. A plan is a diff between two views, but the system has three. Reading all three before you plan is what turns the green diff from a guess into a decision.

terraform helm eks state troubleshooting devops
$ cat CDKTF .md
· 6 min read

The Bug I Blamed on the Provider

For weeks a terraform plan showed drift I couldn't fix on an S3 bucket that was already configured correctly. I filed it under 'AWS provider quirk' and worked around it. The provider was innocent. My own code was lying to me before Terraform ever saw it.

cdktf terraform iac typescript devops troubleshooting