The Config Was Byte-Identical. That Proved the Wrong Thing.

opentelemetry kafka terraform devops troubleshooting

We were codifying an OpenTelemetry Collector that had been installed by hand and never captured in any repo. The safe way to do that is not to invent a new config, it’s to reproduce the one that is already running. So we re-derived it, and then we checked our work the strictest way we knew: we compared the result byte-for-byte against the live production config. It matched, exactly. YAML parsed. terraform validate was clean, lint was clean, and several independent reviews came back with nothing.

Then it started up and crash-looped.

Everything that was supposed to catch a bad config had run and passed, and the config was bad anyway. The gap between those two facts is the whole point of this post, and it is not specific to collectors or Kafka. It is about what “byte-identical” actually proves, which turns out to be a different thing from what everyone in the room assumed it proved.

What “byte-identical” verified

The comparison was real, and it was useful. It proved that the transformation from the running config to the codified one was faithful. Nothing was dropped, nothing was reworded, no key drifted in translation. If your worry is “did we accurately reproduce the source,” a byte-for-byte diff is a complete and honest answer.

But that is a statement about the copy. It relates the new config to the old one and says they are the same. It says nothing at all about a second, entirely separate property: whether that config is valid for the thing it now has to run against. Those are two different questions. Fidelity of the copy, and validity for the target. They happen to produce the same shade of green, so it is easy to check the first and believe you have answered the second.

Here they had diverged, badly.

The gap the copy could not see

The config had been written for one version of the collector. The image the unit pinned was roughly two dozen minor releases newer. A config-driven binary and its config are a contract, and that contract had moved underneath a config that stayed frozen. Three things had broken across those releases:

  • A telemetry field the config still referenced had been removed. To the newer binary it is simply an unknown key, and an unknown key at startup is fatal, not ignored.
  • The Kafka receiver’s topic had gone from a single string to a list, while the Kafka exporter kept the singular form. That asymmetry is the nastiest kind: a naive find-and-replace across the file “fixes” both receiver call sites and breaks all eight exporter ones, or the reverse, and looks tidy either way.
  • A producer message-size setting now exceeded a newer broker-side write ceiling, so even a clean start would have failed the first time it tried to send.

None of this was drift in our copy. Every one of these was faithfully, correctly reproduced from a source that was itself no longer valid for the target. The diff was doing exactly its job and telling us nothing we needed to hear.

Why every green check was green

This is the part worth internalizing, because the checks were not negligent. They were answering their own questions correctly.

terraform validate checks that your configuration and plan are well-formed. It does not instantiate the component the config describes, it does not download the pinned image, and it has no idea what fields that image’s binary will accept. YAML validation checks structure: is this a map, is that a list, are the quotes balanced. It cannot know that topic needs to be a list in this version when it was a string in the last one, because “which version” is not in the YAML. The reviews checked that the codified config matched production, which it did, flawlessly.

And the deepest reason sits one layer down. This config is read by the collector binary, at startup, not by an admission webhook and not by any control-plane check. So a config that is invalid for that binary still applies cleanly. The resource is created, the object is accepted, everything reports success, and then the process reads its own config a second later and dies. The failure lives past every gate that had a chance to stop it, because none of those gates ever start the actual pinned component.

Byte-identical, valid YAML, clean plan, approved review. Four true statements, none of which is “this will start.”

The only check that means anything

The fix was not a better linter. It was to run the real thing: obtain the actual pinned binary, the correct distribution (a different build of the same collector ships a different set of components, so the wrong one gives you a confident false pass), point it at the real config with the real environment, and let it start. That is the first check in the whole chain that instantiates the components the config names and enforces the contract the pinned version actually defines. It catches the removed field, the receiver-list change, and the config-language errors that validate structurally cannot, because validate never builds a component.

So “start the pinned binary with this config” became a required pre-apply step, and it applies to more than first-time codification. Any time you bump the image, you have moved the contract and frozen the config. Any time you copy a spec file from one place to another, you have carried a config across a version boundary you may not have noticed. Both are the same setup as the original bug.

There is a smaller version of this that shows up as a warning rather than a crash. When the newer binary logs that some field or component name is deprecated in favour of its current form, that is not noise to scroll past. It is the identical failure, one release early, announcing itself while it is still free to fix. Rename it while the thing is idle and a restart costs nothing; leave it, and the release that finally removes the alias turns a config nobody edited into a config that no longer starts.

What to actually take from this

Pin a version and copy a config, and you have quietly created two things that must agree and were verified separately: a copy that is faithful to its source, and a source that may no longer be valid for its target. A byte-for-byte diff, a passing validate, a green lint, and a clean review can all confirm the first while every one of them stays blind to the second.

Verify against the thing you pinned, not the thing you copied from. The most convincing green light in the pipeline was an honest answer to a question nobody in the room was actually asking.

$ cat KAFKA .md
· 8 min read

I Inferred Three Things About a Live System. Two of Them Weren't True.

Three times in one week I made a claim about a live migration by reading a config file, a name prefix, or a template, and twice the claim was wrong. The artefact and the live system are two views of the same thing, and they can agree for reasons the artefact can't tell you. Only one of the views is the truth.

kafka opentelemetry terraform troubleshooting devops
$ cat TERRAFORM .md
· 9 min read

The Plan Was Green Because Two of Three Views Agreed. The Third One Was the Truth.

A terraform plan came back clean for a managed observability stack after an incident fix had been applied by hand. The fix was live, the fix was in the file, and state alone was behind. A plan is a diff between two views, but the system has three. Reading all three before you plan is what turns the green diff from a guess into a decision.

terraform helm eks state troubleshooting devops
$ cat CDKTF .md
· 6 min read

The Bug I Blamed on the Provider

For weeks a terraform plan showed drift I couldn't fix on an S3 bucket that was already configured correctly. I filed it under 'AWS provider quirk' and worked around it. The provider was innocent. My own code was lying to me before Terraform ever saw it.

cdktf terraform iac typescript devops troubleshooting