The Constraint Was Satisfied. The Zone Was Empty Anyway.

kubernetes scheduling opentelemetry finops mechanism

I was restarting a collector pod by pod to converge some zone-aware routing, one at a time so nothing lost capacity. Three safe deletions later, two pods of the same Deployment sat on a single node and its own zone was empty.

The namespace has a hard spread constraint on every one of these Deployments. Its entire job is to prevent the state I had just produced. And it was satisfied the whole time.

What the constraint actually promises

A topologySpreadConstraint is evaluated once, when a pod is scheduled, and never again. Kubernetes does not rebalance. When a pod comes up, the scheduler counts, per value of the topology key, the pods that match the constraint’s labelSelector, computes the skew against maxSkew, and places accordingly. After that moment nothing re-evaluates; a pod stays where it landed until something deletes it.

That makes “satisfied” a claim about two things at once: a count, and the population it was counted over. The count was right in my case. The population was not.

It is tempting to reach for whenUnsatisfiable: DoNotSchedule and call the constraint stricter. That is not the issue, and it was not mine. The strictness was never the input. The population was.

Defect one: the rollout stacks

The first way the population lies goes through time.

During a rolling update there are briefly two ReplicaSets: the outgoing one being drained and the incoming one being filled. A bare labelSelector matches both. The scheduler is therefore placing new pods against a population that includes pods about to disappear, sees a skew it cannot satisfy, and packs two replicas onto one node. They then stay there indefinitely, because nothing re-evaluates, and the next rollout does it again.

The fix is one line:

matchLabelKeys:
- pod-template-hash

The scheduler appends the pod’s own pod-template-hash to the selector before counting, which scopes the population to the pod’s own ReplicaSet. The dying siblings stop being counted and the rollout spreads normally. GA since Kubernetes 1.30, and cheap enough that it belongs on every spread constraint that has a selector with no per-workload identity in it.

Defect two: the shared label pools three Deployments

The second way the population lies goes sideways, and it is the severe one.

An operator manages three collector Deployments in this namespace, three, two and two replicas. It sets app.kubernetes.io/component: opentelemetry-collector on all of them, because that is what the label means: what a thing is, not which thing it is.

Every one of the three Deployments carries a spread constraint selecting on that component label. So every one of them computes its skew over the union of all seven pods. Three constraints, one shared population, none of them counting its own replicas.

During my pod-by-pod restart, the combined distribution across the seven pods was 3/1/2 by zone. From the union’s point of view, the one-pod zone was the minimum, so the skew filter kept admitting pods there. From the three-replica Deployment’s point of view, that zone already had a pod, and the pod I had just deleted belonged to it, so its replacements were being steered away from its own empty zone. Two of its replicas ended up stacked on one node.

Here is the part worth sitting with. The constraint was satisfied throughout, and the scheduler was correct throughout. It was told to count pods it should not have counted, it counted them, and it honoured the skew it computed. The defect was never in the mechanism. It was in the population declaration.

The fix is per-workload identity in the selector. The operator sets an instance label per Deployment, so the constraint becomes:

topologySpreadConstraints:
- maxSkew: 1
  topologyKey: topology.kubernetes.io/zone
  whenUnsatisfiable: DoNotSchedule
  labelSelector:
    matchLabels:
      app.kubernetes.io/component: opentelemetry-collector
      app.kubernetes.io/instance: <the deployment's own instance>   # scoped
  matchLabelKeys:
  - pod-template-hash                                               # added

Read the instance label off a live pod rather than assuming the key. Operators label by kind, and a label that identifies a kind is exactly the wrong key to spread on.

The repair that converges on the wrong answer

The standard repair for a stacked Deployment is to delete the doubled pods one at a time. The skew filter only admits a node whose count equals the current global minimum, so each replacement is forced onto the emptiest legal node. Deterministic, not luck, and I had used it the day before on a differently-shaped version of the same problem.

With an entangled selector, that repair is a trap. Every deletion recomputes the same wrong minimum, over the same wrong population, and lands the replacement in the same wrong place. I watched the same pod come back to the same node twice before accepting that the repair itself assumed something the constraint had quietly broken: that the filter counts what you think it counts.

The way out is to make the wrong choice illegal for one scheduling cycle. Cordon the node it keeps picking, delete the pod, so the only legal placements left are the right ones, and uncordon afterwards. That is not a fix; it is an escape. The fix is the selector.

One more trap in the same family, because it bites during exactly this kind of investigation: a Helm apply returns before the rollout it triggered finishes, because maxUnavailable: 1 makes two-of-three count as available. The apply said done while a pod was still Pending. Check rollouts yourself; do not inherit the apply’s optimism.

The fix, verified by the case that failed

Both lines went in: the per-instance label in matchLabels, the pod-template-hash in matchLabelKeys. I did not repair the stacked Deployment by hand afterwards, and that was deliberate.

The rolling update that applied the fix re-ran the exact case that had failed. Same constraint semantics, same namespace, same starting skew, the one situation that had produced two pods on a node. The Deployment came back one pod per zone, unaided. When the fix’s deployment is the failure’s reproduction, you get verification and regression test in the same event.

The same rollout incidentally took the same-zone share of the estate’s connections to 99 percent, because every replaced pod resolved into its own zone. That is the stake: with a zonal load balancer, cross-zone off, externalTrafficPolicy: Local, a zone with no pod is a load balancer node with no healthy target, and zone-affine routing pays nothing for the clients in that zone. A scheduling bug that looks cosmetic is a routing bug is a cost bug. An ordinary deploy fixed all three at once.

Where this generalises

A selector is a population declaration, not a name. Whenever you write one, the question is not “does it match my workload” but “what is the full set of things it matches”, and the two answers diverge in ways that never show up in the object you are editing.

The mild version of the lie is the sidecar: a chart’s single-replica metrics pod usually carries the same name label as the thing you meant, so the count includes it, and a lopsided spread satisfies maxSkew: 1 while leaving a zone genuinely empty. If the pods you want carry no distinguishing label of their own, the exclusion has to be expressed, not implied:

labelSelector:
  matchLabels:
    app.kubernetes.io/name: <the chart>
  matchExpressions:
  - key: app.kubernetes.io/component
    operator: DoesNotExist

The severe version is sibling Deployments pooled by a kind label, and no amount of strictness fixes it, because the strictness applies to a count over the wrong set.

The diagnostic habit that catches both: before trusting a spread constraint, enumerate what the scheduler actually counts, and enumerate it from live pods, kubectl get pods --show-labels, not from the spec you wrote. The spec tells you what you meant. The pods tell you what you got, and those are the only two facts in the system, and they only agree when the selector is right.

A satisfied constraint proves a count. Only intent decides whether the count was over the thing you meant.

$ cat CLICKHOUSE .md
· 8 min read

136 Million PUTs for 17 GiB of Data

Object storage bills per operation, and a ClickHouse part on an S3 disk is not one object but one per column. So the cost of a cold tier is a function of how many parts exist, not how many bytes they hold, and every setting that starves merges becomes a line on the bill. Two chart defaults did exactly that, and the fix that stopped it had never been committed.

clickhouse s3 finops observability mechanism
$ cat GIT .md
· 7 min read

84 Repositories Vanished. The Fix Was mkdir.

S3 has no empty directories, and in a fully packed git repository the refs directories are exactly that: empty. A file-by-file sync carried every object intact and dropped the two directories git requires to call something a repository, so 84 of 517 came back unreadable while the database still said they had commits. The migration's health checks stayed green the whole time, because none of them ever open a repository.

git aws s3 gitlab mechanism
$ cat KAFKA .md
· 8 min read

The Error Was 25 Hours Old. The Client Was Healthy the Whole Time.

A Kafka client wrapper that recovers by counting errors and panicking past a threshold is recovery proportional to traffic: stream processors trip it in seconds, quiet request-driven producers never do, and so they latch the last error and serve it indefinitely while every health signal stays green. The tell is in the error's own digits: an elapsed-time value that is byte-identical across occurrences is one cached event, not a recurring failure.

kafka resilience mechanism reliability observability