September 18, 2026 · 7 min read
The Constraint Was Satisfied. The Zone Was Empty Anyway.
I was restarting a collector pod by pod to converge some zone-aware routing, one at a time so nothing lost capacity. Three safe deletions later, two pods of the same Deployment sat on a single node and its own zone was empty.
The namespace has a hard spread constraint on every one of these Deployments. Its entire job is to prevent the state I had just produced. And it was satisfied the whole time.
What the constraint actually promises
A topologySpreadConstraint is evaluated once, when a pod is scheduled, and never again. Kubernetes does not rebalance. When a pod comes up, the scheduler counts, per value of the topology key, the pods that match the constraint’s labelSelector, computes the skew against maxSkew, and places accordingly. After that moment nothing re-evaluates; a pod stays where it landed until something deletes it.
That makes “satisfied” a claim about two things at once: a count, and the population it was counted over. The count was right in my case. The population was not.
It is tempting to reach for whenUnsatisfiable: DoNotSchedule and call the constraint stricter. That is not the issue, and it was not mine. The strictness was never the input. The population was.
Defect one: the rollout stacks
The first way the population lies goes through time.
During a rolling update there are briefly two ReplicaSets: the outgoing one being drained and the incoming one being filled. A bare labelSelector matches both. The scheduler is therefore placing new pods against a population that includes pods about to disappear, sees a skew it cannot satisfy, and packs two replicas onto one node. They then stay there indefinitely, because nothing re-evaluates, and the next rollout does it again.
The fix is one line:
matchLabelKeys:
- pod-template-hash
The scheduler appends the pod’s own pod-template-hash to the selector before counting, which scopes the population to the pod’s own ReplicaSet. The dying siblings stop being counted and the rollout spreads normally. GA since Kubernetes 1.30, and cheap enough that it belongs on every spread constraint that has a selector with no per-workload identity in it.
Defect two: the shared label pools three Deployments
The second way the population lies goes sideways, and it is the severe one.
An operator manages three collector Deployments in this namespace, three, two and two replicas. It sets app.kubernetes.io/component: opentelemetry-collector on all of them, because that is what the label means: what a thing is, not which thing it is.
Every one of the three Deployments carries a spread constraint selecting on that component label. So every one of them computes its skew over the union of all seven pods. Three constraints, one shared population, none of them counting its own replicas.
During my pod-by-pod restart, the combined distribution across the seven pods was 3/1/2 by zone. From the union’s point of view, the one-pod zone was the minimum, so the skew filter kept admitting pods there. From the three-replica Deployment’s point of view, that zone already had a pod, and the pod I had just deleted belonged to it, so its replacements were being steered away from its own empty zone. Two of its replicas ended up stacked on one node.
Here is the part worth sitting with. The constraint was satisfied throughout, and the scheduler was correct throughout. It was told to count pods it should not have counted, it counted them, and it honoured the skew it computed. The defect was never in the mechanism. It was in the population declaration.
The fix is per-workload identity in the selector. The operator sets an instance label per Deployment, so the constraint becomes:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app.kubernetes.io/component: opentelemetry-collector
app.kubernetes.io/instance: <the deployment's own instance> # scoped
matchLabelKeys:
- pod-template-hash # added
Read the instance label off a live pod rather than assuming the key. Operators label by kind, and a label that identifies a kind is exactly the wrong key to spread on.
The repair that converges on the wrong answer
The standard repair for a stacked Deployment is to delete the doubled pods one at a time. The skew filter only admits a node whose count equals the current global minimum, so each replacement is forced onto the emptiest legal node. Deterministic, not luck, and I had used it the day before on a differently-shaped version of the same problem.
With an entangled selector, that repair is a trap. Every deletion recomputes the same wrong minimum, over the same wrong population, and lands the replacement in the same wrong place. I watched the same pod come back to the same node twice before accepting that the repair itself assumed something the constraint had quietly broken: that the filter counts what you think it counts.
The way out is to make the wrong choice illegal for one scheduling cycle. Cordon the node it keeps picking, delete the pod, so the only legal placements left are the right ones, and uncordon afterwards. That is not a fix; it is an escape. The fix is the selector.
One more trap in the same family, because it bites during exactly this kind of investigation: a Helm apply returns before the rollout it triggered finishes, because maxUnavailable: 1 makes two-of-three count as available. The apply said done while a pod was still Pending. Check rollouts yourself; do not inherit the apply’s optimism.
The fix, verified by the case that failed
Both lines went in: the per-instance label in matchLabels, the pod-template-hash in matchLabelKeys. I did not repair the stacked Deployment by hand afterwards, and that was deliberate.
The rolling update that applied the fix re-ran the exact case that had failed. Same constraint semantics, same namespace, same starting skew, the one situation that had produced two pods on a node. The Deployment came back one pod per zone, unaided. When the fix’s deployment is the failure’s reproduction, you get verification and regression test in the same event.
The same rollout incidentally took the same-zone share of the estate’s connections to 99 percent, because every replaced pod resolved into its own zone. That is the stake: with a zonal load balancer, cross-zone off, externalTrafficPolicy: Local, a zone with no pod is a load balancer node with no healthy target, and zone-affine routing pays nothing for the clients in that zone. A scheduling bug that looks cosmetic is a routing bug is a cost bug. An ordinary deploy fixed all three at once.
Where this generalises
A selector is a population declaration, not a name. Whenever you write one, the question is not “does it match my workload” but “what is the full set of things it matches”, and the two answers diverge in ways that never show up in the object you are editing.
The mild version of the lie is the sidecar: a chart’s single-replica metrics pod usually carries the same name label as the thing you meant, so the count includes it, and a lopsided spread satisfies maxSkew: 1 while leaving a zone genuinely empty. If the pods you want carry no distinguishing label of their own, the exclusion has to be expressed, not implied:
labelSelector:
matchLabels:
app.kubernetes.io/name: <the chart>
matchExpressions:
- key: app.kubernetes.io/component
operator: DoesNotExist
The severe version is sibling Deployments pooled by a kind label, and no amount of strictness fixes it, because the strictness applies to a count over the wrong set.
The diagnostic habit that catches both: before trusting a spread constraint, enumerate what the scheduler actually counts, and enumerate it from live pods, kubectl get pods --show-labels, not from the spec you wrote. The spec tells you what you meant. The pods tell you what you got, and those are the only two facts in the system, and they only agree when the selector is right.
A satisfied constraint proves a count. Only intent decides whether the count was over the thing you meant.